The Lock File¶
ditto.lock is a tool-generated, committed, PR-reviewed file that records which
snapshots your test suite legitimately owns. It is the source of truth behind
ditto verify, ditto prune, and the credential-free CLI inventory.
It is modelled on package-lock.json, Cargo.lock, and poetry.lock:
generated rather than hand-edited, deterministic, merge-friendly, and reviewed
as part of the diff. The one exception is
retiring a target.
Why a committed file¶
A snapshot is "legitimate" only if your suite is supposed to own it. That fact
cannot be derived from which tests happened to run in one session, and a
throwaway cache cannot survive a fresh CI checkout. ditto.lock persists it: it
is committed, so a clean checkout — or a partial, failed, or parallel run — still
knows the full, authoritative set.
ditto.lock lives in pytest's rootdir, usually the project root. Commit it.
Do not add it to .gitignore; ditto warns when .gitignore in the rootdir
has a ditto.lock or /ditto.lock line (it doesn't detect broader patterns
such as *.lock).
What it records¶
For each resolved target (a backend URI such as the local .ditto/
directory or redis://…), the lock stores one entry per snapshot: the test
nodeid, the snapshot key, and the recorder. It records the targets your tests
actually used — including per-test record(target=…) marks — because it is
written by real runs. It never stores storage_options, but it does store each
target URI verbatim, so ditto refuses a target URI that contains a password or
a secret query parameter; pass credentials as
storage options,
or in a profile's storage_options, instead.
How it is produced and maintained¶
| Command | Effect on ditto.lock |
|---|---|
pytest (normal run) |
Appends entries for any snapshots recorded this run. If it can't write the lock, it warns and the run still passes. |
ditto lock (pytest --ditto-lock) |
Rebuilds the lock from a full, passing run: in each target the run used, drops entries for tests and keys that no longer exist. Targets the run didn't use are kept as they were. Refuses a filtered, narrowed or failing run. If it can't write the lock, the run fails. |
ditto update (pytest --ditto-update) |
On a full, passing run, rebuilds the lock as ditto lock does, and the run fails if it can't write the lock. On a filtered, narrowed or failing run, only appends. |
ditto prune (pytest --ditto-prune) |
Does not write the lock; deletes backend snapshots absent from it. |
A rebuild works test by test. A test that passed this run has its entries
replaced by the snapshots it used. A test that didn't run its body to a pass
keeps its entries: one that was skipped (for example by a platform skipif, or
a module-level pytest.skip or pytest.importorskip), xfailed, deselected with
--deselect, or in a path pytest didn't collect (--ignore, --ignore-glob,
a conftest.py collect_ignore or collect_ignore_glob, or norecursedirs).
Entries for tests that no longer exist are dropped, which is what cleans up
after a renamed or deleted test. A skip on one machine therefore never removes
a snapshot that another machine still runs.
Sharing a target¶
A target holds one set of snapshots. Every checkout that uses it, whether another branch of the project, another project with the same test paths, CI or a developer's machine, reads and writes that same set:
- Writes are shared. Recording a snapshot, or running
ditto update, on one branch changes what every other branch compares against. - Orphans are ambiguous.
ditto verifyandditto pruneonly look at keys under the test modules your suite owns, so another suite with different module paths is left alone. But a snapshot that only another branch, or another project with the same test paths, has recorded looks like an orphan:ditto verifyreports it as drift, andditto prune --sharedwould delete it.
So:
- Give each project its own target path, such as
s3://bucket/<project>/. - If branches change snapshots independently, keep the snapshots in the
repository, in the default
.dittodirectories. Git then keeps each branch's snapshots and lock together, and merges them like any other file. - Use a remote target for snapshots that change through one branch, such as baselines updated only on the main branch.
A separate remote path per branch doesn't work yet. The lock records a remote
target by its URI, so on a branch whose path differs, ditto verify finds no
entries for it and fails, even after the snapshots are copied there. Recording
the lock under the branch's path instead puts that path into the lock, which
the merge then carries into the main branch.
Issue #251 tracks
making this possible.
Because ditto can't check that a target is used by one checkout only, ditto
prune deletes nothing from one that might be shared unless you pass
--shared (pytest --ditto-prune-shared). Without it, prune says how many
snapshots it left in each such target and the run fails. A target might be
shared when it is anything other than a file:// path inside the project,
such as the default .ditto: a remote URI, or a file:// path outside the
project. Symlinks are followed first, so a .ditto inside the project that
links to a directory outside it counts as shared. ditto prune --check lists
what would be deleted either way.
How the lock is used¶
ditto verify— diffs the live backend against the lock and fails on drift (missing, orphan, or unsynced snapshots). See ditto verify.ditto prune— deletes backend snapshots that are not in the lock, never a snapshot created during the same run. See ditto prune.- CLI inventory (
list/status/stats/lint) — reads the lock for a fast, credential-free remote inventory (below).
Under pytest-xdist distribution (-n) the lock isn't updated, and verify,
lock and prune are refused; see
Running in CI.
Declared vs physical state (and the inventory trade-off)¶
The lock is a declared record — what is legitimate — not a physical one
— what is actually stored. They can diverge: a snapshot deleted from the backend
but still in the lock (missing), or an orphan in the backend not in the lock.
Detecting that divergence is exactly what ditto verify is for.
This shapes how the read-only inventory commands source their data:
flowchart TD
C["ditto list / status / stats / lint"] --> L{"--live?"}
L -- "yes" --> P["pytest --setup-only pass<br/>(physical, all targets, needs credentials)"]
L -- "no (default)" --> T{"per target"}
T -- "local file://" --> F["filesystem walk<br/>(real size + mtime, shows orphans)"]
T -- "remote (redis://, s3://, …)" --> K["ditto.lock<br/>(declared entries, size = —)"]
- Local file targets are read from disk — real sizes and mtimes, including on-disk orphans not in the lock — instantly and without credentials.
- Remote targets are read from the lock — credential-free, but size and
mtime are unknown (shown as
—). --liveruns the pytest introspection pass for authoritative physical state across every target. It imports your test modules and needs the same credentials your test run needs.
The reason for the split is a hard asymmetry: reading a local file's true state
is free (a stat), but reading a remote backend's true state requires connecting
to it, which is credential-gated. You can have at most two of {uniform
behaviour, true physical state, fast & credential-free} — ditto's default keeps
inventory fast and credential-free, and offers truth on demand via --live.