Skip to content

The Lock File

ditto.lock is a tool-generated, committed, PR-reviewed file that records which snapshots your test suite legitimately owns. It is the source of truth behind ditto verify, ditto prune, and the credential-free CLI inventory.

It is modelled on package-lock.json, Cargo.lock, and poetry.lock: generated rather than hand-edited, deterministic, merge-friendly, and reviewed as part of the diff. The one exception is retiring a target.

Why a committed file

A snapshot is "legitimate" only if your suite is supposed to own it. That fact cannot be derived from which tests happened to run in one session, and a throwaway cache cannot survive a fresh CI checkout. ditto.lock persists it: it is committed, so a clean checkout — or a partial, failed, or parallel run — still knows the full, authoritative set.

ditto.lock lives in pytest's rootdir, usually the project root. Commit it. Do not add it to .gitignore; ditto warns when .gitignore in the rootdir has a ditto.lock or /ditto.lock line (it doesn't detect broader patterns such as *.lock).

What it records

For each resolved target (a backend URI such as the local .ditto/ directory or redis://…), the lock stores one entry per snapshot: the test nodeid, the snapshot key, and the recorder. It records the targets your tests actually used — including per-test record(target=…) marks — because it is written by real runs. It never stores storage_options, but it does store each target URI verbatim, so ditto refuses a target URI that contains a password or a secret query parameter; pass credentials as storage options, or in a profile's storage_options, instead.

How it is produced and maintained

Command Effect on ditto.lock
pytest (normal run) Appends entries for any snapshots recorded this run. If it can't write the lock, it warns and the run still passes.
ditto lock (pytest --ditto-lock) Rebuilds the lock from a full, passing run: in each target the run used, drops entries for tests and keys that no longer exist. Targets the run didn't use are kept as they were. Refuses a filtered, narrowed or failing run. If it can't write the lock, the run fails.
ditto update (pytest --ditto-update) On a full, passing run, rebuilds the lock as ditto lock does, and the run fails if it can't write the lock. On a filtered, narrowed or failing run, only appends.
ditto prune (pytest --ditto-prune) Does not write the lock; deletes backend snapshots absent from it.

A rebuild works test by test. A test that passed this run has its entries replaced by the snapshots it used. A test that didn't run its body to a pass keeps its entries: one that was skipped (for example by a platform skipif, or a module-level pytest.skip or pytest.importorskip), xfailed, deselected with --deselect, or in a path pytest didn't collect (--ignore, --ignore-glob, a conftest.py collect_ignore or collect_ignore_glob, or norecursedirs). Entries for tests that no longer exist are dropped, which is what cleans up after a renamed or deleted test. A skip on one machine therefore never removes a snapshot that another machine still runs.

Sharing a target

A target holds one set of snapshots. Every checkout that uses it, whether another branch of the project, another project with the same test paths, CI or a developer's machine, reads and writes that same set:

  • Writes are shared. Recording a snapshot, or running ditto update, on one branch changes what every other branch compares against.
  • Orphans are ambiguous. ditto verify and ditto prune only look at keys under the test modules your suite owns, so another suite with different module paths is left alone. But a snapshot that only another branch, or another project with the same test paths, has recorded looks like an orphan: ditto verify reports it as drift, and ditto prune --shared would delete it.

So:

  • Give each project its own target path, such as s3://bucket/<project>/.
  • If branches change snapshots independently, keep the snapshots in the repository, in the default .ditto directories. Git then keeps each branch's snapshots and lock together, and merges them like any other file.
  • Use a remote target for snapshots that change through one branch, such as baselines updated only on the main branch.

A separate remote path per branch doesn't work yet. The lock records a remote target by its URI, so on a branch whose path differs, ditto verify finds no entries for it and fails, even after the snapshots are copied there. Recording the lock under the branch's path instead puts that path into the lock, which the merge then carries into the main branch. Issue #251 tracks making this possible.

Because ditto can't check that a target is used by one checkout only, ditto prune deletes nothing from one that might be shared unless you pass --shared (pytest --ditto-prune-shared). Without it, prune says how many snapshots it left in each such target and the run fails. A target might be shared when it is anything other than a file:// path inside the project, such as the default .ditto: a remote URI, or a file:// path outside the project. Symlinks are followed first, so a .ditto inside the project that links to a directory outside it counts as shared. ditto prune --check lists what would be deleted either way.

How the lock is used

  • ditto verify — diffs the live backend against the lock and fails on drift (missing, orphan, or unsynced snapshots). See ditto verify.
  • ditto prune — deletes backend snapshots that are not in the lock, never a snapshot created during the same run. See ditto prune.
  • CLI inventory (list/status/stats/lint) — reads the lock for a fast, credential-free remote inventory (below).

Under pytest-xdist distribution (-n) the lock isn't updated, and verify, lock and prune are refused; see Running in CI.

Declared vs physical state (and the inventory trade-off)

The lock is a declared record — what is legitimate — not a physical one — what is actually stored. They can diverge: a snapshot deleted from the backend but still in the lock (missing), or an orphan in the backend not in the lock. Detecting that divergence is exactly what ditto verify is for.

This shapes how the read-only inventory commands source their data:

flowchart TD
    C["ditto list / status / stats / lint"] --> L{"--live?"}
    L -- "yes" --> P["pytest --setup-only pass<br/>(physical, all targets, needs credentials)"]
    L -- "no (default)" --> T{"per target"}
    T -- "local file://" --> F["filesystem walk<br/>(real size + mtime, shows orphans)"]
    T -- "remote (redis://, s3://, …)" --> K["ditto.lock<br/>(declared entries, size = —)"]
  • Local file targets are read from disk — real sizes and mtimes, including on-disk orphans not in the lock — instantly and without credentials.
  • Remote targets are read from the lock — credential-free, but size and mtime are unknown (shown as —).
  • --live runs the pytest introspection pass for authoritative physical state across every target. It imports your test modules and needs the same credentials your test run needs.

The reason for the split is a hard asymmetry: reading a local file's true state is free (a stat), but reading a remote backend's true state requires connecting to it, which is credential-gated. You can have at most two of {uniform behaviour, true physical state, fast & credential-free} — ditto's default keeps inventory fast and credential-free, and offers truth on demand via --live.