Skip to content

delete-index-atomicity-rocksdb flakes: the raw-RocksDB oracle's read-only open races the server's compactions #2500

Description

@kriszyp

Symptom

integrationTests/database/delete-index-atomicity-rocksdb.test.ts intermittently reddened Integration Tests 3/6 on main with:

✖ oracle proof (ItemF): sees a seeded index entry, and sees it vanish on a normal delete
Error: Corruption: Corruption: IO error: No such file or directory:
  While open a file for random read: /tmp/harper-integration-test-XXXXXX/database/data/000021.sst:
  No such file or directory

failing the parent suite #1854 audit:false delete must not orphan secondary-index entries (raw RocksDB oracle).

Observed on main (workflow Integration Tests), all before the fix below landed:

run when (UTC) shard
33560941525 2026-09-01T21:25Z 3/6 (Bun)
33561117013 2026-09-01T21:27Z 3/6 (uWS HTTP)
33584091644 2026-09-02T02:40Z 3/6 (uWS HTTP)

It appeared on both the Bun and uWS shard-3 variants, so it was never runtime-specific.

Root cause: the test's own oracle, not the product

The throw comes from the test process, not from Harper. The suite's oracle opened its own
RocksDatabase.open(dir, { readOnly: true }) handles against the live database directory while
the Harper process under test was writing and compacting it.

readOnly: true maps to rocksdb::DB::OpenForReadOnly, which replays the MANIFEST into a file list
and then opens each of those files holding no reference on any of them. A compaction in the
writing process can unlink an input SST inside that window, and the open fails naming the file
that vanished — RocksDB's generic "the MANIFEST may be corrupted" wording. Nothing is corrupt: the
keys live on in the compaction's output file. Tracked engine-side as HarperFast/rocksdb-js#812.

Nothing here is an SST disappearing from under the server's reader; the server holds proper
references. This is a harness race, and no product code is implicated.

Reproduction (mechanism, isolated from Harper)

Two processes against one RocksDB directory, @harperfast/rocksdb-js 2.8.0, Linux:

  • writer — 3,000 puts then flush() in a loop (each flush makes a new L0 file, which is exactly
    what keeps L0→L1 compaction running), plus a createCheckpoint() per iteration.
  • readerRocksDatabase.open(target, { readOnly: true }) + a short range scan, in a loop.
reader target attempts failed opens
the live directory (pre-fix oracle) 3,804 273 (7.2%)
a checkpoint (post-fix oracle) 5,804 0 genuine (7 were the harness's own non-atomic cursor file)

The local failures surface as rocksdb-js's Database does not exist, which is what it reports for a
read-only IO error once the file name has been discarded; CI's build surfaced the same race with the
file name intact. Same window, same cause.

Fix

  • Deflake the #1854 RocksDB oracle: reopen when a live compaction removes an SST mid-open #2429 (2026-09-02) — bounded reopen for that one error signature. A mitigation.
  • Give the #1854 RocksDB oracle a checkpoint it cannot race #2452 (2026-09-02) — the actual fix: the oracle now reads a RocksDB
    checkpoint published by the fixture (/Snapshot/), never the live directory. Nothing writes to a
    checkpoint, so the window does not exist there, and createCheckpoint() flushes the memtable —
    which the oracle needs anyway, since Harper opens table/index column families with disableWAL
    defaulting to true. refreshOracle() asserts the published path is exactly the per-sequence
    checkpoint path, so a fixture regression that handed back a live directory fails closed instead of
    being opened.

The #1854 guarantee the suite pins is unchanged: an audit:false delete must not orphan
secondary-index entries, proven against raw index column families with both a monitor-fired abort
(Arm A) and a request-thrown abort (Arm B).

No recurrence on main since #2452 landed (2026-09-02T22:58Z).

Why this issue is open

The flake itself is closed. What stays open is the follow-up: the checkpoint is a workaround for
HarperFast/rocksdb-js#812 (readOnly:true has no safe form against a live writer, and there is no
OpenAsSecondary in the binding). When #812 ships a follower-mode open that pins the files it lists,
the oracle can drop the checkpoint dance and read the live directory directly.

Until then the rule is a class rule, not a rule about this one file: no test may open the server's
live storage directory
. integrationTests/README.md gets that written down, with the checkpoint
recipe, so the next storage-oracle test does not rediscover this.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Fields

    Priority

    P3

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions