Skip to content

db/state: replace the .vi perfect hash with two Elias-Fano sequences (40.11GB -> 549MB) - #23844

Draft
AskAlexSharov wants to merge 24 commits into
mainfrom
alex/vi_ef_accessor_37
Draft

AskAlexSharov wants to merge 24 commits into
mainfrom
alex/vi_ef_accessor_37

Conversation

@AskAlexSharov

@AskAlexSharov AskAlexSharov commented Sep 7, 2026 •

Copy link
Copy Markdown
Collaborator

.vi is 40.11 GB of the 83.5 GB of accessor files on a mainnet archive — the largest single item, and 95.4% of it is a recsplit ordinal array over 8.19 billion keys. This replaces it with two Elias-Fano sequences: 40.11 GB → 549 MB.

The offset is already known when .vi is consulted

historySeekInFiles did two lookups for one read. The first, iit.seekInFiles, resolves the key in .efi, which computes the key's ordinal in .ef on the way to the offset and then discards it; it then decodes that key's txNum list and calls seq.Seek, whose second return value is the rank of the txNum in the list — dropped as _. The second was a full perfect-hash lookup over 8.19 billion keys, to get an offset.

But buildVI writes .v in exactly that order — .ef in key order, each key's txNums in order, valOffset advancing on one global counter. So:

ordinal = starts[keyNum] + rank
offset  = pages[ordinal / pageValuesCount]

buildVI also advances valOffset only once per CompressedPageValuesCount, so the old .vi stored the same page offset 64 times over on most files (16 on the rest).

db/datastruct/posidx

The structure is not history-specific, so it lands as its own package. It maps a position to a value for data written in a single pass as consecutive runs of items:

starts[r] = items before run r   ->  ordinal = starts[r] + item
pages[p]  = value of page p      ->  value   = pages[ordinal / pageSize]

A run is one key's consecutive values in .v; an item is one of them. starts counts items, pages holds the values, and the two have different lengths — 2,784,209 against 489,483 for accounts.1152-1184.

Two known pieces: starts is a CSR row-pointer array (Arrow's ListArray offsets, ClickHouse's Array offsets column), pages is sampled-per-page addressing (Parquet's OffsetIndex, Lucene skip lists), both Elias-Fano encoded. The pairing is what a quasi-succinct index does. db/state keeps only what is genuinely history's: which order .v is written in, and reading v1 files.

Get(run, item) maps a position to its value — the only operation this PR needs. The reverse direction (which run owns an ordinal) lands with the transactions-to-block index that wants it.

Change

  • buildVI emits both sequences as it walks. It no longer hashes, so it cannot collide, and the retry loop is gone.
  • TwoLayerLookupByHashWithOrdinal surfaces the ordinal the .efi lookup already computes; seekInFiles returns it, the rank, and which .ef file matched.
  • A positional .vi only answers for the .ef it was built from, and the two file lists are searched separately, so the caller takes the .v paired with the file the seek matched rather than the one covering the txNum. The pairing holds by construction; a missing pair is an error.
  • All five readers converted: historySeekInFiles, the two merge-heap iterators, HistoryTraceKeyFiles, HistoryDump. Plus the merge path and FilesItem lifecycle.
  • erigon seg integrity --check=HistoryVi re-derives every offset from the .ef/.v pair the way buildVI does and compares it with what the .vi answers. It asks a v1 .vi the same question by txNum+key, so it covers both formats.

Verification

The new writer builds an index from real mainnet .ef/.v, and every offset is compared against what the existing .vi returns for the same txNum+key:

domain page values checked mismatch .vi → new
commitment 64 165,604,124 0 865.56 → 12.73 MB 68x
accounts 64 33,301,783 0 140.75 → 2.89 MB 49x
storage 1 24,955,229 0 105.47 → 19.41 MB 5.4x
receipt 1 18,631,837 0 78.74 → 9.34 MB 8.4x
rcache 16 12,500,000 0 52.83 → 1.19 MB 44x
code 1 478,329 0 1.54 → 0.34 MB 4.6x

255.5M real values, zero mismatches, every domain and both page sizes. Whole-archive sizing over all 47 files: 40.11 GB → 549 MB (73x).

db/, execution/, cmd/, rpc/ green; lint clean; both rpc-integration jobs pass.

Notes for review

  • v1 .vi files are still read. Requiring a rebuild (min-supported 2.0) was the first attempt, and the rpc-integration jobs caught it: MustSupport panics on a datadir holding v1 files and rpcdaemon cannot rebuild accessors, so port 8545 never opened. HistoryValueIndex now reads either format, picked by the version openDirtyAccessor already parses. Lookup takes the txNum and key alongside the position — v2 ignores them, v1 addresses by txNum+key — keeping the branch in one place rather than at all five call sites.
  • Existing datadirs keep their v1 .vi until it is replaced. missedMapAccessors matches any .vi version, so a present v1 file is never missed and never rebuilt, and min stays v1.0 so it opens. The saving lands on freshly built or re-downloaded files; to get it on an existing datadir, delete *.vi and restart. Treating a below-current .vi as missed would rebuild automatically, but that is a full pass over every .v on the first start after upgrade — a separate decision.
  • The lo field in the seek cache is an independent fix. The cache is keyed on the first half of the key hash alone, so a colliding hi returned another key's txNum rather than a miss, before this PR too (~4096/2^64 per lookup). What changes is the consequence: a bogus txNum used to miss the perfect hash and fall through to the DB, where a positional lookup returns a plausible wrong value. Same for the miss/hit branch order at txNum == 0.
  • A positional miss is an error, not a fallback. The caller took (keyOrdinal, rank) out of the .ef it is holding, so a miss cannot mean "key absent" — only that the .v disagrees with the .ef indexing it. History seek and both range iterators report it instead of falling through to the DB or dropping the key, which is what HistoryTraceKeyFiles already did.
  • ReadEliasFanoChecked replays the jump table against the upper bits rather than only counting them — the same single pass, since it already walked upperBits. Over every single-bit mutation of an index, Open accepted 192 files that then panicked in a lookup; now none, with 528 of 1216 still accepted (flips in lowerBits, which give a wrong value and cannot run off the end).
  • EliasFano.Seek reports a hit at the constructor's maxOffset even when the sequence holds no such value ({0,3,4} built with maxOffset=6 returns val=6, pos=2, ok=true). Nothing here calls it — existing callers pass the true last element, so it has never bitten. Filed separately; the transactions-to-block index has to handle it.
  • Dropped TestHistoryBuildVI_PageCounterResetOnCollisionRetry — it pinned collision-retry behavior that no longer exists.
  • The 5.4x on storage and 4.6x on code are the unpaged (page=1) files, where there is one offset per value and only the Elias-Fano compression applies. The paged domains, which are the bulk of the bytes, are 44–68x.

…quences

.vi is 40.11GB of the 83.5GB of accessor files on a mainnet archive, and 95.4%
of it is the recsplit ordinal array. The offset it stores is a pure function of
the key ordinal in .ef and the rank of the txNum in that key's list - both are
computed by the inverted-index lookup that runs one statement earlier and then
discarded. buildVI also advances the offset only once per compressed page, so
the same offset is written 64 times over on most files.

Verified on a mainnet archive: 670M derived lookups against the existing .vi,
no misses, offsets monotone in ordinal order, and the derived value count equal
to .vi's own key count. Building both replacement sequences for real over all
47 files gives 40.11GB -> 549MB.

No implementation yet - this is the spec for review.
.vi mapped txNum+key -> offset in .v with a perfect hash over every history
value: 8.19 billion keys and 40.11GB on a mainnet archive, 95.4% of it the
recsplit ordinal array.

That offset is a pure function of two numbers the lookup path already has.
seekInFiles resolves the key in .efi, which computes the key's ordinal in .ef
on the way to the offset, and Seek on the key's txNum list returns the rank of
the txNum - both were discarded. buildVI writes .v in exactly that order, so
the value's position is cumValues[keyOrdinal]+rank, and the offset only
advances once per compressed page. Both sequences are monotone.

.vi v2 is those two Elias-Fano sequences and no perfect hash. Measured over all
47 files of a mainnet archive: 40.11GB -> 549MB. Build no longer hashes, so it
also cannot collide, and the retry loop goes away.

AccessorVI min-supported moves to 2.0: v1 files are rejected and rebuilt from
.v and .ef by the usual missed-accessor path.
Bumping AccessorVI min-supported to 2.0 made MustSupport panic on any datadir
still holding v1 files, and rpcdaemon cannot rebuild accessors - it just fails
to start. The mainnet and gnosis rpc-integration jobs caught this: port 8545
never opened.

Min-supported goes back to 1.0 and HistoryValueIndex reads either format. The
version already parsed by openDirtyAccessor picks the reader, so no magic byte
is needed - a v1 .vi is a recsplit index whose first byte is a small version
number and would not be distinguishable otherwise.

Lookup takes the txNum and key alongside the position; a v2 index ignores them,
a v1 index addresses by txNum+key as before. That keeps the branch in one place
instead of at all five call sites.
The structure has nothing to do with history: it indexes values that grow along
a known order and repeat in runs, addressed by (group, member) or by a flat
ordinal. Nothing in it mentions txNums, keys or .v files.

db/state keeps only the history-specific part - which order .v is written in,
and reading v1 files that still hold a perfect hash.
Ungrouped mode, Ordinal/Value/Get as three methods, PageSize, HasGroups,
GroupCount and FilePath were all written for callers that do not exist. Get is
the whole read API; the history adapter keeps its own path and drops a KeyCount
only its tests read.
…nicking

Get took the position from a separate file and indexed straight into the
Elias-Fano sequences, which panic with a bare 'index out of range' when the two
files do not belong together - inside an rpc read, with nothing naming the
files. It now returns false, which the caller already handles, and that also
subsumes the empty-index check at the call site.

Format version starts at 1: it is this package's own, not the .vi file version
it was extracted from.
…to page

Get answers 'where is this item'. Group and Page answer the opposite: which
group owns an ordinal, which page holds a value. Both are predecessor searches
over the same two sequences, and Group is what an index like
transactions-to-block needs - it maps a position back to the block owning it.

Empty groups own no ordinal, so Group looks for the first group starting after
the ordinal and steps back, which skips them.

The header gains the item count, so an ordinal past the end reports not-found
rather than resolving to the last group.

Neither reverse call passes maxOffset to EliasFano.Seek: Seek reports a hit
there even when the sequence holds no such value, and this index builds its
sequences with a bound rather than the true maximum.
The code said groups/values, the docs said cumValues/pageOffsets and the API
said group/member - three names for two arrays, which is what made the
structure hard to read.

One set now, everywhere:

  starts[r] = items before run r     ordinal = starts[r] + item
  pages[p]  = value of page p        value   = pages[ordinal / pageSize]

A run is one key's consecutive values; an item is one of them. starts counts
items and pages holds the values, so no word appears in both.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Positional lookups are unsafe when history and inverted-index ranges diverge, with additional parser and compatibility defects.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Replaces large history .vi perfect-hash accessors with compact Elias-Fano positional indexes while retaining v1 compatibility.

Changes:

  • Adds the reusable pagedidx index and tests.
  • Propagates key ordinals and ranks through history readers.
  • Updates accessor versions, lifecycle handling, and fixtures.
File summaries
File Description
db/datastruct/pagedidx/paged_index.go Implements the new index.
db/datastruct/pagedidx/paged_index_test.go Tests index operations.
db/recsplit/index_reader.go Exposes lookup ordinals.
db/state/history_vi.go Supports v1 and v2 .vi.
db/state/history_vi_test.go Tests both formats.
db/state/history.go Builds and reads positional indexes.
db/state/history_stream.go Migrates streaming readers.
db/state/inverted_index.go Returns ordinal and rank metadata.
db/state/cache.go Caches positional metadata.
db/state/state_recon.go Carries key ordinals.
db/state/dirty_files.go Manages new accessor lifecycle.
db/state/merge.go Integrates indexes into merges.
db/state/statecfg/version_schema_gen.go Declares v2 .vi versions.
db/state/history_test.go Updates history tests.
db/state/snap_repo_test.go Generates v2 test accessors.
db/state/dirty_files_test.go Updates accessor-opening tests.
db/state/aggregator_test.go Uses v2 accessor schemas.
cmd/utils/app/snapshots_publishable_state_test.go Updates publishability fixtures.
Review details

Files not reviewed (1)

  • db/state/statecfg/version_schema_gen.go: Generated file
  • Files reviewed: 17/18 changed files
  • Comments generated: 8
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread db/datastruct/posidx/index.go
Comment thread db/datastruct/pagedidx/paged_index.go
Comment thread db/state/history.go Outdated
Comment thread db/state/history.go
Comment thread db/state/history_stream.go
Comment thread db/state/history_vi.go
Comment thread cmd/utils/app/snapshots_publishable_state_test.go
Comment thread db/state/statecfg/version_schema_gen.go
A v2 .vi is addressed by (keyOrdinal, rank) alone. Those come from the .ef
file seekInFiles landed in, while the .v file is chosen by txNum from an
independent list, so a divergence answered another key's offset. The seek
result now carries its source range, through the seek cache too, and
FilesItem.LookupHistoryValue refuses the mismatch in one place for all three
callers.

Also moves the .vi.torrent fixtures to v2.0 with their primaries, and takes
hist.vi to v2.0 in versions.yaml so the generated schema matches its source.
@AskAlexSharov
AskAlexSharov force-pushed the alex/vi_ef_accessor_37 branch 3 times, most recently from 1938134 to a4a9b42 Compare September 8, 2026 04:09
…ct corrupt accessors at open

The paged index header was 1+8+8 bytes, so both Elias-Fano sequences started at
an offset 1 byte off an 8-byte boundary. They are read as []uint64 straight out
of the mapping, which made every lookup an unaligned load. Pad the header to 24
bytes while no v2 .vi file exists outside a test datadir.

Open now refuses a file it cannot trust instead of crashing on it. A fuzz target
over Open found three ways to panic, all of them in eliasfano32 rather than in
the writer:

- a forged count/universe reached make([]uint64, totalWords) inside
  ReadEliasFano, before any caller could inspect the result
- count == MaxUint64 wrapped count+1 to a zero divisor in computeLayout
- an upper-bits stream shorter than its header claims ran the decode walk off
  the end

ReadEliasFanoChecked covers all three: it derives the layout from the header,
refuses one that does not fit the bytes, and counts the set bits. It is used
only by the paged index, so .ef and multiencseq keep their unchecked path.
Validating the jump table would cost what rebuilding it costs, so a lookup can
still leave the sequence on a forged table; the fuzz target records that
boundary rather than implying a stronger one.

Build validates the counts it was sized for, since AddOffset writes into the
sequences with no bounds of its own and a miscount would ship a wrong index.

An accessor that fails to open is left nil and rebuilt by BuildMissedAccessors,
and history falls back to the DB meanwhile, so ErrCorrupt degrades to a rebuild
rather than a failed start.

The package is posidx, not pagedidx: values are optional, so "paged" named the
half a caller can leave out, while every caller asks whether an index is
positional.
…g its range

A positional .vi only answers for the .ef it was built from, and the two file
lists are searched separately. Instead of threading the .ef range through the
seek result, the LRU entry and a guard method, seekInFiles reports which file
matched and the caller takes the .v paired with it. The mismatch it guarded
against is now a pairedFile miss.

Also: BinarySearch returns the ordinal it searched for, so the range iterator
no longer re-hashes its first key; the ordinal folds into recsplit's existing
two-layer helper instead of a second pair of exported methods; and Index.Page
goes, it has no caller.
…files

`erigon seg integrity --check=HistoryVi` walks each .ef with the .v paired to
it, re-derives every value's offset the way buildVI does, and compares it with
what the .vi answers. A v1 .vi is asked the same question by txNum+key, so the
check covers both formats.

@awskii awskii left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two things that aren't line-anchored.

Existing datadirs never see the saving. missedMapAccessors (history.go:186) matches any .vi version through MatchVersionedFile, so a present v1.2 file is never "missed" and never rebuilt — min stays v1.0, OpenHistoryValueIndex reads it as legacy, and it stays v1 forever. Is the path "delete *.vi and rebuild", or "wait for the next snapshot release"? Worth a line in the description, since 40.11 GB -> 549 MB only lands for freshly built or re-downloaded files.

The lo field is a fix in its own right. The seek cache is keyed on hi alone, so a colliding hi returned another key's txNum rather than a miss, before this PR too. cache.go:104 frames it as protecting the positional result; it also closes that. About 4096/2^64 per lookup, so nothing urgent — it just isn't in the description.

Reviewed at 3816f9d. db/state, posidx, recsplit and cmd/utils/app are green here including -race; the go vet unkeyed-field noise in version_schema_gen.go is pre-existing on main.

Comment thread db/state/inverted_index.go Outdated
Comment thread db/datastruct/posidx/index.go Outdated
Comment thread db/datastruct/posidx/index.go Outdated
Comment thread db/datastruct/posidx/index_test.go
Comment thread db/state/history.go Outdated
Comment thread db/state/dirty_files.go
Comment thread db/state/merge.go Outdated
Comment thread db/datastruct/posidx/index.go Outdated
Comment thread db/recsplit/eliasfano32/elias_fano.go
- A cached seek miss is stored as found=0, and the hit branch was tested first,
  so txNum=0 returned the miss as a hit with an all-zero position. With a
  positional value index that resolves to another key's value.
- ReadEliasFanoChecked replays the jump table against the upper bits instead of
  only counting them, in the same pass. Over every single-bit mutation of an
  index, Open used to accept 192 files that then panicked in a lookup; now none.
  FuzzOpen and a deterministic sweep call Get on what Open accepts.
- A positional lookup that misses means the .v disagrees with the .ef indexing
  it, not that the key is absent, so history seek and the range iterators report
  it instead of falling through to the DB or dropping the key.
- The .vi is madvised again; it stopped being hinted when the accessor moved off
  FilesItem.index.
- OpenHistoryValueIndex takes AccessorVI.Current, the version the file name is
  built from, rather than a literal.
- NewWriter rejects a run count without an item count, which left nil builders
  for AddRun to dereference.
- Index.Run goes: it is wrong for a sequence ending in two or more empty runs,
  and has no caller until the txn->block index.
- ReadEliasFano keeps its doc comment.
@AskAlexSharov

Copy link
Copy Markdown
Collaborator Author

@awskii both confirmed, and both are in the description now.

Existing datadirs. You read it right. missedMapAccessors asks MatchVersionedFile for any .vi, so a present v1 file is never missed and never rebuilt; min stays v1.0 so MustSupport lets it open, and OpenHistoryValueIndex reads it as legacy forever. The path is delete *.vi and restart — that makes the files missed, and BuildMissedAccessors rebuilds them at v2.0 — or wait for the next snapshot release. 40.11 GB -> 549 MB only lands after one of those.

I did not make a below-current .vi count as missed. That would rebuild automatically, but it is a full pass over every .v on the first start after upgrade, which is a release-note decision rather than something to slip into this PR. Say the word if you want it here.

The lo field. Also right, and it is now called out separately. The seek cache is keyed on hi alone, so a colliding hi returned another key's txNum rather than a miss, before this PR too — about 4096/2^64 per lookup.

What changes is the consequence. Under a txNum+key .vi a bogus txNum missed the perfect hash and the query fell through to the DB. Under a positional .vi it resolves that key's ordinal against this key's file and returns a plausible wrong value. Same for the ordering bug you found one comment over: pre-PR it fed Lookup2 a bogus key and missed; now it returns run 0 item 0.

Pushed as 9133a0c on top of the two from the earlier round. Every other thread answered inline.

…go.mod

gofmt aligns the entries around a long key differently in 1.27 than in 1.26,
and go.mod names 1.26, so reformatting with a newer local toolchain is what CI
rejects.
# Conflicts:
#	db/integrity/integrity_action_type.go
#	db/state/history.go
# Conflicts:
#	db/recsplit/index_reader.go
#	db/state/history.go
#	db/state/inverted_index.go
#	db/state/merge.go
# Conflicts:
#	db/state/inverted_index_test.go

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants