Skip to content

events: seal the hot index one slab at a time - #1097

Draft
tamirms wants to merge 1 commit into
tamirms/rocksdb-bulk-loadfrom
tamirms/events-hot-index-slabs
Draft

tamirms wants to merge 1 commit into
tamirms/rocksdb-bulk-loadfrom
tamirms/events-hot-index-slabs

Conversation

@tamirms

@tamirms tamirms commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

The problem

Indexing events for getEvents was both the biggest memory risk in ingestion and most of the cost of every ledger commit.

The hot events index kept one bitmap per term (every contract id, topic and event type seen in the chunk) in memory, listing the events that carry it, for the whole chunk, in every open hot database. That grows with the number of distinct terms, and a chunk can hold tens of millions of them under adversarial traffic, so there was no bound. And every ledger's commit wrote one RocksDB key per posting (one term on one event), about 6,000 keys per ledger on pubnet, through the write-ahead log, the memtable and later compaction. Those keys were most of each commit.

The idea

Keep the index out of RocksDB's write path entirely, and keep only a fixed-size slice of it in memory.

Event ids are cut into slabs of 65,536. Postings for the slab being filled go into a fixed in-memory buffer. When the slab is full, it is sealed: its terms are written in key order as one SST file, RocksDB's own sorted file format, and handed to RocksDB with IngestExternalFile, which adds the file to the index column family as it is. Every key in the file starts with the slab number, so no two slabs overlap, and RocksDB places each file straight at the bottom level. Nothing goes through the write-ahead log or the memtable, and nothing is ever flushed or compacted. Readers see the whole slab at once, and it is durable the moment the call returns.

Why ingestion gets faster

The ledger commit carries no index bytes at all. Indexing a ledger's events is appending postings to the in-memory buffer, and the only disk work for the index is one sequential file write per 65,536 events, done in the background. In the measurements below, the commit phase of a ledger went from 72 ms to 21 ms at p50.

Why memory is bounded

A slab is 65,536 events and an event carries at most 7 terms, so the live buffer holds at most 458,752 postings, whatever the number of distinct terms in the chunk or events in a ledger. At most two slabs are in memory, the one being filled and the one being sealed: the writer waits for the previous seal before starting the next, so a slow disk stalls ingestion at a slab boundary instead of letting slabs pile up. The sealed files' index and filter blocks live in RocksDB's block cache, which has a fixed size, not on the heap.

What it costs

Queries read sealed slabs from RocksDB instead of from memory. A lookup over a whole chunk of 91 sealed slabs takes 0.2 to 0.6 ms per term with a warm block cache, where the in-memory index answered in microseconds. Query latency was the lowest priority here, after memory and ingestion latency.

Failure and recovery

A seal runs in the background, so its failure is reported through ingestion. If a seal fails, the slab's postings are still in memory, so queries keep working, and the next ledger's apply returns the error, so ingestion stops rather than continuing with a slab missing from the index. On restart, the slabs that were not sealed are indexed again from the stored event data and sealed again.

Alternatives measured

The engine in #902 bounds memory only as a function of ledger density and term count, and spends memory on query speed. One in-memory bitmap per slab is bounded but about five times slower to update. Writing sealed slabs through RocksDB's normal write path stalls commits on flushes, and term-first keys need compaction. The sealed-slab file was the only option that is both bounded by construction and faster to ingest.

Results

bench-ingest hot on synthetic ledgers of 6,000 events with 7 terms each, 600 ledgers at a 600 ms cadence, two runs each, before and after:

before after
ingest per ledger, p50 139 and 144 ms 70 and 71 ms
ingest per ledger, p99 288 and 319 ms 82 and 88 ms
commit phase, p50 72 ms 21 ms
peak RSS 1,373 and 1,379 MB 549 and 545 MB
peak RSS, unpaced over 2,000 ledgers 3,018 MB 655 MB

On a real pubnet chunk of 3,000 ledgers, p50 went from 19 to 9 ms and p99 from 38 to 16 ms. A kill -9 harness that ingests in a child process, kills it at a random moment, reopens and verifies every term passed 100 of 100 rounds, a third of them with the kill landing during a seal.

🤖 Generated with Claude Code

The hot events index held every term of a chunk as an in-memory bitmap
and wrote one events_index key per posting into each ledger's batch. Its
memory grew with the number of distinct terms in the chunk, and the
per-posting keys were most of each commit.

Event ids are now cut into slabs of 65,536, the matcher's unit. The slab
being filled is a fixed array of postings with per-bucket chains, which
one writer appends to and readers walk without a lock. When the next
slab starts, the full one is sealed in the background: its terms are
written in key order as one file and loaded into events_index under
slab || term. events_index has auto compaction off, and its files never
overlap so compaction would have nothing to do; their index blocks live
in the block cache. The writer waits for the previous seal before it
starts the next one, so the index holds at most two slabs in memory
whatever the chunk holds, and the ledger commit carries no index bytes
any more.

A failed seal keeps its slab readable from memory and fails the next
ledger's apply, through the apply hook, which can now fail; the ledger
report then names PhaseApply, and the write phases carry their item
counts only once the commit landed. The next open indexes the unsealed
events again from events_data, through the same path ingestion uses,
and seals a full slab among them as soon as the next one starts.

LookupKeys reads the slabs the window touches, from memory or from
events_index, and returns a bitmap per key, empty rather than nil, since
it has not seen the rest of the chunk. The bench sink counts an ingest
total only for a ledger whose apply succeeded. The design documents
describe the index as it is now.

Measured with bench-ingest hot on the same synthetic ledgers (6,000
events of 7 terms each per ledger) at a 600 ms cadence, 600 ledgers,
two runs each: per-ledger ingest p50 139 and 144 ms before, 70 and 71 ms
after; p99 288 and 319 ms before, 82 and 88 ms after; the events, commit
and apply phases went from 12, 72 and 11 ms to 4, 21 and 1 ms at p50;
peak RSS 1,373 and 1,379 MB before, 549 and 545 MB after.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BiCL19kUrno3Z4SAYgpAfh
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant