Repository navigation
Conversation
The hot events index held every term of a chunk as an in-memory bitmap and wrote one events_index key per posting into each ledger's batch. Its memory grew with the number of distinct terms in the chunk, and the per-posting keys were most of each commit. Event ids are now cut into slabs of 65,536, the matcher's unit. The slab being filled is a fixed array of postings with per-bucket chains, which one writer appends to and readers walk without a lock. When the next slab starts, the full one is sealed in the background: its terms are written in key order as one file and loaded into events_index under slab || term. events_index has auto compaction off, and its files never overlap so compaction would have nothing to do; their index blocks live in the block cache. The writer waits for the previous seal before it starts the next one, so the index holds at most two slabs in memory whatever the chunk holds, and the ledger commit carries no index bytes any more. A failed seal keeps its slab readable from memory and fails the next ledger's apply, through the apply hook, which can now fail; the ledger report then names PhaseApply, and the write phases carry their item counts only once the commit landed. The next open indexes the unsealed events again from events_data, through the same path ingestion uses, and seals a full slab among them as soon as the next one starts. LookupKeys reads the slabs the window touches, from memory or from events_index, and returns a bitmap per key, empty rather than nil, since it has not seen the rest of the chunk. The bench sink counts an ingest total only for a ledger whose apply succeeded. The design documents describe the index as it is now. Measured with bench-ingest hot on the same synthetic ledgers (6,000 events of 7 terms each per ledger) at a 600 ms cadence, 600 ledgers, two runs each: per-ledger ingest p50 139 and 144 ms before, 70 and 71 ms after; p99 288 and 319 ms before, 82 and 88 ms after; the events, commit and apply phases went from 12, 72 and 11 ms to 4, 21 and 1 ms at p50; peak RSS 1,373 and 1,379 MB before, 549 and 545 MB after. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BiCL19kUrno3Z4SAYgpAfh
tamirms
force-pushed
the
tamirms/events-hot-index-slabs
branch
from
October 6, 2026 05:18
0718fc5 to
c2bdf2c
Compare
tamirms
added this pull request to stack #1089
October 6, 2026 09:22
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The problem
Indexing events for getEvents was both the biggest memory risk in ingestion and most of the cost of every ledger commit.
The hot events index kept one bitmap per term (every contract id, topic and event type seen in the chunk) in memory, listing the events that carry it, for the whole chunk, in every open hot database. That grows with the number of distinct terms, and a chunk can hold tens of millions of them under adversarial traffic, so there was no bound. And every ledger's commit wrote one RocksDB key per posting (one term on one event), about 6,000 keys per ledger on pubnet, through the write-ahead log, the memtable and later compaction. Those keys were most of each commit.
The idea
Keep the index out of RocksDB's write path entirely, and keep only a fixed-size slice of it in memory.
Event ids are cut into slabs of 65,536. Postings for the slab being filled go into a fixed in-memory buffer. When the slab is full, it is sealed: its terms are written in key order as one SST file, RocksDB's own sorted file format, and handed to RocksDB with
IngestExternalFile, which adds the file to the index column family as it is. Every key in the file starts with the slab number, so no two slabs overlap, and RocksDB places each file straight at the bottom level. Nothing goes through the write-ahead log or the memtable, and nothing is ever flushed or compacted. Readers see the whole slab at once, and it is durable the moment the call returns.Why ingestion gets faster
The ledger commit carries no index bytes at all. Indexing a ledger's events is appending postings to the in-memory buffer, and the only disk work for the index is one sequential file write per 65,536 events, done in the background. In the measurements below, the commit phase of a ledger went from 72 ms to 21 ms at p50.
Why memory is bounded
A slab is 65,536 events and an event carries at most 7 terms, so the live buffer holds at most 458,752 postings, whatever the number of distinct terms in the chunk or events in a ledger. At most two slabs are in memory, the one being filled and the one being sealed: the writer waits for the previous seal before starting the next, so a slow disk stalls ingestion at a slab boundary instead of letting slabs pile up. The sealed files' index and filter blocks live in RocksDB's block cache, which has a fixed size, not on the heap.
What it costs
Queries read sealed slabs from RocksDB instead of from memory. A lookup over a whole chunk of 91 sealed slabs takes 0.2 to 0.6 ms per term with a warm block cache, where the in-memory index answered in microseconds. Query latency was the lowest priority here, after memory and ingestion latency.
Failure and recovery
A seal runs in the background, so its failure is reported through ingestion. If a seal fails, the slab's postings are still in memory, so queries keep working, and the next ledger's apply returns the error, so ingestion stops rather than continuing with a slab missing from the index. On restart, the slabs that were not sealed are indexed again from the stored event data and sealed again.
Alternatives measured
The engine in #902 bounds memory only as a function of ledger density and term count, and spends memory on query speed. One in-memory bitmap per slab is bounded but about five times slower to update. Writing sealed slabs through RocksDB's normal write path stalls commits on flushes, and term-first keys need compaction. The sealed-slab file was the only option that is both bounded by construction and faster to ingest.
Results
bench-ingest hoton synthetic ledgers of 6,000 events with 7 terms each, 600 ledgers at a 600 ms cadence, two runs each, before and after:On a real pubnet chunk of 3,000 ledgers, p50 went from 19 to 9 ms and p99 from 38 to 16 ms. A kill -9 harness that ingests in a child process, kills it at a random moment, reopens and verifies every term passed 100 of 100 rounds, a third of them with the kill landing during a seal.
🤖 Generated with Claude Code