fix(rlm): honest retained-bytes accounting for eager documents - #111
Merged
Merged
Conversation
The eager _Document.retained_bytes() estimated retention as sys.getsizeof(text) * 2, which ignores the per-line string object overhead in .lines and the per-int object overhead in .line_starts. Both scale with line count, not text size, so many-short-line documents under-counted real retention 3-5x in production (a live daemon capture showed ~4.7GB of split-line strings on a store that believed it was within budget). Result: the DocumentStore LRU under-evicted and the store_budget_mb cap lied. Replace the flat estimate with an honest sum computed once at ingest: text + (.lines list + every line's sizeof) + (.line_starts list + every int's sizeof). Lazy-mode accounting (packed offset-index arrays, no text) was already honest and is unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The eager
_Documentmode (files ≤rlm.max_document_mb) retainstext,lines = text.split("\n"), andline_starts— but theDocumentStoreLRU accounted its retained bytes assys.getsizeof(text) * 2, which ignores per-line string object overhead, the list objects, and (as it turns out)line_startsbeing a plainlist[int]with per-int object overhead. On short-lined content the real retention is several times the accounted value (11.7× on the test input; a live capture showed ~4.7GB of split-line strings on a store that believed it was within budget), sorlm.store_budget_mbunder-evicted and the budget lied.Changes
src/simba/rlm/context.py— eager docs compute an honest retained-bytes sum once at ingest (O(n_lines)):getsizeof(text) + getsizeof(lines) + Σ getsizeof(line) + getsizeof(line_starts) + Σ getsizeof(n). Lazy-mode accounting (packedarray('Q'), no text) was verified already honest and left untouched. Docstrings updated from the old "~2× text size" description.tests/rlm/test_context.py— red-first: an accounting-honesty test (50k×3-char lines; old estimate 400,080 B vs true 4,689,592 B) and an eviction test where the old estimate admitted two docs but honest accounting forces the first to demote to lazy.Notes
No new config; this makes the existing
rlm.store_budget_mbknob enforce what it claims.