fix(hooks): bound pre_tool transcript reads to a tail window - #110
Merged
Merged
Conversation
PreToolUse is the highest-frequency hook (fires on every tool call, and the same code runs daemon-side via /hook/pre_tool across sessions). _extract_thinking and _post_compaction_tail_bytes both read the ENTIRE transcript to find something that always lives near EOF, so five concurrent requests against ~2.1GB Codex rollouts produced ~30GB of in-flight allocations. Both now share a bounded-tail helper (hooks.pre_tool_tail_mb, default 16MB): stat the file, seek to size-cap, read to EOF, discard a partial leading line. A thinking block or compaction marker older than the tail window degrades to not-found, which is the correct behavior for both callers (most-recent-only, and warn-once-at-offset-0 respectively).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Root cause of the recurring multi-GB daemon balloons:
src/simba/hooks/pre_tool_use.pyperformed two unbounded whole-file transcript reads on the highest-frequency hook (PreToolUse fires on every tool call, and the same code runs daemon-side via/hook/pre_toolconcurrently across sessions):_extract_thinking:read_text().strip().split("\n")of the entire transcript — to find the last thinking block, which lives at the end of the file. On multi-GB rollouts this read gigabytes per tool call (and on non-Claude-format transcripts, read them to return"")._post_compaction_tail_bytes:read_bytes()of the entire file torfindthe last compaction marker. Its caller's "cheap path" skips the read only for small files — inverted for exactly the files that hurt.A
malloc_historycapture of a live 30.9GB balloon attributed ~29.7GB to five concurrent in-flight reads of ~2GB transcripts through this path. No config cap governed it.Changes
hooks.pre_tool_tail_mb(default 16.0): maximum transcript bytes read from the END of the file by pre_tool inspectors._read_tail_bytes(path, cap)helper: stat → seek(size−cap) → read to EOF, discarding a partial leading line;cap<=0degrades to uncapped (matching sibling cap idioms)._extract_thinkingscans only the tail (documented tradeoff: a thinking block older than the window is not found — the block of interest is the most recent one)._post_compaction_tail_bytesscans only the tail; marker-not-in-window degrades to the never-compacted(total, 0)case, and the context-low warning still fires correctly there since tail ≥ cap ≥ threshold. Absolute offsets preserved exactly when the marker is in the window.Notes
Follow-up (separate PR): the same bounded-tail treatment for
tailor/hook.py's Stop-hook transcript read andusage_signals.py, sharing this helper.