Skip to content

Stop Clemory views yielding bytes when iterated - #821

Open
zardus wants to merge 1 commit into
masterfrom
feature/clemory-view-iter
Open

zardus wants to merge 1 commit into
masterfrom
feature/clemory-view-iter

Conversation

@zardus

@zardus zardus commented Sep 6, 2026 •

Copy link
Copy Markdown
Member

THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS

Problem

Iterating any of cle's three Clemory views yields the byte stored at each
address instead of the address. On binaries/tests/x86_64/fauxware:

>>> ld = cle.Loader("../binaries/tests/x86_64/fauxware", auto_load_libs=False)
>>> lo = ld.memory.min_addr
>>> list(ld.memory.load(lo, 6))
[127, 69, 76, 70, 2, 1]
>>> list(itertools.islice(cle.ClemoryView(ld.memory, lo, lo + 0x100), 6))
[127, 69, 76, 70, 2, 1]
>>> list(itertools.islice(cle.ClemoryTranslator(ld.memory, lambda k: k + lo), 6))
[127, 69, 76, 70, 2, 1]

Those six values are the first six bytes of the image -- \x7fELF and the class
and data fields after it -- not addresses. ClemoryReadOnlyView, which is what
Loader.memory_ro_view hands out, raises KeyError: 0 instead, and so does a
ClemoryView whose address space does not start at zero.

Root cause

None of the three defines __iter__, and neither does ClemoryBase. Python
therefore falls back to the legacy sequence protocol, which calls obj[0],
obj[1], ... and stops only on IndexError. These classes raise KeyError,
so the loop either runs off the end of the address space or dies on its first
step.

ClemoryBase declares seven methods and raises NotImplementedError in each:
__getitem__, __setitem__, __contains__, load, store, backers and
find. That list has been one short since b75fb69 created the class in 2020 --
__iter__ went to Clemory alone and was never declared on the base, so the
protocol fallback has stood in for it ever since.

Fix

Declare __iter__ on ClemoryBase beside the other seven, so a class that does
not implement it refuses instead of falling through. Clemory keeps its own
implementation and is unaffected.

ClemoryView(mem, min_addr, min_addr + 0x100)
    -> raises NotImplementedError
ClemoryTranslator(mem, lambda a: a + min_addr)
    -> raises NotImplementedError
ClemoryReadOnlyView(arch, mem)
    -> raises NotImplementedError

The other shape is to implement address iteration on the views, and I did not,
for a measured reason rather than a preference. The three observations below are
taken at master 0e77ade3. ClemoryReadOnlyView.backers() is a point query:
asked to enumerate a whole clemory it returns nothing, while
_flattened_backers holds three entries for this binary.
ClemoryView.backers() mis-clamps -- for the view above, 256 addresses wide, it
yields one backer of 2420 bytes starting at 0. And ClemoryTranslator refuses
backers() and find() outright, because address translation gives it no space
to walk. Say the word if you would rather have iteration. #818 fixes the same
class of defect on Clemory itself; the two do not overlap and neither depends
on the other.

Testing

tests/test_clemory_views.py loads binaries/tests/x86_64/fauxware, builds all
three views over loader.memory and asserts that iterating each one raises
NotImplementedError, then checks that Clemory still yields one value per
backed byte, so the base declaration cannot shadow the subclass. It fails on
master, where the first view iterates instead.

Validation: #821 (comment)

session: sharpen

ClemoryView, ClemoryTranslator and ClemoryReadOnlyView define no __iter__, and
neither does ClemoryBase, which declares __getitem__, __setitem__, __contains__,
load, store, backers and find and raises NotImplementedError in each. Python
therefore falls back to the legacy __getitem__ sequence protocol, which calls
obj[0], obj[1], ... and stops only on IndexError. These classes raise KeyError,
so iterating a view whose address space starts at 0 yields the byte stored at
each address, and iterating one that starts anywhere else raises KeyError: 0. On
binaries/tests/x86_64/fauxware the first six values a ClemoryView over
loader.memory yields are [127, 69, 76, 70, 2, 1], which is the start of the ELF
header rather than any address in it.

Declare __iter__ on ClemoryBase beside the other seven, so a class that does not
implement it refuses instead of falling through. Clemory keeps its own
implementation.

The alternative is to implement address iteration on the views, and it is a
larger change than it looks. ClemoryReadOnlyView.backers() is a point query that
yields nothing when asked to enumerate a whole clemory, ClemoryView.backers()
mis-clamps a backer that runs past the end of the view, and ClemoryTranslator
refuses backers() and find() outright because it has no address space to walk.
That belongs in its own change.
@zardus

zardus commented Sep 6, 2026

Copy link
Copy Markdown
Member Author

THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS

What each of the three views yields when iterated, before and after this change.
binaries/tests/x86_64/fauxware loaded with auto_load_libs=False, mem is
loader.memory and min_addr is 0x400000.

Before — two of the three walk off into the image's bytes; the third dies on
its first step:

angr/cle master 0e77ade
loader.memory.min_addr           : 0x400000
first 6 bytes of the image       : [127, 69, 76, 70, 2, 1]
first 6 addresses in that range  : [4194304, 4194305, 4194306, 4194307, 4194308, 4194309]
ClemoryView(mem, min_addr, min_addr + 0x100)
    -> [127, 69, 76, 70, 2, 1]
ClemoryTranslator(mem, lambda a: a + min_addr)
    -> [127, 69, 76, 70, 2, 1]
ClemoryReadOnlyView(arch, mem)
    -> raises KeyError: 0

After — all three refuse, and say so:

with this change
loader.memory.min_addr           : 0x400000
first 6 bytes of the image       : [127, 69, 76, 70, 2, 1]
first 6 addresses in that range  : [4194304, 4194305, 4194306, 4194307, 4194308, 4194309]
ClemoryView(mem, min_addr, min_addr + 0x100)
    -> raises NotImplementedError
ClemoryTranslator(mem, lambda a: a + min_addr)
    -> raises NotImplementedError
ClemoryReadOnlyView(arch, mem)
    -> raises NotImplementedError

@zardus

zardus commented Sep 6, 2026

Copy link
Copy Markdown
Member Author

THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS

Validation record for head 91ed3fa8f20eadd115ca158ee6a3ed1df5ea3a59 against
baseline 0e77ade3c39a3cee05f65051e57955675e1ac21b.

  • Regression: python -m pytest tests/test_clemory_views.py — fails on the
    baseline cle/memory.py (ClemoryView still falls back to the sequence protocol), passes on head
  • Full suite: python -m pytest tests — baseline 261 passed, 9 skipped; head 262
    passed, 9 skipped; 0 failed either side
  • Lint/type: ruff check 0.16.5 clean and black 26.5.1, the formatter
    .pre-commit-config.yaml pins, leaves both files unchanged; pylint per changed
    file 10.00 -> 10.00 on cle/memory.py and 10.00 on the new test file; pyright
    errors per changed file 0 -> 0 on both
  • Fixture check: check-test-inputs.py and check-stale-pins.py both exit 0
    against the head; angr/binaries at 003e82a2, no fixture added

Two mutations of the fix each fail the new test: returning an empty iterator from
ClemoryBase.__iter__ instead of raising, and deleting Clemory.__iter__ so the
new declaration shadows it.

No caller iterates any of the three views. Neither ClemoryView nor
ClemoryTranslator is constructed anywhere in cle 0e77ade3, angr 87411a71 or
angr-management aa843e5c. ClemoryReadOnlyView is constructed in exactly one
place, Loader.gen_ro_memview, which CFGFast calls, and nothing that handles
the object it caches iterates it: Loader.fast_memory_load_pointer calls
unpack_word, Block.bytes calls load, Block._lift_nocache passes it to the
lifter, and Clinic and both lifters call next(clemory.backers(addr)).
Replacing ClemoryBase.__iter__ with a recorder gives 2 calls on a calibration
that really does iterate two views, and 0 across Project + gen_ro_memview() +
CFGFast(normalize=True) + 30 steps on binaries/tests/x86_64/fauxware, which
recovers 40 functions; that run imported angr 854269112.

Caveats: the workspace's complete gate could not be run for this branch, because
entering its shell rebuilds the native libraries under six live corpus sweep
lanes. The suites above were run against a standalone worktree on the shared
interpreter. No corpus sweep was run, because no caller reaches the changed
code.

@angr-bot

angr-bot commented Sep 6, 2026

Copy link
Copy Markdown
Member

Corpus decompilation diffs can be found at angr/dec-snapshots@master...angr/cle_821

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants