Raised in CNCF TAG-Security review of the Hive security self-assessment (cncf/toc#2286), on the ioscan section:
have you red teamed to know how well this works? What should a potential user be aware of here? Is this perfect defense, works decently well, a partial mitigation, etc.?
The honest current answer
No red-teaming has been done. A search of the repository for red-team, adversarial-evaluation or penetration-test artifacts returns nothing relevant. src/pkg/ioscan has 56 unit tests across its rule, Unicode, classifier and canary paths, but unit tests establish that known shapes are caught — they say nothing about unknown ones, and they cannot produce a detection rate.
This matters because ioscan is the control most likely to be read as a prompt-injection defense, when the thing actually containing the consequence of a successful injection is the proxy's hard-deny of POST /pulls and PUT /pulls/{n}/merge for every ACMM mode.
Task
Run and publish an adversarial evaluation of the injection path:
- Build a corpus of prompt-injection attempts against the kick path — at minimum: direct instruction override, Unicode steganography (invisible controls, tag characters, variation selectors), homoglyph substitution, base64 and multi-layer encoding, instruction smuggling via code blocks/diffs/labels/author fields, and multi-turn/split-payload attempts across issue + comments.
- Measure what the deterministic rules catch, what the optional classifier adds, and what gets through.
- For anything that gets through, determine whether the network-layer deny rules still contain it — that distinction is the real finding.
- Publish the methodology and results, including the miss rate.
Done when
- A red-team evaluation document exists in
src/docs/
security-self-assessment.md cites measured results instead of describing efficacy as unmeasured
- Any rule gaps found are filed as their own issues
Raised in CNCF TAG-Security review of the Hive security self-assessment (cncf/toc#2286), on the
ioscansection:The honest current answer
No red-teaming has been done. A search of the repository for red-team, adversarial-evaluation or penetration-test artifacts returns nothing relevant.
src/pkg/ioscanhas 56 unit tests across its rule, Unicode, classifier and canary paths, but unit tests establish that known shapes are caught — they say nothing about unknown ones, and they cannot produce a detection rate.This matters because
ioscanis the control most likely to be read as a prompt-injection defense, when the thing actually containing the consequence of a successful injection is the proxy's hard-deny ofPOST /pullsandPUT /pulls/{n}/mergefor every ACMM mode.Task
Run and publish an adversarial evaluation of the injection path:
Done when
src/docs/security-self-assessment.mdcites measured results instead of describing efficacy as unmeasured