Human-driven security review workflow using an agentic coding tool as an analysis assistant, with deterministic scanning as the detection layer.
Mode: manual. The engineer initiates every analysis. Nothing runs unattended, nothing is triggered by repository events, and no agent holds repository write access or CI credentials.
A repeatable procedure, not a pipeline. It gives you:
- A deterministic scan producing SARIF (Semgrep). SCA and SBOM tools are
optional and are run separately;
make scandoes not invoke them. - A structured triage procedure where an agent explains and classifies a finding.
- A separate adversarial validation pass that tries to refute the verdict.
- A normalized verdict record in
findings/reviewed.sarif, ready to import into a findings store. The DefectDojo importer is a placeholder; see Status. - Metrics that tell you whether this is working before you commit to automating it.
- Not a CI integration. See
docs/07-graduation.mdfor the criteria to move there. - Not a guarantee of coverage. Nothing is reviewed unless a human asks for it.
- Not a replacement for the deterministic scanners. It sits on top of them.
The agent never processes untrusted repository content (issue bodies, PR
descriptions, comments from external contributors), so the prompt-injection and
credential-exposure surface that affects event-triggered CI agents does not apply
here. The only channel out of the perimeter is what the engineer explicitly puts
in context, which is auditable and enforced by scripts/check_scope.py.
# 0. Read these first
# docs/02-setup.md - authentication, which account, what to verify
# docs/03-data-handling.md - what leaves the perimeter and what must not
make install SEMGREP_VERSION=1.2.3 # install deterministic tooling
make scan # semgrep -> findings/raw.sarif
make triage N=0 # build a review packet for finding #0make triage writes a self-contained packet to findings/packets/. You open the
agent, load prompts/triage-finding.md together with the packet, and get back a
verdict. Then:
make validate ID=F-abc123 # build the adversarial challenge packet
make record ID=F-abc123 AGENT=<triage model> VALIDATOR=<validation model> MINUTES=12
make decide ID=F-abc123 DECISION=accepted BY=appsec-oncall
make metrics # how the programme is actually performingmake record refuses when AGENT and VALIDATOR name the same model. That is
the independence mechanism, not a preference — a second pass on the same model
inherits the same blind spots.
Working: scan, triage queue, packet builder, scope enforcement, verdict recording, human decisions, metrics, tests.
Not implemented. These are placeholders and will tell you so when run:
goldenset/run_goldenset.py— no regression cases are configured. Until they are, the graduation criterion that depends on them cannot be met; seedocs/07-graduation.md.scripts/import_defectdojo.py— no DefectDojo credentials or import settings. Reviewed findings are infindings/reviewed.sarifand can be imported by hand.
make testStandard library only, no network, no scanner required. If it needs a
pip install to run, that is a bug.
make list shows every finding. On a real repository that is hundreds, so it
filters:
make list LEVEL=error UNREVIEWED=1
make list RULE=injection LIMIT=20
make triage ID=F-abc123Filters work on what the SARIF carries — scanner level, rule id, path. They
deliberately do not filter on policy severity or priority class: those are
assigned by the agent during triage and do not exist yet at listing time.
Indices shift when filters change, so use ID= for anything you want to
reproduce or put in a ticket.
| Path | Contents |
|---|---|
CLAUDE.md |
Project context the agent reads at session start |
docs/ |
Procedures, data handling, design decisions, controls, graduation |
prompts/ |
The five review procedures, one file each |
policy/ |
Excluded paths and severity policy. Both are enforced by scripts |
scripts/ |
Scan, scope check, packet builder, verdict recorder, metrics |
rules/ |
Local Semgrep rules. Empty by default; make scan refuses without them |
goldenset/ |
Intended for regression cases. Placeholder, see Status |
templates/ |
Verdict and challenge schemas, threat-model template |
findings/ |
Local working directory. Git-ignored |
Documentation, in reading order:
| File | Read it when |
|---|---|
docs/01-getting-started.md |
First: the ordered path to your first review |
docs/02-setup.md |
Before the first run: accounts, retention, scanner pin |
docs/03-data-handling.md |
Before the first run: what leaves the perimeter |
docs/04-procedures.md |
Choosing which of the four procedures fits |
docs/05-design-decisions.md |
Changing a control, to see what you would undo |
docs/06-controls.md |
Auditing: what each control refuses and what proves it |
docs/07-graduation.md |
Deciding whether to keep going or automate |
The numbering has gaps where documents were never needed.
- Python 3.9+
- Semgrep at the version recorded in
docs/02-setup.md - Optional: Trivy or Grype, Syft, DefectDojo instance
- An agentic coding tool authenticated under commercial terms (see
docs/02-setup.md)
- Identity change: a scanner fingerprint that cannot carry identity — an empty
value, a non-string, or a message such as Semgrep's
requires loginwhen it runs without an account — is now rejected, and identity falls back to file content. Findings recorded before this change under such a fingerprint had every location for one rule in one file collapsed into a single id, so their ids move.ID_SCHEMEis deliberately not bumped, because the log's scheme guard cannot tell the two apart; if you recorded verdicts against a scanner that emitted placeholder fingerprints, archivefindings/and retriage. - Breaking audit-log change:
recorded_atnow followsreviewed_atin the CSV header. Writer-stamped UTC recording time controls revision ordering, decision supersession and--since; the model'sreviewed_atis retained as self-reported history. Recorders refuse older headers. Archive the oldfindings/directory or retriage before recording with the new contract. - Breaking challenge-schema change: every challenge must carry the exact
16-hex
verdict_digestfrom its challenge packet. Old challenges without the field are refused; a challenge for another revision is refused before writing. Rebuild the packet and rerun validation whenever the verdict changes. findings/reviewed.sarifaccumulates and is never pruned. Over a long programme it grows without bound; rotate it yourself if that becomes a problem.make check-templatefails in this repository on purpose:CLAUDE.mdhere is the template you copy and fill in per project. The check proves only that no template placeholders remain; it does not judge whether the replacement text meaningfully describes the project.- A scanner upgrade can change finding ids. See
docs/02-setup.md.
Not yet chosen. The prompts, docs and scripts here are original work; nothing is vendored from another project. Pick a license before anyone else depends on this.