Autonomous malware analysis pipeline for Linux, Windows, and Android samples. It orchestrates triage, static analysis, deep reverse engineering, optional sandbox execution, and final report synthesis into markdown/PDF deliverables.
- Identifies sample type/OS and routes through a domain-aware stage pipeline.
- Runs intelligence + reverse engineering helpers (VirusTotal, Ghidra, JADX, entropy/crypto scans).
- Executes Linux samples in a hardened Docker sandbox for runtime evidence.
- Produces a consolidated report with classification, ATT&CK mapping, IOCs, findings, detections, and appendix.
Curated public outputs are stored in examples/. Runtime outputs are written to reports/.
./setup.shThis script:
- creates/activates
.venv - installs Python dependencies
- prompts for required env vars and writes
.env - clones/updates TheZoo into
samples/theZoo - attempts
git lfs pullwhengit-lfsis installed - builds the Linux sandbox image
malware-sandbox-linux
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtRequired environment variables:
export MALREV_API_KEY="<your-openai-compatible-key>"
export VT_API_KEY="<your-virustotal-key>"Optional:
export MALREV_MODEL="gpt-5.4-mini"
export GHIDRA_HOME="/opt/ghidra"
export OLLAMA_HOST="http://localhost:11434"python main.py /absolute/path/to/sample --verboseThe repository uses a stage-based pipeline (new system only):
triagenano_tagstaticdeepdeep_behavioral(static-inferred behavior for non-Linux)dynamicdebugcross(correlation/intermediate synthesis)- final synthesis + optional PDF render
main.pycallspipeline.runner.run_new_pipeline_to_report()- pipeline stages update a shared
SampleState - final synthesis agent consumes stage outputs and emits canonical report markdown
tools/pdf_report.pyrenders branded PDF from markdown
main.py: CLI entrypoint and runtime orchestration kickoff.pipeline/runner.py: stage execution order, synthesis wiring, report/meta writing.
pipeline/sample_state.py: shared state model (SampleState, findings, tags, outputs).pipeline/pipeline_config.py: model/API config and stage-level model wrappers.pipeline/stages/:triage.py: file identification + initial classification context.nano_tag.py: compact semantic tagging.static_analysis.py: VT/static enrichment.deep_analysis.py: deep static with Ghidra/JADX workflows.deep_behavioral.py: static-inferred runtime chain when dynamic is limited.dynamic_analysis.py: sandbox-based behavior capture.debug_analysis.py: debugger-oriented deep inspection stage.cross_reference.py: correlation and intermediate synthesis support.
agents/src/triage_agent.py: triage agent configuration.agents/src/static_agent.py: static-analysis agent + tool set.agents/src/deep_static_agent.py: deep static agent + Ghidra/JADX tools.agents/src/dynamic_agent.py: dynamic-analysis agent + sandbox tools.agents/src/debug_agent.py: debugger-focused agent.agents/src/final_report_agent.py: final canonical report synthesis agent.agents/prompts/*.txt: stage/final instructions templates.
tools/virustotal.py: VT scanning/intel retrieval.tools/file_inspector.py: file type/hash/PE entropy/crypto constant analysis.tools/ghidra.py: Ghidra-powered static extraction/decompile functions.tools/jadx_tools.py: APK decompilation/search/manifest extraction.tools/docker_runner.py: sandbox execution wrappers and artifact collection.tools/pdf_report.py: markdown -> styled PDF renderer.
samples/: local sample corpus (including TheZoo checkout).reports/: generated runtime reports from new analyses.examples/: curated, public-safe example report sets.
tests/unit/: unit coverage for config/tool/pipeline behavior.tests/e2e/: corpus-driven end-to-end tests (env-gated).
- Python 3.10+
- Docker (required for Linux dynamic sandbox stage)
- Ghidra (for deep static native analysis)
- JADX (for APK deep static workflows)
- VirusTotal API key (
VT_API_KEY) - OpenAI-compatible API key (
MALREV_API_KEY)