Skip to content

Add fuzz-testing for Oak's analysis engine - #1418

Open
lionel- wants to merge 44 commits into
oak-query/hardeningfrom
oak-query/fuzz
Open

lionel- wants to merge 44 commits into
oak-query/hardeningfrom
oak-query/fuzz

Conversation

@lionel-

@lionel- lionel- commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Closes #1409

This PR adds structured fuzz testing of R programs that might be susceptible to cause cyclic panics in the analysis engine. It generates small R workspaces, chooses an initial analysis query, and executes a history of queries and edits. The main goal is to increase our confidence that we have implemented the necessary cycle recovery handlers for our Salsa queries, and uncover cases where the cycle depends on a particular starting point or succession of edits. The fuzz tests also generally increase our test coverage for panics and for Salsa invariants.

Much of the size of this PR is in the modelling of scenarios, and R code generation and mutations, all of it gated behind the fuzz feature, so should be pretty low-risk overall.

The harness detects:

  • All panics, including failures unrelated to Salsa cycles.

  • Inconsistencies in symbol resolution between an incrementally updated database with queries hitting the Salsa cache and a fresh database. This consistency check can catch various issues in the implementation of Salsa queries (in dependency tracking, identity bookkeeping, etc).

  • Timeouts, which can expose severe performance regressions, e.g. if we introduce non-linear complexity that only surfaces depending on query/edit order.
    How it works:

  • Test scenarios model an initial workspace configuration, an initial cold query to kick off an analysis, and a history of subsequent queries and edits. The cold query runs before any other analysis has warmed the database.

  • Initial scenarios include patterns/motifs of source dependencies: chains, self-loops, mutual pairs, rings, and overlapping cycles. These represent relationships between generated R files, intended to exercise different dependency/graph patterns in the underlying Salsa queries. The workspace generator also models package metadata and various file loaders (R/ directory in a package, and testthat and Shiny layouts).

  • The R generator models the various effects supported by the analysis engine, in particular Attach and Source effects which are the main source of cycles. It also includes effects such as Assign and Quote, as well as shadowing and nesting. These are rendered as real R programs and passed through the whole Oak pipeline.

  • A mutator (fuzzing jargon) then randomly varies the scenarios: adding source calls, renaming or shadowing identifiers, changing package exports, adding or removing files, and changing the cold query or edit history.

  • We explore these scenarios in two ways. Fixed-seed runs provide reproducible checks in development and PR tests. Coverage-guided exploration uses cargo fuzz and libFuzzer to retain inputs that reach new code paths and mutate them further. Both these approaches use the same scenario model, runner, and structured mutation logic.

Here is an example of a generated program and edit/query history. An edit introduces a self-source inside a function, and the following imports query invokes cycle recovery:

  op 1: edit [0] -> "val_0 <- 1\nread <- function(...) {\n  library <- function(...) val_4 <- 1\n  library(pkgz)\n  val_0\n  base::source(\"a.R\")\n}\n"
  op 2: imports[0]
      recovered: semantic_index(w/a.R)
      recovered: exports(w/a.R)
      recovered: attached_packages(w/a.R)
      ...

You can inspect generated scenarios such as this one with just fuzz-replay-seed 0.

Workflow:

  • just fuzz runs six fixed-seed "mutation blocks". A block is a reproducible batch of tests. A set of starting scenarios are generated from a fixed random seed, then randomly mutated and run.

  • just fuzz-explore runs "coverage-guided exploration": it observes which code paths each scenario exercises, keeps only scenarios that add coverage (in terms of Rust code coverage), and mutates them further. These scenarios accumulate in a local corpus (a saved collection of inputs for future runs). The CI fuzzing campaign uses the same exploration approach and retain the corpus between successful runs, so later CI runs build on earlier discoveries.

  • just fuzz-sweep-semantic path/to/corpus checks saved scenarios for incremental/fresh consistency.

  • just fuzz-replay-scenario path/to/scenario.json replays a saved scenario with operation and recovery tracing. Use just fuzz-replay-semantic path/to/scenario.json to reproduce a consistency mismatch.

In CI:

  • Pull requests check for panics with the fixed-seed mutation blocks and by replaying the saved corpus when available. Corpus replay also checks incremental/fresh consistency.

  • Pushes to main and manual runs also run a five-minute coverage-guided exploration campaign. Scheduled campaigns explore for thirty minutes twice a week.

CI exploration resumes from the accumulated corpus, which is stored in the GitHub Actions cache. We run the campaign twice a week to keep that cache alive, but it could still expire or be evicted. In the future we'll likely want to store the corpus in an oak-ci or oak-qa repository to ensure we don't lose it.

The fuzzing guide at https://github.com/posit-dev/ark/blob/oak-query/fuzz/crates/oak_db/fuzz/README.md documents the detailed commands and CI setup.

Positron Release Notes

New Features

  • N/A

Bug Fixes

  • N/A

To cancel the effect of edge removal, which tends to remove source cycles.

Every block improves, and package-graph recursion is now reached by
mutated scenarios routinely rather than almost exclusively by the seeds.

| block | recovery single | recovery restarts | `Package::resolve`
single | restarts |
|---|---|---|---|---|
| 0 | 32.5% | 67.8% | 0.1% | 3.3% |
| 1 | 48.5% | 74.8% | 0.4% | 2.7% |
| 2 | 43.6% | 68.9% | 0.1% | 2.0% |
| 3 | 47.9% | 74.3% | 0.2% | 2.0% |
| 4 | 37.2% | 69.8% | 1.2% | 6.0% |
| 5 | 37.1% | 79.8% | 1.4% | 3.7% |
| **all** | **41.2%** | **72.5%** | **0.6%** | **3.3%** |
@lionel-
lionel- added this pull request to stack #1412 September 23, 2026 06:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Harden cycle recovery in Oak analysis

1 participant