Conversation
To cancel the effect of edge removal, which tends to remove source cycles. Every block improves, and package-graph recursion is now reached by mutated scenarios routinely rather than almost exclusively by the seeds. | block | recovery single | recovery restarts | `Package::resolve` single | restarts | |---|---|---|---|---| | 0 | 32.5% | 67.8% | 0.1% | 3.3% | | 1 | 48.5% | 74.8% | 0.4% | 2.7% | | 2 | 43.6% | 68.9% | 0.1% | 2.0% | | 3 | 47.9% | 74.3% | 0.2% | 2.0% | | 4 | 37.2% | 69.8% | 1.2% | 6.0% | | 5 | 37.1% | 79.8% | 1.4% | 3.7% | | **all** | **41.2%** | **72.5%** | **0.6%** | **3.3%** |
lionel-
added this pull request to stack #1412
September 23, 2026 06:51
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1409
This PR adds structured fuzz testing of R programs that might be susceptible to cause cyclic panics in the analysis engine. It generates small R workspaces, chooses an initial analysis query, and executes a history of queries and edits. The main goal is to increase our confidence that we have implemented the necessary cycle recovery handlers for our Salsa queries, and uncover cases where the cycle depends on a particular starting point or succession of edits. The fuzz tests also generally increase our test coverage for panics and for Salsa invariants.
Much of the size of this PR is in the modelling of scenarios, and R code generation and mutations, all of it gated behind the
fuzzfeature, so should be pretty low-risk overall.The harness detects:
All panics, including failures unrelated to Salsa cycles.
Inconsistencies in symbol resolution between an incrementally updated database with queries hitting the Salsa cache and a fresh database. This consistency check can catch various issues in the implementation of Salsa queries (in dependency tracking, identity bookkeeping, etc).
Timeouts, which can expose severe performance regressions, e.g. if we introduce non-linear complexity that only surfaces depending on query/edit order.
How it works:
Test scenarios model an initial workspace configuration, an initial cold query to kick off an analysis, and a history of subsequent queries and edits. The cold query runs before any other analysis has warmed the database.
Initial scenarios include patterns/motifs of source dependencies: chains, self-loops, mutual pairs, rings, and overlapping cycles. These represent relationships between generated R files, intended to exercise different dependency/graph patterns in the underlying Salsa queries. The workspace generator also models package metadata and various file loaders (
R/directory in a package, and testthat and Shiny layouts).The R generator models the various effects supported by the analysis engine, in particular Attach and Source effects which are the main source of cycles. It also includes effects such as Assign and Quote, as well as shadowing and nesting. These are rendered as real R programs and passed through the whole Oak pipeline.
A mutator (fuzzing jargon) then randomly varies the scenarios: adding source calls, renaming or shadowing identifiers, changing package exports, adding or removing files, and changing the cold query or edit history.
We explore these scenarios in two ways. Fixed-seed runs provide reproducible checks in development and PR tests. Coverage-guided exploration uses cargo fuzz and libFuzzer to retain inputs that reach new code paths and mutate them further. Both these approaches use the same scenario model, runner, and structured mutation logic.
Here is an example of a generated program and edit/query history. An edit introduces a self-source inside a function, and the following imports query invokes cycle recovery:
You can inspect generated scenarios such as this one with
just fuzz-replay-seed 0.Workflow:
just fuzzruns six fixed-seed "mutation blocks". A block is a reproducible batch of tests. A set of starting scenarios are generated from a fixed random seed, then randomly mutated and run.just fuzz-exploreruns "coverage-guided exploration": it observes which code paths each scenario exercises, keeps only scenarios that add coverage (in terms of Rust code coverage), and mutates them further. These scenarios accumulate in a local corpus (a saved collection of inputs for future runs). The CI fuzzing campaign uses the same exploration approach and retain the corpus between successful runs, so later CI runs build on earlier discoveries.just fuzz-sweep-semantic path/to/corpuschecks saved scenarios for incremental/fresh consistency.just fuzz-replay-scenario path/to/scenario.jsonreplays a saved scenario with operation and recovery tracing. Usejust fuzz-replay-semantic path/to/scenario.jsonto reproduce a consistency mismatch.In CI:
Pull requests check for panics with the fixed-seed mutation blocks and by replaying the saved corpus when available. Corpus replay also checks incremental/fresh consistency.
Pushes to main and manual runs also run a five-minute coverage-guided exploration campaign. Scheduled campaigns explore for thirty minutes twice a week.
CI exploration resumes from the accumulated corpus, which is stored in the GitHub Actions cache. We run the campaign twice a week to keep that cache alive, but it could still expire or be evicted. In the future we'll likely want to store the corpus in an
oak-cioroak-qarepository to ensure we don't lose it.The fuzzing guide at https://github.com/posit-dev/ark/blob/oak-query/fuzz/crates/oak_db/fuzz/README.md documents the detailed commands and CI setup.
Positron Release Notes
New Features
Bug Fixes