Skip to content

Make tape recording thread-safe (fixes #218) - #311

Merged
devmotion merged 6 commits into
masterfrom
dmw/thread-safe-tape
Oct 2, 2026
Merged

devmotion merged 6 commits into
masterfrom
dmw/thread-safe-tape

Conversation

@devmotion

@devmotion devmotion commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Fixes #218.

Problem

InstructionTape was Vector{AbstractInstruction}, and record! called push! on it without synchronization. When the differentiated function spawns tasks (Threads.@spawn, @threads, …) that do tracked arithmetic, several threads push onto the same tape concurrently:

Fix

Lock-free recording. record! wraps the instruction in a node and links it into a backward-linked list with a single @atomicswap on the tape's last field. The swap returns the predecessor, so there is no lock and no second atomic, and recording tasks never wait for each other.

Recording and replay are separate phases. finish! ends recording: it swaps last to a Finished() sentinel and moves the list into a Vector once. After that:

  • reading a tape that is still recording (length, the passes, compiling) throws;
  • recording onto a finished tape throws and leaves the tape unchanged (record! sees Finished() as its predecessor at no extra cost, and restores it before throwing);
  • empty! clears the tape and starts recording again, except for NULL_TAPE (the tape of untracked values), which never records and stays finished.

The tape constructors (GradientTape, JacobianTape, HessianTape) call finish!, so the public API is unchanged. Code that records onto a bare InstructionTape (internal API) now has to call ReverseDiff.finish!(tape) before replaying it.

InstructionTape is no longer an AbstractVector; it implements only length, and the recorded instructions are read with ReverseDiff.instructions(tape), which checks the phase once.

Some packages use this internal API directly: Lux.jl (ReverseDiff training extension), AbstractDifferentiation.jl (derivative) and Gen.jl (backprop) record onto a bare InstructionTape() and call reverse_pass! or iterate it, so they will need a finish! call. Packages that go through GradientTape and friends (e.g. SciMLSensitivity.jl) are unaffected.

Printing. The two-argument show of tapes and instructions prints a single line (7-element InstructionTape, InstructionTape (recording), ScalarInstruction(+)). The text/plain form of a tape lists its instructions and is truncated like an array, instead of printing every instruction of a long tape; that of an instruction shows its input, output and cache.

Why the order is still valid. The swaps put all recordings in one total order that respects each task's program order and every synchronization between tasks. Every rule records its instruction before its output becomes visible. So for each dependency, the instruction producing a value comes before the one consuming it: by program order within a task, and through the synchronization needed to hand the value over across tasks. That holds per dependency, so it doesn't matter whether a task produces, consumes or both. The same argument orders in-place writes against reads in race-free programs. The order isn't deterministic across recordings, so gradients may differ in the last bits.

Performance

Apple Silicon, minimum times unless noted. "64 tasks" is a function that spawns 64 tasks, each doing 200 tracked scalar ops on one tape (median of 100 gradient calls).

master this PR
1.10 tracked x*y 57 ns, 160 B 53 ns, 192 B
1.10 scalar Rosenbrock (n = 10⁴), gradient 8.5 ms 8.8 ms
1.10 replay of a prerecorded tape: Rosenbrock / 270k-instruction tape 8.0 / 8.1 ms 7.8 / 10.7 ms
1.10 64 tasks, 1 / 4 / 8 threads 7.2 ms / crash / crash 7.8 / 7.9 / 8.4 ms
1.13 tracked x*y 249 ns, 160 B 233 ns, 192 B
1.13 scalar Rosenbrock (n = 10⁴), gradient 25.5 ms 25.8 ms
1.13 replay: Rosenbrock / 270k-instruction tape 8.5 / 9.0 ms 8.9 / 9.4 ms
1.13 64 tasks, 1 / 4 / 8 threads 19.5 ms / crash / crash 19.8 / 13.5 / 12.6 ms

The extra 32 bytes per op are the list node. They cost about 10% replay time for very large tapes on 1.10 (allocation layout) and nothing measurable on 1.13.

Alternatives I benchmarked and rejected:

  • Threads.SpinLock around push!: 72 ns per op on 1.10, and 28 / 32 ms for 64 tasks at 8 threads on 1.10 / 1.13, i.e. slower than serial.
  • ReentrantLock: about 2× slower than SpinLock under contention on 1.13.
  • Base.Channel, with or without a consumer task: slower than a lock, since put! takes a lock itself, and a consumer task competes with the recorders for that lock.
  • A doubly linked list as the tape (no Vector): 1.3–2× slower replay due to pointer chasing.
  • Storing the link inside the instructions: dynamic setfield! (the instruction type isn't fully inferred at the call site).
  • Slots claimed with an atomic counter in chunked buffers: false sharing on 1.10.
  • Weaker memory orderings for the swap in record!: :acquire_release performs the same as the default, and :monotonic is 2.5× slower for 64 tasks at 8 threads on 1.10 (4.1–4.4 vs 1.6 ms, reproducible) and the same on 1.13. The contended swap costs the cache line transfer, not the fence.
  • finish! counting the nodes, then resize! and filling from the back (instead of push! + reverse!): about 2× slower on a 270k-instruction tape, since it walks the list twice.

A shared tape still writes one cache line per op from all threads, so recording doesn't reach the speed of fully private per-task tapes (2.8 / 5.2 ms at 8 threads). I'll open a follow-up issue for that.

Tests

  • New testset in test/api/GradientTests.jl: compares the gradient of a function that spawns 64 tasks with a serial reference (repeated, plus a compiled tape). CI now runs with JULIA_NUM_THREADS=4 so this is actually exercised.
  • test/TapeTests.jl pins the phase rules (reading while recording throws, recording after finish! throws, empty! restarts, NULL_TAPE stays finished).
  • Tests that inspect bare tapes finish them first (helper take_recorded! in test/utils.jl).

🤖 Generated with Claude Code

devmotion and others added 2 commits October 1, 2026 09:42
`InstructionTape` was a plain `Vector`, so recording from several threads
(e.g. `Threads.@spawn` inside the differentiated function) raced on `push!`.
On Julia 1.10 this corrupts the heap ("double free"), on 1.11+ it throws a
`ConcurrencyViolationError` or silently yields wrong gradients.

`InstructionTape` is now an `AbstractVector` wrapping the instructions and a
`SpinLock` that guards appending and emptying. Reads stay unlocked.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tape is an append-only log written by `record!`. Implement only the
iteration interface plus `getindex`, `length` and `isempty`, so generic array
functions (`copy`, `similar`, broadcasting, …) don't silently produce unlocked
`Vector`s, and the custom `show` is also used for REPL display.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.52941% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 89.38%. Comparing base (b144aa2) to head (cf0f647).

Files with missing lines Patch % Lines
src/api/tape.jl 85.71% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##           master     #311      +/-   ##
==========================================
+ Coverage   89.18%   89.38%   +0.19%     
==========================================
  Files          19       19              
  Lines        1933     1978      +45     
==========================================
+ Hits         1724     1768      +44     
- Misses        209      210       +1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

devmotion and others added 4 commits October 1, 2026 15:36
Recording takes one atomic swap per instruction on a backward-linked list
instead of a lock, so tasks recording concurrently don't wait for each other.
`finish!` turns the list into the instruction vector; reading a tape that is
still recording throws, and so does recording onto a finished tape, until
`empty!` starts a new recording. The tape constructors finish their tapes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…compactly

`record!` swapped out `Finished()` before throwing, so a rejected record put
the tape back into the recording state. Restore it before throwing.

The two-argument `show` of tapes and instructions now prints one line, and the
verbose `text/plain` form of a tape is truncated like an array. Showing a tape
that is still recording no longer throws.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolve the conflicts with the Runic reformat (#312) by keeping this
branch's changes, formatting them with Runic 1.11.1, and ending functions
with an explicit `return nothing` like master does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`empty!(NULL_TAPE)` (e.g. via `empty!(input.tape)` in `jacobian`) put the
shared null tape back into the recording state, after which reading it
threw. It now leaves `NULL_TAPE` unchanged.

Drop `iterate` and `eltype` for `InstructionTape`: the passes, `compile`
and `show` already read `instructions(tp)`, which checks the phase once.

Compare the threaded gradients with `≈` instead of an absolute tolerance.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@devmotion
devmotion merged commit 089b22b into master Oct 2, 2026
9 checks passed
@devmotion
devmotion deleted the dmw/thread-safe-tape branch October 2, 2026 11:01
devmotion added a commit that referenced this pull request Oct 2, 2026
Finishes the tape before counting its instructions in the elementwise tests,
as recording and reading are now separate phases (#311).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
devmotion added a commit that referenced this pull request Oct 2, 2026
…d, and finish tapes before reading them

- `copyto!` into a 0-dimensional `TrackedArray` was ambiguous with `Base`'s
  0-dimensional method instead of throwing the destination error.
- Replace the tests of `trackresults` and `getpartial` with nestings of
  ForwardDiff inside and around the broadcast.
- The broadcast tests finish their tapes before reading or replaying them (#311).
- Name the replayed function `replayf` and clarify two comments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

double free crash with multi-threaded code only when using multiple threads

1 participant