Make tape recording thread-safe (fixes #218) - #311
Merged
Merged
Conversation
`InstructionTape` was a plain `Vector`, so recording from several threads
(e.g. `Threads.@spawn` inside the differentiated function) raced on `push!`.
On Julia 1.10 this corrupts the heap ("double free"), on 1.11+ it throws a
`ConcurrencyViolationError` or silently yields wrong gradients.
`InstructionTape` is now an `AbstractVector` wrapping the instructions and a
`SpinLock` that guards appending and emptying. Reads stay unlocked.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tape is an append-only log written by `record!`. Implement only the iteration interface plus `getindex`, `length` and `isempty`, so generic array functions (`copy`, `similar`, broadcasting, …) don't silently produce unlocked `Vector`s, and the custom `show` is also used for REPL display. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #311 +/- ##
==========================================
+ Coverage 89.18% 89.38% +0.19%
==========================================
Files 19 19
Lines 1933 1978 +45
==========================================
+ Hits 1724 1768 +44
- Misses 209 210 +1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Recording takes one atomic swap per instruction on a backward-linked list instead of a lock, so tasks recording concurrently don't wait for each other. `finish!` turns the list into the instruction vector; reading a tape that is still recording throws, and so does recording onto a finished tape, until `empty!` starts a new recording. The tape constructors finish their tapes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…compactly `record!` swapped out `Finished()` before throwing, so a rejected record put the tape back into the recording state. Restore it before throwing. The two-argument `show` of tapes and instructions now prints one line, and the verbose `text/plain` form of a tape is truncated like an array. Showing a tape that is still recording no longer throws. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolve the conflicts with the Runic reformat (#312) by keeping this branch's changes, formatting them with Runic 1.11.1, and ending functions with an explicit `return nothing` like master does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`empty!(NULL_TAPE)` (e.g. via `empty!(input.tape)` in `jacobian`) put the shared null tape back into the recording state, after which reading it threw. It now leaves `NULL_TAPE` unchanged. Drop `iterate` and `eltype` for `InstructionTape`: the passes, `compile` and `show` already read `instructions(tp)`, which checks the phase once. Compare the threaded gradients with `≈` instead of an absolute tolerance. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
devmotion
added a commit
that referenced
this pull request
Oct 2, 2026
Finishes the tape before counting its instructions in the elementwise tests, as recording and reading are now separate phases (#311). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
devmotion
added a commit
that referenced
this pull request
Oct 2, 2026
…d, and finish tapes before reading them - `copyto!` into a 0-dimensional `TrackedArray` was ambiguous with `Base`'s 0-dimensional method instead of throwing the destination error. - Replace the tests of `trackresults` and `getpartial` with nestings of ForwardDiff inside and around the broadcast. - The broadcast tests finish their tapes before reading or replaying them (#311). - Name the replayed function `replayf` and clarify two comments. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #218.
Problem
InstructionTapewasVector{AbstractInstruction}, andrecord!calledpush!on it without synchronization. When the differentiated function spawns tasks (Threads.@spawn,@threads, …) that do tracked arithmetic, several threads push onto the same tape concurrently:reallocof the buffer, i.e.double free or corruptionas reported in double free crash with multi-threaded code only when using multiple threads #218, orUndefRefErrorduring the reverse pass.ConcurrencyViolationError("Vector can not be resized concurrently"), or silently wrong gradients.Fix
Lock-free recording.
record!wraps the instruction in a node and links it into a backward-linked list with a single@atomicswapon the tape'slastfield. The swap returns the predecessor, so there is no lock and no second atomic, and recording tasks never wait for each other.Recording and replay are separate phases.
finish!ends recording: it swapslastto aFinished()sentinel and moves the list into aVectoronce. After that:length, the passes, compiling) throws;record!seesFinished()as its predecessor at no extra cost, and restores it before throwing);empty!clears the tape and starts recording again, except forNULL_TAPE(the tape of untracked values), which never records and stays finished.The tape constructors (
GradientTape,JacobianTape,HessianTape) callfinish!, so the public API is unchanged. Code that records onto a bareInstructionTape(internal API) now has to callReverseDiff.finish!(tape)before replaying it.InstructionTapeis no longer anAbstractVector; it implements onlylength, and the recorded instructions are read withReverseDiff.instructions(tape), which checks the phase once.Some packages use this internal API directly: Lux.jl (ReverseDiff training extension), AbstractDifferentiation.jl (
derivative) and Gen.jl (backprop) record onto a bareInstructionTape()and callreverse_pass!or iterate it, so they will need afinish!call. Packages that go throughGradientTapeand friends (e.g. SciMLSensitivity.jl) are unaffected.Printing. The two-argument
showof tapes and instructions prints a single line (7-element InstructionTape,InstructionTape (recording),ScalarInstruction(+)). Thetext/plainform of a tape lists its instructions and is truncated like an array, instead of printing every instruction of a long tape; that of an instruction shows its input, output and cache.Why the order is still valid. The swaps put all recordings in one total order that respects each task's program order and every synchronization between tasks. Every rule records its instruction before its output becomes visible. So for each dependency, the instruction producing a value comes before the one consuming it: by program order within a task, and through the synchronization needed to hand the value over across tasks. That holds per dependency, so it doesn't matter whether a task produces, consumes or both. The same argument orders in-place writes against reads in race-free programs. The order isn't deterministic across recordings, so gradients may differ in the last bits.
Performance
Apple Silicon, minimum times unless noted. "64 tasks" is a function that spawns 64 tasks, each doing 200 tracked scalar ops on one tape (median of 100
gradientcalls).x*ygradientx*ygradientThe extra 32 bytes per op are the list node. They cost about 10% replay time for very large tapes on 1.10 (allocation layout) and nothing measurable on 1.13.
Alternatives I benchmarked and rejected:
Threads.SpinLockaroundpush!: 72 ns per op on 1.10, and 28 / 32 ms for 64 tasks at 8 threads on 1.10 / 1.13, i.e. slower than serial.ReentrantLock: about 2× slower thanSpinLockunder contention on 1.13.Base.Channel, with or without a consumer task: slower than a lock, sinceput!takes a lock itself, and a consumer task competes with the recorders for that lock.Vector): 1.3–2× slower replay due to pointer chasing.setfield!(the instruction type isn't fully inferred at the call site).record!::acquire_releaseperforms the same as the default, and:monotonicis 2.5× slower for 64 tasks at 8 threads on 1.10 (4.1–4.4 vs 1.6 ms, reproducible) and the same on 1.13. The contended swap costs the cache line transfer, not the fence.finish!counting the nodes, thenresize!and filling from the back (instead ofpush!+reverse!): about 2× slower on a 270k-instruction tape, since it walks the list twice.A shared tape still writes one cache line per op from all threads, so recording doesn't reach the speed of fully private per-task tapes (2.8 / 5.2 ms at 8 threads). I'll open a follow-up issue for that.
Tests
test/api/GradientTests.jl: compares the gradient of a function that spawns 64 tasks with a serial reference (repeated, plus a compiled tape). CI now runs withJULIA_NUM_THREADS=4so this is actually exercised.test/TapeTests.jlpins the phase rules (reading while recording throws, recording afterfinish!throws,empty!restarts,NULL_TAPEstays finished).take_recorded!intest/utils.jl).🤖 Generated with Claude Code