Skip to content

Measure through a benchmark package of its own - #15

Draft
sinoru wants to merge 3 commits into
developfrom
feature/benchmarks
Draft

sinoru wants to merge 3 commits into
developfrom
feature/benchmarks

Conversation

@sinoru

@sinoru sinoru commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

The performance suites move out of the test targets into Benchmarks, a package of its own on package-benchmark, so that nothing a client resolves includes the harness: the library's manifest still has no dependencies. swift test no longer measures anything, and no XCTest case is left in the package.

The benchmark package

Every case comes across, 69 of them, each with its comparison: the standard library's Mutex, DispatchSemaphore, pthread_rwlock_t, a concurrent DispatchQueue, and an actor. The pointer chase, the shared turn budget, and the pinned threads carry over from the old harness; what the XCTest meter and the corelibs clock did is the harness's job now.

  • The library is measured as it ships, with no -enable-testing.
  • A sample is a million turns, so the scaled output reads per turn; --scale prints the raw sample.
  • The default metrics are the wall clock and instructions retired, with CPU time and context switches under contention. Allocation and retain counts are a separate pass (--metric mallocCountTotal --metric retainCount --metric releaseCount): the harness counts them by hooking every allocation, retain, and release, which put about an eighth on the clock of an asynchronous handoff (AsyncMutex on a short queue, 2,029 → 2,324 ns) and instructions of its own on every path that retains.
  • Baselines: baseline update before a change and baseline compare after, in place of running each side fifteen times by hand.

Workers live on a pool of threads kept for the whole run, because the allocation counter sets up a per-thread record on a thread's first free and traps when that free is inside the thread's exit. A contended sample runs in stretches of at most a tenth of a second, every worker parked between them.

The harness runs on macOS and Linux only, so Windows, Android, and WASI are no longer measured. They still run the stress suites in release.

Tests and CI

README

The Performance tables are measured again, as the median of seven runs of the whole suite, each run's figure the median of its samples. A single run was not enough: one had concurrent reads at 5.8 ns on RWLock against 4.7 on Mutex, where five reruns of that pair and the seven-run median (4.8 against 5.0) have RWLock ahead. The prose is checked against the new figures. Two sentences change: the read-mostly async figures, and "readers at once cost no more than a Mutex", where the old harness measured less. Its contended Mutex was half again slower than now, with fresh threads each sample. How to run the benchmarks, keep a baseline, and take the counts gets a section of its own.

The CHANGELOG is untouched: nothing a client receives changes.

Checked

  • macOS 26.7.1, Swift 6.4: the benchmark package and the library build in release without warnings. The debug unit tests (--disable-xctest --skip StressTests) and the release stress suites (--disable-xctest --disable-testable-imports --filter StressTests) pass.
  • The full benchmark suite, seven times over, with no failure and no stall.
  • Not checked locally: the Android row's test step with --disable-xctest, and Run benchmarks on a runner. This PR's CI is where both are seen first.

The performance suites were XCTest cases inside the test targets, timed by
XCTest's meter on Apple platforms and by a clock of the harness's own
elsewhere, read as a mean of ten runs, and compared by running each side
fifteen times over. They move to `Benchmarks`, a package of their own on
the package-benchmark harness, so that nothing a client resolves includes
the harness: the library's manifest still has no dependencies.

Every case comes across, each comparison with it: the standard library's
`Mutex`, `DispatchSemaphore`, `pthread_rwlock_t`, a concurrent
`DispatchQueue`, and an `actor`. What is measured is the library as it
ships, with no `-enable-testing`; a sample is a million turns, so the
scaled output reads per turn; and the harness reports percentiles of the
clock and of instructions retired, keeps baselines, and compares a run
against one. Allocation and retain counts are there for the asking, and
not by default, since the hooks that count them slow the paths they count.

Workers live on a pool of threads kept for the whole run: the allocation
counter sets up a per-thread record on a thread's first free, which for a
thread that frees nothing before it exits is from inside the exit, where
the allocator traps. A contended sample runs in stretches of at most a
tenth of a second, every worker parked between them.
The XCTest performance suites, and the harness they measured through, go:
the benchmark package measures the same cases now. What they leave behind
in the test utilities goes with them, the Atomic dependency the harness's
counter needed and the AddressSanitizer flag it skipped on. No XCTest case
is left in the package.

CI follows. The release run is the stress suites alone, on both
workflows, and the tests build and run without XCTest everywhere, which
the WebAssembly row alone did before: the corelibs runner traps on an
empty list, and Android's runner loop takes Swift Testing only. The
benchmarks run on one macOS and one Linux row per Swift minor, the
platforms the harness supports, and put their table in the job's summary
without failing on a number; their package is in the path filter, and the
job's timeout makes room for them.

Two rows stay marked experimental for failures that showed in the
measurements: the arm64 Windows actor increment (#6) and the arm64 Linux
dispatch crash (#11). Neither row runs the measurements any more, so each
is kept marked until a run without them shows it clear.
Running the tests no longer measures anything, so the README's account of
how to measure moves to a section of its own: how to run the benchmark
package, how to keep a baseline and compare against it, and why the
allocation and retain counts are a pass of their own.

The performance tables are measured again by the benchmark package, as
the median of seven runs of the whole suite, each run's figure the median
of its samples, where the old ones were XCTest's means over seven runs.
The prose beside them is checked against the new figures. Two sentences
move with them: a turn among one task in eight writing costs what it now
measures, and readers at once cost no more than a `Mutex`, where they
measured as costing less under the old harness, whose contended `Mutex`
ran half again slower than it does with its threads kept warm.
@sinoru
sinoru force-pushed the feature/benchmarks branch from f226218 to 2903175 Compare October 10, 2026 18:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant