Skip to content

Release v1.3.0 — lifecycle correctness - #7

Merged
CBEPX merged 36 commits into
mainfrom
release/v1.3.0
Sep 27, 2026
Merged

CBEPX merged 36 commits into
mainfrom
release/v1.3.0

Conversation

@CBEPX

@CBEPX CBEPX commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Summary

Release v1.3.0 of the CBEPX fork: lifecycle correctness on posix, stop-gate controls, data-driven model aliases, and two security fixes. Full notes in CHANGELOG.md (section 1.3.0).

Highlights:

Known limitations (documented in CHANGELOG/README): Windows kills from stored records are refused until process identity lands in v1.4.0.

Validation

  • npm test: 313 tests, 0 failures, 0 leaked processes; npm run build, check-version, claude plugin validate . --strict, npm audit --omit=dev, npm pack --dry-run all green.
  • Per-task Claude reviews, whole-branch review, five Codex adversarial passes (final verdict: ship), five fix waves each re-reviewed.
  • Live smoke against real Codex: setup, awaited task, status, SessionEnd teardown with identity match.

Ported with reference to upstream PRs by ALV0612, Soumya95, kevin9327, mzl9039, sylvesterkaczmarek, mittalpk, SomSamantray, weivwang.

🤖 Generated with Claude Code

CBEPX and others added 30 commits September 27, 2026 18:42
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rn failure, fileChange guard

- An `error` notification without `willRetry: true` now ends the turn as
  failed instead of waiting for a turn/completed that never comes (openai#698).
- A task whose turn fails without an error payload records an errorMessage
  and a summary naming the turn status, not the first line of rawOutput (openai#757).
  The job status report shows the error for failed jobs.
- `fileChange` started items without `changes` no longer crash progress (openai#775).
- Test helper `run()` now forwards `timeout` to spawnSync.

Co-authored-by: ALV0612 <ALV0612@users.noreply.github.com>
Co-authored-by: Soumya95 <Soumya95@users.noreply.github.com>
Co-authored-by: kevin9327 <kevin9327@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…or only when it adds information

Co-authored-by: ALV0612 <ALV0612@users.noreply.github.com>
Co-authored-by: Soumya95 <Soumya95@users.noreply.github.com>
Co-authored-by: kevin9327 <kevin9327@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: ALV0612 <ALV0612@users.noreply.github.com>
Co-authored-by: Soumya95 <Soumya95@users.noreply.github.com>
Co-authored-by: kevin9327 <kevin9327@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…turn id

A turn/start response without turn.id left state.turnId null, so every
notification stayed buffered and the capture never completed (openai#781).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A broker socket stuck in `connecting` fires neither connect nor error, so
waitForBrokerEndpoint and the broker client's initialize could hang forever
(openai#773). Each probe attempt is now bounded by min(500 ms, remaining budget)
and destroys its socket on expiry; the broker client rejects with ETIMEDOUT
after connectTimeoutMs (default 2000), and withAppServer falls back to a
direct app-server on ETIMEDOUT when a broker was requested.

`status <id> --wait` reported a timed-out wait with exit 0 and no hint
(openai#774). It now prints "Timed out after <N>s while the job was still
running." and exits 1; the --json snapshot is unchanged but exits 1 too.

Co-authored-by: kevin9327 <kevin9327@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…s probe, never signal a stale pid

ensureBrokerSession now defaults killProcess to terminateProcessTree, so a
replaced or never-ready broker is no longer leaked (openai#753/openai#762). A live,
owned broker that misses the 150 ms probe gets a full retry window before
it is treated as wedged (openai#768). A dead or foreign pid from a stale record
is never signalled; only its files are cleared (openai#749). A probe socket that
connected but closes slowly now reads as ready.

Co-authored-by: Soumya95 <Soumya95@users.noreply.github.com>
Co-authored-by: mzl9039 <mzl9039@users.noreply.github.com>
Co-authored-by: sylvesterkaczmarek <sylvesterkaczmarek@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…al and escape hatch; drop hooks.json description

- setup --review-gate-model/--review-gate-effort (inherit clears); the stop
  gate forwards them to its review task (openai#769)
- CODEX_REVIEW_GATE_MAX_ROUNDS defaults to 3; explicit 0 keeps unbounded (openai#548)
- failure reasons name the kill signal and end with the disable hint (openai#589, openai#483)
- hooks.json drops the top-level description key (openai#459)

Co-authored-by: mittalpk <mittalpk@users.noreply.github.com>
Co-authored-by: SomSamantray <SomSamantray@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
/codex:transfer hardcoded ~/.claude/projects as the only accepted
transcript root, so a Claude Code install using CLAUDE_CONFIG_DIR to
relocate its config directory could never pass the source-path check
(openai#721). resolveClaudeProjectsDir(env) now derives the projects dir from
CLAUDE_CONFIG_DIR (resolved via path.resolve when set) and falls back to
~/.claude/projects otherwise; resolveClaudeSessionPath threads options.env
through to it and to the TRANSCRIPT_PATH_ENV lookup, defaulting to
process.env.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…model catalogue

Model aliases now resolve against the local Codex catalogue
($CODEX_COMPANION_MODEL_CATALOG -> $CODEX_HOME/models_cache.json ->
`codex debug models --bundled` -> hardcoded fallback), picking the listed
model whose slug ends in -<alias> (lowest priority, newest family on
ties). Exact slugs pass through; hidden models never match an alias.
--effort is rejected when the catalogued model does not list it. Adds the
astra alias. Tests pin the catalogue to a fixture. (openai#468/openai#703/openai#485/openai#128)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e broker.json before use

Co-authored-by: weivwang <weivwang@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: weivwang <weivwang@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…n pid sidecar, job records and broker.json

A recorded pid is signalled, reaped as dead, or evicted from the state lock
only once its start identity proves it is still the recorded process (openai#743):
linux /proc/<pid>/stat starttime, darwin `ps -o lstart=,comm=` pinned to
LC_ALL=C/TZ=UTC. Unavailable identity never authorises a kill; win32 reports
identity-unavailable (CIM identity is v1.4.0); posix records without an
identity fall back to the command-line check. runCommand reports a timed-out
command as status null instead of 0. Cancel says when it left a worker
running, and the SessionEnd teardown line prints its reason.

Co-authored-by: sylvesterkaczmarek <sylvesterkaczmarek@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…led regardless of identity

The not-ready branch holds the child handle it just spawned, so its pid cannot
have been recycled: kill it directly on every platform instead of routing it
through the stored-record identity check (which cannot answer on win32).

Co-authored-by: sylvesterkaczmarek <sylvesterkaczmarek@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Bump package.json, package-lock.json, plugins/codex/.claude-plugin/plugin.json
and .claude-plugin/marketplace.json to 1.3.0 via scripts/bump-version.mjs.
Add the 1.3.0 CHANGELOG.md section (Fixed/Added/Changed/Known limitations,
upstream #/PR citations) and sync plugins/codex/CHANGELOG.md to match.
Add a README FAQ note on the Windows kill limitation, resolved in v1.4.0.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ies no id (openai#781)

The timeout path could not send turn/interrupt without a turnId. Also correct
the stale comments in failTurnOnTimeout and on the error notification.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…cy pid

A fresh broker that fails readiness is killed through its child handle, and
not at all once it has exited; the numeric process-group kill is only a
fallback behind identity or command-line proof. A legacy broker.json pid is
re-verified at kill time after the readiness retry, and the record is only
cleared while it still names the endpoint we tore down.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A worker cancel may not signal but that is still alive would later overwrite
the cancelled record. The job now stays running with its pid sidecar, the log
and output say cancellation is not confirmed, and cancel exits 1.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The deadline now starts before the self-probe; every probe gets at most
min(500 ms, time left) and is skipped (entry held) under 50 ms, and a failed
self-probe is cached per process, so slow ps calls cannot push a lock wait
past the SessionEnd budget.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…y write

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…n, triage)

The fallback-root limitation now says state is orphaned on every plugin
update when CLAUDE_PLUGIN_DATA is unset; the cancel entry describes the
pending-cancellation outcome; rescue docs stop hardcoding the spark slug;
triage marks the issues and PRs this release closes as fixed-in v1.3.0.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The test relied on cancel recording cancelled after a refused kill, which is
the behaviour 41ca814 removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… record

A group kill raises ESRCH for a live pid that leads no process group; fall
back to the pid itself before reporting not delivered. Cancel treats an
undelivered kill of a live worker as pending, and a worker that outlives an
acknowledged cancel no longer overwrites the cancelled record (checked under
the state lock).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A broker stuck in connect has no cleanup handlers yet, so signalling only the
broker left its app-server child behind. The live, unreaped child is a detached
group leader, so its group is killed; the handle is the fallback. An exited
child is still never signalled.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A running record without an identity whose pid was recycled by a long-lived
unrelated process was never reaped, so a refused cancel stayed pending and the
thread stayed in use. A readable command line without codex-companion.mjs now
marks the job dead; nothing is signalled, and an unreadable one proves nothing.
Test fixtures standing in for live workers now carry a companion command line.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A model-only --review-gate-model change skipped effort validation and could
persist a model with an effort it cannot run. The stored effort is now checked
against the new model before any write; on failure nothing is written.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…iation

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ion pid

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…no group (H1)

terminateProcessTree no longer falls back to kill(pid) after the group
kill raised ESRCH; it reports groupGone instead. terminateRecordedProcess
re-proves identity (or the legacy command line) immediately before the
bare-pid SIGTERM and refuses when the pid was recycled in between.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… (H2)

processCommandLine reads /proc/<pid>/cmdline on linux and runs
`ps -ww -o command=` with COLUMNS=10000 elsewhere; empty or unreadable
output is null (unknown), which the reaper and broker ownership check
already treat as no action. Legacy command-line fixtures that faked ps
now say platform darwin.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
CBEPX and others added 6 commits September 27, 2026 22:58
…stop (H3)

cleanupSessionJobs now reads each kill outcome and removes a running
foreground job's record only when the kill was delivered, the pid is
provably gone, or the job has no pid. A refused, undelivered or failed
kill, or a job the budget never reached, keeps its record and files,
is logged as "[codex] SessionEnd left <id> running: <reason>", and keeps
the broker up through activeWorkspaceJobs.

Each identity probe now gets half of what is left (worker cleanup) or of
the step budget (broker teardown): the H1 re-verify can add a second
probe per kill, and both must fit inside the 12 s SessionEnd budget.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Halved SessionEnd probe budgets could be fractional (e.g. 500.5); spawnSync
threw ERR_OUT_OF_RANGE, the hook logged kill-failed and kept a live worker.
Floor both halving sites, and clamp in runCommand so no caller can pass a
fractional, zero (unbounded) or negative timeout. Same clamp on the stop
review gate's env override.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A dead legacy worker left as a zombie has an empty /proc/<pid>/cmdline while
kill 0 still sees it, so processCommandLine returned null (unknown) and the
reaper kept the job running forever. On an empty cmdline read /proc/<pid>/stat:
state Z or X yields "<defunct>" (as ps prints it, no companion marker), so the
reaper's legacy rule reconciles the job without signalling. Any other state or
an unreadable stat stays null.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ring teardown

The SessionEnd hook's clearBrokerSession(cwd) guarded the unlink with an
existsSync pre-check, but the broker's own SIGTERM handler
(clearOwnSessionRecord) can delete the same broker.json concurrently: the
check and the unlink are not atomic, so the hook could still hit
ENOENT and exit 1. Replace the pre-check with try/unlink/catch-ENOENT in
clearBrokerSession and in every unlinkSync site of teardownBrokerSession
(pidFile, logFile, the unix socket path) so a concurrently self-cleaning
broker can remove any of them without failing the hook. Other error codes
(e.g. EPERM) still surface exactly as before.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
teardownBrokerSession's pidFile/logFile unlinks were only tolerant of ENOENT,
so a locked or otherwise unremovable file (EPERM, notably reported on Windows
upstream in openai#633/openai#626) still escaped and failed the SessionEnd hook. These two
unlinks are best-effort cleanup, same as the socket-path unlink beside them
(already a catch-all) and rmdirSync below them: swallow every error, not just
ENOENT. clearBrokerSession is unchanged and stays ENOENT-only, since it is the
one unlink whose target (broker.json) is the ownership record itself.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… subagent labels survive

The replay of notifications buffered before the turn/start response ran
every message through belongsToTurn, dropping a subagent's thread/started
(its thread id is not yet known) and losing the label. Route live and
replayed notifications through one function. Adds the
FAKE_CODEX_SUBAGENT_EARLY_STARTED fixture knob and a deterministic test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@CBEPX
CBEPX merged commit 78576a5 into main Sep 27, 2026
17 of 27 checks passed
@CBEPX
CBEPX deleted the release/v1.3.0 branch September 27, 2026 21:37
CBEPX added a commit that referenced this pull request Sep 28, 2026
- broker-reuse E2E reads errorMessage from the stored job, not the foreground JSON
- explicit fileURLToPath/pathToFileURL import in tests/process.test.mjs
- export cleanupSessionJobs and guard main() behind direct execution in the hook
- renderCancelPending wired into the cancel caller with diagnostic/json/log split
- SURVIVOR-before-finally assertion checks both markers exist first
- spec: post-snapshot descendants have no lifetime bound or later cleanup
- inline cancel example reads broker.json only on win32

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant