Repository navigation
fix(deploy): give Compose its own wait deadline before the SIGKILL - #95
Conversation
Every `up --wait` is bounded only by the external `timeout` command, so the way a slow stack fails is a SIGKILL delivered wherever Compose happens to be. Killed between the rename and the removal of a recreate, that strands a `<shortid>_<service>` container still holding the service's fixed `container_name` — which is why "Cleanup failed existing stacks on failure" exists at all. Add `--wait-timeout` at 90% of `service-startup-timeout` to all five `up --wait` sites (three deploy, two rollback). Compose now hits its own deadline first and exits on its own terms, naming the services that never became healthy instead of dying silently, and the external `timeout` becomes a backstop for a genuinely wedged Compose. This is a mitigation, not a cure, and the cleanup step stays: `--wait-timeout` bounds the final `--wait` phase, not the `depends_on: service_healthy` waits Compose performs mid-sequence, so a stack that stalls on a dependency can still burn the whole budget and take the SIGKILL with later services half-recreated. 90% rather than a fixed subtrahend so the margin scales with whatever callers configure (300s -> 270s) without a second input to keep in sync. Computed in shell because GitHub expressions have no arithmetic operators. Verified: yamllint --strict clean; actionlint clean apart from the pre-existing job.workflow_sha finding; `up --dry-run -d --quiet-pull --wait --wait-timeout 270 --remove-orphans` confirmed rc=0 against Compose v5.5.1 on the piwine runner, whose plan shows the very `<shortid>_<service>` rename this guards.
Reviewer's GuideThe workflow now computes a 90%-of-budget Compose wait timeout and applies it to every deploy and rollback Sequence diagram for Compose wait timeout handlingsequenceDiagram
participant Workflow
participant Compose
participant Cleanup
Workflow->>Workflow: Compute Compose wait budget
Workflow->>Compose: docker compose up --wait --wait-timeout COMPOSE_WAIT_TIMEOUT
alt Compose reaches wait timeout
Compose-->>Workflow: Exit with unhealthy services
Workflow->>Cleanup: Run failure cleanup
else Compose is genuinely wedged
Workflow->>Compose: External timeout sends SIGKILL
Workflow->>Cleanup: Remove stranded recreate containers
else Compose completes in time
Compose-->>Workflow: Exit successfully
end
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
Hey - I've reviewed your changes and they look great!
Sourcery assessment
Needs a human reviewer. This changes production deployment timing and can cause Compose to stop during a recreate, leaving a partially updated stack or causing an outage before the external timeout and cleanup paths run. Reverting prevents future occurrences, but it cannot undo any partial deployment or service disruption that already happened.
Replaces #94, which GitHub auto-closed when its base branch (
fix/retry-transient-registry-throttles, now merged as #93) was deleted. Same commit, rebased ontomain; a closed PR's base cannot be retargeted, so it needed a new number. The review comments on #94 still apply.The problem
Every
up --waitin this workflow is bounded only by the externaltimeoutcommand. That means the way a slow stack fails is a SIGKILL delivered wherever Compose happens to be — including between the rename and the removal of a recreate, which strands a<shortid>_<service>container still holding the service's fixedcontainer_name.That isn't hypothetical; the workflow already documents it, in the comment on the step that exists to mop it up:
So today the cleanup step is load-bearing for damage caused by how the timeout fires.
The change
--wait-timeoutat 90% ofservice-startup-timeouton all fiveup --waitsites — three indeploy, two inrollback. Compose hits its own deadline first and exits on its own terms, naming the services that never became healthy rather than dying silently mid-operation. The externaltimeoutis demoted to a backstop for a genuinely wedged Compose.90% rather than a fixed subtrahend so the margin scales with whatever callers configure (300s → 270s) without a second input to keep in sync. Computed in shell because GitHub expressions have no arithmetic operators.
What this does not fix
The cleanup step stays, and should.
--wait-timeoutbounds the final--waitphase, not thedepends_on: service_healthywaits Compose performs mid-sequence. A stack that stalls on a dependency can still burn the whole budget and take the SIGKILL with later services half-recreated. This makes the stranding rare; it does not make it impossible. The code comment says so explicitly so nobody deletes the cleanup step on the strength of this PR.Verification
yamllint --strictcleanactionlintclean apart from the pre-existingjob.workflow_shafinding (present onmaintoo)shellcheck -xclean over all 16 scripts underscripts/docker compose up --dry-run -d --quiet-pull --wait --wait-timeout 270 --remove-orphansconfirmedrc=0against Compose v5.5.1 on the piwine runner — non-mutating, and its plan happens to show the exacta774a6ece886_network-optimizerrename this guards againstSummary by Sourcery
Give every deployment and rollback Compose update its own scaled wait deadline before the external timeout terminates it.
Bug Fixes:
Enhancements:
Documentation: