Repository navigation
Conversation
Every `up --wait` is bounded only by the external `timeout` command, so the way a slow stack fails is a SIGKILL delivered wherever Compose happens to be. Killed between the rename and the removal of a recreate, that strands a `<shortid>_<service>` container still holding the service's fixed `container_name` — which is why "Cleanup failed existing stacks on failure" exists at all. Add `--wait-timeout` at 90% of `service-startup-timeout` to all five `up --wait` sites (three deploy, two rollback). Compose now hits its own deadline first and exits on its own terms, naming the services that never became healthy instead of dying silently, and the external `timeout` becomes a backstop for a genuinely wedged Compose. This is a mitigation, not a cure, and the cleanup step stays: `--wait-timeout` bounds the final `--wait` phase, not the `depends_on: service_healthy` waits Compose performs mid-sequence, so a stack that stalls on a dependency can still burn the whole budget and take the SIGKILL with later services half-recreated. 90% rather than a fixed subtrahend so the margin scales with whatever callers configure (300s -> 270s) without a second input to keep in sync. Computed in shell because GitHub expressions have no arithmetic operators. Verified: yamllint --strict clean; actionlint clean apart from the pre-existing job.workflow_sha finding; `up --dry-run -d --quiet-pull --wait --wait-timeout 270 --remove-orphans` confirmed rc=0 against Compose v5.5.1 on the piwine runner, whose plan shows the very `<shortid>_<service>` rename this guards.
Reviewer's GuideThe workflow now gives Compose its own wait deadline at 90% of service-startup-timeout across all deployment and rollback up --wait calls, allowing Compose to report unhealthy services and exit cleanly before the external timeout becomes a SIGKILL backstop; existing cleanup remains because dependency waits can still exceed the Compose wait phase. Sequence diagram for Compose wait deadline and timeout backstopsequenceDiagram
participant Workflow
participant Compose
participant Cleanup
Workflow->>Workflow: Compute Compose wait budget
Workflow->>Compose: docker compose up --wait --wait-timeout COMPOSE_WAIT_TIMEOUT
alt services become healthy
Compose-->>Workflow: Success
else Compose wait deadline reached
Compose-->>Workflow: Failure naming unhealthy services
Workflow->>Cleanup: Run existing failed-stack cleanup
else Compose remains wedged
Workflow--xCompose: External timeout sends SIGKILL
Workflow->>Cleanup: Run existing failed-stack cleanup
end
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
Hey - I've reviewed your changes and they look great!
Sourcery assessment
Needs a human reviewer. This changes production deployment timing and can cause Compose to stop during a recreate before all services are healthy, leaving a partially updated stack or triggering rollback/cleanup paths. Reverting prevents future occurrences, but any interrupted deployment or resulting outage must be repaired separately.
|
Note on CI: Ran the same three checks locally against this branch in the meantime:
|
The problem
Every
up --waitin this workflow is bounded only by the externaltimeoutcommand. That means the way a slow stack fails is a SIGKILL delivered wherever Compose happens to be — including between the rename and the removal of a recreate, which strands a<shortid>_<service>container still holding the service's fixedcontainer_name.That isn't hypothetical; the workflow already documents it, in the comment on the step that exists to mop it up:
So today the cleanup step is load-bearing for damage caused by how the timeout fires.
The change
--wait-timeoutat 90% ofservice-startup-timeouton all fiveup --waitsites — three indeploy, two inrollback. Compose hits its own deadline first and exits on its own terms, naming the services that never became healthy rather than dying silently mid-operation. The externaltimeoutis demoted to a backstop for a genuinely wedged Compose.90% rather than a fixed subtrahend so the margin scales with whatever callers configure (300s → 270s) without a second input to keep in sync. Computed in shell because GitHub expressions have no arithmetic operators.
What this does not fix
The cleanup step stays, and should.
--wait-timeoutbounds the final--waitphase, not thedepends_on: service_healthywaits Compose performs mid-sequence. A stack that stalls on a dependency can still burn the whole budget and take the SIGKILL with later services half-recreated. This makes the stranding rare; it does not make it impossible. The code comment says so explicitly so nobody deletes the cleanup step on the strength of this PR.Verification
yamllint --strictcleanactionlintclean apart from the pre-existingjob.workflow_shafinding (present onmaintoo)docker compose up --dry-run -d --quiet-pull --wait --wait-timeout 270 --remove-orphansconfirmedrc=0against Compose v5.5.1 on the piwine runner — non-mutating, and its plan happens to show the exacta774a6ece886_network-optimizerrename this guards againstSummary by Sourcery
Give every deployment and rollback Compose wait a scaled internal deadline so failures occur cleanly before the external timeout intervenes.
Bug Fixes:
Enhancements:
Documentation: