Skip to content

🎟️ feat: give each job and deploy its own run token - #306

Open
erick-GeGe wants to merge 3 commits into
feat/oidc-subfrom
feat/run-tokens
Open

erick-GeGe wants to merge 3 commits into
feat/oidc-subfrom
feat/run-tokens

Conversation

@erick-GeGe

@erick-GeGe erick-GeGe commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #305. GitHub will retarget it to main once #305 merges.
Independent of single sign-on: it only changes what a job's or a deploy's container carries.

Why

To report back, every container estela launches gets a token in JOB_INFO:

Container Token it carries today What that token can do
Spider, queued (started by hand or by a cron job, the usual path) The DRF token of the project's owner Anything the owner can
Spider started with ?sync=true, which skips the queue The DRF token of whoever started it Anything that person can
Deploy build deploy_manager's DRF token, the same one for every deploy Report on any deploy of any project

None of them ever expire, and all sit in plain text in the pod's environment, readable by anyone who can read the pod. A container needs exactly two endpoints; these tokens open all of them.

What

Each run gets its own token, good only for reporting on itself:

  • One per run. Launching a job or a deploy creates a random estela-run_… token. Only its SHA-256 is stored (RunToken, migration 0045_runtoken).
  • Two endpoints accept it: PATCH …/jobs/{jid} and PUT …/deploys/{did}, the only ones containers call. Everywhere else it is not recognised at all, so it cannot create API keys or edit a profile (RUN_AUTHENTICATION_CLASSES).
  • Only its own run, only to update it. The token of job 3549 cannot touch job 3550 or read anything (IsOwnRun, opted into per view with run_token_target).
  • It dies with the run. Once the job is COMPLETED, ERROR or STOPPED, or the deploy SUCCESS, FAILURE or CANCELED, it stops working. No fixed expiry, because jobs have no maximum duration.
  • It acts as the same user as before: a queued job (by hand or by a cron job) as the project's owner, a ?sync=true job as whoever started it (person or API key owner), a deploy as deploy_manager. Existing permission checks apply unchanged.
  • A hyphen in the prefix, not the API keys' underscore, so nothing that recognises keys by estela_ mistakes one for the other.

What does not change

  • The containers. The entrypoint and the build still send Authorization: Token <value>; only the value is different. No image needs rebuilding.
  • DRF tokens still authenticate, so runs started before this deploy finish normally. Nothing hands out a new one for a run any more; the fallback can go once those have drained.
  • People, the web and the CLI.

How it was checked

End to end on a local cluster, with the same mechanism on the single sign-on working branch: real spider jobs (manual and queued) and two real deploys reported their status with their own token, and every token was dead once its run finished. That branch differs only in also listing the gateway's JWT class on these two views.

Deploy notes

Migration 0045_runtoken. No configuration change.

A RunToken belongs to one job or one deploy. Only the two endpoints a
container reports to accept it, only to update that run, and it stops
working once the run is over. Only its hash is stored.
Manual jobs, cron jobs and deploys used to carry a DRF token in JOB_INFO:
the launcher's, the project owner's, or deploy_manager's shared one. None
of them expire and all can do anything their user can. Each run now gets
a token that dies with it. DRF tokens still authenticate, so runs started
before this deploy finish normally.
Not only scheduled ones: a job started by hand goes through the queue too,
unless it asks for sync=true.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants