Skip to content

feat: serve a turn through the gateway - #55

Merged
Lawhy merged 1 commit into
mainfrom
feat/gateway-chat
Oct 1, 2026
Merged

Lawhy merged 1 commit into
mainfrom
feat/gateway-chat

Conversation

@Lawhy

@Lawhy Lawhy commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Step 5 of the server: one turn, end to end, in the HTTP-free core.

Gateway.chat(messages, tools, params, session_id) -> (message, generation, prompt_tokens)

  • Renders the first turn in full, later turns as the increment past what the session served; the engine gets the rollout's ids plus the increment.
  • max_tokens capped to the context left; a prompt that leaves no room is PromptTooLongError (400).
  • Refuses, before the engine is called, anything that would make the prompt differ from the model's context: history that does not extend what was served (non-assistant messages exactly, the assistant's by role), an assistant message tokin did not generate, extending a non-stop generation — all HistoryConflictError (409).
  • A stop that does not end on a stop id is a GenerationError (502). An engine abort is GenerationAbortedError (503) with nothing recorded, following miles' session server (#3337 there); the harness's retry policy decides.
  • No apis import: the gateway speaks messages.py and GenerationParams' names, so Anthropic later adds apis/anthropic.py only.

apis/openai_chat.py

  • ChatRequest declares only model, messages, tools, max_tokens, temperature, top_p; everything else is refused by name. The alignment test now lists the 30 unserved OpenAI keys.
  • tools keep the key order the harness sent: pydantic reordered them to {"function", "type"}, which tojson printed into every system prompt (found in an end-to-end run).
  • params gives the harness's ask under GenerationParams' names; respond(message, generation, prompt_tokens) builds OpenAI's own ChatCompletion, reporting a stop with tool calls as tool_calls.

Tests: FakeTokenizer and ChatML move to tests/conftest.py; TestChat covers each path above; a tool call in unclosed reasoning stays a draft.

Integration test with OpenAI Agents SDK harness and 1024 concurrency per model:

model ok rollout = full render logprobs aligned uncached prompt tokin p99 / turn
Qwen3-30B-A3B 1024 1014 1024 3.4% 1.6 ms
Qwen3.6-35B-A3B 1023 1012 1024 16.1% 2.8 ms

The divergences are the templates' own reasoning edits (trimmed whitespace, an unclosed think block), where a text re-render would train on tokens the model never saw. httpx2's default pool (100 connections, 5 s wait) failed 761 of 1024 sessions in the same setup; the CLI has to lift it.

uv run pytest tests -q 103 passed; tokenizer suite unchanged; ruff, mypy, pre-commit clean.

@Lawhy
Lawhy force-pushed the feat/gateway-chat branch from aea28d9 to ac69d07 Compare October 1, 2026 04:44
`Gateway.chat(messages, tools, params, session_id)` runs one turn in
tokin's own vocabulary: the first turn renders in full, later turns
render only the messages past what the session served, the prompt is
checked against the context length and `max_tokens` capped to what is
left, and the reply is parsed from the sampled ids. It returns the
parsed message, the generation and the prompt length; `ChatRequest`
turns its body into those arguments and `respond` turns the result
into OpenAI's `ChatCompletion`, so another API adds a module under
`apis/` and leaves the gateway alone.

What would make the prompt differ from the model's context is refused
before the engine is called: history that does not extend what was
served (non-assistant messages compared exactly, the assistant's by
role, since SDKs reshape it), an assistant message tokin did not
generate, and extending a generation that did not end on a stop token
(409). A stop reported off a stop id is an engine error (502). An
engine abort is a 503 with nothing recorded, as miles' session server
does, so the harness's own retry policy decides.

`ChatRequest` declares only what harnesses were seen sending, plus
`temperature` and `top_p`; every other OpenAI key is refused by name,
and a test lists them against the installed SDK so a new upstream key
gets a decision. Tools are kept in the key order they came in, since
templates print them with `tojson` and the order reaches the prompt.

Checked end to end with an OpenAI Agents SDK harness at 1024 concurrent
sessions against Qwen3-30B-A3B and Qwen3.6-35B-A3B: every rollout a
token prefix of the conversation's full render bar the template's own
reasoning edits, prefix cache hits on all but the increment, and
tokin's own cost per turn under 3 ms.
@Lawhy
Lawhy force-pushed the feat/gateway-chat branch from ac69d07 to c21fd83 Compare October 1, 2026 04:46
@Lawhy
Lawhy merged commit d526cb0 into main Oct 1, 2026
5 checks passed
@Lawhy
Lawhy deleted the feat/gateway-chat branch October 1, 2026 04:48
@Lawhy Lawhy mentioned this pull request Oct 1, 2026
Lawhy added a commit that referenced this pull request Oct 1, 2026
Supersedes #52 and #51.

| Package | From | To | Group |
|---|---|---|---|
| `openai` | 3.14.0 | 3.18.0 | runtime |
| `ruff` | 0.16.7 | 0.16.8 | dev |

Both are minor/patch. The `openai` 3.14.1–3.18.0 notes add API surface
and fix bugs; nothing touches `openai.types.chat` params or
`ChatCompletion`, the only parts tokin imports. `ruff` 0.16.8 changes no
config keys.

The lock was regenerated with uv 0.12.5 (the `uv-lock` pre-commit hook's
rev), pinning each package to the version Dependabot chose so the
cooldown holds. It is byte-identical to applying both Dependabot diffs
onto `main`.

## Verification

- `uv sync --locked`
- `pre-commit run --all-files` — all hooks pass, incl. `uv-lock`,
`mypy`, fast `pytest`
- `pytest tests -q` — 103 passed
- venv `ruff` 0.16.8: `check` and `format --check` clean

Both Dependabot PRs were tested on `cd934c7`, before #55; this is the
first run of `openai` 3.18.0 against the gateway path and
`test_every_openai_key_is_served_or_listed_unserved`.

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant