feat: serve a turn through the gateway - #55
Merged
Merged
Conversation
Lawhy
force-pushed
the
feat/gateway-chat
branch
from
October 1, 2026 04:44
aea28d9 to
ac69d07
Compare
`Gateway.chat(messages, tools, params, session_id)` runs one turn in tokin's own vocabulary: the first turn renders in full, later turns render only the messages past what the session served, the prompt is checked against the context length and `max_tokens` capped to what is left, and the reply is parsed from the sampled ids. It returns the parsed message, the generation and the prompt length; `ChatRequest` turns its body into those arguments and `respond` turns the result into OpenAI's `ChatCompletion`, so another API adds a module under `apis/` and leaves the gateway alone. What would make the prompt differ from the model's context is refused before the engine is called: history that does not extend what was served (non-assistant messages compared exactly, the assistant's by role, since SDKs reshape it), an assistant message tokin did not generate, and extending a generation that did not end on a stop token (409). A stop reported off a stop id is an engine error (502). An engine abort is a 503 with nothing recorded, as miles' session server does, so the harness's own retry policy decides. `ChatRequest` declares only what harnesses were seen sending, plus `temperature` and `top_p`; every other OpenAI key is refused by name, and a test lists them against the installed SDK so a new upstream key gets a decision. Tools are kept in the key order they came in, since templates print them with `tojson` and the order reaches the prompt. Checked end to end with an OpenAI Agents SDK harness at 1024 concurrent sessions against Qwen3-30B-A3B and Qwen3.6-35B-A3B: every rollout a token prefix of the conversation's full render bar the template's own reasoning edits, prefix cache hits on all but the increment, and tokin's own cost per turn under 3 ms.
Lawhy
force-pushed
the
feat/gateway-chat
branch
from
October 1, 2026 04:46
ac69d07 to
c21fd83
Compare
Lawhy
added a commit
that referenced
this pull request
Oct 1, 2026
Supersedes #52 and #51. | Package | From | To | Group | |---|---|---|---| | `openai` | 3.14.0 | 3.18.0 | runtime | | `ruff` | 0.16.7 | 0.16.8 | dev | Both are minor/patch. The `openai` 3.14.1–3.18.0 notes add API surface and fix bugs; nothing touches `openai.types.chat` params or `ChatCompletion`, the only parts tokin imports. `ruff` 0.16.8 changes no config keys. The lock was regenerated with uv 0.12.5 (the `uv-lock` pre-commit hook's rev), pinning each package to the version Dependabot chose so the cooldown holds. It is byte-identical to applying both Dependabot diffs onto `main`. ## Verification - `uv sync --locked` - `pre-commit run --all-files` — all hooks pass, incl. `uv-lock`, `mypy`, fast `pytest` - `pytest tests -q` — 103 passed - venv `ruff` 0.16.8: `check` and `format --check` clean Both Dependabot PRs were tested on `cd934c7`, before #55; this is the first run of `openai` 3.18.0 against the gateway path and `test_every_openai_key_is_served_or_listed_unserved`. Co-authored-by: Claude <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Step 5 of the server: one turn, end to end, in the HTTP-free core.
Gateway.chat(messages, tools, params, session_id) -> (message, generation, prompt_tokens)max_tokenscapped to the context left; a prompt that leaves no room isPromptTooLongError(400).stopgeneration — allHistoryConflictError(409).stopthat does not end on a stop id is aGenerationError(502). An engineabortisGenerationAbortedError(503) with nothing recorded, following miles' session server (#3337 there); the harness's retry policy decides.apisimport: the gateway speaksmessages.pyandGenerationParams' names, so Anthropic later addsapis/anthropic.pyonly.apis/openai_chat.pyChatRequestdeclares onlymodel,messages,tools,max_tokens,temperature,top_p; everything else is refused by name. The alignment test now lists the 30 unserved OpenAI keys.toolskeep the key order the harness sent: pydantic reordered them to{"function", "type"}, whichtojsonprinted into every system prompt (found in an end-to-end run).paramsgives the harness's ask underGenerationParams' names;respond(message, generation, prompt_tokens)builds OpenAI's ownChatCompletion, reporting astopwith tool calls astool_calls.Tests:
FakeTokenizerandChatMLmove totests/conftest.py;TestChatcovers each path above; a tool call in unclosed reasoning stays a draft.Integration test with OpenAI Agents SDK harness and 1024 concurrency per model:
The divergences are the templates' own reasoning edits (trimmed whitespace, an unclosed think block), where a text re-render would train on tokens the model never saw. httpx2's default pool (100 connections, 5 s wait) failed 761 of 1024 sessions in the same setup; the CLI has to lift it.
uv run pytest tests -q103 passed; tokenizer suite unchanged; ruff, mypy, pre-commit clean.