Skip to content

feat: read the harness's chat completions request - #54

Merged
Lawhy merged 2 commits into
mainfrom
feat/chat-request
Sep 29, 2026
Merged

Lawhy merged 2 commits into
mainfrom
feat/chat-request

Conversation

@Lawhy

@Lawhy Lawhy commented Sep 29, 2026

Copy link
Copy Markdown
Member

Step 4 of the chat block: the harness-facing request.

  • apis/openai_chat.py: ChatRequest, the canonical request; Anthropic Messages and OpenAI Responses will convert into it later, as miles and polar do.
  • Allowlist. Undeclared keys are refused by name. Keys tokin cannot honour (n, stop, tool_choice, parallel_tool_calls, logprobs, response_format) accept only their no-op values, since SDKs send those unasked. OpenAI's bookkeeping (user, metadata, store, stream_options, service_tier, safety_identifier, prompt_cache_*) is accepted and ignored.
  • Upstream alignment. test_every_openai_key_is_served_or_refused compares the fields with the installed SDK's CompletionCreateParams; a Dependabot bump that adds a key fails until it is served or added to REFUSED.
  • Messages are lenient: unknown keys dropped, content and tool_calls read into lists so a bad item fails as a 400 rather than later.
  • params(max_tokens, stop_ids) passes every field named as a GenerationParams key; the caller's max_tokens wins so the gateway can cap it to the context left.
  • ChatTemplate.conform now joins text parts and refuses non-text content. Which content a template reads is a family fact (Qwen3 renders a parts list as nothing; Qwen3.5 and GLM read it), and it is where a vision family will keep images.
  • pydantic>=2 declared, since it is imported directly.

uv run pytest tests -q 97 passed; tokenizer suite 943 passed, 11 skipped; ruff, mypy, pre-commit clean.

`apis/openai_chat.py` holds `ChatRequest`, the gateway's canonical
request: other APIs will convert into it, as miles and polar do.

It is an allowlist. A key not declared is refused by name, since
dropping it would change what the harness gets back without telling
it. Keys `tokin` cannot honour (`n`, `stop`, `tool_choice`,
`parallel_tool_calls`, `logprobs`, `response_format`) are narrowed to
the values that ask for nothing, because SDKs send those unasked, and
the bookkeeping OpenAI accepts is declared and ignored. A test compares
the fields against the installed SDK's own parameter list, so a bump of
`openai` that adds a key fails until the key is served or refused.

Messages are read leniently: unknown keys dropped, OpenAI's lazily
validated `Iterable` fields read into lists so a bad item is a 400.
`params` hands the engine every field named as a `GenerationParams`
key, with the caller's `max_tokens` winning so it can be capped to the
context left.

Joining text parts moves to `ChatTemplate.conform`, where non-text
content is refused: which content a template can read is the family's
fact (Qwen3 renders a list of parts as nothing, Qwen3.5 reads it), and
it is where a vision family will keep images.
@Lawhy
Lawhy enabled auto-merge (squash) September 29, 2026 23:24
@Lawhy
Lawhy merged commit cd934c7 into main Sep 29, 2026
5 checks passed
@Lawhy
Lawhy deleted the feat/chat-request branch September 29, 2026 23:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant