A curated list of OpenAI-compatible APIs, gateways, tools, SDK patterns, test harnesses, and production notes.
OpenAI-compatible APIs are becoming the common interface for model access. The hard part is not the interface. The hard part is choosing, testing, routing, observing, and controlling real production traffic.
- Gateways and API Layers
- Testing and Benchmarking
- Cost and Budget Control
- SDK Examples
- Local and Self-Hosted Servers
- Agent Workflows
- Production Checklist
- Contributing
- QuotaCheap - OpenAI-compatible API gateway with server-side upstream credentials, user API keys, quotas, logs, usage tracking, balances, and billing visibility.
- LiteLLM - Proxy and SDK for calling many LLM providers with an OpenAI-compatible interface.
- OpenRouter - Unified API for accessing multiple models through an OpenAI-compatible style interface.
- Portkey - AI gateway focused on routing, observability, and reliability.
- llm-gateway-benchmark - Benchmark OpenAI-compatible gateways for latency, success rate, and repeatable scenarios.
- openai-compatible-healthcheck - Proposed healthcheck pattern for endpoint compatibility, auth, chat completions, streaming, and error formats.
- ai-agent-budget-guard - CLI guardrail that stops runaway AI agents before they burn a configured budget.
- Track prompt tokens, completion tokens, cache reads, request count, latency, and retries separately.
- Use per-user and per-task budgets for agentic workflows, not just one global monthly cap.
import OpenAI from 'openai';
const client = new OpenAI({
apiKey: process.env.QUOTACHEAP_API_KEY,
baseURL: 'https://api.quota.cheap/v1',
});
const response = await client.chat.completions.create({
model: 'gpt-5.4-mini',
messages: [{ role: 'user', content: 'Hello from an OpenAI-compatible API.' }],
});
console.log(response.choices[0].message.content);from openai import OpenAI
client = OpenAI(
api_key=os.environ['QUOTACHEAP_API_KEY'],
base_url='https://api.quota.cheap/v1',
)
response = client.chat.completions.create(
model='gpt-5.4-mini',
messages=[{'role': 'user', 'content': 'Hello from an OpenAI-compatible API.'}],
)
print(response.choices[0].message.content)- Ollama - Local model runner with OpenAI-compatible endpoints for some workflows.
- LocalAI - Self-hosted OpenAI-compatible API for local models.
- vLLM - High-throughput inference server with OpenAI-compatible server support.
OpenAI-compatible APIs are especially useful for agents because SDKs and tools can switch endpoints without rewriting the whole integration. Production agent workflows still need:
- budget guards
- request logs
- token usage tracking
- concurrency limits
- retry policies
- model routing rules
- human-review gates for risky actions
Before sending real traffic through any OpenAI-compatible endpoint:
- Confirm auth errors are clear.
- Confirm chat completions work for your SDK.
- Confirm streaming behavior if you use streaming.
- Measure p50 and p95 latency with your prompts.
- Check token usage fields in responses.
- Define request, user, project, and daily budget limits.
- Keep API keys out of client-side code.
- Log request IDs and failures.
- Avoid claiming exact model limits unless verified from official docs or API responses.
Pull requests are welcome for tools, docs, examples, and production lessons.
Rules:
- no affiliate spam
- no unverifiable pricing claims
- no fake benchmarks
- no secrets or private payloads
- explain why a resource is useful
CC0-1.0