Skip to main content

POST /v1/chat/completions

OpenAI-compatible chat completions. Same request shape, same SSE streaming contract.

Request

prompt_tokens on the request is ignored — the server re-estimates.

Non-streamed response

  • content is null when the turn is tool-calls only (OpenAI behavior).
  • tool_calls appears with finish_reason: tool_calls.
  • reasoning_content appears on reasoning models (also split out of <think>…</think> content) — see Reasoning.
  • usage.prompt_tokens_details.cached_tokens appears when prompt caching applies.

Example

Errors

400 missing model · 401 bad key · 402 out of credits · 404 unknown model · 429 model busy / warming (honor retry-after) · 502/503 upstream or gateway trouble. Full table: Errors.