Skip to main content

Streaming

Streaming returns tokens as the model generates them, reducing time-to-first-token. The stream format is identical to OpenAI and Anthropic.

OpenAI format

Set stream: true in your Chat Completions request.

Anthropic format

Streaming works the same way in the Messages API. Set stream: true.

Limitations

Tool call streaming follows the same partial-delta format as OpenAI. Function arguments arrive as incremental JSON deltas in the tool_calls array. The final chunk contains the complete tool call with finish_reason: "tool_calls".