Chat completions
POST /v1/chat/completions follows the OpenAI format. Below are the Beatrice specifics; the full schema is in the reference.
Parameters
| Parameter | Notes |
|---|---|
model, messages | Required |
max_tokens or max_completion_tokens | Output, reasoning included. Default 8,192, maximum 32,768 |
stream, stream_options.include_usage | SSE streaming; with include_usage a last chunk carries the tokens |
temperature, top_p, seed, stop | As in OpenAI |
presence_penalty, frequency_penalty | As in OpenAI |
tools, tool_choice, parallel_tool_calls | Function calling, more than one at a time too |
response_format | For example JSON |
logprobs, top_logprobs, reasoning_effort | Forwarded to the model |
n | Only 1 |
Other fields are ignored without an error, so clients that send their own options keep working.
The response
choices[0].message.content: the answer.choices[0].message.reasoning_content: the model’s reasoning, when present.choices[0].finish_reason:stop,length(out of output room),tool_calls, orcontent_filterwhen the request touches a red line.usage:prompt_tokens,completion_tokens(reasoning included) andprompt_tokens_details.cached_tokens(input read from the cache).
Streaming
With "stream": true the answer arrives as server-sent events, in the OpenAI format, ending with data: [DONE].
Reasoning comes in delta.reasoning_content, the answer in delta.content.
stream = client.chat.completions.create(
model="beatrice-flash",
messages=[{"role": "user", "content": "Explain recursion."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")
elif chunk.usage:
print("\n", chunk.usage) If the client closes the connection, generation stops at once on our side too.
Queues and timing
Everyone’s requests share the same capacity: when it is full, your request waits in a queue for a few seconds before starting.
If the queue is full or the wait exceeds the limit we answer 503 server_overloaded with Retry-After. The OpenAI SDKs retry by themselves.
Every response carries an x-request-id header: include it when you contact support.