BEATRICE Docs
Sections

Chat completions

POST /v1/chat/completions follows the OpenAI format. Below are the Beatrice specifics; the full schema is in the reference.

Parameters

ParameterNotes
model, messagesRequired
max_tokens or max_completion_tokensOutput, reasoning included. Default 8,192, maximum 32,768
stream, stream_options.include_usageSSE streaming; with include_usage a last chunk carries the tokens
temperature, top_p, seed, stopAs in OpenAI
presence_penalty, frequency_penaltyAs in OpenAI
tools, tool_choice, parallel_tool_callsFunction calling, more than one at a time too
response_formatFor example JSON
logprobs, top_logprobs, reasoning_effortForwarded to the model
nOnly 1

Other fields are ignored without an error, so clients that send their own options keep working.

The response

  • choices[0].message.content: the answer.
  • choices[0].message.reasoning_content: the model’s reasoning, when present.
  • choices[0].finish_reason: stop, length (out of output room), tool_calls, or content_filter when the request touches a red line.
  • usage: prompt_tokens, completion_tokens (reasoning included) and prompt_tokens_details.cached_tokens (input read from the cache).

Streaming

With "stream": true the answer arrives as server-sent events, in the OpenAI format, ending with data: [DONE]. Reasoning comes in delta.reasoning_content, the answer in delta.content.

stream = client.chat.completions.create(
    model="beatrice-flash",
    messages=[{"role": "user", "content": "Explain recursion."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
    elif chunk.usage:
        print("\n", chunk.usage)

If the client closes the connection, generation stops at once on our side too.

Queues and timing

Everyone’s requests share the same capacity: when it is full, your request waits in a queue for a few seconds before starting. If the queue is full or the wait exceeds the limit we answer 503 server_overloaded with Retry-After. The OpenAI SDKs retry by themselves. Every response carries an x-request-id header: include it when you contact support.

l'amor che move il sole e l'altre stelle