BEATRICE Docs
Sections

Limits and plans

Every organization has a plan. Your plan, its limits and up-to-date usage are in the console.

What is counted

Two separate numbers:

  • Uncached input tokens. When you continue a conversation the model has already seen, the repeated part is read from the cache and doesn’t count: it costs almost nothing.
  • Output tokens, reasoning included.

The windows

WindowHow it works
Last 5 hoursRolling: usage leaves the count 5 hours after the request
MonthCalendar month, Italian time: it resets at midnight on the 1st
ConcurrencyRequests running or queued at the same time for your organization

The check happens before every request. A request that starts under the limit runs to the end, even if it goes over.

When you reach a limit

  • 5-hour window or too many requests at once: 429 rate_limit_exceeded, with Retry-After.
  • Month: 429 insufficient_quota, with Retry-After until the 1st of the month.
  • For quotas we add x-should-retry: false, so the SDKs don’t insist for nothing.

Agentic tools like Cline or OpenCode sometimes send several requests at once. If your plan allows one at a time, set them to work sequentially.

Plans

During the beta we assign plans. Public prices and plans will come when the beta ends. If you need more capacity, write to us.

l'amor che move il sole e l'altre stelle