Limits and plans
Every organization has a plan. Your plan, its limits and up-to-date usage are in the console.
What is counted
Two separate numbers:
- Uncached input tokens. When you continue a conversation the model has already seen, the repeated part is read from the cache and doesn’t count: it costs almost nothing.
- Output tokens, reasoning included.
The windows
| Window | How it works |
|---|---|
| Last 5 hours | Rolling: usage leaves the count 5 hours after the request |
| Month | Calendar month, Italian time: it resets at midnight on the 1st |
| Concurrency | Requests running or queued at the same time for your organization |
The check happens before every request. A request that starts under the limit runs to the end, even if it goes over.
When you reach a limit
- 5-hour window or too many requests at once:
429 rate_limit_exceeded, withRetry-After. - Month:
429 insufficient_quota, withRetry-Afteruntil the 1st of the month. - For quotas we add
x-should-retry: false, so the SDKs don’t insist for nothing.
Agentic tools like Cline or OpenCode sometimes send several requests at once. If your plan allows one at a time, set them to work sequentially.
Plans
During the beta we assign plans. Public prices and plans will come when the beta ends. If you need more capacity, write to us.