Defining 1 Prompt
Coding Plan quotas operate on a prompt basis rather than token volume. One prompt corresponds to one user interaction turn in your coding tool, regardless of the underlying model queries initiated by the agent.A single user instruction such as “refactor this database repository and update the test suite” may trigger 15 to 20 background model calls (file reads, tool calls, and test executions). That entire interaction turn is counted as 1 prompt.
Quota Windows
The gateway enforces two separate time windows:- Zero Overage Charges: When quota in either window is depleted, requests hard-stop until the next window reset.
- Deposit Isolation: Coding Plan usage never draws from or impacts your PAYG deposit credit balance.
Plan Tier Limits
In-Flight Concurrency
Concurrent request limits control how many active agent requests can run simultaneously. Extra requests sent while at the concurrency ceiling are temporarily rejected until running queries finish.Load-Based Multipliers
Each prompt consumes quota scaled by the model multiplier coefficient. Flagship models dynamically adjust their multiplier when upstream capacity experiences high load:“Under load” status is evaluated in real-time based on live upstream provider error rates rather than rigid hourly schedules.
Best Practices for Quota Efficiency
- Use Lightweight Models for Routine Work: Route file searches, lint fixes, and boilerplate generation through fast models like
gemini-3-flash(1× multiplier). - Reserve Flagship Models for Complex Tasks: Reserve flagship models for multi-file refactoring, systemic debugging, and architecture design.
- Batch Small Edits Together: Group related requests into a single prompt turn to minimize prompt quota burn.