Docs/Core concepts
Usage, cost & limits
Track tokens and spend per session, watch prompt-cache warmth, and cap an agent run by turns, dollars or wall-clock time.
Source: crates/gray/src/turn_caps.rs
/usage (alias /cost) shows the session's tokens and cost, priced from the LiteLLM rate table. The footer keeps a running total, and the session stores it so a resumed session continues the count.
Caps
Three optional bounds, checked once per prompt turn before any model work. Flags beat env; neither is persisted.
| Flag | Env | Stops after |
|---|---|---|
| --max-turns N | GRAY_MAX_TURNS | N agent turns this process runs |
| --max-cost-usd X | GRAY_MAX_COST_USD | X dollars of priced spend |
| --max-wall-secs S | GRAY_MAX_WALL_SECS | S seconds since the process started |
Prompt cache
Providers keep a prompt cache warm for roughly five idle minutes. The footer's ◷ countdown shows how long is left, and Gray warns in the transcript when a request re-billed tokens that should have been cached — an idle gap, a model switch, or an eviction on the provider's side.
Join the Discord
Chat with the people building and running Gray. Show what you made with it.
Join Discord →