Docs/Core concepts

Usage, cost & limits

Track tokens and spend per session, watch prompt-cache warmth, and cap an agent run by turns, dollars or wall-clock time.

Source: crates/gray/src/turn_caps.rs

/usage (alias /cost) shows the session's tokens and cost, priced from the LiteLLM rate table. The footer keeps a running total, and the session stores it so a resumed session continues the count.

Caps

Three optional bounds, checked once per prompt turn before any model work. Flags beat env; neither is persisted.

FlagEnvStops after
--max-turns NGRAY_MAX_TURNSN agent turns this process runs
--max-cost-usd XGRAY_MAX_COST_USDX dollars of priced spend
--max-wall-secs SGRAY_MAX_WALL_SECSS seconds since the process started

Prompt cache

Providers keep a prompt cache warm for roughly five idle minutes. The footer's ◷ countdown shows how long is left, and Gray warns in the transcript when a request re-billed tokens that should have been cached — an idle gap, a model switch, or an eviction on the provider's side.

Join the Discord

Chat with the people building and running Gray. Show what you made with it.

Join Discord →

Need help?

Open an issue on GitHub, run /feedback from the REPL, or ask in Discord.

Open an issue →