Using Gray
Context window & auto-compact
How Gray resolves an AI agent's context window, and how auto-compact summarizes older turns before you hit the limit and lose the thread.
Source: crates/gray/src/compact.rs
The context window resolves in order: the --context-window flag or GRAY_CONTEXT_WINDOW, then the value auto-fetched from the provider, then the LiteLLM model table, then a hardcoded per-model fallback.
/context # inspect /context 128k # set — 128000, 128k and 1m all parse /context auto # clear the override
When usage crosses the window minus a 16k reserve, Gray compacts before the next turn: history is summarized into a two-message summary, the same flow as a manual /compact. On a context_length or max_tokens overflow error it compacts and retries once. Auto is the default; no flag needed.