Using Gray

Context window & auto-compact

How Gray resolves an AI agent's context window, and how auto-compact summarizes older turns before you hit the limit and lose the thread.

Source: crates/gray/src/compact.rs

The context window resolves in order: the --context-window flag or GRAY_CONTEXT_WINDOW, then the value auto-fetched from the provider, then the LiteLLM model table, then a hardcoded per-model fallback.

/context            # inspect
/context 128k       # set — 128000, 128k and 1m all parse
/context auto       # clear the override

When usage crosses the window minus a 16k reserve, Gray compacts before the next turn: history is summarized into a two-message summary, the same flow as a manual /compact. On a context_length or max_tokens overflow error it compacts and retries once. Auto is the default; no flag needed.