Added
- Keep-warm for Anthropic prompt caches, on by default. When a conversation goes quiet (a subagent running a long build or test, or you stepping away), the gateway refreshes its 5-minute cache 4 minutes 45 seconds after the last request by sending the same request again and closing it as soon as Anthropic reports usage, so you pay one cache read (0.05x to 0.1x the input price) instead of writing the whole conversation again at 1.25x when work resumes. It keeps going for up to
gateway.cache.keepWarmMinutesidle minutes (35 by default;0turns it off), only for conversations with at least 20,000 cached tokens, and stops for a conversation after any miss or provider error. A harness that warms its own cache (omp's main session does) resets the timer, so nothing is warmed twice. The requests it repeats are held in memory only (at most 64 MB), never on disk or in logs;x-superpi-cache: offon a request opts that conversation out. On the founder's own traffic since 1 October, re-caching after 5 to 60 minute pauses was $576 of $1,458 cache-write spend, almost all in omp subagents, which omp never warms; replaying that traffic, keep-warm would have saved about $327 for $71 of refreshes. The Usage page's Cache writes card shows refreshes sent, their cost, and what they saved.
Fixed
- The gateway answered a harness's cache-refresh request from its own response cache. omp repeats a request about 4.5 minutes later to keep Anthropic's prompt cache alive; the gateway replied with its stored copy, so the request never reached Anthropic, the provider's cache expired, the next turn wrote the whole conversation again, and the replay was even counted as gateway savings. Since 1 October this happened to 502 of omp's 1,421 refreshes. The gateway now answers an Anthropic request from its cache only when the stored response is under a minute old; later repeats go to Anthropic. OpenAI and ChatGPT requests are unchanged.