# Cache-friendly context compaction for OpenCode: why prompt-cache stability matters
Direct answer
Cache-friendly compaction keeps OpenCode sessions usable near a model's context limit without forcing the model to reprocess the entire conversation. The plugin appends a summary request to the existing cached history, then sends later requests with the stable system-and-tools prefix, the summary and new turns. It can reduce local-model prefill delays, but it does not preserve every detail or guarantee cache hits.
What people searching for OpenCode context compaction need to know
Search results around OpenCode compaction repeatedly converge on four questions: what compaction removes, why a long session becomes slow before the first new token, how an alternative plugin differs from OpenCode's built-in mechanism, and which settings are safe for local models. Those questions are related, but they are not identical.
Context compaction solves a capacity problem. Prompt or prefix caching solves repeated-computation work. A summary may make the next prompt shorter while still destroying reuse of a previously cached prefix. Conversely, an append-only conversation can preserve cache reuse while continuing to grow toward the context limit. A useful design needs to manage both effects rather than treating “fewer tokens” as the only objective.
`opencode-cache-compact` is an MIT-licensed OpenCode plugin aimed primarily at locally hosted models, where prefilling a large prompt can dominate the wait before generation begins. Its central technique is simple: ask for the handoff summary by appending a normal turn to the conversation, then cut only the context sent to the model on subsequent requests. The full session remains on disk and visible in the OpenCode interface.
The project is young. Its repository was created on 21 September 2026 and, when reviewed on 1 October 2026, documented support for OpenCode V1 and the V2 beta plugin surface. That makes it an interesting operational tool, not a mature default that should be installed without testing.
Context length, compaction and prompt caching are different layers
A language model receives a finite sequence of tokens. OpenCode adds system instructions, tool schemas, conversation messages, tool results, files and the current request. As an agent works, that sequence grows. When it approaches the model's context window, older material must be removed, summarized or moved into some other memory system.
OpenCode's official documentation describes compaction as replacing older conversation with a summary while retaining recent work. This creates space so a long session can continue. It is inherently lossy: the summary is a representation of the earlier work, not the full record.
Prompt caching addresses a different cost. Many inference systems can reuse computation for a token-identical prefix that has already been processed. If the next request begins with the same system prompt, tools and history, the server may reuse cached key/value states instead of prefilling all those tokens again. Appending a new turn preserves the earlier prefix; inserting, deleting or rewriting material near the beginning can invalidate much of that reuse.
This distinction matters most for local inference. A cloud API can hide prefill behind a fast fleet and may offer discounted cached tokens. A single workstation has fixed memory bandwidth and compute. Re-reading a very long prompt can make time to first token much longer even if output generation remains fast.
For other ways to run models without paid inference, see the FreeAI Tokens guides to free local AI tools, free GPU credits and free AI API access.
How the cache-friendly plugin works
The plugin's README describes a three-stage sequence.
1. Measure usage and trip at a threshold
After each completed step, the plugin reads OpenCode's reported `tokens.total` and compares it with the active model's `limit.context`. The default threshold is 68 percent. If the provider does not report a context limit, the plugin has a configurable fallback of 131,072 tokens.
The percentage is not a universal recommendation. It leaves room for the summary, tool calls and a long agent step, but the right margin depends on the model, output budget and workload. A model with a 32k window and large tool results needs a different safety margin from a 128k model used mainly for short edits.
2. Summarize by appending
When the threshold is crossed, the plugin appends an ordinary user message asking the model for a handoff summary. It does not first rebuild a shortened version of the old conversation. If the inference server still holds the prefix cache, most of the existing history can be reused and only the new request plus generated summary need fresh work.
The summary output is capped at 1,200 tokens by default. That cap is a trade-off: a very short summary is cheap to prefill but may lose decisions, file paths, unfinished steps or verification evidence. A long summary retains more but becomes part of the new baseline on every later request.
3. Cut outgoing model context and resume
After receiving the summary, the plugin remembers the boundary and rewrites later outgoing requests to a compact shape: system instructions, tools, the summary and messages created after it. It can then send a short “continue” turn automatically.
The important word is outgoing. The stored OpenCode session is not erased. The terminal interface can still display the earlier conversation, while the model receives the reduced view. This separation improves recoverability and inspection, although the model can no longer reason over omitted details unless the summary retained them or a tool retrieves them again.
Why preserving the prefix can reduce delay
Autoregressive inference has two visibly different phases. During prefill, the model processes the supplied prompt and creates attention state. During decode, it generates new tokens one at a time. On long agent sessions, prefill can be the slow phase even when the model produces output at an acceptable tokens-per-second rate.
Official prompt-caching guidance from OpenAI states the general rule clearly: keep reusable instructions and conversation history stable, append new messages, and remember that summarization or truncation can reset cache reuse by changing the prefix. The exact cache implementation and lifetime vary by provider and local server, but the prefix-stability principle is broader than one API.
The plugin author reported that compaction on a Strix Halo system dropped from more than ten minutes to roughly one or two minutes. That is a project-author observation, not an independently reproduced benchmark. Hardware, model size, context length, quantization, server cache policy and whether the cache was evicted can change the result dramatically.
The defensible conclusion is therefore not “this plugin makes compaction five times faster.” It is that append-first summarization gives a cache-capable server the opportunity to reuse the already processed prefix, while a reconstructed summary request may force a cold prefill.
Installation for OpenCode V1 and V2
The repository documents different configuration shapes for the two OpenCode generations.
For V1, the npm package can be added to the `plugin` array with options such as a 68 percent threshold. V1 depends on the experimental `experimental.chat.messages.transform` hook. For V2 beta, the package is placed in the `plugins` array and uses `session.hook("context")` plus the event stream.
The project also documents loading from a local checkout. V1 can use a `file://` path. The current V2 beta build rejects file paths in config, so its documented workaround is a small re-export file in the global OpenCode plugins directory. Auto-discovered plugins receive default options; use the npm form if custom options are required.
If this plugin should be the only compactor, the README says to disable OpenCode's native automatic compaction. Running two independent compactors creates a race: the native path may summarize first, changing the context before the cache-friendly plugin acts.
A cautious installation sequence is:
- Confirm the exact OpenCode version and which plugin API it exposes.
- Record the model's real context limit and output budget.
- Confirm that the local inference server actually supports prefix or KV-cache reuse across requests.
- Pin the package version or Git commit instead of following an unreviewed moving target.
- Scope the plugin to the intended local `providerID/modelID` values.
- Disable native auto-compaction only after verifying that the plugin loads.
- Begin with `autoResume` disabled so the first cut can be inspected manually.
- Use debug logging during a disposable session and verify one summary, one cut and one resume.
- Re-enable automatic resume only after checking that unfinished work continues correctly.
- Keep an export or backup of the session before relying on the plugin for important work.
Settings that deserve deliberate choices
| Setting | Default | Practical question |
|---|---|---|
| `threshold` | 68 | How much margin is needed for tools, summary and the longest expected step? |
| `models` | all | Should cloud or non-caching providers be excluded? |
| `summaryMaxTokens` | 1200 | How much state must survive without making the new prefix too large? |
| `contextLimit` | 131072 | Is the fallback accurate for the configured model? |
| `autoResume` | true | Should work continue without human review after a lossy summary? |
| `abortOnTrip` | true | Is stopping the active agent step acceptable when usage crosses the threshold? |
| `disablePrune` | true | Does the relevant OpenCode version still expose this behavior? |
| `debug` | false | Can temporary logs be enabled without retaining sensitive content longer than intended? |
The `models` option is especially important. The README says an empty list applies to all models. A user who alternates between a local llama.cpp endpoint and a hosted provider may want cache-friendly compaction only for the local model. Different providers expose context usage and cache behavior differently, so one policy is unlikely to be optimal everywhere.
What the plugin preserves—and what it cannot
It preserves the on-disk conversation, system prompt, tool definitions, generated summary and new turns after the cut. It also preserves a chance of using the existing prefix cache during the summary request.
It cannot guarantee that the server still has that cache. A restart, eviction, changed model, different worker, expired cache or modified prompt serialization can cause a miss. The first request after the cut must still prefill the new baseline containing system instructions, tools and summary.
It cannot make summarization lossless. Exact command output, failed hypotheses, rejected designs and subtle user preferences may disappear. For software work, the summary should explicitly retain current objectives, changed files, commands already run, test results, unresolved errors and the next safe action.
It also does not create persistent memory across unrelated sessions. The cut boundary is held in memory; restarting OpenCode means the plugin will not apply the previous cut until a later threshold trip. The full transcript remains available to the human, but the model's active context behavior restarts.
When cache-friendly compaction is a good fit
It is a strong candidate when all of the following are true:
- OpenCode is connected to a locally hosted, OpenAI-compatible model server.
- Long prefill time is a measured problem, not merely an assumption.
- The server reuses prompt prefixes across sequential requests.
- Sessions contain long, append-only tool conversations.
- The operator can inspect logs and recover if a summary omits needed state.
- A young third-party plugin is acceptable in the threat and maintenance model.
It may be unnecessary when sessions rarely approach the context window, the provider already performs efficient native compaction, the server does not preserve a reusable cache, or latency is dominated by decode rather than prefill. It may also be the wrong choice for highly regulated work where a lossy, model-generated summary cannot become the authoritative state.
How to test it without trusting impressions
Create a repeatable A/B test using the same model, quantization, context, prompt history and server process. Warm the session with a long but non-sensitive transcript. Record the server's cached-token or prefix-hit metric if available, time to first token, total compaction time, prompt tokens processed and the factual completeness of the resulting summary.
Run the native compactor and the plugin in separate fresh sessions. Repeat each condition because the first run may include model loading, filesystem cache or shader compilation. A useful result includes distributions or at least several runs, not one stopwatch measurement.
Then test semantic continuity. Ask the agent to list the active task, edited files, verification already completed, known failures and next step. Compare it with the full transcript. A fast summary that loses a destructive-operation warning or an uncommitted change is not an improvement.
Finally, test failure paths: empty summary output, server restart, cache eviction, OpenCode restart, tool-call loop at the threshold and switching to an unlisted model. The project's own test suite covers its state machine and mock end-to-end request shapes, but it cannot certify every OpenCode build or inference server.
Security and privacy considerations
The plugin runs inside an AI coding agent with access to conversation content and potentially powerful tools. Review the package source and dependency lock before installation. Pin a version, restrict filesystem permissions and treat debug logs as potentially sensitive.
Because the full session remains on disk, compacting what the model sees is not data deletion. If the goal is privacy or retention reduction, the OpenCode storage location and backups must be handled separately. Likewise, sending a summary request to a hosted provider still sends the relevant conversation prefix to that provider; cache friendliness does not change the data boundary.
For broader project evaluation, use the FreeAI Tokens editorial hub and verify provider conditions in the free AI tiers directory.
Verdict
`opencode-cache-compact` applies a sound systems idea to a real local-LLM bottleneck: do not rewrite a reusable prefix before asking the model to summarize it. Its append, cut and resume flow can reduce cold-prefill work while keeping the human-visible session intact.
The benefit is conditional. It depends on cache-capable serving, stable serialization, accurate context reporting and a summary that preserves operational state. The plugin is also recent and integrates with experimental or beta OpenCode surfaces. Adopt it as a measured optimization: pin it, scope it to local models, test continuity and failure recovery, and compare cache metrics plus latency before replacing native compaction.
Sources and evidence
- opencode-cache-compact repository and README: design, installation, defaults, compatibility, tests and caveats.
- OpenCode compaction documentation: native compaction purpose, context shape and configuration.
- OpenCode plugin documentation: supported plugin installation and lifecycle concepts.
- OpenAI prompt-caching guide: independent primary-source explanation of stable prefixes, append-only history and cache invalidation by compaction.
Sources were reviewed on 1 October 2026. Timing claims from the plugin author are clearly identified as project-reported and were not independently reproduced by FreeAI Tokens.