GUIDE

Cache-Friendly Context Compaction for OpenCode

A source-grounded guide to OpenCode context compaction, stable prompt prefixes, local-model prefill latency, installation, testing and operational trade-offs.

OpenCode conversation flowing through a stable prompt prefix, appended summary request and compacted local-model context

Built from multiple verified sources.

Quick answer

Cache-friendly compaction keeps OpenCode sessions usable near a model's context limit without forcing the model to reprocess the entire conversation. The plugin appends a summary request to the existing cached history, then sends later requests with the stable system-and-tools prefix, the summary and new turns. It can reduce local-model prefill delays, but it does not preserve every detail or guarantee cache hits.

# Cache-friendly context compaction for OpenCode: why prompt-cache stability matters

Direct answer

Cache-friendly compaction keeps OpenCode sessions usable near a model's context limit without forcing the model to reprocess the entire conversation. The plugin appends a summary request to the existing cached history, then sends later requests with the stable system-and-tools prefix, the summary and new turns. It can reduce local-model prefill delays, but it does not preserve every detail or guarantee cache hits.

What people searching for OpenCode context compaction need to know

Search results around OpenCode compaction repeatedly converge on four questions: what compaction removes, why a long session becomes slow before the first new token, how an alternative plugin differs from OpenCode's built-in mechanism, and which settings are safe for local models. Those questions are related, but they are not identical.

Context compaction solves a capacity problem. Prompt or prefix caching solves repeated-computation work. A summary may make the next prompt shorter while still destroying reuse of a previously cached prefix. Conversely, an append-only conversation can preserve cache reuse while continuing to grow toward the context limit. A useful design needs to manage both effects rather than treating “fewer tokens” as the only objective.

`opencode-cache-compact` is an MIT-licensed OpenCode plugin aimed primarily at locally hosted models, where prefilling a large prompt can dominate the wait before generation begins. Its central technique is simple: ask for the handoff summary by appending a normal turn to the conversation, then cut only the context sent to the model on subsequent requests. The full session remains on disk and visible in the OpenCode interface.

The project is young. Its repository was created on 21 September 2026 and, when reviewed on 1 October 2026, documented support for OpenCode V1 and the V2 beta plugin surface. That makes it an interesting operational tool, not a mature default that should be installed without testing.

Context length, compaction and prompt caching are different layers

A language model receives a finite sequence of tokens. OpenCode adds system instructions, tool schemas, conversation messages, tool results, files and the current request. As an agent works, that sequence grows. When it approaches the model's context window, older material must be removed, summarized or moved into some other memory system.

OpenCode's official documentation describes compaction as replacing older conversation with a summary while retaining recent work. This creates space so a long session can continue. It is inherently lossy: the summary is a representation of the earlier work, not the full record.

Prompt caching addresses a different cost. Many inference systems can reuse computation for a token-identical prefix that has already been processed. If the next request begins with the same system prompt, tools and history, the server may reuse cached key/value states instead of prefilling all those tokens again. Appending a new turn preserves the earlier prefix; inserting, deleting or rewriting material near the beginning can invalidate much of that reuse.

This distinction matters most for local inference. A cloud API can hide prefill behind a fast fleet and may offer discounted cached tokens. A single workstation has fixed memory bandwidth and compute. Re-reading a very long prompt can make time to first token much longer even if output generation remains fast.

For other ways to run models without paid inference, see the FreeAI Tokens guides to free local AI tools, free GPU credits and free AI API access.

How the cache-friendly plugin works

The plugin's README describes a three-stage sequence.

1. Measure usage and trip at a threshold

After each completed step, the plugin reads OpenCode's reported `tokens.total` and compares it with the active model's `limit.context`. The default threshold is 68 percent. If the provider does not report a context limit, the plugin has a configurable fallback of 131,072 tokens.

The percentage is not a universal recommendation. It leaves room for the summary, tool calls and a long agent step, but the right margin depends on the model, output budget and workload. A model with a 32k window and large tool results needs a different safety margin from a 128k model used mainly for short edits.

2. Summarize by appending

When the threshold is crossed, the plugin appends an ordinary user message asking the model for a handoff summary. It does not first rebuild a shortened version of the old conversation. If the inference server still holds the prefix cache, most of the existing history can be reused and only the new request plus generated summary need fresh work.

The summary output is capped at 1,200 tokens by default. That cap is a trade-off: a very short summary is cheap to prefill but may lose decisions, file paths, unfinished steps or verification evidence. A long summary retains more but becomes part of the new baseline on every later request.

3. Cut outgoing model context and resume

After receiving the summary, the plugin remembers the boundary and rewrites later outgoing requests to a compact shape: system instructions, tools, the summary and messages created after it. It can then send a short “continue” turn automatically.

The important word is outgoing. The stored OpenCode session is not erased. The terminal interface can still display the earlier conversation, while the model receives the reduced view. This separation improves recoverability and inspection, although the model can no longer reason over omitted details unless the summary retained them or a tool retrieves them again.

Why preserving the prefix can reduce delay

Autoregressive inference has two visibly different phases. During prefill, the model processes the supplied prompt and creates attention state. During decode, it generates new tokens one at a time. On long agent sessions, prefill can be the slow phase even when the model produces output at an acceptable tokens-per-second rate.

Official prompt-caching guidance from OpenAI states the general rule clearly: keep reusable instructions and conversation history stable, append new messages, and remember that summarization or truncation can reset cache reuse by changing the prefix. The exact cache implementation and lifetime vary by provider and local server, but the prefix-stability principle is broader than one API.

The plugin author reported that compaction on a Strix Halo system dropped from more than ten minutes to roughly one or two minutes. That is a project-author observation, not an independently reproduced benchmark. Hardware, model size, context length, quantization, server cache policy and whether the cache was evicted can change the result dramatically.

The defensible conclusion is therefore not “this plugin makes compaction five times faster.” It is that append-first summarization gives a cache-capable server the opportunity to reuse the already processed prefix, while a reconstructed summary request may force a cold prefill.

Installation for OpenCode V1 and V2

The repository documents different configuration shapes for the two OpenCode generations.

For V1, the npm package can be added to the `plugin` array with options such as a 68 percent threshold. V1 depends on the experimental `experimental.chat.messages.transform` hook. For V2 beta, the package is placed in the `plugins` array and uses `session.hook("context")` plus the event stream.

The project also documents loading from a local checkout. V1 can use a `file://` path. The current V2 beta build rejects file paths in config, so its documented workaround is a small re-export file in the global OpenCode plugins directory. Auto-discovered plugins receive default options; use the npm form if custom options are required.

If this plugin should be the only compactor, the README says to disable OpenCode's native automatic compaction. Running two independent compactors creates a race: the native path may summarize first, changing the context before the cache-friendly plugin acts.

A cautious installation sequence is:

  1. Confirm the exact OpenCode version and which plugin API it exposes.
  2. Record the model's real context limit and output budget.
  3. Confirm that the local inference server actually supports prefix or KV-cache reuse across requests.
  4. Pin the package version or Git commit instead of following an unreviewed moving target.
  5. Scope the plugin to the intended local `providerID/modelID` values.
  6. Disable native auto-compaction only after verifying that the plugin loads.
  7. Begin with `autoResume` disabled so the first cut can be inspected manually.
  8. Use debug logging during a disposable session and verify one summary, one cut and one resume.
  9. Re-enable automatic resume only after checking that unfinished work continues correctly.
  10. Keep an export or backup of the session before relying on the plugin for important work.

Settings that deserve deliberate choices

SettingDefaultPractical question
`threshold`68How much margin is needed for tools, summary and the longest expected step?
`models`allShould cloud or non-caching providers be excluded?
`summaryMaxTokens`1200How much state must survive without making the new prefix too large?
`contextLimit`131072Is the fallback accurate for the configured model?
`autoResume`trueShould work continue without human review after a lossy summary?
`abortOnTrip`trueIs stopping the active agent step acceptable when usage crosses the threshold?
`disablePrune`trueDoes the relevant OpenCode version still expose this behavior?
`debug`falseCan temporary logs be enabled without retaining sensitive content longer than intended?

The `models` option is especially important. The README says an empty list applies to all models. A user who alternates between a local llama.cpp endpoint and a hosted provider may want cache-friendly compaction only for the local model. Different providers expose context usage and cache behavior differently, so one policy is unlikely to be optimal everywhere.

What the plugin preserves—and what it cannot

It preserves the on-disk conversation, system prompt, tool definitions, generated summary and new turns after the cut. It also preserves a chance of using the existing prefix cache during the summary request.

It cannot guarantee that the server still has that cache. A restart, eviction, changed model, different worker, expired cache or modified prompt serialization can cause a miss. The first request after the cut must still prefill the new baseline containing system instructions, tools and summary.

It cannot make summarization lossless. Exact command output, failed hypotheses, rejected designs and subtle user preferences may disappear. For software work, the summary should explicitly retain current objectives, changed files, commands already run, test results, unresolved errors and the next safe action.

It also does not create persistent memory across unrelated sessions. The cut boundary is held in memory; restarting OpenCode means the plugin will not apply the previous cut until a later threshold trip. The full transcript remains available to the human, but the model's active context behavior restarts.

When cache-friendly compaction is a good fit

It is a strong candidate when all of the following are true:

It may be unnecessary when sessions rarely approach the context window, the provider already performs efficient native compaction, the server does not preserve a reusable cache, or latency is dominated by decode rather than prefill. It may also be the wrong choice for highly regulated work where a lossy, model-generated summary cannot become the authoritative state.

How to test it without trusting impressions

Create a repeatable A/B test using the same model, quantization, context, prompt history and server process. Warm the session with a long but non-sensitive transcript. Record the server's cached-token or prefix-hit metric if available, time to first token, total compaction time, prompt tokens processed and the factual completeness of the resulting summary.

Run the native compactor and the plugin in separate fresh sessions. Repeat each condition because the first run may include model loading, filesystem cache or shader compilation. A useful result includes distributions or at least several runs, not one stopwatch measurement.

Then test semantic continuity. Ask the agent to list the active task, edited files, verification already completed, known failures and next step. Compare it with the full transcript. A fast summary that loses a destructive-operation warning or an uncommitted change is not an improvement.

Finally, test failure paths: empty summary output, server restart, cache eviction, OpenCode restart, tool-call loop at the threshold and switching to an unlisted model. The project's own test suite covers its state machine and mock end-to-end request shapes, but it cannot certify every OpenCode build or inference server.

Security and privacy considerations

The plugin runs inside an AI coding agent with access to conversation content and potentially powerful tools. Review the package source and dependency lock before installation. Pin a version, restrict filesystem permissions and treat debug logs as potentially sensitive.

Because the full session remains on disk, compacting what the model sees is not data deletion. If the goal is privacy or retention reduction, the OpenCode storage location and backups must be handled separately. Likewise, sending a summary request to a hosted provider still sends the relevant conversation prefix to that provider; cache friendliness does not change the data boundary.

For broader project evaluation, use the FreeAI Tokens editorial hub and verify provider conditions in the free AI tiers directory.

Verdict

`opencode-cache-compact` applies a sound systems idea to a real local-LLM bottleneck: do not rewrite a reusable prefix before asking the model to summarize it. Its append, cut and resume flow can reduce cold-prefill work while keeping the human-visible session intact.

The benefit is conditional. It depends on cache-capable serving, stable serialization, accurate context reporting and a summary that preserves operational state. The plugin is also recent and integrates with experimental or beta OpenCode surfaces. Adopt it as a measured optimization: pin it, scope it to local models, test continuity and failure recovery, and compare cache metrics plus latency before replacing native compaction.

Sources and evidence

Sources were reviewed on 1 October 2026. Timing claims from the plugin author are clearly identified as project-reported and were not independently reproduced by FreeAI Tokens.

Frequently asked questions

What is context compaction in OpenCode?

Context compaction replaces older conversation with a shorter summary so the session can continue within the model's context window.

How is cache-friendly compaction different?

It requests the summary by appending to the existing history before cutting later outgoing context, giving a cache-capable server a chance to reuse the already processed prefix.

Does opencode-cache-compact delete the session history?

No. The repository says it rewrites only model-bound requests; the full OpenCode session remains stored and visible in the interface.

Does the plugin guarantee faster compaction?

No. The result depends on server-side cache reuse, eviction, model size, context length, hardware and prompt serialization, so it should be measured on the intended stack.

Sources and evidence