A cache miss is now an evidence question
Prompt caching is often described as a pricing feature: send the same long prefix again and the provider can reuse prior computation. That description is correct but operationally incomplete. A production request is assembled from model choice, service tier, system policy, tool schemas, tool order, conversation state, retrieved material, user input, and runtime defaults. Any one of those can change the reusable prefix. Until September, many teams could see fewer cached tokens but could not distinguish their own payload drift from routing, expiration, eviction, or an unsupported configuration.
OpenAI's Prompt Cache Diagnostics, generally available for supported GPT-5.6-and-later Responses API models, adds a comparison primitive. A request can name a recent completed response as its baseline through prompt_cache_options.comparison_response_id. The response then reports whether the comparison found a hit, a miss with a reason, a missing diagnostic record, or an inconclusive result. The diagnostic may estimate how many baseline tokens were reusable and how many missed. Actual reuse and billing still come from the current response's usage fields.
That separation matters. Diagnostics do not load the old conversation, force a cache entry to remain warm, or change which entry the request uses. They explain one comparison. The production objective is therefore not “make diagnostics say hit.” It is to keep the stable part of the request intentionally stable, make every change attributable, and prove that reuse improves the economics of work the application actually accepts.
A cache hit is a property of one request. A cache contract is the system that explains why the hit should exist, what may invalidate it, and whether the application benefited.
Why aggregate hit rate misleads
A dashboard can show a high percentage of requests with some cached tokens while the expensive portion still changes every turn. A hit can reuse 2,000 tokens and process 50,000 new ones. Another workflow can record fewer hit events yet reuse a much larger policy and tool prefix. Retries, extra tool calls, failed evaluations, and longer completions can erase the savings. The relevant numerator is accepted work, not requests.
Current community reports illustrate the observability gap without proving universal behavior. A September r/OpenAI discussion attracted hundreds of votes around improved caching economics, while commenters also complained about unexplained misses. A measured openai/codex issue reports 252 cold requests among 100,475 mid-turn requests, or 0.25 percent, with clustering across independent threads. That is a developer-reported trace, not an OpenAI service benchmark. Its value is the method: preserve request evidence, identify when unchanged local prefixes go cold, and separate those events from client-controlled drift.
Exact-prefix reuse makes request construction part of the runtime
For supported OpenAI models, a reusable prefix must meet the documented minimum length; the current diagnostics guide uses 1,024 tokens for GPT-5.6-and-later supported models. Reuse depends on matching content from the beginning of the request and on compatible processing settings. A different model, service tier, tool definition, tool order, instruction, or early message can move or break the prefix. A cache key can influence affinity, but it does not make non-identical content identical.
The consequence is architectural. Prompt assembly cannot remain an incidental string concatenation hidden inside a framework. It should produce a typed artifact with a version, canonical serialization, content hashes, tool-catalog version, model contract, and declared dynamic boundary. The same artifact should be available to telemetry and evaluation so an incident can be replayed.
policy + reference + base tools
| canonicalize and hash
v
versioned stable prefix -----------+
| |
dynamic task + history + retrieval | expected reuse contract
| |
v v
Responses request -----> prompt-cache diagnostics
| |
+------ usage + latency ----+
|
quality evaluator
|
cost per accepted task
| Request element | Placement | Change policy | Cache consequence |
| Safety and business policy | Stable prefix | Versioned, reviewed release | Intentional invalidation |
| Base tool schemas | Stable prefix | Canonical order and schema hash | Rename/reorder may miss |
| Optional tools | Append after base set | Append-only when possible | Limits damage to earlier prefix |
| User/task facts | Dynamic suffix | Per request | Should not disturb earlier prefix |
| Retrieved documents | Dynamic suffix or versioned block | Stable IDs and explicit versions | Changed content invalidates its segment |
| Conversation compaction | Dynamic state | Recorded compactor/version | Expected boundary shift |
Some invalidations are correct. A revised compliance rule should replace the old cached prefix even if it increases cost. A model fallback may be necessary during an outage. Compaction may be required to stay within a context limit. The contract should label these as planned transitions rather than treating every miss as a defect.
Version the prefix before optimizing it
A useful manifest describes both content and compatibility. Hash the canonical form sent to the API, not a source directory whose build step can reorder JSON keys or tools. Record the code version that produced the request. If a framework injects hidden instructions, expose its version and observed token boundary in the run record even when the exact hidden text is unavailable.
apiVersion: llm.example/v1
kind: PromptPrefixContract
metadata:
id: support-agent-prefix-v18
spec:
model_contract: gpt-6-astra@2026-09
service_tier: standard
minimum_cacheable_tokens: 1024
canonicalization: json-c14n-v1
blocks:
- {name: safety_policy, version: policy-42, mutable: false}
- {name: support_manual, version: sha256:9f2..., mutable: false}
- {name: base_tools, version: tools-17, mutable: append_only}
dynamic_boundary: after_base_tools
expected_invalidations: [model_change, policy_release, tool_breaking_change]
owner: llm-platform
quality_fixture_set: support-acceptance-v9
rollback_prefix: support-agent-prefix-v17
Do not place volatile timestamps, trace IDs, randomized tool descriptions, per-request user facts, or unordered maps before the boundary. If the platform needs the data for logging, carry it in request metadata rather than prompt text when the API permits. When a tool catalog is large, keep frequently used base tools in a canonical order and append optional tools after the stable set. The OpenAI diagnostics documentation specifically calls out changed tools and ordering as miss causes.
A stable prefix is not automatically a small prefix. A long manual may be eligible for reuse but still slow evaluation, hide conflicting instructions, or make every policy edit expensive. Apply the same discipline used for code dependencies: delete dead material, separate modules by authority, test precedence, and make ownership explicit. The related instruction-debt guide covers that cleanup problem; caching should not become an excuse to preserve it.
Pair diagnostics with actual usage and a local diff
The API comparison should be one layer in a diagnostic stack. First compare the locally serialized prefix and settings. Then request provider diagnostics against a known recent baseline. Finally read actual cached-token usage, latency, error, retry, tool-call, and evaluator outcomes. A local match plus a provider-side miss is a different incident from a renamed tool or reordered schema.
import hashlib, json, time
from openai import OpenAI
client = OpenAI()
def canonical_hash(instructions, tools):
payload = {"instructions": instructions, "tools": tools}
wire = json.dumps(payload, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(wire.encode()).hexdigest()
baseline = client.responses.create(
model="gpt-6-astra",
instructions=POLICY,
input=FIXTURE_INPUT,
tools=BASE_TOOLS,
)
started = time.perf_counter()
current = client.responses.create(
model="gpt-6-astra",
instructions=POLICY,
input=LIVE_INPUT,
tools=BASE_TOOLS,
prompt_cache_options={"comparison_response_id": baseline.id},
)
record = {
"prefix_contract": "support-agent-prefix-v18",
"prefix_hash": canonical_hash(POLICY, BASE_TOOLS),
"comparison_response_id": baseline.id,
"diagnostic": current.prompt_cache_diagnostics.model_dump()
if current.prompt_cache_diagnostics else None,
"input_tokens": current.usage.input_tokens,
"cached_tokens": current.usage.input_tokens_details.cached_tokens,
"latency_ms": round((time.perf_counter() - started) * 1000),
"accepted": run_acceptance_checks(current),
}
emit_telemetry(record)
Protect response IDs and telemetry as operational data. OpenAI states that diagnostic records contain configuration metadata, token estimates, and hashes rather than raw prompts or outputs, are scoped to the organization, and expire after a short period. Your own logs may be more sensitive because they connect customer, workload, policy version, and usage. Apply retention, access, and deletion rules to the local evidence.
Use evidence states instead of guesswork
| State | Local prefix | Provider diagnostic | Operator action |
| CLIENT_DRIFT | Changed | Miss with matching reason | Fix ordering or accept version change |
| EXPECTED_TRANSITION | Changed | Model/settings/compaction change | Record release and reset baseline |
| UNEXPLAINED_COLD | Unchanged | Miss or unavailable | Correlate routing, age, region, transport, incident |
| RECORD_EXPIRED | Unknown or same | Comparison not found | Select a newer baseline; do not infer cause |
| REUSE_CONFIRMED | Unchanged | Hit | Verify actual cached tokens and economics |
Optimize cost per accepted task
Let R be cache-read tokens, W cache-write or newly processed reusable tokens when priced separately, U other uncached input, O output, and T tool cost. Use the provider's current model prices rather than copying a rate into application code. Add retries and all turns for one task. Then divide by tasks that pass the same frozen acceptance criteria.
task_cost = Σ(read_price*R + write_price*W + input_price*U
+ output_price*O + tool_cost*T)
cost_per_accepted_task = Σ(task_cost) / accepted_tasks
reused_token_ratio = Σ(R) / Σ(R + W + U)
acceptance_rate = accepted_tasks / attempted_tasks
Consider two variants over the same 200 support cases. Variant A reports cache hits on 94 percent of calls but makes 1,000 calls, including many retries, and accepts 150 cases. Variant B hits on 82 percent, makes 720 calls, and accepts 170. Variant A may look better on a request dashboard while costing more per accepted case. The evaluator, not the cache chart, decides the winner.
Track time to first token and end-to-end completion separately. Prompt reuse can reduce prefill work, but tool latency, queueing, model output, validation, and retries may dominate the task. Also segment by prefix version, model, service tier, region, transport, session age, compaction state, and tool-catalog version. A global average hides the exact cohort whose prefix drifted.
Failure modes span payload, platform, economics, and privacy
| Failure | Evidence | Wrong response | Better response |
| Dynamic prefix data | Local hash changes every request | Blame provider cache | Move volatile data after boundary |
| Tool churn | Diagnostics report tools changed | Remove useful tools indiscriminately | Canonicalize base set; append optional tools |
| Session-affinity shift | Fresh/forked session loses reuse | Assume cache key is a guarantee | Record session, transport, key, and baseline |
| Provider-side cold event | Same local hash; misses cluster across threads | Rewrite prompts repeatedly | Correlate and escalate with bounded trace |
| Expired diagnostic record | Comparison response not found | Call it a cache miss | Use a recent baseline and label unknown |
| Quality regression | Reuse rises while acceptance falls | Celebrate savings | Rollback prefix or model change |
| Isolation assumption | Shared gateway credentials or scope unclear | Treat cache as harmless optimization | Review tenant scope and side-channel risk |
Research on prompt-cache timing and the 2026 CacheProbe work are reminders that reuse can reveal information through latency or shared gateway behavior if isolation is weak. Do not use timing tests against data you are not authorized to probe. For third-party gateways, document whose organization or account owns the cache scope, whether credentials are shared, and which provider data controls apply. The site's ZDR and private-processing guide explains why caching and retention features must be checked against the actual data-control contract.
A 14-day rollout that can disprove the optimization
- Days 1-2 - inventory: capture the exact serialized request, framework version, model, tier, tools, session metadata, usage, latency, retries, and quality result for representative tasks.
- Days 3-4 - contract: split stable policy/reference/tools from dynamic task, retrieval, and conversation content. Canonicalize and hash the stable prefix.
- Days 5-6 - fixtures: freeze success, failure, tool-choice, long-context, compaction, fallback, and privacy fixtures. Define accepted-task criteria before tuning.
- Days 7-8 - diagnostics: add recent-baseline comparisons to canary traffic. Classify client drift, expected transitions, expired records, and unexplained cold events.
- Days 9-10 - experiment: compare the current builder with stable tool ordering and a deliberate dynamic boundary. Keep workload and evaluator fixed.
- Days 11-12 - stress: test session forks, reconnects, compaction, tool appends, model fallback, idle periods, and concurrent threads.
- Days 13-14 - decide: promote only if cost and latency per accepted task improve without quality, privacy, or reliability regression. Keep the old prefix version as rollback.
Alert on unexplained shifts in reused tokens, not single misses. Include a minimum sample size and cohort. When a provider incident is suspected, retain a small reproducible trace with response IDs, timestamps, model, region when available, transport, prefix hash, session metadata, and diagnostic result. Do not attach customer prompt content to a public issue.
Frequently asked questions
Does Prompt Cache Diagnostics force a cache hit?
No. It compares a current response with a recent baseline and explains the observed relationship when possible. The baseline does not become conversation input, and the comparison request does not pin or restore a cache entry.
Is cache hit rate enough to measure savings?
No. Record tokens reused, newly processed input, cache writes when applicable, output, tools, latency, retries, evaluator result, and total cost per accepted task.
What belongs in the reusable prefix?
Stable policy, shared reference material, and canonical base-tool definitions. Put user-specific, task-specific, retrieved, temporal, and randomized content after the stable boundary unless there is a deliberate versioned reason not to.
Should we keep tool definitions fixed forever?
No. Stability is not stagnation. Release tool changes as explicit versions, validate them, and reset the diagnostic baseline. Append optional tools after a stable base set when that preserves correct behavior.
Can every miss be fixed in client code?
No. Local drift is fixable. Provider routing, eviction, transient infrastructure, and diagnostic-record availability are not. The evidence contract tells you which class you are dealing with.
Sources and further reading
Current product and community facts were checked on September 30, 2026. Community reports are evidence of practitioner problems, not service-wide performance claims.