LLM operations | September 30, 2026

Prompt caching needs a contract, not a hit-rate victory lap

OpenAI's new cache-miss diagnostics turn a silent cost and latency change into inspectable evidence. The useful operating model is broader: version the reusable prefix, test its boundaries, separate client drift from provider behavior, and optimize cost per accepted task.

Exact-prefix reuse Per-request evidence Accepted-task economics Sources checked Sep 30
Prompt prefix contract flowing through diagnostics, usage telemetry, and quality evaluation

A cache miss is now an evidence question

Prompt caching is often described as a pricing feature: send the same long prefix again and the provider can reuse prior computation. That description is correct but operationally incomplete. A production request is assembled from model choice, service tier, system policy, tool schemas, tool order, conversation state, retrieved material, user input, and runtime defaults. Any one of those can change the reusable prefix. Until September, many teams could see fewer cached tokens but could not distinguish their own payload drift from routing, expiration, eviction, or an unsupported configuration.

OpenAI's Prompt Cache Diagnostics, generally available for supported GPT-5.6-and-later Responses API models, adds a comparison primitive. A request can name a recent completed response as its baseline through prompt_cache_options.comparison_response_id. The response then reports whether the comparison found a hit, a miss with a reason, a missing diagnostic record, or an inconclusive result. The diagnostic may estimate how many baseline tokens were reusable and how many missed. Actual reuse and billing still come from the current response's usage fields.

That separation matters. Diagnostics do not load the old conversation, force a cache entry to remain warm, or change which entry the request uses. They explain one comparison. The production objective is therefore not “make diagnostics say hit.” It is to keep the stable part of the request intentionally stable, make every change attributable, and prove that reuse improves the economics of work the application actually accepts.

A cache hit is a property of one request. A cache contract is the system that explains why the hit should exist, what may invalidate it, and whether the application benefited.

Why aggregate hit rate misleads

A dashboard can show a high percentage of requests with some cached tokens while the expensive portion still changes every turn. A hit can reuse 2,000 tokens and process 50,000 new ones. Another workflow can record fewer hit events yet reuse a much larger policy and tool prefix. Retries, extra tool calls, failed evaluations, and longer completions can erase the savings. The relevant numerator is accepted work, not requests.

Current community reports illustrate the observability gap without proving universal behavior. A September r/OpenAI discussion attracted hundreds of votes around improved caching economics, while commenters also complained about unexplained misses. A measured openai/codex issue reports 252 cold requests among 100,475 mid-turn requests, or 0.25 percent, with clustering across independent threads. That is a developer-reported trace, not an OpenAI service benchmark. Its value is the method: preserve request evidence, identify when unchanged local prefixes go cold, and separate those events from client-controlled drift.

Exact-prefix reuse makes request construction part of the runtime

For supported OpenAI models, a reusable prefix must meet the documented minimum length; the current diagnostics guide uses 1,024 tokens for GPT-5.6-and-later supported models. Reuse depends on matching content from the beginning of the request and on compatible processing settings. A different model, service tier, tool definition, tool order, instruction, or early message can move or break the prefix. A cache key can influence affinity, but it does not make non-identical content identical.

The consequence is architectural. Prompt assembly cannot remain an incidental string concatenation hidden inside a framework. It should produce a typed artifact with a version, canonical serialization, content hashes, tool-catalog version, model contract, and declared dynamic boundary. The same artifact should be available to telemetry and evaluation so an incident can be replayed.

policy + reference + base tools
          | canonicalize and hash
          v
  versioned stable prefix -----------+
          |                           |
dynamic task + history + retrieval   | expected reuse contract
          |                           |
          v                           v
     Responses request -----> prompt-cache diagnostics
          |                           |
          +------ usage + latency ----+
                         |
                 quality evaluator
                         |
              cost per accepted task
Request elementPlacementChange policyCache consequence
Safety and business policyStable prefixVersioned, reviewed releaseIntentional invalidation
Base tool schemasStable prefixCanonical order and schema hashRename/reorder may miss
Optional toolsAppend after base setAppend-only when possibleLimits damage to earlier prefix
User/task factsDynamic suffixPer requestShould not disturb earlier prefix
Retrieved documentsDynamic suffix or versioned blockStable IDs and explicit versionsChanged content invalidates its segment
Conversation compactionDynamic stateRecorded compactor/versionExpected boundary shift

Some invalidations are correct. A revised compliance rule should replace the old cached prefix even if it increases cost. A model fallback may be necessary during an outage. Compaction may be required to stay within a context limit. The contract should label these as planned transitions rather than treating every miss as a defect.

Version the prefix before optimizing it

A useful manifest describes both content and compatibility. Hash the canonical form sent to the API, not a source directory whose build step can reorder JSON keys or tools. Record the code version that produced the request. If a framework injects hidden instructions, expose its version and observed token boundary in the run record even when the exact hidden text is unavailable.

apiVersion: llm.example/v1
kind: PromptPrefixContract
metadata:
  id: support-agent-prefix-v18
spec:
  model_contract: gpt-6-astra@2026-09
  service_tier: standard
  minimum_cacheable_tokens: 1024
  canonicalization: json-c14n-v1
  blocks:
    - {name: safety_policy, version: policy-42, mutable: false}
    - {name: support_manual, version: sha256:9f2..., mutable: false}
    - {name: base_tools, version: tools-17, mutable: append_only}
  dynamic_boundary: after_base_tools
  expected_invalidations: [model_change, policy_release, tool_breaking_change]
  owner: llm-platform
  quality_fixture_set: support-acceptance-v9
  rollback_prefix: support-agent-prefix-v17

Do not place volatile timestamps, trace IDs, randomized tool descriptions, per-request user facts, or unordered maps before the boundary. If the platform needs the data for logging, carry it in request metadata rather than prompt text when the API permits. When a tool catalog is large, keep frequently used base tools in a canonical order and append optional tools after the stable set. The OpenAI diagnostics documentation specifically calls out changed tools and ordering as miss causes.

A stable prefix is not automatically a small prefix. A long manual may be eligible for reuse but still slow evaluation, hide conflicting instructions, or make every policy edit expensive. Apply the same discipline used for code dependencies: delete dead material, separate modules by authority, test precedence, and make ownership explicit. The related instruction-debt guide covers that cleanup problem; caching should not become an excuse to preserve it.

Pair diagnostics with actual usage and a local diff

The API comparison should be one layer in a diagnostic stack. First compare the locally serialized prefix and settings. Then request provider diagnostics against a known recent baseline. Finally read actual cached-token usage, latency, error, retry, tool-call, and evaluator outcomes. A local match plus a provider-side miss is a different incident from a renamed tool or reordered schema.

import hashlib, json, time
from openai import OpenAI

client = OpenAI()

def canonical_hash(instructions, tools):
    payload = {"instructions": instructions, "tools": tools}
    wire = json.dumps(payload, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(wire.encode()).hexdigest()

baseline = client.responses.create(
    model="gpt-6-astra",
    instructions=POLICY,
    input=FIXTURE_INPUT,
    tools=BASE_TOOLS,
)

started = time.perf_counter()
current = client.responses.create(
    model="gpt-6-astra",
    instructions=POLICY,
    input=LIVE_INPUT,
    tools=BASE_TOOLS,
    prompt_cache_options={"comparison_response_id": baseline.id},
)

record = {
    "prefix_contract": "support-agent-prefix-v18",
    "prefix_hash": canonical_hash(POLICY, BASE_TOOLS),
    "comparison_response_id": baseline.id,
    "diagnostic": current.prompt_cache_diagnostics.model_dump()
        if current.prompt_cache_diagnostics else None,
    "input_tokens": current.usage.input_tokens,
    "cached_tokens": current.usage.input_tokens_details.cached_tokens,
    "latency_ms": round((time.perf_counter() - started) * 1000),
    "accepted": run_acceptance_checks(current),
}
emit_telemetry(record)

Protect response IDs and telemetry as operational data. OpenAI states that diagnostic records contain configuration metadata, token estimates, and hashes rather than raw prompts or outputs, are scoped to the organization, and expire after a short period. Your own logs may be more sensitive because they connect customer, workload, policy version, and usage. Apply retention, access, and deletion rules to the local evidence.

Use evidence states instead of guesswork

StateLocal prefixProvider diagnosticOperator action
CLIENT_DRIFTChangedMiss with matching reasonFix ordering or accept version change
EXPECTED_TRANSITIONChangedModel/settings/compaction changeRecord release and reset baseline
UNEXPLAINED_COLDUnchangedMiss or unavailableCorrelate routing, age, region, transport, incident
RECORD_EXPIREDUnknown or sameComparison not foundSelect a newer baseline; do not infer cause
REUSE_CONFIRMEDUnchangedHitVerify actual cached tokens and economics

Optimize cost per accepted task

Let R be cache-read tokens, W cache-write or newly processed reusable tokens when priced separately, U other uncached input, O output, and T tool cost. Use the provider's current model prices rather than copying a rate into application code. Add retries and all turns for one task. Then divide by tasks that pass the same frozen acceptance criteria.

task_cost = Σ(read_price*R + write_price*W + input_price*U
              + output_price*O + tool_cost*T)

cost_per_accepted_task = Σ(task_cost) / accepted_tasks

reused_token_ratio = Σ(R) / Σ(R + W + U)
acceptance_rate = accepted_tasks / attempted_tasks

Consider two variants over the same 200 support cases. Variant A reports cache hits on 94 percent of calls but makes 1,000 calls, including many retries, and accepts 150 cases. Variant B hits on 82 percent, makes 720 calls, and accepts 170. Variant A may look better on a request dashboard while costing more per accepted case. The evaluator, not the cache chart, decides the winner.

Track time to first token and end-to-end completion separately. Prompt reuse can reduce prefill work, but tool latency, queueing, model output, validation, and retries may dominate the task. Also segment by prefix version, model, service tier, region, transport, session age, compaction state, and tool-catalog version. A global average hides the exact cohort whose prefix drifted.

Failure modes span payload, platform, economics, and privacy

FailureEvidenceWrong responseBetter response
Dynamic prefix dataLocal hash changes every requestBlame provider cacheMove volatile data after boundary
Tool churnDiagnostics report tools changedRemove useful tools indiscriminatelyCanonicalize base set; append optional tools
Session-affinity shiftFresh/forked session loses reuseAssume cache key is a guaranteeRecord session, transport, key, and baseline
Provider-side cold eventSame local hash; misses cluster across threadsRewrite prompts repeatedlyCorrelate and escalate with bounded trace
Expired diagnostic recordComparison response not foundCall it a cache missUse a recent baseline and label unknown
Quality regressionReuse rises while acceptance fallsCelebrate savingsRollback prefix or model change
Isolation assumptionShared gateway credentials or scope unclearTreat cache as harmless optimizationReview tenant scope and side-channel risk

Research on prompt-cache timing and the 2026 CacheProbe work are reminders that reuse can reveal information through latency or shared gateway behavior if isolation is weak. Do not use timing tests against data you are not authorized to probe. For third-party gateways, document whose organization or account owns the cache scope, whether credentials are shared, and which provider data controls apply. The site's ZDR and private-processing guide explains why caching and retention features must be checked against the actual data-control contract.

A 14-day rollout that can disprove the optimization

  1. Days 1-2 - inventory: capture the exact serialized request, framework version, model, tier, tools, session metadata, usage, latency, retries, and quality result for representative tasks.
  2. Days 3-4 - contract: split stable policy/reference/tools from dynamic task, retrieval, and conversation content. Canonicalize and hash the stable prefix.
  3. Days 5-6 - fixtures: freeze success, failure, tool-choice, long-context, compaction, fallback, and privacy fixtures. Define accepted-task criteria before tuning.
  4. Days 7-8 - diagnostics: add recent-baseline comparisons to canary traffic. Classify client drift, expected transitions, expired records, and unexplained cold events.
  5. Days 9-10 - experiment: compare the current builder with stable tool ordering and a deliberate dynamic boundary. Keep workload and evaluator fixed.
  6. Days 11-12 - stress: test session forks, reconnects, compaction, tool appends, model fallback, idle periods, and concurrent threads.
  7. Days 13-14 - decide: promote only if cost and latency per accepted task improve without quality, privacy, or reliability regression. Keep the old prefix version as rollback.

Alert on unexplained shifts in reused tokens, not single misses. Include a minimum sample size and cohort. When a provider incident is suspected, retain a small reproducible trace with response IDs, timestamps, model, region when available, transport, prefix hash, session metadata, and diagnostic result. Do not attach customer prompt content to a public issue.

Frequently asked questions

Does Prompt Cache Diagnostics force a cache hit?

No. It compares a current response with a recent baseline and explains the observed relationship when possible. The baseline does not become conversation input, and the comparison request does not pin or restore a cache entry.

Is cache hit rate enough to measure savings?

No. Record tokens reused, newly processed input, cache writes when applicable, output, tools, latency, retries, evaluator result, and total cost per accepted task.

What belongs in the reusable prefix?

Stable policy, shared reference material, and canonical base-tool definitions. Put user-specific, task-specific, retrieved, temporal, and randomized content after the stable boundary unless there is a deliberate versioned reason not to.

Should we keep tool definitions fixed forever?

No. Stability is not stagnation. Release tool changes as explicit versions, validate them, and reset the diagnostic baseline. Append optional tools after a stable base set when that preserves correct behavior.

Can every miss be fixed in client code?

No. Local drift is fixable. Provider routing, eviction, transient infrastructure, and diagnostic-record availability are not. The evidence contract tells you which class you are dealing with.

Sources and further reading

Current product and community facts were checked on September 30, 2026. Community reports are evidence of practitioner problems, not service-wide performance claims.

Related guides