On-device inference | October 10, 2026

ML Drift makes GPU portability a conformance problem

Google's new GPU engine can place one compute graph across OpenCL, OpenGL ES, Metal, and WebGPU. That is a valuable integration layer, not proof that every backend produces the same output, latency, memory use, fallback behavior, or observability. Migrate with a device-level acceptance contract.

LiteRT migrationBackend conformanceApp-level benchmarksApache 2.0Sources checked Oct 10
A device matrix comparing multiple GPU inference backends

The release changes the default GPU path, not the laws of heterogeneous hardware

Google open-sourced ML Drift on October 8 as the GPU compute engine behind LiteRT acceleration and as a standalone C++ library. The project targets Android, iOS and macOS, web, Windows, and Linux through several lower-level APIs. Google also said the legacy TensorFlow Lite GPU delegate will receive no new features and encouraged applications to move toward ML Drift. That makes this an operational migration, not merely another repository launch.

The interesting abstraction is a unified kernel language and compute graph that can be lowered to OpenCL, OpenGL ES, Metal, or WebGPU. Application developers get a higher-level LiteRT path with graph partitioning, tensor virtualization, zero-copy buffers, and asynchronous execution. Runtime engineers can use the lower-level graph and shader interfaces directly. Both paths aim to reduce the amount of platform-specific inference code a team owns.

The current developer signal is real but bounded. At our October 10 check, the official repository had 70 stars and six open issues. A focused r/LocalLLaMA thread had 103 votes and 13 comments. The most useful comments did not celebrate “GPU acceleration” in the abstract. They asked whether Google's order-of-magnitude claim transfers beyond the published mobile workloads, how the engine compares with llama.cpp on desktop GPUs, why Android emphasizes OpenCL, and which runtime problem justifies another engine.

Those questions point to the right thesis: source-level portability is not runtime equivalence. A common API can dispatch different kernels, accumulations, memory layouts, caches, drivers, and fallbacks. An Android device may choose OpenCL, another may fall back to OpenGL ES, macOS may use Metal, and a Windows build may reach WebGPU through Direct3D. The application still owns proof that each supported cohort behaves acceptably.

ML Drift can make one implementation portable. Only a conformance matrix can make the release decision portable.

Follow the graph from model to the backend that actually ran

A migration begins with a model artifact, operator set, shapes, quantization policy, and runtime version. LiteRT compiles that model for requested accelerators. Supported subgraphs move to the GPU path; unsupported work may remain on CPU or use another accelerator. ML Drift then lowers GPU work through its kernel and backend layers. Buffers, synchronization, shader compilation, cache state, and driver behavior determine what happens on the device.

model.tflite + immutable hash
  -> LiteRT CompiledModel graph partition
  -> requested accelerator: GPU
  -> delegated GPU subgraphs + explicit fallback subgraphs
  -> ML Drift unified kernel / graph representation
  -> OpenCL | OpenGL ES | Metal | WebGPU
  -> device driver + shader cache + memory allocator
  -> tensors, timings, counters, errors, and app-visible result

The chain contains two different compatibility promises. API compatibility means the application can construct and invoke a model. Behavioral compatibility means the deployed path still meets output, latency, memory, energy, crash, and observability requirements. The first is documented by packages and interfaces. The second belongs in your tests.

Platform cohortLikely GPU pathEvidence to retain
AndroidOpenCL with OpenGL ES fallbackSoC, OS, driver, chosen backend, delegated ops, fallback count
iOS / macOSMetal; some product paths may also expose WebGPUOS and device family, Metal feature set, memory pressure, profiler availability
WebWebGPU in a supported browserBrowser build, adapter, feature flags, shader compilation and cache state
LinuxOpenCL or WebGPU over VulkanGPU and driver, ICD or Vulkan stack, selected path, headless/app context
WindowsWebGPU over Direct3DGPU, driver, DirectX Shader Compiler files, selected adapter, cache state

Do not trust a requested accelerator field as evidence that the whole graph ran there. Capture the compiled partition and actual backend at runtime. A model that delegates 96% of arithmetic can still lose its expected latency benefit if a small unsupported operator forces repeated device-to-host transfers.

Version a conformance contract per model and device family

Freeze the artifacts before measuring. Record the model and tokenizer hashes, LiteRT and ML Drift revisions, build flags, precision, delegate options, expected backend, input corpus, reference outputs, tolerance policy, benchmark method, and acceptance thresholds. If any of these changes, the release needs a new result rather than an inherited green badge.

schema: gpu-conformance/v1
model:
  uri: app://models/vision-v14.tflite
  sha256: "..."
runtime:
  litert: "2.2.0"
  mlDriftCommit: "..."
  requestedAccelerator: gpu
cohort:
  platform: android
  osRange: "15-17"
  socFamilies: ["adreno-7xx", "mali-g7xx"]
acceptance:
  output:
    maxAbsError: 0.002
    top1Agreement: 0.999
  delegation:
    minimumGpuOpsPct: 95
    unexpectedFallbacks: 0
  latencyMs:
    warmP50: 18
    warmP95: 27
    coldP95: 80
  memoryPeakMb: 420
  crashFreeSessions: 0.9995
  thermalRunsBeforeP95Breach: 20
observability:
  backendRecorded: required
  partitionRecorded: required
  perOpProfile: preferred

Choose tolerances before looking at the new output. The policy should match the application's risk. A camera effect may tolerate small pixel drift while a biometric or safety-related model needs a stricter task-level validation. For language models, exact-token agreement can be too strict for benign floating-point variation yet too weak for hidden logit drift. Retain tensor-level or metric-level comparisons for selected layers and adversarial cases.

Make fallback part of the contract. “Correct output” from a CPU fallback is not a successful GPU migration if the product promise is frame-time or battery improvement. Conversely, a fully delegated graph that changes the result beyond tolerance is not acceptable because it looks fast.

Choose the narrow migration path before the ambitious one

Google documents two LiteRT routes. Existing TensorFlow Lite applications can first change packages while retaining the Interpreter-style API. New work can adopt the v2 CompiledModel API for explicit accelerator selection, zero-copy tensor buffers, NPU support, and asynchronous execution. The second path offers more capability but changes more of the application's runtime contract.

// Android: pin a tested LiteRT release.
dependencies {
  implementation("com.google.ai.edge.litert:litert:2.2.0")
}

// Pseudocode: request the GPU, then record what compiled.
val options = CompiledModel.Options(
  accelerators = listOf(Accelerator.GPU)
)
val compiled = CompiledModel.create(modelBytes, options)
telemetry.record(compiled.backendInfo())
telemetry.record(compiled.partitionSummary())
val result = compiled.run(reusableInputBuffers)

Do not combine package migration, API migration, quantization change, model update, and new asynchronous scheduling in one release. Start with one pinned model and a representative cohort. Keep the legacy path available behind a server-controlled flag. After output and fallback gates pass, introduce buffer reuse and asynchronous execution with separate UI-frame and concurrency tests.

PathGood fitMain review
Interpreter-compatible package changeLowest-risk first step for an existing TFLite appPackaging, delegate selection, output and performance regression
CompiledModel synchronousNew runtime with explicit accelerator and buffer controlPartitioning, tensor ownership, fallback and memory behavior
CompiledModel asynchronous / zero-copyCamera, media, and interactive pipelinesSynchronization, lifetime, contention, frame pacing and cancellation
Standalone ML Drift C++Custom graphics or inference runtimeGraph construction, shaders, backend coverage and full ownership burden

Compare outputs before comparing speed

Build a frozen corpus that represents normal inputs, shape boundaries, empty or maximum-size inputs, low-contrast or noisy media, language edge cases, and previously failing examples. Run a stable CPU or known-good production path as the reference. Compare the new backend at the tensor or task level and report the first layer where divergence exceeds policy.

def compare_backend(reference, candidate, cases, policy):
    report = []
    for case in cases:
        expected = reference.run(case.input, capture=policy.capture_layers)
        actual = candidate.run(case.input, capture=policy.capture_layers)
        report.append({
            "case": case.id,
            "backend": candidate.actual_backend,
            "fallbacks": candidate.fallback_ops,
            "output": compare_task_metric(expected, actual, policy.task),
            "layers": compare_tensors(expected, actual, policy.tensor),
        })
    return gate(report, policy.thresholds)

Repeat tests because shader compilation, cache warming, dynamic shapes, allocator reuse, and thermal state alter execution. Include cold install, first model load, first inference, warm steady state, application background and foreground transitions, orientation changes, resource pressure, and device sleep/wake. When the application uses the GPU for rendering, test contention rather than benchmarking inference in isolation.

Observability itself needs acceptance criteria. A September LiteRT issue demonstrates the point: per-operation profiling succeeded on Android OpenCL and macOS Metal in the reporter's setup but returned an unimplemented error on the macOS WebGPU path. That does not prove inference is broken. It proves a team cannot assume the same diagnostic depth across backends. Define which missing counters block rollout and which can be replaced with trace markers or application-level measurements.

Benchmark the application boundary, not only a shell binary

Google's benchmark documentation warns that running a binary through adb shell can produce different behavior from a foreground Android application because scheduling and process priority differ. Use the official benchmark tool to isolate runtime behavior, then validate the same workload inside the shipping application. Both views matter.

Report cold compilation separately from warm invocation. For generative workloads, separate prefill from decode and state the prompt length, generated tokens, KV-cache format, sampling, synchronization, and quantization. For media pipelines, report frame latency, dropped frames, upload and readback time, and the competing render workload. Always include device, OS, driver, power mode, temperature, repetitions, and percentile method.

MetricWhy it mattersCommon misleading result
Cold P95Captures shader compile and cache creationOnly warm averages are published
Warm P50 / P95 / P99Shows both typical and tail experienceOne fastest run
GPU delegation percentageExplains transfer and fallback costsGPU requested is reported as GPU used
Peak memoryProtects mobile multitasking and low-memory devicesModel size is used as memory estimate
Energy and thermal decayTests sustained interactive workloadsA five-second benchmark
Task-quality deltaConnects numerical change to product behaviorLatency wins without output checks

Failure modes a unified API does not remove

FailureWhat it looks likeControl
Silent fallbackCorrect result, little or no speed improvementRecord partition and actual backend; gate unexpected CPU ops
Backend numerical driftRare task regressions on one GPU familyFrozen edge corpus, layer captures, predeclared tolerances
Cold-start regressionFirst interaction stalls while shaders compileMeasure clean-cache startup; precompile or warm only with product approval
Driver fragmentationCrash or wrong output on a narrow OS/device groupCohort rollout, denylist, server kill switch, legacy fallback
Memory-path regressionExtra copies or allocations erase compute gainsTrace buffer ownership, reuse, upload/readback, and peak memory
Render contentionInference improves while UI frames dropBenchmark inside the complete graphics workload
Missing diagnosticsProduction slowdown cannot be attributedRequire backend IDs and app spans; qualify profiler gaps
Benchmark scope errorMobile vendor result becomes a desktop universal claimPublish the exact workload and cohort; retest target hardware

Promote per model and device cohort

Stage 1: laboratory parity. Run reference and candidate paths on a compact device matrix. Block on output drift, unexpected fallback, crashes, or missing required telemetry. Keep the model unchanged.

Stage 2: employee or synthetic shadow. Execute the candidate without serving its result where product constraints allow. Compare output, latency, memory, and energy against the active path. Include cold starts and lifecycle events.

Stage 3: small cohort. Enable one model on a narrowly defined device and OS group. Monitor crash-free sessions, initialization failures, fallback ratio, latency percentiles, memory pressure, thermal decay, task quality, and user-visible errors. Keep a remote kill switch.

Stage 4: controlled expansion. Add device families separately. A passing Adreno cohort does not approve Mali, Apple, Intel, or browser WebGPU. Reopen the gate for a model, quantization, runtime, driver, or material application-pipeline change.

ML Drift migration checklist

Pin the stackModel, tokenizer, quantization, LiteRT, ML Drift, app build, driver, and OS.
Observe the pathActual backend, compiled partition, fallback ops, shader cache, and memory transfers.
Freeze output policyReference, edge corpus, task metrics, tensor captures, and tolerances chosen in advance.
Measure real conditionsCold and warm percentiles, app priority, graphics contention, memory, energy, and thermal decay.
Qualify diagnosticsKnow which backends provide per-op profiles and which require app-level substitutes.
Release narrowlyGate each model-device cohort, retain the legacy path, and rehearse the kill switch.

FAQ

What is ML Drift?

ML Drift is Google's Apache-2.0 GPU compute engine for on-device machine learning. It powers the newer LiteRT GPU path and also exposes standalone C++ APIs for teams building custom compute graphs and shaders.

Is it a drop-in replacement for the TFLite GPU delegate?

There is a low-change LiteRT migration route that retains the Interpreter-style API, but the modern CompiledModel route adds different buffer, accelerator, and asynchronous-execution behavior. Even a package-only change deserves output and performance regression tests.

Will one model be identical on OpenCL, Metal, and WebGPU?

Do not assume so. Precision, accumulation order, kernels, fallback, drivers, and supported features differ. Define task and tensor tolerances, then test each supported device family.

Does the Google benchmark prove ML Drift is faster for my app?

No. It is useful first-party evidence for the tested devices and workloads. Your release decision should use the shipping model, application pipeline, devices, scheduler context, and quality thresholds.

Sources and further reading

All current facts were checked on October 10, 2026. Vendor performance claims are identified as first-party and should be reproduced on target devices.