ML Drift makes GPU portability a conformance problem
Google's new GPU engine can place one compute graph across OpenCL, OpenGL ES, Metal, and WebGPU. That is a valuable integration layer, not proof that every backend produces the same output, latency, memory use, fallback behavior, or observability. Migrate with a device-level acceptance contract.
LiteRT migrationBackend conformanceApp-level benchmarksApache 2.0Sources checked Oct 10
The release changes the default GPU path, not the laws of heterogeneous hardware
Google open-sourced ML Drift on October 8 as the GPU compute engine behind LiteRT acceleration and as a standalone C++ library. The project targets Android, iOS and macOS, web, Windows, and Linux through several lower-level APIs. Google also said the legacy TensorFlow Lite GPU delegate will receive no new features and encouraged applications to move toward ML Drift. That makes this an operational migration, not merely another repository launch.
The interesting abstraction is a unified kernel language and compute graph that can be lowered to OpenCL, OpenGL ES, Metal, or WebGPU. Application developers get a higher-level LiteRT path with graph partitioning, tensor virtualization, zero-copy buffers, and asynchronous execution. Runtime engineers can use the lower-level graph and shader interfaces directly. Both paths aim to reduce the amount of platform-specific inference code a team owns.
The current developer signal is real but bounded. At our October 10 check, the official repository had 70 stars and six open issues. A focused r/LocalLLaMA thread had 103 votes and 13 comments. The most useful comments did not celebrate “GPU acceleration” in the abstract. They asked whether Google's order-of-magnitude claim transfers beyond the published mobile workloads, how the engine compares with llama.cpp on desktop GPUs, why Android emphasizes OpenCL, and which runtime problem justifies another engine.
Those questions point to the right thesis: source-level portability is not runtime equivalence. A common API can dispatch different kernels, accumulations, memory layouts, caches, drivers, and fallbacks. An Android device may choose OpenCL, another may fall back to OpenGL ES, macOS may use Metal, and a Windows build may reach WebGPU through Direct3D. The application still owns proof that each supported cohort behaves acceptably.
ML Drift can make one implementation portable. Only a conformance matrix can make the release decision portable.
Follow the graph from model to the backend that actually ran
A migration begins with a model artifact, operator set, shapes, quantization policy, and runtime version. LiteRT compiles that model for requested accelerators. Supported subgraphs move to the GPU path; unsupported work may remain on CPU or use another accelerator. ML Drift then lowers GPU work through its kernel and backend layers. Buffers, synchronization, shader compilation, cache state, and driver behavior determine what happens on the device.
The chain contains two different compatibility promises. API compatibility means the application can construct and invoke a model. Behavioral compatibility means the deployed path still meets output, latency, memory, energy, crash, and observability requirements. The first is documented by packages and interfaces. The second belongs in your tests.
OS and device family, Metal feature set, memory pressure, profiler availability
Web
WebGPU in a supported browser
Browser build, adapter, feature flags, shader compilation and cache state
Linux
OpenCL or WebGPU over Vulkan
GPU and driver, ICD or Vulkan stack, selected path, headless/app context
Windows
WebGPU over Direct3D
GPU, driver, DirectX Shader Compiler files, selected adapter, cache state
Do not trust a requested accelerator field as evidence that the whole graph ran there. Capture the compiled partition and actual backend at runtime. A model that delegates 96% of arithmetic can still lose its expected latency benefit if a small unsupported operator forces repeated device-to-host transfers.
Version a conformance contract per model and device family
Freeze the artifacts before measuring. Record the model and tokenizer hashes, LiteRT and ML Drift revisions, build flags, precision, delegate options, expected backend, input corpus, reference outputs, tolerance policy, benchmark method, and acceptance thresholds. If any of these changes, the release needs a new result rather than an inherited green badge.
Choose tolerances before looking at the new output. The policy should match the application's risk. A camera effect may tolerate small pixel drift while a biometric or safety-related model needs a stricter task-level validation. For language models, exact-token agreement can be too strict for benign floating-point variation yet too weak for hidden logit drift. Retain tensor-level or metric-level comparisons for selected layers and adversarial cases.
Make fallback part of the contract. “Correct output” from a CPU fallback is not a successful GPU migration if the product promise is frame-time or battery improvement. Conversely, a fully delegated graph that changes the result beyond tolerance is not acceptable because it looks fast.
Choose the narrow migration path before the ambitious one
Google documents two LiteRT routes. Existing TensorFlow Lite applications can first change packages while retaining the Interpreter-style API. New work can adopt the v2 CompiledModel API for explicit accelerator selection, zero-copy tensor buffers, NPU support, and asynchronous execution. The second path offers more capability but changes more of the application's runtime contract.
// Android: pin a tested LiteRT release.
dependencies {
implementation("com.google.ai.edge.litert:litert:2.2.0")
}
// Pseudocode: request the GPU, then record what compiled.
val options = CompiledModel.Options(
accelerators = listOf(Accelerator.GPU)
)
val compiled = CompiledModel.create(modelBytes, options)
telemetry.record(compiled.backendInfo())
telemetry.record(compiled.partitionSummary())
val result = compiled.run(reusableInputBuffers)
Do not combine package migration, API migration, quantization change, model update, and new asynchronous scheduling in one release. Start with one pinned model and a representative cohort. Keep the legacy path available behind a server-controlled flag. After output and fallback gates pass, introduce buffer reuse and asynchronous execution with separate UI-frame and concurrency tests.
Path
Good fit
Main review
Interpreter-compatible package change
Lowest-risk first step for an existing TFLite app
Packaging, delegate selection, output and performance regression
CompiledModel synchronous
New runtime with explicit accelerator and buffer control
Partitioning, tensor ownership, fallback and memory behavior
CompiledModel asynchronous / zero-copy
Camera, media, and interactive pipelines
Synchronization, lifetime, contention, frame pacing and cancellation
Standalone ML Drift C++
Custom graphics or inference runtime
Graph construction, shaders, backend coverage and full ownership burden
Compare outputs before comparing speed
Build a frozen corpus that represents normal inputs, shape boundaries, empty or maximum-size inputs, low-contrast or noisy media, language edge cases, and previously failing examples. Run a stable CPU or known-good production path as the reference. Compare the new backend at the tensor or task level and report the first layer where divergence exceeds policy.
def compare_backend(reference, candidate, cases, policy):
report = []
for case in cases:
expected = reference.run(case.input, capture=policy.capture_layers)
actual = candidate.run(case.input, capture=policy.capture_layers)
report.append({
"case": case.id,
"backend": candidate.actual_backend,
"fallbacks": candidate.fallback_ops,
"output": compare_task_metric(expected, actual, policy.task),
"layers": compare_tensors(expected, actual, policy.tensor),
})
return gate(report, policy.thresholds)
Repeat tests because shader compilation, cache warming, dynamic shapes, allocator reuse, and thermal state alter execution. Include cold install, first model load, first inference, warm steady state, application background and foreground transitions, orientation changes, resource pressure, and device sleep/wake. When the application uses the GPU for rendering, test contention rather than benchmarking inference in isolation.
Observability itself needs acceptance criteria. A September LiteRT issue demonstrates the point: per-operation profiling succeeded on Android OpenCL and macOS Metal in the reporter's setup but returned an unimplemented error on the macOS WebGPU path. That does not prove inference is broken. It proves a team cannot assume the same diagnostic depth across backends. Define which missing counters block rollout and which can be replaced with trace markers or application-level measurements.
Benchmark the application boundary, not only a shell binary
Google's benchmark documentation warns that running a binary through adb shell can produce different behavior from a foreground Android application because scheduling and process priority differ. Use the official benchmark tool to isolate runtime behavior, then validate the same workload inside the shipping application. Both views matter.
Report cold compilation separately from warm invocation. For generative workloads, separate prefill from decode and state the prompt length, generated tokens, KV-cache format, sampling, synchronization, and quantization. For media pipelines, report frame latency, dropped frames, upload and readback time, and the competing render workload. Always include device, OS, driver, power mode, temperature, repetitions, and percentile method.
Metric
Why it matters
Common misleading result
Cold P95
Captures shader compile and cache creation
Only warm averages are published
Warm P50 / P95 / P99
Shows both typical and tail experience
One fastest run
GPU delegation percentage
Explains transfer and fallback costs
GPU requested is reported as GPU used
Peak memory
Protects mobile multitasking and low-memory devices
Model size is used as memory estimate
Energy and thermal decay
Tests sustained interactive workloads
A five-second benchmark
Task-quality delta
Connects numerical change to product behavior
Latency wins without output checks
Failure modes a unified API does not remove
Failure
What it looks like
Control
Silent fallback
Correct result, little or no speed improvement
Record partition and actual backend; gate unexpected CPU ops
Measure clean-cache startup; precompile or warm only with product approval
Driver fragmentation
Crash or wrong output on a narrow OS/device group
Cohort rollout, denylist, server kill switch, legacy fallback
Memory-path regression
Extra copies or allocations erase compute gains
Trace buffer ownership, reuse, upload/readback, and peak memory
Render contention
Inference improves while UI frames drop
Benchmark inside the complete graphics workload
Missing diagnostics
Production slowdown cannot be attributed
Require backend IDs and app spans; qualify profiler gaps
Benchmark scope error
Mobile vendor result becomes a desktop universal claim
Publish the exact workload and cohort; retest target hardware
Promote per model and device cohort
Stage 1: laboratory parity. Run reference and candidate paths on a compact device matrix. Block on output drift, unexpected fallback, crashes, or missing required telemetry. Keep the model unchanged.
Stage 2: employee or synthetic shadow. Execute the candidate without serving its result where product constraints allow. Compare output, latency, memory, and energy against the active path. Include cold starts and lifecycle events.
Stage 3: small cohort. Enable one model on a narrowly defined device and OS group. Monitor crash-free sessions, initialization failures, fallback ratio, latency percentiles, memory pressure, thermal decay, task quality, and user-visible errors. Keep a remote kill switch.
Stage 4: controlled expansion. Add device families separately. A passing Adreno cohort does not approve Mali, Apple, Intel, or browser WebGPU. Reopen the gate for a model, quantization, runtime, driver, or material application-pipeline change.
ML Drift migration checklist
Pin the stackModel, tokenizer, quantization, LiteRT, ML Drift, app build, driver, and OS.
Observe the pathActual backend, compiled partition, fallback ops, shader cache, and memory transfers.
Freeze output policyReference, edge corpus, task metrics, tensor captures, and tolerances chosen in advance.
Measure real conditionsCold and warm percentiles, app priority, graphics contention, memory, energy, and thermal decay.
Qualify diagnosticsKnow which backends provide per-op profiles and which require app-level substitutes.
Release narrowlyGate each model-device cohort, retain the legacy path, and rehearse the kill switch.
FAQ
What is ML Drift?
ML Drift is Google's Apache-2.0 GPU compute engine for on-device machine learning. It powers the newer LiteRT GPU path and also exposes standalone C++ APIs for teams building custom compute graphs and shaders.
Is it a drop-in replacement for the TFLite GPU delegate?
There is a low-change LiteRT migration route that retains the Interpreter-style API, but the modern CompiledModel route adds different buffer, accelerator, and asynchronous-execution behavior. Even a package-only change deserves output and performance regression tests.
Will one model be identical on OpenCL, Metal, and WebGPU?
Do not assume so. Precision, accumulation order, kernels, fallback, drivers, and supported features differ. Define task and tensor tolerances, then test each supported device family.
Does the Google benchmark prove ML Drift is faster for my app?
No. It is useful first-party evidence for the tested devices and workloads. Your release decision should use the shipping model, application pipeline, devices, scheduler context, and quality thresholds.
Sources and further reading
All current facts were checked on October 10, 2026. Vendor performance claims are identified as first-party and should be reproduced on target devices.