The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Vulkan can provide a low-level route to GPU capabilities on Android, but it is not Android’s machine-learning runtime. For custom on-device inference, Android’s current documented path is LiteRT with hardware delegates; the available Android guidance does not establish that every LiteRT GPU delegate uses Vulkan internally. Treat Vulkan as part of the GPU platform, and LiteRT as the documented ML runtime.
What Vulkan does—and what it does not do
Android describes Vulkan as a low-overhead, cross-platform API for high-performance 3D graphics. It gives software a way to manage GPU work, with features such as reduced CPU overhead and SPIR-V support. Those are GPU-programming capabilities, not a model-loading or inference runtime.
In a machine-learning application, the runtime handles the model and inference workflow. It may delegate supported operations to specialized hardware. Vulkan is relevant to Android’s broader GPU landscape and to native GPU or graphics/compute implementations, but the cited Android documentation does not say that Vulkan is the universal low-level backend for Android ML delegates.
Which Android ML stack to use today
LiteRT and hardware delegates
Android’s custom-ML guidance identifies LiteRT as its official ML inference runtime. It documents LiteRT delegates distributed through Google Play services for accelerated execution on hardware such as GPUs or NPUs. Android also describes an Acceleration Service API that can help an app select an acceleration configuration at runtime.
#1 Best Overall
A delegate is an acceleration option offered through the ML stack; it is not a guarantee that a particular model will run entirely on a GPU. Availability depends on the device, runtime support, model operations, and the selected configuration. The Android documentation establishes GPU delegates, but does not identify one Vulkan backend that applies across devices.
NNAPI and migration
NNAPI was deprecated in Android 15. Android’s NDK guidance recommends migrating performance-critical workloads to alternatives, giving the TensorFlow Lite GPU runtime as an example. The migration guidance describes TensorFlow Lite in Google Play services, with an optional GPU delegate. Deprecation is not the same as immediate removal: it means developers should not treat NNAPI as Android’s preferred path for new performance-critical work.
Rank #2
Does LiteRT use Vulkan for GPU inference?
The Android documentation covered here says that LiteRT delegates can use specialized hardware such as GPUs, but it does not establish that every LiteRT GPU delegate executes through Vulkan. The safe answer is therefore: LiteRT can offer GPU acceleration, while its universal low-level backend is not specified by these sources.
Do not infer a Vulkan-specific speedup from the fact that an app uses a GPU delegate. Results depend on model operators, input sizes, device and driver behavior, runtime and delegate support, and precision. No Vulkan-specific Android ML benchmark is established here; performance claims should be based on measurements for the app’s actual models and target devices.
Check Android and Vulkan compatibility
Android’s Vulkan overview says Vulkan is available starting with Android 7.0 (API level 24). It also says all 64-bit devices running Android 10.0 (API level 29) or later support Vulkan 1.1. These platform statements are useful filters, not proof that a specific ML workload, delegate, or driver will behave as required.
The same overview reports that 85% of active Android devices support Vulkan, but the retrieved page statement does not identify a measurement date. Treat it as Android’s published availability claim, not a current 2026 measurement or a statement about ML acceleration.
| Vulkan profile | Support among active Vulkan-supporting devices | Data date |
|---|---|---|
| AVP 2025 | 80.1% | October 2025 |
| AVP 2022 | 86.5% | October 2025 |
| AVP 2021 | 95.5% | October 2025 |
These Android Vulkan Profile figures describe support for profile feature sets among active Vulkan-supporting devices, not the share of all Android devices and not ML performance. A profile or Vulkan version can narrow compatibility questions, but testing on representative target devices remains necessary.
Plan for device and driver variation
For broad device targeting, verify the actual hardware and software combination rather than relying only on Android version. Android’s native engine guidance recommends considering OpenGL ES support as a fallback for older devices where Vulkan implementations may not run an app reliably. That is graphics compatibility guidance; it does not define an equivalent ML-specific fallback mechanism.
Best Value
- Check whether the target devices expose the Vulkan version or profile your native implementation needs.
- Confirm that the LiteRT runtime and chosen delegate support the device and the model’s operations.
- Measure latency and throughput on representative devices, including any fallback path.
- Validate driver behavior and app reliability across the device set you intend to support.
Weigh on-device benefits against costs
Android’s on-device inference guidance identifies potential benefits such as lower network latency, offline availability, keeping data on-device, and less server-side computation. These are properties of on-device processing, not guarantees provided by Vulkan or by GPU execution alone.
On-device inference can also consume battery, and models may occupy multiple megabytes. When choosing an implementation, assess model size and workload alongside latency, offline requirements, privacy needs, device coverage, and measured power behavior. The cited material provides no Vulkan-specific battery or inference-speed figure.
Quick Recap
A practical decision sequence
- Choose the ML runtime: For custom Android inference, start with LiteRT and evaluate its documented delegates.
- Check acceleration eligibility: Determine whether the target runtime, device, and model support the GPU or other delegate you want; do not assume every operation will be accelerated.
- Use Vulkan where your implementation needs it: Vulkan is a low-level GPU API, especially relevant to native GPU work, not a substitute for the ML runtime.
- Test target devices and fallbacks: Measure the real model and verify driver reliability; consider graphics fallbacks such as OpenGL ES where appropriate for older-device support.
- Review legacy NNAPI integrations: Because NNAPI is deprecated in Android 15, assess Android’s migration guidance for performance-critical workloads.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




