LiteRT: High-Performance On-Device Machine Learning Framework  |  Google AI Edge  |  Google for Developers Skip to main content / English Deutsch Español Français Indonesia Português – Brasil Русский 中文 – 简体 日本語 한국어 Sign in Google AI Edge LiteRT Android Desktop Web LiteRT-LM MediaPipe MediaPipe Solutions MediaPipe Framework Model Explorer AI Edge Portal Google Tensor SDK Apps AI Edge Gallery AI Edge Eloquent MediaPipe Demos API Reference AI Edge LiteRT More LiteRT-LM MediaPipe More Model Explorer AI Edge Portal Google Tensor SDK Apps More API Reference Home Overview Migrating from TensorFlow Lite LiteRT CLI Overview Installation Common Commands Troubleshooting & Resources GenAI Deployments Overview Convert PyTorch GenAI models Run LLMs using LiteRT-LM Model Conversion & Optimization Overview Convert PyTorch models Overview Convert PyTorch GenAI models Convert TensorFlow models Overview Use pre-trained models Convert TensorFlow models Add Signatures Model compatibility Overview Select operators Select operators Allowlist Advanced Fused operators Operator versions RNN models Optimize models Overview Post-training quantization Post-training dynamic range quantization Post-training integer quantization Post-training float16 quantization Post-training integer quantization with int16 activations Quantization specification Inspecting quantization errors Add model metadata Overview Metadata Writer API Design and build models Overview Performance best practices On-device training Convert JAX models Overview Inference & Hardware Acceleration Overview GPU acceleration NPU acceleration Overview Run LLMs using LiteRT-LM Google Tensor Intel Qualcomm MediaTek Samsung Creating a new accelerator Implementing a Compiler Plugin Implementing the Dispatch API Accelerator Test Suite (ATS) Benchmark & Profiling Benchmark CompiledModel API Benchmark Interpreter API Create custom operators Custom ops of CompiledModel API Custom ops of Interpreter API Run on Android Overview Run with CompiledModel API (accelerated on GPU/NPU) Kotlin API C++ API Use prebuilt C++ library Hardware acceleration GPU acceleration NPU acceleration Google Tensor MediaTek NPU Qualcomm NPU Intel NPU Run with Interpreter API Google Play services runtime Overview Java API C API Hardware acceleration Acceleration service GPU with Interpreter API GPU with C/C++ API NPU delegates Overview Qualcomm NPUs for Mobile AI Development Development tools Models with metadata Overview Generate model interfaces Customize data input and output Run on iOS / macOS Run with CompiledModel API (accelerated on GPU) C++ API Use prebuilt C++ library Run with Interpreter API Swift API Core ML delegate GPU delegate Run on Web with LiteRT.js Run with CompiledModel API (accelerated on GPU) Overview JavaScript API Run on Desktop (Linux, Windows) Run with CompiledModel API (accelerated on GPU) C++ API Use prebuilt C++ library Python API Run on Embedded & IoT Run with CompiledModel API C++ API Use prebuilt C++ library Run with Interpreter API Overview Get started Linux-based devices with Python Understand the C++ library Build and convert models Build from Source Build Compiled Model API Build with CMake Build Interpreter API Build for Android Build for iOS Build for Linux-based IoT Build with CMake Overview Cross compilation for ARM Build Python Wheel Reduce binary size Libraries and tools Task Library Overview ImageClassifier ObjectDetector ImageSegmenter ImageEmbedder ImageSearcher NLClassifier BertNLClassifier BertQuestionAnswerer TextEmbedder TextSearcher AudioClassifier Customized API Model Maker Overview Images & video Image classification Object detection Text Text classification BERT question & answer Text search Audio Audio classification Speech recognition Android Desktop Web MediaPipe Solutions MediaPipe Framework AI Edge Gallery AI Edge Eloquent MediaPipe Demos Introducing Google AI Edge Portal: Benchmark Edge AI at scale. Sign-up to request access during private preview. Home Products Google AI Edge LiteRT Send feedback Stay organized with collections Save and categorize content based on your preferences. LiteRT is Google's on-device framework for high-performance ML & GenAI deployment on edge platforms. Efficient conversion, runtime, and optimization for on-device machine learning. Built on the battle-tested foundation of TensorFlow Lite LiteRT isn't just new; it's the next generation of the world's most widely deployed machine learning runtime. It powers the apps you use every day, delivering low latency and high privacy on billions of devices. Trusted by the most critical Google apps 100K+ applications, billions of global users LiteRT Highlights Cross Platform Ready Unleash GenAI Simplified hardware acceleration Multi-framework support Deploy via LiteRT Streamline your deep learning workflow from training to on-device deployment. 1.Obtain a model Use .tflite pre-trained models or convert PyTorch, JAX or TensorFlow models to .tflite. Explore models Learn about conversion 2.Optimize Use the LiteRT optimization toolkit to quantize your models post-training. Explore optimization 3.Run Deploy your model with LiteRT and pick the optimal accelerator for your app. View deployment targets Choose Your Development Path Use LiteRT to deploy AI anywhere—from high-performance mobile apps to resource-constrained IoT devices. sync Existing TFLite User Transitioning to LiteRT to leverage enhanced performance and unified APIs across platforms (Android, Desktop, Web). camera BYOM : Bring your own Models Have a PyTorch model, looking to implement on-device vision or audio experiences. auto_awesome Deploying Generative AI Models Creating sophisticated on-device chatbots using optimized open-weight GenAI models like Gemma or another open-weight model with LiteRT-LM. developer_board [Advanced] Model Expert Authoring custom models or performing deep hardware-specific CPU/GPU/NPU optimizations for peak performance. Samples, models, and demo See LiteRT sample app on GitHub Complete, end-to-end sample apps. See sample apps See genAI models Pre-trained, out-of-the-box Gen AI models. Go to HuggingFace See demos - Google AI Edge Gallery App A gallery that showcases on-device ML/GenAI use cases using LiteRT. Open Play Store Blogs and Announcements Stay up to date with the latest announcements, technical deep dives, and performance benchmarks from the LiteRT team. Explore more blogs LiteRT.js, Google's high performance Web AI Inference LiteRT.js, Google's high-performance web AI inference library for running models using WebAssembly, WebGPU, and WebNN. LiteRT unlocks Core Ultra NPU performance for AI PC Learn how Google and Intel integrate LiteRT with OpenVINO to offload AI inference to Intel NPUs for high-performance AI PCs. Edge AI from the Trenches: A practical guide to LiteRT Community insights on untangling LiteRT vs LiteRT-LM and deploying real-world LLMs on edge devices. Accelerating on-device AI: A look at Arm and Google AI Edge optimization Arm Scalable Matrix Extension 2 (SME2) and Google AI Edge accelerate on-device generative AI using LiteRT and XNNPACK. Building real-world on-device AI with LiteRT and NPU Learn how industry leaders build real-world, high-performance on-device AI applications using LiteRT and NPUs. Bring state-of-the-art agentic skills to the edge with Gemma 4 Deploy agentic and multi-step planning capabilities entirely on-device with the new Gemma 4 family and LiteRT. LiteRT: The universal framework for on-device AI Google's unified on-device ML framework, evolving from TFLite for high-performance deployment. MediaTek NPU and LiteRT: Powering the next generation of on-device AI Expanding NPU acceleration support to MediaTek chipsets for high-efficiency AI. Unlocking Peak Performance on Qualcomm NPU with LiteRT Unlocking breakthrough performance for generative AI on Qualcomm Neural Processing Units. LiteRT: Maximum Performance, Simplified Introducing the CompiledModel API for automated hardware selection and async execution. On-device GenAI in Chrome, Chromebook Plus, and Pixel Watch with LiteRT-LM Deploy language models on wearables and browser-based platforms using LiteRT-LM. Google AI Edge small language models, multimodality, and function calling Latest insights on RAG, multimodality, and function calling for edge language models Join the Community LiteRT GitHub Community Contribute directly to the project and collaborate with core developers. Hugging Face Hub Access optimized open-weight models on the Hugging Face Hub. lightbulb LiteRT Collaboration and Feature Intake Submit feature requests and Collaboration with LiteRT team. Start Your LiteRT Journey Ready to take your on-device ML to the next level? Explore the documentation and start building today. Explore the Docs Except as otherwise noted, the content of this page is licensed under the Creative Commons Attribution 4.0 License, and code samples are licensed under the Apache 2.0 License. For details, see the Google Developers Site Policies. Java is a registered trademark of Oracle and/or its affiliates. Last updated 2026-07-17 UTC. Need to tell us more? [[["Easy to understand","easyToUnderstand","thumb-up"],["Solved my problem","solvedMyProblem","thumb-up"],["Other","otherUp","thumb-up"]],[["Missing the information I need","missingTheInformationINeed","thumb-down"],["Too complicated / too many steps","tooComplicatedTooManySteps","thumb-down"],["Out of date","outOfDate","thumb-down"],["Samples / code issue","samplesCodeIssue","thumb-down"],["Other","otherDown","thumb-down"]],["Last updated 2026-07-17 UTC."],[],[]] Connect Blog Bluesky Instagram LinkedIn X (Twitter) YouTube Programs Google Developer Program Google Developer Groups Google Developer Experts Accelerators Google Cloud & NVIDIA Developer consoles Google API Console Google Cloud Platform Console Google Play Console Firebase Console Actions on Google Console Cast SDK Developer Console Chrome Web Store Dashboard Google Home Developer Console Android Chrome Firebase Google Cloud Platform Google AI All products Terms Privacy Manage cookies English Deutsch Español Français Indonesia Português – Brasil Русский 中文 – 简体 日本語 한국어