Build once, run everywhere
Cross-platform deployment
Deploy across mobile, web, desktop, and IoT.
Generative AI
Engineered with superior support for LLMs, Diffusion, and more generative models.
Hardware acceleration
CPU, GPU, and NPU hardware acceleration with SOTA performance across leading edge platforms.
Multi-framework conversion
Direct export and quantization from PyTorch, TensorFlow, and JAX into .tflite format.
Built on the battle-tested foundation of TensorFlow Lite
Trusted by prominent Google apps
250K+ applications, billions of global users
What you can build with LiteRT
Local GenAI & assistants
Run open-weight LLMs like Gemma on-device with LiteRT-LM, without a server round trip.
Real-time image segmentation
Run real-time, multi-class image segmentation from the camera on CPU, GPU, or NPU.
Live speech recognition
Transcribe speech on-device with open-weight ASR models like Whisper, Moonshine, and Parakeet on CPU, GPU, or NPU.
On-device vision-language
Run the FastVLM vision-language model on-device with NPU acceleration on Qualcomm devices.
Streaming text-to-speech
Synthesize natural speech fully offline, with playback starting after the first sentence while the rest streams in.
IoT & edge robotics
Deploy lightweight models, such as anomaly detection and sensor fusion, to microcontrollers and embedded devices.
Deploy with LiteRT
Streamline your AI workflow from training to on-device deployment.
1. Convert
2. Quantize
3. Deploy & Accelerate
Choose LiteRT Runtime or LiteRT-LM
Use LiteRT Runtime to run models with CPU, GPU, and NPU acceleration. For LLMs, use LiteRT-LM, which builds on LiteRT to manage conversations, the KV cache, and streaming output.
LiteRT Runtime for ML
High-performance runtime for ML workloads. Tailored for hardware-accelerated execution across CPU, GPU, and NPU backends with advanced features such as zero-copy buffer interoperability and async execution.
LiteRT-LM for LLM
LiteRT-LM is an orchestration layer built on the LiteRT Runtime that powers the LLM execution loop—seamlessly integrating hardware-accelerated inference, multimodal inputs, constrained decoding, and agentic tool calling.
Samples, models, and demos
Explore open-source code samples, pre-trained models, and interactive demos.
End-to-end sample applications
Complete, runnable open-source sample applications across Android, iOS, Web, desktop, and IoT for LiteRT and LiteRT-LM.
Explore open-weight LLMs & models
Pre-converted, ready-to-deploy models on Hugging Face including Gemma 4, Llama 3.2, Qwen 3.5, and MobileNet.
Try on-device Gemma: AI Edge Gallery app
Experience offline LLMs running entirely on your mobile device with LiteRT-LM. Download the open-source sample app on Google Play and GitHub.
Blogs and announcements
Stay up to date with the latest announcements, technical deep dives, and performance benchmarks on the LiteRT news and blogs page.
Join the community
LiteRT GitHub community
Contribute directly to the project and collaborate with core developers.
Hugging Face Hub
Access optimized open-weight models on the Hugging Face Hub.
Collaboration & feature intake
Submit feature requests and collaborate with the LiteRT team.