The production-ready orchestration layer to run LLMs with LiteRT, engineered for high-performance, cross-platform execution.

Spotlight

Check out our latest blog to discover how EmbeddingGemma 2 brings multimodal semantic search to the edge, unifying text, images, video, and audio into a single vector space directly on-device with high performance and efficiency.

What's New (v0.18.0)

New Model Release (EmbeddingGemma 2): Shipped multimodal EmbeddingGemma 2 supporting text, vision, and audio embeddings with Matryoshka dimension truncation across Python, Kotlin, Swift, Web (JavaScript), and C++. Check out our Embedding Models documentation and try the web demo!

CLI & Developer Experience: Added fast model imports (litert-lm import) and an OpenAI-compatible /v1/embeddings endpoint (litert-lm serve) for multimodal embedding models.

Model Info API: Added ModelInfo and litert-lm describe for full model introspection—covering metadata, capabilities, and runtime requirements before loading.

NPU & GPU Acceleration: Enabled dynamic on-demand KV cache growth on NPU and attention mask pruning optimizations on GPU.

Why LiteRT-LM?

Deploy LLMs across Android, iOS, Web, and Desktop.
Maximize performance with GPU and NPU acceleration.
Support for popular LLMs as well as multi-modality (Vision, Audio) and Tool Use.

Start building

Python APIs with hardware acceleration on Linux, MacOS, Windows, Android, and Raspberry Pi.
Native Android apps and JVM-based desktop tools.
Native iOS (macOS coming soon) Swift APIs.
JavaScript and TypeScript APIs for browser-based web apps with WebGPU acceleration.
Build cross-platform Flutter apps using the community-maintained flutter_gemma package.
x-platform C++ APIs .
Build .litertlm files from converted LiteRT models.

Join the Community

Contribute to the open-source project, report issues, and see examples.
Download pre-converted models (Gemma, Qwen and more), and join the discussion.

Blogs and Announcements

Experience >2x faster decode speeds on mobile GPUs with zero quality degradation.
Discover how LiteRT-LM supercharges on-device GenAI deployments, unlocking Gemma 4's full potential with Swift, JavaScript, and Flutter APIs.
Deploy Gemma 4 in-app and across a broader range of devices with stellar performance and reach using LiteRT-LM.
Deploy language models on wearables and browser-based platforms using LiteRT-LM at scale.
Explore how to fine-tune FunctionGemma and enable function calling capabilities powered by LiteRT-LM Tool Use APIs.
Latest insights on RAG, multimodality, and function calling for edge language models.