The following are the release notes for the LiteRT October release.
LiteRT Release Notes: Version 2.3.0
Summary
The LiteRT 2.3.0 release represents a leap in on-device AI efficiency, particularly for Large Language Models (LLMs) and cross-platform GPU/NPU utilization.
API and Platform updates
- Swift Package for iOS & macOS - New type-safe Swift APIs to run LiteRT models with hardware acceleration (developer guide).
- Support Library - Migrated from TensorFlow Lite Support to LiteRT Support (migration guide).
- Android - minSdk raised to 24; GPU (litert-gpu) and vendor-specific NPU libraries now ship as separate Maven packages (NPU sample).
Hardware acceleration
- Standalone Open Source LiteRT ML Drift GPU accelerator
- NPU Scaling - Samsung NPU (Exynos AI LiteCore) with LiteRT.
- Intel NPU/GPU - Weight sharing roughly halves AOT-compiled LLM size.
- MediaTek NPU - Support expanded to all Android S+ versions.
Performance and Memory optimizations
- FlashAttention on Metal - Up to 38% faster decode, up to 2.98x faster prefill, and >4 GB less peak RAM at 32K context.
- Fused LLM kernels - New QKV-RMSNorm-RoPE and SwiGLU GPU kernels for ~12% faster decode.
- GPU memory - Metal memory residency on iOS/macOS avoids swap paging; zero-copy mmap for the program cache lowers peak RAM.
Validation / Testing Infrastructure
- ATS - Coverage expanded to all core single ops and to macOS (Metal and WebGPU).