LiteRT October Release

The following are the release notes for the LiteRT October release.

LiteRT Release Notes: Version 2.3.0

Summary

The LiteRT 2.3.0 release represents a leap in on-device AI efficiency, particularly for Large Language Models (LLMs) and cross-platform GPU/NPU utilization.

API and Platform updates

  • Swift Package for iOS & macOS - New type-safe Swift APIs to run LiteRT models with hardware acceleration (developer guide).
  • Support Library - Migrated from TensorFlow Lite Support to LiteRT Support (migration guide).
  • Android - minSdk raised to 24; GPU (litert-gpu) and vendor-specific NPU libraries now ship as separate Maven packages (NPU sample).

Hardware acceleration

Performance and Memory optimizations

  • FlashAttention on Metal - Up to 38% faster decode, up to 2.98x faster prefill, and >4 GB less peak RAM at 32K context.
  • Fused LLM kernels - New QKV-RMSNorm-RoPE and SwiGLU GPU kernels for ~12% faster decode.
  • GPU memory - Metal memory residency on iOS/macOS avoids swap paging; zero-copy mmap for the program cache lowers peak RAM.

Validation / Testing Infrastructure

  • ATS - Coverage expanded to all core single ops and to macOS (Metal and WebGPU).