Embedding Models

LiteRT-LM provides lightweight, on-device vector embedding generation designed for edge AI scenarios including semantic search, Retrieval-Augmented Generation (RAG), text classification, and intent clustering. By running directly on client devices, LiteRT-LM eliminates network latency and preserves user privacy.

Key Capabilities

  • High-density vector representations tailored for mobile and edge runtimes.
  • Low-latency CPU, GPU, and NPU acceleration options.
  • Unified API paradigm across major desktop, mobile, and web ecosystems.
  • Built-in support for cosine similarity and vector distance evaluation.

Learn more about EmbeddingGemma 2 support in our blog post. Visit LiteRT Community on Hugging Face for models optimized for on-device execution.

You can experience EmbeddingGemma 2 in action directly in your browser with the Multimodal Search Web Demo, try Instant Media Search and Video Moments Finder on your own device with the Google AI Edge Gallery App on Android and iOS, and see personal knowledge base retrieval in action with AI Edge Foresight.

Task Instruction Prefixes

EmbeddingGemma 2 uses asymmetric task-specific instruction prefixes to optimize embeddings for target semantic tasks. Applying the correct task prefix before the prompt ensures optimal cosine similarity calculation.

Task Type Instruction Prefix Primary Use Case
Document Indexing task: search result | text: Ingesting documents, web pages, or notes into a local vector database.
Search Query task: search query | text: User search queries matched against indexed documents.
Classification task: classification | text: Zero-shot intent matching or document category tagging.
Clustering task: clustering | text: Grouping similar texts, chat histories, or notes.
Symmetric Similarity task: sentence similarity | text: Pairwise semantic comparison between two equivalent text snippets.

Always trim leading and trailing whitespace from input strings before prepending the instruction prefix.

Cross-Platform Code Examples With EmbeddingGemma 2

Use the following tabs to see how to generate embeddings across platforms using the LiteRT-LM SDKs:

Python

from litert_lm import Backend, Content, EmbeddingEngine

gpu = Backend.GPU()
with EmbeddingEngine(
    "embedding-gemma-v2.litertlm",
    backend=gpu,
    vision_backend=gpu,
    audio_backend=gpu,
) as engine:
  inputs = [
      "A photo of a red apple",
      Content.ImageFile("apple.png"),
      Content.AudioFile("audio_sample.wav"),
  ]

  for item in inputs:
    emb = engine.compute_embedding(item).embedding
    print(
        f"✓ Success! Output dimension: {len(emb)} | First 5 values: {emb[:5]}..."
    )

Kotlin

import com.google.ai.edge.litertlm.*
import java.io.File

fun main() {
  val config = EmbeddingEngineConfig(
      modelPath = "embedding-gemma-v2.litertlm",
      backend = Backend.GPU(),
      visionBackend = Backend.GPU(),
      audioBackend = Backend.GPU(),
  )

  EmbeddingEngine(config).use { engine ->
    engine.initialize()

    val inputs = listOf(
        InputData.Text("A photo of a red apple"),
        InputData.Image(File("apple.png").readBytes()),
        InputData.Audio(File("audio_sample.wav").readBytes())
    )

    for (item in inputs) {
      val emb = engine.computeEmbedding(listOf(item)).embedding
      println("✓ Success! Output dimension: ${emb.size} | First 5 values: ${emb.take(5)}...")
    }
  }
}

Swift

import LiteRTLM

func main() async throws {
  let engine = EmbeddingEngine(config: .init(
    modelPath: "embedding-gemma-v2.litertlm",
    backend: .gpu, visionBackend: .gpu, audioBackend: .gpu
  ))
  try await engine.initialize()
  defer { Task { await engine.close() } }

  for item in [Content.text("Red apple"), .imageFile("apple.png"), .audioFile("audio.wav")] {
    let resp = try await engine.computeEmbedding(contents: [item])
    print("Vector dimension: \(resp.embedding.count) | First 5 values: \(resp.embedding.prefix(5))...")
  }
}

JavaScript

import { EmbeddingEngine } from '@litert-lm/core';

async function initEmbeddings() {
  const engine = await EmbeddingEngine.create({
    model: 'https://huggingface.co/litert-community/embeddinggemma-2-740m-litert-lm/resolve/main/embeddinggemma-2-740m.litertlm'
  });

  const queryText = 'task: search query | text: Top places to visit in Kyoto';
  const { embedding } = await engine.computeEmbedding(queryText, {normalize: true});

  console.log(`Generated embedding with dimension: ${embedding.length}`);
  return engine;
}

For a complete browser-based multimodal search experience, try out our web demo in Chrome.

CLI

You can also serve embedding models locally using the LiteRT-LM CLI's OpenAI-Compatible Server and the /v1/embeddings endpoint:

Linux/macOS

# Install litert-lm
pip install litert-lm==0.18.0

# Open terminal and launch the server
litert-lm serve --host 127.0.0.1 --port 9379

# Text Embedding:
curl -X POST http://127.0.0.1:9379/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "model": "embedding-gemma-v2.litertlm,gpu",
    "input": "Hello world from the LiteRT server"
  }'

Windows

# Install litert-lm
pip install litert-lm==0.18.0

# Open terminal and launch the server
litert-lm serve --host 127.0.0.1 --port 9379

# Text Embedding:
Invoke-RestMethod -Uri "http://127.0.0.1:9379/v1/embeddings" `
  -Method Post `
  -ContentType "application/json" `
  -Body '{"model": "embedding-gemma-v2.litertlm,gpu", "input": "Hello world from the LiteRT server"}'