Universal embedder guide

The MediaPipe Universal Embedder task combines the features of individual audio, text, and image embedders into a single unified service, while also adding support for video embeddings. It wraps the built-in LiteRT-LM EmbeddingEngine directly, providing real on-device, high-performance multimodal embeddings.

You can use the Universal Embedder to vectorize text, images, and audio samples, or run multimodal embedding extraction on lists of composite contents to capture semantic meaning across modalities.

Get Started

Start using this task by following one of these implementation guides for your target platform. These platform-specific guides walk you through a basic implementation of this task, including recommended configuration options and code examples:

Models

We support the EmbeddingGemma V2 model for use with this task, which is available under the Gemma license.

EmbeddingGemma V2

EmbeddingGemma V2 is a multimodal embedding model capable of projecting text, images, and audio into a single shared vector space. This unified representation is what allows for accurate cross-modal capabilities, such as semantic image search based on natural language concepts.

You can find the model cards for the different modalities from the LiteRT community on Hugging Face:

Modality Size Model Card
Text 270M Hugging Face
Text-Vision 440M Hugging Face
Omnimodal 740M Hugging Face

Task details

This section describes the capabilities, inputs, outputs, and configuration options of this task.

Features

  • Unified Multimodal Embedding API - Extract embeddings from text strings, raw or compressed images, and audio data within a single task lifecycle.
  • Direct LiteRT-LM Wrapper - Runs completely on-device using a shared built-in EmbeddingEngine for optimal memory footprint and performance.
  • Similarity Comparison Utilities - Built-in static utility functions to compute the cosine similarity between raw float arrays or embedding results.
  • Semantic Memory Bridge - Integrates directly with SemanticRetriever by exposing an EmbeddingProvider implementation.
Task inputs Task outputs
Accepts inputs of various modalities, either individually or as a mixed list:
  • Text (String)
  • Image (MPImage / UIImage / Compressed bytes)
  • Audio (AudioData / Float samples)
  • Composite Content (A mixed list of any of the preceding objects)
Outputs an EmbeddingResult containing:
  • Embedding: The high-dimensional feature vector, either as raw floating-point numbers or scalar-quantized.

Configuration options

This task has the following configuration options:

Option Name Description Value Range Default Value
baseOptions Specifies the model asset path or file descriptor, along with backend acceleration options (CPU, GPU, NPU). BaseOptions object Required
l2Normalize Whether to normalize the returned embedding feature vector using L2 normalization. Boolean False
activationDataType Specifies the precision type to use for activations during inference. ActivationDataType (e.g., FLOAT32, FLOAT16, INT16, INT8) FLOAT32
cacheDir The absolute path to the directory for storing compiled model cache artifacts. String Not set
maxInputLength The maximum input signature sequence length to load from the multisignature model. Integer Not set
visionTokensPerImage The number of vision soft tokens to generate per image. Integer Not set