The MediaPipe Universal Embedder task combines the features of individual audio, text, and image embedders into a single unified service, while also adding support for video embeddings. It wraps the built-in LiteRT-LM EmbeddingEngine directly, providing real on-device, high-performance multimodal embeddings.
You can use the Universal Embedder to vectorize text, images, and audio samples, or run multimodal embedding extraction on lists of composite contents to capture semantic meaning across modalities.
Get Started
Start using this task by following one of these implementation guides for your target platform. These platform-specific guides walk you through a basic implementation of this task, including recommended configuration options and code examples:
- Android - Code example - Guide
- Python - Guide
- iOS - Code example - Guide
- Web - Code example - Guide
Models
We support the EmbeddingGemma V2 model for use with this task, which is available under the Gemma license.
EmbeddingGemma V2
EmbeddingGemma V2 is a multimodal embedding model capable of projecting text, images, and audio into a single shared vector space. This unified representation is what allows for accurate cross-modal capabilities, such as semantic image search based on natural language concepts.
You can find the model cards for the different modalities from the LiteRT community on Hugging Face:
| Modality | Size | Model Card |
|---|---|---|
| Text | 270M | Hugging Face |
| Text-Vision | 440M | Hugging Face |
| Omnimodal | 740M | Hugging Face |
Task details
This section describes the capabilities, inputs, outputs, and configuration options of this task.
Features
- Unified Multimodal Embedding API - Extract embeddings from text strings, raw or compressed images, and audio data within a single task lifecycle.
- Direct LiteRT-LM Wrapper - Runs completely on-device using a shared built-in EmbeddingEngine for optimal memory footprint and performance.
- Similarity Comparison Utilities - Built-in static utility functions to compute the cosine similarity between raw float arrays or embedding results.
- Semantic Memory Bridge - Integrates directly with
SemanticRetrieverby exposing anEmbeddingProviderimplementation.
| Task inputs | Task outputs |
|---|---|
Accepts inputs of various modalities, either individually or as a mixed list:
|
Outputs an EmbeddingResult containing:
|
Configuration options
This task has the following configuration options:
| Option Name | Description | Value Range | Default Value |
|---|---|---|---|
baseOptions |
Specifies the model asset path or file descriptor, along with backend acceleration options (CPU, GPU, NPU). | BaseOptions object |
Required |
l2Normalize |
Whether to normalize the returned embedding feature vector using L2 normalization. | Boolean |
False |
activationDataType |
Specifies the precision type to use for activations during inference. | ActivationDataType (e.g., FLOAT32, FLOAT16, INT16, INT8) |
FLOAT32 |
cacheDir |
The absolute path to the directory for storing compiled model cache artifacts. | String |
Not set |
maxInputLength |
The maximum input signature sequence length to load from the multisignature model. | Integer |
Not set |
visionTokensPerImage |
The number of vision soft tokens to generate per image. | Integer |
Not set |