The MediaPipe Semantic Retriever task manages vector embedding generation, on-device vector storage, and query retrieval for text, image, and audio content. You can use this task to build advanced search systems, personalized content retrieval, or on-device knowledge retrieval pipelines (such as RAG) that operate completely on-device.
Get Started
Start using this task by following one of these implementation guides for your target platform. These platform-specific guides walk you through a basic implementation of this task, including recommended configuration options and code examples:
- Android - Code example - Guide
- Python - Guide
- iOS - Code example - Guide
- Web - Code example - Guide
Models
We support the EmbeddingGemma V2 model for use with this task, which is available under the Gemma license.
EmbeddingGemma V2
EmbeddingGemma V2 is a multimodal embedding model capable of projecting text, images, and audio into a single shared vector space. This unified representation is what allows for accurate cross-modal capabilities, such as semantic image search based on natural language concepts.
You can find the model cards for the different modalities from the LiteRT community on Hugging Face:
| Modality | Size | Model Card |
|---|---|---|
| Text | 270M | Hugging Face |
| Text-Vision | 440M | Hugging Face |
| Omnimodal | 740M | Hugging Face |
Task details
This section describes the capabilities, inputs, outputs, and configuration options of this task.
Features
- On-Device Storage and Search - Leverages a built-in vector database (such as SQLite) to store, query, and manage embeddings locally on the device without cloud round-trips.
- Multimodal Ingestion - Supports inserting diverse content types, including text documents, local images, and local audio files, converting them into embedded retrieval records.
- Automated Text Chunking - Divides long text documents into smaller chunks automatically using character-based boundaries, ensuring optimal embedding sequence lengths and maintaining parent-child relationships.
- Flexible Retrieval - Queries stored records using either a query string or multimodal query parts to retrieve the topK most semantically similar parent records or chunks.
| Task inputs | Task outputs |
|---|---|
Accepts the following inputs for ingestion or search:
|
Outputs a sorted list of retrieval results containing:
|
Configuration options
This task has the following configuration options:
| Option Name | Description | Value Range |
|---|---|---|
vectorStore |
The vector store instance used to persist embeddings, such as an on-device SQLite database or an in-memory database. | VectorStore object |
providers |
A list of embedding providers (such as a UniversalEmbedder instance) used to generate high-dimensional vectors for text, images, or audio. |
List<EmbeddingProvider> |
textChunker |
An optional chunker implementation used to partition long text files. Defaults to character-based chunking with size 512 and overlap 100. | TextChunker instance |