Semantic retriever guide

The MediaPipe Semantic Retriever task manages vector embedding generation, on-device vector storage, and query retrieval for text, image, and audio content. You can use this task to build advanced search systems, personalized content retrieval, or on-device knowledge retrieval pipelines (such as RAG) that operate completely on-device.

Get Started

Start using this task by following one of these implementation guides for your target platform. These platform-specific guides walk you through a basic implementation of this task, including recommended configuration options and code examples:

Models

We support the EmbeddingGemma V2 model for use with this task, which is available under the Gemma license.

EmbeddingGemma V2

EmbeddingGemma V2 is a multimodal embedding model capable of projecting text, images, and audio into a single shared vector space. This unified representation is what allows for accurate cross-modal capabilities, such as semantic image search based on natural language concepts.

You can find the model cards for the different modalities from the LiteRT community on Hugging Face:

Modality Size Model Card
Text 270M Hugging Face
Text-Vision 440M Hugging Face
Omnimodal 740M Hugging Face

Task details

This section describes the capabilities, inputs, outputs, and configuration options of this task.

Features

  • On-Device Storage and Search - Leverages a built-in vector database (such as SQLite) to store, query, and manage embeddings locally on the device without cloud round-trips.
  • Multimodal Ingestion - Supports inserting diverse content types, including text documents, local images, and local audio files, converting them into embedded retrieval records.
  • Automated Text Chunking - Divides long text documents into smaller chunks automatically using character-based boundaries, ensuring optimal embedding sequence lengths and maintaining parent-child relationships.
  • Flexible Retrieval - Queries stored records using either a query string or multimodal query parts to retrieve the topK most semantically similar parent records or chunks.
Task inputs Task outputs
Accepts the following inputs for ingestion or search:
  • Text Documents (String)
  • Local Images (Uri / File path)
  • Local Audio (Uri / File path)
  • Multimodal Parts (A list of parts combining text, image, and audio)
Outputs a sorted list of retrieval results containing:
  • Record ID: The unique identifier of the matching record.
  • Content: The original matching content or document chunks.
  • Metadata: Key-value string properties associated with the record.
  • Score: The similarity score (e.g. cosine similarity) compared to the query.

Configuration options

This task has the following configuration options:

Option Name Description Value Range
vectorStore The vector store instance used to persist embeddings, such as an on-device SQLite database or an in-memory database. VectorStore object
providers A list of embedding providers (such as a UniversalEmbedder instance) used to generate high-dimensional vectors for text, images, or audio. List<EmbeddingProvider>
textChunker An optional chunker implementation used to partition long text files. Defaults to character-based chunking with size 512 and overlap 100. TextChunker instance