Semantic retriever guide for Python

The MediaPipe Semantic Retriever task manages on-device vector embedding generation, vector storage, and search retrieval for text, images, and audio files. These instructions show you how to use the Semantic Retriever within Python.

For more information about the capabilities, models, and configuration options of this task, see the Overview.

Setup

This section describes key steps for setting up your development environment and code projects specifically to use Semantic Retriever. For general information on setting up your development environment for using MediaPipe tasks, including platform version requirements, see the Setup guide for Python.

Packages

The Semantic Retriever requires the mediapipe pip package:

$ python -m pip install mediapipe

Imports

Import the following modules to access the Semantic Retriever task functions:

from mediapipe.tasks import python
from mediapipe.tasks.python.retrieval import semantic_retriever
from mediapipe.tasks.python.retrieval import universal_embedder

Create the task

The Semantic Retriever uses a UniversalEmbedder to generate embeddings for the content you ingest and for your queries. First create the embedder, then pass it to a SemanticRetrieverOptions object and create the retriever with create_from_options(). Both objects are context managers, so you can use with statements to release their resources automatically.

# 1. Create the embedder that generates the vector embeddings
base_options = python.BaseOptions(
    model_asset_path='embeddinggemma-2-text-vision-440m.litertlm'
)
embedder_options = universal_embedder.UniversalEmbedderOptions(
    base_options=base_options,
    l2_normalize=True,
)

with universal_embedder.UniversalEmbedder.create_from_options(
    embedder_options
) as embedder:
  # 2. Configure the retriever
  retriever_options = semantic_retriever.SemanticRetrieverOptions(
      embedder=embedder,
      embedding_dimension=768,
      chunk_size=512,
      chunk_overlap=100,
      chunking_mode=semantic_retriever.ChunkingMode.CHARACTER,
  )

  # 3. Create the Semantic Retriever
  with semantic_retriever.SemanticRetriever.create_from_options(
      retriever_options
  ) as retriever:
    ...

Configuration options

SemanticRetrieverOptions accepts the following options:

Option Name Description Value Range Default Value
embedder The UniversalEmbedder instance used to generate embeddings for ingested content and queries. UniversalEmbedder Required
embedding_dimension The dimension of the embedding vectors produced by the embedder. Positive integer 768
database_path The file path of the on-device SQLite database used to persist the vector store. If omitted, an in-memory store is used. Path as a string None
chunk_size The maximum size of each chunk that long text documents are split into before embedding. Positive integer 512
chunk_overlap The number of characters or words that consecutive chunks overlap by. Non-negative integer 100
chunking_mode Whether chunk_size and chunk_overlap are measured in characters or words. ChunkingMode.CHARACTER, ChunkingMode.WORD ChunkingMode.CHARACTER

The remaining examples on this page assume they run inside the with block that creates retriever.

Ingest content

SemanticRetriever provides ingestion APIs for text documents, images, audio, and multimodal content. Each record is identified by a unique ID and can carry an optional dictionary of string metadata, which you can later use to filter retrieval results or delete records.

Ingest text documents

Insert documents with or without accompanying metadata:

# Ingest text without metadata
retriever.insert_document(
    doc_id='doc_1',
    text='The capital of France is Paris.',
)

# Ingest text with metadata
retriever.insert_document(
    doc_id='doc_burger',
    text='A grilled beef burger topped with melted cheddar, lettuce, and '
         'tomato on a toasted sesame bun.',
    metadata={'category': 'food', 'type': 'text'},
)

Ingest images and audio

Insert images and WAV audio files by their file paths. You can optionally pass the already-decoded data with image_bytes or audio_data to avoid reading the file from disk:

retriever.insert_image(
    image_id='img_cat',
    image_path='cat.jpg',
    metadata={'category': 'animals', 'type': 'image'},
)

retriever.insert_audio(
    audio_id='aud_birds',
    audio_path='birds_singing.wav',
    metadata={'category': 'animals', 'type': 'audio'},
)

Ingest multimodal content

To store a single record that combines several modalities, pass a list of part dictionaries to insert_content(). Each part sets a kind from TaskPartKind (TEXT, IMAGE, or AUDIO) along with the matching payload key: text, image_bytes or file_path, or audio_data or audio_path.

retriever.insert_content(
    record_id='combo_burger',
    parts=[
        {
            'kind': semantic_retriever.TaskPartKind.TEXT,
            'text': 'Juicy beef burger with fries',
        },
        {
            'kind': semantic_retriever.TaskPartKind.IMAGE,
            'file_path': 'burger.jpg',
        },
    ],
    metadata={'category': 'food'},
)

List stored records

Call get_all_record_ids() to inspect which records are currently indexed:

print('Indexed record IDs:', retriever.get_all_record_ids())

Retrieve records

To query the database, call retrieve() with a query, the maximum number of records to return (limit), and optionally a min_similarity threshold. The query can be a plain string or, for multimodal queries, a list of part dictionaries in the same format used by insert_content(). The method returns a RetrievalResult whose records list contains the matching records ordered by score.

result = retriever.retrieve(
    query='feline pet',
    limit=4,
    min_similarity=0.0,
)

for record in result.records:
  print(
      f'[{record.score:.4f}] id={record.id} '
      f'metadata={record.metadata} text={record.text!r}'
  )

Filter by metadata

Pass a metadata_filter dictionary to restrict the results to records whose metadata contains all of the given key-value pairs:

filtered_result = retriever.retrieve(
    query='something delicious to eat',
    limit=5,
    min_similarity=0.0,
    metadata_filter={'category': 'food'},
)

Query with an image

Use a list of parts to query with an image or audio instead of text:

with open('cat.jpg', 'rb') as f:
  cat_bytes = f.read()

image_query_result = retriever.retrieve(
    query=[
        {
            'kind': semantic_retriever.TaskPartKind.IMAGE,
            'image_bytes': cat_bytes,
        }
    ],
    limit=5,
    min_similarity=0.0,
)

Delete records and cleanup

You can remove a single record by ID, remove every record matching a metadata filter, or clear the database entirely:

# Delete a single record
retriever.delete_record('doc_1')

# Delete all records whose metadata matches the filter
retriever.delete_with_metadata_filter({'category': 'food'})

# Delete every record
retriever.delete_all()

The retriever and embedder release their resources when their with blocks exit. If you don't use with statements, call close() on each of them when you are done.