The MediaPipe Semantic Retriever task manages on-device vector embedding generation, vector storage, and search retrieval for text, images, and audio files. These instructions show you how to use the Semantic Retriever within Python.
For more information about the capabilities, models, and configuration options of this task, see the Overview.
Setup
This section describes key steps for setting up your development environment and code projects specifically to use Semantic Retriever. For general information on setting up your development environment for using MediaPipe tasks, including platform version requirements, see the Setup guide for Python.
Packages
The Semantic Retriever requires the mediapipe pip package:
$ python -m pip install mediapipe
Imports
Import the following modules to access the Semantic Retriever task functions:
from mediapipe.tasks import python
from mediapipe.tasks.python.retrieval import semantic_retriever
from mediapipe.tasks.python.retrieval import universal_embedder
Create the task
The Semantic Retriever uses a UniversalEmbedder to generate embeddings for the
content you ingest and for your queries. First create the embedder, then pass it
to a SemanticRetrieverOptions object and create the retriever with
create_from_options(). Both objects are context managers, so you can use
with statements to release their resources automatically.
# 1. Create the embedder that generates the vector embeddings
base_options = python.BaseOptions(
model_asset_path='embeddinggemma-2-text-vision-440m.litertlm'
)
embedder_options = universal_embedder.UniversalEmbedderOptions(
base_options=base_options,
l2_normalize=True,
)
with universal_embedder.UniversalEmbedder.create_from_options(
embedder_options
) as embedder:
# 2. Configure the retriever
retriever_options = semantic_retriever.SemanticRetrieverOptions(
embedder=embedder,
embedding_dimension=768,
chunk_size=512,
chunk_overlap=100,
chunking_mode=semantic_retriever.ChunkingMode.CHARACTER,
)
# 3. Create the Semantic Retriever
with semantic_retriever.SemanticRetriever.create_from_options(
retriever_options
) as retriever:
...
Configuration options
SemanticRetrieverOptions accepts the following options:
| Option Name | Description | Value Range | Default Value |
|---|---|---|---|
embedder |
The UniversalEmbedder instance used to generate embeddings for ingested content and queries. |
UniversalEmbedder |
Required |
embedding_dimension |
The dimension of the embedding vectors produced by the embedder. | Positive integer | 768 |
database_path |
The file path of the on-device SQLite database used to persist the vector store. If omitted, an in-memory store is used. | Path as a string | None |
chunk_size |
The maximum size of each chunk that long text documents are split into before embedding. | Positive integer | 512 |
chunk_overlap |
The number of characters or words that consecutive chunks overlap by. | Non-negative integer | 100 |
chunking_mode |
Whether chunk_size and chunk_overlap are measured in characters or words. |
ChunkingMode.CHARACTER, ChunkingMode.WORD |
ChunkingMode.CHARACTER |
The remaining examples on this page assume they run inside the with block
that creates retriever.
Ingest content
SemanticRetriever provides ingestion APIs for text documents, images, audio,
and multimodal content. Each record is identified by a unique ID and can carry
an optional dictionary of string metadata, which you can later use to filter
retrieval results or delete records.
Ingest text documents
Insert documents with or without accompanying metadata:
# Ingest text without metadata
retriever.insert_document(
doc_id='doc_1',
text='The capital of France is Paris.',
)
# Ingest text with metadata
retriever.insert_document(
doc_id='doc_burger',
text='A grilled beef burger topped with melted cheddar, lettuce, and '
'tomato on a toasted sesame bun.',
metadata={'category': 'food', 'type': 'text'},
)
Ingest images and audio
Insert images and WAV audio files by their file paths. You can optionally pass
the already-decoded data with image_bytes or audio_data to avoid reading the
file from disk:
retriever.insert_image(
image_id='img_cat',
image_path='cat.jpg',
metadata={'category': 'animals', 'type': 'image'},
)
retriever.insert_audio(
audio_id='aud_birds',
audio_path='birds_singing.wav',
metadata={'category': 'animals', 'type': 'audio'},
)
Ingest multimodal content
To store a single record that combines several modalities, pass a list of part
dictionaries to insert_content(). Each part sets a kind from TaskPartKind
(TEXT, IMAGE, or AUDIO) along with the matching payload key: text,
image_bytes or file_path, or audio_data or audio_path.
retriever.insert_content(
record_id='combo_burger',
parts=[
{
'kind': semantic_retriever.TaskPartKind.TEXT,
'text': 'Juicy beef burger with fries',
},
{
'kind': semantic_retriever.TaskPartKind.IMAGE,
'file_path': 'burger.jpg',
},
],
metadata={'category': 'food'},
)
List stored records
Call get_all_record_ids() to inspect which records are currently indexed:
print('Indexed record IDs:', retriever.get_all_record_ids())
Retrieve records
To query the database, call retrieve() with a query, the maximum number of
records to return (limit), and optionally a min_similarity threshold. The
query can be a plain string or, for multimodal queries, a list of part
dictionaries in the same format used by insert_content(). The method returns a
RetrievalResult whose records list contains the matching records ordered by
score.
result = retriever.retrieve(
query='feline pet',
limit=4,
min_similarity=0.0,
)
for record in result.records:
print(
f'[{record.score:.4f}] id={record.id} '
f'metadata={record.metadata} text={record.text!r}'
)
Filter by metadata
Pass a metadata_filter dictionary to restrict the results to records whose
metadata contains all of the given key-value pairs:
filtered_result = retriever.retrieve(
query='something delicious to eat',
limit=5,
min_similarity=0.0,
metadata_filter={'category': 'food'},
)
Query with an image
Use a list of parts to query with an image or audio instead of text:
with open('cat.jpg', 'rb') as f:
cat_bytes = f.read()
image_query_result = retriever.retrieve(
query=[
{
'kind': semantic_retriever.TaskPartKind.IMAGE,
'image_bytes': cat_bytes,
}
],
limit=5,
min_similarity=0.0,
)
Delete records and cleanup
You can remove a single record by ID, remove every record matching a metadata filter, or clear the database entirely:
# Delete a single record
retriever.delete_record('doc_1')
# Delete all records whose metadata matches the filter
retriever.delete_with_metadata_filter({'category': 'food'})
# Delete every record
retriever.delete_all()
The retriever and embedder release their resources when their with blocks
exit. If you don't use with statements, call close() on each of them when
you are done.