Universal embedder guide for Python

The MediaPipe Universal Embedder task provides on-device generation of high-dimensional embedding vectors across multiple input modalities including text, images, and audio. These instructions show you how to use the Universal Embedder within Python.

For more information about the capabilities, models, and configuration options of this task, see the Overview.

Setup

This section describes key steps for setting up your development environment and code projects specifically to use Universal Embedder. For general information on setting up your development environment for using MediaPipe tasks, including platform version requirements, see the Setup guide for Python.

Packages

The Universal Embedder requires the mediapipe pip package:

$ python -m pip install mediapipe

Imports

Import the following classes to access the Universal Embedder task functions:

import mediapipe as mp
from mediapipe.tasks import python
from mediapipe.tasks.python import retrieval

Create the task

You initialize a UniversalEmbedder by passing a UniversalEmbedderOptions object containing the model asset path and other configuration options to the create_from_options function.

from mediapipe.tasks.python.retrieval import UniversalEmbedder, UniversalEmbedderOptions

# Configure options
options = UniversalEmbedderOptions(
    base_options=mp.tasks.BaseOptions(model_asset_path="embeddinggemma-2-text-vision-440m.litertlm"),
    l2_normalize=True
)

# Initialize the Universal Embedder
universal_embedder = UniversalEmbedder.create_from_options(options)

Generate embeddings

The UniversalEmbedder provides distinct inference methods for individual input modalities as well as a generic list content API.

Embed Text

To extract the high-dimensional feature vector for a text string:

text_result = universal_embedder.embed_text("The quick brown fox jumps over the lazy dog")
float_vector = text_result.embeddings[0].float_embedding

Embed Image

To extract embeddings from an image:

image = mp.Image.create_from_file("monument.jpg")
image_result = universal_embedder.embed_image(image)

Embed Audio

To extract embeddings from an audio file:

audio_data = mp.tasks.audio.AudioData.create_from_wav_file("birds.wav")
audio_result = universal_embedder.embed_audio(audio_data)

Embed Multimodal Content

You can pass a mixed list of content objects (strings, audio data buffers, images, or byte arrays) to generate a composite multimodal embedding representation:

contents = [
    "A majestic landmark in Paris",
    image
]
multi_result = universal_embedder.embed_content(contents)

Compute similarity

You can calculate the semantic cosine similarity between any two resulting embedding vectors:

embedding_a = text_result.embeddings[0]
embedding_b = multi_result.embeddings[0]

# Compute similarity
score = UniversalEmbedder.cosine_similarity(embedding_a, embedding_b)
print(f"Similarity Score: {score}")

Cleanup

Release the system resources and close the underlying LiteRT-LM inference engine when finished:

universal_embedder.close()