The MediaPipe Universal Embedder task provides on-device generation of high-dimensional embedding vectors across multiple input modalities including text, images, and audio. These instructions show you how to use the Universal Embedder within Android apps.
For more information about the capabilities, models, and configuration options of this task, see the Overview.
Code example
The MediaPipe Tasks example code is a basic implementation of a Universal Embedder app for Android. You can use the app as a starting point for your own Android app, or refer to it when modifying an existing app. The Universal Embedder example code is hosted on GitHub.
Download the code
The following instructions show you how to create a local copy of the example code using the git command line tool.
To download the example code:
- Clone the git repository using the following command:
git clone https://github.com/google-ai-edge/mediapipe-samples
- Optionally, configure your git instance to use sparse checkout,
so you have only the files for the Universal Embedder example app:
cd mediapipe-samples git sparse-checkout init --cone git sparse-checkout set examples/universal_embedder/android
After creating a local version of the example code, you can import the project into Android Studio and run the app. For instructions, see the Setup Guide for Android.
Setup
This section describes key steps for setting up your development environment and code projects specifically to use Universal Embedder. For general information on setting up your development environment for using MediaPipe tasks, including platform version requirements, see the Setup guide for Android.
Dependencies
Universal Embedder uses the com.google.mediapipe:tasks-retrieval library. The
retrieval package directly depends on LiteRT-LM, so you don't need to specify
the LiteRT-LM dependency manually. Add the following dependency to the
build.gradle file of your Android app:
dependencies {
implementation 'com.google.mediapipe:tasks-retrieval:latest.release'
}
Create the task
You initialize a UniversalEmbedder using the createFromOptions() factory
function, which accepts an Android Context and a UniversalEmbedderOptions
configuration object containing the model asset path and inference settings. To
use hardware acceleration, specify options to set a GPU delegate.
import android.content.Context;
import com.google.mediapipe.tasks.core.BaseOptions;
import com.google.mediapipe.tasks.core.Delegate;
import com.google.mediapipe.tasks.retrieval.universalembedder.UniversalEmbedder;
import com.google.mediapipe.tasks.retrieval.universalembedder.UniversalEmbedderOptions;
// Configure the embedder options
UniversalEmbedderOptions options = UniversalEmbedderOptions.builder()
.setBaseOptions(BaseOptions.builder()
.setModelAssetPath("embeddinggemma-2-text-vision-440m.litertlm")
.build())
.setTextDelegate(Delegate.GPU)
.setVisionDelegate(Delegate.GPU)
.setL2Normalize(true)
.build();
// Initialize the Universal Embedder
UniversalEmbedder universalEmbedder = UniversalEmbedder.createFromOptions(context, options);
Generate embeddings
The UniversalEmbedder provides distinct inference methods for individual input
modalities as well as a generic multimodal list ingestion API.
Embed Text
To extract the high-dimensional feature vector for a text string:
import com.google.mediapipe.tasks.components.containers.EmbeddingResult;
import com.google.mediapipe.tasks.components.containers.Embedding;
EmbeddingResult textResult = universalEmbedder.embedText("a yellow piece of fruit");
Embedding textEmbedding = textResult.embeddings().get(0);
Embed Image
To extract embeddings from a raw image using BitmapImageBuilder:
import com.google.mediapipe.framework.image.BitmapImageBuilder;
import com.google.mediapipe.framework.image.MPImage;
MPImage image = new BitmapImageBuilder(bitmap).build();
EmbeddingResult imageResult = universalEmbedder.embedImage(image);
Embedding imageEmbedding = imageResult.embeddings().get(0);
Embed Audio
To extract embeddings from a segmented audio buffer:
import com.google.mediapipe.tasks.components.containers.AudioData;
AudioData audioData = ...; // Your AudioData object
EmbeddingResult audioResult = universalEmbedder.embedAudio(audioData);
Embedding audioEmbedding = audioResult.embeddings().get(0);
Embed Multimodal Content
You can pass a mixed list of content objects (strings, audio data buffers, MPImages, or byte arrays) to generate a composite multimodal embedding representation:
List<Object> contents = new ArrayList<>();
contents.add("A majestic landmark in Paris");
contents.add(image);
EmbeddingResult multiResult = universalEmbedder.embedContent(contents);
Compute cosine similarity
You can calculate the semantic similarity between two resulting embeddings using built-in utility functions mapping both items into the shared vector space:
// Compute similarity directly using Embedding objects
double score = UniversalEmbedder.cosineSimilarity(textEmbedding, imageEmbedding);
Integrate with Semantic Retriever
The UniversalEmbedder is fully compatible with the SemanticRetriever
ecosystem. You can expose its embedding generator as a standard provider:
import com.google.mediapipe.tasks.core.EmbeddingProvider;
EmbeddingProvider provider = universalEmbedder.getProvider();
Cleanup
Release the underlying system resources and close the LiteRT-LM inference engine when finished:
universalEmbedder.close();