The MediaPipe Universal Embedder task combines the features of individual audio, text, and image embedders into a single unified service. This guide shows you how to use the Universal Embedder for web and JavaScript applications.
For more information about the capabilities, models, and configuration options of this task, see the Overview.
Code example
The example code for Universal Embedder provides a complete implementation of this task in JavaScript for your reference. This code helps you test this task and get started on building your own multimodal embedding app. You can view the Universal Embedder example code on GitHub.
Setup
This section describes key steps for setting up your development environment and code projects specifically to use Universal Embedder. For general information on setting up your development environment for using MediaPipe tasks, including platform version requirements, see the Setup guide for Web.
JavaScript packages
Universal Embedder code is available through the
@mediapipe/tasks-retrieval
package. You can find and download these libraries from links provided in the
platform Setup guide.
You can install the required packages with the following code for local staging using the following command:
npm install @mediapipe/tasks-retrieval
If you want to deploy to a server, you can use a content delivery network (CDN) service, such as jsDelivr, to add code directly to your HTML page, as follows:
<head>
<script src="https://cdn.jsdelivr.net/npm/@mediapipe/tasks-retrieval@latest/retrieval_bundle.mjs" crossorigin="anonymous"></script>
</head>
Model
The MediaPipe Universal Embedder task requires a trained model that is compatible with this task. For more information on available models, see the Models section of the task overview.
You can download a model and host it alongside your application, or reference
it directly by URL in the modelAssetPath option.
Create the task
Use one of the Universal Embedder createFrom...() functions to prepare the task for
running inferences. The following code example demonstrates using the
createFromOptions() function.
import { UniversalEmbedder, FilesetResolver } from "@mediapipe/tasks-retrieval";
let universalEmbedder;
async function createEmbedder() {
const retrieval = await FilesetResolver.forRetrievalTasks("https://cdn.jsdelivr.net/npm/@mediapipe/tasks-retrieval@latest/wasm");
universalEmbedder = await UniversalEmbedder.createFromOptions(
retrieval,
{
baseOptions: {
modelAssetPath: "https://huggingface.co/litert-community/embeddinggemma-2-text-vision-440m-litert-lm/resolve/main/embeddinggemma-2-text-vision-440m.litertlm"
}
}
);
}
createEmbedder();
Prepare data & Run inference
The UniversalEmbedder provides distinct inference methods for individual input
modalities (embedText, embedImage, embedAudio), as well as a generic
embedContent API when processing multimodal content natively on the Web. All
of these methods are asynchronous and return a Promise.
Text embeddings
Format your text natively as a String and invoke the text embedding method:
const textEmbeddingResult = await universalEmbedder.embedText("The quick brown fox jumps over the lazy dog");
console.log("Text embedding:", textEmbeddingResult.embeddings[0].floatEmbedding);
Image embeddings
Pass the encoded image bytes (for example, the contents of a JPEG or PNG file)
as a Uint8Array into the image embedding method:
const response = await fetch("monument.jpg");
const imageBytes = new Uint8Array(await response.arrayBuffer());
const imageEmbeddingResult = await universalEmbedder.embedImage(imageBytes);
console.log("Image embedding:", imageEmbeddingResult.embeddings[0].floatEmbedding);
Audio embeddings
Pass a Float32Array containing the raw audio samples into the audio embedding
method:
// Assuming you have extracted audio data into a Float32Array (e.g. from an AudioContext)
const audioSamples = new Float32Array([...]);
const audioEmbeddingResult = await universalEmbedder.embedAudio(audioSamples);
console.log("Audio embedding:", audioEmbeddingResult.embeddings[0].floatEmbedding);
Handle and display results
The UniversalEmbedder returns an EmbeddingResult object that typically
contains a single Embedding. Each Embedding includes the high-dimensional
feature vector, which may be raw floats or scalar quantizations depending on the
configuration.
function processEmbeddingResult(result) {
if (result.embeddings.length > 0) {
const floatVector = result.embeddings[0].floatEmbedding;
console.log(`Generated an embedding with ${floatVector.length} dimensions.`);
}
}
processEmbeddingResult(textEmbeddingResult);
Use the UniversalEmbedder.cosineSimilarity() utility to compute the semantic
similarity between two embedding results on the fly:
const similarity = UniversalEmbedder.cosineSimilarity(
textEmbeddingResult.embeddings[0],
imageEmbeddingResult.embeddings[0]
);
console.log("Cosine similarity:", similarity);