Universal embedder guide for iOS

The MediaPipe Universal Embedder task provides on-device generation of high-dimensional embedding vectors across multiple input modalities including text, images, and audio. These instructions show you how to use the Universal Embedder within iOS apps.

For more information about the capabilities, models, and configuration options of this task, see the Overview.

Setup

This section describes key steps for setting up your development environment and code projects specifically to use Universal Embedder. For general information on setting up your development environment for using MediaPipe tasks, including platform version requirements, see the Setup guide for iOS.

Libraries

Add MediaPipeTasksRetrieval to your project with Swift Package Manager:

// Swift Package Manager
.package(url: "https://github.com/google-ai-edge/mediapipe.git", from: "1.1.0")

Then, import the module in your Swift or Objective-C files:

import MediaPipeTasksRetrieval

Create the task

You initialize a UniversalEmbedder using the model path directly, or by passing a UniversalEmbedderOptions object configured with custom backend settings or L2 normalization.

// Configure options
let options = UniversalEmbedderOptions()
options.baseOptions.modelAssetPath = Bundle.main.path(forResource: "embeddinggemma-2-text-vision-440m", ofType: "litertlm")!
options.l2Normalize = true

// Initialize the Universal Embedder
let universalEmbedder = try UniversalEmbedder(options: options)

Generate embeddings

The UniversalEmbedder offers specialized inference methods for text, images, and audio, as well as a generic list content API.

Embed Text

To extract the high-dimensional feature vector for a text string:

let textResult = try universalEmbedder.embed(text: "The quick brown fox jumps over the lazy dog")
let floatVector = textResult.embeddings[0].floatEmbedding

Embed Image

You can embed image data by passing an MPImage object or the raw image bytes:

// Using an MPImage object (e.g. UIImage wrapper)
let image = try MPImage(uiImage: UIImage(named: "monument.jpg")!)
let imageResult = try universalEmbedder.embed(image: image)

// Or using raw NSData compressed image bytes
let data = try Data(contentsOf: imageURL)
let bytesResult = try universalEmbedder.embed(imageBytes: data)

Embed Audio

You can embed audio data by passing a float buffer or an AudioData object:

let audioResult = try universalEmbedder.embed(audio: audioData)

Embed Multimodal Content

You can supply an array of diverse content pieces (strings, image objects, audio objects) to generate a composite representation:

let contentArray: [Any] = [
    "A majestic landmark in Paris",
    image
]
let compositeResult = try universalEmbedder.embed(content: contentArray)

Compute similarity

You can compute the semantic cosine similarity between any two returned embedding vectors:

let embeddingA = textResult.embeddings[0]
let embeddingB = compositeResult.embeddings[0]

let similarityNumber = try UniversalEmbedder.cosineSimilarity(
    embedding1: embeddingA,
    embedding2: embeddingB
)
print("Similarity Score: \(similarityNumber.doubleValue)")