Semantic retriever guide for iOS

The MediaPipe Semantic Retriever task manages on-device vector embedding generation, vector storage, and search retrieval for text, images, and audio files. These instructions show you how to use the Semantic Retriever within iOS applications.

For more information about the capabilities, models, and configuration options of this task, see the Overview.

Setup

This section describes key steps for setting up your development environment and code projects specifically to use Semantic Retriever. For general information on setting up your development environment for using MediaPipe tasks, including platform version requirements, see the Setup guide for iOS.

Libraries

Add MediaPipeTasksRetrieval to your project with Swift Package Manager:

// Swift Package Manager
.package(url: "https://github.com/google-ai-edge/mediapipe.git", from: "1.1.0")

Then, import the module in your Swift or Objective-C files:

import MediaPipeTasksRetrieval

Create the task

You initialize a SemanticRetriever using the SemanticRetriever(components:) initializer or the SemanticRetriever.create(fromComponents:) factory method. These accept a SemanticRetrieverComponents object configuring the underlying vector store and embedding providers.

// 1. Set up the embedding provider
let embedderOptions = UniversalEmbedderOptions()
embedderOptions.baseOptions.modelAssetPath = Bundle.main.path(forResource: "embeddinggemma-2-text-vision-440m", ofType: "litertlm")!
let embedder = try UniversalEmbedder(options: embedderOptions)

// 2. Set up the vector store (SQLite or In-Memory) and components
let vectorStore = SqliteVectorStore(
    databasePath: databasePath,
    embeddingDimension: MPPSemanticRetrieverDefaultEmbeddingDimension
)
let components = try SemanticRetrieverComponents(
    vectorStore: vectorStore,
    chunker: nil,
    providers: [embedder]
)

// 3. Initialize the Semantic Retriever
let semanticRetriever = try SemanticRetriever(components: components)

Ingest content

SemanticRetriever provides convenient ingestion APIs for multiple content modalities.

Ingest Text Documents

You can insert text documents with or without accompanying metadata:

// Ingest text without metadata
try semanticRetriever.insertDocument(withId: "doc_1", text: "The capital of France is Paris.")

// Ingest text with metadata
let metadata = ["category": "geography", "language": "en"]
try semanticRetriever.insertDocument(withId: "doc_2", text: "The capital of Spain is Madrid.", metadata: metadata)

Ingest Images and Audio

You can ingest local files by supplying their file paths:

let imagePath = Bundle.main.path(forResource: "eiffel_tower", ofType: "jpg")!
try semanticRetriever.insertImage(withId: "img_1", filePath: imagePath)

let audioPath = Bundle.main.path(forResource: "birds_singing", ofType: "wav")!
try semanticRetriever.insertAudio(withId: "aud_1", filePath: audioPath)

Ingest Multimodal Content

You can ingest complex content blocks by passing a list of TaskPart elements:

let textPart = TextPart(text: "An iconic iron monument in Paris")
let imagePart = ImagePart(filePath: imagePath)

try semanticRetriever.insertContent(withId: "composite_1", parts: [textPart, imagePart])

Retrieve records

To query your stored on-device vector database, call retrieve with a query string or parts, and the number of results to fetch (topK):

let results = try semanticRetriever.retrieve(withText: "What is the capital of France?", topK: 3)

for result in results {
  let recordId = result.recordId
  let score = result.score
  let meta = result.metadata
  print("Record: \(recordId), Match Score: \(score), Metadata: \(meta)")
}

Delete records

You can remove elements from the database using their record IDs:

try semanticRetriever.delete(withIds: ["doc_1", "img_1"])