The MediaPipe Semantic Retriever task manages on-device vector embedding generation, vector storage, and search retrieval for text, images, and audio files. These instructions show you how to use the Semantic Retriever within iOS applications.
For more information about the capabilities, models, and configuration options of this task, see the Overview.
Setup
This section describes key steps for setting up your development environment and code projects specifically to use Semantic Retriever. For general information on setting up your development environment for using MediaPipe tasks, including platform version requirements, see the Setup guide for iOS.
Libraries
Add MediaPipeTasksRetrieval to your project with Swift Package Manager:
// Swift Package Manager
.package(url: "https://github.com/google-ai-edge/mediapipe.git", from: "1.1.0")
Then, import the module in your Swift or Objective-C files:
import MediaPipeTasksRetrieval
Create the task
You initialize a SemanticRetriever using the SemanticRetriever(components:)
initializer or the SemanticRetriever.create(fromComponents:) factory method.
These accept a SemanticRetrieverComponents object configuring the underlying
vector store and embedding providers.
// 1. Set up the embedding provider
let embedderOptions = UniversalEmbedderOptions()
embedderOptions.baseOptions.modelAssetPath = Bundle.main.path(forResource: "embeddinggemma-2-text-vision-440m", ofType: "litertlm")!
let embedder = try UniversalEmbedder(options: embedderOptions)
// 2. Set up the vector store (SQLite or In-Memory) and components
let vectorStore = SqliteVectorStore(
databasePath: databasePath,
embeddingDimension: MPPSemanticRetrieverDefaultEmbeddingDimension
)
let components = try SemanticRetrieverComponents(
vectorStore: vectorStore,
chunker: nil,
providers: [embedder]
)
// 3. Initialize the Semantic Retriever
let semanticRetriever = try SemanticRetriever(components: components)
Ingest content
SemanticRetriever provides convenient ingestion APIs for multiple content
modalities.
Ingest Text Documents
You can insert text documents with or without accompanying metadata:
// Ingest text without metadata
try semanticRetriever.insertDocument(withId: "doc_1", text: "The capital of France is Paris.")
// Ingest text with metadata
let metadata = ["category": "geography", "language": "en"]
try semanticRetriever.insertDocument(withId: "doc_2", text: "The capital of Spain is Madrid.", metadata: metadata)
Ingest Images and Audio
You can ingest local files by supplying their file paths:
let imagePath = Bundle.main.path(forResource: "eiffel_tower", ofType: "jpg")!
try semanticRetriever.insertImage(withId: "img_1", filePath: imagePath)
let audioPath = Bundle.main.path(forResource: "birds_singing", ofType: "wav")!
try semanticRetriever.insertAudio(withId: "aud_1", filePath: audioPath)
Ingest Multimodal Content
You can ingest complex content blocks by passing a list of TaskPart elements:
let textPart = TextPart(text: "An iconic iron monument in Paris")
let imagePart = ImagePart(filePath: imagePath)
try semanticRetriever.insertContent(withId: "composite_1", parts: [textPart, imagePart])
Retrieve records
To query your stored on-device vector database, call retrieve with a query
string or parts, and the number of results to fetch (topK):
let results = try semanticRetriever.retrieve(withText: "What is the capital of France?", topK: 3)
for result in results {
let recordId = result.recordId
let score = result.score
let meta = result.metadata
print("Record: \(recordId), Match Score: \(score), Metadata: \(meta)")
}
Delete records
You can remove elements from the database using their record IDs:
try semanticRetriever.delete(withIds: ["doc_1", "img_1"])