Semantic retriever guide for Android

The MediaPipe Semantic Retriever task manages on-device vector embedding generation, vector storage, and search retrieval for text, images, and audio files. These instructions show you how to use the Semantic Retriever within Android apps.

For more information about the capabilities, models, and configuration options of this task, see the Overview.

Code example

The MediaPipe Tasks example code is a basic implementation of a Semantic Retriever app for Android. You can use the app as a starting point for your own Android app, or refer to it when modifying an existing app. The Semantic Retriever example code is hosted on GitHub.

Download the code

The following instructions show you how to create a local copy of the example code using the git command line tool.

To download the example code:

  1. Clone the git repository using the following command:
    git clone https://github.com/google-ai-edge/mediapipe-samples
    
  2. Optionally, configure your git instance to use sparse checkout, so you have only the files for the Semantic Retriever example app:
    cd mediapipe-samples
    git sparse-checkout init --cone
    git sparse-checkout set examples/semantic_retriever/android
    

After creating a local version of the example code, you can import the project into Android Studio and run the app. For instructions, see the Setup Guide for Android.

Setup

This section describes key steps for setting up your development environment and code projects specifically to use Semantic Retriever. For general information on setting up your development environment for using MediaPipe tasks, including platform version requirements, see the Setup guide for Android.

Dependencies

Semantic Retriever uses the com.google.mediapipe:tasks-retrieval library. The retrieval package directly depends on LiteRT-LM, so you don't need to specify the LiteRT-LM dependency manually. AppSearch dependencies are optional and only required if you plan to use AppSearchVectorStore. Add the dependencies to the build.gradle file of your Android app:

dependencies {
    implementation 'com.google.mediapipe:tasks-retrieval:latest.release'

    // Optional: Only required if using the AppSearch backend
    implementation 'androidx.appsearch:appsearch:1.2.0-alpha02'
    implementation 'androidx.appsearch:appsearch-local-storage:1.2.0-alpha02'
}

Create the task

You initialize a SemanticRetriever using the createFromComponents() factory function, which accepts an Android Context and a SemanticRetrieverComponents configuration block. This configuration defines the vector database instance and embedding providers.

import android.content.Context;
import com.google.mediapipe.tasks.retrieval.semanticretriever.SemanticRetriever;
import com.google.mediapipe.tasks.retrieval.semanticretriever.SemanticRetrieverComponents;
import com.google.mediapipe.tasks.retrieval.universalembedder.UniversalEmbedder;
import com.google.mediapipe.tasks.retrieval.components.SqliteVectorStore;

// 1. Initialize the embedding engine / provider
UniversalEmbedder embedder = UniversalEmbedder.createFromOptions(context, embedderOptions);

// 2. Build the components list
SemanticRetrieverComponents components = new SemanticRetrieverComponents()
    .setVectorStore(new SqliteVectorStore(context, "semantic_db"))
    .addProvider(embedder.getProvider());

// 3. Create the Semantic Retriever
SemanticRetriever semanticRetriever = SemanticRetriever.createFromComponents(context, components);

Ingest content

The SemanticRetriever provides dedicated on-device ingestion APIs for documents, images, audio, and multimodal chunks.

Ingest Text Documents

You can insert text documents. If the document is large, the retriever will automatically chunk it into smaller segments:

// Ingest text without metadata
semanticRetriever.insertDocument("doc_1", "The capital of France is Paris.");

// Ingest text with metadata
Map<String, String> metadata = new HashMap<>();
metadata.put("category", "geography");
semanticRetriever.insertDocument("doc_2", "The capital of Germany is Berlin.", metadata);

Ingest Images and Audio

You can ingest local media using their Android Uri paths:

Uri imageUri = Uri.parse("file:///android_asset/images/sunny_beach.jpg");
semanticRetriever.insertImage("img_1", imageUri);

Uri audioUri = Uri.parse("file:///android_asset/ambient_birds.wav");
semanticRetriever.insertAudio("aud_1", audioUri);

Ingest Multimodal Content

You can assemble custom lists of components using Part primitives:

import com.google.mediapipe.tasks.core.Part;
import com.google.mediapipe.tasks.core.TextPart;

List<Part> contentParts = new ArrayList<>();
contentParts.add(new TextPart("An iconic iron monument in Paris"));

semanticRetriever.insertContent("composite_1", contentParts);

Retrieve records

To query your stored data, call retrieve with a query string or query parts and the number of matching results (topK):

import com.google.mediapipe.tasks.retrieval.semanticretriever.RetrievalResult;

// Query the database using a query string
List<RetrievalResult> results = semanticRetriever.retrieve("What is the capital of France?", 3);

for (RetrievalResult result : results) {
    String id = result.id();
    double score = result.score();
    System.out.println("Match ID: " + id + ", Similarity Score: " + score);
}

Delete records and cleanup

You can remove elements from your vector database using their original record IDs, and release system resources when finished:

// Delete specific records by ID
List<String> idsToDelete = Arrays.asList("doc_1", "img_1");
semanticRetriever.delete(idsToDelete);

// Close and release database resources
semanticRetriever.close();