
Welcome back to the blog post series “Build intelligent Android apps” where we take a basic Android app and transform it into a personalized, intelligent, and agentic experience. In our previous post we introduced Jetpacker, the demo app we’ll use throughout this series.
In this blog post, we will share how you can use Gemini Nano through ML Kit’s Prompt API to build intelligent on-device features.
Building intelligent on-device features refers to the ability to process prompts and data directly on a device without sending data to a server. This offers a few advantages:
- User data can be processed locally on the device, preserving user privacy
- Functionality of the model is reliable even with spotty or no internet connection
- No additional cloud inference cost, since everything runs on the user’s hardware
With the benefits of on-device in mind, we identified three features to add in Jetpacker that can improve the user experience: summarizing trip itineraries, managing expenses, and capturing voice notes.
On-device features in Jetpacker: Summarizing trip itineraries, managing expenses, and voice notes
High quality tailored summarization of short texts
The itinerary screen gives users a quick overview of all activities for a given trip. Since this screen contains a lot of information, it can quickly become overwhelming. To help users prepare without feeling overwhelmed, we can add a ‘Get ready for your trip’ section at the top.
The romantic Paris trip is summarized as a classic Parisian adventure blending art, sights, and delicious food. A tip and some useful phrases are also added.
By inputting a trip itinerary and asking an LLM to summarize it, we can generate a quick summary of the trip along with packing tips and useful local phrases. This is a great use case for an on-device model for several reasons:
- Performance and quality: Both the input and output text are relatively short. With that, we can expect the performance and quality of an on-device solution to be on par with more powerful cloud models.
- Scalability: Shifting inference on-device allows us to scale this feature from a few users to millions without worrying about managing increasing cloud inference costs.
- Low latency and reliability: On-device inference guarantees low latency, providing a reliable experience even when users are offline.
To build with on-device, we use Gemini Nano, Google’s most efficient model optimized for mobile devices. Gemini Nano was first introduced a few years ago, and is now running on over 140 million devices. The latest version of the model, Gemini Nano 4, is built on the architecture foundation of the recently released Gemma 4 model, and is further optimized for maximum battery and performance efficiency.
Using ML Kit’s Prompt API, we can take advantage of Gemini Nano 4’s new model capabilities to prototype our on-device features. We’ll create a prompt that includes the itinerary of a trip and ask the model to generate a summary along with any preparation tips.
// implementation("com.google.mlkit:genai-prompt:1.0.0-beta3")
// Define the configuration for Gemini Nano 4 E2B preview model
val previewFastConfig = generationConfig {
modelConfig = modelConfig {
releaseStage = ModelReleaseStage.PREVIEW
preference = ModelPreference.FAST
}
}
val geminiNano2BPreviewModel = Generation.getClient(previewFastConfig)
val tripItinerary = ...
val getReadyForYourTripSummary = geminiNano2BPreviewModel
.generateContent("Given this trip itinerary: $tripItinerary,
generate the following: overall vibe, tips on how to prepare for this
trip, and common short phrases to learn for the trip.")Finding the optimal prompt usually requires some iteration, and the AICore app is perfect for this step in the process. After opting into the developer preview option for AICore, we can download preview models such as Gemini Nano 4 to test prompts and see the model’s expected outputs. With a few iterations on the prompt, we were able to improve the speed of the response from 13 seconds to under 2 seconds! Check out the final code implementation and prompt here.
The first iteration of our prompt generated way too many tokens, and optimizing it helped keep responses quick and to the point.
Local processing for sensitive user input
Next, to help users enjoy their trip even more, we’ll build a simple expense manager that takes the manual work out of sorting through receipts and calculating budgets.
Taking a photo of a restaurant bill, data is parsed and shown in the expense overview screen of the app.
Since receipts might contain sensitive information like credit card number and addresses, this is another great use case for an on-device solution. With on-device, users can be confident that private information will be processed locally on the device without any of their data being sent to the cloud.
In addition, Gemini Nano 4 has improved model capabilities for multimodality, especially for image understanding tasks like OCR and visual data extraction, making it a great solution for tasks like extracting information from receipts.
For this use case, the prompt will analyze an image of the receipt, and output information such as: a generated title, amount spent and category of the expense. To ensure the model outputs the information in the preferred format, we can use ML Kit’s Structured Output API to seamlessly output a Kotlin data object that we define.
// implementation("com.google.mlkit:genai-prompt:1.0.0-beta3")
// ksp("com.google.mlkit:genai-schema-compiler:1.0.0-alpha1")
@Generable("Information extracted from an expense receipt")
data class ParsedReceipt(
@Guide("Generated title for the expense less than 6 words. Based on restaurant or activity name.")
val title: String,
@Guide("Total amount of the expense. Look for values at the bottom and words like total or balance due.")
val amount: Double,
@Guide("Type of expense", enumValues = ["travel", "food", "shopping", "entertainment", "other"])
val category: String,
)
val prompt = "Determine if the image is a receipt or expense.
If it is NOT a receipt or expense, output the text 'NOT_A_RECEIPT'.
Otherwise, parse the receipt information."
val request = generateContentRequest(ImagePart(bitmap), TextPart(prompt)) {}
val requestWithStructuredOutput = generateTypedContentRequest(request, ParsedReceipt::class)
// Define the configuration for Gemini Nano 4 E4B preview model
// When selecting models, you can specify which performance charactertists are most important
// for your use case. Use ModelPreference.FULL when you want to prioritize reasoning power over speed.
// Use ModelPreference.FAST when complex logic is not required and latency is a priority.
val previewFullConfig = generationConfig {
modelConfig = modelConfig {
releaseStage = ModelReleaseStage.PREVIEW
preference = ModelPreference.FULL
}
}
val geminiNano4BPreviewModel = Generation.getClient(previewFullConfig)
val response = geminiNano4BPreviewModel.generateContent(requestWithStructuredOutput)
val parsedReceipt: ParsedReceipt? = response.candidates.firstOrNull()?.responseMultimodal input
Lastly, to help users record audio memos during the trip, let’s build a fully on-device voice notes feature. Using ML Kit’s Speech Recognition API, we’ll enable users to record short voice notes that are automatically transcribed to text. With the transcribed text, we’ll use ML Kit’s Prompt API to identify which trip activity is associated with the recorded voice note, letting users easily recap their trip as they scroll through the trip’s itinerary.
The Roman holiday itinerary shows voice note extracts.
The ML Kit GenAI Speech Recognition API allows you to transcribe audio content to text fully on-device using two distinct modes. Basic mode uses a traditional on-device speech recognition model and is available on most Android devices with API level 31 and higher. Advanced mode uses Gemini Nano to offer broader language coverage and better quality, and is currently supported on Pixel 10 devices.
For our feature we combine the Speech Recognition API with the ML Kit GenAI Prompt API:
// implementation("com.google.mlkit:genai-prompt:1.0.0-beta3")
// implementation("com.google.mlkit:genai-speech-recognition:1.0.0-alpha1")
val tripEvents = ...
// Set up speech recognition
val speechRecognizerOptions =
speechRecognizerOptions {
locale = Locale.US
preferredMode = SpeechRecognizerOptions.Mode.MODE_ADVANCED
}
val speechRecognizer: SpeechRecognizer = SpeechRecognition.getClient(speechRecognizerOptions)
suspend fun transcribeVoiceNote(recognizer: SpeechRecognizer) {
// Display partial text as the user is recording audio
var partialTextResponse = ""
// Display the full text once user is finished recording audio
var transcription = ""
val request: SpeechRecognizerRequest
= speechRecognizerRequest { audioSource = AudioSource.fromMic() }
recognizer.startRecognition(request).collect { response ->
when (response) {
is SpeechRecognizerResponse.PartialTextResponse -> {
partialTextResponse = response.text
}
is SpeechRecognizerResponse.FinalTextResponse -> {
transcription = response.text
processAndCategorizeVoiceNote(transcription, tripEvents)
}
}
}
}
fun processAndCategorizeVoiceNote(transcribedVoiceNote: String, events: List) {
val prompt = "Given the voice note $transcribedVoiceNote
and the following events for this trip: $events, rewrite this transcription
to remove filler words. Then, identify which events from the
list this rewritten transcription matches to."
// Utilize ML Kit's Prompt API to process voice note and tag it with the relevant trip activities
Generation.getClient().generateContent(prompt)
} Conclusion
Using ML Kit’s GenAI APIs, we were able to take advantage of Gemini Nano to develop fully on-device intelligent features for the JetPacker app, and provide an improved user experience without any additional cloud costs.
Check out the full source code for Jetpacker on Github, and watch the video Build Intelligent Android apps with Google’s AI to learn more about how to integrate intelligent features directly into your app using on-device models, cloud-powered reasoning, and the latest agentic frameworks.
Learn more
Check out the other parts of this blog post series:
Part 1: Introduction of the app and a high-level overview.
Part 2 (this post!): On-device intelligence. Deep-dive into ML Kit’s GenAI APIs and Gemini Nano to build privacy-first features like itinerary summarization, receipt parsing, and local audio processing.
Part 3: Hybrid and cloud reasoning. Explore how to use Firebase AI Logic to ground LLM answers in real-world data like Google Maps and web context.
Part 4: System integration. Integrating with the Android intelligence system using AppFunctions.
Part 5 (coming soon): In-app agentic workflows. Extend the app with an end-to-end booking assistant powered by A2UI and ADK.
Interested in more on Android Development? Follow Android Developers on YouTube or LinkedIn!
All code snippets in this blog post follow the following copyright notice:
Copyright 2026 Google LLC.
SPDX-License-Identifier: Apache-2.0








