Dev
September 17, 2026
1 views
2 min read

Build intelligent Android apps: On-device inference

Curated by Patrick
Source: Android Developers
Build intelligent Android apps: On-device inference
Tech Daily Byte Analysis

In a new post, Android Developer Relations engineer Caren Chang walks developers through integrating Gemini Nano 4, Google’s most efficient mobile LLM, into Jetpacker. By calling the ML Kit Prompt API, the app generates concise trip summaries, extracts structured data from receipt photos, and links transcribed voice memos to itinerary events—all without leaving the handset. The implementation shows a drop in response time from 13 seconds to under two after prompt tuning, and it leverages the Structured Output API to map receipt fields directly into a Kotlin data class. Gemini Nano 4 runs on more than 140 million devices and offers two preference modes—FAST for low latency and FULL for deeper reasoning—allowing developers to balance speed against capability.

The showcase aligns with Google’s broader push to shift AI workloads from the cloud to the edge. By embedding a model derived from the Gemma 4 architecture, Google positions Gemini Nano as a competitor to Apple’s on‑device Core ML models and Meta’s upcoming LLaMA‑mobile efforts, promising comparable quality for short‑text tasks while eliminating per‑inference cloud costs. The emphasis on multimodal OCR and advanced speech recognition (available on Pixel 10 devices) reflects a market trend where developers demand on‑device privacy for sensitive data such as receipts and voice recordings, especially in regulated regions.

Looking ahead, the success of Jetpacker’s on‑device features will hinge on broader device support and the stability of preview APIs. Developers must monitor the rollout of Gemini Nano 4 beyond the preview stage, watch for battery‑impact reports, and evaluate how the model scales with longer prompts or richer multimodal inputs. If Google expands the model’s vocabulary and reasoning depth while maintaining sub‑second latency, on‑device LLMs could become the default for many consumer apps, but premature adoption may expose apps to inconsistent performance across the Android ecosystem.

Key Takeaways

Gemini Nano 4 can generate useful textual summaries and structured data locally, cutting inference latency to under two seconds after prompt optimization.

The ML Kit Prompt and Structured Output APIs let developers map LLM outputs directly to Kotlin objects, simplifying integration.

On‑device processing safeguards receipt and voice‑note privacy while avoiding cloud compute fees, a compelling proposition for data‑sensitive apps.

Adoption will depend on expanded device compatibility and real‑world battery impact as the preview moves to general availability.

About the Source

This analysis is based on reporting by Android Developers. Here is a short excerpt for context:

Posted by Caren Chang, Developer Relations Engineer, Android Developer Relations Welcome back to the blog post series "Build intelligent Android apps" where we take a basic Android app and transform it into a personalized, intelligent, and agentic experience. In our previous post we introduced Jetpacker, the demo app we'll use throughout this series. In this blog post, we will share how you can use Gemini Nano through ML Kit’s Prompt API to build intelligent on-device features. Building intelligent on-device features refers to the ability to process prompts and data directly on a device without sending data to a server. This offers a few advantages: User data can be processed locally on the device, preserving user privacy Functionality of the model is reliable even with spotty or no internet connection No additional cloud inference cost, since everything runs on the user’s hardware With the benefits of on-device in mind, we identified three features to add in Jetpacker that can improve the user experience: summarizing trip itineraries, managing expenses, and capturing voice notes. On-device features in Jetpacker: Summarizing trip itineraries, managing expenses, and voice notes High quality tailored summarization of short texts The itinerary screen gives users a quick overview of all activities for a given trip. Since this screen contains a lot of information, it can quickly become overwhelming. To help users prepare without feeling overwhelmed, we can add a ‘Get ready for your trip’ section at the top. The romantic Paris trip is summarized as a classic Parisian adventure blending art, sights, and delicious food. A tip and some useful phrases are also added. By inputting a trip itinerary and asking an LLM to summarize it, we can generate a quick summary of the trip along with packing tips and useful local phrases. This is a great use case for an on-device model for several reasons: Performance and quality: Both the input and output text are relatively short. With that, we can expect the performance and quality of an on-device solution to be on par with more powerful cloud models. Scalability: Shifting inference on-device allows us to scale this feature from a few users to millions without worrying about managing increasing cloud inference costs. Low latency and reliability: On-device inference guarantees low latency, providing a reliable experience even when users are offline. To build with on-device, we use Gemini Nano, Google’s most efficient model optimized for mobile devices. Gemini Nano was first introduced a few years ago, and is now running on over 140 million devices. The latest version of the model, Gemini Nano 4, is built on the architecture foundation of the recently released Gemma 4 model, and is further optimized for maximum battery and performance efficiency. Using ML Kit’s Prompt API, we can take advantage of Gemini Nano 4’s new model capabilities to prototype our on-device features. We’ll create a prompt that includes the itinerary of a trip and ask the model to generate a summary along with any preparation tips. // implementation("com.google.mlkit:genai-prompt:1.0.0-beta3") // Define the configuration for Gemini Nano 4 E2B preview model val previewFastConfig = generationConfig { modelConfig = modelConfig { releaseStage = ModelReleaseStage.PREVIEW preference = ModelPreference.FAST } } val geminiNano2BPreviewModel = Generation.getClient(previewFastConfig) val tripItinerary = ... val getReadyForYourTripSummary = geminiNano2BPreviewModel .generateContent("Given this trip itinerary: $tripItinerary, generate the following: overall vibe, tips on how to prepare for this trip, and common short phrases to learn for the trip.") Finding the optimal prompt usually requires some iteration, and the AICore app is perfect for this step in the process. After opting into the developer preview option for AICore, we can download preview models such as Gemini Nano 4 to test prompts and see the model’s expected outputs. With a few iterations on the prompt, we were able to improve the speed of the response from 13 seconds to under 2 seconds! Check out the final code implementation and prompt here. The first iteration of our prompt generated way too many tokens, and optimizing it helped keep responses quick and to the point. Local processing for sensitive user input Next, to help users enjoy their trip even more, we’ll build a simple expense manager that takes the manual work out of sorting through receipts and calculating budgets. Taking a photo of a restaurant bill, data is parsed and shown in the expense overview screen of the app. Since receipts might contain sensitive information like credit card number and addresses, this is another great use case for an on-device solution. With on-device, users can be confident that private information will be processed locally on the device without any of their data being sent to the cloud. In addition, Gemini Nano 4 has improved model capabilities for multimodality, especially for image understanding tasks like OCR and visual data extraction, making it a great solution for tasks like extracting information from receipts. For this use case, the prompt will analyze an image of the receipt, and output information such as: a generated title, amount spent and category of the expense. To ensure the model outputs the information in the preferred format, we can use ML Kit’s Structured Output API to seamlessly output a Kotlin data object that we define. // implementation("com.google.mlkit:genai-prompt:1.0.0-beta3") // ksp("com.google.mlkit:genai-schema-compiler:1.0.0-alpha1") @Generable("Information extracted from an expense receipt") data class ParsedReceipt( @Guide("Generated title for the expense less than 6 words. Based on restaurant or activity name.") val title: String, @Guide("Total amount of the expense. Look for values at the bottom and words like total or balance due.") val amount: Double, @Guide("Type of expense", enumValues = ["travel", "food", "shopping", "entertainment", "other"]) val category: String, ) val prompt = "Determine if the image is a receipt or expense. If it is NOT a receipt or expense, output the text 'NOT_A_RECEIPT'. Otherwise, parse the receipt information." val request = generateContentRequest(ImagePart(bitmap), TextPart(prompt)) {} val requestWithStructuredOutput = generateTypedContentRequest(request, ParsedReceipt::class) // Define the configuration for Gemini Nano 4 E4B preview model // When selecting models, you can specify which performance charactertists are most important // for your use case. Use ModelPreference.FULL when you want to prioritize reasoning power over speed. // Use ModelPreference.FAST when complex logic is not required and latency is a priority. val previewFullConfig = generationConfig { modelConfig = modelConfig { releaseStage = ModelReleaseStage.PREVIEW preference = ModelPreference.FULL } } val geminiNano4BPreviewModel = Generation.getClient(previewFullConfig) val response = geminiNano4BPreviewModel.generateContent(requestWithStructuredOutput) val parsedReceipt: ParsedReceipt? = response.candidates.firstOrNull()?.response Multimodal input Lastly, to help users record audio memos during the trip, let’s build a fully on-device voice notes feature. Using ML Kit’s Speech Recognition API, we’ll enable users to record short voice notes that are automatically transcribed to text. With the transcribed text, we’ll use ML Kit’s Prompt API to identify which trip activity is associated with the recorded voice note, letting users easily recap their trip as they scroll through the trip’s itinerary. The Roman holiday itinerary shows voice note extracts. The ML Kit GenAI Speech Recognition API allows you to transcribe audio content to text fully on-device using two distinct modes. Basic mode uses a traditional on-device speech recognition model and is available on most Android devices with API level 31 and higher. Advanced mode uses Gemini Nano to offer broader language coverage and better quality, and is currently supported on Pixel 10 devices. For our feature we combine the Speech Recognition API with the ML Kit GenAI Prompt API: // implementation("com.google.mlkit:genai-prompt:1.0.0-beta3") // implementation("com.google.mlkit:genai-speech-recognition:1.0.0-alpha1") val tripEvents = ... // Set up speech recognition val speechRecognizerOptions = speechRecognizerOptions { locale = Locale.US preferredMode = SpeechRecognizerOptions.Mode.MODE_ADVANCED } val speechRecognizer: SpeechRecognizer = SpeechRecognition.getClient(speechRecognizerOptions) suspend fun transcribeVoiceNote(recognizer: SpeechRecognizer) { // Display partial text as the user is recording audio var partialTextResponse = "" // Display the full text once user is finished recording audio var transcription = "" val request: SpeechRecognizerRequest = speechRecognizerRequest { audioSource = AudioSource.fromMic() } recognizer.startRecognition(request).collect { response -> when (response) { is SpeechRecognizerResponse.PartialTextResponse -> { partialTextResponse = response.text } is SpeechRecognizerResponse.FinalTextResponse -> { transcription = response.text processAndCategorizeVoiceNote(transcription, tripEvents) } } } } fun processAndCategorizeVoiceNote(transcribedVoiceNote: String, events: List) { val prompt = "Given the voice note $transcribedVoiceNote and the following events for this trip: $events, rewrite this transcription to remove filler words. Then, identify which events from the list this rewritten transcription matches to." // Utilize ML Kit's Prompt API to process voice note and tag it with the relevant trip activities Generation.getClient().generateContent(prompt) } Conclusion Using ML Kit’s GenAI APIs, we were able to take advantage of Gemini Nano to develop fully on-device intelligent features for the JetPacker app, and provide an improved user experience without any additional cloud costs. Check out the full source code for Jetpacker on Github, and watch the video Build Intelligent Android apps with Google’s AI to learn more about how to integrate intelligent features directly into your app using on-device models, cloud-powered reasoning, and the latest agentic frameworks. Learn more Check out the other parts of this blog post series: Part 1: Introduction of the app and a high-level overview. Part 2 (this post!): On-device intelligence. Deep-dive into ML Kit’s GenAI APIs and Gemini Nano to build privacy-first features like itinerary summarization, receipt parsing, and local audio processing. Part 3: Hybrid and cloud reasoning. Explore how to use Firebase AI Logic to ground LLM answers in real-world data like Google Maps and web context. Part 4: System integration. Integrating with the Android intelligence system using AppFunctions. Part 5 (coming soon): In-app agentic workflows. Extend the app with an end-to-end booking assistant powered by A2UI and ADK. Interested in more on Android Development? Follow Android Developers on YouTube or LinkedIn! All code snippets in this blog post follow the following copyright notice: Copyright 2026 Google LLC. SPDX-License-Identifier: Apache-2.0
Read the original at Android Developers

More in Dev