Android Developers Blog: Inside Android Skills


Posted by Jose Alcérreca, Developer Relations Engineer, Android Developer Relations

We released the official Android Skills in April, and the response surpassed all our expectations. In this blog post, I’ll address some of the feedback we received, explaining the philosophy and methodology behind the project. Hopefully, this will also help you understand what happens behind the scenes when you install and use skills, allowing you to make better use of tokens and your own time.

Why are there so few official skills?

Currently, we only consider new skills when there’s a verifiable knowledge gap in state-of-the-art (SOTA) models. Put simply: you don’t need to teach the model what it already knows. (Though there are a few exceptions—read on!)

We’ve released around 20 official skills so far, and they intentionally target highly specific, fast-moving areas that standard models aren’t fully grounded on yet—things like AGP 9, Navigation 3, advanced Camera APIs, and Perfetto SQL.

What about core, more general, skills? Every installed skill injects 100–200 tokens into the baseline context of every task you start. If that skill actually activates, that count can quickly jump into the thousands. In most cases, hoarding basic skills is both counterproductive and expensive. Before installing a skill for writing basic Kotlin or Compose, consider if your LLM of choice really needs it, or if it knows those topics well enough already.

Evaluating skills

Before their release, each skill is tested against a comprehensive set of evals that prove that the skill delivers clear value. These evals should pass when the skill is active, and fail otherwise. Evals are to skills what integration tests are to code.

timeout_s: 1200
repository:
  url: [redacted - internal git repo]
  working_dir: wear_compose_m3_empty_app
category_ids:
  - wear
prompt: |-
  Add a horizontal pager to MainActivity.kt. Have three pages in the pager. Each page should contain
  the text "Page 1", "Page 2", and "Page 3" respectively in the center of the screen.
commands:
  build:
    - ./gradlew assembleDebug
acceptance_criteria:
  project_builds: true
  llm_diff_judge:
    - Must use `HorizontalPagerScaffold`.
    - Each page should use `AnimatedPage` to wrap a `ScreenScaffold`.

Example eval that checks the correct implementation of a horizontal pager on a wear app

At a minimum, we test the skill in Android Studio using the latest Gemini Flash model. Depending on the skill, we also ensure compatibility with other models such as Gemini Pro and other agents such as Antigravity, and third-party systems.

All of the evals run with access to the Knowledge Base, so if the information is in the documentation, and models decide to search for it, we don’t publish a skill for it.

Using the Android Knowledge Base (Android Studio or Android CLI)

If you develop Android apps, you should always use the Android Knowledge Base to have access to the official documentation. If you use the agent in Android Studio, it’s already available as a tool, but if you use another agent, install Android CLI. Among other things, it contains the docs command, which gives your agent access to the official Android documentation. Having a single tool is much more efficient than installing hundreds of skills.

If your model is acting overconfident, and you want it to consult the documentation more often, a very common way to motivate it is to add “Always consult the official Android documentation when dealing with Android APIs” to your AGENTS.md file or equivalent. Of course, you can also force this by asking the agent to check the documentation directly in your prompts.

Why are pull requests disabled?

Because our evaluation framework depends on internal infrastructure that cannot be open-sourced, we are unable to accept direct pull requests for new skills—without this infrastructure, we would have no way to re-evaluate incoming PR changes. However, we actively monitor community feedback. If you want to report a bug, suggest an optimization, or request a new official skill, please file an issue!

When do core or basic skills make sense?

While SOTA models generally don’t need basic skills, there are some scenarios where enabling core or community-built skills adds real value. For example:

  • You’re using vague prompts: Skills amplify your intent. If you give a loose prompt like “add animations to this screen,” a specific Compose animation skill can inspire the model, pushing it toward modern APIs or screenshot testing patterns it might not have otherwise considered.
  • You want to use smaller, cheaper models: Frontier LLMs are expensive. If you are offloading routine tasks to smaller open-weight models like Gemma 4, enabling basic skills fills the knowledge gaps that smaller parameters miss.
  • You’re refactoring or reviewing legacy code: Models excel at generating code that works, but when editing old codebases, they often prioritize staying consistent with the surrounding legacy patterns over rewriting things with modern accuracy. A specialized reviewer agent equipped with core skills can help break that habit.
  • You deviate from the norm: LLMs love the standard “Google way” of architecting Android apps. If your team uses a highly customized view-layer architecture, the model will struggle to stay aligned. A custom skill explicitly describing your architecture goes a long way.

Where can I find core skills?

The Android community has your back. Chris Banes has a comprehensive collection of skills for Compose and Kotlin, Ivan Morgillo published a skill that audits Compose projects, and Jaewoong Eum created two on testing and performance.

Always download skills from reputable sources! I personally wouldn’t trust repositories containing dozens or hundreds of Android skills as they’re probably AI-generated and untested, and they could even contain malicious or biased instructions. Also, don’t install general software engineering skills blindly; a lot of them are tailored for web development.

Goal: deprecation

Loosely paraphrasing Karpathy: Skills of today will be in the models of tomorrow. As SOTA models keep improving, we expect skills to be obsolete, especially those built around new APIs. To figure out when to retire them, we run our evals when new models drop. If they pass, we’ll keep them around for a few months until most users have transitioned over.



Source link

  • Related Posts

    Delivering safer, age-appropriate experiences on Google Play

    Providing a safe online experience and protecting users from harm is a top priority at Google Play. We take this responsibility seriously and have been investing continuously to offer baseline…

    Celebrating 5 years of Jetpack Compose

    Posted by Rebecca Franks, Developer Relations Engineer, Nick Butcher, Product Manager, Loryn Hairston, Product Marketing Manager, Android Today, we officially celebrate five years since the release of Jetpack Compose 1.0.…

    Leave a Reply

    Your email address will not be published. Required fields are marked *