Note #35 •

An AI Agent Found 24 Android Bugs. Sorting Them Was the Hard Part.

An agent audited Android apps and reported 24 bugs

GitHub Security Lab built an open source framework called the Taskflow Agent so security researchers can package the AI prompts they find effective and share them. A taskflow is that package, a small workflow rather than one enormous instruction. One of the team's researchers wrote Android auditing taskflows, ran them against shipping applications, and has reported 24 vulnerabilities so far, a handful of them critical.

What matters is which parts of the job the model did well and which parts it did badly, because the writeup is direct about both.

What the taskflows change

Two of the taskflows carry most of the result.

The first is new and it splits entry points into mobile and non-mobile. Entry points are the places attacker-controlled data can reach, and a repository often holds a mobile app alongside a web server or a desktop tool. Separating them lets the model reason about the right attack surface instead of a blended one.

The second is an edit to an existing classification step. It now carries an explicit list of vulnerability classes and asks the model to consider them against each entry point and component. Mobile bug classes are less widely known and models are non-deterministic, so the prompt names them. If an earlier step found an intent-based entry point, the list includes confused deputy and insecure broadcasts.

The writeup also reports that a strict prompt and a broad prompt across multiple runs beat either prompt alone. The strict one catches the obvious classes. The broad one leaves the model room to notice something nobody listed.

Two of the findings

OsmAnd is a navigation app with more than 10 million downloads on Android. It exports an activity called MapActivity, which handles settings imports and reads four intent extras to control the import. Those extras are named settings_version, silent_import, replace, and export_type_list_key, and they were meant to arrive from an in-process service. Android does not restrict the extras an external caller can attach to an intent aimed at an exported activity, so any installed app can set them.

With those flags under control, an app with no permissions can import settings into OsmAnd without a notification and without a confirmation prompt. From there the attacker rewrites the map tile URL template to point at their own server and serves the real tiles back, so the map still renders. What comes back to the attacker is the x and y coordinate of every tile the user had loaded. The same bug exposes the origin and destination of planned routes. The user sees nothing change.

The second finding is in the Wikipedia Android app, which registers the wikipedia:// deeplink. Its handler checks the URI authority with an endsWith call against the base domain, so a host like evil-wikipedia.org passes. The app then opens an attacker-controlled page in a WebView and runs the attacker's JavaScript. The same suffix comparison appears a second time in the app's cookie handling, where it leaks the wikipedia.org cookie list, including a token valid across Wikimedia projects. Chained together, the two make an account takeover.

Where the model still needs a human

The writeup does not pretend otherwise. The model finds low severity bugs reliably and reports them even when explicitly told not to. Severity estimates are frequently wrong, because the factors that lower severity are hard to see without running the code. A path traversal confined to external storage is a weak finding. If the app's internal storage takes priority over the attacker-written external data, there is no vulnerability at all, and the model can report one anyway. Every finding should be reviewed by a researcher who knows mobile.

The suggested remedy is to make the model write a proof of concept and run it. That costs more time and more tokens, and the writeup notes the model is strong here. It understood the behavior of security-relevant APIs across languages without access to the language source, and most of the proof-of-concept code it produced needed little correction.

What to copy

Structure the prompt around the target. Splitting entry points by platform and enumerating vulnerability classes per component is portable to any codebase, and it is what separates a model that audits from a model that summarizes.

Require a working proof of concept. The writeup's own position is that a security researcher has to review each finding, and the cheapest way to make that review real is to ask for exploit code rather than a plausible paragraph.

Budget for it. A GitHub Copilot license is required, the prompts consume premium model requests, a run on a medium sized repository can take one to two hours, and the tool calls add up quickly.

The framework's reason for existing is that a taskflow is a file. One that found a confused deputy bug in one Android app can be run against another.