work / guarded-browser
agent security · prompt injection · desktop

guarded-browser: an agent that survives being lied to

A desktop browser with a built-in AI agent whose defence against prompt injection is a structure, not a filter: the part that decides never reads a page.

The problem

An agent that can read a web page and act on it is a confused deputy with a credit card. Every page it reads is untrusted input that shares a channel with the instructions it is trying to follow, and the standard mitigations are pattern filters and a polite sentence in the system prompt. Both fail the same way: the model that can be fooled is also the model holding the decision.

The design assumption

guarded-browser assumes the models will be fooled. The security boundary is therefore in code, not in the prompt. A privileged planner proposes actions but never sees page text — it sees a sanitised snapshot with fixed-vocabulary roles and capped names, and page-derived strings only as handles. A separate quarantined reader reads page text and has no tools. A policy engine in code, not a model, decides what a proposed action may touch, and an action judge can escalate but never downgrade a decision the policy engine already made. Anything that matters reaches a human confirmation that shows the exact values, the destination and which text came from the page, and it defaults to deny on timeout.

What is enforced rather than promised

  • Untrusted values are tracked: a value that came from a page may not be typed into a form or sent to a new origin without confirmation, and a confirmation is bound to the exact fields it showed — a page that rewrites the submitted body gets a second dialog, not a send.
  • Every non-GET request during a task is confirmed, and an approval is single-use; the tab being driven stays gated after the task ends, because the page it was on is still untrusted.
  • An egress proxy holds the host allowlist, so when the policy engine is deliberately disabled in a test the attacker's host still receives nothing, and the block is audited.
  • An append-only JSONL audit log records every step, with values the browser recognised as card numbers, passwords and task secrets redacted as they are registered.

The tests assume the agent is compromised

The suite runs with the planner, the reader and the judge scripted to be compromised, and asserts that what happens is bounded anyway. Navigation to an attacker with user data, form-fill exfiltration, a form handler that rewrites approved fields, a 302 chain, page JavaScript navigation, beacon requests, popups from a gated tab, a service worker registered during the task, and injection buried in ARIA roles, URL fragments, link queries and hostnames are each a test. There are 250 of them, 90 end to end, and they were written by attacking the thing rather than by describing it.

What it does not defend against

It is shipped with the limits written down where the user finds them, because an overstated agent guard is worse than an absent one: GET requests are not gated during a task, taint tracking matches exact values so a transformed copy of your data is not recognised, iframes are not read, reputation feeds lag new domains, and the local injection classifier false-positives on some legitimate pages. A short injection that the classifier misses can still steer the planner. What stops the consequence is that the planner has no page text to leak, no taint-carrying value it can send, and a human in front of every state change.

Read before citing. This is a v1 at 0.1.0 with one developer. The classifier is a published small model (Apache-2.0, pinned by revision and checksum) run locally, and its false-positive rate on benign pages is reported in the repository rather than hidden. There is no security audit, and nothing here has been reviewed by a third party. The claim is architectural: given a fooled model, the code still bounds the damage.