Back

Scarecrow

A usability test that needs no recruiting. Pick a persona and a goal, and it drives your site in a real browser, then reports the friction it hit.

Role
Designer and sole builder
Year
2026
Domain
Developer tooling · LLM agents
Status
Public on GitHub, Apache 2.0

01Context

I am the sole designer across seven products. The answer to “has anyone tested this flow?” is usually no.

The method is well understood. What costs is the logistics: recruiting five people, scheduling them, sitting with them and writing it up takes about a week, and there is rarely a week between a build and a demo. So the test gets skipped and the client sees the flow first.

02The problem

The gap I wanted to close sits before user research: you have built a flow, you suspect it is confusing, and every way of checking costs about a week.

What you want at that moment is a stranger: someone who has not sat in the planning meetings, does not know what the button is supposed to do, and gets stuck where a real person gets stuck. What makes them useful is ignorance of the design, and the designer cannot get it back.

03How it works

Persona + goalOpen the pageNumber the elementsLook, decide, actFindings report

You give it a URL, one of the personas below, and a goal in plain language: sign up for a trial, find pricing. It opens the site in a real browser, visible by default so you can watch it work, and then loops: number every control on the page, screenshot it, decide the next action in character, click or type or scroll, repeat. It stops when the goal is met or the step budget runs out, then writes up what got in the way. Each finding carries a severity and a suggested fix.

The numbering is the part that makes the loop work at all. Before each decision the tool walks the page for everything visible and interactive, stamps a number onto each one, and takes the shot with those numbers showing. The model receives the picture and a list reading [#3] button: Sign up, and answers with one action and one number.

One step of the loop. The agent answers with an action and a number; the numbers are stamped onto the page it was shown.
Two crops from one run. The first is the terminal, where the agent works: a header reading Margaret, cautious first-timer, novice, with the goal, the URL and the provider beneath it. Then step one of eight, click hash three, followed by Clicked hash three, Sign In. Then step two of eight, type margaret.t.smith@example.com into hash two, followed by the confirmation that it was typed. A line reads running, Esc to stop. The second crop is the page the agent was shown, a sign-in form, with a small numbered badge stamped on every interactive control: one on Continue with Google, two on the email field, three on the password field, four on the reveal-password control, five on Sign In and six on Create an account. The email field already contains the address the agent typed.

Everything else is organised around the persona. A generic “test my site” agent produces generic notes, because what counts as friction depends on who is looking. Each persona is written in second person with its own priorities, so each one reports different problems from the same page.

The skeptic, thinking in character before touching anything.
A crop of the terminal during a run with the skeptic persona. It reads: Priya, 41, a product manager who was auto-charged 120 dollars seven years ago, is browsing the vithos.in website. She needs to sign in and upload a medical report. She starts by scrolling to the footer to check for trust signals like privacy policy, terms of service, and company address. She will then scan the page for any dark patterns such as pre-checked boxes or urgency claims.
The five built-in personas. Custom personas are added through a wizard.
A table of the five built-in personas and what each one notices. Novice, Margaret, a cautious first-timer, who notices confusing jargon, buttons that do not say what they do, and missing help. Power, Dev, an impatient power user, who notices slow workflows, hidden pricing and docs, and cluttered UI. Skeptic, Priya, privacy-conscious, who notices hidden fees, dark patterns, missing trust signals and pre-checked opt-ins. Rushed, Marco, mobile-native and thumb-driven, who notices desktop-only elements, tiny tap targets and anything slow on a phone. Access, Sam, low vision and scanning for barriers, who notices low contrast, missing labels and keyboard traps.

04What I decided, and why

The model judges from the picture and acts through a number. A user cannot read markup: if a label is present but set in faint grey at 11px below the fold, the accessibility tree reports a working flow while the person is lost. So the screenshot is what the model reasons about. But a model that describes where it wants to click is a model that can miss, so every visible control is numbered before the shot is taken and the reply has to name one of those numbers. The model sees the same numbers it has to answer with, so a reply can be checked against the list before anything touches the page.

Rejected Letting the model return a CSS selector or a pair of coordinates. Free-form, and nothing to validate it against until it has already acted.

The vocabulary is closed, and one word in it is giving up. The model may answer with click, type, scroll, scroll up, wait, done or give up, and nothing else. Give up is in the list on purpose: a real person who cannot find the button abandons the task, and that is the most useful finding a run can produce. A reply the parser cannot read resolves to the same thing, so a confused model fails the way a confused user does rather than crashing the run.

Severity is rated by the persona, not in general. Every finding carries the issue, a severity, the usability principle it violates and a concrete fix. The severity is scored by how badly it hurt this user, so the same missing label is high for Sam and barely worth noting for Dev. Without that the personas would be voice only, five ways of narrating one identical list.

Two speeds, because they answer different questions. A full walkthrough is a session: several steps, a transcript and a findings list. A glance is one screen and one reaction. I separated them because the quick question came up far more often than the slow one, and paying session cost for it meant skipping the check entirely.

Rejected One mode with a step budget.

A local model provider. It runs against hosted models and against a local daemon that costs nothing and sends nothing off the machine. The tool screenshots an unreleased product in full on every step, so the local path lets someone try it before deciding what they are willing to send.

Rejected Hosted providers only.

Hosted and local in one list, with the local sizes stated because that is what the choice costs you.
The model picker in the terminal, headed Choose a model with a filter field. Under a heading reading gemini, key set, three hosted models: gemini-2.5-flash marked default, gemini-2.5-pro, and gemini-2.0-flash. Under a heading reading ollama, local, eight local models each with its download size: qwen3-vl marked default at about 8.5 gigabytes, qwen3 at 2.5, llama3.2-vision at 7.9, llama3.2 at 2.0, llava at 4.5, moondream at 1.7 and gemma3 at 2.3, with one more below. A footer reads type to filter, arrows to move, enter to select, escape to cancel.

Findings come out in four shapes. A readable report for a person, JSON for anything downstream, an HTML file for sending to somebody, and JUnit XML so a run can fail a build. The XML is what lets it run in CI instead of on request.

Rejected A readable report and nothing else.

One run, as it comes back. Recreated session with invented findings; the output shape is the real one.
Recreated Scarecrow terminal session. The command runs the privacy-conscious skeptic persona against a pricing page with the goal of starting a trial. The run did not meet its goal and stopped at step three of eight. Three steps are logged: landing on the pricing page looking for a total, opening the plan comparison and finding no annual price, and reaching the trial signup where a card form appears with no cost stated. Three findings follow, each with a severity and a suggested fix. High: card details are requested before any price is shown, fixed by stating the charge and the date above the card field. Medium: the start-trial button does not say what happens at day fourteen, fixed by naming the renewal in the supporting line. Low: the newsletter opt-in is pre-checked at signup, fixed by shipping it unchecked. The run closes by reporting three findings over nine model calls in forty-one seconds, written to a transcript, JSON and JUnit XML.

The cost is bounded and stated. One vision call per step plus a debrief, so about nine calls for an eight-step run. The step budget is a flag with a low default and a hard ceiling, and the agent has no say in it.

Rejected Letting the agent decide when it is done. Better runs, unbounded cost.

05Meeting the tool

A command-line tool usually makes you learn its whole surface before you can run it once. I wanted the opposite: type as little as you know, and it asks for the rest.

The interface is chosen by how much you already typed, not by a mode flag.
A table pairing five ways of invoking the tool with the interface each one produces. The command crow on its own opens the full terminal interface. Crow with a persona flag asks for the URL and the goal, the two things it is missing. Crow with only a URL asks who is testing and what they are trying to do. Crow with a URL, a persona and a task has nothing missing, so it shows the run back to you and waits for confirmation. Adding the yes flag skips that confirmation, and it is the flag continuous integration uses.

There is no mode flag anywhere in that table. The interface is inferred from how complete the invocation already is, so the tool never asks for something you have given it and never assumes something you have not.

The goal is already typed, so the only thing it asks for is the persona.
A picker in the terminal headed Choose a persona, listing five: novice, Margaret, a cautious first-timer, currently selected; power, Dev, an impatient power user; skeptic, Priya, privacy-conscious; rushed, Marco, mobile-native and thumb-driven; and access, Sam, low vision, scanning for barriers. A footer reads arrows to move, enter to select, escape to cancel.

A finished command still stops and shows you the run first. Typos are cheap in most tools and expensive in this one: a mistyped URL spends real minutes and real API credit before anything tells you it was wrong. So the last step before a run is the run described back to you. --yes skips it, which is what automation wants and what a person almost never does.

The last thing before a run: the run, described back to you.
A confirmation panel in the terminal headed CONFIRM YOUR SESSION. It lists the goal, sign in to vithos.in and upload a medical report; the URL, https://vithos.in; the persona, novice; the steps, eight; the mode, walkthrough; and the device, desktop. Below it a line reads: press Enter to start the run, or N to go home.

Rejected Running immediately once the arguments parse, the way most CLIs do.

06Telling whether it worked

The model's verdict and the fact are kept apart. At the end of a run the model reports whether this persona would have completed the task: yes, unsure, or no, labelled as the model's opinion. Given a piece of text or a fragment of a URL, --success checks for it independently, so one answer in the report is not up to the model.

Runs are comparable, which is what makes it a loop. crow diff lists past runs with a score, lower being better, and compares any two: how many findings each produced, how the score moved, and whether that counts as an improvement. One run only ever describes a page. Two runs either side of a change are evidence about the change, which is the only claim this kind of tool can honestly make.

07What the bugs changed

One revision fixed three bugs, and two of them changed a rule rather than a line. A mistyped step count could bill nine hundred model calls with nothing standing in front of it, so the tool caps its own spend now and stops trusting the number it was handed. And a dead domain produced a confident report about Chrome's own error page and exited clean, so a run that cannot load the site now fails: a tool that cannot see must not answer. The third was an upload flag with two parse paths and a guard on only one of them. The mechanics of all three are in the repository.

The budget running out, and the tool asking rather than spending.
A prompt in the terminal reading: step eight of eight, ran out of steps, add five more? Two options are offered, Add 5 steps and No, finish, with the first selected. A footer reads: arrows to navigate, enter to confirm, escape to cancel.

08What it does not do

A synthetic tester is easy to oversell, and a team that thinks it has tested something will stop looking. This section is in the README for the same reason it is here.

The docs say it is a first pass and a hypothesis generator, not real data. Synthetic testers are over-agreeable, cannot be bored, and sometimes misread a page. It catches obvious copy, flow and layout problems, and stops well short of evidence about your users.

Two limitations are specific to this tool. The agent reads page content and acts on it, so a hostile page can try to steer it; every run gets a fresh browser with no cookies, uploads attach only to a dialog the agent itself opened, and the numbered answer is checked before it is used, which narrows the problem without closing it. And to keep a run inside one tab it rewrites links that would open a new one, so it never finds friction caused by a popup.

09Where it is

Node with Playwright driving Chromium, a terminal interface built in Ink and React, axe-core behind the optional accessibility audit, and the Anthropic and OpenAI SDKs covering four providers between them. Eleven test files and about 1,200 lines of them, with lint and the suite gating publication. The command is crow, and crow doctor checks the runtime, the browser and the key before you waste a run finding out.

  • A usability pass with nobody to recruit
  • 5 personas, 4 model providers, one needing no key
  • Public source, Apache 2.0

github.com/abhyuday1602/scarecrow

No usage numbers: the source is public, which is not the same as adopted.

Other case studies

Back to work