litmusly
PIPELINE ONE PASSDRIVER STAGEHANDSCHEMA 0-5 FINDINGS

How a reagent
reads your preview.

Litmusly is a GitHub App. When a pull request opens, it finds the preview deploy, sends a grounded reagent to walk it in a real browser, and reads the reaction against a strict schema. Here is the whole pass, step by step.

PIPELINE · ONE PASSPR#247
IN
trigger
pull_request · opened / synchronize
01
discover
GitHub Deployments API → preview URL
02
walk
Stagehand on Browserbase · real browser
03
review
Vercel AI Gateway · Gemini · 0-5 schema
04
tag
blocker · friction · nit
05
synthesis
one verdict · Pro and up
OUT
comment
one PR comment · edits itself
BLOCKERFRICTIONNIT→ ONE COMMENT
READSGITHUB DEPLOYMENTSSTAGEHAND · BROWSERBASEVERCEL AI GATEWAYNEVER YOUR SOURCE · ONLY THE RENDERED PREVIEW
THE PIPELINE

From pull request to reaction.

Every review follows the same grounded path. No configuration branches, no per-repo tuning. Five steps, one comment.

STEP 01

Discover the preview deploy

A pull request opens or updates. Litmusly finds its preview through the GitHub Deployments API, so it works with Vercel or any host that reports a deployment. No manual URLs, no dashboard to babysit.

STEP 02

Walk it in a real browser

Each enabled reagent drives a live browser session with Stagehand on Browserbase. It lands on the homepage, reads the above-the-fold value proposition, finds the primary call to action, clicks it, and observes what happens, capturing screenshots along the way.

STEP 03

Review the walk against a strict schema

The walk trace goes to a model through the Vercel AI Gateway. The review is bound to a schema: between 0 and 5 findings, each one specific enough to act on. Nothing free-form, nothing unbounded.

STEP 04

Tag and post one comment

Every finding is tagged blocker, friction, or nit, with a short title and a quote from the walk. It lands as a single PR comment that edits itself in place on the next push. Many PRs deserve zero findings, and get zero.

STEP 05

Synthesize across reagents

On Pro and higher tiers, a synthesis pass folds every reagent's findings into one collective verdict, so you read one reaction to the change instead of parallel opinions.

STEP 02 · THE WALK

A reagent walks the live preview, not the code.

The reagent drives a real browser. It reads what a visitor reads, in the order a visitor reads it, and records the reaction as it goes.

01Lands on the homepage.
02Reads the above-the-fold value proposition.
03Finds and clicks the primary call to action.
04Observes the result, capturing screenshots along the way.
SPECIMEN · LIVE WALKRUN_a8f3
0:00reagent · viewport 1440×900 · screenshot captured
0:02GET / → 200 · above-the-fold in view
0:11"Okay, what does this actually do?"
0:24locate primary CTA · click
0:29"That scrolled me past three sections first."
0:41observe result · screenshot captured
0:52walk complete · trace → review
STEP 03 · THE REVIEW

The review is bound to a schema.

The walk trace is read against a strict shape: between 0 and 5 findings, each tagged, titled, and quoted. The ceiling is what keeps the comment sharp instead of an endless list. Severity reads like a reaction scale.

REACTION SCALE
BLOCKERNITCLEAN
BLOCKER

Something a real user in this reagent's role would hit and be stopped by. Worth fixing before you merge.

FRICTION

Not a hard stop, but a rough edge that slows the user down or muddies the path to the call to action.

NIT

A small, optional polish item. Worth noting, safe to defer.

GROUNDING

A reagent is a written profile.

Not a random sample and not a mood. Every reagent is an explicit profile that shapes how the model reads the same walk. Because the profile is written down, the review stays anchored to what a real person in that role would care about.

INSIDE A REAGENT PROFILE
BACKSTORY

Who the reagent is and what brought them to the page. Sam is a first-timer on desktop. Marcus is on a 390px phone. Elena is watching for anything that leaks.

PRIORITIES

An ordered list of what this reagent cares about most, first to last. The order decides which moments in the walk carry weight.

BLIND TO

An explicit list of what this reagent ignores, so it stops chasing issues outside its lane and the comment stays quiet.

“If I have to pinch-zoom your page, we’re done.” — MARCUS · MOBILE·390PX
THE MACHINERY

Built to run on every push.

No accuracy scores, no benchmarks to quote. The reagent reads what a real person in that role would run into, and stops. What keeps it sustainable is the routing, not a cheaper read.

Routed through the AI Gateway

Every review goes through the Vercel AI Gateway and runs on Gemini. One place to route, meter, and cap spend.

Volume is the tier line, not quality

Model quality is not a paywall. What scales with your plan is how many reviews you run per day, not how carefully each one is read.

The static prefix is cached

The reagent profile and system prefix stay stable across runs, so repeated reviews reuse the same context instead of re-paying for it.

Put a reaction on your
next pull request.

Install the GitHub App, enable the reagents you want, and open a PR. The review shows up where your team already reviews code.

Get startedSee pricing ↗