How it works

From push to report

What actually happens between a commit landing and evidence arriving on your pull request, and what we need from you, which is very little.

The lifecycle of a run

A change lands: a push to a pull request, a change on a branch you selected, or you asking SuperGorilla for a test in GitHub, the same way you’d ask a teammate. We get your app running, agents use it the way real users would, and a report lands back on the pull request, or on the commit’s check for a branch push: what was tried, what happened, and the evidence.

You configure which of those triggers apply, per repository. A common starting point is every pull request into your main branch.

Getting your app running

Whatever your stack, SuperGorilla figures out how to run your app on its own. There is no staging URL to point at, no preview deploy to wire up, and nothing for you to stand up first. If your app won’t start, the run tells you why, with the logs that explain it.

How it decides what to test

Two sources. The change itself: the agent reads what the change is trying to do, works out which user flows it touches, and tests them the way a customer will. Your critical flows, coming soon: show the agent how to test a flow by recording it with SuperGorilla, walk through it once, and it tests that flow automatically on every change from then on.

Steering and guardrails, in plain language

Steering is optional. When the agent needs a hint, you brief it the way you’d brief a new tester: “Use the seeded account demo@acme.dev. Skip the onboarding tour. The payment form only takes the 4242 test card.” No selectors, no test code, just instructions in English.

Guardrails work the same way, but they outrank everything else the agent is told, including replies on the pull request: pages it stays out of, actions it never takes, data it never touches. “Never place a real order. Don’t email anyone but the seeded accounts. Leave the admin panel alone.” Steering shapes a run; guardrails come first in every run.

What a report contains

Four kinds of evidence: a video of the testing session, screenshots of each step’s outcome, console and network logs captured while it happened, and for anything that broke, steps to reproduce with a severity rating. Enough to judge a finding without reproducing it yourself.

What it remembers

Runs keep what they learn about running your app: how it starts, the test data that works, the quirks. Memory operates at five levels: organisation, project, repository, branch and pull request. It knows this codebase’s gotchas and what already passed or broke on this very PR, so later runs pick up where earlier ones left off.

What it can't do yet

Honest limits, current as of now: web applications only, with mobile and desktop on the roadmap. Flows that require a real phone, real payment rails, or a human on the other end (live chat, phone verification) need test-mode equivalents. And like any tester, it can miss things; reports are evidence, not a guarantee.

Send the gorillas in.

Get started

Free to start, no card needed. Your first organisation gets $10 of credits.