Skate is an assistant you add to your web app with one script tag. Your user types a task, like “give Sam editor access to this space”. Skate does it in your real interface, then reads the page again and says done only if the change shows.
That is a claim, and a claim should come with a test. This post says how we test it, what the scores are today, and what the misses look like. We are not at our target yet.
The scores
| App | What it is | Goals completed | False claims |
|---|---|---|---|
| Docmost | docs | 33 of 37 | 0 |
| Vikunja | tasks | 26 of 30 | 0 |
| Documenso | e-signature | 27 of 31 | 2 |
| Gitea | code hosting | 25 of 30 | 1 |
Read those as 89%, 87%, 87% and 83%. They are single runs, from 29 September 2026. Earlier runs in the same month scored 25 of 37, 22 of 30 and 22 of 31 on the first three. A single run is a single draw, so a run on another day could land a goal or two either way. Read the next section before you rely on them.
These are apps we have worked on
This matters more than the scores. Docmost, Vikunja, Documenso and Gitea are apps we used while building Skate, and we fixed failures we found on them. Our own benchmark notes call their results “regression evidence, not evidence about an unfamiliar site.” So the table above tells you where we are on apps we know. It does not tell you what happens on yours.
We have run Skate cold on two apps it had never seen. On Gitea, the first run completed 11 of 30 goals, with 3 false claims. On Vikunja, before it became a familiar app, the first run completed 10 of 22 valid cases (45%), with 0 false claims; that run stopped at case 24 of 34 because of a test-harness fault. After we worked through the Gitea failures, a later run on the same app scored 25 of 30. That 25 is what an app we have tuned on gets. It is not what a new app gets.
So a realistic first-day expectation on your app is closer to 4 tasks in 10 than to 8 in 10. That is why we suggest trying your most common tasks first, and reading the misses, before you switch Skate on for all your users.
How it is graded
We run Skate on self-hosted copies of open-source apps. Because we host them, we can reset them between runs and read their saved data.
Each goal is a thing a real user would type: add a task, change a priority, add a description, give someone access. A separate checker decides if the goal was met. It reads the app’s saved data. It does not read what Skate said. So “Skate says it is done” counts for nothing if the data did not change.
The graders are frozen. When a run fails because the grader is strict, we do not edit the grader. We note it and move on. That keeps us from tuning the test to the score.
What a false claim is
A false claim is Skate saying a task is done when it is not. It is the failure we care about most, because it is the one that makes a user stop trusting the assistant. A task that fails and says so is a bad day. A task that fails and says it worked is a broken promise.
In the September 29 runs, Docmost and Vikunja had none, Documenso had two and Gitea had one. In an earlier Documenso run the checker flagged five, and when we read them, two were real: the assistant typed a new title into a field, never saved it, and said done. The other three were correct answers that the checker marked wrong because the last message on screen was a “Not confirmed” line. That is a fault in how we ground answers and in the checker, and we have not fixed it. Our checker is frozen, so we count it as a known miss and leave it.
After those runs we added a rule so the two real cases are refused. “Now refused” means we replayed the recorded runs and the rule stops them. It does not mean we have a live run with the fix.
We also test the opposite mistake: refusing something that was fine. We replayed our recorded runs through the current false-claim checks. They refused 0 of 1,589 correct cases and 44 of 91 false claims, which is fewer than half. So the checks are safe on what we recorded and not yet complete. This is a replay, not live traffic.
What the misses look like
The misses are not exotic. These come from our last cycle on an open-source online store (EverShop), run four times. We set aside 20 case-runs that hit a password. Of the rest, 27 failed. The top three causes:
- No route to the category (8 case-runs). The search for a category returned nothing. The assistant went back to the home page, searched again, and then either gave up or was stopped by the guard that prevents endless repeats.
- The search missed the product (5). The assistant searched for a named product, did not find it, and gave up. Two never submitted the search and used only the suggestions drop-down.
- A model call timed out (4). One slow call ended the whole task. We now retry a timeout once, because nothing has reached the page at that point.
This is why the website says online stores are not ready. They are not. The tasks that make a store work, like finding the right product among many and going through a cart, are harder than “give Sam access”. We will not show a store demo until we can show a store score.
What we do with a score
Each cycle we take the largest group of failures, find the shared cause, fix it in the runtime with no app-specific text, replay the safety checks to confirm no correct claim is newly refused, and write down what we predicted apart from what we measured. The changelog lists the results by date.
What this means if you run a SaaS app
Skate is strongest today on SaaS apps with settings, records and exports. The tasks users ask support about are often like the ones in these benchmarks, but your app is new to us, so expect a lower first score and a tuning step.
If you want to see it on your app, try it with the extension on your own site: 10 requests a day, nothing to install on the site. Or sign up and add the tag. Try the 5 to 10 tasks your users ask about most, and look at the misses first.