A merchant wrote this on the Shopify Community forum about the store’s AI chat: “Until AI can accurately make changes (ie. actually do things / take action), and not just reply with text and links.” He added that the bot does not “actually solve the issue I have” (Shopify Community).
It is one person’s post. It does describe something a lot of users feel. They ask for a thing to be done, and they get a link to an article about how to do it.
What the data says
Gartner’s numbers on this are the best we found.
- In a Gartner survey of 5,728 customers, fielded in December 2023, only 14% of issues were fully resolved in self-service. In 43% of the failures, customers “couldn’t find content relevant to their issue” (CX Today).
- In a Gartner survey of 3,566 customers, run in February and March 2026, 58% of people who use generative AI had used it to complete a task on their behalf. In B2B it was 74%. A Gartner analyst said customers “increasingly expect AI to help them take action, such as booking an appointment, placing an order” (IT-Online).
- In the same survey, 87% said a company that uses generative AI for service must give access to a human (Gartner).
So people want the task done, and they want a person available if it goes wrong. Those two do not conflict. They describe an assistant that acts, stays inside limits, and hands over cleanly.
Three kinds of assistant
It helps to separate them, because each is good at something.
- It answers. You ask, it replies from your help center. This is the right tool for “what does this plan include?”
- It guides. A tour or a checklist walks the user through the screens. This is the right tool for learning a product you will use every day.
- It does. The user says “give Sam editor access” and the work happens. This is the right tool for a chore the user wants over with.
Most support AI today is the first kind. Intercom’s Fin can take actions, but through APIs, data connectors and MCP, not by clicking your interface (see our comparison, checked against fin.ai). If the action already sits behind an API you can connect, that is a fine route. If it lives in your app’s interface and has no API, someone has to click.
Why doing is harder than answering
A wrong answer wastes a minute. A wrong action changes data.
Two incidents show what people worry about. A support bot at Cursor invented a one-device policy that did not exist. The company’s co-founder wrote “We have no such policy,” and said AI replies in email support would now be labeled (Slashdot). And in July 2025 a founder said a coding agent deleted his production database after he “explicitly told it eleven times in ALL CAPS not to do this” (The Register).
We do not think these are reasons to avoid acting assistants. They are reasons to build four things in from the start.
- Ask before anything irreversible. Sends, deletes and payments wait for the user, and the card says what will happen.
- Check the page afterward. The assistant reads the screen again and says done only when the change shows. If it cannot tell, it says partial or unconfirmed.
- Act as the user. Same session, same permissions. If the user could not do it, the assistant cannot.
- Let the owner switch it off, and let the owner see what users asked for and did not get.
What Skate does, and what it does not
Skate is one script tag. Your user types a task, Skate does it in your real interface, and then checks.
We test it on open-source apps and publish the score. The latest single runs, on 29 September 2026: Docmost 33 of 37 goals, Vikunja 26 of 30, Documenso 27 of 31, Gitea 25 of 30. These are apps we built against; on an app we had never seen, the first run completed 10 of 22. A separate checker reads the saved data to decide if each goal was met. The tasks that pass most reliably are the ordinary ones: create and rename things, change a setting, set a priority or a due date, comment, export a file. We are not at our target yet.
It does not do online stores yet. Questions that need counting across a long list are weak. Our checker flagged a number of those as wrong, and in the Documenso case, three of five flags were correct answers marked wrong by a strict checker. One role restriction case failed too. The benchmark post has the details.
How to tell whether your users need an assistant that does
You do not need us to find out. Take the last 50 support tickets and sort each into one of three piles.
- Questions. The user wants to know something. An answer fixes it.
- Tasks. The user wants something done and the way to do it exists in your app. Someone could do it for them in two minutes.
- Bugs and requests. The app cannot do what they want.
If the task pile is the largest, an assistant that does has something to take off your plate. If the bug pile is largest, fix the product first. No assistant fixes that.
Write down the exact words users used in the task pile. Those are the sentences your assistant has to understand, and they are the sentences worth testing first.
If you want to see how Skate handles your list, try it on your own site with the extension, 10 requests a day, or sign up and add the tag.