Briefing · 3 September 2026
We cloned it, ran it, and put both sides online so you can try them. Every screenshot below is from a real run, not a marketing image.
Live demo
Both sides are running for real. Ask them anything. Nothing is charged and no order is ever placed.
Fictional company, fictional catalog. Usage is capped, so if the assistant stops replying it has hit its daily limit rather than broken.
This is what a customer sees. It sits inside the shop rather than beside it, and it can act on the cart, not just talk about it.
The storefront
Nothing exotic. A catalog, a cart, an orders page. The assistant is one tab, not a bolted-on chat bubble.
Turn one
We asked for a gift for a nine year old who likes building sets, under $45. It answered with real product cards pulled from the catalog, not a paragraph of text.
Turn two
It built a comparison table on the spot. The follow up needed no product names, because it held the context.
Turn three
The cart on the right updated for real. The returns answer came from a policy lookup, not from memory of training data.
Under the hood
The panel lists the exact tools each reply called, with timings, and the facts it has learned about this shopper. Nothing appears that a tool did not return.
This is the half most people have not seen. Same technology, pointed at the back office, for the people who run the store.
The portal
Sales, orders, conversion, and a queue of things that need a decision today. The assistant sits on the right.
The ask
It did not do either. It wrote two proposed changes and stopped. Note the line under the button: "Nothing applies until you approve."
The approval
Once approved, stock moved from 3 to 53 and the low stock count dropped from 9 to 8. The second change is still waiting.
Under the hood
Every figure the assistant quoted maps back to a query it ran. A staged change records who proposed it and when.
The agent never changes anything on its own. It prepares the change, shows the before and after, and waits for a person.
That single design decision is what makes this sellable to a cautious client. It is enforced in the code, not asked for in a prompt.
Five minutes of clicking turned up a real mistake, and it is more useful than anything in the marketing.
On a checkout card, the agent wrote that express delivery had been selected, at $9.99, arriving in two days. The card underneath showed standard shipping, free, three to five days, and a total with no fee added.
We read the code. The card was right. The demo store lets the agent look up shipping options but gives it no way to choose one, so it described a choice it had not made.
The agent is allowed to say something wrong. It is not allowed to make something wrong happen.
The card is built by the server from the real cart. The agent cannot touch it. That is the safety model working exactly as designed, caught live.
Anthropic gave away the recipe, not the meal. Connecting it to a real store, with real data, and getting it live is the part they left out. That is our work, and this release creates demand for it.
Anthropic does not publish blueprints for niche ideas. Shopify, Visa, Mastercard, Intuit and Square are named partners. Our clients will start asking about this.
We no longer have to be the only voice saying agentic commerce is real. There is now a reference architecture with a link.
Their rules are enforced in code and written down in one file. Ours live in prompts and rehearsal. A written safety page is something a client security review can actually read.
We sell a shopping assistant. This points at a second product for the same client, aimed at a different buyer inside the business.
Anthropic states plainly that it is unmaintained and accepts no contributions. It is a design reference and a checklist, never a dependency.
Live demo
Both sides are running for real. Ask them anything. Nothing is charged and no order is ever placed.
Fictional company, fictional catalog. Usage is capped, so if the assistant stops replying it has hit its daily limit rather than broken.