Skip to content
KameleonAI
All posts
Product14 July 2026

"What if it says the wrong thing?"

The most common objection to putting AI in front of your customers is also the most reasonable one. The honest answer, including the parts that are still risks.

Every merchant asks a version of this, usually within the first two minutes.

What if it invents a product? What if it promises next-day shipping I don't offer? What if it makes up a discount code? What if it's rude to someone at 3am and I find out from a review?

These are not naive objections. They're the correct ones. Anyone who has watched a language model confidently state something false understands the risk straight away, and the risk is real. Your storefront agent talks to your customers in your voice. One confident lie can cost you a customer and a chunk of trust that took years to build.

Rather than reassure you, it's more useful to pull the failure modes apart, because each one needs a different thing.

The fear is really four different fears

Hallucination is the one people name first: it will invent things that don't exist. Over-promising is a separate problem, and it covers shipping times, return windows, stock, a discount that isn't running. There the agent isn't inventing a product, it's making a commitment on your behalf. Then there's the aggressive salesperson worry, since nobody wants their store to feel like a timeshare presentation.

The quietest one is that you won't know what it did, and in our experience that's the fear that actually blocks the decision. Not that something will go wrong, but that something could go wrong for weeks without you finding out.

They need different answers because they have different causes.


Where the answers come from

The first two fears get the same answer, and it isn't a better prompt.

An agent can only invent products if it's allowed to answer from memory. So it doesn't have one. It reads the real product record, live price, live stock, the actual variants, the policy text you wrote, and it answers from that. If a product isn't in your catalog, there's nothing for it to describe. If something's out of stock, it knows the way your own cart knows.

Discounts work the same way with one extra step. Every synced discount is hidden from the agent by default. You turn one on, per discount, with a confirmation. Your staff codes and your friends-and-family codes stay invisible. A seasonal promo is visible for exactly the window you choose. The agent cannot offer a discount you haven't handed it, because it doesn't have one.

When it genuinely doesn't know something, the right behaviour is to say so and point the shopper at your team. An agent that says "I don't know, but here's how to reach the people who do" is doing its job.

That's the honest ceiling on this. Grounding removes the category of error where the agent makes things up about your store. It does not make the agent omniscient, and it can still misjudge a nuance in a policy you wrote ambiguously. Which is why the fourth fear matters most.


Pushiness is a settings problem

You write the sales instructions. Brand voice, what to emphasise, what it must never say, how hard to push, when to back off. If your brand is understated, it's understated. If you never want it to create urgency, it doesn't.

There's a broader point underneath this one worth sitting with. Right now, how many of your visitors get any personal attention? For almost every online store, the honest answer is none. A static page shows the same thing to a first-time visitor and a returning customer, to someone who knows exactly what they want and someone who has no idea. That's an impersonal experience, even though we've all agreed to call it normal.

So the useful comparison is the agent against the nothing that's there now, not the agent against a good human salesperson.


The fourth fear is the real one

You will know what it did. That's what we built the product around, and it's the part most worth checking before you trust any tool like this.

Every conversation is stored and readable. Not a summary. The actual transcript, what the shopper said and what the agent said back, with the outcome attached. You can filter by how the conversation ended, by whether it led to a cart, by whether the shopper seemed happy.

Every number in the dashboard is a link, which matters more. When it says a conversation put value in a cart, you click through and read the conversation that did it. When it says shoppers keep asking about sensitive skin, you open the forty chats where they asked. Every figure in the product traces back to the source text it came from.

And the counting is deliberately pessimistic. An add-to-cart only counts as ours when the agent's own action put it there. When attribution is uncertain, it counts as organic, against us. Orders join to conversations through Shopify's own cart token rather than a guess from timing. The number you see is a floor, and the real contribution is higher than what we report.

That decision costs us on paper every single month. We made it because a number you can't check is worth nothing, and a number that flatters us is worth less than nothing the first time you check it and find it inflated.


How to evaluate this on any tool, including ours

Don't take the above on faith. Test it the way a suspicious person would.

Install it on your own store during a trial and then try to break it. Ask about a product you don't sell and see whether it invents one or tells you the truth. Ask a genuinely complex question, the kind your customers actually ask, with the messy real-world context attached. Ask in a second language if you sell across markets. Ask about a policy edge case you know is ambiguous on your site, and see whether it guesses or defers.

Then ask for a discount you never enabled.

Then go into the dashboard and try to find the conversation you just had. If you can't reconstruct exactly what happened from the merchant side, that's your answer, regardless of how well it performed.

The risk of putting AI in front of your customers is real, and anybody who tells you it isn't is selling something. What you should insist on is being able to watch it. If the grounding slips somewhere, or the tone comes out wrong, or something gets promised that shouldn't have been, you want to find that out from your own dashboard in the same week rather than from a review six weeks later.