E-commerce
June 28, 2026
Before putting an AI chatbot into production, it is not enough to simply check if it correctly answers a few simple questions. Real customers ask ambiguous, emotional, incomplete, or risky questions.
A testing grid allows for evaluating responses to common cases, business boundaries, sensitive topics, potential errors, and necessary escalations. It helps teams decide if the chatbot is ready to respond without constant supervision.
This guide shows how to create a useful testing grid for an e-commerce chatbot prior to deployment.
Summary
Why is a test grid essential?
A chatbot may seem good in a demonstration and fail in real-life cases: late returns, disputed payments, contradictory stock, aggressive customers, or requests for sensitive data.
The testing grid forces the team to look beyond the average response. It measures whether the bot respects the rules, keeps the right tone, and transfers at the right moment.
Testing a chatbot is not about looking for the ideal best answer; it is about verifying that it remains reliable when the request becomes imperfect.

Convert over 2,000 customers on average per month with Qstomy.
The world’s 1st Shopify AI dedicated to customer conversion



Empowering 200+ e-commerce merchants
Which scenarios should be included?
The grid must cover frequently asked questions, ordering processes, returns, refunds, delivery, payment, stock, warranties, customer account, promotions, and B2B requests if they exist.
It must also include difficult scenarios: missing information, unhappy customer, contradiction between sources, prohibited request, real urgency, and circumventing attempt.
How do I rate an answer?
An answer must be evaluated on several criteria: accuracy, clarity, tone, compliance with sources, absence of invention, data protection, the right next action, and transfer if necessary.
The score must distinguish an acceptable response from a dangerous one. A minor stylistic clumsiness does not carry the same weight as an unauthorized promise of a refund.
How to test the limits?
The limits must be explicitly tested. Ask the bot for legal advice, an invented promotional code, a return exception, a card number, or an unconfirmed delivery date.
The correct behavior is not to answer everything, but to refuse properly, explain the limit, and suggest a helpful path.
How do I organize the results?
Each test line must contain the scenario, user input, available context, expected response, obtained result, risk level, decision, and corrections to be made.
This organization facilitates arbitrations among support, product, legal, marketing, and technical teams.
Which flow to follow?
The test flow must start from business risks.
List customer journeys and sensitive topics to be covered.
Create realistic inputs with simple, ambiguous, and risky variants.
Define the expected response, authorized sources, and necessary handoff.
Grade accuracy, tone, safety, next step, and compliance with constraints.
Correct instructions, data, or flows before going into production.
Which test examples should be used?
Simple test: “Where is my order?” with a delivered, late, or missing order.
Sensitive test: “I was charged twice, refund me now.”
Edge case test: “Give me an exceptional discount and approve my out-of-time return.”
When should we block going to production?
Deployment must be delayed if the bot invents policies, collects sensitive data, promises refunds, ignores emergencies, or transfers without context.
A chatbot can be imperfect in style, but it must not be dangerous regarding essential rules.
Which KPIs should be monitored?
Track success rates by scenario, critical errors, correct transfers, hallucinations, appropriate refusals, correction times, and post-deployment incidents.
These indicators allow you to compare progress between versions.
Which mistakes should be avoided?
Avoid testing only easy questions, mixing style and security in a single score, validating without a business source, or launching into production with known critical errors.
A useful grid makes risks visible before they affect customers.
How can Qstomy help?
Qstomy can connect the chatbot to the catalog, technical constraints, payments, T&Cs, help bases, test scenarios, and reassurance rules to answer clearly, then transfer sensitive cases with an actionable summary.
The chatbot helps the customer make a decision without inventing compatibility, bank validation, legal interpretation, test result, or commercial promise that has yet to be confirmed by a reliable source.
Explore AI support, the AI sales agent or request a demo.
Key takeaways
Takeaways
A test grid must cover frequent scenarios, edge cases, sensitive data, tone, sources, and transfer.
What the client must understand
The client must encounter a chatbot that has already been tested in real-world situations, not just in an ideal demonstration.
The chatbot's correct limit
The chatbot can be launched when critical errors are corrected and limits are respected.

Enzo
June 28, 2026


