E-commerce
June 28, 2026
Many shops measure support by ticket volume, response time, or bot deflection rate. Useful, but not enough. A bot that deflects 70% of requests with inaccurate answers increases returns, negative reviews, and silent churn.
The quality of support answers is broken down into four dimensions: accuracy, tone, resolution, satisfaction. Lorikeet reminds us in 2026 that the resolution rate matters more than deflection alone: avoiding a ticket does not mean solving the problem (Lorikeet, 2026 support metrics).
This guide #116 covers QA scorecards, quality KPIs, and improvement loops. Distinct from chatbot KPIs (#11) (adoption and ROI): here we focus on the quality of the response content, both bot and agents.
Summary
Why measure the quality of support responses?
Measuring e-commerce support answer quality goes beyond volume and speed.
Limits of volume-only KPIs
Closed tickets: does not tell you if the customer is satisfied
First response time: fast but wrong = worse than slow but correct
Bot deflection: hides hallucinations or outdated macros
Global CSAT: average hides weak agents or problematic intents
Cost of poor quality
Avoidable returns (wrong size info), negative post-chat reviews, repeat contacts on the same topic, silent churn, double work for agents to correct bot errors. Bookbag notes that a fast bot with a poor response drops the CSAT (Bookbag, 2025 e-commerce support metrics).
When to start
Starting at 500 bot conversations or 200 agent tickets/month. Before scaling ads or expanding markets. Lagging signals: increase in returns, "lie" reviews, WISMO repeat contacts, post-bot rage escalation.

Convert over 2,000 customers on average per month with Qstomy.
The world’s 1st Shopify AI dedicated to customer conversion



Empowering 200+ e-commerce merchants
What are the 4 pillars of response quality?
The 4 pillars of support response quality structure any evaluation.
1. Accuracy
Factually correct answer according to policy and data. OK: exact delivery time, up-to-date return. KO: promising 48h when it takes 5 days, announcing an expired promo code as valid.
2. Tone
Brand voice, empathy, clarity. OK: empathy + solution + CTA. KO: robotic, condescending, too colloquial/off-brand. Both bot and agent share the same tonal register.
3. Resolution
Problem solved without unnecessary back-and-forth. FCR (First Contact Resolution), justified escalation, no dangling "I will get back to you" without follow-up.
4. Satisfaction
Post-interaction CSAT, verbatims, proxies: recontact rate, pre-purchase post-chat return.
Priorities by intent
WISMO: accuracy is key. Angry customer: tone is the priority. Complex dispute: resolution + satisfaction.
How to measure the accuracy of the responses?
Measuring accuracy protects the brand and reduces returns.
Sources of truth
Up-to-date official return/delivery policy page
Shopify catalog stock sync and variants
Versioned Gorgias macros with review date
Bot corpus aligned with help center policies and product sheets
Promotions: centralized marketing start/end dates
Accuracy audit grid (1-5)
5: 100% accurate, sources cited if needed
3: partially accurate, minor correction required
1: incorrect, risk of dispute or return
Common e-commerce errors
Return window 14 days vs 30 days depending on the market. Bot says available, variant OOS at checkout. International vs domestic delivery times mixed up. Non-cumulative promo code announced as cumulative.
Formula and process
Factual error rate = conversations containing an error / audited sample × 100. Bot target < 3%, agents < 1%. Error detected: correct macro/bot within 24 hours, contact customer if purchase is impacted. Classify cause: stale corpus, poorly routed intent, obsolete agent macro.
How to evaluate brand tone and voice?
Measuring tone ensures brand consistency across all channels.
Support tone guidelines
Register: informal vs formal address, formal vs warm
Empathy: acknowledging emotion before offering a solution
Clarity: short sentences, avoiding jargon
Proactivity: anticipating the next question
Tone audit scale (1-5)
5: perfect brand voice. 3: acceptable but generic. 1: inappropriate or condescending. Target average of 4.2+ out of 5. Segment bot vs agent, email vs chat.
Critical tone situations
Customer anger: empathy before policy. Delivery delay: sincere apology + action. Refusing late return: firm but respectful, alternative proposed.
Example and anti-patterns
Delay: "I understand your impatience, [First Name]. Order #1234 is in transit, delivery on Thursday. I'm keeping an eye on it with you." Avoid: "As stated on the website" without a link. "It's not our responsibility, but the carrier's" without a solution.
How to measure first contact resolution?
Measuring resolution distinguishes between a sent response and a resolved issue.
First Contact Resolution (FCR)
Share of interactions resolved without customer bounceback or ticket reopening within 48-72 hours. Bookbag: typical e-commerce FCR 65-80%, top performers 82-88%, with AI on automated categories 85-92% (Bookbag, 2026 FCR benchmarks).
Complementary metrics
Recontact rate: same intent within 7 days
Reopen rate: ticket closed then reopened
Escalation rate: bot or L1 to senior agent
Time to resolution: first message to closure
FCR by ticket type
WISMO: target 90%+. Return initiation: 75%+ via self-service portal. Product sizing: 70%+ with guide. Dispute: 40% acceptable, normal escalation.
Gorgias tags
`resolved_first_contact`, `reopened`, `escalated` for quality reporting. Customer returns within 48 hours on the same topic = resolution failure even if the ticket was closed.
How to measure satisfaction beyond CSAT?
Measuring satisfaction captures the overall perception of response quality.
Post-interaction CSAT
1-5 scale or emoji survey after closing chat or email. E-commerce target is 4.3+ (85%+ according to Lorikeet). Segment by channel and intent. 2 questions max: rating + optional mobile-friendly verbatim.
When CSAT is misleading
Response bias: unhappy customers do not respond
Low volume: 30 responses/month is not representative
Timing: survey before actual resolution
Supplementing CSAT
Monthly tagged verbatim analysis. Recontact rate as a proxy for silent dissatisfaction. Return rate post-chat pre-purchase. Monitoring reviews mentioning support. Optional question: "Was the answer accurate?" isolates correctness vs. tone.
Which QA methods and scorecards should be adopted?
The support QA methods: audits, sampling, team calibration.
Monthly audit process
Extract 5% of bot conversations + 10% of agent tickets
4-dimension scorecard: accuracy, tone, resolution, satisfaction
Double review of a 10% sample for calibration
Export scores by agent, intent, channel
Action plan for the top 3 gaps the following month
Standard scorecard
Each dimension rated 1-5 + comment. Overall verdict: pass / coaching / critical fail. Quarterly 1-hour calibration session: 10 conversations rated together.
Automatic signals to prioritize
Unmatched bot: intent not understood
Negative sentiment: NLP anger detection
Long thread: 5+ messages = possible resolution failure
Low CSAT: automatic conversation review
Office Gurus recommends Contact Quality via conversation audits in 2025-2026 (Office Gurus, support trends 2026). See support automation errors.
How do you compare the quality of bots and human agents?
Bot vs. Agent Quality: different metrics and levers.
Bot: strengths and risks
Strengths: consistency, 24/7, real-time inventory
Risks: hallucination, rigidity, poor escalation
Key metric: factual accuracy + qualified deflection
Agents: strengths and risks
Strengths: empathy, complex cases, judgment
Risks: inconsistency, obsolete macro, fatigue
Key metric: tone + FCR + CSAT
Qualified deflection
Deflection × accuracy = true bot value. 70% deflection × 95% accuracy = 66.5% actual value. Measure handoff quality: full context, maintained tone, no customer re-questioning.
Quality sprints
Post-launch bot week: 100 conversations reviewed, top 10 unmatched fixed. Same period for agents: review macros, remove obsolete ones, roleplay 5 difficult scenarios.
Which tools and dashboard for Shopify?
Shopify support quality tools stack: Gorgias + bot analytics + dashboard.
Gorgias
Native CSAT: post-ticket survey
Quality tags: intent, resolution, escalation
Macro analytics: usage and freshness
Rules: low CSAT alert
Monthly Quality Dashboard
4 quadrants: accuracy %, average tone, FCR %, CSAT. 6-month trend. By chat and email channel. Structured 45-minute weekly review: see weekly quality review (#277) for agenda, sample, and action backlog.
Data Pipeline
Gorgias export + bot analytics + Google Sheet. Shopify sidebar agent order context: reduces status and address accuracy errors. Anonymize PII for BPO shared samples.
Which quality KPIs and benchmarks should be targeted?
E-commerce support quality KPIs in a unified dashboard.
Leading KPIs
Audit accuracy: bot 97%+, agents 99%+
Average tone: 4.2+/5
Blended FCR: 70%+ (82-88% top performers)
Justified escalation: 95%+ valid handoffs
Lagging KPIs
CSAT: 4.3+ (85%+)
7-day recontact: < 8%
Pre-purchase post-chat returns: downward trend
QQS composite score
(Accuracy % × 0.35) + (Tone/5 × 20) + (FCR % × 0.25) + (CSAT/5 × 20). Score 0-100, target 80+. eDesk recommends not sacrificing accuracy for FRT: speed without correctness costs more in the long run (eDesk, Support Metrics 2026).
Slack alert if audit accuracy < 95% weekly or CSAT < 4.0 rolling 7-day. Segment by WISMO intent, return, pre-purchase, agent, FR/EN market.
How does Qstomy measure and improve quality?
Qstomy measures and improves support response quality continuously.
Quality capabilities
Confidence score: flags uncertain responses before sending
Source citation: bot cites terms hub or product sheet
Tone guardrails: brand voice settings
Contextual escalation: enriched Gorgias handoff
QA Export: conversations for monthly audit
Integrated CSAT: post-conversation feedback
Quantified DTC Scenario
Cosmetics 800 conv/month bot. M1 Audit: accuracy 91%, tone 3.8, CSAT 4.1. Actions: return corpus update, empathy prompt, 12 product intents. M3: accuracy 97%, tone 4.3, CSAT 4.5, recontact -35%.
Improvement loop
Unmatched + low confidence → weekly review → corpus update → accuracy re-audit → 1-page monthly management report. Auto tag `low_confidence` → Gorgias priority review queue.
Explore AI support, AI sales agent, request a demo.
Which operational playbooks should be launched this month?
Playbook 1: 4-dimension scorecard
Create a Notion grid: accuracy, tone, resolution, satisfaction, ratings 1-5, comment. Share with the entire support team.
Playbook 2: 50-conversation audit
This month: 30 bot + 20 agent randomly drawn. Rate the 4 dimensions. Top accuracy error = fix corpus or macro within 48 hours.
Playbook 3: CSAT + FCR tags
Activate Gorgias post-chat CSAT. Tags `resolved_first_contact`, `reopened`. Baseline FCR and CSAT 14 days before actions.
Playbook 4: weekly quality review
30 min Friday: 5 low scores, 3 corpus/macro fixes, 1 micro-training if tone pattern.
Playbook 5: monthly QQS report
Calculate composite score. Trend arrows vs. M-1. Top 3 fixes, top 3 wins, next month's actions.
Useful linking
Measured quality improves. Unmeasured quality degrades silently with volume.

Enzo
June 28, 2026


