E-commerce
June 29, 2026
Friday at 5 p.m., the dashboard shows a stable CSAT at 4.1 and a decreasing ticket volume. On Monday, three clients tweet the same confusion regarding return return times. The aggregated numbers didn't lie: they had simply masked a qualitative pattern that was visible in the transcripts, not in the averages.
Lorikeet points out that traditional manual auditing only covers 1 to 3% of tickets: most quality issues remain hidden until they blow up (Lorikeet, QA volume 2026). Supp estimates that AI scoring on 100% of threads, coupled with a targeted human review, is a game-changer for the same budget (Supp, AI QA 2026).
This guide, #259, deals with the monthly customer conversation audit: a 2-hour qualitative ritual, distinct from continuous dashboards. It
Summary
Why a monthly ritual rather than a dashboard alone?
A monthly conversation audit does not replace your KPIs. It reads between the lines of what the average CSAT erases.
What the dashboard shows
Volume, FRT, overall CSAT, bot deflection
Trends by channel and by macro tag
Threshold alerts (WISMO spike, CSAT drop)
What only the audit reveals
Recurring untagged phrases (« it's confusing on the site »)
Gaps between brand tone vs policy (unauthorized promise)
Poorly documented bot-to-human friction
Product signals: same SKU mentioned 8× without a dedicated tag
Actionable marketing and merchandising verbatims
SupportBench sets the industry IQS (Internal Quality Score) around 88%: the monthly audit verifies that your conversations contribute to it, not just your macros (SupportBench, QA 2026 scorecard).
Fashion DTC Example
Brand with 1,400 tickets/month, stable CSAT of 4.2. M-3 Audit: 12 threads mention "size guide unreadable on mobile" without a sizing tag. PDP fix + size bot macro. Return size tickets M+1: −19%, return segment CSAT +0.4 pt.

Convert over 2,000 customers on average per month with Qstomy.
The world’s 1st Shopify AI dedicated to customer conversion



Empowering 200+ e-commerce merchants
How does it differ from neighboring measurement guides?
Five tailored content pieces, five different rhythms and depths.
Analytics conversations
Analytics conversations: taxonomy, continuous collection, volume KPI by intent. The #259: monthly qualitative reading on a stratified sample.
Weekly bot audit (#143)
Bot audit (#143): bot accuracy grid, 60 min/week session. The #259: all channels (bot + human + social), global voice of the customer. For the 45-minute ops ritual between two audits: weekly QA review (#277). Friction report: weekly friction (#281).
Response quality (#116)
Quality (#116): continuous accuracy and FCR KPIs. The #259 feeds #116 with tagged root causes.
Product insights (#33)
Product insights: verbatim mining to catalog. The #259: ritual cadence + cross-team owners.
Legacy bot tags (#258)
Legacy bot (#258): history_exposed tags to audit. The #259 integrates these tags into the monthly sample.
Conversion signals (#260)
CRO signals (#260): quantitative alerts support → conversion. The #259 provides the verbatims, the #260 the thresholds and actions.
Who is participating and what monthly schedule should be adopted?
The ritual conversation audit takes 2 hours if the scope is set in advance.
Monthly RACI
Audit Owner: support lead or CX manager (facilitator)
Lead Auditor: senior agent or QA (scores the threads)
Rotating Participants: 1 field agent, 1 ops/merch, 1 marketing (30 min each max)
Product Owner: receives top 3 product signals, not the entire session
Standard Schedule (last Thursday or first Tuesday)
D-5: sample export + pre-filling Sheet
D-3: auditor reads 50% of threads asynchronously (1 h)
D0: 2 h session (sections 8-9)
D+2: deliverables to owners + Slack recap
D+7: at least 1 fix deployed (macro, PDP, bot chunk)
Quarterly Calibration
Every 3 months, two auditors score the same 5 threads. Gap > 1 point on any dimension = recalibrate section 5 grid. SupportBench recommends reviewing the scorecard every 3 to 6 months.
What data should be exported before the session?
The export audit conversations must be reproducible month after month.
Required fields per thread
conversation_id, date, channel (chat, email, IG, bot)
existing intent tags, resolution (FCR yes/no)
CSAT/CES if available, agent vs bot, handoff yes/no
order_id, main SKU if known
full transcript (customer messages + replies)
repeat_contact_7d, history_exposed if bot (#258)
Gorgias / helpdesk Export
Tickets closed, M-1 period, CSV export. Include Customer Timeline summary if available. Separate bot logs: merge by conversation_id before audit.
Target period and volume
Full M-1 (avoid promo weeks alone unless typical). Final sample: 30 to 50 threads for DTC < 2,000 tickets/month, 50 to 80 if higher volume. eesel AI: trends matter more than an isolated rating (eesel AI, QA support 2026).
Tool
Google Sheet or Notion database. One row = one thread. Direct transcript link in helpdesk. No manual copy-pasting of messages if API export is possible.
How to build the qualitative audit grid?
The voice of the customer audit grid blends QA scoring and business insights capture.
Six dimensions (scale 1-5)
Accuracy (25%): policy, delay, price, correct inventory
Resolution (25%): request handled without unnecessary recontact
Brand voice (15%): empathy, clarity, no robotic script
Process (15%): tags, escalation, auth if required
Voice of the customer (10%): notable verbatim captured
Opportunity (10%): assisted sale or avoided friction
Auto-fail compliance
Automatic score of 1 if: unauthorized refund promise, exposed PII, medical/regulated advice, false return policy. SupportBench: compliance categories on auto-fail.
Insights columns (free text)
verbatim_client, root_cause, fix_type (macro / bot / PDP / ops / policy), owner, priority P1-P3. Align tags with taxonomy (#135).
Conversation IQS Score
Weighted average of dimensions. Monthly IQS = average of audited threads. DTC target: ≥ 88%. Alert if < 85% or drop > 3 pts vs M-1.
Which five axes of qualitative analysis by thread?
Beyond the score, five conversation reading areas structure the note-taking.
Area 1: Real intent vs tag
Did the customer want a return or an exchange? Does the helpdesk tag reflect the request? Discrepancy = bot or agent triage issue.
Area 2: Friction point
Where did the customer get stuck? Site, policy, delay, product, bot, handoff. Note the customer's exact phrase.
Area 3: Response quality
Complete response? Copy-pasting macro out of context? Contradiction between bot and agent? See brand voice.
Area 4: Product / ops signal
Recurring mention of SKU, carrier, packaging, promo? Escalate to merchandising (#108) if dynamic pattern.
Area 5: Reputational risk
Customer tone turning harsh, threat of public review, creepy historical complaint (#257). Tag reputation_risk for immediate lead review.
Auditor note mini-template
"Intent: [X]. Friction: [Y]. Verbatim: "...". Proposed fix: [Z]. Owner: [name]." Max 3 lines per thread in live session.
How do you sample 30 to 50 representative conversations?
The stratified monthly audit sample beats a pure random draw.
Typical DTC Breakdown
35 % top 3 volume intents (WISMO, return, product)
15 % bot-only resolved (verify accuracy)
15 % handoff bot → human
10 % CSAT 1-2 or high CES
10 % repeat_contact_7d (#256)
10 % risk intents (promo, dispute, VIP, regulated)
5 % secondary channels (IG, WhatsApp)
Gorgias filters ready to paste
Date = last 30 days, status = closed. Sub-samples: tag wismo + CSAT < 3; tag bot_resolved; tag escalated; tag history_exposed + complaint; created_via instagram.
Avoid bias
Do not audit only the threads of your best agent. Quarterly auditor rotation. Include 5 threads where you corrected the bot live (override).
AI Acceleration (optional)
Add-on: score 100% of threads by LLM on section 5 topic, then human audit the 200 flagged threads (low score, compliance, declining sentiment). Estimated cost $0.01 to $0.05/ticket.
How to conduct the 2-hour audit session?
The monthly audit session follows a strict agenda to stay on schedule.
0-15 min: month context
Owner presents KPI M-1 vs M-2: volume, CSAT, FCR, top 5 intents delta. 1 slide, no dashboard debate.
15-75 min: ticket review (12 to 15 tickets deep dive)
Read full transcript aloud or screen shared. Auditor announces scores + verbatim. Participants add root_cause. Prioritize auto-fail and low CSAT tickets first.
75-105 min: patterns and clustering
Group identical root_causes together. Example: 4 tickets "unclear return delay on PDP" → 1 merchandising action. Omind: detecting policy deviation across 12 agents is worth more than an isolated ticket (Omind, QA retail 2026).
105-120 min: top 3 actions + owners
Choose exactly 3 P1 actions with a D+7 deadline. Postpone the rest to the Notion backlog. Record 2 min of customer verbatims for the "voice of the month" all-hands.
What deliverables should be produced after the audit?
An audit without a deliverable is a reading meeting. Five mandatory monthly audit deliverables.
Deliverable 1: 1-page summary sheet
Monthly IQS, delta vs M-1, top 3 root_causes, top 3 verbatims, 3 P1 actions. Slack #support + @product tag if catalog signal.
Deliverable 2: patch backlog
Notion or Linear: macro to rewrite, bot chunk to sync, PDP to enrich, escalation rule. Link conversation_id proof.
Deliverable 3: bot regression cases
Each serious bot error → 1 case added to dataset #143. Replay before next corpus deployment.
Deliverable 4: merchandising / marketing brief
If ≥ 3 threads same website friction: 5-line email to owner with verbatims. Link questions → blog (#127) if SEO intent.
Deliverable 5: documented decision
Ambiguous policy detected → support decision ticket (#237). Avoid the following month reproducing the same error.
How to combine a monthly audit with continuous dashboards?
The audit vs dashboard must feed into each other, not duplicate each other.
Dashboard feeds the audit
Spike intent → over-sample this intent in the following month
Low CSAT segment → pull 5 additional threads for this segment
New bot flow → audit 10 dedicated threads in M+1
Audit feeds the dashboard
New root_cause tag → add to Gorgias reporting
Verbatim cluster → monthly "top website friction" widget
IQS audit → complementary KPI to CSAT (more discriminating)
Recommended monthly dashboard
Columns: IQS audit | Global CSAT | FCR | top intent delta | top root_cause | actions closed D+7 (target 3/3) | IQS bot subset | IQS human subset. Separate bot and human: a blended CSAT masks a bot drift.
Loop with post-ticket feedback
Cross-reference audit and post-support feedback (#239): CSAT 1-2 threads audited priority-wise even if volume is low.
How does Qstomy facilitate the monthly audit of conversations?
Qstomy exports structured transcripts, intents, handoffs, and tags to power the audit without manual export.
Audit features
Monthly CSV export: transcript, intent, confidence, RAG sources
Audit tags: history_exposed, reco_declined, handoff_reason
AI-assisted score: pre-score 6 dimensions section 5
Flag queue: auto-fail compliance threads at the top of the session
Regression pack: Month-1 cases exported to test set
Verbatim extract: top customer phrases of the month
Quantified DTC Scenario
Beauty brand, 950 conv/month (52% bot, 48% human). Before ritual #259: ad hoc audit of 8 threads/quarter, unknown CSAT, 2 reactive fixes/month. After monthly ritual of 40 threads + Qstomy pre-score: CSAT 91%, 3 actions P1 closed D+7 at 100%, repeat sizing tickets −24% over 90 days, audit prep time −60% (auto export vs copy-paste).
Recommended Stack
Qstomy logs + Gorgias tickets + Sheet grid section 5. Async reviews D-3, live session D0. No enterprise QA tool required to start.
Explore AI customer support, Shopify, request a demo.
Which playbooks should be used to launch the ritual this month?
Playbook 1: Sheet grid (2 h, week 1)
Duplicate columns section 5-6. Add weighted IQS formulas. Share with support lead + 1 senior agent.
Playbook 2: first export (1 h)
Pull 40 M-1 threads according to stratification section 7. Pre-fill metadata, leave scores blank.
Playbook 3: pilot session (2 h)
Agenda section 8. Score 12 threads minimum. Produce summary sheet section 9. Deadline 3 actions D+7.
Playbook 4: bot loop (1 h post-session)
Bot errors → regression case #143. Sync corpus within 48 h. Replay 5 cases before closing P1 action.
Playbook 5: cross-team sharing (30 min)
Email product/marketing: 3 verbatims + 1 intent delta graph. Invite to M+1 session of 15 min if catalog signal is present.
Playbook 6: M+2 iteration (1 h)
Compare IQS M+1 vs pilot. Adjust grid weights if a dimension is not very discriminating. Activate AI pre-scoring if volume > 800 conv/month.
Useful links
A monthly conversation audit is not about rereading tickets to check boxes. It is about listening to your market directly, once a month, with enough structure to act before the dashboard turns red. Brands that ritualize this review get six months ahead of those that only listen to their averages.

Enzo
June 29, 2026


