Skip to content

An AI copilot that clears the support queue faster

An AI copilot for support teams that drafts replies and surfaces relevant context, so agents resolve tickets with less repetitive typing.

Open queue · Tier 2
Tier 2 · Orders & warranty · every ticket carries the copilot's own verdict · signed in as Rosa Delgado
Tier 2 · 11 agents online Thu, 09/17/2026RD
34Open in Tier 2Orders & warranty
6Breaching todaySLA passes before noon
21Draft readyCited, waiting for review
8Low confidenceOne thin source; read closely
5No matchDeclined to draft

Tickets

12 of 34 · 6 breaching shown in red
TicketSubjectSLACopilot
#48207Jordan Pike · Boots coming apart at the sole after 5 months 1:12 PMDraft ready
#48199Amara Osei · Tent pole snapped on first setup, WT-TN-0418 10:40 AMDraft ready
#48213Devin Marsh · Can I get the Kestrel pack in a torso size you don't list? 11:05 AMNo match
#48188Helen Brandt · Refund shows pending for 9 days 10:52 AMDraft ready
#48221Luis Carrera · Error E-214 when applying gift card at checkout 12:30 PMDraft ready
#48174Priya Raman · Rain shell seam tape lifting after a wash 11:18 AMLow confidence
#48230Sam Whitaker · Wrong color stove fuel canister delivered 1:45 PMDraft ready
#48196Grace Liu · Price adjustment after sale started the next day 11:32 AMDraft ready
#48238Owen Fitch · Headlamp won't hold charge, bought in store 2:02 PMLow confidence
#48241Nadia Haddad · Ship a warranty replacement to a different address? 2:15 PMDraft ready
#48179Theo Grant · Rental return late fee on a canceled trip 11:50 AMNo match
#48246Ivy Kowalski · Sleeping bag zipper snagging, WT-SB-7730 2:40 PMDraft ready
Verdicts are set when a ticket arrives and again on every customer reply22 more

Copilot · #48207

Reads the open ticket
Draft ready3 passages cleared 0.78 · every claim cites one
Sources retrievedscore · threshold tick
1Footwear warranty: sole and upper defectsHC-0412 · Help center · updated 06/02/20260.91
2Ridgeback Mid sole separating at toe#41877 · Resolved · replacement shipped · 07/18/20260.87
3Returns vs warranty claims for worn itemsPOL-118 · Policy · updated 03/20/20260.82
2 more read, below threshold, not cited
Draft · not sentHi Jordan, I'm sorry your Ridgeback Mids are coming apart. Sole separation within 12 months of purchase is covered as a manufacturing defect under our footwear warranty1, and because the boots have been worn this goes through a warranty claim rather than a return, so the unworn rule doesn't apply3. If you reply with a photo of the toe on each boot, we'll approve the claim and ship a replacement pair2. No need to send the old pair back first.
Insert into composerOpen sources

Who, what, and how long

Industry

E-commerce & Retail

Duration

11 weeks

Cooperation model

Fixed price

Services
AI copilotKnowledge retrievalHelpdesk integration
Integrations
ShopifyStripeKlaviyoShipStation
Technologies
PythonFastAPIOpenAI GPT-4oPinecone Vector DBZendesk APIReact
Team
1 Project lead1 Product designer1 ML engineer1 Backend engineer1 Data engineer1 QA engineer

Client name withheld under NDA. Engagement details are shown to the extent our agreement permits.

Contextual Reply Drafter

Generates grounded reply suggestions based on ticket history and documentation with source citations.

Fifty thousand resolved tickets and the product documentation were chunked, embedded and indexed in Pinecone, along with the metadata that matters for filtering: product area, ticket outcome, article age. A draft is generated only from passages the retriever actually returned, and every claim in it carries the citation it came from. When nothing clears the retrieval threshold, the copilot declines to draft. It won't write something merely plausible.

What shipped
  • 50,000 resolved tickets and the docs, chunked and embedded
  • Drafts written only from retrieved passages, each cited
  • Declines to draft when nothing clears the threshold
Ticket #48207 · Jordan PikeDraft ready
Boots coming apart at the sole after 5 months · received 9:12 AM · Due 1:12 PM
Order WT-311842 Thu, 09/17/2026RD

Customer message

Email · 9:12 AM
JPJordan Pike3 earlier orders · 1 earlier ticket
ProductRidgeback Mid WPSKUWT-BT-2291OrderWT-311842Bought04/11/2026
Hi, my Ridgeback Mids (order WT-311842) have started separating at the toe, the sole is peeling off the upper on both boots. I've only had them since April and mostly used them on weekend trails. Is this something you'll replace, or do I need to send them back as a return? The return page says items must be unworn.

#48213 · declined to draft

Same queue, 9:41 AM
Devin Marsh · Kestrel pack in a torso size you don’t list?
No matchNothing cleared 0.78HC-0520Kestrel 48 pack: sizing and fit0.69#43318Custom harness request (no resolution recorded)0.61No draft written. Nearest passages shown so the agent can see what was read.
Logged as an answer gap for the knowledge team

Suggested reply

Written only from the passages below · not sent
Hi Jordan, I'm sorry your Ridgeback Mids are coming apart. Sole separation within 12 months of purchase is covered as a manufacturing defect under our footwear warranty1, and because the boots have been worn this goes through a warranty claim rather than a return, so the unworn rule doesn't apply3. If you reply with a photo of the toe on each boot, we'll approve the claim and ship a replacement pair2. No need to send the old pair back first.
3 claims, 3 citations: each sentence that states policy points at a passage Nothing in the draft comes from outside the 3 cited passages Lands in the composer as editable text. The agent sends it, or doesn't

Retrieval for this ticket

Citation threshold 0.78 · below it a passage is read, never cited
CitePassageSourceMatched onScore
1HC-0412 Footwear warranty: sole and upper defectsDocsMeaning + keyword0.91
2#41877 Ridgeback Mid sole separating at toeTicketKeyword WT-BT-22910.87
3POL-118 Returns vs warranty claims for worn itemsDocsMeaning0.82
—#39120 Ridgeback laces fraying at eyeletsTicketKeyword WT-BT-22910.64
—HC-0097 Caring for waterproof leather bootsDocsMeaning0.58
On screen

A suggested reply written only from the three passages that cleared the citation threshold, each claim marked with its source, the retrieval table showing two more passages read but not cited, and a ticket in the same queue where nothing cleared and the copilot declined to draft.

Knowledge for #48207
Search runs on the ticket that's open · results arrive before anyone thinks to look
Search manually Thu, 09/17/2026RD
Query built from the open ticket · no search box to fill“boot sole separating from upper at toe after five months, replace or return, worn outdoors”
Dense vectors · meaning1 phrase, embedded
Keywords · codes embeddings blurWT-BT-2291WT-311842delaminat*

Documentation

Help center, policy, product docs · 4 of 11
HC-0412Footwear warranty: sole and upper defects0.91Help center · 06/02/2026·meaning + delaminat*Sole separation or delamination within 12 months of purchase is treated as a manufacturing defect, whatever the use.
POL-118Returns vs warranty claims for worn items0.82Policy · 03/20/2026·meaningThe unworn condition applies to returns only. Worn items with a defect are handled as a warranty claim.
HC-0433Shipping a warranty replacement0.74Help center · 08/14/2026·meaningReplacements ship once a claim is approved. Customers keep or recycle the original pair.
HC-0097Caring for waterproof leather boots0.58Help center · 11/04/2024·meaningClean off mud after each use and let boots dry away from direct heat to protect the membrane.

Resolved tickets

What an agent actually did · 4 of 17
#41877Ridgeback Mid sole separating at toe0.87Resolved · replacement shipped·keyword WT-BT-2291Asked for photos of both toes, approved warranty claim same day, replacement pair shipped.
#44310Ridgeback Mid toe cap peeling0.81Resolved · replacement shipped·keyword WT-BT-2291Toe cap lifting at 7 months. Treated as defect under footwear warranty, no return needed.
#45502Sole separating on Talus trail runner0.71Resolved · store credit·meaningPast 12 months; offered store credit as goodwill rather than a warranty replacement.
#39120Ridgeback laces fraying at eyelets0.64No resolution recorded·keyword WT-BT-2291Customer sent photos. Ticket closed without a note on what was offered.

What the copilot reads

52,418 documents across six sources · product area, outcome and article age kept as filters
Resolved tickets50,2121 in 3 without a recorded resolution
Help center1,184Filtered by article age
Product docs612By product area
Macros214Common cases only
Carrier & shipping notes108By product area
Returns & warranty policy88Filtered by article age
On screen

Knowledge search run on the open ticket with no query typed: a meaning phrase plus SKU and order keywords, results grouped into documentation and resolved tickets, and the 52,418 documents the copilot reads across six sources.

Inline Knowledge Search

Surfaces relevant help articles and past ticket resolutions right inside the helpdesk.

Search runs on the ticket the agent is already looking at, so there's no query to compose and results are there before anyone thinks to look. It's hybrid: dense vectors for meaning, keyword matching for the error codes and SKUs that embeddings blur. Results are grouped by whether they came from documentation or from a ticket someone actually resolved.

What shipped
  • Queries the open ticket; no search box to fill
  • Hybrid retrieval: vectors for meaning, keywords for codes
  • Docs and resolved tickets grouped separately
Reply to Jordan Pike · #48207
The draft is in the composer as editable text · it sends when an agent sends it
Edited by RD Thu, 09/17/2026RD

Composer

Draft inserted 9:26 AM · not sent
ToJordan Pike <j.pike@…>SubjectRe: Boots coming apart at the sole after 5 months
Hi Jordan, I'm sorry your Ridgeback Mids are coming apart. thanks for the detail, and sorry the Ridgebacks let you down this early. Sole separation within 12 months of purchase is covered as a manufacturing defect under our footwear warranty1, and because the boots have been worn this goes through a warranty claim rather than a return, so the unworn rule doesn't apply3. If you reply with a photo of the toe on each boot, we'll approve the claim and ship a replacement pair2 in your usual size 10. No need to send the old pair back first.
Citations attached to this reply · internal, not shown to the customer1HC-04122#418773POL-118
Edit distance9%A light edit: tone and a size, no facts changed. Logged for the Tier 2 queue with acceptance and post-send CSAT. A heavy rewrite goes to the weekly retrieval review.
Send replyDiscard draft Only an agent’s click sends

Copilot cohort vs control

Same queues · same 11 weeks
Ticket handle time−43%
First-response time−58%
CSAT (post-resolution survey)Points on the post-resolution survey+11 pts
Copilot Control = 100Differences, not trends

Heavily rewritten drafts

Weekly retrieval review · a search failure, not a bad day
TicketDraft · likely causeEdit
#47912Stove regulator recall eligibilityRecall notice not yet indexed71%
#47988Kids' sizing exchange across brandsRetrieved an outdated size chart64%
#48031Gift card balance after partial refundMacro outranked the policy page58%
#48066Bike rack fit for 2019 hatchbackFit table chunked across rows52%
#48102Loyalty points on a price adjustmentTicket with no resolution recorded49%
#48140Canceled backorder, card still heldKeyword matched the wrong order code46%
Flagged above 45% edit distance · reviewed with the knowledge team

Human-in-the-Loop Review

On screen

The draft in the composer after an agent's edits, with its citations attached and edit distance logged, beside the copilot cohort's results against control and the heavily rewritten drafts queued for the weekly retrieval review.

Agents review, tweak and send suggestions with one click.

Nothing sends itself. The draft lands in the composer as editable text with its citations attached, and an agent's edits are captured as a signal: a heavily rewritten draft is a retrieval failure worth investigating, not just a bad day. Acceptance rate, edit distance and post-send CSAT are tracked per queue, which is how the pilot could prove the 43% against a control instead of asserting it.

What shipped
  • Drafts land editable in the composer, never sent
  • Edit distance captured as a retrieval-quality signal
  • Acceptance, edits and CSAT tracked per queue against a control
Introduction

What we were brought in to do

Agents retyped the same answers and dug through old tickets for context. We embedded a copilot that drafts and retrieves right inside the flow of work.

A retail support team of around thirty agents, handling ticket volume that had doubled in a year while headcount hadn't. Handle time was the constraint, and the team had already tried macros, which covered the common cases and left the long tail, where the time actually goes. The engagement was scoped as a pilot with a control group from the first week.

AI Product

The problem

Where the old way broke

Agents spent most of their time drafting repetitive replies and hunting through past tickets and docs, slowing every customer response.

Agents spent most of a ticket reconstructing context: opening past tickets from the same customer, searching documentation that used different words from the customer's, then writing an answer someone had already written eleven times. Timing a sample of tickets put more than half the handle time before a single word of the reply was typed.

We built a copilot panel that drafts contextual reply suggestions and surfaces related tickets and documentation inline, so agents start from a grounded draft, review it and send.

The solution

What we built together

04
  1. 01

    Indexed the knowledge base and ticket history for retrieval

    Indexing exposed the real constraint: a third of the ticket history had no resolution recorded, so it could be retrieved but not learned from.

  2. 02

    Drafted grounded, citeable reply suggestions in the helpdesk

    A draft is generated only from passages the retriever actually returned, each claim carrying its citation, and nothing is drafted when retrieval clears no threshold.

  3. 03

    Kept a human in the loop: agents review before sending

    Drafts land editable in the composer and never send themselves, with edit distance captured as a signal about retrieval quality. It says nothing about the agent.

  4. 04

    Measured handle time and CSAT against a control group

    A control group of agents ran the whole eleven weeks, so the figures are real differences, with no comparison against last quarter.

Process

Phase by phase

  1. Phase 1: Knowledge Ingestion & Vector RAG

    Ticket History & Doc Vectorization

    Indexed 50,000+ past resolved support tickets and technical product documentation into Pinecone vector storage.

    • Vector RAG Pipeline
    • Knowledge Base Indexer
    • Semantic Search Benchmark
  2. Phase 2: Helpdesk Panel UX Design

    Inline Helpdesk Copilot Interface

    Designed an inline copilot panel inside the helpdesk that drafts grounded responses with cited sources.

    • Helpdesk Sidebar App
    • Figma Design Specs
    • Copilot UI Prototype
  3. Phase 3: Guardrails & Citation Logic

    Human-in-the-Loop Review Controls

    Configured strict citation checks, so agents review, tweak and approve drafted responses before sending.

    • Citation Verification Engine
    • Safety Guardrail Specs
    • Agent Review Flow
  4. Phase 4: Pilot & CSAT Verification

    Support Tier A/B Trial

    Piloted across 3 support tiers, achieving a 43% handle time reduction and +11 point CSAT gain.

    • Handle Time Dashboard
    • CSAT Report
    • Production Rollout
Pilot · copilot cohort vs control
11 weeks · 3 support tiers · ~30 agents · control group from week 1
Rolled out to production Thu, 09/17/2026RD
Duration11 weeksTiers3 support tiersTeam~30 agentsComparisonConcurrent control groupSeasonality and ticket mix move these metrics, so every figure is a difference against control.

Phase 1

Done
Ticket history & doc vectorization50,000+ resolved tickets and the product docs chunked, embedded and indexed. A third of the history had no resolution recorded. Vector RAG pipeline Knowledge base indexer Semantic search benchmark

Phase 2

Done
Inline helpdesk copilot panelA panel inside the helpdesk that drafts grounded replies with cited sources and searches on the open ticket. Helpdesk sidebar app Design specs Copilot UI prototype

Phase 3

Done
Guardrails & citation logicStrict citation checks, a refusal when nothing clears, and a review flow where agents tweak and approve before sending. Citation verification engine Safety guardrail specs Agent review flow

Phase 4

Done
Support tier A/B trialPiloted across 3 support tiers with a control group of agents on the same queues from the first week. Handle time dashboard CSAT report Production rollout

Ticket handle time

vs control
−43%
Copilot cohort 57Control 100
Indexed to control = 100. The saving came mostly from reading.

First-response time

vs control
−58%
Copilot cohort 42Control 100
Indexed to control = 100. First response is almost all context-gathering, so search does most of the work.

CSAT (post-resolution survey)

vs control
+11 pts
Copilot cohort +11 ptsControl baseline
Difference in points on the post-resolution survey, copilot cohort against control.
On screen

The pilot: four phases from indexing to an A/B trial across 3 support tiers over 11 weeks, and the three results as differences against a concurrent control group: −43% handle time, −58% first response, +11 points CSAT.

Operational results after launch

−43%

Ticket handle time

−58%

First-response time

+11 pts

CSAT

All three figures are differences against a concurrent control group of agents on the same queues over the same eleven weeks. Seasonality and ticket mix move these metrics enough to swamp the effect in a before-and-after comparison. CSAT is the post-resolution survey, in points, not percent.

Client name withheld under NDA. Figures are approximate, drawn from the engagement’s own reporting.

Ways of working

About our collaboration

  • A cross-functional team of 6 worked on a fixed price basis over 11 weeks, covering AI copilot, Knowledge retrieval, Helpdesk integration. We shipped in two-week increments, each one releasable and reviewed live before it merged. Decisions were written down as they were made, so the reasoning outlived the people who made it.

    The pilot ran against a control group of agents from the first week, so the CSAT and handle-time claims are measured differences, not trends read off history. Agents' edits to drafts were reviewed weekly as a retrieval signal, because a heavily rewritten draft is a search failure worth investigating.

What we'd carry into the next one

  • RAG-driven draft replies reduced ticket handle time by 43% across support tiers.

    The saving was in the reading. Retrieval removed the reconstruction that took up most of a ticket; drafting saved rather less than expected.

  • Instant documentation search cut customer first-response time by 58%.

    First response fell furthest because it's almost entirely context-gathering. There's no history to re-read, so the copilot's search is the whole task.

  • Human-in-the-loop agent review maintained 100% accuracy with zero hallucinated replies.

    Zero hallucinated replies is a property of the send gate, not the model. Nothing reaches a customer that an agent hasn't read, which makes the accuracy claim structural.

One ticket, three verdicts

The copilot drafts only what it can cite. When it can’t, it says so.

The same customer question against three states of the knowledge base: everything indexed, one thin policy page, and nothing that clears. See which passages came back, which were cited, and why each one is in front of the agent. Switch tabs, or use the arrow keys once one is focused.

Ticket #48207 · Jordan Pike

“Boots coming apart at the sole after 5 months.” Is this a replacement or a return, given the boots have been worn?

Draft ready4 passages retrieved · 3 cleared 0.78

In the index: The warranty article, the worn-items policy and a resolved ticket for the same SKU are all indexed.

Sources retrieved, and why the agent sees each
  • 1 HC-0412Footwear warranty: sole and upper defects

    Cleared the threshold on meaning and on the keyword delaminat*. Cited for the coverage rule.

    0.91
  • 2 #41877Ridgeback Mid sole separating at toe

    Same SKU WT-BT-2291, matched by keyword. Resolved with a replacement, so the next step can be learned from it.

    0.87
  • 3 POL-118Returns vs warranty claims for worn items

    Answers the customer's worry about the unworn rule. Cited for that sentence only.

    0.82
  • not cited #39120Ridgeback laces fraying at eyelets

    Below the threshold, and no resolution was recorded. Shown so the agent sees it was read; never cited.

    0.64
What lands in the composer
Sole separation within 12 months is covered as a manufacturing defect1, and worn boots go through a warranty claim, not a return3. Send a photo of each toe and we'll ship a replacement pair2.

Three passages cleared, so the copilot drafts from those three and nothing else. Each claim carries the citation it came from.

The agent: Reads the draft against three citations, edits if needed, and sends. Nothing sends itself.

Illustrative: the retrieval scores and the 0.78 citation threshold are chosen to show the mechanism, not measured. The white tick on each bar is the threshold.

Architecture

From an open ticket to a reply an agent chose to send

Most of a ticket used to go on reconstructing context. Now the pipeline does the reading, and every step leaves something an agent can check: which passages came back, which cleared, and which sentence cites which.

  1. 01 · Trigger
    The ticket an agent opensThe open ticket is the query. There is no search box to fill, so results are there before anyone thinks to look.
  2. 02 · Ingest
    Tickets and docs, chunked50,000+ resolved tickets and the product docs embedded, with product area, ticket outcome and article age kept for filtering.
  3. 03 · Retrieve
    Hybrid searchDense vectors for meaning, keyword matching for the error codes and SKUs embeddings blur. Docs and resolved tickets kept apart.
  4. 04 · Draft
    Citation checkA draft is written only from passages the retriever returned, each claim cited. When nothing clears the threshold, no draft.
  5. 05 · Deliver
    Editable in the composerNothing sends itself. Edit distance, acceptance and CSAT are tracked per queue against a control group.

No reply the team didn’t read

A send gate, a citation rule & a control group

Nothing reaches a customer unread

Drafts land in the composer as editable text and never send themselves. Zero hallucinated replies is a property of that send gate, not the model: an agent reads every reply before it goes.

Cited, or not drafted

A draft is written only from passages the retriever returned, and every claim carries its citation. When nothing clears the threshold, the copilot declines to draft. It won't write something merely plausible.

Measured against a control

A control group of agents worked the same queues for all eleven weeks, so handle time and CSAT are measured differences. Heavy rewrites are reviewed weekly as retrieval failures.

Want AI drafting in your helpdesk that cites its sources and never answers a customer on its own? Scope your build in 3 minutes.

Scope your build
Have a project?

Let's talk

Running a large platform, shaping a first MVP, or getting a product ready for a funding round? Tell us where you are. We'll shape the process around it, and stay with you after launch.