An AI copilot that clears the support queue faster
An AI copilot for support teams that drafts replies and surfaces relevant context, so agents resolve tickets with less repetitive typing.
Who, what, and how long
E-commerce & Retail
11 weeks
Fixed price
Client name withheld under NDA. Engagement details are shown to the extent our agreement permits.
Contextual Reply Drafter
Generates grounded reply suggestions based on ticket history and documentation with source citations.
Fifty thousand resolved tickets and the product documentation were chunked, embedded and indexed in Pinecone, along with the metadata that matters for filtering: product area, ticket outcome, article age. A draft is generated only from passages the retriever actually returned, and every claim in it carries the citation it came from. When nothing clears the retrieval threshold, the copilot declines to draft. It won't write something merely plausible.
- 50,000 resolved tickets and the docs, chunked and embedded
- Drafts written only from retrieved passages, each cited
- Declines to draft when nothing clears the threshold
A suggested reply written only from the three passages that cleared the citation threshold, each claim marked with its source, the retrieval table showing two more passages read but not cited, and a ticket in the same queue where nothing cleared and the copilot declined to draft.
Knowledge search run on the open ticket with no query typed: a meaning phrase plus SKU and order keywords, results grouped into documentation and resolved tickets, and the 52,418 documents the copilot reads across six sources.
Inline Knowledge Search
Surfaces relevant help articles and past ticket resolutions right inside the helpdesk.
Search runs on the ticket the agent is already looking at, so there's no query to compose and results are there before anyone thinks to look. It's hybrid: dense vectors for meaning, keyword matching for the error codes and SKUs that embeddings blur. Results are grouped by whether they came from documentation or from a ticket someone actually resolved.
- Queries the open ticket; no search box to fill
- Hybrid retrieval: vectors for meaning, keywords for codes
- Docs and resolved tickets grouped separately
Human-in-the-Loop Review
The draft in the composer after an agent's edits, with its citations attached and edit distance logged, beside the copilot cohort's results against control and the heavily rewritten drafts queued for the weekly retrieval review.
Agents review, tweak and send suggestions with one click.
Nothing sends itself. The draft lands in the composer as editable text with its citations attached, and an agent's edits are captured as a signal: a heavily rewritten draft is a retrieval failure worth investigating, not just a bad day. Acceptance rate, edit distance and post-send CSAT are tracked per queue, which is how the pilot could prove the 43% against a control instead of asserting it.
- Drafts land editable in the composer, never sent
- Edit distance captured as a retrieval-quality signal
- Acceptance, edits and CSAT tracked per queue against a control
What we were brought in to do
Agents retyped the same answers and dug through old tickets for context. We embedded a copilot that drafts and retrieves right inside the flow of work.
A retail support team of around thirty agents, handling ticket volume that had doubled in a year while headcount hadn't. Handle time was the constraint, and the team had already tried macros, which covered the common cases and left the long tail, where the time actually goes. The engagement was scoped as a pilot with a control group from the first week.
AI Product
Where the old way broke
Agents spent most of their time drafting repetitive replies and hunting through past tickets and docs, slowing every customer response.
Agents spent most of a ticket reconstructing context: opening past tickets from the same customer, searching documentation that used different words from the customer's, then writing an answer someone had already written eleven times. Timing a sample of tickets put more than half the handle time before a single word of the reply was typed.
We built a copilot panel that drafts contextual reply suggestions and surfaces related tickets and documentation inline, so agents start from a grounded draft, review it and send.
What we built together
- 01
Indexed the knowledge base and ticket history for retrieval
Indexing exposed the real constraint: a third of the ticket history had no resolution recorded, so it could be retrieved but not learned from.
- 02
Drafted grounded, citeable reply suggestions in the helpdesk
A draft is generated only from passages the retriever actually returned, each claim carrying its citation, and nothing is drafted when retrieval clears no threshold.
- 03
Kept a human in the loop: agents review before sending
Drafts land editable in the composer and never send themselves, with edit distance captured as a signal about retrieval quality. It says nothing about the agent.
- 04
Measured handle time and CSAT against a control group
A control group of agents ran the whole eleven weeks, so the figures are real differences, with no comparison against last quarter.
Phase by phase
Phase 1: Knowledge Ingestion & Vector RAG
Ticket History & Doc Vectorization
Indexed 50,000+ past resolved support tickets and technical product documentation into Pinecone vector storage.
- Vector RAG Pipeline
- Knowledge Base Indexer
- Semantic Search Benchmark
Phase 2: Helpdesk Panel UX Design
Inline Helpdesk Copilot Interface
Designed an inline copilot panel inside the helpdesk that drafts grounded responses with cited sources.
- Helpdesk Sidebar App
- Figma Design Specs
- Copilot UI Prototype
Phase 3: Guardrails & Citation Logic
Human-in-the-Loop Review Controls
Configured strict citation checks, so agents review, tweak and approve drafted responses before sending.
- Citation Verification Engine
- Safety Guardrail Specs
- Agent Review Flow
Phase 4: Pilot & CSAT Verification
Support Tier A/B Trial
Piloted across 3 support tiers, achieving a 43% handle time reduction and +11 point CSAT gain.
- Handle Time Dashboard
- CSAT Report
- Production Rollout
The pilot: four phases from indexing to an A/B trial across 3 support tiers over 11 weeks, and the three results as differences against a concurrent control group: −43% handle time, −58% first response, +11 points CSAT.
Operational results after launch
−43%
Ticket handle time
−58%
First-response time
+11 pts
CSAT
All three figures are differences against a concurrent control group of agents on the same queues over the same eleven weeks. Seasonality and ticket mix move these metrics enough to swamp the effect in a before-and-after comparison. CSAT is the post-resolution survey, in points, not percent.
Client name withheld under NDA. Figures are approximate, drawn from the engagement’s own reporting.
About our collaboration
A cross-functional team of 6 worked on a fixed price basis over 11 weeks, covering AI copilot, Knowledge retrieval, Helpdesk integration. We shipped in two-week increments, each one releasable and reviewed live before it merged. Decisions were written down as they were made, so the reasoning outlived the people who made it.
The pilot ran against a control group of agents from the first week, so the CSAT and handle-time claims are measured differences, not trends read off history. Agents' edits to drafts were reviewed weekly as a retrieval signal, because a heavily rewritten draft is a search failure worth investigating.
What we'd carry into the next one
RAG-driven draft replies reduced ticket handle time by 43% across support tiers.
The saving was in the reading. Retrieval removed the reconstruction that took up most of a ticket; drafting saved rather less than expected.
Instant documentation search cut customer first-response time by 58%.
First response fell furthest because it's almost entirely context-gathering. There's no history to re-read, so the copilot's search is the whole task.
Human-in-the-loop agent review maintained 100% accuracy with zero hallucinated replies.
Zero hallucinated replies is a property of the send gate, not the model. Nothing reaches a customer that an agent hasn't read, which makes the accuracy claim structural.
One ticket, three verdicts
The copilot drafts only what it can cite. When it can’t, it says so.
The same customer question against three states of the knowledge base: everything indexed, one thin policy page, and nothing that clears. See which passages came back, which were cited, and why each one is in front of the agent. Switch tabs, or use the arrow keys once one is focused.
“Boots coming apart at the sole after 5 months.” Is this a replacement or a return, given the boots have been worn?
In the index: The warranty article, the worn-items policy and a resolved ticket for the same SKU are all indexed.
- 1 HC-0412Footwear warranty: sole and upper defects
Cleared the threshold on meaning and on the keyword delaminat*. Cited for the coverage rule.
0.91 - 2 #41877Ridgeback Mid sole separating at toe
Same SKU WT-BT-2291, matched by keyword. Resolved with a replacement, so the next step can be learned from it.
0.87 - 3 POL-118Returns vs warranty claims for worn items
Answers the customer's worry about the unworn rule. Cited for that sentence only.
0.82 - not cited #39120Ridgeback laces fraying at eyelets
Below the threshold, and no resolution was recorded. Shown so the agent sees it was read; never cited.
0.64
Three passages cleared, so the copilot drafts from those three and nothing else. Each claim carries the citation it came from.
The agent: Reads the draft against three citations, edits if needed, and sends. Nothing sends itself.
Illustrative: the retrieval scores and the 0.78 citation threshold are chosen to show the mechanism, not measured. The white tick on each bar is the threshold.
From an open ticket to a reply an agent chose to send
Most of a ticket used to go on reconstructing context. Now the pipeline does the reading, and every step leaves something an agent can check: which passages came back, which cleared, and which sentence cites which.
- 01 · TriggerThe ticket an agent opensThe open ticket is the query. There is no search box to fill, so results are there before anyone thinks to look.
- 02 · IngestTickets and docs, chunked50,000+ resolved tickets and the product docs embedded, with product area, ticket outcome and article age kept for filtering.
- 03 · RetrieveHybrid searchDense vectors for meaning, keyword matching for the error codes and SKUs embeddings blur. Docs and resolved tickets kept apart.
- 04 · DraftCitation checkA draft is written only from passages the retriever returned, each claim cited. When nothing clears the threshold, no draft.
- 05 · DeliverEditable in the composerNothing sends itself. Edit distance, acceptance and CSAT are tracked per queue against a control group.
No reply the team didn’t read
A send gate, a citation rule & a control group
Nothing reaches a customer unread
Drafts land in the composer as editable text and never send themselves. Zero hallucinated replies is a property of that send gate, not the model: an agent reads every reply before it goes.
Cited, or not drafted
A draft is written only from passages the retriever returned, and every claim carries its citation. When nothing clears the threshold, the copilot declines to draft. It won't write something merely plausible.
Measured against a control
A control group of agents worked the same queues for all eleven weeks, so handle time and CSAT are measured differences. Heavy rewrites are reviewed weekly as retrieval failures.
Want AI drafting in your helpdesk that cites its sources and never answers a customer on its own? Scope your build in 3 minutes.
Scope your buildNearby engagements
Data & AnalyticsA checkout that stopped losing sales to a 6-second load
Profiling found the real bottleneck behind a slow checkout (an N+1 query and an oversized bundle) and tuned both, with before-and-after metrics locked in as a baseline against future regressions.
E-commerce & Retail · 5 weeks
Web PlatformsA reader people finish, and a library that remembers where they stopped
A reading platform for a comics catalog: a browse surface people can actually navigate, a reader that gets out of the way, and a history that puts everyone back on the page they left.
E-commerce & Retail · 18 weeks
Web PlatformsOne number, a whole handset, and a database that keeps up with 125 brands
A metered IMEI lookup service that turns fifteen digits into a device, its specifications and its status. It's sold three ways to three audiences and backed by a device database that maintains itself.
Consumer Electronics · 22 weeks
Let's talk
Running a large platform, shaping a first MVP, or getting a product ready for a funding round? Tell us where you are. We'll shape the process around it, and stay with you after launch.














