Skip to content

Keeping AI in production healthy, month after month

A managed retainer that monitors, tunes, and improves live AI systems, from support agents to forecasting, so they keep getting better after launch.

Fleet · production
6 models across 4 systems · today since 12:00 AM · answers held to a 3-second target
Export Thu, 09/17/2026RA
Uptime · per system
99.9%SLA met on 4 of 4 systems this month
Answers under 3s · p95
All modelsSlowest today: store associate, 2.7s
Quality drift caught
Pre-impact2 regressions held by last night's eval
Model cost · monthly
−21%Against the three months before the retainer

Models in production

p95 and volume since midnight · error budget for the month
System · modelRoleChannelp95 todayVolumeError budget leftStatus
Support agentVerith M3AnswersOnline chat2.4s18,42071%Healthy
Support agentGuardrail classifierSign-offOnline chat0.3s18,42088%Healthy
Store associate assistantVerith M3AnswersIn store2.7s6,91554%Input drift
Product Q&AEmbedding modelRetrievalOnline0.4s31,20893%Healthy
Product Q&AVerith M3AnswersOnline2.1s12,76480%Healthy
Demand forecastingForecast model · in-houseNightly batchBuyingBatch · ran 1:40 AM1 run97%Healthy

Uptime against the SLA · per system

September to date · not aggregated
Support agent99.9% SLA met
Store associate assistant99.9% SLA met
Product Q&A99.9% SLA met
Demand forecasting99.9% SLA met

Since midnight

Wardlight log
3:52 AMsupport-agent/answer v48 held · 2 regressions
3:52 AM3 other versions passed the frozen set
7:05 AMStore associate · question mix shifting
9:00 AMDaily cost & latency review signed off

The brief, in specifics

Industry

E-commerce & Retail

Duration

Ongoing

Cooperation model

Ongoing retainer, phased

Services
MonitoringModel tuningContinuous improvement
Integrations
ShopifyStripeKlaviyoShipStation
Technologies
PythonLangChainVerith M3PrometheusGrafanaFastAPI
Team
1 Project lead1 Product designer1 ML engineer1 Backend engineer1 Data engineer1 QA engineer

Client name withheld under NDA. Engagement details are shown to the extent our agreement permits.

What went wrong, and when

  • AI systems drift, prompts go stale, and edge cases pile up. Without ownership, quality quietly erodes after launch.

    Every one of the four had the same shape of gap. Latency and error rates were monitored because that's what infrastructure monitoring does; response quality wasn't, because nobody had defined what quality meant for these systems. Drift was therefore invisible by construction: the systems kept answering, quickly and successfully, with answers that had quietly got worse.

    We put monitoring, evaluation, and a tuning cadence around every live AI system, with a monthly review of quality, cost, and new edge cases to fix.

Process

Phase by phase

  1. Phase 1: Model Drift & Prompt Audit

    Production Evaluation Baseline

    Audited production AI performance, hallucination rates, and prompt degradation across live retail customer service agents.

    • Model Drift Audit
    • SLA Monitor Architecture
    • Quality Baseline Report
  2. Phase 2: Automated Eval & Guardrails

    Regression Suite & Safety Filters

    Engineered continuous evaluation pipelines testing every agent output against regression suites and safety guardrails.

    • Eval Suite Pipeline
    • Hallucination Detector
    • Safety Filter Layer
  3. Phase 3: Cost Optimization & Caching

    Token Compression & Semantic Cache

    Optimized prompt tokens and implemented semantic caching, reducing monthly API costs by 32%.

    • Token Optimization Engine
    • Semantic Cache Layer
    • Cost Reduction Report
  4. Phase 4: Ongoing SLA & Monthly Tuning

    Continuous Improvement Cadence

    Delivered monthly performance reports, prompt updates, and new feature integrations under guaranteed SLA response times.

    • Monthly SLA Reports
    • Continuous Improvement Log
    • Prometheus Dashboard
Monthly report · August 2026
For Fernbrook's commercial team: quality, spend and edge cases, and what to fix, what to cap, what to retire
Send to business owners Thu, 09/17/2026RA
Uptime · per system
99.9%SLA met on 4 of 4 · not aggregated
Quality drift caught
Pre-impactBlocked or flagged before customers saw it
Model cost
−21%Monthly inference vs three months prior
Tuning sprints
2Every 14 days · against that fortnight's edge cases

Incidents · each on its own clock

August 1–31
08/0108/0808/1508/2208/31
INC-0804Question mix shifted after the patio range launchInput driftStore associate assistantDetected 08/04 9:12 AMAcknowledged 9:20 AMFixed in sprint 08/07
INC-0819Candidate v45 regressed on returns wordingBlocked pre-releaseSupport agentFlagged 08/19 3:48 AMHeld 3:48 AMShipped as v46 08/21
INC-0823Retrieval hit rate under floor after catalog importSLA metricProduct Q&ADetected 08/23 2:05 PMAcknowledged 2:11 PMResolved 4:40 PM
INC-0827Provider-side latency, p95 approaching targetWithin SLAAll answer models · Verith M3Detected 08/27 11:31 AMAcknowledged 11:36 AMResolved 12:14 PM
Before the retainer: a deprecated model version, unnoticed for 11 daysThe replacement behaved differently on Fernbrook’s edge cases. Uptime monitoring stayed green throughout; nothing was watching quality.

Retainer hours · where they went

Shares, August
Continuous evaluation
Tuning sprints
Daily cost & latency review
Incidents
Monthly report

Decisions for the business

Commercial, not engineering
FixPrice-match answersPolicy source is stale; the question mix is rising. Next sprint.
CapNightly forecast notesCap tokens per run; the notes run longer than buyers use.
RetireAssociate catalog-PDF lookupSuperseded by Product Q&A retrieval; still billed on every miss.
On screen

The monthly report for the business: uptime per system, every incident on its own clock, how the retainer hours were spent, and what to fix, cap or retire.

The numbers, before and after

99.9%

Uptime

Pre-impact

Quality drift caught

−21%

Model cost

Uptime is per system against the agreed SLA, never aggregated. Quality drift caught before impact is a count of regressions the evaluation pipeline blocked or flagged before they reached customers, which is the claim this retainer exists to make. Cost is monthly inference spend, compared against the three months before the retainer began.

Client name withheld under NDA. Figures are approximate, drawn from the engagement’s own reporting.

Introduction

The engagement

Several AI systems were live and needed watching and improving. We run a managed retainer that keeps everything healthy and getting better.

A retailer with four AI systems already in production, each built by a different team or vendor and none of them owned after launch. The engagement started from an incident: a model provider deprecated a version, the replacement behaved differently on the retailer's edge cases, and nobody noticed for eleven days because nothing was watching quality. Only uptime.

Managed Operations

How it was handled

  1. 01

    Set up continuous evaluation pipelines for production prompts

    The evaluation set had to be built from scratch, because none of the four live systems had a definition of what a good answer looked like.

  2. 02

    Monitored response quality, latency, and cost per query daily

    Daily review was set as a cadence, not a dashboard, because the four systems had all had monitoring and none of it had been looked at.

  3. 03

    Ran prompt tuning sprints every two weeks to catch new edge cases

    Tuning runs every two weeks against the edge cases those two weeks actually produced, so nothing waits for a complaint to get scheduled.

  4. 04

    Reported business impact, SLA compliance, and recommendations monthly

    Reporting monthly instead of continuously was deliberate: a dashboard nobody opens had already been tried on all four systems and had changed nothing.

Continuous LLM Evaluation Pipeline

Automated regression tests and hallucination detectors running against live production traffic.

A frozen evaluation set is replayed against every prompt or model change before it ships, and a sample of live traffic is scored continuously afterward. Hallucination checks are grounded wherever possible, never just model-graded: an assertion is verified against the source the system retrieved, so the detector isn't another model's opinion. A regression blocks the release outright; it doesn't just file a ticket.

What shipped
  • Frozen eval set replayed before every prompt or model change
  • Live traffic sampled and scored continuously
  • Regressions block the release outright
Overnight evaluation · 09/17/2026
Every changed prompt replayed against the frozen set of 2,400 recorded conversations · started 1:15 AM
Compare versions Thu, 09/17/2026RA
support-agent/answer · v48Release blocked by the gate, not a ticket — v47 stays liveReplayed2,400v47 live95.2%v48 candidate94.5%Regressions2

Pass rate by intent Illustrative counts

Block if any intent falls > 2 pts
IntentConversationsv47 livev48 candidateΔ pts
Order status62096.1%97.3%+1.1
Returns & exchanges54095.9%97.0%+1.1
Store stock41095.4%96.3%+1.0
Delivery windows33094.5%95.5%+0.9
Price match28093.9%86.1%−7.9
Warranty claims22092.7%85.9%−6.8
All intents2,40095.2%94.5%−0.7
The total moves under a point; two intents fall several. The gate reads each intent, not the average.

Scored overnight

5 changes
support-agent/answerv48Support agent · Shorter policy preambleHeld
associate/answerv23Store associate · Aisle lookup wordingPassed
product-qa/answerv31Product Q&A · Spec table formattingPassed
guardrail/sign-offv9Support agent · No change · nightly replayPassed
Verith M3 updateproviderAll answer prompts · Model change replayPassed

Grounded check · conversation PM-1184

Price match · verified against retrieved source, not model-graded
“Will you match the price I found on a marketplace seller?”
v48 answerYes — we match any online price, including marketplace sellers, within 14 days.
Retrieved source · price-match-policy.md §2Matches apply to named national retailers. Third-party marketplace sellers are excluded.

Release gate

v48 → production
Frozen set replayed2,400 / 2,400
Ungrounded assertions17 vs 3 live
Per-intent rule2 intents > 2 pts
Rollback targetv47 stays live
On screen

A prompt version scored overnight against 2,400 recorded conversations: pass rate by intent with the two regressions picked out in red, the grounded check that failed against its retrieved source, and the gate that held the release.

Token spend · August 2026
Where the month's inference spend went, feature by feature, after the cache and prompt compression took their cut
Export CSV Thu, 09/17/2026RA
Model cost · monthly inference
−21%Against the three months before the retainer
Where the cut came from
Cache + prefixesSemantic cache and cached prompt prefixes
Answer model
Verith M3Unchanged · no cheaper-model trade

Spend by feature Three months before August

Bars to shape · figures in the monthly report
Feature · systemMonthly spendLevers applied
Customer chat answersSupport agentCachePrefixFew-shot cut
Product questionsProduct Q&ACachePrefix
Aisle & stock answersStore associatePrefix
Answer sign-offSupport agentPrefix
Catalog retrievalProduct Q&A—
Nightly forecast notesDemand forecasting—

Semantic cache · similarity threshold per surface

A support answer tolerates near-matching; a legal one does not
SurfaceThresholdBehaviour
Order status & delivery0.92Near-matches reuse the answer
Store stock & aisles0.94Keyed per store as well
Product questions0.95Keyed per product
Returns policy wording0.98Almost exact only
Warranty & legal termsOffNever served from store

Prompt compression

support-agent/answer
Stable instructions → cached prefixTone, policy, format · cachedQuestion + context · billed
Few-shot examplesex1ex2ex3ex4ex5ex6ex7ex85 cut · replay showed no effect on the frozen set

Token & Cost Optimizer

On screen

Where the month's token spend went, feature by feature, after the semantic cache and prompt compression took their cut: similarity thresholds per surface, stable instructions in a cached prefix, and the few-shot examples that were cut.

Semantic caching layer and prompt compression that lowers monthly inference bills.

A semantic cache sits in front of the model, so a question already answered this hour returns from store instead of being billed again, with the similarity threshold tuned per surface, because a support answer tolerates near-matching and a legal answer doesn't. Prompts were compressed by moving stable instructions into cached prefixes and cutting few-shot examples that measurement showed were doing nothing.

What shipped
  • Semantic cache with per-surface similarity thresholds
  • Stable instructions moved into cached prefixes
  • Few-shot examples cut where measurement showed no effect

24/7 Model SLA Monitoring

Prometheus & Grafana dashboards tracking latency, drift, and accuracy metrics in real time.

Latency, cost per request, retrieval hit rate and eval scores are exported to Prometheus as first-class metrics and alerted on with thresholds agreed as an SLA, never chosen by whoever built the dashboard. Drift is watched on the input distribution as well as the output, because the first sign of trouble is usually that people started asking a different kind of question.

What shipped
  • Latency, cost, retrieval hit rate and eval scores as SLA metrics
  • Alert thresholds agreed contractually, never picked ad hoc
  • Input drift watched alongside output quality
Latency & SLA · live
Answers against the 3-second target hour by hour · exported to Prometheus, alerted on thresholds agreed as the SLA
All SLA metrics within threshold Thu, 09/17/2026RA

p95 answer time by hour · all answer models

Scale 0–4s · bars to shape
3.0s target
12 AM4 AM8 AM12 PM4 PM8 PM11 PM

Agreed SLA metrics

Schedule 2 of the retainer
Answer latency · p95≤ 3.0s
Uptime · per system≥ 99.9%
Eval score · frozen setFloor per intent
Retrieval hit rateFloor per surface
Cost per requestCeiling per feature

Questions as they land

Store and online · last 40 seconds
TimeChannelQuestionAnswered
10:14:52 StoreDo we carry the 20V hedge trimmer at Maple Grove?2.3s
10:14:49 OnlineWhere is my order FB-448210?Cache0.2s
10:14:47 OnlineIs the teak bench finish safe for covered porches?2.0s
10:14:41 StoreWhich aisle are the raised-bed kits in?1.6s
10:14:38 OnlineCan I return a patio umbrella without the box?Cache0.3s
10:14:33 OnlineWill you price match a marketplace listing?2.6s
10:14:30 StoreCustomer wants delivery windows for Saturday1.9s
10:14:26 OnlineWhat does the grill warranty cover on burners?2.4s
10:14:21 OnlineIs the 6-person tent in stock near 55311?1.2s
10:14:17 StorePrice match request for a competitor flyer2.2s

Input drift · question mix

This week vs four-week baseline
Order status
Returns
Store stock
Price match
Warranty
Price-match questions rising before any answer score has moved

Alerts today

Grafana → on-call
Input drift · price matchOpened 7:05 AM · into sprint
Eval gate · v48 held3:52 AM · no customer impact
p95 approaching targetNone today
On screen

Answers against the three-second target hour by hour, the SLA-agreed metrics, the live stream of questions as they land in store and online, and drift watched on the question mix.

Working inside their operation

  1. 01

    A cross-functional team of 6 worked on an ongoing retainer, covering Monitoring, Model tuning, Continuous improvement. We ran a standing mid-week checkpoint and written decisions in place of status meetings. Nothing shipped without a live demo first.

    The cadence is the product: continuous evaluation, daily cost and latency review, a tuning sprint every two weeks, and a monthly report on quality, spend and edge cases. The monthly report goes to the business, not to engineering, because the decisions that matter (what to fix, what to cap, what to retire) are commercial ones.

What changed in the runbook

  • Continuous prompt tuning and evaluation maintained 99.8% production uptime.

    Uptime was never the hard part. The retainer earns its place by catching the quality regressions that leave uptime untouched, which is exactly what the deprecation incident was.

  • Semantic caching and token optimization reduced monthly LLM API costs by 32%.

    Most of the cost reduction came from the cache and from prompt prefixes, not from a cheaper model. The quality trade nobody wanted was never necessary.

  • Proactive model drift detection prevented quality decay as underlying models updated.

    Watching the input distribution is what makes drift detectable early: people start asking a different kind of question before the answers visibly get worse.

One version, one night

The average said ship it. Two intents said no.

A candidate prompt for the support agent, replayed overnight against 2,400 recorded conversations: the run, the regressions the per-intent rule caught, and what the gate did about them. Switch tabs, or use the arrow keys once one is focused.

support-agent/answer · v48 · overnight 09/17/2026Release held
Intentv47 livev48 candidate
Order status620 / 62096.1%97.3%
Returns & exchanges540 / 54095.9%97.0%
Store stock410 / 41095.4%96.3%
Delivery windows330 / 33094.5%95.5%
Price match280 / 28093.9%86.1%
Warranty claims220 / 22092.7%85.9%
Frozen set replayed2,400of 2,400 recorded conversationsReplayed 1:15–3:52 AM. Replay speed here is illustrative; pass counts per intent are illustrative too.
Architecture

Every change passes the same set; every answer stays on a metric someone agreed to

The four systems all had infrastructure monitoring, and it watched uptime. This is the loop wrapped around them so quality is watched too, from the question as it lands to the decision the business makes each month.

  1. 01 · Traffic
    Live questions, store and onlineA sample of live traffic is scored continuously after every release, and the input distribution is watched as well as the answers.
  2. 02 · Gate
    Frozen set · 2,400 conversationsReplayed against every prompt or model change before it ships. A regression blocks the release outright.
  3. 03 · Serve
    Semantic cache + cached prefixesA question already answered returns from store, with the similarity threshold tuned per surface; stable instructions sit in cached prefixes.
  4. 04 · Watch
    Prometheus + GrafanaLatency against the 3-second target, cost per request, retrieval hit rate and eval scores exported as metrics, alerted on SLA-agreed thresholds.
  5. 05 · Decide
    Fortnightly sprint · monthly reportTuning runs against the edge cases each two-week cycle produced; the monthly report on quality, spend and edge cases goes to the business.
So quality can’t erode unseen

Blocked releases, grounded checks & drift on the inputs

A regression can't ship quietly

Every prompt or model change replays the frozen evaluation set first, and a regression blocks the release instead of opening a ticket someone might read later.

Hallucination checks are grounded

Where possible an assertion is verified against the source the system retrieved, so the detector isn't another model's opinion of the answer.

Drift shows before answers degrade

The input distribution is watched alongside output quality. The incident that started this went unnoticed for 11 days because only uptime was monitored.

Would you know if your AI’s answers got worse last week? Scope your build in 3 minutes.

Scope your build
Have a project?

Let's talk

Running a large platform, shaping a first MVP, or getting a product ready for a funding round? Tell us where you are. We'll shape the process around it, and stay with you after launch.