[ Case study ]
Tier-1 support was drowning in the same forty questions — billing cycles, seat management, integrations — while genuinely complex tickets waited days. The team considered a chatbot and feared the classic outcome: a wall of confident nonsense.
CLIENT a project-management SaaS vendor — FOCUS Ground it in a cleaned knowledge base
Representative examplesEvery case study in this library is an illustrative composite of the kind of engagement we deliver — written to show our method and standards, not to name clients.
Software is the product and support is the promise: the client sells a project-management tool to small teams, and its reputation lives or dies on how fast a confused customer gets unstuck. The support team is small, senior, and proud of the complex tickets they resolve — the migrations, the broken integrations, the edge cases. Below that sits a mountain of routine volume that grew with every pricing-page change and every new integration shipped.
Tier-1 support was drowning in the same forty questions — billing cycles, seat management, integrations — while genuinely complex tickets waited days. The team considered a chatbot and feared the classic outcome: a wall of confident nonsense.
We proposed a grounded assistant that answers only from a consolidated, versioned corpus of their own docs, cites what it used, and escalates billing changes, contract questions, and anything account-specific to humans by rule. The refusal rules were written and tested before the personality was, because their constraint was trust, not capability. An eval set with expected answers ran before launch and runs weekly after, so quality drift is measured rather than discovered by complaint. The Zendesk desk stays the system of record; the assistant files context-rich tickets rather than replacing the tool the team lives in.
Just as important is what we ruled out, and why:
The three doc sources were consolidated, deduplicated, and versioned; the assistant answers only from that corpus and cites what it used.
Billing changes, account-specific questions, and anything about contracts escalate to humans by rule — the assistant knows what it doesn't know.
A question set with expected answers ran before launch and runs weekly after, so quality drift is measured rather than discovered by complaint.
Delivered by the systems pod — automation specialist + engineer over 8 weeks, with working increments reviewed with the client every week.
Obstacle
Two conflicting billing articles — one pre-dating a pricing change — produced the only seriously wrong answer in testing, and nothing in the corpus marked one as authoritative.
Handled: We consolidated the corpus before further configuration: deduplicated, versioned, and given a source-of-truth rule per topic, then re-ran the eval set clean.
Obstacle
The support lead's weekly transcript review was proving impossible — she was skimming, not reviewing, because transcripts arrived as raw dumps every Monday.
Handled: We reshaped the export around her actual questions — refusals, escalations, low-confidence answers first — and the review became fifteen focused minutes.
The headline: tier-1 tickets resolved without human touch, month two post-launch versus month before — 44% → 71%, read from Support desk reports. A second check: median response time for escalated (complex) tickets at 2.1 → 0.4 days.
The senior team got their complex tickets back, and the mood in the support channel shifted from triage to craft. Customers quote the assistant's cited answers back accurately now, which quietly improved the questions they ask. The weekly eval has become the support lead's early-warning system — she flagged a doc gap to product twice before any customer complained. Nobody on the team is guarding their job from the assistant; they treat it like a junior colleague who cites their sources.
The result was read from Support desk reports against the pre-engagement baseline over the stated window, with a guardrail check on median response time for escalated (complex) tickets. Where platform-reported numbers and business outcomes differ, this record says which layer it is quoting.
What we would do differently
We would have cleaned the docs before configuring anything — two conflicting billing articles were the source of the only serious wrong answer in testing.
[ Related service ]
[ Related builds ]
45 min 4 minAverage handling time per invoice batch (human review only), measured over the first full month
4 hrs 25 minAverage policy-comparison preparation per client file, verified over 40 files
[ Next step ]
Next case study