We scored thousands of Fin answers: 5 patterns that tank quality
An in-house benchmark of Fin Agent answer quality, reply by reply. The 5 most common failure patterns — and how to catch them before your CSAT does.
Everyone watches Fin's resolution rate. Almost no one watches the quality of each answer. We did the opposite: scored replies one by one, on a constant rubric. Here's what surfaced.
At ResponZ we score every reply — Fin and human — out of 100 across four axes: relevance, accuracy, tone & policy, completeness. Not an absolute score: one relative to what the support team declares it expects, with universal heuristics as a fallback. Aggregating those scores across analyzed accounts, five failure patterns keep coming back.
Method in one line
Score = 100 − Σ penalties (relevance 0→−40, accuracy 0→−30, tone 0→−20, completeness 0→−10). Same rubric for Fin and humans, to compare like with like.1. The plausible… but stale answer
Pattern #1, by a wide margin. Fin cites the knowledge base with confidence — except the article is 14 months old and the product has changed since. The customer gets a confident, wrong answer: the worst of both worlds. Invisible in the resolution rate (the conv closes) but paid for in day-7 reopens.
Signal to hunt: replies whose accuracy score drops right after a product change. That's exactly where cross-referencing recent commits/PRs against Fin answers pays off.
2. The generic answer that ignores the real question
The customer asks a precise three-part question; Fin answers beside the point with a close-but-not-exact article. Relevance score in free fall. This pattern explodes on multi-intent questions — the ones a human would have split up.
3. Right content, wrong tone
A factually correct answer, delivered in a robotic tone to a visibly frustrated customer, without the slightest acknowledgment of the friction. Technically “resolved,” relationally blown. This is the penalty that separates support that retains from support that churns.
We thought Fin handled billing well. The score said otherwise: correct on substance, ice-cold on delivery — on exactly the most sensitive topics.
4. The unkeepable promise
“A refund will be processed within 48h.” Except no concrete action was triggered and internal policy doesn't allow it. Fin commits the company to a promise it can't keep — the most expensive accuracy penalty downstream.
5. The loop: the same answer, again and again
Across three successive exchanges, Fin returns a variation of the same answer while the customer rephrases, increasingly annoyed. Each repeat digs the score deeper. It's the most reliable signal that escalation was due two messages ago.
What these patterns share
None shows up in the resolution rate. All show up in a per-answer score. Fin quality isn't a reassuring average — it's a distribution with a long tail of toxic answers nobody reads.
How to bring it under control
- Declare your reference. Tone, rules, sensitive topics: without them, you score blind on generic heuristics.
- Audit the tail, not the average. The 5% of replies under 30/100 cause 80% of the relational damage.
- Cross-reference with the product. A recent code change touching a topic = suspicion of stale answers.
The 30-second test
Connect Intercom to ResponZ and let it score your last 50 conversations live. You'll immediately see your tail — the Fin answers dragging your CSAT down with nothing to flag them.The resolution rate tells you Fin closed the conversation. The per-answer score tells you whether it closed it well. The two don't tell the same story — and it's the second one that decides whether your customers come back.