
What We Changed After Reading Every Detractor Comment for a Month
A month ago I stopped looking at our NPS dashboard and started reading every detractor comment individually, one at a time, instead of letting a summary tell me what they said. We are still early enough that this was possible to do by hand, and doing it by hand is exactly why it was worth doing.
Most teams never do this. The dashboard shows a score, maybe a trend line, maybe an AI summary if the tool has one, and the individual sentences that produced that score get skimmed at best. I wanted to know what we would learn if nobody skimmed anything for thirty days.
What I expected to find
I expected the comments to be mostly about missing features. That is the easy story every founder tells themselves, that detractors are unhappy because the product does not do enough yet. It is a comfortable story because the fix is always "build more," which is a familiar problem with a familiar kind of solution.
That is not what the comments mostly said.
What the comments actually said
Reading them one by one, a different pattern showed up. A meaningful share of the detractor comments were not about anything the product could not do. They were about something the product did not make clear, a setup step that was ambiguous, a result that showed up but was not explained well enough for the person to trust it, an expectation we had set in a landing page or an email that the actual product experience did not quite match.
Feature gaps were in there too, but they were not the majority of what people were actually frustrated about. The bigger pattern was a gap between what someone expected to happen and what they understood to be happening once they were inside the product. That is a communication and clarity problem, not a roadmap problem, and it is a much cheaper thing to fix.
I want to be careful here about what I am and am not claiming. Our user base right now is still small, measured in dozens of active accounts, not thousands, so this is a pattern from a real but limited sample, not a statistically rigorous study. Treat it as a founder's honest read of a small dataset, not a benchmark anyone else should cite.
A rough taxonomy, for anyone doing this themselves
Reading comment by comment, most of what we saw sorted into a few buckets. This is not a scientific framework, just what became visible after enough individual reads.
Expectation mismatch. The comment describes something that did not match what marketing, onboarding, or a sales conversation implied would happen. This was the largest bucket by feel, not by precise count.
Missing clarity, not missing feature. The capability existed, but the person did not realize it, could not find it, or did not trust the output enough to use it. Several of these were one clarifying sentence away from being a non-issue.
Genuine feature gap. The thing the person wanted to do was simply not possible yet. This was real, but smaller than I expected going in.
Price or plan friction. A handful of comments were really about cost relative to perceived value, not about the product experience itself.

What we actually changed
We did not add a single new feature because of this exercise. Every change was to clarity. We rewrote two onboarding steps that were technically correct but assumed knowledge a new user did not have yet. We added a short explanation next to a result in the product that people had been reading correctly but not trusting, because nothing told them how it was calculated. We adjusted language on a landing page that was setting an expectation the first-week product experience did not fully deliver on yet, closing that gap from the marketing side instead of only from the product side.
None of these were big engineering efforts. All of them came directly from sentences a dashboard summary would have compressed into a generic line like "users want clearer onboarding," which is technically true and almost useless, because it tells you a category exists without telling you which two sentences actually explain it.
Why this is worth doing even at small scale
The instinct at any stage is to wait until you have enough volume to justify a "real" analysis. I think that instinct is backwards. At small scale, reading every comment by hand is still possible, and the patterns you find are still real, just from a smaller and more honest sample than a growth-stage team could manage. Waiting for scale to do this properly just means shipping on a dashboard summary for longer, which is exactly the thing that misses the two sentences that would have told you what to fix.
We built the AI Insights feature in Elvan because most teams will never have the time to read every comment by hand the way I did this month, and a plain-English summary of open-text feedback gets you most of the way there without the manual effort. But doing it manually once, even briefly, changes what you look for afterward. I would not have known what to ask the summary to surface if I had not read the raw sentences first.
