Key takeaways
- Record source, date, product state, locale, rating, incentive disclosure, verification status, and retrieval method with each review.
- Separate direct observation from interpretation and promotional copy; do not paraphrase criticism into praise.
- Deduplicate syndicated reviews, detect suspicious patterns, and preserve disagreement instead of reporting a synthetic consensus.
- Review text can generate hypotheses and language; platform samples rarely justify population percentages.
01
The corpus was shaped before the analyst downloaded it
A platform’s users, moderation, ranking, verification, incentive rules, product mix, and time period determine which voices appear. Then the research team filters by rating and language. A theme count from that corpus is not automatically a measure of the whole customer base.
Create a source record before coding: platform, URL or stable ID, retrieval date, review date, market, language, product version, rating, verified status where defined, incentive disclosure, response from the business, and collection method. Respect platform terms and avoid collecting profile details that the decision does not need.
Swipe to compare every column
| Layer | Preserve | Risk if omitted |
|---|---|---|
| Platform | Rules, ranking and audience | Cross-platform counts treated equally |
| Review | Exact excerpt and context | Paraphrase changes meaning |
| Commercial link | Incentive or relationship disclosure | Promotion mistaken for independent experience |
| Product state | Version, date and market | Old failure presented as current |
02
Remove duplicates without erasing repeated experience
Syndication can place the same review on several sites. Near-identical campaigns can create many entries that are not independent. Use stable IDs, exact and fuzzy text matching, timestamps, and source relationships to mark likely duplicates. Keep the raw record and document the rule.
Repeated complaints from different identifiable transactions may be an important signal; repeated copies of one complaint are not. Suspicious reviews belong in a separate status, not silently deleted because they complicate the story.
03
Code claims, contexts, and counterexamples
Separate what the reviewer reports happened, the interpretation they attach to it, the emotion, the context, and the desired resolution. Keep positive and negative cases inside the same theme. A slow setup for one integration and a fast setup for another may reveal a dependency that “easy onboarding” hides.
CMA guidance now treats fake and concealed incentivized reviews as banned practices in the UK and expects publishers to take effective steps. FTC endorsement guidance similarly makes material connections relevant. Laws differ by market; provenance and visible disclosure are useful operational controls everywhere.
04
Publish the method beside the insight
State sources, retrieval window, filters, languages, number of records before and after deduplication, coding method, incentive handling, and known gaps. Use review evidence to improve a brief, support script, product queue, or test—not to fabricate a “voice of the market.”
If AI helps cluster text, validate clusters against sampled reviews and keep the model, prompt, version, and review notes. Never publish a generated quotation. A real awkward sentence with permission is better evidence than a clean line no customer said.
Primary sources and further reading
Use the source material to validate details against your own context and current platform configuration.
- CMA: Fake reviews guidance
- FTC: Endorsement Guides—What People Are Asking
- FTC: How to evaluate online reviews
This field note follows the XenGrowth editorial policy: primary sources where available, visible limitations, material review dates, and no invented first-hand experience.
Stay with the problem



