AI Review Highlights
Amazon's AI Review Highlights were built to turn thousands of opinions into clarity. Instead, customers stared at ambiguous icons and scrolled past interactive elements they didn't know were there. I rebuilt the experience so customers could understand it, trust it, and act on it, and turned the process into a framework Amazon now uses to evaluate AI experiences more broadly.
- Three-state sentiment systemDistinct icons for positive, mixed, and negative make sentiment legible at a glance.
- Mention counts on every tagNumbers signal weight and volume, so customers know what most reviewers actually said.
- Clear affordance to exploreBlue link styling and a 'Select to learn more' cue signal that tags are tappable.
- Evidence one tap awayA bottom sheet reveals sentiment split and source quotes, grounding the AI summary in real reviews.
- Three-state sentiment systemDistinct icons for positive, mixed, and negative make sentiment legible at a glance.
- Mention counts on every tagNumbers signal weight and volume, so customers know what most reviewers actually said.
- Clear affordance to exploreBlue link styling and a 'Select to learn more' cue signal that tags are tappable.
- Evidence one tap awayA bottom sheet reveals sentiment split and source quotes, grounding the AI summary in real reviews.
AI was generating insights customers didn't understand
Amazon's AI review summaries helped customers digest thousands of reviews at a glance, but a critical layer of that experience was failing quietly underneath it.
Customers struggled to:
- Interpret sentiment states
- Recognize interactive review filters
- Understand how the AI arrived at its conclusions
This limited both trust and engagement with one of the feature's most valuable capabilities, the ability to go deeper than the summary itself.
Accelerate customer decision-making by surfacing trustworthy insights that improve purchase confidence and overall shopping satisfaction.
Three signals were breaking customer trust
Legacy CX (Control) · Customer Reviews section of the Product Detail Page (PDP)

| # | Signal | Problem | Customer impact |
|---|---|---|---|
| 1 | Low icon comprehension | The previous experience only represented positive sentiment. With no mixed or negative signaling, customers frequently conflated the two states or assumed information was missing. | Lower trust in AI |
| 2 | Low interactivity | Aspect tags looked like metadata instead of tappable controls. | Missed review exploration |
| 3 | Low discoverability | Customers rarely used the AI tags to explore deeper review content, leaving the summary's richest value untouched. | Less engagement |
Designing inside multiple constraints
Before designing anything, I validated the existing system

Baseline user study
I designed and executed a user research study on usertesting.com to establish ground truth on current icon comprehension. I documented confusion rates around the absent mixed and negative sentiment states, then synthesized the findings into an actionable report, the evidentiary foundation for every downstream design decision and stakeholder alignment conversation that followed. The study was structured as a scenario-based usability test, asking participants to evaluate an air fryer purchase while I collected both quantitative (Likert-scale) and qualitative (open-ended) feedback.
Customers frequently misinterpreted or overlooked mixed and negative sentiment states because only positive sentiment had a clear icon, revealing that the icon system lacked a clear, cohesive mental model.
- The previous experience only represented positive sentiment. With no mixed or negative signaling, customers had no way to distinguish between those two states. Why it mattered: customers could form incorrect impressions of a product's strengths and weaknesses, directly affecting purchase confidence.
- The absence of mixed and negative icons wasn't read as 'mixed' or 'negative'. Both states were read as missing information, a loading error, or the same as each other. Why it mattered: customers decode icons as a system, not independently, so missing states undermined the one that was present.
- Only positive sentiment (the green checkmark) was reliably understood, since it leveraged an existing, universal convention. This confirmed the failure was inconsistency across the system, not sentiment communication itself. Why it mattered: ambiguity in two of three states was enough to erode trust in the AI-generated summaries as a whole.
Customers perceived aspect tags as passive labels rather than interactive controls, resulting in low engagement with deeper review exploration.
- Aspect tags were not recognized as interactive controls. Participants frequently read them as descriptive metadata rather than something tappable. Why it mattered: one of the feature's primary value propositions, drilling into reviews by attribute, was largely hidden from customers.
- Participants lacked clear affordance cues. The visual treatment didn't sufficiently signal that the tags were clickable, so few participants expected tapping to do anything, and several were surprised to learn the tags opened filtered reviews. Why it mattered: customers can't use functionality they don't perceive exists.
Most customers consumed the AI summary without exploring the underlying review evidence, limiting transparency and reducing trust in the AI's insights.
- Customers rarely explored beyond the summary itself, tending to read it and continue their purchase decision without investigating supporting review content. Why it mattered: the richest part of the experience, validating AI insights against real reviews, went underutilized.
- The feature lacked transparency and evidence. Because customers didn't discover the deeper review pathways, many were left wondering how the AI arrived at its conclusion. Why it mattered: limited discoverability directly contributed to lower trust.
- Customers needed stronger signals that additional information was available. Deeper review exploration needed to feel intentional, accessible, and obviously connected to the AI summary. Why it mattered: better discoverability would increase both engagement and confidence in the AI's recommendations.
Competitive audit
I conducted a competitive audit of 20+ leading e-commerce platforms to understand customer expectations, identify successful review patterns, and establish a UX foundation for introducing AI-generated insights in a way that felt intuitive and trustworthy.




No established industry standard exists for communicating positive, mixed, and negative sentiment as a cohesive system.
Many platforms leaned on text labels, stars, or ratings rather than iconography alone, suggesting sentiment is difficult to communicate through abstract visuals by themselves.
Common metaphors (thumbs, emojis, plus/minus) carried multiple meanings across contexts, creating semantic ambiguity and room for misinterpretation.
Five decisions that shaped a two-week sprint into a system
Defined the problem before designing the solution
The initial request was straightforward: introduce mixed and negative sentiment icons. Rather than optimizing single states, I stepped back to evaluate the sentiment system as a whole. Because only positive sentiment had a clear, established icon, simply adding two more icons risked reinforcing an inconsistent experience. I recommended validating customer comprehension first so the solution addressed the root problem instead of isolated symptoms.
Outcome: Shifted the project from designing one icon to designing a cohesive sentiment system grounded in research.Reduced the solution space through research and systems thinking
I explored multiple approaches for representing mixed and negative sentiment, including flat, squiggly, plus/minus, thumbs, and facial expressions. Each was evaluated for customer comprehension, cultural scalability, and compatibility with Amazon's design system. Rather than treating the question as "Should we add mixed and negative icons?" I reframed it as "Which complete sentiment system is easiest for customers to understand?" The strongest candidates were incorporated into the same comprehension study, allowing the team to evaluate the entire system instead of individual symbols.
Outcome: Established a research-backed process for evaluating complete sentiment systems rather than isolated icons.Balanced customer comprehension with design system consistency
One leading concept reused the Rio checkmark for positive sentiment. While it had performed well previously, I identified a semantic conflict with Rio's existing success-alert pattern that could introduce ambiguity as the feature scaled. Rather than treating this as a design preference, I documented the tradeoff, surfaced the risk to stakeholders, and ensured it was considered alongside usability findings during decision-making.



Structured experiments to maximize learning
The team initially considered testing iconography and interactivity together. I recommended separating these questions into sequential experiments so each result could be attributed to a single variable. This approach prioritized learning over speed, ensuring every experiment produced actionable evidence that could inform future iterations.
Outcome: Established a repeatable experimentation framework that was later reused across AI review experiences.Expanded the experiment to answer the right questions
The original testing plan couldn't accommodate every cohesive icon system within the standard treatment limit. I developed a data-backed recommendation outlining the statistical and design tradeoffs and partnered with stakeholders to expand the experiment. This allowed the complete icon system comparison to run in a single study while preserving methodological rigor.
Outcome: Increased experiment coverage without compromising result quality and accelerated subsequent design decisions.Three experiments, each building on the last
Icon comprehension
Tested which icon system was most cohesive and comprehensible to customers. Positive and negative icon treatments (checkmark, checkmark with circle, diagonal arrows, diagonal arrows with circle, vertical arrows, vertical arrows with circle) were crossed with the two mixed sentiment finalists (flat and squiggly), all presented with the same interactive style from the legacy cx to isolate the icon variable alone.
| Control | T1 | T2 | T3 | T4 | T5 | T6 | T7 | T8 | T9 | weblabT10 | T11 | T12 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Positive | |||||||||||||
| Mixed | — | ||||||||||||
| Negative |
Diagonal arrows paired with the squiggly mixed sentiment icon won (T10). This produced the most cohesive and comprehensible system overall, outperforming the checkmark despite leadership's early confidence in it, and validating the systems conflict concern raised in Decision 3.
Interactivity
Using the winning icon set from Experiment 1 as the control, tested whether presenting aspect tags as links (the current, universally understood interactive pattern) or as buttons (the pattern most common in the competitive audit) drove stronger engagement.



The, link-style interactivity (T1), won. Customers responded to the established, universally understood affordance of a link more than a button-style treatment, closing the door on a planned follow-up experiment that would have tested color as a discoverability lever for buttons.
Transparency
Using the winning treatment from Experiment 2 as the control, tested whether adding a visible mention count (for example, surfacing how many reviews referenced a given attribute) increased customer trust, reinforcing that the summary was grounded in real review volume, not fabricated.
The mention count treatment won. Making visible evidence counts increased both transparency and trust signals, confirming that discoverability of underlying evidence mattered as much as icon clarity.
Outcomes were evaluated against Amazon's standard set of engagement, conversion, and satisfaction metrics for the Customer Reviews surface.
The procedural outcome
The biggest takeaway from this project was the importance of disciplined experimentation when navigating ambiguity. Given the business constraints, the fastest path would have been to move forward with the strongest-performing treatment, but isolating variables allowed us to understand what was actually driving comprehension, trust, and user behavior. The final solution was only one outcome. The more lasting impact was the decision-making approach behind it: using research and evidence to move conversations beyond opinions and toward clearer product decisions. This mindset continues to shape how I approach AI experiences, where building trust requires understanding not just what works, but why it works.
Two week timeline
Four treatment experiment limit