*Source: Out of the Blue, published case study.
Overview
Out of the Blue is an eCommerce observability platform. It watches a brand's entire stack, storefront, ads, martech, checkout, and tells growth teams the moment something changes, whether that's a business metric moving (Northstar) or the site breaking (Pulse).
A GMV spike and a broken checkout pixel are genuinely different problems, different audiences, different urgency, different next actions. What they share is a harder problem sitting underneath both: how do you represent anomalous data so a human can understand it fast, regardless of what kind of anomaly it is. I designed the foundational taxonomy that solved that shared representation problem once. Insights and Errors each built their own product decisions, their own hierarchy, their own content, on top of it, as two genuinely separate products that happen to inherit a consistent, legible way of representing data underneath.
That taxonomy is still shaping both products today, visible in the live Insights and Errors surfaces. The business outcomes tied to this kind of clarity, with full numbers and sourcing, are covered in Impact below.
My role: End-to-end Product Designer
(Research through shipped UI, across both surfaces, with decisions tied directly to the business outcomes covered in Impact below.)
Impact Snapshot
{Systems}
One taxonomy, reused: a shared foundation for representing metrics and events, adopted by two independently designed products, Northstar and Pulse, instead of each inventing its own from scratch.
{Research}
Grounded in direct research, sitting in on CX conversations, reading through support tickets, and scheduling my own UXR calls with early customers, including the one behind the $19K/month figure below.
{Craft}
Replaced a single flat paragraph carrying five kinds of information at equal visual weight with a structured five-part hierarchy, still live in Out of the Blue's product today, serving as a foundation and shaping both Northstar and Pulse independently.
{UX Outcome}
Metric movement made legible before a word is read: percentage change and dollar impact pulled out and bolded above the summary sentence, not buried inside it.
{Business Impact}
The business case this foundation was built to serve, per Out of the Blue's published results: $19K/month recovered for one customer, 400% more monitoring coverage for another, ~$2M per incident cited industry wide as the cost of an undetected pixel failure. (source: outoftheblue.ai/case-study/case-study-tbs)
The Problem
Two products, same underlying failure. A GMV anomaly on the Insights side and a JavaScript error on the Errors side both landed on screen as one flat block, forecast, cause, and consequence, all carrying equal visual weight. A user scanning a feed of fifteen cards couldn't tell, in half a second, which one actually needed their attention right now.
I didn't arrive at that diagnosis by guessing. Sitting in on CX conversations and reading through support tickets surfaced the same complaint on both surfaces, worded differently each time. On Insights, growth leads kept asking some version of "why did this move, and is it actually bad." On Errors, ops managers kept asking "is this live right now, and do I need to drop what I'm doing." Different products, same underlying question: what am I looking at, and does it need me right now. That pattern repeating across two otherwise unrelated products is what told me the real issue wasn't the copy. It was that nothing had decided, ahead of the words, what a user needs to see first, second, and third.
My own working notes from that period name the diagnosis directly:
Fix the flow of information. Add contextualization to the data. Improve the interconnectivity between components. Bring consistency across every metric under a grouped insight.
Neither product is named in that list, and that's not an accident. That's the moment the problem stopped being "the Insight card is confusing" and became "we don't yet have a foundation for representing anomalous data that a human can understand fast." Fixing that foundation didn't mean solving both products at once. It meant building something solid enough that Insights and Errors could each design their own hierarchy, tone, and next actions on top of it.
Finding the One Question Underneath Both Products
I did not have a formal research budget for this, and I want to say that plainly rather than dress it up. What I had was direct, active involvement: sitting in on CX conversations, reading through support tickets, and scheduling my own calls with early customers, including two who later became the subjects of Out of the Blue's own published case studies.
The harder part wasn't collecting input, it was arriving anywhere useful from it. A support ticket about a confusing GMV card and an offhand comment on a call about a checkout error looked like two unrelated complaints from two unrelated products. My job was treating them as the same signal wearing different clothes, not as two separate feature requests to log and prioritize separately.
The clearest example of that synthesis is the Impact Type table, covered in full under Design Decisions. Customers weren't consistently asking "what kind of error is this," a technical framing that would have organized the table by category, code type, source system. They kept asking a plainer question, worded differently every time but pointing at the same thing: "do I need to do something right now." That repeated question, heard across support tickets on Errors and offhand comments on Insights calls, is what pushed the table away from a technical taxonomy and toward one ordered by urgency and actionability instead, Biz-Ops, Technical, Warning, each carrying a predetermined answer to "can I act on this."
That reframe, from "what kind of thing is this" to "do I need to act," is the one piece of process work underneath everything else in this case study. Every hierarchy decision that follows is built around that one question, because customers kept asking it, over and over, in their own words.
The Foundation: Building Blocks
Before either card could be designed well, the underlying data had to be structured well. This is the part of the work that doesn't show up on screen, and it's the part that made everything downstream possible.
I built a taxonomy, what I called the Building Blocks, that defined, once, how any metric or event gets represented regardless of which product surface it appears on:
Metric Hierarchy
Every metric gets ranked by dollar value and percent contribution to total impact, Metric M1, M2, M3, in descending order of business weight. A metric's Category, Funnel stage, and Source are attached at this level, so a card never has to guess how to label or group a number, it inherits that from the taxonomy.
Numeric Representation
Every value gets compared against the same set of reference points: what actually happened (actual value), what happened last time (previous value), and what was expected to happen (expected or threshold value). The gap between actual and expected (deviation) is what tells the system something moved, and a defined upper and lower limit (range) flags when that gap is too large in either direction, not just too low, but also suspiciously too high. This is what makes "GMV increased 102% from expected value" and "8 infrastructure errors, duration 7h 15m" possible to render with the same underlying logic, both are just this framework applied to a different metric type.
Impact Type Logic
Every anomaly gets classified as Biz-Ops, Technical, or Warning, and each classification carries a predefined answer to two questions: can the business user act on this directly, and what is the recommended next action. This single table is what lets the Errors card show a concrete "Pause TikTok ads" suggested action on one card and a "no action, it's just a warning" note on another, without a designer hand crafting that judgment call per instance.
Event Lifecycle Fields
Start time, end time, duration, status, active or closed, provider, source. Defined once, and reused as a starting point whether the event is a marketing metric anomaly or a site error, though how each field actually surfaces was still decided separately on each card. The live product shows the difference plainly: an Errors card displays Start of Event, End of Event, and Duration as separate labeled fields, because a technical responder needs forensic precision. An Insights card collapses that same set of fields down to a single date and a Live or Closed tag, because a business user only needs to know whether it's still happening, not the full timestamp breakdown.
Persona Routing
The taxonomy explicitly branches into a Biz Person View and a Tech Person View, with further splits for whether a dollar value crosses a defined threshold or falls under a non-business, technical, or warning impact type. That branching logic, business view versus technical view, only had to be invented once. What each surface did with it, the actual hierarchy, tone, and layout for its own audience, was still designed independently on Insights and on Errors.
Two Products, Two Kinds of Uncertainty
With the taxonomy in place, I explored how it should render as an actual card. The two surfaces needed genuinely separate exploration, not one process run twice, because the uncertainty each one was designed around was different.
Insights: three variations, one clear reader
Variation 1, the quick story view. One dominant number pulled straight from the Building Blocks' numeric representation layer, one short causal sentence pulled from the Impact Type's next-action logic, almost no new UI. Built for a user scanning a feed, not studying one card.
Variation 2, the game of numbers. Primary and secondary metrics from the M1/M2/M3 hierarchy shown side by side, more precision, more new components, a heavier engineering ask.
Variation 3, open to discussion. A hybrid, deliberately left unresolved internally rather than forced to a premature answer, combining pieces of both.
I carried Variation 1's minimal footprint forward as the base, with Variation 2's dollar impact treatment layered in, because the taxonomy's persona routing had already told me who this card's primary reader was: someone scanning, not studying.
Errors: a different uncertainty, a different question
The Errors surface needed its own process, not a copy of the Insights one, because the underlying uncertainty was different. A GMV anomaly can be measured precisely against an expected value. A JavaScript error's actual business severity often can't be, not every error type has a clean, known dollar cost the moment it's detected. The real choice on this surface wasn't between layout variations, it was between two honesty postures: wait until severity scoring was fully precise before showing any number at all, or ship an estimated severity score immediately with the uncertainty stated plainly rather than hidden. I chose the second.
The live product still carries that decision visibly, a Total Estimated Severity figure shown alongside a caveat, rather than a falsely confident number with no acknowledgment that the model was still maturing:
"We're working diligently to quantify severity."
Errors landed on a structurally parallel but visually distinct card, severity counts and a timeline chart instead of a summary sentence, arrived at through its own tradeoff, because a technical responder's first read, and the uncertainty they're reading through, is genuinely different from a business user's.
Same foundation, two independently reasoned outputs.
Design Decisions
Visual Hierarchy: What Earns Weight, and Why
On Insights: only numbers, metric names, and dimension names are bolded, pulled directly from the taxonomy's actual-versus-expected framing, because a growth lead scanning the feed needs to catch "how much" before anything else.
On Errors: severity tier and duration carry the equivalent weight instead, because a technical responder scanning for something urgent needs to catch "how bad, how long" first, the two facts that decide whether they drop what they're doing right now.
Different content earns the weight on each card, but the underlying rule holds across both: whatever answers "how much" or "how urgent" gets bolded, everything else recedes.
Progressive Disclosure: Summary State Versus Full Detail
On Insights: a teammate, JK, proposed a hover to expand interaction, compact summary on default, full breakdown on hover, so a business user isn't reading root cause and benchmarking language before they've decided the number is worth their time.
On Errors: the same two-state shape solved a different problem, a technical responder shouldn't have to scroll past raw status codes and payload strings just to confirm severity and duration. Collapsing that detail behind an expand action kept the fast-triage read clean on both cards.
Two separate reasons, not one pattern copied over, produced the same collapsed-by-default shape on both cards.
Componentization: What's Shared, What's Surface Specific
On Insights: the dollar impact chip and category tags were designed once and carry a consistent visual language, shape, colour logic, placement, that a returning user learns and doesn't have to relearn card to card.
On Errors: that same visual language carries over into the severity tier and duration display, same shape conventions, same colour logic, applied to different content, whether or not the underlying implementation was literally shared code.
The specific content inside each card, the sentence grammar on Insights, the error log table on Errors, stayed surface specific, because forcing those into one shared shape would have degraded both.
Content Structure as an Execution Layer, Tuned Per Persona
On Insights: this became a five part sentence grammar, forecasting, period over period, root cause, co-related events, industry benchmarking.
On Errors: the equivalent execution layer became structured fields rather than prose, what does this mean, what can you do, suggested action, pulled directly from the Impact Type table's predefined next-action logic.
Different surface, same underlying decision about what a reader needs first, second, third.
Severity and Category Tagging
On Insights: tags read Green Shoots, Performance Alert, Slow Moving, Outage, business language describing the shape of a trend.
On Errors: tags read High, Medium, Low, operational language describing urgency.
Different vocabularies, same job: letting a user triage a feed by category before investing any reading time in a single card.
Both reading paths, category tag then headline metric then dollar impact on Insights, severity tier then duration then chart on Errors, are different sequences of the same underlying question: what happened, how much does it matter, what do I do. That structural sameness, not visual sameness, is the actual deliverable.
Shipped, Impacted, and Beyond
The clearest proof this work landed isn't a metric I own, it's two screenshots from Out of the Blue's live product today, taken years after the design work was done.
The current Insights feed shows a funnel collapse described in bolded metric names and bolded percentage deltas, chained causally from one event to its downstream consequence, exactly the period-over-period-plus-root-cause sentence structure defined in the Building Blocks content layer.
The current Errors surface shows an actual value plotted against a dotted expected range line, and a severity count broken into High, Medium, Low tiers, drawn directly from the numeric representation framework and Impact Type classification defined in the taxonomy.
I want to stay precise about what this proves. I'm not asserting authorship of every pixel live today, and I'm not claiming Insights and Errors run on one shared technical backend. What the two screenshots together show is that the underlying design pattern, one taxonomy informing two structurally consistent but independently designed surfaces, is still shaping how the product is built, years later, on both product lines.
That pattern is what made the business outcomes possible. Out of the Blue's published case study on The Beard Struggle credits resolving a Klaviyo data issue with recovering over $19,000 a month in previously lost revenue, once the problem became clear enough to act on the same day it surfaced. A separate published case study on Abbott Lyon reports 400% more comprehensive monitoring coverage than the brand had relying on customers to self-report issues. Out of the Blue's own product messaging separately cites an expired ad pixel costing brands roughly $2 million per incident industry wide, and up to 30% of users never returning after a broken experience.
Source: outoftheblue.ai/case-study/case-study-tbs, Out of the Blue's published case study, accessed 2026.
These are company level, customer specific results, not a controlled measurement of this taxonomy in isolation.
But they describe exactly the outcome class this foundation was built to produce: the gap between an anomaly existing in the data and a human acting on it, closed faster, because both surfaces state what happened, how much it costs, and what to do, in that order, before a reader gets to a full sentence.
Tradeoffs
One reused design pattern over two custom-built ones.
A single Building Blocks taxonomy meant neither the Insights nor the Errors design had to invent its own information hierarchy from scratch, but it also meant every future addition to either surface has to fit inside a shared shape rather than being free to diverge. I accepted a real ceiling on per-surface flexibility in exchange for long-term consistency, and I'd make that trade again for these two products, since their audiences overlap heavily.
Minimal visual weight over maximum precision, on both cards.
Both final cards default to the lightest version of themselves, one headline number, one dollar figure, rather than the fuller numeric picture Variation 2 offered. That's the right call for a feed built for scanning, but it means a deeply engaged user has to actively hover or expand to get the complete picture rather than seeing it by default.
Shipping four sentence types over delaying for a fifth.
The industry benchmarking sentence, peer comparison against similar brands, was specified and ready in the taxonomy but never shipped during my time there. The cost of that cut is real: without it, a growth lead has no way to know whether their anomaly is unusual for their industry or entirely expected. I prioritised the four types answering urgent questions over holding the release for one that added context. That was the right call for the timeline.
Reflection
What I'd want to test next: whether the shared taxonomy still holds as more metric types and error categories get added, or whether it eventually needs to fork into surface specific variants. And separately: whether the persona routing logic, Biz Person View versus Tech Person View, correctly predicts what a real user wants to see first as the customer base diversifies beyond early adopters, or whether it needs more granularity than a binary split.
The principle underneath all of this: building the taxonomy before building either card was the actual design decision that mattered most. Both cards were built on top of it, independently, not dictated by it. The foundation is the work.
That work, the root cause methodology and the interaction design behind it, is covered in the next case study (link).