← Back to projects
AUTOPSY BRO / PROTOTYPECASE 02 / 03

When a feature dies while the dashboard still looks healthy, what actually killed it?

Autopsy Bro scrutinises existing product evidence across metrics, segments, time, causality and trade-offs to reconstruct what actually happened.

Product AnalyticsRCAEvidence ScrutinyDecision SupportAI Product Concept
Product AnalyticsRoot Cause AnalysisEvidence EvaluationExperiment DesignMetric Integrity

Run the autopsy

WORKING PROTOTYPE · SYNTHETIC DEMO DATAOpen full screen ↗

Explore the actual prototype here. No live integrations; the demonstration uses synthetic evidence.

A SOURCE-TRACEABLE REVIEW
EvidenceMetricPopulationTimeCausalityTrade-offsFinal ruling

One claim, examined from seven connected perspectives.

1. How and where this started

While working on products, I noticed something odd: sometimes the metrics improve, but the product does not. A feature can look healthy on the dashboard while retention, support or a key segment quietly gets worse. If an unexpected death needs an autopsy, an unexpected product failure should too.

2. Root cause identification

The gap was between the evidence and the conclusion.

Teams already have dashboards, cohorts, experiments, support data and research, but those signals rarely explain themselves. A metric can improve for the wrong reason, an average can hide segment damage, and two simultaneous changes can make one look guilty.

What I see What I still need to know
Activation increased Did the metric still represent the same behaviour?
The average improved Which users actually produced the lift?
Retention fell after a release What else changed at the same time?
An experiment won Did the effect survive downstream?
Support increased Where was the increase concentrated?
Revenue kept growing Was product health weakening underneath?

The question became: does the story we are telling actually survive the evidence around it?

3. The solution

Autopsy Bro reconstructs the failure before the wrong fix makes it worse.

The user brings the problem and the evidence they already have, then follows one connected case through Evidence → Metric → Population → Time → Causality → Trade-offs → Final ruling. Every major finding keeps its source, population, period, definition and calculation attached, so the explanation can be checked before the team acts on it.

4. Work mode / Bro mode

Same case. Different vocabulary.

Work mode keeps the language professional and stakeholder-ready; Bro mode translates the same analysis into plain, direct language for early PMs and founders. The numbers, sources, calculations and ruling never change.

Work: “The aggregate uplift masks a negative outcome within Enterprise.”

Bro: “The average is hiding the problem. Overall looks better. Enterprise doesn’t.”

5. Outputs and why

Investigate first. Share once the case holds up.

The Interactive Autopsy lets the user follow the evidence, inspect contradictions and see how the explanation develops. The Full Report packages the final findings, unknowns, contributors, provenance and ruling for people who need the conclusion without rerunning the investigation.

Output Job
Interactive Autopsy Understand what happened and verify the evidence
Full Report Share what happened, why we believe it and what remains unresolved

6. Guardrails and evals

A confident wrong diagnosis can do more damage than the original problem.

Before shipping against real company data, I would test whether every material claim stays grounded, calculations reconcile, contradictions remain visible, weak evidence reduces confidence and correlation never quietly becomes causation. Microsoft has documented how valid metric movements can still lead teams to incorrect experiment conclusions, which makes interpretation quality a product requirement here, not polish. Microsoft Research: metric interpretation pitfalls.

Eval Pass condition
Source grounding Every major finding has inspectable evidence
Numerical consistency Numbers reconcile everywhere
Causality Claims never exceed the evidence
Contradiction handling Conflicting evidence changes the ruling
Unknown handling Missing evidence stays unresolved
Metric integrity Definition or instrumentation changes surface
Segment check Aggregate results are tested against key populations
Time order Claimed causes fit the chronology
Work/Bro parity Both modes reach the same conclusion

7. Proposed testing

The first metric I care about is diagnostic lift.

I would give PMs the raw case first, record what they think happened, then let them run the same case through Autopsy Bro. The prototype earns another round only if users catch something material they missed, drop an unsupported conclusion, or reach a better-supported explanation.

Test What I want to learn
Raw case vs Autopsy Did understanding materially improve?
Experienced PM benchmark Is this useful beyond basic PM hygiene?
Missing source Does confidence fall correctly?
Conflicting source Does the ruling change?
Work vs Bro Does simpler language improve comprehension without hurting trust?
Different case types Does the product generalise beyond one demo?

8. Trade-offs

The product only works if it stays disciplined about what it refuses to do.

Trade-off Decision
Existing evidence vs prediction Analyse what happened
Accuracy vs neatness Keep multiple contributors when needed
Traceability vs magic Let users inspect important claims
Unknown vs guess Leave gaps unresolved
Investigation vs instant advice Recommend after scrutiny
One case vs more dashboards Every view advances the same explanation
Embedded vs standalone Bring the product to existing data
Personality vs credibility Keep Bro useful, not gimmicky

Scrutiny takes longer than asking “why did retention fall?”, but that time only earns its place if it stops the team from fixing the wrong thing.

9. Differentiation

The useful part starts after the dashboard has already told you what moved.

Analytics platforms already handle events, funnels, cohorts, experiments and retention; Amplitude even exposes those objects to AI clients through MCP. Autopsy Bro takes the messy case that remains when those signals disagree and asks whether the explanation built from them actually holds together. Amplitude MCP documentation.

Its unit of work is a case. Its output is a source-backed explanation.

10. Proposed production requirements

If it works, I would ship it where the evidence already lives.

The intended direction is a plugin or MCP-enabled layer inside tools such as Amplitude, Claude, ChatGPT/Codex and relevant work systems, rather than another analytics destination. Amplitude already exposes analytics content through MCP, while plugins in ChatGPT and Codex can package reusable workflow logic with connected apps and their permissions. Amplitude MCP · OpenAI plugin documentation.

A production version would need authenticated source access, permission-aware retrieval, persistent case state, reproducible calculations, provenance, evidence-quality checks, privacy controls and reasoning evals.

11. Market risks and failure conditions

The product fails if it only tells a good PM what they could already see in five minutes.

Other risks follow from that: analytics platforms may absorb the behaviour, source setup may cost too much effort, weak instrumentation may limit the answer, and generic rulings would kill trust quickly. Bro mode can help comprehension, but if the tone becomes more memorable than the diagnosis, the core product has failed.

The strongest validation signal would be: “I had the data. I didn’t see that.”

12. Conclusion

An unexplained death gets an autopsy because the visible symptoms rarely tell the whole story. Product failures deserve the same discipline.

Autopsy Bro explores whether teams can trace what actually went wrong while there is still time to correct the real failure mode, recover customer trust and stop the damage from compounding into churn, revenue loss or lost market share.

Scrutinising the evidence before the wrong diagnosis becomes the next product decision.

NEXT PROJECT / PRODUCT TEARDOWN + STRATEGYFigma Dev Mode↗