$npx -y skills add product-on-purpose/pm-skills --skill measure-experiment-resultsDocuments the results of a completed experiment or A/B test with statistical analysis, learnings, and recommendations. Use after experiments conclude to communicate findings, inform decisions, and build organizational knowledge.
| 1 | <!-- PM-Skills | https://github.com/product-on-purpose/pm-skills | Apache 2.0 --> |
| 2 | # Experiment Results |
| 3 | |
| 4 | An experiment results document captures what happened when you tested a hypothesis, including statistical outcomes, segment analysis, learnings, and clear recommendations. Good results documentation turns individual experiments into organizational knowledge that improves future decision-making. |
| 5 | |
| 6 | ## When to Use |
| 7 | |
| 8 | - After an A/B test or experiment reaches statistical significance |
| 9 | - When an experiment is ended early (for any reason) |
| 10 | - To communicate findings to stakeholders who weren't involved |
| 11 | - During decision-making about whether to ship, iterate, or kill a feature |
| 12 | - To build a repository of learnings that inform future experiments |
| 13 | |
| 14 | ## When NOT to Use |
| 15 | |
| 16 | - The experiment is not designed or run yet -> use `measure-experiment-design` |
| 17 | - The results demand a direction decision -> use `iterate-pivot-decision`; this skill reports the evidence, that one decides |
| 18 | - You want the transferable learning banked for the organization -> follow up with `iterate-lessons-log` |
| 19 | - Your data is survey responses, not a controlled experiment -> use `measure-survey-analysis` |
| 20 | |
| 21 | ## Instructions |
| 22 | |
| 23 | When asked to document experiment results, follow these steps: |
| 24 | |
| 25 | 1. **Summarize the Experiment** |
| 26 | Provide context: what was tested, when it ran, how much traffic it received. Link to the original experiment design document if one exists. |
| 27 | |
| 28 | 2. **Restate the Hypothesis** |
| 29 | Remind readers what you believed would happen and why. This frames the results interpretation. |
| 30 | |
| 31 | 3. **Present Primary Results** |
| 32 | Show the primary metric outcome clearly: what were the values for control and treatment? Include statistical significance (p-value), confidence intervals, and sample sizes. Be honest about whether results are conclusive. |
| 33 | |
| 34 | 4. **Analyze Secondary Metrics** |
| 35 | Present guardrail metrics that ensure you didn't cause unintended harm. Note any secondary metrics that moved unexpectedly.both positive and negative. |
| 36 | |
| 37 | 5. **Segment the Data** |
| 38 | Look for differential effects across user segments (platform, tenure, plan type, etc.). Sometimes overall results mask important segment-level insights. |
| 39 | |
| 40 | 6. **Extract Learnings** |
| 41 | What did you learn beyond the numbers? Include surprising findings, questions raised, and implications for the product hypothesis. Negative results are valuable learnings. |
| 42 | |
| 43 | 7. **Make a Recommendation** |
| 44 | Be clear: should we ship, iterate, or kill? Support the recommendation with the evidence. If the decision is nuanced, explain the trade-offs. |
| 45 | |
| 46 | 8. **Define Next Steps** |
| 47 | Specify what happens now.engineering work to ship, follow-up experiments, metrics to continue monitoring, or documentation to update. |
| 48 | |
| 49 | ## Output Format |
| 50 | |
| 51 | Use the template in `references/TEMPLATE.md` to structure the output. A complete readout fills every template section: Summary; Hypothesis Recap; Results; Segment Analysis; Visualization; Learnings; Recommendation; Next Steps; and Appendix. |
| 52 | |
| 53 | ## Quality Checklist |
| 54 | |
| 55 | Before finalizing, verify: |
| 56 | |
| 57 | - [ ] Statistical methods and significance are clearly stated |
| 58 | - [ ] Confidence intervals are included (not just p-values) |
| 59 | - [ ] Segment analysis checked for differential effects |
| 60 | - [ ] Secondary/guardrail metrics are reported |
| 61 | - [ ] Learnings go beyond just the numbers |
| 62 | - [ ] Recommendation is clear and actionable |
| 63 | - [ ] Negative or inconclusive results are reported honestly |
| 64 | |
| 65 | ## Examples |
| 66 | |
| 67 | See `references/EXAMPLE.md` for a completed example. |