Experimentation Platform


Catapult, Etsy’s internal platform for running, measuring, and evaluating product experiments, was used constantly but had never had dedicated design investment. The insight driving our work was simple but consequential: hygiene, evaluation, and historical context are a loop: better documentation and metric selection feed better learnings, better learnings feed a trustworthy history, and richer history feeds smarter investment in the next generation of experiments.

For part of the year, the team lacked a reliable coordination point, and needed both a product vision and someone to hold the operating structure together.

 

Role

  • Staff Product Designer, functioning as squad Product Manager

  • Research lead across users, initiative leads, and executive leadership

  • Systems design — site map, service design blueprint

  • Team operations — sprint cadence, check-ins, work tracking

Impact

  • Scope of research: Research conducted across 3 organizational altitudes from daily users to executive leadership

  • Scope of delivery: 4 features shipped to the experiment details page in a single quarter, built directly from that research

  • Continuity: Sprint cadence and team structure adopted as standard practice beyond the project

 

Tooling

Utilized Claude to turn raw research notes into testable hypotheses tied to company outcomes, and to define measurable benchmarks like engagement.

Leveraged a Figma-based design system in combination with Cursor to rapid-prototype the platform vision and feature concepts.

 

Challenges & Solutions

Experiment Hygiene

Many experiments launched with poorly written hypotheses and a wide variety of metrics, leading to confusion about experiment intent.

  1. We used AI to grade hypotheses and conclusions, against parameters the product team defined, as a nudge for improvement.

  2. Introduced functionality to pre-set metrics based on surface, platform, audience, etc.

Evaluation

Key data points were buried in hover states, and the metrics table couldn't be filtered, sorted, or searched.

  1. I rebuilt the table in plain language, added filter/sort/search, and moved secondary detail into accordions.

  2. We reformatted the top of the page around what users actually needed first, and moved secondary detail into accordions.

 

Historical Context

Bad metadata meant users hunting for past learnings faced an arduous task — it took precious time and there was no certainty they'd found every relevant experiment.

  1. We linked related experiments into a parent-child chain, so a team could see the story that led to the current test.

  2. Search was rudimentary. Tagging and natural-language search are next. (On roadmap, not shipped in Q3)

leading through ambiguity

The team lacked a reliable coordination point, resulting in work going untracked and decisions getting re-litigated. Leadership decisions leaned on past assumptions and preconceived ideas of user needs, rather than current evidence.

  1. I introduced a more rigorous cadence, detailed check-ins with squad and leadership, and tracked work streams to give the team the structure it was missing.

  2. I ran research across three organizational altitudes and used it to build the Q3 roadmap the team could act on and to align stakeholders.

 

What shipped

Following the Q3 roadmap, the team shipped a redesigned experiment details page — filterable metrics, parent-child experiment linking, an AI-generated summary of product documents, an AI-generated score for the hypothesis and conclusion, and a reworked page hierarchy to surface what actually mattered most.