Statistical Analysis Blueprint for Experimental Design

Planning experiments

Saved articles
Hides the site's navigation and the panels around the article; press Escape to leave.

Online at https://diogoribeiro7.github.io/analytics-blog-jekyll/statistics/2024/02/20/statistical-analysis-experimental-design/

Topics

Planning experiments

  • Define primary outcome metrics before collecting data.
  • Establish guardrail metrics for operational health.
  • Use stratified randomization when heterogeneous subgroups exist.

Power calculations

1
2
3
4
5
6
7
8
9
10
11
12
13
14
library(pwr)

detectable_effect <- 0.03
baseline_rate <- 0.18
power_target <- 0.8
alpha <- 0.05

pwr_result <- pwr.2p.test(
  h = ES.h(baseline_rate, baseline_rate + detectable_effect),
  power = power_target,
  sig.level = alpha
)

ceiling(pwr_result$n)

Analyzing outcomes

1
2
3
4
5
6
7
library(broom)
library(sandwich)
library(lmtest)

model <- glm(conversion ~ treatment + device + country, family = binomial(), data = experiment)
robust <- coeftest(model, vcov = sandwich)
tidy(robust)

Communicating uncertainty

Metric Estimate 95% CI Notes
Lift 2.9% [1.1%, 4.7%] Practical significance achieved
p-value 0.004 – Meets alpha threshold
Sample ratio mismatch 0.6% – Within tolerance

Recommendations

  1. Roll out treatment to 45% of traffic while monitoring device-specific effects.
  2. Launch follow-up experiment measuring lifetime value after 90 days.
  3. Share raw data and analysis scripts in the open science workspace.

Download the power analysis workbook above to adapt these calculations for your experimentation roadmap.

Your private highlights

Kept in this browser only and never sent anywhere. These are your notes, not comments. With text selected, Alt+Shift+H highlights it and Alt+Shift+N adds a note.

Select a passage of the article to highlight it.

    Your reading data

    Revision history

    • Correction Corrected the power calculation: pwr.2p.test returns the sample size per arm, which the text read as the total. Details
    • Editorial Reworded the planning checklist.
    • Published

    Embed interactive plots, widgets, and demos using <figure>, <iframe>, or <div class="interactive-embed"> containers. Ensure each embed includes descriptive captions for accessibility.

    © 2024 Diogo Ribeiro. Text and figures under CC BY 4.0. Code samples under MIT.

    How to cite

    Use the quick export buttons to save citations for reference managers or copy the formatted text directly.

    Diogo Ribeiro (2024). Statistical Analysis Blueprint for Experimental Design. DataLog | Data Science & Research Theme. https://diogoribeiro7.github.io/analytics-blog-jekyll/statistics/2024/02/20/statistical-analysis-experimental-design/.

    BibTeX

    RIS

    EndNote

    Open science & reproducibility badges

    These badges highlight the transparency practices applied to this work. Hover or focus on each badge to learn more about the criteria.

    • Open Data Dataset and code repository published with permissive license. Public repository, DOI issued, README with reproduction steps.
    • Reproducible Workflow Containerized environment and automated tests provided. Continuous integration pipeline with reproducibility checks.
    • Transparent Peer Review Peer review reports archived with DOI and linked to article. Open peer review statement and archived reports on Zenodo.

    Related posts

    Loading mathematical content