Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
jurgendn avatar

Experiment Design

  • 41 installs
  • 1 repo stars
  • Updated July 31, 2026
  • jurgendn/agent-skills

Helps with design & ui/ux tasks.

About

experiment-design is a Claude Code skill for design & ui/ux. It helps solo builders move faster with AI-assisted development.

  • experiment-design
  • Design & UI/UX
  • AI-coding skill

Experiment Design by the numbers

  • 41 all-time installs (skills.sh)
  • Ranked #1,275 of 1,880 Design & UI/UX skills by installs in the Skillselion catalog
  • Data as of Aug 2, 2026 (Skillselion catalog sync)
npx skills add https://github.com/jurgendn/agent-skills --skill experiment-design

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs41
repo stars1
Last updatedJuly 31, 2026
Repositoryjurgendn/agent-skills

What it does

Helps with design & ui/ux tasks.

Files

SKILL.mdMarkdownGitHub ↗

Experiment Design

Design the smallest experiment that can meaningfully reduce uncertainty.

The goal is not to maximize benchmark numbers. The goal is to determine whether the claimed mechanism actually works.

---

Procedure

1. State the Hypothesis

Write the claim in falsifiable form.

Bad:

Our method improves graph learning.

Better:

Adaptive random-walk length based on spectral gap improves ONMI
over fixed walk-length baselines in dynamic community detection.

The hypothesis should specify:

  • what changes;
  • compared against what;
  • under which conditions;
  • measured by which metric.

---

2. Define the Minimal Experiment

Specify the smallest setup capable of testing the hypothesis.

Dataset / Environment

State:

  • dataset name;
  • synthetic or real-world;
  • task type;
  • graph size or sample size;
  • relevant constraints.

Example:

Dataset:
Dynamic LFR benchmark with controlled overlap and fragmentation.

Split Protocol

Explicitly define:

  • train / validation / test split;
  • temporal split if applicable;
  • leakage prevention.

Example:

Train:
first 60% of snapshots

Validation:
next 20%

Test:
final 20%

No future edges visible during training.

Intervention

Describe exactly what changes.

Bad:

Use our method.

Better:

Replace fixed walk length t=5 with adaptive
t selected from estimated mixing time.

Compute Budget

Record:

  • GPU/CPU budget;
  • wall-clock limit;
  • number of seeds;
  • search budget.

Example:

3 random seeds
Single RTX 4090
Maximum training time: 6 hours

---

3. Define Fair Baselines

Baselines must be:

  • competitive;
  • properly tuned;
  • compute-matched where possible.

Weak baseline selection invalidates conclusions.

Example:

Baselines:
- Static Louvain
- DF-Louvain
- RW refinement with fixed t
- Spectral clustering

Also define:

  • which hyperparameters are tuned;
  • whether all methods receive equal search budget.

---

4. Define Metrics

Separate:

  • primary metric;
  • secondary diagnostics.

Example:

Primary:
ONMI

Secondary:
Modularity
Runtime
Community fragmentation rate

If metrics can disagree, state this explicitly.

---

5. Identify Confounders

List alternative explanations for the result.

Common confounders:

  • data leakage;
  • larger parameter count;
  • longer training;
  • larger receptive field;
  • preprocessing artifacts;
  • seed sensitivity;
  • prompt contamination;
  • hidden filtering effects.

Example:

Potential confounder:
Adaptive walk length may simply increase effective neighborhood size,
not improve mixing-time alignment.

---

6. Add Mandatory Sanity Checks

Before scaling, verify correctness.

Overfit-Small-Batch Test

The model should overfit a tiny subset.

Failure here usually indicates:

  • implementation bug;
  • metric issue;
  • optimization failure.

Metric Verification

Manually verify:

  • metric direction;
  • normalization;
  • averaging logic;
  • edge-case behavior.

Seed Stability

Run at least 3 seeds for unstable systems.

Large variance may invalidate small gains.

---

7. Define Decision Criteria

Specify what outcome changes the conclusion.

ResultInterpretation
Large, stable improvement across seedsSupports hypothesis
Small gain within varianceLikely null result
Gain disappears under fair tuningBaseline unfairness
Improvement only on one datasetWeak generalization
Higher accuracy but extreme runtimeTradeoff, not clear win
Regression on controlled benchmarksEvidence against hypothesis

Avoid post-hoc reinterpretation.

---

8. Recommend Run Order

Run cheapest experiments first.

Example:

1. Verify data pipeline
2. Overfit 1 batch
3. Run baseline sanity benchmark
4. Run toy synthetic dataset
5. Run ablation study
6. Run full benchmark
7. Run multi-seed evaluation

Do not scale before correctness is verified.

---

Rules

  • Prefer toy experiments before large-scale sweeps.
  • Every experiment should answer a decision-relevant question.
  • Baselines must receive fair tuning effort.
  • Record what is intentionally not tested.
  • Avoid benchmark inflation through hidden compute advantages.
  • Small reproducible experiments are more valuable than large noisy ones.

---

Output Format

Hypothesis

State the falsifiable claim.

Minimal Experiment

  • Dataset:
  • Split protocol:
  • Intervention:
  • Compute budget:

Baselines

  • Baseline 1:
  • Baseline 2:
  • Baseline 3:

Metrics

Primary

  • ...

Secondary

  • ...

Risks / Confounders

  • ...
  • ...
  • ...

Sanity Checks

  • ...
  • ...
  • ...

Run Order

1. ... 2. ... 3. ...

Decision Criteria

OutcomeInterpretation
......

Scope Limits

Explicitly state:

  • what is not being tested;
  • assumptions held fixed;
  • conclusions that cannot yet be made.

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.