
Ccf Experiment Designer
- 27 installs
- 1.5k repo stars
- Updated July 8, 2026
- mikubaka88/ccfa-skills
Design CCF paper evidence packages (datasets, baselines, metrics, ablations, robustness tests) and build result tables and figures only from supplied real values.
About
This skill designs experiments that test a research paper's central claims, producing datasets, baselines, metrics, ablations, robustness tests, and publication-ready result tables and figures from supplied values. A developer uses it for benchmark planning and result presentation, and it never fabricates numbers or outcomes.
- Covers datasets, baselines, metrics, ablations, and robustness tests
- Builds LaTeX tables and figures only from supplied real values or explicit placeholders
Ccf Experiment Designer by the numbers
- 27 all-time installs (skills.sh)
- Ranked #1,135 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mikubaka88/ccfa-skills --skill ccf-experiment-designerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 27 |
|---|---|
| repo stars | ★ 1.5k |
| Last updated | July 8, 2026 |
| Repository | mikubaka88/ccfa-skills ↗ |
What it does
Design CCF paper evidence packages (datasets, baselines, metrics, ablations, robustness tests) and build result tables and figures only from supplied real values.
Files
CCF Experiment Designer
Core Rule
Design experiments that test the paper's central claims. Build result tables and publication figures only from supplied real values or explicit placeholders. Never fabricate numbers, improvements, significance, benchmark ranks, or user-study outcomes. Follow the user's requested output shape: experiment plan, table, LaTeX table, figure spec, ablation list, or execution queue.
Modes
design: datasets, baselines, metrics, ablations, robustness, efficiency, failure analysis, and execution priority.result-template: fill-in tables withTBDplaceholders.result-presentation: LaTeX tables, figure plans, SVG/PDF-ready chart specs, captions, and QA from supplied real results.
Workflow
1. Identify target venue, paper type, central claims, available results, and whether the task is planning or presenting results. 2. Extract the storyline from the idea or draft. Use ../ccf-paper-writer/references/storyline-blueprint.md only as a schema, not as a writing handoff. 3. Map every major claim to required evidence, reviewer question, dataset/workload, baseline, metric, ablation, and robustness/failure test. 4. If datasets or baselines are unknown, use public-safe search or hand off to ccf-literature-searcher; mark uncertainty instead of guessing. 5. Load references/evidence-design.md for venue-family expectations and references/result-templates.md for result tables. 6. For result presentation, preserve units, seeds, confidence intervals, dataset names, and metric direction. Mark missing values explicitly. 7. Hand off to ccf-paper-writer for manuscript prose, ccf-integrity-auditor for number/claim consistency, and ccf-submission-checker for package or artifact readiness.
Adaptive Output Contract
Return the requested artifact first. For a result table request, output the table. For a figure request, output the figure spec/caption/QA notes. For a full experiment-design request, use this default structure:
Mode:
Venue and assumptions:
Claim-evidence matrix:
Dataset / benchmark needs:
Baseline matrix:
Main experiments:
Ablations:
Robustness / failure / efficiency:
Result tables or figure specs:
Missing values:
Execution priority:
No-fabrication status:
Next CCFA owner:References
references/evidence-design.md: experiment and benchmark design.references/result-templates.md: fill-in result tables and presentation scaffolds.
interface:
display_name: "CCF Experiment Designer"
short_description: "Design experiments and build result tables/figures from supplied real values; never invent results."
default_prompt: "Use $ccf-experiment-designer for experiment design, baselines, ablations, result templates, and real-result figure/table presentation."
Evidence Design
Use this file to design CCF-A evidence packages without fabricating results.
Evidence Principle
Start from the claim, not the table. Each experiment should answer one of these reviewer questions:
- Does the method solve the stated problem?
- Why does the mechanism work?
- Is the comparison fair and current?
- Does the result generalize across settings?
- What fails, when, and why?
- Can the work be reproduced or audited?
- Does the evidence match the venue's expectations?
Venue-Family Evidence
AI/ML:
- Strong close baselines, simple baselines, ablations, seeds/statistics, compute, hyperparameters, robustness, scaling, and diagnostic analysis.
CV/multimedia/graphics:
- Visual comparison, qualitative failure cases, per-category or hard-case analysis, fair image/video settings, user or perceptual studies when needed.
NLP:
- Data quality, annotation agreement, leakage checks, automatic and human evaluation validity, error taxonomy, ethics, and task assumptions.
DB/KDD/IR:
- Realistic workload, scale, latency, throughput, memory, ranking metrics, indexing or pipeline cost, ablations, and deployment constraints.
Systems/networks/architecture/storage:
- Real bottleneck, implementation detail, end-to-end results, microbenchmarks, sensitivity, overhead, workload variation, operational boundaries.
Security/crypto:
- Threat model, attacker capabilities, adaptive attacks, bypass tests, guarantees, false positives/negatives, disclosure, ethics, and assumptions.
HCI/CSCW/UbiComp:
- Research questions, participants, procedure, tasks, measures, statistics, qualitative coding, triangulation, ethics, and claim scope.
SE/PL/FM:
- Tool benchmarks, real programs, formal statements, proof sketches, soundness/completeness tradeoffs, threats to validity, usability where relevant.
Theory:
- Formal model, theorem statement, proof roadmap, examples, lower/upper-bound relationship, relation to barriers and open problems.
Baseline Matrix
Use this before proposing experiments:
Baseline:
Why included:
Source / citation needed:
Implementation source:
Fairness constraints:
Expected metric:
Can run? yes / no / unknown
If missing, reason:Baseline categories:
- Closest prior method.
- Current strong method.
- Simple sanity baseline.
- Ablated version of the user's method.
- Human, oracle, random, heuristic, or classical baseline when meaningful.
- Deployed or default system baseline for systems/HCI/security tasks.
Ablation Logic
Each ablation should test a mechanism:
Component or assumption:
Why it matters:
Replacement or removal:
Metric affected:
Expected interpretation if it fails:Useful ablations:
- Remove a module.
- Replace a specialized component with a generic alternative.
- Vary a key hyperparameter.
- Change data scale, workload, or domain.
- Test hard cases and failure cases.
- Analyze compute, memory, latency, or cost.
Benchmark Or Dataset Design
For a new benchmark:
Task definition:
Data source:
Annotation or generation process:
Splits:
Metrics:
Baseline suite:
Leakage checks:
Difficulty and diversity:
Human or expert validation:
License and ethics:
Maintenance plan:Benchmark papers should be evaluated as benchmarks, not only as methods. Do not score them as weak because they lack a method, and do not invent adoption claims.
Minimum Convincing Package
For most CCF-A submissions:
1. Main comparison against close and strong baselines. 2. Mechanism ablation or proof. 3. Robustness/generalization/stress test. 4. Failure analysis or limitation study. 5. Reproducibility details.
If this package is infeasible, narrow the central claim before adding weaker experiments.
Result Templates
Use these templates to produce fill-in result tables. Keep all unknown numbers blank, TBD, or bracketed placeholders.
Claim-Evidence Matrix
| Claim | Reviewer question | Evidence needed | Dataset/benchmark | Baselines | Metrics | Result placeholder | Status |
| --- | --- | --- | --- | --- | --- | --- | --- |
| | | | | | | TBD | planned / running / done |Main Comparison Table
| Method | Source | Setting | Metric 1 ↑/↓ | Metric 2 ↑/↓ | Metric 3 ↑/↓ | Notes |
| --- | --- | --- | --- | --- | --- | --- |
| User method | this paper | | TBD | TBD | TBD | |
| Baseline A | citation needed | | TBD | TBD | TBD | |
| Baseline B | citation needed | | TBD | TBD | TBD | |Ablation Table
| Variant | Component changed | Mechanism tested | Metric 1 ↑/↓ | Metric 2 ↑/↓ | Interpretation after user fills result |
| --- | --- | --- | --- | --- | --- |
| Full method | none | full mechanism | TBD | TBD | |
| w/o component A | remove A | necessity of A | TBD | TBD | |
| generic replacement | replace A | specificity of design | TBD | TBD | |Robustness / Stress Test Table
| Stress condition | Why it matters | Dataset/workload | Metric | User result | Failure threshold | Reviewer concern answered |
| --- | --- | --- | --- | --- | --- | --- |
| | | | | TBD | | |Qualitative / Case Study Table
| Case | Selection rule | Expected observation | User-provided result | What it demonstrates | Failure or limitation |
| --- | --- | --- | --- | --- | --- |
| | representative / hard / failure | | TBD | | |Execution Priority Table
| Priority | Experiment | Claim defended | Cost | Dependency | Appendix/main | Stop condition |
| --- | --- | --- | --- | --- | --- | --- |
| P0 | | | low/medium/high | | main | |No-Fabrication Reminder
Use this note when returning templates:
No experimental result has been generated here. All TBD cells must be filled from user-run experiments, paper-provided numbers, or verified public baseline reports with matching protocol.