
Statspai
- 2 installs
- 3.2k repo stars
- Updated August 4, 2026
- brycewang-stanford/awesome-agent-skills-for-empirical-research
statspai is a Claude skill for StatsPAI, an agent-native Python package with 390+ functions for causal inference and applied econometrics returning structured, exportable result objects.
About
A skill for StatsPAI, an agent-native Python package for causal inference and applied econometrics exposing 390+ functions behind a single import and a unified API. Developers use it to run OLS, IV, DID, staggered DID, RDD, propensity score matching, synthetic control, double machine learning, causal forest, and neural causal models. It returns self-describing result objects with publication-ready Word, Excel, and LaTeX export.
- Agent-native causal inference toolkit with 390+ functions behind one import
- Self-describing API (list_functions, describe_function, function_schema) for LLM workflows
- Every function returns a structured result object with LaTeX/Word/Excel export
Statspai by the numbers
- 2 all-time installs (skills.sh)
- Ranked #1,759 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
statspai capabilities & compatibility
- Capabilities
- causal inference · regression modeling · did analysis
- Use cases
- data analysis
What statspai says it does
Agent-native causal inference & econometrics toolkit for Python. 390+ functions, one import, unified API.
StatsPAI is the agent-native Python package for causal inference and applied econometrics.
Every function returns a `CausalResult` with `.summary()`, `.plot()`, `.to_latex()`, `.to_word()`, `.to_excel()`, `.cite()`
npx skills add https://github.com/brycewang-stanford/awesome-agent-skills-for-empirical-research --skill statspaiAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 2 |
|---|---|
| repo stars | ★ 3.2k |
| Last updated | August 4, 2026 |
| Repository | brycewang-stanford/awesome-agent-skills-for-empirical-research ↗ |
What it does
Run causal-inference and econometrics analyses in Python (OLS, IV, DID, RDD, PSM, SCM, DML, causal forest) with structured, exportable results.
Who is it for?
Running a full causal-inference workflow in Python with agent-discoverable functions and publication-ready output.
Skip if: Stata-based analysis or simple regressions where a lightweight library suffices.
When should I use this skill?
The user asks to implement causal inference, run DID/IV/RDD/PSM/SCM/DML, or do econometric analysis in Python.
What you get
Causal estimates from a unified API with structured result objects exportable to Word, Excel, and LaTeX.
- causal estimates
- LaTeX/Word/Excel regression tables
- structured result objects
By the numbers
- 390+ functions
- one import (import statspai as sp)
- 29 academic themes in interactive editor
Files
StatsPAI: Agent-Native Causal Inference & Econometrics
StatsPAI is the agent-native Python package for causal inference and applied econometrics. One import statspai as sp, 390+ functions, covering the complete empirical research workflow.
Source: https://github.com/brycewang-stanford/StatsPAI PyPI: pip install statspai Paper: Published in Journal of Open Source Software (JOSS)
Why StatsPAI for Agents?
StatsPAI is the first econometrics toolkit purpose-built for LLM-driven research workflows:
1. Self-describing API: sp.list_functions(), sp.describe_function("did"), sp.function_schema("rdrobust") — agents can discover and understand functions without documentation lookup 2. Unified result objects: Every function returns a CausalResult with .summary(), .plot(), .to_latex(), .to_word(), .to_excel(), .cite() 3. One import: No need to juggle 20+ packages — import statspai as sp covers everything 4. Publication-ready output: Word, Excel, LaTeX, HTML export in every function
Core Methods
Classical Econometrics
sp.regress(df, "y ~ x1 + x2", cluster="firm_id") # OLS
sp.ivreg(df, "y ~ x1 | z1 + z2", cluster="state") # IV/2SLS
sp.panel(df, "y ~ x1 + x2", entity="firm", time="year", model="fe") # Panel FE
sp.heckman(df, "y ~ x1", "select ~ z1 + z2") # Heckman selection
sp.qreg(df, "y ~ x1 + x2", quantile=0.5) # Quantile regressionDifference-in-Differences
sp.did(df, "y", "treated", "post") # Auto-dispatch (2x2 or staggered)
sp.callaway_santanna(df, "y", "group", "time") # Staggered DID (CS 2021)
sp.sun_abraham(df, "y", "cohort", "time") # Interaction-weighted event study
sp.bacon_decomposition(df, "y", "treated", "time") # TWFE diagnostic
sp.honest_did(result, method="smoothness") # Sensitivity to PT violations
sp.continuous_did(df, "y", "dose", "time") # Continuous treatmentRegression Discontinuity
sp.rdrobust(df, "y", "running_var", cutoff=0) # Sharp RD (CCT 2014)
sp.rdrobust(df, "y", "running_var", fuzzy="treatment") # Fuzzy RD
sp.rddensity(df, "running_var") # McCrary density test
sp.rdmc(df, "y", "running_var", cutoffs=[0, 5, 10]) # Multi-cutoff RD
sp.rkd(df, "y", "running_var", cutoff=0) # Regression kink designMatching & Reweighting
sp.match(df, "treatment", covariates, method="psm") # Propensity score matching
sp.match(df, "treatment", covariates, method="cem") # Coarsened exact matching
sp.ebalance(df, "treatment", covariates) # Entropy balancingSynthetic Control
sp.synth(df, "y", "unit", "time", treated_unit=1, treated_period=2000) # ADH SCM
sp.sdid(df, "y", "unit", "time", treated_units, treated_periods) # Synthetic DIDMachine Learning Causal Inference
sp.dml(df, "y", "treatment", controls, model="PLR") # Double/Debiased ML
sp.causal_forest(df, "y", "treatment", controls) # Causal Forest (GRF)
sp.metalearner(df, "y", "treatment", controls, learner="dr") # DR-Learner
sp.tmle(df, "y", "treatment", controls) # Targeted MLE
sp.aipw(df, "y", "treatment", controls) # Augmented IPWNeural Causal Models
sp.tarnet(df, "y", "treatment", controls) # TARNet
sp.cfrnet(df, "y", "treatment", controls) # CFRNet
sp.dragonnet(df, "y", "treatment", controls) # DragonNetRobustness & Workflow
sp.spec_curve(df, "y", "treatment", controls, specs) # Specification curve
sp.robustness_report(result) # Automated robustness report
sp.subgroup_analysis(df, "y", "treatment", subgroups) # Heterogeneity with Wald test
result.to_latex() # Export to LaTeX
result.to_word("output.docx") # Export to Word
result.cite() # Auto-generate citationInteractive Visualization (v0.6+)
fig = result.plot()
sp.interactive(fig) # Stata Graph Editor-style WYSIWYG editing, 29 academic themesAgent Integration Pattern
import statspai as sp
# Step 1: Discover available functions
functions = sp.list_functions()
# Step 2: Understand a specific function
info = sp.describe_function("callaway_santanna")
# Step 3: Get JSON schema for structured calls
schema = sp.function_schema("callaway_santanna")
# Step 4: Execute and get structured results
result = sp.callaway_santanna(df, "y", "group", "time")
print(result.summary())
result.to_latex("tables/did_results.tex")When to Use StatsPAI vs Other Packages
| Scenario | Use StatsPAI | Alternative |
|---|---|---|
| Agent-driven analysis pipeline | ✅ Best choice — self-describing API | pyfixest (no agent API) |
| Full causal inference workflow | ✅ 390+ functions, one import | Assemble 10+ R/Python packages |
| Publication-ready output needed | ✅ Word/Excel/LaTeX/HTML built-in | statsmodels (no export) |
| Staggered DID with diagnostics | ✅ CS + SA + Bacon + HonestDID | differences (partial) |
| Neural causal models | ✅ TARNet/CFRNet/DragonNet | econml (partial) |
| Stata users migrating to Python | ✅ Stata-equivalent function names | linearmodels (limited) |
StatsPAI - Agent-Native Causal Inference & Econometrics Toolkit
Source: https://github.com/brycewang-stanford/StatsPAI
The agent-native Python package for causal inference and applied econometrics. 390+ functions, one import, designed for both AI agents and human researchers.
Published in Journal of Open Source Software (JOSS). Built by the CoPaper.AI team at Stanford REAP.
Key Features
- One import, unified API:
import statspai as sp - Agent-native:
list_functions(),describe_function(),function_schema()for LLM integration - 390+ functions: OLS, IV, DID (staggered), RDD, PSM, SCM, DML, Causal Forest, Meta-Learners, TMLE, neural causal models
- Publication-ready: Word, Excel, LaTeX, HTML export in every function
- Interactive visualization: Stata Graph Editor-style WYSIWYG plot editing (v0.6+)
Related skills
FAQ
How many functions does StatsPAI expose?
390+ functions behind a single import, covering the complete empirical research workflow.
How do agents discover functions?
Via a self-describing API using sp.list_functions(), sp.describe_function(), and sp.function_schema().