
Soc Policy Analysis
- 35 installs
- 223 repo stars
- Updated June 6, 2026
- asgard-ai-platform/skills
soc-policy-analysis is a skill that conducts structured policy analysis from problem definition through evidence-based recommendation.
About
This skill conducts structured policy analysis including problem definition, alternative evaluation, and evidence-based recommendation. A team uses it to evaluate policy options, compare interventions, or make public-sector recommendations. It follows a six-step process with an evaluation matrix and an implementation plan, applicable to government, corporate, and organizational decisions.
- Runs structured policy analysis from problem definition to evidence-based recommendation
- Uses a six-step process with an effectiveness/efficiency/equity/feasibility evaluation matrix
- Includes a worked example on delivery-rider injury policy in Taipei
Soc Policy Analysis by the numbers
- 35 all-time installs (skills.sh)
- Ranked #1,773 of 3,282 Productivity & Planning skills by installs in the Skillselion catalog
- Data as of Aug 2, 2026 (Skillselion catalog sync)
soc-policy-analysis capabilities & compatibility
- Capabilities
- soc cognitive bias · ops okr planning
- Use cases
- research · planning
What soc-policy-analysis says it does
IRON LAW: Problem Definition Determines Everything
Always include status quo
Implementation kills good policy
npx skills add https://github.com/asgard-ai-platform/skills --skill soc-policy-analysisAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 35 |
|---|---|
| repo stars | ★ 223 |
| Last updated | June 6, 2026 |
| Repository | asgard-ai-platform/skills ↗ |
What it does
Evaluate policy alternatives against criteria and recommend an evidence-based option.
Who is it for?
Teams evaluating policy options or comparing interventions for a public or organizational problem.
Skip if: Analyses built on strawman alternatives, which the docs call intellectually dishonest.
When should I use this skill?
When evaluating policy options, comparing interventions, or making public-sector recommendations.
What you get
A policy analysis with problem framing, alternatives, an evaluation matrix, a recommendation, and an implementation plan.
- policy analysis
- evaluation matrix
- implementation plan
By the numbers
- 6-step analysis process
- 3-5 genuine alternatives per analysis
- 5 common evaluation criteria
Files
Policy Analysis
Overview
Policy analysis is a systematic method for evaluating alternative courses of action to address public problems. It follows a structured process: define the problem → identify alternatives → establish criteria → evaluate → recommend. It applies to government policy, corporate policy, and organizational decision-making.
Framework
IRON LAW: Problem Definition Determines Everything
How you define the problem determines which solutions are considered.
"Traffic congestion" suggests road-building. "Excessive car dependency"
suggests public transit. "Inefficient land use" suggests zoning reform.
The same observable situation can be framed as different problems,
leading to completely different policy responses. Make the framing
explicit and examine alternatives.The Six Steps
1. Define the Problem
- What is the problem? (observable evidence, not just symptoms)
- Who is affected? How severely?
- What causes it? (root cause, not proximate cause)
- How is the problem framed? Are there alternative framings?
2. Identify Alternatives
- Status quo (do nothing — always include as baseline)
- Incremental options (modify existing policy)
- Transformative options (fundamentally new approach)
- Aim for 3-5 genuine alternatives, not strawmen
3. Establish Evaluation Criteria Common criteria:
| Criterion | Question |
|---|---|
| Effectiveness | Does it solve the problem? |
| Efficiency | Benefit relative to cost? |
| Equity | Who bears the costs? Who gets the benefits? |
| Feasibility | Political, administrative, and technical viability? |
| Sustainability | Can it be maintained long-term? |
4. Evaluate Alternatives Score each alternative against each criterion. Use evidence (data, case studies, research) wherever possible. Acknowledge uncertainty.
5. Recommend Select the alternative with the best overall profile. Justify the trade-offs explicitly — no alternative will score highest on every criterion.
6. Implementation Considerations
- Political feasibility: Who needs to approve? Who might oppose?
- Administrative capacity: Can existing institutions implement this?
- Timeline and phasing
- Monitoring and evaluation plan
Output Format
# Policy Analysis: {Problem}
## Problem Definition
- Problem: {description}
- Affected population: {who}
- Root cause: {analysis}
- Alternative framings: {other ways to define this problem}
## Alternatives
1. Status quo
2. {Option A}
3. {Option B}
4. {Option C}
## Evaluation Matrix
| Criterion | Status Quo | Option A | Option B | Option C |
|-----------|-----------|----------|----------|----------|
| Effectiveness | L/M/H | L/M/H | L/M/H | L/M/H |
| Efficiency | L/M/H | ... | ... | ... |
| Equity | L/M/H | ... | ... | ... |
| Feasibility | L/M/H | ... | ... | ... |
## Recommendation
**{Selected option}** — {justification including trade-off acknowledgment}
## Implementation Plan
- Political path: {approval process}
- Timeline: {phases}
- Monitoring: {how to measure success}Examples
Correct Application
Scenario: Policy analysis for reducing food delivery rider injuries in Taipei
- Problem: 45% increase in delivery rider traffic injuries (2023-2024)
- Alternatives: (1) Status quo, (2) Mandatory insurance + training, (3) Speed limits on delivery apps during peak hours, (4) Platform liability for rider injuries
- Evaluation: Option 2 scores highest on feasibility (incremental) and effectiveness (directly addresses risk); Option 4 is most effective but politically difficult (platform lobbying)
- Recommendation: Option 2 as immediate action, with Option 4 as medium-term legislative goal ✓
Incorrect Application
- "The problem is delivery riders drive too fast" → Symptom, not root cause. WHY do they drive fast? Because platform algorithms reward speed, per-delivery pay incentivizes rushing, and there's no penalty for unsafe driving. Different root causes lead to different solutions. Violates Iron Law: problem definition determines everything.
Gotchas
- Always include status quo: "Do nothing" is a valid option and the baseline for comparison. Sometimes it's the best option if alternatives are worse.
- Avoid strawman alternatives: Including obviously bad options to make your preferred option look good is intellectually dishonest. All alternatives should be genuine.
- Equity is often traded for efficiency: Policies that are most efficient often have unequal distributional impacts. Make this trade-off explicit.
- Implementation kills good policy: A brilliant policy that can't be implemented is worthless. Feasibility is not a secondary criterion — it's a prerequisite.
- Evidence hierarchy: RCTs > quasi-experiments > case studies > expert opinion > anecdote. Use the strongest evidence available, and be explicit about evidence quality.
References
- For cost-benefit analysis methodology, see
references/cost-benefit.md - For evidence-based policy frameworks, see
references/evidence-hierarchy.md
Example: 台北市青年住宅可負擔性政策分析
Scenario
台北市政府住宅及都市發展局委員會,2025 年底面臨以下壓力:
- 台北市 25–39 歲青年的租金佔稅後所得比中位數已達 42%(2024 年內政部不動產資訊平台)
- 社會住宅存量約佔全市住宅 1.8%(目標為 5%,但進度落後)
- 2024 年「青年安心成家」購屋補貼方案申請人數是核准量的 6.3 倍,排隊等待者逾 8 萬戶
- 民調顯示 61% 台北市 30 歲以下居民考慮遷往他縣市或出走海外
局長需在下次市議會前提出具體政策建議,預算規模限制在每年 NTD 30 億以內。
---
Analysis
Step 1 — 定義問題
可觀察的現象:台北市青年租金負擔過重,居住不穩定性升高。
受影響對象:
- 主要:25–39 歲受薪青年,尤其月薪 NTD 40,000 以下者
- 次要:雇主(人才外流)、生育率(高租金壓縮成家誘因)
根本原因分析:
- 供給側:都更速度慢,新增住宅供給每年約 7,000 戶但需求缺口估計 15,000 戶
- 需求側:低利率時代形成的投資性持有(空屋率約 11%)
- 制度側:租賃市場資訊不透明,房東偏好短租,租客保障薄弱
替代框架(Iron Law:定義問題決定解法方向):
| 框架 | 政策導向 |
|---|---|
| 「供給不足」 | 加速興建社會住宅 |
| 「投機性囤房」 | 空屋稅 / 囤房稅二稅合一 |
| 「租客保護不足」 | 強化租賃法規,限制漲租幅度 |
| 「補貼需求」 | 租金補貼直接發給青年 |
本分析採供給 + 制度並重框架,因數據顯示供給缺口與租客保護缺失同時存在。
---
Step 2 — 識別替代方案
1. 現狀維持(基準線):維持現行社宅申請 + 青安補貼,不新增政策工具 2. Option A:加速社會住宅興建 增加市有土地 BOT 合作,目標 2028 年前新增 12,000 戶社宅 3. Option B:推行租金補貼擴大版 將現行「青年租金補貼」受惠戶從 3 萬戶擴增至 8 萬戶,每戶每月最高補貼 NTD 3,500 4. Option C:推動租賃市場管制 強制登記租賃契約、限制年漲租幅度不超過 CPI+1%,並建立公開租金指數
---
Step 3 — 評估標準
| 標準 | 問題 |
|---|---|
| 有效性 | 能否實質降低青年租金負擔率(目標:5 年內降至 35% 以下)? |
| 效率 | 每 NTD 1 億預算能改善多少戶家庭? |
| 公平性 | 誰受益?誰承擔成本? |
| 政治可行性 | 市議會通過難度?房東/建商利益是否受損? |
| 永續性 | 政策效果是否依賴持續補貼,或能創造長期結構改變? |
---
Step 4 — 替代方案評估
現狀維持
- 有效性:低。現行方案每年幫助約 3 萬戶,缺口 8 萬戶等待中,問題持續惡化。
- 效率:預算 NTD 10 億/年,但排隊效應表示邊際效益下降。
- 公平性:中。資源集中在通過審查的申請人,最弱勢者(無固定地址)反而排除在外。
- 政治可行性:高。無需新立法。
- 永續性:低。未改變供需結構。
Option A:加速社宅興建
- 有效性:高(長期)。12,000 戶直接增加可負擔供給,且社宅可作為市場錨定價格。
- 效率:中低。每戶興建成本約 NTD 200–250 萬,加計土地成本每年需 NTD 25–30 億,接近預算上限,且效益 3–5 年後才顯現。
- 公平性:高。直接服務低收入與中低收入青年,非持有資產者受益。
- 政治可行性:中。需市議會通過都市計畫變更,可能遭鄰避效應(NIMBY)抵制;建商利益中立至正向。
- 永續性:高。實體資產長期存在,不依賴年度預算。
Option B:租金補貼擴大
- 有效性:中(短期)。立即改善 8 萬戶現金流,但不增加供給;若房東因此調漲租金,補貼效益部分被抵銷(文獻顯示房東吸收率約 20–40%)。
- 效率:高(每戶年成本 NTD 42,000)。8 萬戶共需 NTD 33.6 億,略超預算;可縮至 7 萬戶控制在 NTD 29.4 億。
- 公平性:中。補貼直達個人,但僅限正式租約持有者,地下租賃市場(估計佔 30%)排除在外。
- 政治可行性:高。直接利益輸送給選民,議員普遍支持。
- 永續性:低。補貼停止即失效,且財政壓力隨通膨增加。
Option C:租賃市場管制
- 有效性:中(需長期觀察)。租金上漲管制可防止進一步惡化,但不解決現有高基期問題;租約登記可改善市場透明度,降低資訊不對稱。
- 效率:極高(成本極低)。主要為行政成本,每年 NTD 1–2 億可執行。
- 公平性:中高。現有租客受益,但新租客(若房東因管制縮減供給)可能更難找到房子(管制的標準副作用)。
- 政治可行性:低。房東族群(亦為重要選民)強烈反對;台灣現無城市層級漲租管制先例,需修法。
- 永續性:中。法規長期有效,但可能壓制新供給進入。
---
綜合評估矩陣
| 標準 | 現狀維持 | Option A 社宅 | Option B 補貼 | Option C 管制 |
|---|---|---|---|---|
| 有效性 | L | H(長期) | M | M |
| 效率 | M | L | H | H |
| 公平性 | M | H | M | M |
| 政治可行性 | H | M | H | L |
| 永續性 | L | H | L | M |
| 綜合 | 弱 | 強(遲延) | 中強(短期) | 弱(政治障礙) |
---
Step 5 — 建議
主方案:Option A + Option B 並行(雙軌策略)
- 立即(2026 Q1):啟動租金補貼擴大至 7 萬戶(預算 NTD 29.4 億/年),解決短期民生壓力
- 中期(2026–2028):以省下的行政預算推進社宅興建先導計畫(3 個場址,目標 3,000 戶),為後續擴大奠基
- Option C 做法降階為行政措施:推動租約強制登記(不含漲租管制),提升市場透明度,降低政治阻力
取捨說明:Option B 補貼效率高但不可永續;Option A 社宅永續但短期無感。並行策略以補貼爭取政治時間,以社宅建立長期結構。Option C 漲租管制暫緩,因政治可行性不足且可能抑制新供給,風險大於收益。
---
Result
# Policy Analysis: 台北市青年住宅可負擔性
## Problem Definition
- Problem: 台北市 25–39 歲青年租金負擔率中位數達 42%,遠超國際警戒線 30%
- Affected population: 估計 22 萬青年租屋戶;間接影響雇主人才留存與市整體生育率
- Root cause: 供給缺口(每年需求超出供給約 8,000 戶)+ 投資性持有造成 11% 空屋 + 租賃制度保護不足
- Alternative framings:
(1) 投機囤房問題 → 空屋稅(需中央修法,市府權限不足)
(2) 租客保護問題 → 租賃管制(短期政治阻力大)
(3) 補貼需求問題 → 租金直補(即時有效但不永續)
(4) 供給不足問題 → 社宅興建(永續但遲延)
## Alternatives
1. 現狀維持(基準線)
2. Option A:加速社會住宅興建(12,000 戶 / 3 年)
3. Option B:租金補貼擴大版(擴至 7 萬戶)
4. Option C:租賃市場管制(強制登記 + 漲租上限)
## Evaluation Matrix
| Criterion | 現狀維持 | Option A | Option B | Option C |
|--------------|---------|----------|----------|----------|
| Effectiveness | L | H | M | M |
| Efficiency | M | L | H | H |
| Equity | M | H | M | M |
| Feasibility | H | M | H | L |
| Sustainability | L | H | L | M |
## Recommendation
**Option A + Option B 雙軌並行** — 以補貼(B)換取短期政治穩定與即時民生改善,
以社宅(A)建立長期結構性解方。Option C 租賃管制降階為行政措施(強制租約登記),
迴避政治高風險的漲租上限條款。主要取捨:接受補貼方案財政不永續性,以 3 年內社宅
先導計畫的進展作為政策轉型的依據。
## Implementation Plan
- Political path: 補貼擴大案列入 2026 預算案(市議會支持度高);社宅先導計畫以既有
都市更新條例推動,選擇政治阻力較小的 3 個市有地場址,避免 NIMBY 集中爆發
- Timeline:
- 2026 Q1:補貼擴大方案上線,7 萬戶受惠
- 2026 Q2:社宅先導計畫 3 場址動工
- 2028 Q4:首批 3,000 戶社宅落成,評估是否擴大第二期
- Monitoring:
- 主指標:每半年追蹤青年租金負擔率(目標 2030 年降至 37% 以下)
- 次指標:社宅申請等待名單長度、補貼退出率(升遷離補貼者比例)
- 反指標:監測私人租賃市場供給量,確認管制未引發供給萎縮
將上述內容寫入檔案:
Cost-Benefit Analysis
Cost-benefit analysis (CBA) quantifies the net social value of a policy alternative by converting all expected effects into monetary terms. It answers one question: do total benefits exceed total costs? When used alongside the evaluation matrix in the parent skill, CBA provides the Efficiency criterion score with numerical backing.
---
Core Formula
NPV = Σ (Bₜ - Cₜ) / (1 + r)ᵗ for t = 0 to T| Symbol | Meaning |
|---|---|
NPV | Net Present Value — the headline result |
Bₜ | Total benefits in year t |
Cₜ | Total costs in year t |
r | Discount rate (social rate of time preference) |
T | Time horizon of the analysis |
t = 0 | Present year (no discounting) |
Benefit-Cost Ratio (BCR):
BCR = PV(Benefits) / PV(Costs)BCR > 1 means benefits outweigh costs. BCR and NPV can rank alternatives differently — use NPV to compare mutually exclusive options, BCR when budget is constrained.
---
Step-by-Step Procedure
Step 1: Define the Scope
Before any calculation, fix three boundaries:
- Perspective: Whose costs and benefits count? Government budget only? All affected parties in Taiwan? Global externalities?
- Time horizon: Long enough to capture major effects. A transit project might need 30 years; a training program might need 5-10.
- Counterfactual: Compare against the status quo alternative (do nothing), not against an ideal. This aligns with the parent skill's requirement to always include status quo as the baseline.
Step 2: Enumerate Effects
List every consequence of the policy, positive and negative:
For each alternative:
Benefits:
- Direct outputs (e.g., injuries prevented, time saved)
- Indirect/spillover effects (e.g., reduced healthcare burden)
- Option value (future flexibility preserved)
Costs:
- Implementation (capital, setup)
- Operating (recurring annual)
- Compliance burden on regulated parties
- Unintended negative effectsTransfer payments (taxes, subsidies, transfers between parties) are NOT social benefits or costs — they are distributional. A subsidy shifts money; it doesn't create value. Exclude from NPV, but note in the equity analysis.
Step 3: Monetize Effects
This is the hardest step. Use the following hierarchy:
| Evidence Type | Method | Example |
|---|---|---|
| Market price exists | Use market price directly | Cost of materials, labor wages |
| No market but close proxy | Revealed preference | Hedonic pricing for noise pollution (property value drop) |
| Willingness to pay elicited | Stated preference | Contingent valuation surveys |
| Physical unit with literature value | Benefit transfer | VSL (Value of Statistical Life) from existing studies |
| Completely uncertain | Sensitivity analysis range | Use lower/upper bounds |
Common monetization anchors used in Taiwan policy:
- Value of Statistical Life (VSL): Taiwan's official figure for road safety analyses is approximately NT$23–27 million per fatality (varies by year; always cite source and year).
- Value of Time (VOT): Ministry of Transportation uses approximately NT$200–350/hour depending on trip purpose.
- DALY (Disability-Adjusted Life Year): For health interventions, WHO-CHOICE threshold for Taiwan is roughly 1–3× GDP per capita per DALY averted ≈ NT$800,000–2,400,000.
Step 4: Discount Future Values
Choose the discount rate explicitly and defend it:
| Rate | Rationale | Use case |
|---|---|---|
| 3% | Social rate of time preference (long-run growth + pure time preference) | Infrastructure, environment, public health |
| 5% | Government borrowing cost proxy | Budget-constrained government projects |
| 8–10% | Opportunity cost of capital | Projects competing with private investment |
For climate and multi-generational impacts, declining discount rates (3% near-term, 1% long-term) are increasingly standard (Stern Review approach).
Present Value factor table (for quick estimates):
| Year | r=3% | r=5% | r=8% |
|---|---|---|---|
| 1 | 0.971 | 0.952 | 0.926 |
| 5 | 0.863 | 0.784 | 0.681 |
| 10 | 0.744 | 0.614 | 0.463 |
| 20 | 0.554 | 0.377 | 0.215 |
| 30 | 0.412 | 0.231 | 0.099 |
Step 5: Calculate NPV
Sum discounted net benefits across all years.
Step 6: Sensitivity Analysis
CBA results depend heavily on assumptions. Always test:
1. Discount rate: run at r−2%, baseline r, r+2% 2. Key monetization values: run at 70%, 100%, 130% of central estimate 3. Time horizon: shorter and longer 4. Participation/take-up rate: if policy depends on behavior change
Report results as a range, not a single number. A policy whose NPV is positive across all sensitivity runs is robust. A policy that flips from positive to negative under plausible assumptions should be flagged.
---
Worked Example: Delivery Rider Mandatory Insurance + Training (Option 2)
Using the scenario from the parent SKILL.md.
Scope: Government + riders + platforms, 5-year horizon, r = 3%
Problem baseline (Year 0):
- 8,000 injury incidents/year involving delivery riders in Taipei
- Average direct cost per incident: NT$180,000 (medical + lost earnings)
- Annual social cost of status quo: 8,000 × NT$180,000 = NT$1.44 billion/year
Policy: Mandatory insurance + 8-hour safety training
Costs (annual, NT$ million):
Year 0 (setup):
- Regulatory framework + administration: 50
- Platform compliance systems: 120
Total Year 0 costs: 170
Year 1-5 (recurring):
- Training delivery (~60,000 riders/yr): 90
- Insurance premium subsidy (partial): 200
- Enforcement: 30
Total annual operating cost: 320Benefits (annual, from Year 1, NT$ million):
Injury reduction estimate: 30% (conservative; comparable Singapore program: 35%)
Injuries avoided: 8,000 × 30% = 2,400/year
Monetized: 2,400 × NT$180,000 = 432
Severity reduction (surviving serious injuries become minor):
Additional benefit: ~80
Total annual benefit: 512NPV Calculation:
r = 0.03
T = 5
# Costs
costs = [170, 320, 320, 320, 320, 320] # Year 0 to Year 5
# Benefits (0 in Year 0, 512M from Year 1)
benefits = [0, 512, 512, 512, 512, 512]
pv_costs = sum(costs[t] / (1 + r)**t for t in range(T+1))
pv_benefits = sum(benefits[t] / (1 + r)**t for t in range(T+1))
npv = pv_benefits - pv_costs
bcr = pv_benefits / pv_costsResults:
PV(Costs) = NT$1,637M
PV(Benefits) = NT$2,343M
NPV = NT$706M ← positive: policy passes CBA
BCR = 1.43 ← NT$1.43 returned per NT$1 spentSensitivity check:
| Scenario | Injury reduction | NPV (NT$M) | BCR |
|---|---|---|---|
| Pessimistic | 15% | −NT$158M | 0.90 |
| Central | 30% | +NT$706M | 1.43 |
| Optimistic | 45% | +NT$1,570M | 1.96 |
Interpretation: Policy is NPV-positive under central assumptions but flips negative if the injury reduction effect is below ~22%. This is the key uncertainty to resolve with a pilot program before full rollout.
---
Common Mistakes
Double-counting benefits: If you include "reduced healthcare costs" AND "increased productivity from healthy workers" AND "VSL for injuries prevented" — some of these overlap. Map the causal chain; count each effect once.
Ignoring distributional effects: A policy with positive NPV can still be regressive. NPV is aggregate. If NT$700M net benefit flows entirely to platform shareholders while riders bear training costs, NPV is the wrong headline. Report distributional breakdown separately (see Equity criterion in parent skill).
Optimism bias in costs: Government project cost estimates are systematically low. Apply a reference class adjustment: infrastructure projects in Taiwan average 30-50% cost overruns. Consider adjusting upward or discounting cost estimates.
Wrong counterfactual: Comparing against an impossible ideal ("zero injuries") rather than the realistic status quo inflates apparent benefits. Always compare to what would actually happen without the policy.
Attributing all correlation to causality: If injuries decline by 30% after the policy, not all 30% may be caused by the policy (secular trends, other concurrent changes). Use quasi-experimental evidence where possible; be explicit when you can't.
---
When CBA Is Not Sufficient
CBA struggles with:
- Rights and dignity: monetizing harm to a person's autonomy is contested
- Irreversible harms: extinction, environmental tipping points — standard discounting understates these
- Distributional justice: aggregate NPV can mask who wins and loses
- Uncertainty about effects: if the mechanism is poorly understood, the numbers are fictional
In these cases, use CBA as one input alongside multi-criteria analysis (the evaluation matrix), not as the decision rule. A policy analyst who presents only NPV and ignores equity and rights has done incomplete work.
---
Quick Reference: CBA Checklist
□ Perspective defined (whose costs/benefits?)
□ Time horizon justified
□ Status quo is the counterfactual baseline
□ Transfer payments excluded from NPV
□ Monetization method stated for each major effect
□ VSL / VOT source and year cited
□ Discount rate chosen and justified
□ Sensitivity analysis run on top 3 uncertain parameters
□ Distributional impact noted separately
□ Optimism bias addressed (cost estimates)
□ Conclusion states NPV range, not a single numberEvidence Hierarchy for Policy Analysis
The Hierarchy
Evidence quality runs from highest to lowest internal validity — the degree to which you can attribute observed outcomes to the policy itself rather than confounding factors.
Level 1 Systematic review / meta-analysis of RCTs
Level 2 Randomized Controlled Trial (RCT)
Level 3 Quasi-experiment (DiD, RD, IV, ITS)
Level 4 Observational study with controls (regression, matching)
Level 5 Case study / comparative case study
Level 6 Expert opinion / Delphi panel
Level 7 Anecdote / stakeholder testimonyHigher levels answer "does this policy cause the outcome?". Lower levels answer "does this outcome correlate with the policy?" — useful for generating hypotheses, not confirming causality.
---
Level Definitions and Identifying Marks
Level 1 — Systematic Review / Meta-Analysis
A structured search of all studies meeting pre-specified criteria, combined to produce a pooled estimate.
Identifying marks:
- Protocol registered before search (PROSPERO, OSF)
- Explicit inclusion/exclusion criteria
- PRISMA flow diagram
- Heterogeneity statistics (I², τ²)
When to cite it: When you find one, prefer it over any individual study. Check the search cutoff date — a 2015 meta-analysis may miss the best recent trials.
Key journals: Campbell Collaboration (social policy), Cochrane (health), J3P (3ie development).
---
Level 2 — Randomized Controlled Trial (RCT)
Random assignment eliminates selection bias: treatment and control groups are identical in expectation on all observed and unobserved characteristics.
Core estimator:
ATE = E[Y | T=1] - E[Y | T=0]where T=1 is treated, T=0 is control, and Y is the outcome.
Threats to internal validity:
| Threat | Description | Detection |
|---|---|---|
| Attrition bias | Differential dropout | Compare dropout rates by arm |
| Contamination | Control group receives treatment | Check spillover; cluster RCT |
| Non-compliance | Not all treated units take up treatment | Report ITT and LATE (IV) |
| Hawthorne effect | Behavior changes from being observed | Blind if possible |
Threats to external validity (the harder problem for policy):
- Sample recruited from willing participants → results may not generalize
- Lab or pilot scale differs from full rollout (SUTVA violations at scale)
- Context specificity: an RCT in Kenya may not transfer to Taiwan
Practical note: RCTs in public policy are rare and expensive. When you find one, note: who funded it, what population, what scale, what time horizon.
---
Level 3 — Quasi-Experiments
When randomization is impossible or unethical, quasi-experiments exploit natural variation to approximate a counterfactual.
Difference-in-Differences (DiD)
Setup: Treatment group exposed to policy at time T; control group not exposed.
DiD = (Ȳ_treat,post - Ȳ_treat,pre) - (Ȳ_control,post - Ȳ_control,pre)Identifying assumption: Parallel trends — in the absence of the policy, treatment and control would have moved together. Test by plotting pre-treatment trends visually and with event-study coefficients.
Worked example: Taipei introduces mandatory helmet fines for scooters in Q1 2023. Control group: New Taipei (no change). Outcome: ER admissions for head injuries.
Taipei: pre=120/month → post=80/month (−40)
New Taipei: pre=110/month → post=105/month (−5)
DiD = (−40) − (−5) = −35 admissions/month attributable to policyPitfall: Parallel trends fails if something else changed in the treatment area simultaneously (e.g., a road safety campaign launched the same month).
---
Regression Discontinuity (RD)
Setup: Units just above and below a threshold are treated differently; near the threshold they are otherwise identical.
Estimator:
LATE = lim(x→c⁺) E[Y|X=x] - lim(x→c⁻) E[Y|X=x]where c is the cutoff.
Validity checks:
- No bunching at cutoff (McCrary density test)
- Covariates smooth through cutoff
- Bandwidth sensitivity — results should be stable across bandwidth choices
Policy example: Firms with ≥50 employees must provide parental leave (a threshold). Compare firms with 48–49 vs. 50–51 employees on female hiring rates.
---
Instrumental Variables (IV)
Setup: An instrument Z affects treatment T but affects outcome Y only through T.
LATE = Cov(Y, Z) / Cov(T, Z) [Wald estimator, binary Z]Validity conditions (must be argued, not tested empirically): 1. Relevance: Z predicts T (testable — F > 10 is the rule of thumb) 2. Exclusion restriction: Z affects Y only via T (untestable — requires theory) 3. Monotonicity: Z affects all compliers in the same direction
Classic policy instrument: Draft lottery number as IV for military service → effect on earnings (Angrist 1990).
Warning: IV estimates the Local Average Treatment Effect (LATE) — the effect for compliers only. This may not generalize to the whole population.
---
Interrupted Time Series (ITS)
Setup: Long time series of outcome, intervention at known point. Estimate level change and slope change.
Y_t = β₀ + β₁·t + β₂·D_t + β₃·(t - T₀)·D_t + ε_t
where D_t = 1 if t ≥ T₀ (post-intervention)β₂: immediate level change at interventionβ₃: change in slope after intervention
Minimum data requirement: ≥12 pre-intervention time points; more is better to model seasonality.
Pitfall: Any concurrent event (economic shock, media campaign) confounds the estimate. A comparison series (control region) strengthens the design.
---
Level 4 — Observational with Controls
Regression, propensity score matching, or synthetic control without a clean natural experiment.
Honest framing: These establish correlation and can control for observed confounders. They cannot rule out omitted variable bias from unobserved confounders. Report them as "associated with" not "caused by".
Sensitivity analysis to report:
- Coefficient stability when adding controls (Oster 2019: bound on bias from unobservables)
- Omitted variable bias calculation: how large would an unobserved confounder need to be to explain the result?
---
Level 5 — Case Studies
When useful:
- Mechanism tracing: how and why did the policy work?
- Plausibility probe before committing to a larger study
- Rare phenomena where N is inherently small
Most rigorous form: John Stuart Mill's methods — Method of Agreement (same outcome across different contexts with one common factor) and Method of Difference (different outcome across similar contexts, one difference).
Honest limitation: Cannot establish generalizability. One successful city ≠ national policy.
---
Levels 6–7 — Expert Opinion and Anecdote
Use for:
- Identifying what questions to ask
- Generating hypotheses
- Filling gaps where no formal evidence exists
- Political feasibility assessment (practitioners know what will work institutionally)
Do not use as primary evidence for causal claims. If a policy recommendation relies primarily on Level 6–7 evidence, say so explicitly.
---
Evidence Quality Scoring Table
When you evaluate evidence for a policy claim, complete this table:
| Dimension | Score 1 | Score 2 | Score 3 |
|---|---|---|---|
| Design | Anecdote/opinion | Observational | RCT/quasi-experiment |
| Sample | Convenience, small N | Representative, moderate N | Population-level, large N |
| Context match | Different country/sector | Similar context | Same context |
| Replication | Single study | 2–3 studies | Meta-analysis |
| Recency | >10 years ago | 5–10 years | <5 years |
Total score interpretation:
- 5–7: Low confidence — treat as hypothesis-generating
- 8–11: Moderate confidence — acknowledge uncertainty in recommendation
- 12–15: High confidence — appropriate for strong recommendation
---
Applying the Hierarchy in Policy Analysis
Step 1: State the causal claim you need evidence for
Be precise: "Policy X causes outcome Y in population Z." Vague claims invite vague evidence.
Step 2: Search for the strongest available evidence first
Search order: 1. Campbell / Cochrane systematic reviews 2. 3ie Development Evidence Portal (for development interventions) 3. What Works Clearinghouse (education) 4. Government evidence repositories (UK What Works Centres, US MDRC) 5. Google Scholar for peer-reviewed quasi-experiments 6. Grey literature (think tanks, government evaluations) — use with caution
Step 3: Assess internal validity of what you find
For each study, ask:
- What is the study design? (Level 1–7)
- What is the identifying assumption, and is it plausible?
- Were there any major threats to validity the authors acknowledge?
Step 4: Assess external validity (context transfer)
Even a perfect RCT in Oslo may not apply to Taipei. Check:
- Population: demographics, behavior, institutions similar?
- Implementation context: administrative capacity comparable?
- Scale: pilot effect vs. system-wide rollout?
Step 5: Document the evidence explicitly in your policy analysis
In the Evaluation Matrix, add an evidence quality notation:
| Criterion | Option A | Evidence |
|-----------|----------|----------|
| Effectiveness | H | DiD study, Taiwan context (Li et al. 2022) |
| Equity | M | Observational only; no distributional analysis |
| Feasibility | H | Expert consensus; 3 comparable cities implemented |---
Common Mistakes When Using Evidence
Mistake 1: Citing the study, not the estimate "Studies show this policy works" — which studies? In what context? What effect size?
Correct: "Li et al. (2022) find a 23% reduction (95% CI: 14–32%) in injuries using a DiD design on Taiwan municipal data."
Mistake 2: Ignoring null results and publication bias Published studies skew toward positive findings. A single positive RCT may be the one lucky trial among several unpublished nulls. Check registries (AEA RCT Registry, ClinicalTrials.gov) for registered-but-unpublished studies.
Mistake 3: Treating correlation as causation to fit a preferred policy If your preferred option only has Level 4 evidence but a rival option has Level 2 evidence, that comparison must appear in the policy analysis. Omitting it is intellectually dishonest.
Mistake 4: Context laundering Taking a study from a high-income, high-capacity context and applying it to a low-capacity context without noting the mismatch. Administrative feasibility often mediates whether evidence from elsewhere transfers.
Mistake 5: Ignoring effect size Statistical significance ≠ policy significance. A statistically significant 0.5% reduction in injuries may not justify the cost of implementation. Always report effect sizes and practical significance alongside p-values or confidence intervals.
---
Evidence Language Guide
Match your language to the evidence you have:
| Evidence level | Appropriate language |
|---|---|
| Level 1–2 | "evidence demonstrates", "has been shown to cause" |
| Level 3 | "quasi-experimental evidence suggests", "associated with a reduction of X% after controlling for..." |
| Level 4 | "correlates with", "is associated with" |
| Level 5 | "consistent with the hypothesis that", "in the case of [city], this approach was followed by..." |
| Level 6–7 | "practitioners report", "stakeholders indicate", "in the view of [expert]" |
Avoid: "evidence shows" for Level 4–7. Avoid: "proven" for anything below Level 1–2 with replication.
Related skills
FAQ
What are the six steps of policy analysis?
Define the problem, identify alternatives, establish criteria, evaluate alternatives, recommend, and plan implementation.
Why include a status-quo option?
The docs say 'do nothing' is always a valid baseline for comparison and is sometimes the best option if alternatives are worse.