
Brainstorm Okrs
- 99 installs
- 451 repo stars
- Updated July 21, 2026
- borghei/claude-skills
brainstorm-okrs is a skill that generates and validates outcome-focused OKR sets using the Radical Focus framework with counter-metrics.
About
This skill generates and validates outcome-focused OKR sets using Christina Wodtke's Radical Focus methodology. It produces qualitative objectives with measurable key results and a counter-metric for each set, then scores quality with a validator script. Teams use it when setting quarterly OKRs or checking existing ones against quality criteria.
- Generates 3 distinct outcome-focused OKR sets using the Radical Focus framework
- Applies a counter-metric test so key results cannot be gamed
- Validates OKR sets with okr_validator.py, requiring a 70% score to commit
Brainstorm Okrs by the numbers
- 99 all-time installs (skills.sh)
- Ranked #1,367 of 3,282 Productivity & Planning skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
brainstorm-okrs capabilities & compatibility
- Capabilities
- okr generation · okr validation · goal setting
- Use cases
- planning · project management
- Pricing
- Free
What brainstorm-okrs says it does
The agent generates and validates outcome-focused OKR sets using Christina Wodtke's Radical Focus methodology.
Generate 3 Distinct OKR Sets
Any OKR set scoring below 70% must be revised before committing.
npx skills add https://github.com/borghei/claude-skills --skill brainstorm-okrsAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 99 |
|---|---|
| repo stars | ★ 451 |
| Last updated | July 21, 2026 |
| Repository | borghei/claude-skills ↗ |
What it does
Set or validate quarterly OKRs with outcome-focused key results and counter-metrics.
Who is it for?
Teams setting quarterly OKRs or auditing existing ones for outcome focus and gaming resistance.
Skip if: Ongoing KPI dashboards or project task tracking.
When should I use this skill?
You are setting quarterly OKRs or validating existing OKRs against quality criteria.
What you get
Three validated OKR sets, each with an inspirational objective, measurable key results, and a counter-metric.
- 3 OKR sets
- counter-metrics
- OKR validation report
By the numbers
- 3 distinct OKR sets generated
- 3 key results per objective
- 70% validator score required to commit
Files
OKR Brainstorming Expert
The agent generates and validates outcome-focused OKR sets using Christina Wodtke's Radical Focus methodology. It produces inspirational objectives with measurable key results, applies counter-metric tests, and scores quality against proven criteria.
Workflow
1. Identify the Theme
The agent asks: "What is the single most important thing this team needs to change this quarter?" The answer becomes the theme. Every OKR must connect back to this theme.
Validation checkpoint: If the user provides more than one theme, the agent pushes back. One theme per team per quarter. Multiple themes means no focus.
2. Generate 3 Distinct OKR Sets
For each set, the agent produces:
1. Objective -- One qualitative, inspirational statement (no numbers) 2. Key Result 1 -- Primary metric proving progress 3. Key Result 2 -- Secondary metric capturing a different dimension 4. Key Result 3 -- Counter-metric preventing gaming of KR1 and KR2 5. Rationale -- 2-3 sentences on why this set matters and how it connects to the theme
Objective quality criteria:
- Qualitative (numbers belong in key results)
- Inspirational (team would be excited to achieve it)
- Time-bound (achievable within one quarter)
- Actionable (team can directly influence the outcome)
Key result quality criteria:
- Measurable (has a metric with a number)
- Outcome-focused (measures results, not activities)
- Set at 60-70% confidence (not sandbagging, not demoralizing)
- Limited to 3 per objective
3. Apply the Counter-Metric Test
For every pair of key results, the agent asks: "Could we hit these numbers by doing something harmful?" If yes, it adds a counter-metric.
Example: If KR1 is "Increase sign-ups by 40%", a counter-metric is "Maintain activation rate above 60%." Without it, the team could game KR1 by lowering sign-up barriers so far that unqualified users flood in.
4. Validate with Tool
python scripts/okr_validator.py --input okrs.jsonThe validator scores each OKR set and surfaces quality issues: disguised tasks, missing metrics, output-framed key results, or missing counter-metrics.
Validation checkpoint: Any OKR set scoring below 70% must be revised before committing.
Example: Quarterly OKR Generation
Input: Theme is "retention" for a SaaS product team.
Output:
OKR Set 1:
Objective: "Become the product teams can't imagine leaving"
KR1: Reduce monthly churn from 4.2% to 2.5%
KR2: Increase 90-day retention cohort from 68% to 82%
KR3 (counter): Maintain NPS score above 45 (prevent forced lock-in tactics)
Rationale: Churn is the top revenue leak. Improving retention directly
increases LTV and reduces pressure on acquisition spend.
OKR Set 2:
Objective: "Make our onboarding so good that users hit value in their first session"
KR1: Increase Day-1 activation rate from 34% to 55%
KR2: Reduce time-to-first-value from 12 minutes to under 4 minutes
KR3 (counter): Maintain support ticket volume below 200/week (don't hide complexity)
Rationale: Users who activate on Day 1 retain at 3x the rate. Onboarding
is the highest-leverage retention lever.
OKR Set 3:
Objective: "Turn our power users into vocal advocates"
KR1: Increase referral-sourced signups from 8% to 20% of new users
KR2: Grow active community members from 500 to 2,000
KR3 (counter): Maintain power user retention above 95% (don't distract them)
Rationale: Advocacy compounds. Referred users have 37% higher retention
than paid-acquisition users.$ python scripts/okr_validator.py --input okrs.json
OKR Validation Results
======================
Set 1: 92/100 - PASS
Objective: Qualitative, inspirational, time-bound
KR1: Measurable, outcome-focused, stretch target
KR2: Measurable, different dimension from KR1
KR3: Valid counter-metric for churn reduction
Set 2: 88/100 - PASS
Objective: Qualitative, inspirational, time-bound
KR1: Measurable, outcome-focused
KR2: Measurable, tracks different dimension
KR3: Valid counter-metric
Note: "under 4 minutes" - verify baseline measurement exists
Set 3: 85/100 - PASS
Objective: Qualitative, inspirational
KR1: Measurable, outcome-focused
KR2: Measurable, but "active" needs precise definition
KR3: Valid counter-metricCommon OKR Mistakes
| Mistake | Example | Fix |
|---|---|---|
| Disguised task | "Launch the mobile app" | Ask "why?" -- measure the outcome the launch enables |
| Too many OKRs | 5 objectives per team | Pick 1, maybe 2. More means no focus |
| 100% confidence | Target you know you will hit | Stretch to 60-70% confidence |
| Activity metric | "Publish 12 blog posts" | Measure impact: "Increase organic traffic by 30%" |
| Set and forget | Review only at quarter end | Weekly check-ins with confidence scoring |
| Top-down only | All OKRs from leadership | Combine top-down direction with bottom-up team insight |
OKRs vs KPIs vs North Star Metric
| Concept | Purpose | Cadence | Example |
|---|---|---|---|
| North Star Metric | Single metric capturing core value delivery | Permanent | Weekly active users completing a workflow |
| KPIs | Health indicators across the business | Ongoing | Revenue, churn rate, response time |
| OKRs | Ambitious quarterly goals that move KPIs | Quarterly | "Become the fastest onboarding in our category" |
Relationship: OKRs are the lever pulled to move KPIs toward the North Star Metric. KPIs indicate business health. The NSM indicates core value delivery. OKRs define what changes this quarter.
Tools
| Tool | Purpose | Command |
|---|---|---|
okr_validator.py | Validate and score OKR sets | python scripts/okr_validator.py --input okrs.json |
okr_validator.py | Run demo validation | python scripts/okr_validator.py --demo |
Troubleshooting
| Symptom | Likely Cause | Resolution |
|---|---|---|
| OKR set scores below 70% consistently | Key results framed as tasks/outputs instead of outcomes, or objective contains numbers | Ask "So what?" for each KR until you reach a measurable outcome; remove numbers from objectives |
| Validator flags "output-oriented language" | KR description starts with verbs like "launch", "build", "implement", "ship" | Reframe: "Launch mobile app" becomes "Increase mobile-originated revenue from 0% to 15%" |
| Team sets 5+ objectives per quarter | Lack of strategic focus or inability to say no | Enforce 1 theme per team per quarter; use the Radical Focus constraint: one objective, maybe two |
| Key results hit 100% every quarter | Targets are sandbagged at 100% confidence | Stretch to 60-70% confidence; if you hit every KR, you are not being ambitious enough |
| Counter-metrics missing from OKR sets | Team did not apply the gaming test to KR pairs | For every pair of KRs, ask: "Could we hit these numbers by doing something harmful?" Add a counter-metric if yes |
| OKRs set and forgotten until quarter end | No weekly check-in rhythm established | Implement weekly confidence scoring (red/yellow/green) per KR; teams with weekly check-ins complete 43% more goals |
| Validator rejects input JSON | Schema mismatch: missing okr_sets key or key_results array per set | Ensure JSON has okr_sets array, each with objective string and key_results array containing description, metric, target_value, current_value |
Success Criteria
- Each OKR set scores above 80/100 on the validator before committing to the quarter
- Maximum 1-2 objectives per team per quarter (focus over breadth)
- Every objective is qualitative and inspirational (no numbers in the objective itself)
- Each objective has exactly 3 key results: primary metric, secondary dimension, and counter-metric
- Key results are set at 60-70% confidence (stretch, not sandbagged)
- Weekly confidence check-ins are conducted, not just end-of-quarter reviews
- OKR retrospectives run at quarter end with structured review of what was learned
Scope & Limitations
In Scope:
- OKR brainstorming using Christina Wodtke's Radical Focus methodology
- Generating 3 distinct OKR sets per theme with counter-metric testing
- Automated validation and scoring of OKR quality (output detection, metric presence, structural checks)
- Guidance on OKR vs. KPI vs. North Star Metric distinctions
- Common OKR mistake identification and remediation
Out of Scope:
- OKR tracking and progress monitoring over the quarter (use dedicated OKR platforms)
- Company-level OKR cascade and alignment across teams (see
senior-pm/for portfolio alignment) - Individual performance-linked OKRs (OKRs should be team goals, not performance reviews)
- Metric instrumentation or analytics setup for measuring key results
Important Caveats:
- OKRs work best when combined with weekly check-ins. Teams that review OKRs only at quarter end see 30-45% lower completion rates.
- The validator catches structural issues but cannot assess strategic quality. A perfectly scored OKR can still be the wrong goal.
- OKRs should be aligned top-down (strategic direction) and bottom-up (team insight). Pure top-down OKRs reduce team ownership.
Integration Points
| Integration | Direction | Description |
|---|---|---|
scrum-master/ | Receives from | Sprint velocity and capacity data inform realistic KR target-setting |
senior-pm/ | Receives from | Portfolio strategic priorities shape quarterly OKR themes |
execution/outcome-roadmap/ | Feeds into | OKR key results become success metrics for roadmap Now/Next items |
execution/prioritization-frameworks/ | Complements | Prioritized initiatives inform which OKR theme to focus on |
discovery/identify-assumptions/ | Receives from | Validated assumptions increase confidence in OKR target feasibility |
discovery/brainstorm-experiments/ | Feeds into | Experiment metrics may become OKR key results when validated |
Tool Reference
okr_validator.py
Validates and scores OKR sets against quality criteria. Checks objectives for qualitative/inspirational language, key results for measurable outcomes, and structural completeness.
| Flag | Type | Default | Description |
|---|---|---|---|
--input | string | (required, mutually exclusive with --demo) | Path to JSON file containing OKR sets |
--demo | flag | off | Run validation on built-in demo data (mix of good and bad OKRs) |
--format | choice | text | Output format: text or json |
Input JSON schema:
{
"okr_sets": [
{
"objective": "string (qualitative, no numbers)",
"key_results": [
{
"description": "string",
"metric": "string (unit of measurement)",
"target_value": "number",
"current_value": "number (baseline)"
}
]
}
]
}References
references/okr-best-practices.md-- Detailed OKR guide with examples and anti-patternsassets/okr_template.md-- OKR document template and quarterly review format
OKR Document: [Team/Company Name]
Quarter: Q[X] [Year] Owner: [Name] Status: Draft | Active | Scored Last Updated: [YYYY-MM-DD]
---
Company Context
Company Objective: [What strategic direction are we supporting?] North Star Metric: [The single metric that captures core value delivery]
---
OKR Set 1
Objective
[Qualitative, inspirational, time-bound statement -- no numbers]
Key Results
| # | Key Result | Metric | Current | Target | Confidence | Score |
|---|---|---|---|---|---|---|
| 1 | /10 | /1.0 | ||||
| 2 | /10 | /1.0 | ||||
| 3 | (counter-metric) | /10 | /1.0 |
Rationale
[2-3 sentences: Why this objective? How does it connect to the company objective?]
---
OKR Set 2 (Optional)
Objective
[Qualitative, inspirational, time-bound statement]
Key Results
| # | Key Result | Metric | Current | Target | Confidence | Score |
|---|---|---|---|---|---|---|
| 1 | /10 | /1.0 | ||||
| 2 | /10 | /1.0 | ||||
| 3 | (counter-metric) | /10 | /1.0 |
Rationale
[2-3 sentences]
---
Weekly Check-In Log
Week [N] - [Date]
| KR | Confidence | Status | This Week's Focus |
|---|---|---|---|
| KR1 | Green/Yellow/Red | [Brief update] | [What will move this forward] |
| KR2 | Green/Yellow/Red | [Brief update] | [What will move this forward] |
| KR3 | Green/Yellow/Red | [Brief update] | [What will move this forward] |
Blockers: [Any impediments]
---
Quarterly Review
Scores
| OKR Set | KR1 | KR2 | KR3 | Average |
|---|---|---|---|---|
| Set 1 | /1.0 | /1.0 | /1.0 | /1.0 |
| Set 2 | /1.0 | /1.0 | /1.0 | /1.0 |
Retrospective
What did we learn?
- [Learning 1]
- [Learning 2]
What would we do differently?
- [Change 1]
- [Change 2]
How should this inform next quarter?
- [Recommendation 1]
- [Recommendation 2]
Example: Q3 OKRs for Acme Analytics Shared Dashboards Launch
Real-world scenario showing how to apply the Radical Focus OKR workflow end-to-end.
Context
Acme Analytics is gearing up for the Q3 GA launch of "Shared Dashboards" -- a feature that lets customers share read-only dashboards with their own clients. The Customer Workflows product team (PM Devi, EM N. Okafor, 4 engineers, 1 designer) needs Q3 OKRs that reflect the launch impact, not the launch itself. The team has been burned before by "ship X" KRs that produced shipped-but-unused features.
The PM is running a Q3 OKR brainstorm on 2026-05-22 (start of quarter is 2026-07-01). She wants three candidate OKR sets, the team picks one, and the team's OKRs map cleanly into the Customer Workflows org OKR ("Make Acme stickier with the customer's own stakeholders").
Inputs
- Org-level OKR for Customer Workflows: "Make Acme stickier with the customer's customers" (single theme)
- Closed beta data: 75% activation, 56% outcome rate, NPS 38, 6 quotable testimonials
- Q3 capacity: ~80 engineering days
- The
okr_validator.pytool for SMART scoring
Applying the skill
1. One theme, locked: "Adoption of Shared Dashboards drives expansion." Rejected secondary themes (mobile parity, design system) -- they're sprints, not OKRs. 2. Generated 3 OKR sets -- adoption-led, expansion-led, retention-led -- each with one objective, three KRs (including a counter-metric), and rationale. 3. Applied the counter-metric test to each set. The adoption-led set was missing a counter-metric on quality; added "weekly NPS of share recipients >= 30". 4. Validated against SMART + Wodtke using the validator tool. The expansion-led set failed Wodtke's "set at 60-70% confidence" (KR1 was sandbagged at 95% confidence). Reset to 60% confidence. 5. Team picked the expansion-led set as primary, with adoption-led as a watch list (informally tracked). 6. Wrote a kill-criteria note: if KR1 misses 50% of target by mid-Q, declare a learning quarter and pivot.
Key decision quoted: "Sandbagged KRs are not safe -- they signal the team is not actually betting on the launch."
The artifact
````markdown
Q3 2026 OKRs -- Customer Workflows squad (Shared Dashboards)
Theme: Adoption of Shared Dashboards drives expansion. Window: 2026-07-01 to 2026-09-30 Confidence ratings: as of 2026-05-22 (planning); revisit weekly during Q3
The chosen OKR (Expansion-led)
Objective
"By end of Q3, Shared Dashboards is the reason mid-market customers expand into the Pro tier."
- Qualitative: yes
- Inspirational: yes (every CSM and AE has a story they will want to repeat)
- Time-bound: yes (Q3)
- Actionable: yes (PM + CSM + Sales own the levers)
Key Results
| # | KR | Target | Baseline (beta) | Confidence | Counter? |
|---|---|---|---|---|---|
| 1 | % of Pro-tier expansions in Q3 attributed to Shared Dashboards (CSM-tagged in HubSpot) | >= 35% | 0% (pre-launch) | 60% | -- |
| 2 | Weekly share-link generation rate, mid-market segment | >= 1.2 links / active workspace / week | 0.7 (beta) | 65% | -- |
| 3 | NPS of share-link RECIPIENTS (external stakeholders, surveyed via in-link survey) | >= 30 | n/a | 60% | yes (counter-metric for KR2) |
Rationale
The team's biggest risk on Shared Dashboards is shipping a feature that is loved by power users but ignored by the median Pro account. KR1 ties the feature to revenue impact (the company-level outcome). KR2 measures real usage (the leading indicator). KR3 is the counter-metric: if recipients hate the experience, KR2 can still climb (customers send links anyway) but the feature is poisoning trust. All three KRs together prevent the team from gaming any one.
Why these confidence levels
KR1 at 60%: this is the first quarter the CSM team is asked to tag the attribution; the data pipeline is unproven. Wodtke's bar. KR2 at 65%: usage in beta was 0.7; the Pro-tier segment has higher meeting cadence; 1.2 is a 70% lift -- ambitious but anchored. KR3 at 60%: we have never measured recipient NPS; the survey instrument is itself an experiment.
Kill criteria
- If KR1 < 17% by mid-Q (week 6), call a learning quarter: feature stays in market, GA already shipped, but Q4 OKRs reset around fixing the attribution gap.
- If KR3 < 15 in any weekly read, raise to a Sev2 design review immediately -- the feature is creating bad impressions externally, which damages the brand beyond Acme product surface.
Weekly health check
| Indicator | Source | Owner |
|---|---|---|
| Share-link generation rate | Amplitude | Devi |
| Pro-tier expansion attribution | HubSpot CSM workflow | Head of CS |
| Recipient NPS | In-link Pendo survey | Devi + Design |
Alternative OKR sets considered (not chosen)
Alternative 1 -- Adoption-led
Objective: "Every Pro-tier workspace experiences Shared Dashboards in Q3."
| # | KR | Target | Confidence |
|---|---|---|---|
| 1 | % of Pro-tier workspaces that have generated >= 1 share link | >= 60% | 70% |
| 2 | % of Pro-tier workspaces with >= 2 active share links at end of quarter | >= 30% | 65% |
| 3 | Counter-metric: support-ticket rate per active workspace (no regression) | <= 1.1x baseline | 75% |
Rejected because: measures adoption breadth, not depth or business impact. A workspace that creates one link and never uses it counts. The expansion-led set forces revenue accountability.
Alternative 2 -- Retention-led
Objective: "Shared Dashboards locks in mid-market customers ahead of renewal."
| # | KR | Target | Confidence |
|---|---|---|---|
| 1 | Net revenue retention, mid-market segment (vs prior quarter) | >= 112% | 60% |
| 2 | Gross churn, mid-market segment (vs prior quarter) | <= 1.4% / mo | 60% |
| 3 | Counter-metric: % of churned customers citing Acme product gaps in exit survey | <= prior quarter | 70% |
Rejected because: retention is influenced by too many factors outside the team's control; the squad cannot honestly say "we moved NRR by 4 points by shipping Shared Dashboards alone." Expansion-led keeps attribution clean.
SMART + Wodtke validation (okr_validator.py output)
$ python scripts/okr_validator.py --input okrs.json --framework wodtke
Set: Expansion-led (CHOSEN)
Objective:
[PASS] Qualitative
[PASS] Inspirational
[PASS] Time-bound (Q3)
[PASS] Actionable
KR1 (35% expansion attributed):
[PASS] Specific
[PASS] Measurable (% with source = HubSpot)
[PASS] Achievable (60% confidence)
[PASS] Relevant (theme)
[PASS] Time-bound (Q3)
KR2 (1.2 links/wk):
[PASS] All SMART
[PASS] 65% confidence
KR3 (NPS >= 30, recipient):
[PASS] All SMART
[PASS] 60% confidence
[PASS] Counter-metric for KR2
Set: Adoption-led
KR3 counter passes; KR1/KR2 high-confidence (>= 70%) -- borderline sandbagged
Set: Retention-led
KR1 confidence 60% but attribution to team's work is weak -- WARNConnection to the company OKR
| Level | Objective | Connection |
|---|---|---|
| Company | Grow ARR to $58M | -- |
| Org (Customer Workflows) | Make Acme stickier with the customer's customers | Theme |
| Squad (Shared Dashboards) | Shared Dashboards is the reason mid-market customers expand | Direct contributor (KR1 -> expansion ARR) |
Q3 weekly Monday cadence
- 15-min weekly OKR check-in at squad standup (Monday 10:00)
- Numbers refreshed Friday afternoon by Devi
- Mid-Q review with VP Product at week 6
- End-Q retro at week 13 using
sprint-retrospective/
Risks to the OKR
| Risk | Impact | Mitigation |
|---|---|---|
| HubSpot CSM tagging is inconsistent | KR1 unreadable | Devi runs a CSM training in week 1; sampling QA weekly |
| Recipient NPS survey has < 30% response rate | KR3 unreadable | Survey appears one time, on third dashboard view; A/B copy |
| Launch slips past 2026-07-14 | All KRs compressed | OKR window starts at GA + 7 days, not 2026-07-01 |
````
Why this works
- One theme, one quarter, one chosen OKR set -- the focus discipline Wodtke insists on.
- KR confidence is honest (60-65%), not sandbagged at 95%. Wodtke's bar.
- KR3 is a real counter-metric (recipient NPS) for KR1+KR2, not a vanity metric.
- The expansion-led objective ties team output to company revenue; the adoption-led alternative was rejected on attribution grounds, not on ambition.
- Kill criteria are written into the OKR document, so a mid-quarter "learning quarter" decision is not a surprise.
What's next
- Feed the OKR into ../outcome-roadmap/ to ensure Now/Next items map to KR levers.
- Use ../north-star-metric/ -- recipient NPS becomes an input metric for the org NSM.
- Pair with ../status-update-generator/ for the Monday cadence.
- Use ../launch-playbook/ to coordinate the GA event that the OKR window depends on.
- Revisit with ../../sprint-retrospective/ at end-of-Q for the OKR retro.
OKR Best Practices Guide
A comprehensive reference for writing, scoring, and managing OKRs using the Radical Focus methodology.
---
OKR Anatomy
Objective
A qualitative, inspirational statement of what you want to achieve this quarter.
Rules:
- No numbers (metrics belong in key results)
- Time-bound to one quarter
- The team can directly influence it
- It is memorable enough to repeat without looking it up
Formula: "[Action verb] [aspirational outcome] [for whom/in what context]"
Examples:
| Strong Objective | Why It Works |
|---|---|
| Become the fastest way for small teams to ship products | Aspirational, specific audience, clear direction |
| Make our customer support so good that users brag about it | Emotional, measurable through KRs, memorable |
| Establish ourselves as the go-to platform for data teams | Strategic, scoped to a segment, time-pressured |
| Weak Objective | Why It Fails |
|---|---|
| Increase revenue by 20% | That is a key result, not an objective |
| Improve things | Too vague to act on |
| Launch the mobile app | That is a project milestone, not an outcome |
| Be the best | Not specific enough to know when achieved |
---
Key Result Examples
Outcome-Focused (Good)
| Key Result | Metric | Why It Works |
|---|---|---|
| Reduce customer onboarding time from 3 days to 4 hours | Hours to first value | Measures user impact, not team activity |
| Increase weekly active users from 10K to 25K | WAU | Measures adoption, not feature count |
| Improve NPS from 32 to 55 | Net Promoter Score | Measures satisfaction, not effort |
| Reduce support ticket volume from 500/week to 200/week | Tickets per week | Measures product quality, not support team size |
Output-Focused (Bad)
| Key Result | Why It Fails | Better Version |
|---|---|---|
| Launch 5 new features | Measures activity, not impact | "Increase feature adoption rate from 15% to 40%" |
| Hire 3 engineers | Hiring is a means, not an end | "Reduce average PR review time from 48h to 12h" |
| Write 20 blog posts | Content quantity is not content impact | "Increase organic search traffic from 5K to 15K monthly visits" |
| Complete the migration | Completion is binary, not measurable progress | "Migrate 95% of users to the new platform with less than 2% support escalation rate" |
---
Common OKR Mistakes
1. Sandbagging
Problem: Setting targets you know you will hit (100% confidence). Why it matters: OKRs are meant to stretch. If you always hit 100%, your targets are too easy. Fix: Aim for 60-70% confidence. Hitting 70% of an ambitious OKR is better than hitting 100% of a safe one.
2. Too Many OKRs
Problem: Team has 5+ objectives with 15+ key results. Why it matters: OKRs create focus. Too many creates the illusion of focus while spreading effort thin. Fix: One objective is ideal. Two is acceptable. Three is the absolute maximum per team per quarter.
3. Confusing OKRs with KPIs
Problem: Using OKRs to track business-as-usual metrics. Why it matters: KPIs track health. OKRs drive change. "Maintain 99.9% uptime" is a KPI, not an OKR. Fix: Ask: "Is this something we want to change, or something we want to maintain?" Change = OKR. Maintain = KPI.
4. No Weekly Check-In
Problem: OKRs are set in January and reviewed in March. Why it matters: Without weekly check-ins, OKRs are forgotten goals. The cadence is the system. Fix: Every Monday, spend 15 minutes asking: "What confidence level are we at for each KR? What will we do this week to move it?"
5. Top-Down Only
Problem: Leadership dictates all OKRs with no team input. Why it matters: Teams closest to the work often have the best insight into what is achievable and impactful. Fix: Leadership sets company-level objectives. Teams propose their own KRs and supporting objectives.
6. Cascading Everything
Problem: Every company OKR must cascade to every team. Why it matters: Forced alignment creates artificial connections and busywork. Fix: Teams should align to 1-2 company objectives. Not every team needs to connect to every company goal.
---
OKR vs KPI vs North Star Metric
North Star Metric (NSM)
- The single metric that best captures the core value your product delivers to customers.
- It is permanent (changes rarely, if ever).
- Example: Airbnb = "Nights booked." Slack = "Messages sent in channels."
KPIs (Key Performance Indicators)
- Ongoing health metrics that you monitor continuously.
- They have targets but are not time-bound initiatives.
- Examples: Revenue, churn rate, page load time, customer satisfaction score.
OKRs (Objectives and Key Results)
- Quarterly ambitious goals designed to drive change.
- They expire at the end of the quarter -- you set new ones.
- They should move your KPIs and ultimately your NSM.
How They Connect
North Star Metric (permanent)
|
v
KPIs (ongoing health)
|
v
OKRs (quarterly change levers)
|
v
Projects / Initiatives (the work you do)Example flow:
- NSM: "Weekly users who complete a workflow"
- KPI: "Activation rate" (currently 45%, healthy range 50-70%)
- OKR: "Objective: Make first-time setup so intuitive that help docs become optional"
- KR1: Increase activation rate from 45% to 62%
- KR2: Reduce support tickets related to setup from 120/week to 30/week
- KR3: Achieve average time-to-first-workflow under 10 minutes (currently 35 min)
---
Scoring and Grading OKRs
At Quarter End
Score each key result on a 0.0 to 1.0 scale:
| Score | Meaning |
|---|---|
| 0.0-0.3 | Failed to make real progress |
| 0.4-0.6 | Made progress but fell short |
| 0.7-0.8 | Strong delivery (this is the sweet spot) |
| 0.9-1.0 | Hit or exceeded -- was the target ambitious enough? |
Interpreting Scores
- Consistently 0.9-1.0: Targets are too easy. Stretch more next quarter.
- Consistently 0.4-0.6: Either targets are too aggressive, or execution needs improvement. Investigate which.
- Consistently 0.0-0.3: Something is broken -- wrong priorities, wrong metrics, or wrong team focus.
The Retrospective Questions
After scoring, ask: 1. What did we learn about our customers/product/market? 2. What would we do differently if we could restart the quarter? 3. How should this inform next quarter's OKRs?
---
OKR Cadence (Quarterly Cycle)
Week -2 to -1: Drafting
- Review previous quarter's scores and learnings.
- Leadership drafts company-level objectives.
- Teams brainstorm their supporting OKRs.
Week 0: Alignment
- Teams present draft OKRs to leadership.
- Resolve conflicts and overlaps.
- Finalize and publish OKRs.
Weeks 1-12: Execution
- Monday check-in (15 min): Review confidence levels (green/yellow/red) for each KR. Identify blockers.
- Mid-quarter review (Week 6): Are we on track? Do any KRs need adjustment? (Adjusting targets is acceptable if you learned something fundamental.)
Week 13: Scoring and Retrospective
- Score all KRs (0.0-1.0).
- Run retrospective.
- Feed learnings into next quarter's drafting.
---
Company to Team Alignment
Alignment Model
Company objectives set the strategic direction. Team OKRs describe how each team contributes.
Company Objective: "Become the market leader in developer tooling"
|
+-- Engineering OKR: "Make our CLI the fastest in the category"
| KR1: Reduce average command execution time from 800ms to 200ms
| KR2: Achieve 95% positive sentiment in developer surveys
| KR3: Reach 10K daily active CLI users (currently 3K)
|
+-- Marketing OKR: "Build developer community that drives organic adoption"
| KR1: Grow community forum from 500 to 5,000 monthly active members
| KR2: Increase organic sign-ups from 20% to 45% of total
| KR3: Achieve 3 community-contributed plugins per month
|
+-- Sales OKR: "Establish enterprise adoption playbook"
KR1: Close 5 enterprise accounts (>$100K ARR each)
KR2: Reduce enterprise sales cycle from 90 to 45 days
KR3: Achieve 90% retention rate in enterprise tierAlignment Rules
1. Not every team must align to every company objective. 2. Teams should propose their own KRs -- leadership should not dictate them. 3. If two teams have conflicting KRs, resolve before the quarter starts. 4. Shared KRs between teams need a single owner (DRI).
Red Flags: Brainstorm OKRs
Common ways this skill's output goes wrong — concrete examples, why they're bad, and how to fix them.
How to use this document
Scan a draft OKR set before committing to the quarter. Each red flag shows the bad version next to the good version, anchored to Christina Wodtke's Radical Focus framework.
---
Red Flag 1: Output disguised as Key Result
Symptom. A KR reads "Launch the mobile app" or "Ship 12 blog posts".
Why it's bad. OKRs measure outcomes (what changed for customers / the business), not activities (what the team did). Output KRs let the team mark "done" while moving no business metric. The Objective gets credit; the company gets nothing.
Bad example:
"Objective: Become the fastest-growing product in our category.
KR1: Launch the new mobile app by August 31.
KR2: Ship 12 blog posts on growth themes.
KR3: Run 4 customer interviews per week."
Good example:
"Objective: Become the fastest-growing product in our category.
KR1: Increase weekly active accounts from 28k to 45k.
KR2: Lift mobile-originated signups from 12% to 30% of total.
KR3 (counter): Maintain D30 retention above 62% (don't trade quality for growth)."
How to catch it. For each KR, ask "could we hit this and have zero impact on a business metric?" If yes, it's output.
---
Red Flag 2: Sandbagging — 100% confidence
Symptom. Team commits to KR targets they are 95-100% certain to hit. Quarter ends; every KR is green.
Why it's bad. Wodtke's discipline targets 60-70% confidence on commitment. 100% confidence means the OKR captures work the team would do anyway — it does not focus or stretch. The team learns nothing about its capacity.
Bad example:
"Current MRR: $480k. KR1 target: $500k MRR by Q3 end (a 4% lift the team would hit on autopilot)."
Good example:
"Current MRR: $480k. KR1 target: $650k MRR by Q3 end (a 35% lift, requiring an activation experiment + a pricing change). Team confidence: 65%. If we hit this, we have learned to ship at this rate; if we miss at 0.6, we have learned where the constraint is."
How to catch it. Ask the team "what is your confidence percentage?" If above 80% on every KR, the bar is too low.
---
Red Flag 3: OKR drift mid-quarter
Symptom. It's week 8 of a 12-week quarter. KR1 looks behind. The team "updates" the KR target downward and calls it red-to-yellow.
Why it's bad. Mid-quarter target changes erase the learning signal. The reason confidence is 60-70% is precisely so half the time the team misses — and the miss teaches what the constraint is. Moving the goal post turns OKRs into theater.
Bad example:
"Week 8 check-in: KR1 was 'Increase MRR by 35%'. We're on track for 18%. Updating KR1 target to 20% to reflect current reality. Now green."
Good example:
"Week 8 check-in: KR1 was 'Increase MRR by 35%'. We're on track for 18%. Status: red. Discussion: the activation experiment underperformed; pricing change is held by legal. Decision: do not move the target. Run the diagnostic: what would unlock the remaining 17%? If nothing realistic, we miss the KR and document why in the retrospective. The miss is signal, not failure."
How to catch it. Compare week-1 KR targets to week-8 KR targets. Any downward shift is a red flag.
---
Red Flag 4: Five+ Objectives per team per quarter
Symptom. The team's OKR doc has 5 Objectives, each with 3 KRs. Total: 15 KRs.
Why it's bad. Radical Focus is in the name — Wodtke's framework demands 1 Objective per team per quarter (2 max). Five objectives equals no focus. The team will do scattered work and feel busy; nothing strategic moves.
Bad example:
"Q3 OKRs:
Obj 1: Improve onboarding.
Obj 2: Reduce churn.
Obj 3: Launch enterprise tier.
Obj 4: Improve performance.
Obj 5: Expand to APAC."
Good example:
"Q3 OKR (single): 'Make our enterprise pricing the obvious choice for mid-market buyers.' KR1: 8 mid-market net-new logos. KR2: enterprise tier ARR from $1.2M to $2.4M. KR3 (counter): SMB churn stays under 4%. The other initiatives (onboarding, performance, APAC) remain on the roadmap but are not Q3 OKRs."
How to catch it. Count Objectives. > 2 is a flag. The conversation is then "which Objective are we actually committing to?"
---
Red Flag 5: Missing counter-metric
Symptom. OKR set has KR1 and KR2 both pulling toward the same direction. No counter-metric.
Why it's bad. Without a counter, the team can game the KRs by doing something harmful. KR1 "increase signups" with no counter on activation means the team lowers the signup bar; numbers go up, business value goes down.
Bad example:
"Objective: Grow the funnel top.
KR1: Signups from 8k/mo to 16k/mo.
KR2: Trial starts from 5k/mo to 10k/mo."
Good example:
"Objective: Grow the funnel top with intent.
KR1: Signups from 8k to 16k/mo.
KR2: Trial starts from 5k to 10k/mo.
KR3 (counter): Day-7 activation rate stays above 38% (no growth at the cost of low-intent users)."
How to catch it. For each KR pair, ask: "Could we hit these by doing something stupid?" If yes, the counter is missing.
---
Red Flag 6: Numbers in the Objective
Symptom. Objective reads "Reach 50k weekly active users by Q3 end".
Why it's bad. Objectives are qualitative and inspirational. They describe the destination, not the speedometer. Numbers belong in Key Results. An Objective with numbers fuses the two and removes the room for KRs to prove progress against the Objective.
Bad example:
"Objective: Reach 50k weekly active users by Q3 end."
Good example:
"Objective: Become the product teams can't imagine leaving.
KR1: Weekly active users from 32k to 50k.
KR2: D30 retention from 58% to 70%.
KR3 (counter): NPS stays above 45."
How to catch it. Read the Objective. If it contains a number, a date, or a percentage, rewrite.
---
Red Flag 7: Activity metrics dressed as KRs
Symptom. KR reads "Publish 12 blog posts" or "Run 8 user interviews".
Why it's bad. Activity ≠ outcome. The team can do 12 posts that nobody reads. The KR validates effort but not impact. (Same root cause as Red Flag 1 but a more subtle phrasing — quantified activities feel measurable.)
Bad example:
"KR1: Publish 12 blog posts on growth themes.
KR2: Run 8 customer interviews."
Good example:
"KR1: Increase organic traffic to growth-themed pages from 4k/mo to 10k/mo (so the blog work is judged by outcome, not output).
KR2: Identify 3 prioritized opportunity areas from customer research (so the interview work is judged by what it produces, not by count of meetings)."
How to catch it. Read each KR. Strip the unit. If "12 posts" becomes "12 X" and the X is any activity unit, it's an activity KR.
---
Red Flag 8: Top-down only, no bottom-up
Symptom. Leadership sets OKRs and hands them to the team. The team had no input.
Why it's bad. OKRs work best when combined top-down (strategic direction) with bottom-up (team insight). Pure top-down OKRs reduce team ownership; the team becomes execution-only and silently disengages.
Bad example:
"VP of Product email: 'Here are the Q3 OKRs for your team. Please confirm by Friday.'"
Good example:
"Q3 OKR process: (1) Leadership shares strategic priorities and constraints (top-down). (2) Each team drafts 2-3 candidate OKR sets aligned to those priorities (bottom-up). (3) Joint review: team defends, leadership challenges, single OKR set committed. (4) Both sides sign the OKR doc."
How to catch it. Ask the team: "did you write this, or did leadership?" If pure top-down, ownership is weak.
---
Red Flag 9: KRs that all hit 100% — every quarter
Symptom. Team completes 100% of KRs three quarters in a row.
Why it's bad. Sustained 100% completion is sandbagging at the team level. Wodtke's frame: confidence at 60-70%, so on average teams hit roughly 70% of KRs. Three quarters at 100% means targets were locked at 100% confidence — the OKR is not stretching the team.
Bad example:
"Q1: 3/3 KRs hit. Q2: 3/3. Q3: 3/3. Team celebrated for consistent delivery."
Good example:
"Q1: 2/3 KRs hit (KR3 missed at 75%). Q2: 3/3 hit. Q3: 1/3 hit (an experiment failed and we learned why). Retrospective: hit-rate of 6/9 (67%) — close to Wodtke target of 60-70%. Team is genuinely stretched. The Q3 miss is the most-discussed item in the retro because it taught us the most."
How to catch it. Look at last 3 quarters of KR hit-rates. > 90% sustained is sandbagging; < 30% is demoralizing. 60-70% is healthy.
---
Red Flag 10: Set-and-forget
Symptom. OKRs are set at quarter start. The next time the team reviews them is at quarter end.
Why it's bad. OKRs without weekly check-ins drift. Teams forget the focus; work fills with non-OKR activities; the quarter ends in surprise. Wodtke and OKR practitioners observe a 30-45% completion-rate uplift from weekly confidence reviews.
Bad example:
"OKR doc created Apr 1. Next review: Jun 25. (Mid-quarter activity: team works on roadmap items, some aligned with OKR, some not.)"
Good example:
"Weekly check-in cadence (every Monday, 15 min): per KR, team rates confidence red/yellow/green and notes the leading indicator. Red KRs require a 1-week action to either unblock or accept the miss. Confidence trend reviewed monthly with the manager. Quarter end is calibration, not surprise."
How to catch it. Open the calendar. If there is no recurring OKR check-in, the OKRs will drift.
---
Red Flag 11: Validator passes but strategy is wrong
Symptom. okr_validator.py scores the OKR set 92/100. The OKR is structurally perfect. But the team is solving the wrong problem.
Why it's bad. The validator grades form (qualitative Objective, measurable KRs, counter-metric present). It cannot grade strategic substance. A 92/100 OKR can still be the wrong commitment for the company.
Bad example:
"OKR scored 92/100. Objective: 'Become the most loved customer-support tool.' KRs measurable, counter present. (Problem: the company strategy this quarter is enterprise expansion, not SMB support tooling. The OKR is aligned to the team's preference, not company strategy.)"
Good example:
"OKR scored 92/100 + alignment review: Q3 company priorities are enterprise expansion. This OKR is SMB-tooling-focused. Conflict. Discussion with VP Product: re-scope the Objective to 'Win the enterprise support buyer in Q3'. Validator re-runs; new score 88/100. Less polished structurally but strategically correct."
How to catch it. Ask: "How does this OKR align with company-level priorities for the quarter?" If the answer is hand-wavy, the OKR is locally optimized.
---
Red Flag 12: OKR = roadmap
Symptom. The Q3 OKR doc lists every roadmap item the team plans to ship.
Why it's bad. OKRs are not a roadmap. The roadmap is the menu; the OKR is the chosen dish. An OKR that lists every initiative is a backlog disguised as goal-setting. The team has not made a focus choice.
Bad example:
"Q3 OKR:
KR1: Ship the dashboard redesign.
KR2: Ship mobile push notifications.
KR3: Ship the Slack integration.
KR4: Ship enterprise SSO.
KR5: Ship the new pricing page."
Good example:
"Q3 OKR: 'Make mid-market the most natural buyer for our product.'
KR1: Mid-market signups (200-1000 employees) from 4% of total to 12%.
KR2: Mid-market trial-to-paid conversion from 18% to 30%.
KR3 (counter): SMB conversion does not drop below 22%.
Initiatives that support this OKR: pricing page redesign, mid-market case studies, sales enablement. (The dashboard redesign and Slack integration are on the roadmap but are not Q3 OKR-driving work.)"
How to catch it. Count KRs. > 4 per Objective is a flag — the team is listing roadmap items, not committing to outcomes.
---
Red Flag Quick Reference
| # | Anti-pattern | One-line check |
|---|---|---|
| 1 | Output disguised as KR | Could we hit this and have zero business impact? |
| 2 | Sandbagging at 100% confidence | What is the team's confidence percentage? |
| 3 | OKR drift mid-quarter | Did week-8 targets shift from week-1 targets? |
| 4 | 5+ Objectives per team | Count Objectives. |
| 5 | Missing counter-metric | Could the team hit these by doing something harmful? |
| 6 | Numbers in the Objective | Read the Objective. Does it contain a number? |
| 7 | Activity metrics dressed as KRs | Is the KR unit an activity count? |
| 8 | Top-down only | Did the team write the OKR, or just receive it? |
| 9 | 100% hit-rate sustained | 3-quarter hit-rate. > 90%? |
| 10 | Set and forget | Is there a recurring weekly check-in? |
| 11 | Validator passes but strategy wrong | How does this align with company priorities? |
| 12 | OKR = roadmap | > 4 KRs per Objective? |
Related Reading
- SKILL.md Troubleshooting
- references/okr-best-practices.md
north-star-metric/(KRs should move input metrics for the NSM)outcome-roadmap/(initiatives that support the OKR live here, not in the OKR)- Christina Wodtke, Radical Focus (2nd ed., 2021)
#!/usr/bin/env python3
"""OKR Validator - Validate and score OKR sets against quality criteria.
Checks that objectives are qualitative and inspirational, key results are
measurable outcomes (not outputs), and each set follows the Radical Focus
framework. Scores each OKR set 0-100 and provides improvement suggestions.
Usage:
python okr_validator.py --input okrs.json
python okr_validator.py --input okrs.json --format json
python okr_validator.py --demo
python okr_validator.py --demo --format json
Input JSON format:
{
"okr_sets": [
{
"objective": "Become the most trusted onboarding experience",
"key_results": [
{
"description": "Reduce time-to-first-value",
"metric": "minutes to first completed task",
"target_value": 5,
"current_value": 18
}
]
}
]
}
Standard library only. No external dependencies.
"""
import argparse
import json
import re
import sys
import textwrap
# Words that suggest an output (task/activity) rather than an outcome
OUTPUT_VERBS = [
"launch", "build", "implement", "create", "develop", "design",
"deploy", "ship", "release", "write", "publish", "deliver",
"migrate", "refactor", "integrate", "set up", "configure",
"hire", "onboard", "train", "document", "automate",
]
DEMO_DATA = {
"okr_sets": [
{
"objective": "Become the most trusted onboarding experience in our category",
"key_results": [
{
"description": "Reduce time-to-first-value for new users",
"metric": "minutes to first completed task",
"target_value": 5,
"current_value": 18,
},
{
"description": "Increase onboarding completion rate",
"metric": "percent of users completing onboarding",
"target_value": 85,
"current_value": 52,
},
{
"description": "Improve new user satisfaction score",
"metric": "CSAT score (1-5) at day 7",
"target_value": 4.5,
"current_value": 3.2,
},
],
},
{
"objective": "Increase revenue by 25%",
"key_results": [
{
"description": "Launch 5 new features",
"metric": "features launched",
"target_value": 5,
"current_value": 0,
},
{
"description": "Build a new billing system",
"metric": "system built",
"target_value": 1,
"current_value": 0,
},
],
},
{
"objective": "Make our platform the fastest way for teams to collaborate on documents",
"key_results": [
{
"description": "Reduce average document collaboration cycle time",
"metric": "hours from draft to final version",
"target_value": 4,
"current_value": 24,
},
{
"description": "Increase concurrent editing adoption",
"metric": "percent of documents with 2+ simultaneous editors",
"target_value": 40,
"current_value": 12,
},
{
"description": "Reduce context-switching during collaboration",
"metric": "average tool switches per collaboration session",
"target_value": 1,
"current_value": 4,
},
],
},
]
}
def validate_objective(objective: str) -> dict:
"""Validate an objective string. Returns issues and score."""
issues = []
score = 100
# Check if objective contains numbers (should be qualitative)
if re.search(r'\d+', objective):
issues.append({
"severity": "warning",
"message": "Objective contains numbers. Objectives should be qualitative and inspirational. Move metrics to key results.",
})
score -= 20
# Check length
if len(objective.split()) < 5:
issues.append({
"severity": "warning",
"message": "Objective is very short. Consider making it more specific and inspirational.",
})
score -= 10
if len(objective.split()) > 25:
issues.append({
"severity": "info",
"message": "Objective is quite long. Aim for a concise, memorable statement.",
})
score -= 5
# Check for vague language
vague_words = ["improve", "better", "enhance", "optimize", "good", "great", "nice"]
objective_lower = objective.lower()
found_vague = [w for w in vague_words if w in objective_lower.split()]
if found_vague and len(objective.split()) < 10:
issues.append({
"severity": "warning",
"message": f"Objective uses vague language ({', '.join(found_vague)}). Be more specific about the desired outcome.",
})
score -= 10
if not issues:
issues.append({
"severity": "pass",
"message": "Objective is qualitative and appears well-formed.",
})
return {"score": max(score, 0), "issues": issues}
def validate_key_result(kr: dict, index: int) -> dict:
"""Validate a single key result. Returns issues and score."""
issues = []
score = 100
description = kr.get("description", "")
metric = kr.get("metric", "")
target_value = kr.get("target_value")
current_value = kr.get("current_value")
# Check for target value
if target_value is None:
issues.append({
"severity": "error",
"message": f"KR{index}: Missing target_value. Key results must be measurable.",
})
score -= 30
# Check for current value
if current_value is None:
issues.append({
"severity": "warning",
"message": f"KR{index}: Missing current_value. Include a baseline to measure progress.",
})
score -= 10
# Check for metric description
if not metric:
issues.append({
"severity": "error",
"message": f"KR{index}: Missing metric. What unit of measurement is this?",
})
score -= 20
# Check for output verbs (suggests task, not outcome)
desc_lower = description.lower()
found_outputs = [v for v in OUTPUT_VERBS if desc_lower.startswith(v) or f" {v} " in f" {desc_lower} "]
if found_outputs:
issues.append({
"severity": "warning",
"message": f"KR{index}: Contains output-oriented language ({', '.join(found_outputs)}). "
f"Key results should measure outcomes, not activities. "
f"Ask: 'What result does this activity produce?'",
})
score -= 25
# Check if target equals current (no stretch)
if target_value is not None and current_value is not None:
if target_value == current_value:
issues.append({
"severity": "error",
"message": f"KR{index}: Target equals current value. There is no improvement to measure.",
})
score -= 30
if not issues:
issues.append({
"severity": "pass",
"message": f"KR{index}: Well-formed measurable key result.",
})
return {"score": max(score, 0), "issues": issues}
def validate_okr_set(okr_set: dict, set_index: int) -> dict:
"""Validate a complete OKR set (objective + key results)."""
objective = okr_set.get("objective", "")
key_results = okr_set.get("key_results", [])
result = {
"set_index": set_index,
"objective": objective,
"objective_validation": validate_objective(objective),
"key_result_validations": [],
"structural_issues": [],
"overall_score": 0,
}
# Validate KR count
kr_count = len(key_results)
if kr_count == 0:
result["structural_issues"].append({
"severity": "error",
"message": "No key results defined. Each objective needs exactly 3 key results.",
})
elif kr_count < 3:
result["structural_issues"].append({
"severity": "warning",
"message": f"Only {kr_count} key result(s). Best practice is exactly 3 per objective.",
})
elif kr_count > 3:
result["structural_issues"].append({
"severity": "warning",
"message": f"{kr_count} key results defined. More than 3 dilutes focus. Pick the 3 most important.",
})
# Validate each KR
kr_scores = []
for i, kr in enumerate(key_results, 1):
kr_validation = validate_key_result(kr, i)
result["key_result_validations"].append(kr_validation)
kr_scores.append(kr_validation["score"])
# Calculate overall score
obj_score = result["objective_validation"]["score"]
kr_avg = sum(kr_scores) / len(kr_scores) if kr_scores else 0
# Structural penalty
structural_penalty = 0
if kr_count == 0:
structural_penalty = 40
elif kr_count != 3:
structural_penalty = 10
result["overall_score"] = max(
int(obj_score * 0.3 + kr_avg * 0.6 + (100 - structural_penalty) * 0.1),
0,
)
# Generate suggestions
result["suggestions"] = _generate_suggestions(result)
return result
def _generate_suggestions(result: dict) -> list[str]:
"""Generate improvement suggestions based on validation results."""
suggestions = []
obj_score = result["objective_validation"]["score"]
if obj_score < 80:
suggestions.append(
"Rewrite the objective to be qualitative and inspirational. "
"Remove numbers and metrics -- those belong in key results."
)
kr_validations = result["key_result_validations"]
output_krs = []
for i, krv in enumerate(kr_validations, 1):
for issue in krv["issues"]:
if "output-oriented" in issue.get("message", ""):
output_krs.append(i)
if output_krs:
kr_list = ", ".join(f"KR{k}" for k in output_krs)
suggestions.append(
f"{kr_list} read like tasks, not outcomes. For each, ask 'So what? What result does "
f"completing this produce?' and use that result as the key result instead."
)
if len(kr_validations) != 3:
suggestions.append(
"Adjust to exactly 3 key results. Include one counter-metric to prevent gaming."
)
if result["overall_score"] >= 80:
suggestions.append("This OKR set is strong. Confirm 60-70% confidence in hitting targets.")
return suggestions
def format_text_report(results: list[dict]) -> str:
"""Format validation results as human-readable text."""
lines = []
lines.append("=" * 60)
lines.append("OKR VALIDATION REPORT")
lines.append("=" * 60)
lines.append("")
for r in results:
lines.append(f"--- OKR Set {r['set_index']} (Score: {r['overall_score']}/100) ---")
lines.append(f"Objective: \"{r['objective']}\"")
lines.append("")
# Objective issues
lines.append(" Objective Validation:")
for issue in r["objective_validation"]["issues"]:
icon = _severity_icon(issue["severity"])
lines.append(f" {icon} {issue['message']}")
lines.append("")
# Structural issues
if r["structural_issues"]:
lines.append(" Structure:")
for issue in r["structural_issues"]:
icon = _severity_icon(issue["severity"])
lines.append(f" {icon} {issue['message']}")
lines.append("")
# KR issues
if r["key_result_validations"]:
lines.append(" Key Results:")
for krv in r["key_result_validations"]:
for issue in krv["issues"]:
icon = _severity_icon(issue["severity"])
lines.append(f" {icon} {issue['message']}")
lines.append("")
# Suggestions
if r["suggestions"]:
lines.append(" Suggestions:")
for s in r["suggestions"]:
wrapped = textwrap.fill(s, width=70, initial_indent=" -> ", subsequent_indent=" ")
lines.append(wrapped)
lines.append("")
# Summary
lines.append("=" * 60)
scores = [r["overall_score"] for r in results]
avg_score = sum(scores) / len(scores) if scores else 0
lines.append(f"Average Score: {avg_score:.0f}/100")
lines.append(f"Sets Evaluated: {len(results)}")
strong = sum(1 for s in scores if s >= 80)
needs_work = sum(1 for s in scores if s < 80)
lines.append(f"Strong: {strong} | Needs Work: {needs_work}")
lines.append("=" * 60)
return "\n".join(lines)
def _severity_icon(severity: str) -> str:
"""Return a text indicator for severity level."""
return {
"error": "[ERROR]",
"warning": "[WARN] ",
"info": "[INFO] ",
"pass": "[OK] ",
}.get(severity, "[????] ")
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
"""Parse command-line arguments."""
parser = argparse.ArgumentParser(
description="Validate and score OKR sets against quality criteria.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=textwrap.dedent("""\
Examples:
python okr_validator.py --demo
python okr_validator.py --input okrs.json
python okr_validator.py --input okrs.json --format json
Input JSON format:
{
"okr_sets": [
{
"objective": "...",
"key_results": [
{"description": "...", "metric": "...", "target_value": N, "current_value": N}
]
}
]
}
"""),
)
group = parser.add_mutually_exclusive_group(required=True)
group.add_argument(
"--input",
help="Path to JSON file containing OKR sets to validate",
)
group.add_argument(
"--demo",
action="store_true",
help="Run validation on built-in demo data (mix of good and bad OKRs)",
)
parser.add_argument(
"--format",
choices=["text", "json"],
default="text",
help="Output format: text (default) or json",
)
return parser.parse_args(argv)
def main(argv: list[str] | None = None) -> None:
"""Main entry point."""
args = parse_args(argv)
if args.demo:
data = DEMO_DATA
else:
try:
with open(args.input, "r", encoding="utf-8") as f:
data = json.load(f)
except FileNotFoundError:
print(f"Error: File not found: {args.input}", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in {args.input}: {e}", file=sys.stderr)
sys.exit(1)
okr_sets = data.get("okr_sets", [])
if not okr_sets:
print("Error: No okr_sets found in input data.", file=sys.stderr)
sys.exit(1)
results = []
for i, okr_set in enumerate(okr_sets, 1):
results.append(validate_okr_set(okr_set, i))
if args.format == "json":
print(json.dumps(results, indent=2))
else:
print(format_text_report(results))
if __name__ == "__main__":
main()
Related skills
FAQ
How does it prevent gaming of OKRs?
For every pair of key results it applies a counter-metric test and adds a counter-metric if the numbers could be hit by doing something harmful.
How are OKR sets validated?
The okr_validator.py script scores each set and any set scoring below 70% must be revised before committing.