
Quality Rubric
- 11 installs
- 30 repo stars
- Updated July 7, 2026
- jeffallan/writing-with-agents
Scores finished content across 10 dimensions on a 1-5 scale and routes any dimension below 4 back to the responsible workflow phase for rework.
About
Evaluates a piece holistically across ten quality dimensions with evidence-based scores and maps gaps back to the phase that should fix them. A writer uses it as a final gate to decide whether content is publishable or needs targeted rework.
- Ten-dimension 1-5 scorecard with a 4+ publishable threshold
- Routes low-scoring dimensions to the responsible phase rather than fixing itself
Quality Rubric by the numbers
- 11 all-time installs (skills.sh)
- +1 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #807 of 1,352 Code Review & Quality skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/jeffallan/writing-with-agents --skill quality-rubricAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 11 |
|---|---|
| repo stars | ★ 30 |
| Last updated | July 7, 2026 |
| Repository | jeffallan/writing-with-agents ↗ |
What it does
Scores finished content across 10 dimensions on a 1-5 scale and routes any dimension below 4 back to the responsible workflow phase for rework.
Files
Role Definition
The Quality Rubric is the evaluation specialist. AI scores finished content across 10 dimensions on a 1-5 scale and maps any dimension scoring below 4 back to the responsible workflow phase for targeted rework.
Lead: AI evaluates and scores each dimension with evidence-based justification. Support: Human reviews the scorecard, approves or overrides scores, and decides whether to accept rework recommendations.
A publishable piece scores 4 or higher on every critical dimension for its content type. The Quality Rubric does not perform rework itself. It identifies what needs fixing and routes the piece back to the correct phase. The responsible phase then executes the fix under its own workflow.
This skill sits downstream of the Judge phase. The Judge handles line-level detection and polish. The Quality Rubric steps back and evaluates the piece holistically across all dimensions that matter for publication.
When to Use This Skill
- After the Judge phase has completed all detection passes and the human has approved edits
- Before publishing, as a final gate to confirm the piece meets minimum standards
- When deciding whether a piece needs rework and which phase should handle it
- When comparing multiple drafts or revisions against a consistent scoring standard
- When a piece feels "off" but the specific weaknesses are hard to articulate
- When establishing quality baselines for a new content type or publication
- When a piece has gone through multiple rework cycles and you need objective evidence of improvement
- When onboarding a new publication standard and need to calibrate expectations with scored examples
- When prioritizing which of several finished drafts to publish first based on quality ranking
Core Workflow
1. Receive polished draft from Judge phase -- Read the full draft, the original Architect blueprint, and any Judge detection reports. Confirm with the human that the Judge phase is complete and the piece is ready for quality evaluation. Do not score a draft that has not been through the Judge.
2. Score each of 10 dimensions on 1-5 scale with justification -- Evaluate the piece against all 10 scoring dimensions. For each dimension, assign a score from 1 (Poor) to 5 (Excellent) and provide a specific justification citing evidence from the text. Do not assign scores without pointing to concrete passages, patterns, or metrics. See references/scoring-dimensions.md for the full rubric.
3. Check scores against minimum publishable thresholds -- Look up the content type in the minimum standards table. Verify that the average score meets the minimum and that all critical dimensions for this content type score 4 or higher. See references/minimum-standards.md for thresholds by content type.
4. For any dimension below 4, identify responsible phase and rework action -- For each dimension that falls below the publishable threshold, identify which workflow phase owns that dimension and specify the rework action. Be precise: "return to Architect to restructure Section 3 and Section 5 ordering" is useful; "needs structural work" is not.
5. Present quality scorecard to human -- Deliver the complete scorecard with pass/fail assessment and rework recommendations. The human decides whether to approve publication, request rework, or override any scores. Do not initiate rework without explicit human approval.
Reference Guide
| Topic | Reference | Load When |
|---|---|---|
| Full 10-dimension rubric with scoring levels and phase mapping | references/scoring-dimensions.md | Scoring any dimension, understanding phase responsibility |
| Minimum publishable thresholds by content type and rework routing | references/minimum-standards.md | Checking pass/fail, routing rework to phases |
Constraints
MUST DO:
- Score all 10 dimensions for every evaluation (mark SEO as N/A for non-SEO content).
- Justify each score with specific evidence from the text -- quote passages, cite section numbers, reference measurable patterns.
- Route rework to the correct phase with a specific action, not a vague directive.
- Present the full scorecard to the human before any rework begins.
- Re-score after rework to confirm the fix actually raised the dimension above threshold.
- Include the content type and its critical dimensions at the top of every scorecard.
- Flag any dimension where the score changed between evaluation rounds and explain why.
- Distinguish between critical and non-critical dimensions for the given content type when reporting pass/fail.
MUST NOT DO:
- Inflate scores to avoid sending a piece back for rework.
- Skip dimensions because "they seem fine" -- every dimension gets a score and justification.
- Attempt rework directly. The Quality Rubric evaluates; other phases execute fixes.
- Approve publication when any critical dimension for the content type scores below 4.
- Change scores after presenting them without explaining the reason for the change.
- Average across dimensions to hide a critical failure -- a single dimension below threshold blocks publication regardless of the average.
- Score based on effort or intent rather than the text as written -- the rubric evaluates the artifact, not the process.
- Combine multiple dimension failures into a single rework action -- each failing dimension gets its own targeted rework directive.
- Use the rubric to evaluate in-progress drafts -- the Quality Rubric applies only to drafts that have completed the Judge phase.
Output Templates
Quality Scorecard
# Quality Scorecard: [Article Title]
## Content Type: [type]
## Minimum Average Required: [n] | Actual Average: [n]
## Status: [PASS / FAIL]
## Dimension Scores
| # | Dimension | Score | Phase Responsible | Justification |
|---|-----------|-------|-------------------|---------------|
| 1 | Thesis and Argument | [1-5] | Architect | [evidence] |
| 2 | Evidence and Depth | [1-5] | Madman + Carpenter | [evidence] |
| 3 | Structure and Flow | [1-5] | Architect + Carpenter | [evidence] |
| 4 | Clarity and Readability | [1-5] | Carpenter + Judge | [evidence] |
| 5 | Voice and Authority | [1-5] | Madman + Carpenter | [evidence] |
| 6 | Opening and Hook | [1-5] | Madman + Carpenter | [evidence] |
| 7 | Conclusion and Takeaway | [1-5] | Architect + Carpenter | [evidence] |
| 8 | Technical Accuracy | [1-5] | Judge | [evidence] |
| 9 | SEO Optimization | [1-5 or N/A] | Architect + Carpenter + Judge | [evidence] |
| 10 | Originality and Value-Add | [1-5] | Madman + Pre-Research | [evidence] |
## Rework Recommendations (if FAIL)
| Dimension | Current Score | Target | Return to Phase | Specific Action |
|-----------|--------------|--------|-----------------|-----------------|
| [name] | [n] | 4 | [phase] | [what to do] |
## Summary
[1-2 sentence overall assessment and next step recommendation]Knowledge Reference
This skill extends the Madman, Architect, Carpenter, Judge framework (Betty S. Flowers, 1981) with a formal quality gate. The original framework implies evaluation at the Judge stage, but treats it as line-level editing rather than holistic scoring. The Quality Rubric adds a structured, dimension-based evaluation that maps failures back to their origin phase, preventing the common mistake of trying to fix structural problems with editorial polish.
The 10 scoring dimensions draw from established content quality frameworks including readability research, SEO best practices, and editorial standards for technical and long-form content. The phase-routing model ensures that rework happens at the right level of abstraction: architectural problems go back to the Architect, voice problems go back to the human and Carpenter, and accuracy problems stay with the Judge.
The distinction between the Judge and the Quality Rubric is scope. The Judge works at the line level: detecting weak verbs, unsupported claims, passive voice, and structural inconsistencies within paragraphs. The Quality Rubric works at the piece level: evaluating whether the overall thesis lands, whether the evidence portfolio is sufficient, and whether the piece delivers on its opening promise. A piece can pass the Judge with clean prose and still fail the Quality Rubric because the argument structure does not hold together.
The re-scoring requirement after rework exists to prevent the common failure mode where a fix in one dimension degrades another. Restructuring to improve flow (dimension 3) can weaken the opening hook (dimension 6) if sections are reordered. Re-scoring catches these regression effects before publication.
Content type determines which dimensions are critical. A technical tutorial requires high scores on Technical Accuracy and Structure but may tolerate a lower Voice score. A thought leadership piece requires high Originality and Voice but may tolerate lighter Technical Accuracy. The minimum standards reference file defines these critical-dimension profiles per content type. Applying the wrong profile leads to false passes or unnecessary rework.
Minimum Publishable Standards
This reference defines the minimum quality thresholds for different content types. A piece passes the quality gate only when it meets both the minimum average score and scores 4 or higher on all critical dimensions for its content type.
Publishable Thresholds by Content Type
| Content Type | Minimum Average | Critical Dimensions (must be 4+) |
|---|---|---|
| Technical documentation | 4.0 | Technical Accuracy (8), Clarity and Readability (4), Structure and Flow (3) |
| Blog post | 3.8 | Opening and Hook (6), Originality and Value-Add (10), Structure and Flow (3) |
| SEO long-form | 3.8 | SEO Optimization (9), Structure and Flow (3), Originality and Value-Add (10) |
| White paper | 4.2 | Evidence and Depth (2), Voice and Authority (5), Technical Accuracy (8) |
| Case study | 3.8 | Evidence and Depth (2), Opening and Hook (6), Conclusion and Takeaway (7) |
| Tutorial | 4.0 | Clarity and Readability (4), Structure and Flow (3), Technical Accuracy (8) |
Reading the table: The number in parentheses is the dimension number from scoring-dimensions.md. A piece of content must meet both conditions to pass: the average across all scored dimensions must meet or exceed the minimum average, AND every critical dimension for that content type must score 4 or higher.
SEO dimension handling: For content types without SEO objectives, dimension 9 (SEO Optimization) is marked N/A and excluded from the average calculation. For SEO long-form content, dimension 9 is critical and included in the average.
---
Pass/Fail Decision Logic
1. Calculate the average score across all scored dimensions (exclude any marked N/A). 2. Check whether the average meets or exceeds the minimum for the content type. 3. Check whether every critical dimension for the content type scores 4 or higher. 4. The piece passes ONLY if both conditions are met.
If the average is met but a critical dimension is below 4: FAIL. The critical dimension must be fixed regardless of how strong the other scores are.
If all critical dimensions are 4+ but the average is below minimum: FAIL. The piece has weak spots across non-critical dimensions that collectively drag it below standard.
If both conditions fail: FAIL. Address critical dimensions first, then re-evaluate the average.
---
Phase-Rework Routing Summary
When a dimension scores below 4, route the piece back to the responsible phase with a specific action. This table provides the default routing. The Quality Rubric should tailor the specific action to the particular weakness found.
| Dimension | Return to Phase | Default Rework Action |
|---|---|---|
| 1. Thesis and Argument | Architect | Re-run throughline identification. Distill single-sentence thesis. Verify every section connects. Cut sections that drift. |
| 2. Evidence and Depth | Madman, then Carpenter | Madman generates specific examples, data, and concrete scenarios for weak claims. Carpenter integrates new evidence into draft. |
| 3. Structure and Flow | Architect | Re-examine blueprint section order, heading hierarchy, and transition logic. Propose alternative structure for human approval. |
| 4. Clarity and Readability | Judge (or Carpenter for systemic issues) | Judge runs additional readability passes on flagged passages. If problems are systemic, Carpenter does a clarity-focused rewrite of affected sections. |
| 5. Voice and Authority | Human + Carpenter | Human rewrites key passages in their own voice. Carpenter preserves human voice while polishing. Cannot be fixed by AI alone. |
| 6. Opening and Hook | Madman, then Carpenter | Review Madman output for unused hook candidates -- anecdotes, data, provocative framings. Carpenter rebuilds opening around strongest candidate. |
| 7. Conclusion and Takeaway | Architect, then Carpenter | Architect defines what insight the conclusion must deliver. Carpenter rebuilds conclusion to synthesize argument into that insight. |
| 8. Technical Accuracy | Judge | Judge runs focused accuracy review, flagging and verifying every factual claim. If source material is flawed, return to research phase. |
| 9. SEO Optimization | seo-writer skill | Run SEO validation from seo-writer. Audit keyword placement, heading hierarchy, and meta description. Carpenter implements fixes. Judge verifies naturalness. |
| 10. Originality and Value-Add | Research + Madman | Research identifies gaps in existing content. Madman generates material filling those gaps using human experience. Requires human input. |
---
Re-Scoring After Rework
After rework is complete and the piece returns from the responsible phase:
1. Re-score only the dimensions that were flagged for rework. 2. Verify the rework did not degrade other dimensions. 3. Recalculate the average with the new scores. 4. Re-check pass/fail conditions.
If the piece still fails, identify whether the rework was insufficient (same phase, deeper fix) or misdirected (different phase needed). Do not loop more than twice on the same dimension without escalating to the human for a decision on whether to accept the current quality level or fundamentally rethink the piece.
Scoring Dimensions Reference
This reference defines the 10 quality dimensions used by the Quality Rubric skill. Each dimension includes its name, responsible phase(s), a 5-level scoring rubric, and rework routing instructions.
Scoring Scale
| Score | Label | Meaning |
|---|---|---|
| 5 | Excellent | Exceeds expectations. Professional publication quality. No improvements needed. |
| 4 | Good | Meets expectations. Ready to publish with minor tweaks at most. |
| 3 | Adequate | Acceptable but noticeable weaknesses. Could publish but quality would be questioned. |
| 2 | Weak | Significant issues that undermine the piece. Not publishable without rework. |
| 1 | Poor | Fundamental problems. Major rework required at the responsible phase. |
Any dimension scoring below 4 triggers a rework recommendation. The piece does not pass the quality gate until all critical dimensions (defined per content type) reach 4 or higher.
---
Dimension 1: Thesis and Argument
Phase Responsible: Architect
| Score | Description |
|---|---|
| 5 | Clear, compelling thesis stated early. Every section advances the argument. Reader can articulate the thesis after reading without referring back. No tangential sections. |
| 4 | Thesis is clear and stated within the introduction. Most sections connect to it directly. One or two minor digressions that do not derail the argument. |
| 3 | Thesis is present but buried or vague. Some sections feel disconnected from the main argument. Reader has to work to identify the central claim. |
| 2 | Thesis is unclear or contradicted by parts of the piece. Multiple sections drift from the main argument. Reader finishes unsure of the point. |
| 1 | No identifiable thesis. The piece reads as a collection of loosely related ideas without a unifying argument. |
If below 4: Return to Architect. Re-run the throughline identification step. The Architect must distill a single-sentence thesis and verify every section connects to it. Cut or restructure sections that drift.
---
Dimension 2: Evidence and Depth
Phases Responsible: Madman + Carpenter
| Score | Description |
|---|---|
| 5 | Every claim backed by specific evidence. Original examples drawn from real experience. Data cited where appropriate. Depth goes beyond surface-level treatment on every key point. |
| 4 | Most claims supported with evidence. Examples are specific rather than generic. One or two points could use deeper support but do not undermine credibility. |
| 3 | Mix of supported and unsupported claims. Some examples are generic or hypothetical when concrete evidence would strengthen the point. Depth is uneven across sections. |
| 2 | Multiple claims lack evidence. Examples feel invented or surface-level. Key arguments rely on assertion rather than demonstration. |
| 1 | Claims are almost entirely unsupported. The piece asserts without demonstrating. No specific evidence, data, or concrete examples. |
If below 4: Return to Madman for more material generation, then Carpenter to integrate new evidence. The Madman should generate specific examples, data points, and concrete scenarios for the weakest claims. The Carpenter then weaves this material into the existing draft.
---
Dimension 3: Structure and Flow
Phases Responsible: Architect + Carpenter
| Score | Description |
|---|---|
| 5 | Logical progression feels inevitable. Each section builds on the previous one. Heading hierarchy tells the complete story on its own. Transitions are seamless and purposeful. |
| 4 | Clear logical structure. Sections follow a coherent order. Heading hierarchy is sound. One or two transitions could be smoother but the reader never gets lost. |
| 3 | Structure is present but some sections feel out of order. A few transitions are abrupt or mechanical. Heading hierarchy has minor inconsistencies. |
| 2 | Significant structural problems. Sections jump between topics without clear logic. Several transitions are missing or forced. Reader has to mentally reorder the content. |
| 1 | No discernible structure. Content is arranged without apparent logic. Headings are misleading or absent. Reader cannot follow the progression. |
If below 4: Return to Architect. Re-examine the blueprint section order, heading hierarchy, and transition logic. The Architect should propose an alternative structure and get human approval before the Carpenter rebuilds transitions.
---
Dimension 4: Clarity and Readability
Phases Responsible: Carpenter + Judge
| Score | Description |
|---|---|
| 5 | Every sentence clear on first read. Complex ideas made accessible without oversimplification. Technical terms defined on first use. Sentence length varies naturally. No jargon without context. |
| 4 | Writing is clear throughout. One or two sentences may benefit from simplification. Technical content is accessible to the target audience. Readability metrics within target range. |
| 3 | Generally clear but several passages require re-reading. Some sentences are overlong or convoluted. Technical terms occasionally used without definition. Readability metrics borderline. |
| 2 | Frequent clarity issues. Multiple passages are confusing on first read. Jargon used without explanation. Sentences are consistently too long or too dense. |
| 1 | Pervasively unclear. Most paragraphs require multiple readings. Writing is inaccessible to the stated target audience. |
If below 4: Continue Judge passes. The Judge should run additional readability and clarity detection passes, flagging specific sentences and passages. If the problem is systemic rather than line-level, return to Carpenter for a clarity-focused rewrite of affected sections.
---
Dimension 5: Voice and Authority
Phases Responsible: Madman (user seeds) + Carpenter
| Score | Description |
|---|---|
| 5 | Confident, authoritative voice throughout. Expertise conveyed through specificity, not assertion. Author's unique perspective is evident. The piece could not have been written by just anyone. |
| 4 | Voice is consistent and confident. Author's perspective comes through in most sections. Expertise is demonstrated rather than claimed. Minor passages where voice flattens. |
| 3 | Voice is present but inconsistent. Some sections sound authoritative while others sound generic. A few passages read as though written by someone without domain expertise. |
| 2 | Voice is weak or inconsistent throughout. The piece reads as a competent summary rather than an authoritative take. Little evidence of the author's unique perspective. |
| 1 | No discernible author voice. The piece reads as generic content that could have been produced by anyone on the topic. No authority or unique perspective. |
If below 4: The human needs to inject more voice. Return to Carpenter with the human rewriting key passages in their own words and the Carpenter preserving that voice while polishing. This dimension cannot be fixed by AI alone -- it requires the human author's direct input, experience, and perspective.
---
Dimension 6: Opening and Hook
Phases Responsible: Madman + Carpenter
| Score | Description |
|---|---|
| 5 | Immediately compelling. Reader hooked within two sentences. The opening creates a question, tension, or promise that demands continued reading. No throat-clearing. |
| 4 | Strong opening that engages quickly. Reader motivated to continue within the first paragraph. Minor throat-clearing that could be trimmed but does not lose the reader. |
| 3 | Adequate opening that states the topic but does not compel. Reader continues out of interest in the topic rather than the writing. Some throat-clearing before the hook arrives. |
| 2 | Weak opening. Generic introduction that could belong to any article on the topic. Reader must push through multiple paragraphs before finding a reason to care. |
| 1 | No hook. The piece opens with background, definitions, or context that gives the reader no reason to continue. Throat-clearing dominates the first several paragraphs. |
If below 4: Return to Madman output for stronger hook candidates. Review the raw Madman material for compelling anecdotes, surprising data points, or provocative framings that were generated but not used. The Carpenter then rebuilds the opening around the strongest candidate.
---
Dimension 7: Conclusion and Takeaway
Phases Responsible: Architect + Carpenter
| Score | Description |
|---|---|
| 5 | Synthesizes the piece into a new insight that could not have been stated at the beginning. Clear, memorable takeaway. Reader knows exactly what to do next or think differently about. |
| 4 | Solid conclusion that ties back to the thesis. Takeaway is clear. Reader leaves with a concrete understanding of the main point and its implications. |
| 3 | Conclusion summarizes but does not synthesize. Takeaway is present but generic. Reader finishes without a strong sense of "so what" or "now what." |
| 2 | Weak conclusion that trails off or introduces new ideas. No clear takeaway. The piece ends without resolution. |
| 1 | No real conclusion. The piece stops rather than ends. No synthesis, no takeaway, no call to action. Reader is left wondering what the point was. |
If below 4: Return to Architect. The conclusion is a structural problem, not a prose problem. The Architect must revisit the throughline and define what insight the conclusion should deliver. Then the Carpenter rebuilds the conclusion to synthesize the argument into that insight.
---
Dimension 8: Technical Accuracy
Phase Responsible: Judge
| Score | Description |
|---|---|
| 5 | All facts verified. Technical details correct and current. Nuances captured accurately. No oversimplifications that mislead. Sources cited where appropriate. |
| 4 | Facts are accurate. Technical details are correct with minor simplifications that do not mislead. One or two points could benefit from additional precision but nothing is wrong. |
| 3 | Mostly accurate but one or two factual errors or misleading simplifications. Technical details are broadly correct but lack precision in places. |
| 2 | Multiple factual errors or significant oversimplifications. Technical details are wrong in ways that would undermine credibility with a knowledgeable reader. |
| 1 | Pervasive inaccuracies. The piece would misinform a reader. Technical details are fundamentally wrong or outdated. |
If below 4: Continue Judge with a focused accuracy review. The Judge should flag every factual claim, verify each against authoritative sources, and correct errors. If the inaccuracies stem from the source material, return to the research phase for better sources.
---
Dimension 9: SEO Optimization
Phases Responsible: Architect + Carpenter + Judge
Note: Score this dimension only for content with SEO objectives. Mark as N/A for content without SEO requirements.
| Score | Description |
|---|---|
| 5 | Primary keyword placed naturally in title, H1, first 100 words, at least two H2s, and conclusion. Semantic variations used throughout. Heading hierarchy follows SEO best practices. Meta description is compelling and within character limits. Internal linking opportunities identified. |
| 4 | Primary keyword present in title, H1, and first 100 words. Most H2s include keyword or semantic variation. Heading hierarchy is sound. Meta description present and adequate. |
| 3 | Primary keyword in title and body but placement is uneven. Some headings miss keyword opportunities. Meta description is generic or missing. Heading hierarchy has minor issues. |
| 2 | Keyword usage is sparse or forced. Headings do not reflect search intent. No meta description. Heading hierarchy is broken or flat. |
| 1 | No evidence of SEO consideration. Keywords absent from key positions. No heading structure. No meta description. Content does not align with search intent. |
If below 4: Run SEO validation from the seo-writer skill. The seo-writer should audit keyword placement, heading hierarchy, and meta description, then provide specific fixes. The Carpenter implements placement changes; the Judge verifies they read naturally.
---
Dimension 10: Originality and Value-Add
Phases Responsible: Madman (user seeds) + Pre-Research
| Score | Description |
|---|---|
| 5 | Contains original analysis, unique examples, and perspectives not found elsewhere. A reader familiar with existing content on this topic would learn something new. The piece advances the conversation. |
| 4 | Offers a distinct angle or original examples on an established topic. Most content adds value beyond what already exists. One or two sections cover well-trodden ground but with fresh framing. |
| 3 | Mix of original and derivative content. Some sections offer new perspectives while others repeat what is widely available. A knowledgeable reader would find partial value. |
| 2 | Mostly derivative. The piece covers the same ground as existing content with little new insight. Examples and analysis are generic rather than original. |
| 1 | Entirely derivative. Nothing in the piece could not be found in existing content on the topic. No original examples, analysis, or perspective. |
If below 4: Return to the research phase or Madman phase. The research phase should identify what existing content covers and where gaps exist. The Madman should then generate material that fills those gaps, drawing on the human's unique experience and perspective. This dimension, like Voice and Authority, requires human input to fix properly.