
Reframe Voice
- 38 installs
- 154 repo stars
- Updated July 30, 2026
- sammcj/agentic-coding
Helps with ai & agent building tasks.
About
reframe-voice is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted development.
- reframe-voice
- AI & Agent Building
- AI-coding skill
Reframe Voice by the numbers
- 38 all-time installs (skills.sh)
- +1 installs in the week ending Jul 26, 2026 (Skillselion tracking)
- Ranked #8,364 of 16,556 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 1, 2026 (Skillselion catalog sync)
npx skills add https://github.com/sammcj/agentic-coding --skill reframe-voiceAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 38 |
|---|---|
| repo stars | ★ 154 |
| Last updated | July 30, 2026 |
| Repository | sammcj/agentic-coding ↗ |
What it does
Helps with ai & agent building tasks.
Files
Reframe Voice
An evidence-led thought-leadership style built around a central reframe: taking a common belief and revealing a deeper, more useful way to think about it. The voice is that of a knowledgeable colleague who has done the work, not a lecturer or salesperson.
Three Core Principles
1. Evidence over assertion. Ground every claim in a named study, researcher, company, personal experience, or specific number. Never ask the audience to trust you without proof. 2. Balanced honesty over tribalism. Praise strengths and call out weaknesses regardless of affiliation. "I give credit where it's due" is a signature phrase. Never pick sides. 3. Practical reframe over surface take. The central move is replacing a widely held belief with a deeper framing. The audience leaves with a shifted mental model, not just new information.
Structural Arc
Every piece follows this order (sections flex in length):
1. Hook - Bold, contrarian opening. Pattern: [Strong claim] + [immediate complication]. Make them stop scrolling. 2. Stakes and Context - Why this matters right now. Cite a specific stat, study, or event. Often includes a personal anchor. 3. The Reframe - The signature move. [Common understanding] is wrong or incomplete. Here is the real issue. Should feel like a lock clicking open. 4. Evidence and Exploration - The longest section. Mix of named studies, specific numbers, real companies, personal anecdotes, and concrete scenarios. 5. Named Framework - Distill into a numbered, memorably named structure. Each component gets: definition, what good looks like, what bad looks like. 6. Practical Application - Advice segmented by audience role ("If you're an engineer...", "If you manage people...", "If you run an organisation..."). 7. Broader Implications - Scale outward. What does this mean for the industry, the economy, the decade? 8. Close - Brief, forward-looking, personal. No summary. End with "Cheers."
For review/comparison pieces, the reframe may be distributed and the framework may be a scorecard. The underlying rhythm still holds.
Voice
- Confident but not arrogant. Freely admit what you don't know.
- Opinionated but evidence-grounded. Earn the right to conclusions.
- Urgent but not alarmist. Thoughtful action, not panic.
- Direct but empathetic. Respect the audience. Never talk down.
- Empathetic toward practitioners doing hard work, even when critiquing their output.
Sentence style: Active voice. Varied length. Short punchy statements for emphasis ("Full stop.") interspersed with longer explanatory sentences. Heavy use of "you" and "your". No jargon without explanation.
Conversational transitions: "Look," / "Here's the thing," / "Let's be honest" / "On we go." / "You get the idea." / "And here's the thing"
Signature phrases: "I give credit where it's due" / "Full stop." / "This is not [X]. It's [Y]." / "Same [X], different [Y]." / "What does bad look like here?" / "I want to be honest with you" / "Cheers."
Never use: Marketing superlatives (game-changing, revolutionary), vague qualifiers without data, tribal dismissals, sycophantic hedging, emoji, filler transitions ("without further ado", "let's dive in").
Seven Rhetorical Techniques
Apply these throughout. Each is explained with examples in references/techniques.md.
1. Specific Analogy - Every major concept gets a vivid, concrete analogy. If they can't picture it, they won't remember it. 2. Paired Contrast - Same input, different human approach, dramatically different outcome. Pattern: "Same [X]. Different [Y]. Dramatically different [Z]." 3. Layered Example - For each framework component: what it is, what good looks like, what bad looks like. 4. Cross-Domain Validation - Prove the same point from 3+ unrelated disciplines. 5. Personal Anchoring - Concrete first-person experience. Not vague claims ("I've worked in this space") but specific scenes (a kitchen table, a particular eval, a conversation). 6. Preemptive Rejection - Name and reject the audience's expected objection before they form it. 7. Honest Concession - Credit the other side before criticising. Acknowledge inconvenient truths.
Quality Checklist
Before finishing, verify:
- [ ] Hook creates tension in the first two sentences
- [ ] Clear reframe shifts the audience's mental model
- [ ] Every major claim grounded in a named source or experience
- [ ] At least one honest concession to the opposing view
- [ ] Framework is named, numbered, and memorable
- [ ] Each framework component has a concrete example
- [ ] Practical advice segmented by audience role
- [ ] Analogies make abstract concepts visual
- [ ] At least one paired contrast
- [ ] Tone: confident, direct, empathetic (not salesy, not tribal)
- [ ] Closes with forward momentum and "Cheers"
- [ ] No marketing superlatives or vague qualifiers
Reference Files
references/techniques.md- Detailed examples of all seven rhetorical techniques, content pillars, and formatting rulesreferences/example.md- A fully worked example piece (RAG pipeline evaluation) demonstrating the style
Your RAG Pipeline Is Lying to You. Here's What's Actually Going Wrong
Every company I talk to right now is building a RAG pipeline. Retrieval augmented generation. You take your company's documents, you chunk them up, you embed them into a vector database, and then when someone asks a question, the system retrieves relevant chunks and feeds them to a language model. It sounds like the right architecture. It is the right architecture. And almost everyone is building it wrong.
Not wrong as in it doesn't work. Wrong as in it works just well enough to be dangerous. Your RAG system returns answers. Those answers sound confident. They cite your documents. And according to research from Databricks and Patronus AI, roughly 15 to 30% of RAG responses in enterprise deployments contain what they call context-grounded hallucinations. Not hallucinated from thin air. Worse. Assembled from real fragments of your actual data into conclusions your documents never supported.
That is a fundamentally different failure mode from what most teams are testing for. Full stop. It is the one that will cost you.
The Problem Isn't Retrieval. It's Assembly
When teams discover their RAG system is producing wrong answers, the instinct is to fix retrieval. Better embeddings. More sophisticated chunking. Reranking layers. Hybrid search combining vector and keyword approaches. And retrieval does matter. A study from Stanford's NLP group found that retrieval quality accounts for roughly 40% of downstream answer quality in RAG systems. That is significant and worth investing in.
But the other 60% is what happens after retrieval. The model receives five, ten, maybe twenty chunks of text and has to synthesise them into a coherent answer. This is where the real failures live, and almost nobody is testing for them systematically.
Think of it like a courtroom. Retrieval is the discovery phase, where lawyers collect all the relevant documents. Assembly is the closing argument, where someone has to take those documents and construct a coherent narrative. You can have perfect discovery, every relevant document in the pile, and still get a closing argument that misrepresents what the evidence actually shows. A lawyer who does that gets disbarred. Your RAG pipeline does it quietly, every single day, and nobody is checking the transcript.
I give credit where it's due. The teams building better retrieval are solving a real problem. Retrieval was genuinely bad eighteen months ago and the improvements have been meaningful. But if you stop there, you are optimising the discovery phase while nobody is watching the closing argument. And it is the closing argument your users actually see.
The Assembly Failure Triad
I've been running evals on RAG systems across multiple domains for the past several months. Legal documents, technical documentation, financial reports, internal knowledge bases. Last month I was evaluating a RAG pipeline for a fintech company. Their retrieval metrics were excellent. Precision above 90%. They were proud of those numbers and they should have been. Then I ran a set of queries that involved policy documents updated in the past year, the kind of questions a compliance officer would ask on a Tuesday morning, and the system confidently cited the superseded policy on four out of ten queries. Same pipeline, same embeddings, same vector store. The only difference was whether the query touched documents that had been updated. That is the moment the team stopped talking about retrieval and started talking about what happens after retrieval.
I'm going to walk you through the three failure patterns that show up everywhere and then lay out a three-layer evaluation architecture for catching them. The failure patterns are structural. They appear regardless of the embedding model, regardless of the vector database, regardless of the LLM doing the synthesis. These are not configuration issues you can tune away.
Failure mode one: temporal collision. The system retrieves chunks from documents written at different points in time and treats them as simultaneously true. Your 2024 pricing policy says enterprise contracts start at $50,000 annually. Your 2025 pricing policy says $75,000. Both are in the vector database. Both are semantically similar to the query "what is our enterprise pricing?" The model retrieves both, and instead of identifying the conflict and surfacing the most current answer, it either picks one without telling you which, averages them into something neither document says, or confidently states the outdated figure because it appeared in a longer, more detailed document that the model weighted more heavily.
This is not a retrieval problem. Both documents were correctly retrieved. It is an assembly problem. The model has no reliable mechanism for temporal reasoning across chunks that lack explicit date metadata in the text itself.
Think about what a good research librarian does when you ask a question and she finds two sources that contradict each other. She does not blend them. She tells you: "The 2024 edition says X. The 2025 edition says Y. Which timeframe are you asking about?" The model skips that step entirely. It merges the sources and hands you a confident answer that neither source supports.
What does good look like for temporal collision? A system that detects date-bearing chunks in the retrieved set and either surfaces the most recent by default with a note that older versions exist, or asks the user to specify a timeframe. What does bad look like? What most teams have today. The model picks a source based on embedding similarity or chunk length, the user has no idea which version they got, and the answer might be eighteen months out of date.
Failure mode two: scope bleed. The system retrieves chunks that are individually correct but apply to different contexts, and the model merges them. Your engineering documentation says the API rate limit is 1,000 requests per minute. That is true for the public API. Your internal documentation says there is no rate limit. That is true for the internal service mesh. A developer asks "what is the API rate limit?" and gets back an answer that combines both contexts into something like "the rate limit is 1,000 requests per minute, though this can be bypassed for internal services." That sounds helpful. It is also a security-relevant disclosure that should never have been assembled from those two sources together in that way.
This is the equivalent of a pharmacist reading two prescription labels, one for you and one for the patient before you, and giving you a combined dosage instruction. Both labels are correct. Combining them is dangerous. And the pharmacist would never do that because pharmacists are trained to verify scope, which patient this applies to, before acting on any piece of information. Models doing RAG synthesis have no equivalent of that verification step.
Again, retrieval worked. Both chunks were relevant to the query. The failure is in assembly. The model cannot reliably distinguish scope boundaries across retrieved chunks.
Failure mode three: confidence inheritance. The model retrieves a mix of authoritative and speculative content and presents the synthesis with the confidence level of the most authoritative source. Your approved product roadmap says feature X ships in Q3. A Slack thread from an engineer says "we might also ship feature Y if we have bandwidth." The RAG system retrieves both and responds: "Features X and Y are planned for Q3." The speculative content inherited the confidence of the authoritative content through the act of assembly.
Yoshua Bengio's group at Mila published research in 2024 showing that language models systematically fail to propagate uncertainty through multi-step reasoning. When one premise is uncertain, the conclusion should inherit that uncertainty. Instead, models tend to converge on confident outputs regardless of input uncertainty. That finding was about reasoning chains, but the mechanism is identical in RAG assembly. Uncertain inputs produce confident outputs because the model has no architecture for tracking provenance and confidence separately.
This one is particularly hard to catch because the answer is not wrong in the way that triggers traditional evals. It is directionally plausible. It contains real information. It just presents speculation as commitment, and that distinction matters enormously when the person reading the answer is a customer, an executive, or an auditor.
And here's the thing. This is not just an AI problem. It is exactly the failure mode that gets journalists fired. A reporter who takes an on-the-record statement from a CEO and a rumour from an anonymous source and presents both with equal confidence in the same paragraph has committed a basic journalistic error. Newsrooms have editorial standards specifically to prevent this. Your RAG pipeline has no equivalent editorial layer. It assembles sources with mixed authority into a single confident voice, every single time.
The RAG Fidelity Stack
Before you jump to the conclusion that I am telling you to abandon RAG, I am not. RAG is the right architecture for grounding language models in your data. The issue is not the pattern. It is the gap between how teams build RAG systems and how they evaluate them.
I call this the RAG Fidelity Stack. It has three layers, and most teams only have the first one.
Layer one: retrieval quality. Are the right chunks being retrieved? This is what most teams measure. Precision, recall, MRR, NDCG. Pinecone published benchmarks last year showing that hybrid search with reranking can push retrieval precision above 90% on well-structured corpora. These metrics are well understood, well tooled, and necessary. They are also insufficient on their own.
Layer two: assembly fidelity. Given correctly retrieved chunks, does the model's synthesised answer faithfully represent what those chunks actually say? This requires comparing the final answer against the source chunks and checking for temporal collisions, scope bleeds, and confidence inheritance. You can build this with an LLM-as-judge pattern, but the judge needs specific rubrics for each failure mode, not a generic "is this answer correct?" prompt. Ragas, the open source RAG evaluation framework, provides faithfulness and answer relevancy metrics that get at this layer, and teams using it report catching assembly errors that their retrieval metrics completely missed.
Layer three: boundary honesty. When the retrieved chunks do not contain enough information to fully answer the question, does the system say so? Or does it fill the gap with plausible-sounding extrapolation and present it as grounded? This is the hardest layer to test and the one with the highest cost when it fails, because the whole point of RAG is to ground the model in your data. If the system silently switches from grounded to ungrounded mid-answer, the user has no way to tell where the boundary is.
What does bad look like? Testing only layer one and celebrating 90% retrieval precision while your assembly layer introduces errors on 20% of multi-source queries. Or worse, testing all three layers once during development and never again, while your document corpus grows and shifts underneath the frozen eval set. That is not evaluation. That is a snapshot.
Same Pipeline, Different Evaluation, Different Outcome
I want to make this concrete. I have seen two teams with nearly identical RAG architectures. Same embedding model, same vector database, same LLM for synthesis. Team A tested retrieval quality obsessively. They built dashboards. They tracked precision and recall weekly. Their retrieval metrics were best in class. But they never tested what the model did with the retrieved chunks. Team B had decent retrieval, nothing special, but they built assembly-level evals. They tested for temporal conflicts, scope boundaries, and confidence mixing using their own documents. Six months in, Team A's system had silently drifted. Updated documents were creating temporal collisions that nobody caught because the retrieval metrics still looked great. Team B caught the same class of errors within days because their evals were designed to surface them.
Same architecture. Same tools. Different evaluation discipline. Dramatically different trust outcomes. The difference was not technology. It was whether the humans building the system understood that retrieval quality and answer quality are not the same measurement.
This Pattern Is Not New
Let's be honest. The assembly problem in RAG is not actually a novel failure mode. It is the same problem that has shown up in every domain where humans or machines have to synthesise information from multiple sources under time pressure.
Intelligence analysis has dealt with this for decades. The 9/11 Commission Report documented how individual intelligence agencies each had correct fragments of information, but the assembly of those fragments into a coherent threat assessment failed catastrophically. The data was retrieved. The synthesis was wrong. The commission's core recommendation was not better collection. It was better integration, a structural fix to the assembly layer.
Medical diagnosis follows the same pattern. A 2023 study in BMJ Quality and Safety found that diagnostic errors in hospitals are most commonly caused not by missing information but by the incorrect integration of available information. The right test results were in the chart. The synthesis into a diagnosis went wrong. Same retrieval, broken assembly.
And it is the same pattern in academic peer review. Reviewers who read five papers on a topic and write a literature review must decide which findings to weight, which contradict, which apply in which context. The quality of the review is not determined by whether they found the papers. It is determined by whether they assembled the findings faithfully. Three independent disciplines. Same conclusion. The hard part is never the retrieval. It is the assembly.
What This Means for Your Team
If you are an engineer building a RAG pipeline, start logging the retrieved chunks alongside every answer, not just the final output. You cannot debug assembly failures without seeing what the model was working with. Build at least five eval cases for each of the three assembly failure modes using your actual documents. Temporal conflicts are the easiest to construct and the most immediately revealing. If you have documents that were updated in the past year, you already have the test data. Use it.
If you manage a team shipping RAG to users, ask your engineers a simple question: what happens when the system retrieves contradictory information? If the answer is "we haven't tested that," you have a gap that matters more than your next feature. Every document update, every policy change, every new data source creates potential temporal collisions in your pipeline, and the failure rate compounds over time as your corpus grows. The system that worked at 500 documents may fail at 5,000. Not because retrieval degraded, but because the assembly problem gets harder as the source material gets more internally contradictory.
If you run an organisation that is investing in RAG as a knowledge management strategy, understand that retrieval is table stakes and assembly quality is the differentiator. The vendors who can show you assembly-level evals, not just retrieval metrics, are the ones building systems that will hold up when the stakes are real. Ask them: "How do you handle temporal conflicts in the retrieved context?" If they talk about embeddings, they have not solved the problem.
The Bigger Picture
RAG is following the same maturity curve as every other AI production pattern. The first generation of adopters built it, shipped it, and measured what was easy to measure. Retrieval metrics are easy. Assembly fidelity is hard. And so the industry optimised for retrieval while the real failure mode sat in the synthesis layer, quietly generating answers that were wrong in ways that standard evals could not catch.
This is not unique to RAG. It is the same pattern in agent evaluation, where the Mount Sinai health study found that ChatGPT Health could correctly identify respiratory failure in its reasoning trace and then recommend waiting 48 hours in its output. It is the same pattern in coding assistants, where the generated code handles the happy path and breaks on edge cases that the model never flagged. Everywhere you look, the gap between "looks right" and "is right" is the gap that determines whether the system creates value or liability. And everywhere you look, that gap lives in the assembly layer, not the retrieval layer.
The teams that figure out how to test for assembly fidelity are the ones whose systems will actually be trusted by the humans who depend on them. And trust, once lost to a confident wrong answer delivered from your own documents, is very expensive to rebuild.
So look at your RAG pipeline. Look at whether you are testing retrieval or testing the whole system. Build the assembly evals. I know they are not glamorous work. But they are the difference between a system that sounds right and a system that is right. And in a world where AI made volume free, correctness is the only thing that matters.
Cheers.
Rhetorical Techniques Reference
Detailed examples for each of the seven techniques. Read this file when you need deeper guidance on executing a specific technique.
1. The Specific Analogy
Every major concept gets a vivid, concrete analogy:
- Expanding bubble for AI capability (surface = frontier, inside = agent territory, outside = still human)
- Calculator ban of 1975 as parallel for AI resistance in education
- Surfing swells for capability forecasting ("A surfer doesn't predict the next wave but reads the sea")
- AI exoskeleton for tools that extend human capability
- Aircraft carriers on a fishing route for misallocated capability
- "Like calling surgery scalpel handling" for reducing frontier ops to prompt engineering
- Courtroom discovery vs closing argument for retrieval vs assembly in RAG pipelines
- Pharmacist combining two patients' prescriptions for scope bleed errors
2. The Paired Contrast
Present two outcomes from identical starting conditions:
- The Maltbot agent that saved $4,200 vs the agent that spammed 500 people. Same technology, different specification quality.
- A kid who asks ChatGPT to write the essay vs one who drafts, uses AI to find weak arguments, strengthens them. Same tool, different metacognition.
- Team of five optimising for correctness vs team of 20 optimising for volume. Same AI tools, different outcomes.
- Claude's one-sentence correct answer vs GPT 5.4's long, carefully structured wrong answer.
- Two teams with identical RAG architectures. One tests retrieval only, one tests assembly. Same pipeline, different evaluation discipline, dramatically different trust outcomes.
Pattern: Same [input/tool]. Different [human skill/approach]. Dramatically different [outcome].
3. The Layered Example
For each framework component, provide three layers:
1. What it is (brief definition) 2. What good looks like (concrete, often role-specific scenario) 3. What bad looks like (concrete failure scenario)
Example from frontier operations: "Boundary sensing is maintaining accurate intuition about where the human-agent boundary sits. Good: a product manager lets an agent draft competitive analysis but reserves stakeholder dynamics for herself. Bad: a marketing director calibrated six months ago and hasn't noticed the boundary moved."
4. Cross-Domain Validation
Show the same pattern across 3+ unrelated fields:
- Team size of five validated across evolutionary psychology (Dunbar), military operations (fire teams), and software engineering (Brooks, Bezos)
- Assembly-over-retrieval validated across intelligence analysis (9/11 Commission), medical diagnosis (BMJ Quality and Safety), and academic peer review
- Foundation-before-leverage validated across education research (Bloom), AI agent performance, and personal parenting experience
Tie them together explicitly: "Three independent disciplines. Same conclusion."
5. Personal Anchoring
Ground arguments in concrete first-person experience:
- "I have three kids. I work in AI every day."
- "My 10-year-old sat at the kitchen table last month working through long division by hand."
- "I ran blind evals. I ran real world tests. And I ran side-by-side comparisons so you don't have to do the guesswork."
- "Last month I was evaluating a RAG pipeline for a fintech company. Their retrieval metrics were excellent. Precision above 90%..."
Pattern: [I did the thing]. [Here is what specifically happened]. [That is why I am telling you this]. The scene must be concrete enough to picture.
6. The Preemptive Rejection
Name and reject the audience's expected objection before they form it:
- "Before you jump to the conclusion that I am recommending firing people, I am not going to do that. And you're going to see why."
- "Instead of making this a session where we talk about the defects of ChatGPT health, I want to make it a conversation about AI agents."
- "I am not in the protect-the-children-from-AI camp."
Pattern: [I know you expect me to say X]. [I'm not going to say that]. [Here's the more interesting thing].
7. The Honest Concession
Credit the other side before criticising:
- "5.4 did an extraordinary job finding and parsing all of those sources... But it also let some dirty data through."
- "If you're teaching someone about Claude, you got to be honest about what ChatGPT does better or there's no point."
- "I give credit where it's due. The teams building better retrieval are solving a real problem. Retrieval was genuinely bad eighteen months ago and the improvements have been meaningful."
Credibility comes from acknowledging inconvenient truths. Never cherry-pick only favourable evidence.
Content Pillars
These recurring themes form the intellectual backbone. Weave them in where relevant:
- Human Specification Quality - The quality of human direction is the single most important variable in AI outcomes
- Foundation Before Leverage - Build the human skill first, then extend with AI
- Correctness Over Volume - AI made volume free; the scarce resource is whether the output is right
- Balanced Assessment - Refuse to pick sides in model wars or adoption debates
- The Expanding Frontier - AI capability is an expanding boundary, not a destination
- Ambition Expansion Over Cost Reduction - "You didn't get a cost reduction. You got an army."
Formatting
- Name every framework. Concrete, visual, two to four words. Not "Best Practices" but "The Assembly Failure Triad."
- Number every component. "Failure mode one...", "Here's the second skill set..."
- Signpost before delivering. "I'm going to walk you through three failure patterns and then lay out a three-layer architecture for catching them."
- Conversational transitions. "On we go." / "So, let's get to it." / "Let's talk about strike teams."
- Soft CTAs. "I've written it all up on the Substack" not aggressive asks. Sign off with "Cheers."
- Length: 2,500 to 7,000 words. Frameworks and examples drive length, not padding.