
Research Intake
- 9 installs
- 30 repo stars
- Updated July 7, 2026
- jeffallan/writing-with-agents
Traverses and indexes source material (vaults, markdown notes, docs, URLs) into a structured knowledge map and identifies research gaps before writing.
About
Ingests arbitrary source material and builds a structured knowledge map of what the corpus contains, lacks, and how it connects. A writer uses it at the start of a project to understand existing research and spot gaps before content creation.
- Handles arbitrary nested notes/vaults with no assumed methodology
- Produces a knowledge map plus gap analysis, not an outline
Research Intake by the numbers
- 9 all-time installs (skills.sh)
- +1 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #1,150 of 1,879 Documentation skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/jeffallan/writing-with-agents --skill research-intakeAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 9 |
|---|---|
| repo stars | ★ 30 |
| Last updated | July 7, 2026 |
| Repository | jeffallan/writing-with-agents ↗ |
What it does
Traverses and indexes source material (vaults, markdown notes, docs, URLs) into a structured knowledge map and identifies research gaps before writing.
Files
Role Definition
The Research Intake specialist traverses, indexes, and maps source material into a structured knowledge map before any content creation begins.
Lead: AI traverses and indexes source material, builds the knowledge map, identifies gaps. Support: Human steers gap-filling priorities, validates the map, and seeds with context about what matters.
This skill handles Obsidian vault files, markdown notes, reference documents, prior research, and web URLs. It assumes arbitrary nested file and folder structures with no predetermined naming conventions or organizational schemes. The skill does not assume Zettelkasten, PARA, or any other specific note-taking methodology.
The knowledge map is a structured inventory of what the research corpus contains, what it lacks, and how its pieces connect. It is not an outline or a content plan -- those belong to downstream skills.
When to Use This Skill
- Starting a writing project and need to understand what material already exists
- Ingesting an Obsidian vault, folder of markdown notes, or collection of reference documents
- Building a research corpus from scattered sources before content planning
- Identifying what gaps exist in existing research before generating new material
- Preparing source material for handoff to the content-strategist or madman phase
- Auditing a knowledge base to understand coverage and depth across topics
- Onboarding to a new domain where prior research exists but is unorganized
- Combining multiple research sources into a unified understanding
- Returning to a dormant project and need to rediscover what research already exists
- Validating that source material has enough depth to support a planned content calendar
- Cross-referencing claims across multiple documents to find contradictions or reinforcement
- Preparing for a content audit where you need to map what topics are already covered and to what depth
- Triaging a large collection of notes to determine which are relevant to a specific writing project
Core Workflow
1. Establish session context -- Use AskUserQuestion to determine the source material location. Ask for the vault path, working directory, or confirmation that no vault exists. If a vault path is provided, confirm it before traversal. If no vault exists, ask the human to provide files, folders, or URLs directly.
2. Traverse and index all source material -- Recursively read every file in the provided path. For each file, extract: title or filename, topics covered, key claims made, sources cited, depth of coverage (deep, moderate, or surface), and any metadata present. Do not skip files based on naming or format assumptions. Read everything.
3. Build knowledge map -- Synthesize the indexed material into a structured knowledge map. Group topics with depth assessments. Extract key claims with their supporting sources. Identify connections between topics across different files. Catalog all existing sources. Surface gaps where coverage is thin or missing. See references/intake-process.md for the full indexing methodology and output format.
4. Present gaps for human steering -- Use AskUserQuestion to present identified gaps as a multi-select list. The human chooses which gaps to investigate. Do not fill gaps without explicit human selection. For each selected gap, propose 3 specific research directions plus a skip option. See references/gap-analysis.md for gap categories and the investigation process.
5. Fill selected gaps and offer vault capture -- Research the human-selected gaps via web search and synthesis. Present findings for validation before any storage. Offer to capture findings as structured notes in the vault path established in step 1. The knowledge map plus any gap-fill results become the input to the content-strategist or madman phase.
Reference Guide
| Topic | Reference | Load When |
|---|---|---|
| Indexing methodology, traversal process, knowledge map format | references/intake-process.md | Traversing sources, building the knowledge map |
| Gap categories, investigation process, vault capture format | references/gap-analysis.md | Presenting gaps, filling selected gaps, capturing findings |
Constraints
MUST DO:
- Traverse all provided material -- do not skip files based on name, size, or assumed relevance
- Build a complete knowledge map before presenting gaps
- Present gaps to the human for steering before investigating any of them
- Offer vault capture for all research findings
- Confirm the vault path or working directory before writing any files
- Flag claims that appear in source material without supporting evidence
- Record the depth assessment (deep, moderate, surface) for every indexed topic
- Distinguish between topics that are discussed and topics that are merely mentioned
- Include file paths in every knowledge map entry so the human can trace claims back to their source
- Present the complete knowledge map before moving to gap identification
MUST NOT DO:
- Assume Zettelkasten, PARA, or any specific note-taking format
- Fill gaps without explicit human approval of which gaps to investigate
- Skip the vault capture offer after gap-filling research
- Impose organizational structure on the existing vault or notes
- Modify existing source files during the intake process
- Proceed to content planning -- that belongs to the content-strategist skill
- Summarize or paraphrase source material during indexing -- extract structure and metadata, preserve original content
- Discard files that appear irrelevant based on filename alone -- open and index every file
- Merge or consolidate source files during intake -- the knowledge map is read-only on the source corpus
- Present gaps without categorizing them -- every gap needs a category to determine the right investigation approach
- Generate content or draw conclusions during the intake process -- intake is indexing and mapping, not synthesis
Output Templates
# Knowledge Map: [Domain]
## Topics Covered
- [Topic A]: [deep / moderate / surface] -- [primary source files]
- [Topic B]: [deep / moderate / surface] -- [primary source files]
## Key Claims and Arguments
- [Claim 1] -- supported by [source/note], strength: [strong / moderate / weak]
- [Claim 2] -- supported by [source/note], strength: [strong / moderate / weak]
## Existing Sources
- [Source 1]: [what it covers], [file location]
- [Source 2]: [what it covers], [file location]
## Connections Identified
- [Topic A] relates to [Topic C] through [mechanism]
- [Claim 2] contradicts [Claim 5] on [specific point]
## Gaps Identified
1. [Gap description] -- [category: undeveloped / unsupported / missing perspective / outdated]
2. [Gap description] -- [category]Knowledge Reference
The knowledge map is the central artifact of this skill. It serves as a structured inventory rather than a content plan. The distinction matters: a knowledge map says "here is what exists, here is what is missing, here is how pieces connect." A content plan says "here is what to write and in what order." The research-intake skill produces the former. Downstream skills like the content-strategist consume the knowledge map and transform it into editorial decisions.
Gap analysis follows a four-category model: undeveloped topics (mentioned but not explored), unsupported claims (asserted without evidence), missing perspectives (one-sided coverage), and outdated material (superseded by newer information). Each category implies a different research action. Undeveloped topics need exploratory research. Unsupported claims need source verification. Missing perspectives need deliberate counter-sourcing. Outdated material needs current-state research. Categorizing gaps before investigating them prevents wasted effort on low-value research directions.
The depth assessment scale (deep, moderate, surface) provides a quick triage of coverage quality. Deep coverage means the source contains detailed evidence, multiple supporting examples, and nuanced argumentation. Moderate coverage means the topic is addressed with some evidence but lacks exhaustive treatment. Surface coverage means the topic is mentioned or referenced without substantive exploration. This three-level scale is deliberately coarse to enable fast indexing across large source collections without getting bogged down in granular scoring.
Connection mapping across source files reveals relationships that no single document contains. When two notes discuss the same concept using different terminology, the knowledge map surfaces this overlap. When one note's conclusion contradicts another's premise, the knowledge map flags the tension. These cross-file connections are among the most valuable outputs of the intake process because they represent insights that exist in the corpus but are invisible to anyone reading files in isolation.
The human steering step for gap-filling prevents wasted research effort. Not every gap is worth investigating. A gap in a peripheral topic may be irrelevant to the planned content. A gap in a core topic may be critical. Only the human can make this judgment because only the human knows the editorial intent. Presenting gaps as a structured multi-select list with categories and proposed research directions gives the human enough information to decide without requiring them to formulate the research plan themselves.
Gap Analysis
This reference details how to identify gaps in a research corpus, present them to the human for steering, investigate selected gaps, and capture findings for vault storage.
---
Gap Identification Categories
After building the knowledge map, scan the corpus for gaps in each of the following categories:
1. Topics Mentioned but Not Developed
A topic appears in the corpus -- referenced in a claim, listed in a connection, or named in a heading -- but has no substantive coverage. The corpus acknowledges the topic exists without explaining, arguing, or providing evidence about it.
Signal: Topic depth rated as "surface" in the knowledge map, or topic appears only in connection mappings without its own coverage.
2. Claims Without Evidence
A claim is stated in the corpus but lacks supporting data, citations, examples, or reasoning. The claim may be correct, but the corpus does not make the case.
Signal: Evidence strength rated as "weak" in the knowledge map. The claim exists as assertion only.
3. Contradictions Between Sources
Two or more files in the corpus make opposing claims about the same topic without resolution. The contradiction may be genuine (the sources disagree) or apparent (the sources address different aspects that appear to conflict).
Signal: Contradictions surfaced during knowledge map synthesis. Both claims exist without a file that resolves or addresses the tension.
4. Missing Perspectives or Counterarguments
The corpus presents a position without engaging opposing views. The coverage is one-sided, which weakens the argument and leaves the content vulnerable to reader objections.
Signal: Key claims exist without corresponding counterarguments. No file in the corpus addresses the strongest objection to the main thesis.
5. Adjacent Topics That Would Strengthen Coverage
Topics not present in the corpus that would meaningfully strengthen the content if included. These are topics the reader would expect to see or that would provide necessary context for existing claims.
Signal: Connection mapping reveals dependencies on topics not covered. Reader questions (anticipated) point to areas the corpus does not address.
6. Outdated Information
Sources or claims in the corpus reference data, statistics, or conditions that may no longer be current. The passage of time has potentially invalidated or weakened the evidence.
Signal: Source dates are more than 2 years old for fast-moving topics, or claims reference specific data points without dates.
---
Gap Presentation
Present identified gaps to the human using AskUserQuestion. Format the presentation as a numbered multi-select list grouped by category.
Presentation Format
I identified the following gaps in your research corpus. Which would you like me to investigate?
**Topics mentioned but not developed:**
1. [Topic X] -- mentioned in [file1.md] but no substantive coverage
2. [Topic Y] -- referenced as a connection to [Topic A] but unexplored
**Claims without evidence:**
3. [Claim Z] -- asserted in [file2.md] without supporting data
4. [Claim W] -- stated in [file3.md], needs citation or example
**Missing perspectives:**
5. No counterargument to [main thesis] found in corpus
6. [Stakeholder group] perspective absent from coverage
**Adjacent topics:**
7. [Adjacent topic] would provide necessary context for [Claim 1]
Select by number (e.g., "1, 3, 5") or say "all" or "none."Do not investigate any gaps until the human responds with their selection.
---
Gap Investigation
For each gap the human selects, propose 3 specific research directions plus a skip option.
Investigation Options Format
For gap [N]: [gap description]
I can investigate via:
A) [Specific research direction 1] -- [what this would cover]
B) [Specific research direction 2] -- [what this would cover]
C) [Specific research direction 3] -- [what this would cover]
D) Skip this gap for now
Which direction?Research Execution
For the selected direction:
1. Search and gather. Use web search to find relevant sources. Prioritize authoritative sources: peer-reviewed research, recognized experts, reputable publications, official documentation.
2. Synthesize findings. Do not dump raw search results. Synthesize what was found into a coherent summary with:
- Key findings relevant to the gap
- Sources consulted (with URLs and dates)
- How the findings connect to existing corpus material
- Any new gaps or questions surfaced by the research
3. Present for validation. Deliver the synthesized findings to the human. The human confirms accuracy, relevance, and whether the gap is adequately filled.
4. Iterate if needed. If the human says the gap is not adequately filled, propose 2 follow-up directions. Do not repeat the same search with different phrasing.
---
Vault Capture
After gap-filling research is validated, offer to capture findings as structured notes in the vault.
Capture Offer Format
Research findings are validated. Would you like me to save these to your vault?
I would create the following notes:
1. [Proposed filename] -- [what it captures]
2. [Proposed filename] -- [what it captures]
Save location: [vault path from session setup]
Proceed with capture? (yes / no / modify filenames)Vault Note Format
Each captured note follows this structure:
---
type: research-source
domain: [domain topic from knowledge map]
date_captured: [ISO 8601 date, e.g., 2026-02-08]
source_url: [URL if applicable]
source_author: [author if known]
gap_filled: [description of the gap this research addresses]
tags:
- [topic tag 1]
- [topic tag 2]
---
# [Source or Finding Title]
**Author:** [author name or "synthesized from multiple sources"]
**URL:** [url or "N/A"]
**Date accessed:** [date]
## Key Claims
- [claim 1]
- [claim 2]
- [claim 3]
## Evidence and Data
[Supporting data, statistics, or examples found during research]
## Connection to Existing Research
- Relates to: [existing vault note or topic from knowledge map]
- Supports: [existing claim in corpus]
- Contradicts: [existing claim, if applicable]
- Fills gap: [original gap description]
## Open Questions
- [Any questions raised by this research that remain unanswered]Filename Conventions
- Use lowercase with hyphens:
api-rate-limiting-research.md - Prefix with domain when the vault covers multiple domains:
fintech-trust-deficit.md - Do not nest in subdirectories unless the vault already uses a folder structure -- match the existing convention
---
Handoff
After intake is complete, the knowledge map plus any gap-fill notes are ready for:
- Content-strategist: If the goal is to plan multiple articles from the research corpus
- Madman: If the goal is to write a single article and the research provides the seed material
State the available next steps to the human and let them choose the path forward.
Research Intake Process
This reference details the step-by-step process for traversing source material, indexing content, and building a knowledge map.
---
Session Setup
Before traversal begins, establish the session context using AskUserQuestion. The question determines the source material location and the working directory for any output files.
Setup Questions
Ask one of the following based on context:
1. Vault path known: "I see you mentioned [path]. Should I traverse this location for source material?" 2. No path provided: "Where is your source material located? Provide a vault path, folder path, or tell me you will paste content directly." 3. Multiple sources: "I can ingest material from multiple locations. What is the primary path, and are there additional folders or URLs to include?"
Store the confirmed path as the session vault path. All vault capture operations later in the workflow will use this path.
---
Source Material Types
The intake process handles all of the following without requiring the user to specify which types are present:
| Source Type | How to Handle |
|---|---|
| Obsidian vault (.md files with frontmatter, wikilinks) | Read all files recursively. Parse YAML frontmatter for metadata. Follow wikilinks to map connections. |
| Plain markdown notes (.md without Obsidian conventions) | Read all files recursively. Use headings and content to infer topics. |
| Reference documents (.md, .txt) | Read full content. Extract claims, sources cited, and topic coverage. |
| Prior research summaries | Read and index as high-value sources. Flag depth of coverage per topic. |
| Web URLs (provided by human) | Fetch and read content. Extract key claims, author, date, and relevance. |
| PDF content (pasted or referenced) | Process provided text. Note that direct PDF reading depends on tooling availability. |
File Traversal Rules
1. Read everything. Do not skip files based on filename, folder name, or file size. A file named "scratch.md" may contain key insights. A file in a folder named "archive" may contain foundational research. 2. Respect depth. Traverse all nested folders to arbitrary depth. Do not stop at a maximum folder depth. 3. Note file metadata. Record filename, path, modification date if available, and file size as context signals. Recently modified files may indicate active research areas. 4. Handle encoding gracefully. If a file cannot be read, log the path and continue. Do not halt traversal on read errors.
---
Indexing Process
For each file read during traversal, extract and record the following:
Per-File Index Entry
### [Filename]
- **Path:** [full path]
- **Topics:** [comma-separated topic list]
- **Depth:** [deep / moderate / surface]
- **Key claims:**
- [claim 1]
- [claim 2]
- **Sources cited:** [any external sources referenced in this file]
- **Connections:** [wikilinks, explicit references to other files, or thematic connections]
- **Notes:** [anything notable -- contradictions, strong passages, gaps within the file]Depth Assessment Criteria
| Depth Level | Criteria |
|---|---|
| Deep | Topic is the primary focus. Multiple claims with supporting evidence. Detailed treatment with nuance. |
| Moderate | Topic receives meaningful coverage but is not the primary focus. Some claims with partial evidence. |
| Surface | Topic is mentioned or touched on briefly. Few or no supporting claims. Quick reference without development. |
---
Building the Knowledge Map
After all files are indexed, synthesize the per-file entries into a unified knowledge map. The knowledge map is the primary output of the intake process.
Synthesis Steps
1. Aggregate topics. Merge topic mentions across all files. For each topic, record the highest depth level found and list all files that cover it.
2. Extract key claims. Pull the strongest claims from across all files. For each claim, note which files support it, whether evidence is provided, and whether any files contradict it.
3. Catalog sources. List all external sources cited across the corpus. Note what each source covers and which files reference it.
4. Map connections. Identify relationships between topics:
- Explicit connections (wikilinks, cross-references, citations)
- Thematic connections (topics that appear together in multiple files)
- Causal or logical connections (topic A depends on or leads to topic B)
- Contradictions (files that make opposing claims about the same topic)
5. Identify gaps. Surface areas where the corpus is thin or missing. See gap-analysis.md for the full gap identification methodology.
Knowledge Map Output Format
# Knowledge Map: [Domain]
## Domain Summary
[2-3 sentences describing the overall scope and character of the research corpus]
## Topics Covered
- [Topic A]: deep -- covered in [file1.md], [file2.md], [file3.md]
- [Topic B]: moderate -- covered in [file2.md], [file4.md]
- [Topic C]: surface -- mentioned in [file1.md]
## Key Claims and Arguments
- [Claim 1] -- supported by [file1.md], [file3.md]; evidence strength: strong
- [Claim 2] -- supported by [file2.md]; evidence strength: moderate
- [Claim 3] -- asserted in [file4.md]; evidence strength: weak (no supporting data)
- [Claim 4] -- contradicted between [file1.md] and [file5.md]
## Existing Sources
- [Source 1]: covers [topics], cited in [file1.md], [file2.md]
- [Source 2]: covers [topics], cited in [file3.md]
## Connections Identified
- [Topic A] relates to [Topic C] through [mechanism]
- [Topic B] depends on [Topic D] (which is a gap -- see below)
- [Claim 2] contradicts [Claim 4] on [specific point]
- [File1.md] and [File3.md] approach [Topic A] from complementary angles
## Gaps Identified
1. [Gap description] -- category: [see gap-analysis.md for categories]
2. [Gap description] -- category: [see gap-analysis.md for categories]
## Corpus Statistics
- Total files indexed: [n]
- Topics identified: [n]
- Key claims extracted: [n]
- External sources cataloged: [n]
- Gaps identified: [n]---
Completion Check
Before delivering the knowledge map, verify:
- [ ] All files in the provided path were read (or logged as unreadable)
- [ ] Every file has a per-file index entry
- [ ] Topics are aggregated with depth assessments
- [ ] Key claims include evidence strength ratings
- [ ] Connections include both explicit and thematic links
- [ ] Contradictions between sources are surfaced, not hidden
- [ ] Gaps are identified and categorized
- [ ] Corpus statistics are accurate
- [ ] The domain summary accurately reflects the corpus scope