
Ccf Literature Searcher
- 23 installs
- 1.5k repo stars
- Updated July 8, 2026
- mikubaka88/ccfa-skills
Search and screen literature, related work, datasets, benchmarks, and citation evidence, and build research-opportunity maps for CCF workflows.
About
This skill searches and screens literature, related work, prior art, datasets, and benchmarks and builds research-opportunity maps for CCF research. A developer uses it for literature search and direction scouting, not for auditing only already-cited references or writing the manuscript.
- Searches literature, related work, datasets, and benchmarks
- Builds research-opportunity maps for direction scouting
Ccf Literature Searcher by the numbers
- 23 all-time installs (skills.sh)
- Ranked #10,032 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mikubaka88/ccfa-skills --skill ccf-literature-searcherAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 23 |
|---|---|
| repo stars | ★ 1.5k |
| Last updated | July 8, 2026 |
| Repository | mikubaka88/ccfa-skills ↗ |
What it does
Search and screen literature, related work, datasets, benchmarks, and citation evidence, and build research-opportunity maps for CCF workflows.
Files
CCF Literature Searcher
Invocation Controls
CCFA Handoff Mode: PARTIAL (Recommended). Follow metadata.ccf_skill_controls.handoff_question_mode and ../ccf-common/references/handoff-modes.md. Use ../ccf-common/references/routing.md to keep literature search separate from idea optimization, manuscript writing, experiment design, paper review, and rebuttal.
Load ../ccf-common/references/task-modes.md before deciding exploratory, quick, or standard mode. Use exploratory mode for early direction scouting, "看看还有没有机会", "这个方向是不是被做完了", or literature search meant to feed idea optimization rather than a final novelty verdict. Use quick mode for a narrow related-work scan or a small set of candidate citations. Use standard mode for Related Work, Introduction, mature idea novelty grounding, benchmark discovery, experiment design, or any task that will feed another CCFA module.
If the user asks for recurring watch, latest-paper monitoring, competitor tracking, "recently any similar idea", arXiv/OpenReview feed scans, or lab/project tracking, route to ccf-literature-monitor by the shared handoff mode. Use this skill for deep retrieval, closest-work clustering, related-work structure, benchmark/dataset discovery, and citation candidates.
Treat user ideas, draft text, unpublished results, and private manuscripts as private material. Load ../ccf-common/references/privacy-and-evidence.md before browsing. Search with public keywords, public titles, venue names, method names, public abstracts, or user-approved query text. Do not paste private draft sentences into a search query unless the user explicitly authorizes it.
Source-quality exclusion: do not search, cite, recommend, or include policy-excluded venues, journals, URLs, or PDFs. The shared policy includes MDPI sources in this exclusion set; record exclusions only in internal screening or the search-notes file.
Core Rule
Ground novelty and positioning using high-quality, inspectable sources. Prefer influential conferences, strong journals, official proceedings pages, archival repositories, and public paper pages. Do not invent papers, citations, venues, links, acceptance status, benchmark status, or numerical results. Separate searched evidence from inference. Literature search is not a kill gate: the presence of related work should produce differentiation options, open gaps, benchmark/evidence choices, and caution labels before any "direction is covered" conclusion. Follow the user's requested output shape: short list, related-work clusters, opportunity map, BibTeX candidates, benchmark table, search folder, or handoff summary.
Mandatory Search Checklist
In standard mode, complete this checklist before final output. In quick mode, run the relevant subset and return a compact checklist status.
1. The user's topic is converted into safe public search queries. 2. Shared source-quality exclusions are applied to search domains, candidates, and final outputs. 3. Sources prioritize primary or high-confidence venues: official proceedings, arXiv/OpenReview when appropriate, ACL Anthology, CVF, PMLR, ACM, IEEE, USENIX, DBLP, Semantic Scholar, OpenAlex, Crossref, and venue or project pages. 4. Candidate papers are deduplicated by title and linked to a stable URL. 5. Each included paper has venue/year/source status, paper type, and relevance rationale. 6. Paper quality is scored on insight, completeness, and experimental numeric evidence. Pure benchmark papers skip the numeric-results score and receive a benchmark-quality note instead. 7. Paper type is one of pure benchmark, pure method, method + benchmark, survey, system/tool, theory/proof, or other. 8. Every claim about a paper is traceable to the linked source or marked as inferred. 9. For idea-stage searches, each closest-work cluster includes what is already covered, what remains under-tested, and at least one possible differentiation or rescue route. 10. A literature-search folder is written when file access is available and the user asked for a reusable report or standard workflow. 11. Optional handoff to ccf-literature-monitor, ccf-paper-writer, ccf-idea-optimizer, ccf-idea-reviewer, ccf-experiment-designer, or ccf-paper-reviewer follows CCFA handoff mode.
Workflow
1. Identify the search purpose: Related Work, Introduction support, novelty check, direction scouting, idea optimization, idea review, experiment design, benchmark/dataset discovery, or reviewer-risk diagnosis. 2. Create public queries from the user's topic. If the topic is too private or underspecified, ask only for non-sensitive keywords or infer broad keywords with lower confidence. 3. Load references/search-and-scoring.md. Search breadth depends on mode:
- Exploratory: 10-20 screened candidates, 5-10 final papers or clusters, plus opportunity gaps.
- Quick: 6-10 screened candidates, 3-6 final papers.
- Standard: 15-30 screened candidates, 8-15 final papers unless the user requests another size.
4. Search discovery indexes first, then verify candidates through stable paper pages or official proceedings when possible. Use broad web search only to find primary links; do not rely on snippets for final claims. 5. Filter by influence and fit. Prefer CCF-A/B conferences, top-field conferences, strong journals, widely used benchmarks, or recent high-signal preprints from credible groups. Exclude low-quality, predatory, inaccessible, or policy-excluded sources. For exploratory searches, include one or two "near miss" or negative-signal clusters if they reveal an open gap, failed assumption, outdated benchmark, missing user group, or neglected system constraint. 6. Classify each paper with the paper-type taxonomy and score it:
insight: how clear and non-obvious the central idea is.completeness: method/evaluation/proof/dataset/reproducibility coverage.experimental numeric evidence: strength and relevance of reported numerical evidence; markN/A benchmarkfor pure benchmark papers.
7. Write the search folder using references/report-template.md. Default folder name:
literature-search-YYYYMMDD-<topic-slug>/
papers.md
papers.csv
search-notes.md8. If the search feeds another module, provide a handoff summary:
- For writing: closest-work groups, novelty gaps, citation cautions.
- For idea optimization: stale/overcrowded directions, open gaps, timely pivots, and minimum viable research questions.
- For idea review: novelty confidence and likely prior-art risks.
- For literature monitoring: watch queries, tracked competitors, and recurring overlap signals.
- For experiment design: datasets, baselines, metrics, benchmark protocols.
- For paper review: missing related work and baseline risks.
Adaptive Output Contracts
Return the requested artifact first. If the user asks for a list of papers, output the list/table directly. If they ask for Related Work material, output clusters and positioning notes. If they ask for a folder, write the folder and summarize it. Use the following defaults for standard or quick search reports.
For standard search, return:
Search purpose:
Queries used:
Source policy:
Folder written:
Top paper table:
Excluded source notes:
Closest-work clusters:
Opportunity map:
Quality-score rationale:
Benchmark/dataset candidates:
Novelty and positioning risks:
Recommended next module:
Checklist status:For quick search, return:
Quick search scope:
Top candidates:
High-risk missing literature:
Opportunity hint:
Folder written:
Compact checklist status:Reference Files
Load only what is needed:
references/search-and-scoring.md: Use for source policy, source-quality exclusions, source tiers, paper-type taxonomy, and scoring anchors.references/report-template.md: Use when writing the literature-search folder files.
interface:
display_name: "CCF Literature Searcher"
short_description: "Search and screen literature, related work, datasets, benchmarks, and citation evidence for CCF research workflows."
default_prompt: "Use $ccf-literature-searcher for its owned CCFA workflow. Respect trigger boundaries, artifact ownership, evidence limits, and handoff rules."
Literature Search Report Template
Use this file when writing a literature-search folder.
Folder Layout
Default:
literature-search-YYYYMMDD-<topic-slug>/
papers.md
papers.csv
search-notes.mdIf the user provides a project directory, write the folder there. Otherwise use the current workspace. If file writing is unavailable, return the same sections in the final answer.
papers.md
# Literature Search: <topic>
Date: YYYY-MM-DD
Search purpose:
Target venue/family:
Source-quality policy: applied
## Summary
- Closest-work clusters:
- Opportunity map:
- Strongest baselines:
- Benchmark/dataset candidates:
- Novelty risks:
- Recommended next action:
## Paper Table
| # | Title | Year | Venue/source | Link | Type | Insight | Completeness | Numeric evidence | Overall | Notes |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | | | | | pure method / pure benchmark / method + benchmark | | | | | |
## Clusters
### Cluster 1: <name>
- Representative papers:
- What this cluster already solves:
- Remaining gap:
- Possible rescue or differentiation route:
- How it affects the user's paper:
## Opportunity Map
| Cluster | Status | Open gap | Possible direction | Evidence needed | Risk |
| --- | --- | --- | --- | --- | --- |
| | crowded but open / covered central claim / benchmark gap / mechanism gap / deployment-system gap / theory-analysis gap / negative-result opportunity | | | | |
## Benchmark And Dataset Candidates
| Name | Link | Task | Metrics | Baselines | Fit | Risks |
| --- | --- | --- | --- | --- | --- | --- |
## Citation And Positioning Cautions
- Claims that need direct citation:
- Papers that may weaken novelty:
- Papers that only support background:papers.csv
Use these columns:
title,year,venue_or_source,link,paper_type,insight_score,completeness_score,numeric_evidence_score,overall_label,relevance_note,quality_noteFor pure benchmark papers, set numeric_evidence_score to N/A benchmark.
search-notes.md
# Search Notes
## Safe Queries Used
-
## Sources Checked
-
## Excluded Sources
- Policy-excluded or low-quality sources: noted in screening notes only.
- Other exclusions:
## Unknowns
- Papers not accessible:
- Venue status not verified:
- Missing benchmark details:
## Handoff Notes
- For writing:
- For idea optimization:
- For direction scouting:
- For experiment design:
- For review:Literature Search And Scoring
Use this file for source selection, source-quality exclusions, paper classification, and quality scoring.
Source Tiers
Preferred primary sources:
- Official proceedings or publisher pages: ACL Anthology, CVF Open Access, PMLR, ACM Digital Library, IEEE Xplore, USENIX, VLDB, OpenReview, AAAI/OJS, SIAM, Springer/LIPIcs when venue-appropriate.
- Stable paper records: arXiv, DBLP, Semantic Scholar, OpenAlex, Crossref, DOI landing pages, project pages, benchmark leaderboards, and dataset pages.
- Strong journals and transactions when the field is journal-centered: JMLR, TPAMI, TOG, TKDE, TDSC, TOSEM, TSE, TODS, CACM, IEEE/ACM Transactions, Nature/Science family only when relevant.
Discovery sources are allowed for finding candidates, but final claims should point to a stable paper page or proceedings page whenever possible.
Hard Exclusions
Exclude:
- MDPI domains, journals, proceedings, PDFs, and special issues.
- Predatory or low-signal venues with no clear review standard.
- Untraceable PDFs without title, authors, venue, or stable link.
- Papers whose only accessible evidence is a search snippet.
If an excluded paper appears important because the user specifically named it, report that it is excluded by policy and ask whether they want a separate user-directed note. Do not include it in the scored candidate table.
Search Strategy
Use several query forms:
"<task>" "<method>" benchmark
"<task>" "<closest baseline>" "dataset"
"<problem phrase>" site:openaccess.thecvf.com OR site:proceedings.mlr.press
"<method phrase>" OpenReview NeurIPS ICLR ICML
"<topic>" DBLP SIGMOD SIGCOMM SOSP CCS CHI PLDI STOCPrefer a mix of:
- recent papers from the last 2-4 years,
- one or two older anchor papers,
- closest baselines,
- benchmark/dataset papers,
- survey or taxonomy papers only when they clarify a field boundary.
Opportunity Mapping
When the search is for early direction scouting, classify each cluster as one or more of:
crowded but open: many papers exist, but a measurable failure mode, setting, or assumption remains weak.covered central claim: the same problem, mechanism, and evidence path are already present; the idea needs a new angle.benchmark gap: the field lacks a dataset, protocol, metric, stress test, workload, or evaluation standard.mechanism gap: papers report performance but do not explain why the method works or fails.deployment/system gap: real constraints, cost, latency, robustness, privacy, security, or user workflow are under-tested.theory/analysis gap: empirical results exist but formal understanding, bounds, or diagnostics are missing.negative-result opportunity: common assumptions fail, and a careful diagnostic or falsification paper may be valuable.
For every covered central claim, still report the best possible rescue route:
What is covered:
Why direct novelty is weak:
Possible rescue route: new problem / new mechanism / new evidence / new venue / stop
Evidence needed to decide:Do not conclude "the direction is dead" merely because the topic is popular. A direction is likely dead only when the user's central problem, mechanism, target setting, and evidence path are all already covered and no credible reframing remains.
Paper-Type Taxonomy
Classify each included paper:
pure benchmark: primary contribution is dataset, benchmark, workload, protocol, evaluation suite, or leaderboard.pure method: primary contribution is method, algorithm, model, system design, proof, or analysis.method + benchmark: introduces a method and a new dataset/benchmark/protocol that both matter.survey: synthesizes prior work.system/tool: primary contribution is implementation, infrastructure, deployment, or usable tool.theory/proof: primary contribution is theorem, bound, proof technique, or formal model.other: use only with a short explanation.
Scoring
Use 1-5 anchors unless the user asks for another scale.
Score paper quality separately from idea viability. A high-quality close paper may reduce direct novelty while also revealing a better problem boundary, baseline, benchmark, or unresolved assumption. A low-quality paper should not be used as strong evidence that a direction is solved.
Insight:
- 5: clear, non-obvious insight that changes how the problem is understood.
- 4: strong idea with a specific mechanism or framing.
- 3: useful but incremental insight.
- 2: mostly engineering assembly or standard framing.
- 1: unclear or weak insight.
Completeness:
- 5: method/protocol/proof, evaluation, limitations, reproducibility, and positioning are all well covered.
- 4: strong coverage with minor missing details.
- 3: adequate but with visible gaps.
- 2: partial; important design or evidence details missing.
- 1: incomplete or hard to audit.
Experimental numeric evidence:
- 5: strong, fair, multi-setting numerical evidence with appropriate baselines and analysis.
- 4: solid numerical evidence with minor weaknesses.
- 3: adequate numbers but limited settings, baselines, or statistics.
- 2: weak numbers, narrow protocol, or unclear fairness.
- 1: numerical evidence absent or not meaningful for claims.
N/A benchmark: use for pure benchmark papers; instead write a benchmark-quality note.
Benchmark-quality note for pure benchmark:
Benchmark scope:
Task realism:
Metric validity:
Baseline coverage:
Adoption or reproducibility signal:
Known limitation:Overall quality label:
A: high-priority close work.B: useful supporting or baseline work.C: background only or weak fit.Risk: must inspect because it may undercut novelty.
Do not convert scores into acceptance probability.