Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
pproenca avatar

Codebase Comprehension Algorithms

  • 81 installs
  • 191 repo stars
  • Updated July 24, 2026
  • pproenca/dot-skills

codebase-comprehension-algorithms is a Claude Code skill in the AI & Agent Building category.

Key points

  • codebase-comprehension-algorithms
  • AI & Agent Building
  • AI-coding skill

Codebase Comprehension Algorithms by the numbers

  • 81 all-time installs (skills.sh)
  • +6 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #5,179 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/pproenca/dot-skills --skill codebase-comprehension-algorithms

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs81
repo stars191
Last updatedJuly 24, 2026
Repositorypproenca/dot-skills

How do I helps with ai & agent building tasks during ai-assisted development?

Helps with ai & agent building tasks during AI-assisted development.

Who is it for?

Best when you're working on ai & agent building and need structured help with codebase-comprehension-algorithms.

Skip if: Teams with no ai & agent building needs, or anyone wanting a generic chat assistant without this specific workflow.

When should I use this skill?

When you need to helps with ai & agent building tasks during ai-assisted development, or when codebase-comprehension-algorithms is a claude code skill in the ai & agent building category.

What you get

Structured output aligned to codebase-comprehension-algorithms: codebase-comprehension-algorithms; AI & Agent Building; AI-coding skill.

Files

SKILL.mdMarkdownGitHub ↗

Community Codebase Comprehension And Domain Mapping Algorithms Best Practices

A practitioner-oriented reference of the algorithms that work for mapping a codebase into understandable feature/business domains. Most of these techniques live in the Software Architecture Recovery and Mining Software Repositories literatures and are invisible to working engineers — yet they're the right tools for the job a coding agent is asked to do every day: "what does this codebase do, and where?"

The 47 rules are organized by execution-lifecycle impact: a wrong decision early in the pipeline (which graph to build, which identifiers to keep) propagates through everything downstream. The three CRITICAL categories (graph-, clust-, valid-) are the ones a wrong call cannot be recovered from later. Read them first.

Scope: proven algorithms with peer-reviewed citations or canonical books — Newman Networks, Leskovec-Rajaraman-Ullman Mining of Massive Datasets, Ganter-Wille Formal Concept Analysis, plus 40+ ICSE / FSE / TSE / PNAS / JMLR papers. No tutorial sites, no Stack Overflow, no marketing posts. Deliberately deferred to a future version: GNN/CodeBERT/code2vec (not "proven over decades" yet) and refactoring-recipe stuff (covered by sibling skills like react-refactor and typescript-refactor).

When to Apply

Use these rules when:

  • Onboarding an agent into an unfamiliar codebase: "explain what this codebase does, by domain"
  • Producing an architecture map: "what are the main subsystems and how do they connect?"
  • Locating a feature: "which files implement payments / authentication / search?"
  • Reviewing a refactor: "did this change respect the architectural boundaries?"
  • Detecting architectural debt: "what files have surprising coupling?"
  • Validating an existing decomposition: "does the README's architecture match the code?"
  • Picking algorithms for any of the above — the user wants something that's proven, not vibes

Rule Categories By Priority

#CategoryPrefixImpactWhat it does
1Graph Construction & Edge Weightinggraph-CRITICALWhich graph to build; omnipresent filter; cycle handling; multilayer
2Community Detection & Clusteringclust-CRITICALLeiden, Infomap, SBM, MCL, Walktrap, spectral, HDBSCAN
3Validation & Quality Metricsvalid-CRITICALMoJoFM, ARI/NMI, resolution limit, consensus, co-change prediction, ablation
4Identifier & Lexical Preprocessinglex-HIGHSamurai splitting, abbreviation expansion, TF-IDF/BM25, stemming, V-O parsing
5Software-Specific Architecture Recoveryarch-HIGHBunch + MQ, ACDC, Limbo, Reflexion, DSM
6Topic Modelling on Source Codetopic-HIGHLDA, LSI/SVD, NMF, HDP, coherence-based K selection
7Evolutionary Coupling & Co-Change Miningevol-HIGHLift / confidence / support, large-commit filter, temporal decay, logical coupling
8Information-Theoretic Methodsinfo-MEDIUM-HIGHNormalized Compression Distance, Mutual Information, MDL, code naturalness
9Centrality, Hierarchy & Labellingrank-MEDIUMPageRank, HITS, betweenness, TextRank/YAKE labels

Quick Reference

1. Graph Construction & Edge Weighting (CRITICAL)

  • `graph-filter-omnipresent-utilities-before-clustering` — Drop the loggers and base classes BEFORE clustering (20-40 MoJoFM points)
  • `graph-pick-edge-type-by-question-asked` — Call, import, co-change, bipartite — the question determines the graph
  • `graph-collapse-sccs-before-clustering` — Tarjan SCC condensation makes cycles explicit and stabilises every algorithm
  • `graph-weight-edges-by-information-content` — IDF / PMI / Jaccard on edges suppresses noise (2-5× MoJoFM)
  • `graph-bipartite-file-term-for-joint-structure` — When DI / dynamic dispatch hides the call graph
  • `graph-combine-signals-in-multilayer-graphs` — Mucha 2010 multilayer modularity over normalised α-weighted layers

2. Community Detection & Clustering (CRITICAL)

  • `clust-leiden-not-louvain` — Louvain produces disconnected communities on 5-25% of nodes (Traag 2019)
  • `clust-infomap-mdl-on-random-walks` — MDL on random walks; the right tool for flow-meaningful graphs
  • `clust-stochastic-block-model` — Bayesian, hierarchical, learns K from data; handles non-assortative structure
  • `clust-mcl-markov-clustering` — Flow simulation; dominant in bioinformatics; robust to noise
  • `clust-walktrap-short-random-walks` — Random-walk distance + hierarchical agglomerative
  • `clust-spectral-laplacian-fiedler` — Optimal k-way normalised cut via Laplacian eigenvectors
  • `clust-hdbscan-density-based` — When clustering on file embeddings, not graphs

3. Validation & Quality Metrics (CRITICAL)

  • `valid-mojofm-as-software-clustering-distance` — The SAR gold-standard distance metric (Wen-Tzerpos 2004)
  • `valid-adjusted-rand-index-and-nmi` — Chance-corrected cross-algorithm comparison
  • `valid-be-aware-of-resolution-limit` — Modularity can't see clusters smaller than √(2m) (Fortunato-Barthélemy PNAS 2007)
  • `valid-consensus-clustering-for-stability` — A single-run answer is unreliable; consensus across runs is the right answer
  • `valid-cochange-prediction-as-ground-truth-proxy` — Temporal held-out co-change replaces missing ground truth
  • `valid-ablate-each-input-signal` — Leave-one-out; reveals which input actually drives the result

4. Identifier & Lexical Preprocessing (HIGH)

  • `lex-split-identifiers-with-samurai` — 87% precision vs 60% for naive camelCase (Enslen MSR 2009)
  • `lex-build-programming-language-stop-words` — Three-layer: keywords + generic + IDF-driven
  • `lex-expand-abbreviations-with-context` — usr → user, ctx → context (Lawrie GenTest 2011)
  • `lex-tf-idf-and-bm25-on-identifiers` — Raw counts are dominated by common terms; TF-IDF / BM25 fix it
  • `lex-stem-versus-subword-tokenization` — Porter stemmer for clustering, BPE for embeddings
  • `lex-extract-verb-object-pattern-from-method-names` — getUserById → (verb=get, object=user); compound concept signal

5. Software-Specific Architecture Recovery (HIGH)

  • `arch-bunch-with-mq-fitness` — MQ fitness function + search; better than Q-maximization on code
  • `arch-acdc-subgraph-patterns` — Subsystem and skeleton patterns; matches architect intuition
  • `arch-limbo-information-bottleneck` — Tishby's IB applied to software (Andritsos-Tzerpos 2005)
  • `arch-reflexion-model` — Compare hypothesized vs actual; the underused gem from Murphy-Notkin 1995
  • `arch-dsm-partitioning` — Design Structure Matrix; 60-year-old engineering technique

6. Topic Modelling on Source Code (HIGH)

  • `topic-lda-on-source-code` — Probabilistic per-file topic distributions over identifier+comment text
  • `topic-lsi-svd-on-term-document` — Deterministic SVD-based semantic embeddings (Maletic-Marcus 2001)
  • `topic-nmf-non-negative-factorization` — Parts-based additive topics, fully reproducible
  • `topic-hdp-for-nonparametric-topic-count` — Hierarchical Dirichlet Process — learns K from data
  • `topic-pick-topic-count-by-coherence-not-perplexity` — Perplexity is anti-correlated with human topic quality

7. Evolutionary Coupling & Co-Change Mining (HIGH)

  • `evol-mine-cochange-with-lift-and-confidence` — Lift > 2 is the cutoff; raw co-change count is noise
  • `evol-filter-large-commits` — A 200-file commit produces 20K spurious pair-counts; filter aggressively
  • `evol-temporal-decay-on-edge-weights` — Exponential decay with 6-month half-life
  • `evol-logical-coupling-as-architectural-signal` — 30-50% of strongest coupling is invisible to static analysis (Gall 1998)

8. Information-Theoretic Methods (MEDIUM-HIGH)

  • `info-normalized-compression-distance` — Cluster without feature engineering; gzip-based universal similarity
  • `info-mutual-information-as-coupling` — Catches non-linear / conditional coupling that lift misses
  • `info-mdl-for-model-selection` — Principled K selection; Occam's razor as a code length
  • `info-naturalness-of-code-as-quality-signal` — Hindle 2012 — code is 30-50% more predictable than English; bugs spike entropy

9. Centrality, Hierarchy & Labelling (MEDIUM)

  • `rank-pagerank-for-module-importance` — Architectural spine via PageRank on the reversed dependency graph
  • `rank-hits-hubs-and-authorities` — Orchestrators vs implementations (Kleinberg 1999)
  • `rank-betweenness-centrality-for-bottlenecks` — Bridges between domains; god-class detection
  • `rank-textrank-for-cluster-labels` — Multi-word keyphrases as cluster labels (Mihalcea-Tarau 2004, YAKE 2020)

How to Use

Start with the question the agent is trying to answer:

  • "What are the main domains in this codebase?"graph- (pick a graph) → clust- (Leiden / Infomap / SBM) → topic- (label them) → valid- (sanity-check stability and ablate)
  • "Which files implement feature X?"topic-lda-on-source-code for theme location; rank-pagerank-for-module-importance with X's files as seed for personalized PageRank
  • "Where is the architectural spine?"rank-pagerank-for-module-importance + rank-hits-hubs-and-authorities on the dependency graph
  • "Does the README's architecture match the code?"arch-reflexion-model is purpose-built for this
  • *"What's the real coupling here (beyond static dependencies)?"* → evol-logical-coupling-as-architectural-signal and evol-mine-cochange-with-lift-and-confidence
  • "How do I cluster without designing features?"info-normalized-compression-distance
  • "How big are the clusters supposed to be?"valid-be-aware-of-resolution-limit and topic-hdp-for-nonparametric-topic-count
  • "How do I know my decomposition is right?" → the entire valid- category; multi-proxy evaluation is mandatory

The skill's worldview: build the right graph first (and filter omnipresent files), pick an algorithm matching the graph and the question, use a code-specific preprocessing pipeline (Samurai + stop-words + stemming + TF-IDF) where lexical signals matter, and always validate — MoJoFM if you have expert ground truth, consensus + co-change prediction + ablation if you don't.

Code examples are in Python because the reference implementations (networkx, igraph, leidenalg, scikit-learn, gensim, graph-tool, hdbscan) all live there. The reasoning generalises to any language.

Reference Files

FileDescription
references/_sections.mdCategory definitions and ordering
assets/templates/_template.mdTemplate for new rules
metadata.jsonVersion and reference information
AGENTS.mdAuto-built TOC navigation

Related Skills

  • computer-science-algorithms — Algorithm-and-data-structure reference (this skill cross-references it for MinHash/LSH, Aho-Corasick, etc.)
  • complexity-optimizer — Static analysis for hot paths the rules here identify
  • design-to-react-algorithms — Companion skill for design-to-code structural recovery

Related skills

FAQ

What does codebase-comprehension-algorithms do?

codebase-comprehension-algorithms is a Claude Code skill in the AI & Agent Building category.

When should I use codebase-comprehension-algorithms?

When you need to helps with ai & agent building tasks during ai-assisted development, or when codebase-comprehension-algorithms is a claude code skill in the ai & agent building category.

What are the main capabilities?

codebase-comprehension-algorithms; AI & Agent Building; AI-coding skill.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.