
Code Copilot Team
- 6 repo stars
- Updated August 5, 2026
- gosha70/code-copilot-team
Code Copilot Team hooks: file protection, auto-format, type verification, context re-injection, git safety, and desktop notifications.
About
code-copilot-team is a Claude Code skill in the AI & Agent Building category. Code Copilot Team hooks: file protection, auto-format, type verification, context re-injection, git safety, and desktop notifications.
- code-copilot-team
- AI & Agent Building
- AI-coding skill
Code Copilot Team by the numbers
- Data as of Aug 5, 2026 (Skillselion catalog sync)
/plugin marketplace add gosha70/code-copilot-team/plugin install code-copilot-team@code-copilot-teamAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| repo stars | ★ 6 |
|---|---|
| Last updated | August 5, 2026 |
| Repository | gosha70/code-copilot-team ↗ |
What it does
Code Copilot Team hooks: file protection, auto-format, type verification, context re-injection, git safety, and desktop notifications.
README.md
Code Copilot Team
Reusable, opinionated configuration for AI-assisted coding with multi-agent team delegation. Ships with templates for ML/AI, Enterprise Java, and Web projects.
Built for Claude Code as the reference implementation, with portable conventions for Cursor, GitHub Copilot, Windsurf, Aider, and local LLMs.
📖 Deep dive: Stop Fighting AI Agents and Build a Reusable Multi-Agent Dev Environment — the full story behind this project, lessons learned from 13+ real build sessions, and why every rule exists.
Why This Exists
Every rule in this repo is failure-driven — it exists because we hit the specific failure it prevents, often more than once. After analyzing 13 sessions of a real project build, we identified six recurring patterns: dependency breaks, agents ignoring conventions, context window exhaustion, schema drift during parallel builds, agents not asking clarifying questions, and commit granularity issues. This setup prevents all of them.
Framework Compliance
Evaluated against the two leading AI coding agent frameworks (February 2026):
OpenAI Harness Engineering — 5.0 / 5.0

Claude Code Best Practice — 10.0 / 10.0

Sources: OpenAI Harness Engineering · Claude Code Best Practice
Further Reading
- Spec-Driven Development vs Code Copilot Team — Side-by-side comparison with GitHub's Spec Kit. TL;DR: SDD defines what to build; Code Copilot Team defines how to behave while building it. They're complementary, not competing.
Spec-Driven Development (SDD)
Code Copilot Team includes a built-in Spec-Driven Development layer that prevents "vibe coding" — the tendency of AI agents to start writing code before requirements are clear. SDD ensures every feature goes through a structured specification process, with the rigor scaled to match the risk.
How It Works
Every task is classified into one of three spec modes based on risk:
| spec_mode | When | What's Required |
|---|---|---|
| full | Security, schema changes, integration, features touching >2 files | plan.md + spec.md + tasks.md |
| lightweight | Features touching 1–2 files, non-critical changes | plan.md + spec.md |
| none | Bug fixes (non-security), docs, trivial changes | plan.md only |
The Plan agent writes plan.md with a YAML frontmatter block that declares spec_mode, feature_id, risk_category, and justification. The Build agent reads this frontmatter and gates itself accordingly — it won't proceed on a full task without a complete spec.md, and it won't proceed on any task that has unresolved [NEEDS CLARIFICATION] markers.
The Four Artifacts
SDD uses exactly four artifact types (no checklists, no extra process):
| Artifact | Purpose | When Created |
|---|---|---|
plan.md |
Implementation plan with frontmatter gating | Always (all modes) |
spec.md |
Requirements, user scenarios, constraints | full and lightweight only |
tasks.md |
Task breakdown with story and priority markers | full only |
lessons-learned.md |
Cross-project learnings for future sessions | End of project |
Templates for all four live in shared/templates/sdd/ and are available across all adapters.
Three-Layer Gating
SDD enforcement operates at three levels:
- Agent-level — The Build agent reads
plan.mdfrontmatter and conditionally requiresspec.mdand resolves[NEEDS CLARIFICATION]markers before proceeding. - CI validation —
scripts/validate-spec.shruns on every PR touchingspecs/. It validates frontmatter fields, checks for required files per spec_mode, and enforces justification forspec_mode: none. - Hooks — Existing hooks remain untouched; SDD gating is additive, not intrusive.
Spec Artifacts Location
All SDD artifacts live in the versioned specs/ directory, organized by feature:
specs/
└── <feature-id>/
├── plan.md ← Always present
├── spec.md ← full / lightweight
├── tasks.md ← full only
├── lessons-learned.md ← End of project
└── collaboration/ ← Peer review artifacts (dual mode)
├── plan-consult.md ← Peer review of plan phase
└── build-review.md ← Peer review of build phase
Risk Classification
The spec-workflow.md rule defines risk categories that map directly to spec_mode:
| Risk Category | spec_mode | Examples |
|---|---|---|
security |
full | Auth changes, secrets handling, permission logic |
schema |
full | Database migrations, API contract changes |
integration |
full | Third-party integrations, cross-service changes |
feature |
full or lightweight | New features (full if >2 files, lightweight if 1–2) |
bug |
none | Non-security bug fixes |
docs |
none | Documentation-only changes |
Adapter Support
SDD rules propagate through the same shared/ → generate.sh → adapters/ pipeline as all other rules. Claude Code gets enforced gating via agent manifests. Other adapters receive advisory content appropriate to their capabilities:
| Adapter | SDD Support Level |
|---|---|
| Claude Code | Enforced — agents gate on frontmatter |
| GitHub Copilot | Full instructions (advisory) |
| Cursor, Windsurf | Always-on rules only (advisory) |
| Aider | Conventions only (advisory) |
Getting Started with SDD
- Start a Plan session — describe your feature to the Plan agent.
- The Plan agent classifies risk, sets
spec_mode, and writes the appropriate artifacts tospecs/<feature-id>/. - Switch to Build — the Build agent reads the frontmatter and gates itself.
- CI validates on PR —
validate-spec.shcatches any missing artifacts or incomplete specs.
No additional setup required — SDD is active by default after installation.
Shape-Up (Product Bets)
SDD answers "how do we know we built the right thing?" — Shape-Up answers "what do we build next, and how big should it be?" The two are complementary: a pitch describes the bet, SDD's plan/spec/tasks describe the implementation underneath one or more scopes of that pitch.
Code Copilot Team ships a local-first Shape-Up implementation: pitches and hill charts as plain files under specs/pitches/<id>/, four agents (pitch-shaper, scope-executor, cycle-retro, cooldown-report), five slash commands (/shape, /bet, /cycle-start, /hill, /cooldown), and validate-pitch.sh enforcing frontmatter (appetite ∈ {2w, 4w, 6w}, bet_status lifecycle, cycle/circuit-breaker conditional rules) on every PR.
Use Shape-Up for product-shaped work — greenfield, ambiguous problem space, multiple possible solutions, time-boxed bets. Use SDD alone for feature-shaped work where the requirement is clear.
📖 Full guide: docs/shape-up-workflow.md — methodology, frontmatter schema, lifecycle diagram, agent reference, install surface, and a worked example.
Peer Review (Multi-Copilot)
Code Copilot Team supports dual-copilot peer review — a second AI provider automatically reviews your work at phase completion. This catches blind spots that a single provider misses, using the same structured collaboration protocol regardless of which providers are involved.
Prerequisites
Install Code Copilot Team — run
setup.sh --claude-code(see Quick Start). This installs all peer review components:peer-review-runner.shandproviders-health.shto~/.local/bin/peer-review-on-stop.shhook to~/.claude/hooks//phase-completecommand to~/.claude/commands/- Provider profile seed to
~/.code-copilot-team/providers.toml
Install the peer provider CLI — the peer provider must be available on your machine. For example, to use OpenAI Codex as a peer reviewer, install the Codex CLI first.
Verify provider availability:
providers-health.sh
Setup — New Projects
# 1. Init project from template
claude-code init ml-rag ~/projects/my-app
# 2. Start session with peer review
claude-code --peer-review codex ~/projects/my-app
Setup — Existing Projects
No project-level changes required. Peer review is driven entirely by session flags and global hooks:
# Just add --peer-review to your usual launch command
claude-code --peer-review codex ~/projects/existing-app
# Or use the default peer from your provider profile
cd ~/projects/existing-app && claude-code --peer-review
How It Works
Start a session with peer review enabled:
claude-code --peer-review codex ~/projects/my-app # explicit peer provider claude-code --peer-review ~/projects/my-app # default peer from profile claude-code --peer-review-off ~/projects/my-app # disable for this session claude-code --peer-review-scope code ~/projects/my-app # scope: code|design|bothWork normally through the Plan → Build phases. Claude detects
CCT_PEER_REVIEW_ENABLED=truein the environment and setscollaboration_mode: dualin the SDD plan.Run
/review-submitafter completing work — the Build agent runs this to start the review loop. The runner spawns a reviewer LLM in a read-only sandbox, captures structured findings, and returns a verdict. On FAIL, the agent addresses findings and resubmits. On PASS, proceed to/phase-complete.Run
/phase-completewhen review passes — validates thatloop-summary.jsonexists, runs the post-phase checklist, and presents the commit for approval.Review the artifact — the collaboration artifact (
build-review.mdorplan-consult.md) is written tospecs/<feature-id>/collaboration/with structured findings and a verdict.
Provider Profile
Peer providers are configured in ~/.code-copilot-team/providers.toml (seeded by setup):
[defaults]
peer_for.claude = "codex"
peer_for.codex = "claude"
[providers.codex]
type = "cli"
command = "codex --quiet --prompt-file {review_request}"
timeout_sec = 300
healthcheck = "codex --version"
[providers.ollama]
type = "ollama"
command = "ollama run {model} < {review_request}"
model = "llama3"
timeout_sec = 600
healthcheck = "ollama list"
Every provider currently requires a command template with {review_request} and {model} placeholders. The type field (cli, openai-compatible, ollama, custom) declares the provider topology and will enable type-aware dispatch and dedicated adapter scripts in a future update. See shared/templates/provider-profile-template.toml for all type-specific fields and commented-out examples.
Safety Model
- Fail-closed — enforced at two levels: (1)
/phase-completerequiresloop-summary.jsonwith PASS or bypass before proceeding, (2) the stop hook blocks session end if review was started but not completed (exit 2). If review was never started, the hook warns but does not block. - Circuit breakers — max rounds (default 5), wall-clock timeout (15 min), stale findings, provider unavailability. All escalate to human via
/review-decide. - Read-only sandbox — reviewer runs in a snapshot copy; real working tree is never modified by the reviewer.
- Escape hatch — set
CCT_PEER_BYPASS=trueto skip validation. CI rejects bypass artifacts. - Identity tracking — collaboration artifacts include
peer_profile(provider name) andrunner_fingerprint(SHA-256 of provider config) for auditability.
Collaboration Modes
| Mode | When | What Happens |
|---|---|---|
| single (default) | No --peer-review flag |
Standard single-provider workflow, no peer review |
| dual | --peer-review [provider] |
Peer reviews at /phase-complete, artifacts written to specs/ |
LLM Wiki Maintainer
code-copilot-team ships a Karpathy-pattern LLM Wiki maintainer that
turns knowledge/raw/ into a curated, cited, agent-readable markdown
layer under knowledge/wiki/. Five operations, one CLI:
./scripts/wiki ingest <source> # multi-page write plan against existing wiki state
./scripts/wiki promote <proposal-dir> # atomic apply (only writer to the canonical wiki content tree, excluding .audit/)
./scripts/wiki query "<question>" # index-first synthesis with citations
./scripts/wiki query --file-back "..." # round-trip the answer back into a patch-set
./scripts/wiki lint # structural lint (frontmatter, links, slugs)
./scripts/wiki lint --health [--strict] # knowledge-health (contradictions, stale claims, weak orphans, missing cross-links)
./scripts/wiki audit-flush # commit pending ingest-log lines (reject-only durability)
./scripts/wiki audit-flush --dry-run # report count + blob SHA without committing
Human approval is always gating, and the source-control boundary
is explicit: the wiki is source-controlled, the proposal workspace
is not. wiki ingest writes draft proposals to a local-only
doc_internal/proposals/ directory (gitignored — proposals are
working drafts, not canonical state). wiki promote is the only
operation that writes to the canonical knowledge/wiki/ content tree;
wiki ingest has one additional tracked write: appending to the
append-only knowledge/wiki/.audit/ingest-log.md audit ledger. The
audit trail under knowledge/wiki/.audit/ records every wiki ingest
decision (timestamp, source SHA, backend, disposition, reason) in
ingest-log.md, and every accepted proposal's original LLM draft in
knowledge/wiki/.audit/proposals/<date>-<slug>/ (applied atomically
by wiki promote). wiki audit-flush (shipped in
gosha70/code-copilot-team#37)
closes the reject-only durability gap: run it after a reject-only session
to commit any pending audit lines in a focused audit: flush N pending ingest-log line(s) commit. Promotion
history is traceable via git on knowledge/wiki/ plus
knowledge/wiki/log.md.
The CLI auto-detects an installed copilot backend in the order
claude → codex → cursor. Override with --backend <name> or
WIKI_INGEST_BACKEND=<name>. Use --backend test for the
deterministic stub backend (no LLM call; this is what CI uses).
For the v1 single-source flow, the legacy invocation
./scripts/wiki-ingest <source> is preserved as a backwards-compat
alias.
Operator docs
- Full operator workflow:
knowledge/README.md§5e. - Workflow page:
knowledge/wiki/workflows/run-wiki-ingest.md. - Design rationale:
specs/wiki-ingest-pipeline/spec.md. - Schema:
knowledge/wiki/schema/— page types, ingest rules, citation rules, lint rules, curator persona.
Benchmark Harness
code-copilot-team ships a benchmark-agnostic harness for evaluating AI
copilots and LLMs on real coding tasks under reproducible isolation —
so you can answer "which copilot/model is actually better on this kind
of work?" with a controlled run record instead of a vibe.
It does not author benchmarks; it runs established public ones (Aider Polyglot, SWE-bench Verified, BigCodeBench) and custom CCT fixtures through one adapter contract. There are two entry points — a terse daily-driver wrapper and the underlying harness CLI:
# Daily driver — safe by default (no-arg run is a free stub smoke + env detection)
./scripts/bench # prove the plumbing, no LLM call, no spend
./scripts/bench sonnet ollama:qwen2.5-coder:7b # compare two models on a coding task
./scripts/bench --preset local-vs-cloud --runs 5 # curated comparison preset
./scripts/bench --list-presets # discovery: available presets
./scripts/bench --list-providers # discovery: detected backends/providers
# Underlying harness
./scripts/benchmark list # adapters + backends + judges
./scripts/benchmark run --benchmark aider-polyglot \
--backend claude-code --model sonnet --runs 3 # one (backend, model) run
./scripts/benchmark compare --config my-compare.json # multi-LLM comparison
./scripts/benchmark report --run-dir runs/<ts>/ --html --csv # rich report (HTML + SVG charts + CSV)
What it measures. Deterministic scoring is the primary signal —
build/test/lint pass, required files present, elapsed time, token usage
— with a calibrated winner-declaration rule (Δ > 2σ AND ≥ threshold)
that refuses to call a winner on noise. A calibrated LLM judge
(issue #34) adds a secondary quality signal (idiomaticity, error
handling, test thoughtfulness, security hygiene), but only after it's
proven to correlate with human reviewers (Spearman ρ ≥ threshold per
dimension); it never overrides the deterministic verdict, and a run
that fails its tests can never win on judge-only criteria. No
dollar-cost estimates are ever reported.
Backends (the agent driving the task): claude-code, codex,
aider, plus a deterministic stub for CI. Local models (vLLM,
Ollama, LM Studio) are reached as providers through the gateway env
vars — ./scripts/bench sonnet vllm:<model>@<endpoint> probes the
endpoint and spawns an ephemeral Anthropic↔OpenAI proxy when needed.
Operator docs
- Full harness guide, CLI reference, adapter/backend/judge contracts:
benchmarks/README.md. - 60-second quickstart:
benchmarks/README.md§ 60-second quickstart. - Design rationale:
specs/benchmark-harness/spec.mdand the per-feature spec bundles underspecs/.
What You Get

- Layered rules — 4 global rules (
~/.claude/rules/) auto-load every session; 15 on-demand skills (~/.claude/skills/*/SKILL.md) loaded by phase agents when needed. - Phase agents (
~/.claude/agents/) — 4 phase agents (research, plan, build, review) plus 5 utility agents (code-simplifier, doc-writer, phase-recap, security-review, verify-app). - Hooks (
~/.claude/hooks/) — 11 lifecycle scripts: test verification, type checking, auto-format, file protection, git safety guards, context re-injection, peer review trigger, desktop notifications, plus 3 self-guarding MemKernel hooks (session recall, pre-compact checkpoint, post-compact recovery) that activate only when MemKernel is installed. - 11 project templates — pre-configured
CLAUDE.mdfiles with stack-specific conventions, slash commands, and agent team roles for each project archetype. - Four-phase workflow — Research → Plan → Build → Review. Plus Ralph Loop for single-agent autonomous iteration.

- Adaptive launcher (
claude-code) — usescmuxon macOS,tmuxelsewhere, with git context display,--peer-reviewflags, andsyncfor keeping projects aligned with template updates.
Quick Start
# 1. Clone
git clone https://github.com/gosha70/code-copilot-team.git
cd code-copilot-team
# 2. Install for your tool(s)
./scripts/setup.sh --claude-code # Claude Code → ~/.claude/
./scripts/setup.sh --codex # OpenAI Codex → ~/.codex/
./scripts/setup.sh --cursor ~/my-project # Cursor → project/.cursor/
./scripts/setup.sh --github-copilot ~/my-project # GH Copilot → project/.github/
./scripts/setup.sh --windsurf ~/my-project # Windsurf → project/.windsurf/
./scripts/setup.sh --aider ~/my-project # Aider → project/CONVENTIONS.md
# Or install everything at once
./scripts/setup.sh --all ~/my-project
# Re-sync after pulling repo updates
git pull && ./scripts/setup.sh --sync --claude-code
The legacy ./claude_code/claude-setup.sh path still works — it delegates to the adapter.
After git pull, run --sync to regenerate configs and re-install.
Alternative: Install as a Claude Code Plugin
For Claude Code users who prefer the plugin system over setup.sh:
# Add the CCT marketplace (one-time)
/plugin marketplace add gosha70/code-copilot-team
# Install the hooks plugin
/plugin install code-copilot-team@code-copilot-team
This installs the same hooks (file protection, auto-format, type verification, context re-injection, git safety, notifications) as setup.sh, but managed through Claude Code's plugin system. Update installed plugins with /plugin marketplace update. The plugin does not include peer-review or memkernel hooks — those are CCT-pipeline-specific and remain in the setup.sh path.
Both install paths coexist. Use setup.sh for the full install (skills, agents, templates, hooks, peer review) or the plugin for hooks only.
Recommended: Install LSP Plugins (Claude Code)
For continuous type-error feedback during edits, install the appropriate code-intelligence plugin. Each requires its language-server binary on $PATH:
# Install the language server first, then the plugin:
pip install pyright && /plugin install pyright-lsp@claude-plugins-official # Python
npm i -g typescript-language-server typescript && /plugin install typescript-lsp@claude-plugins-official # TypeScript
go install golang.org/x/tools/gopls@latest && /plugin install gopls-lsp@claude-plugins-official # Go
These provide native LSP diagnostics and are preferred over the bundled verify-after-edit.sh hook. The hook remains as a fallback for languages without an LSP plugin. See the official plugin catalog for all available languages.
Start a New Project
# Initialize from a template
claude-code init ml-rag ~/projects/my-rag-app
# Start a Claude session in the project
claude-code ~/projects/my-rag-app
Start in an Existing Project
# Just point the launcher at it — global rules load automatically
claude-code ~/projects/existing-api
Sync a Project to Latest Template
After pulling repo updates, sync your project's commands and .claude/ files against the latest template:
# 1. Update global config + templates from repo
git pull && ./scripts/setup.sh --sync --claude-code
# 2. Preview what would change (safe — no files modified)
claude-code sync ~/projects/my-rag-app --dry-run
# 3. Apply the sync
claude-code sync ~/projects/my-rag-app
Sync updates commands and .claude/ contents (e.g. remediation.json) but never overwrites your CLAUDE.md — it shows a diff for manual review instead. Projects initialized with claude-code init have a .claude/template.json that tracks the template; older projects are matched by their CLAUDE.md heading.
Available Templates

| Template | Stack | Agent Team |
|---|---|---|
ml-rag |
Python · FAISS/Chroma · Neo4j/NetworkX | Team Lead, RAG Engineer, KG Engineer, Data Analyst, QA |
ml-langchain |
Python · LangChain/LangGraph/LangSmith | Team Lead, Agent Developer, Integration Engineer, QA & Eval |
ml-app |
Python · FastAPI · LiteLLM · Next.js/React | Team Lead, Backend Dev, Frontend Dev, ML/AI Engineer, QA |
ml-utils |
Python · MCP SDK · Chroma/Qdrant · tree-sitter | Team Lead, MCP Engineer, Retrieval Engineer, Storage Engineer, QA |
ml-n8n |
Python · n8n · REST/webhooks | Team Lead, Workflow Designer, Python Developer, QA & DevOps |
java-enterprise |
Spring Boot · Kafka · GraphQL · React | Team Lead, Backend Dev, Frontend Dev, Data & Messaging, QA, DevOps |
web-static |
Astro/Next.js/Hugo · Tailwind | Team Lead, Frontend Dev, Content & SEO, QA |
web-dynamic |
Next.js/Remix · Node/Python · PostgreSQL | Team Lead, Frontend Dev, Backend Dev, QA, DevOps |
java-tooling |
Java 21 · Gradle · JSR 269 · JavaPoet · Spring AI MCP | Team Lead, APT Engineer, MCP Specialist, Plugin Dev, QA |
gradle-plugin |
Kotlin · Gradle 8 · Plugin<Project> · TestKit matrix · Plugin Portal |
Team Lead, Plugin Eng, Functional Test Eng, Build & Release |
domain-pack |
Versioned content (TBX/JSON-LD/CSV) · Maven Central + PyPI dual publish | Team Lead, Content Curator, JVM Wrapper Eng, Python Wrapper Eng, Release & CI |
Bundled CI Workflows
Each template ships a .github/workflows/ file so CI is wired up the moment the consumer adds their toolchain manifest.
| Stack | Workflow file | What it runs |
|---|---|---|
ml-app, ml-rag, ml-langchain, ml-n8n, ml-utils |
python.yml |
ruff · mypy · pytest --cov · matrix: 3.10, 3.11, 3.12 |
java-enterprise, java-tooling |
gradle.yml |
./gradlew build check test · matrix: JDK 17, 21 · optional publish-staging on tags |
web-static, web-dynamic |
node.yml |
lint · typecheck · test · matrix: Node 20, 22 · auto-detects npm/yarn/pnpm |
domain-pack |
pack-content.yml + pack-publish.yml |
manifest + content schema validation on PR · coordinated Maven Central + PyPI publish on tag |
gradle-plugin |
gradle-plugin.yml |
unit tests · TestKit functional matrix (Gradle 8.5/8.10/current) · sample-consumer smoke · Plugin Portal publish on tag |
Auto-skip on empty project. Each workflow's job is gated on a toolchain marker (pyproject.toml / setup.py / setup.cfg for Python, package.json for Node, gradlew for Gradle — the wrapper, since build steps invoke ./gradlew). A freshly bootstrapped project with no marker yet gets a green skip rather than a red failure. The job activates as soon as the consumer adds the marker file. Gradle projects that have build scripts but no wrapper get a notice nudging them to run gradle wrapper.
Matrix override via workflow_dispatch. Every workflow accepts a manual trigger with an optional version input (e.g. python-version: "3.12" or node-version: "20"). Leave it blank to run the full matrix; set it to a specific version to run that one only.
Dual-branch trigger. All workflows fire on push to master or main — whichever convention a project uses.
Bootstrap path. claude-code init <type> copies .github/workflows/ into the new project automatically. claude-code sync keeps the workflow file up to date alongside remediation.json and commands.
How Configuration Layers Work
~/.claude/CLAUDE.md ← Global agent manifest (base)
~/.claude/rules/*.md ← Global rules (always loaded, 4 files)
├── coding-standards.md SOLID, quality gates, prohibited patterns
├── copilot-conventions.md Cross-tool portable conventions
├── safety.md Destructive action guards, secrets policy
└── copyright-headers.md Copyright header rules for generated source files
~/.claude/skills/*/SKILL.md ← On-demand skills (SKILL.md format, 15 skills)
├── agent-team-protocol/ Three-phase workflow, delegation rules
├── clarification-protocol/ Ask before implementing ambiguous requirements
├── environment-setup/ Environment and config verification
├── infra-verification/ Infrastructure artifact verification ("build it, run it")
├── integration-testing/ Test integration points early
├── memkernel-memory/ MemKernel persistent memory protocol (self-guarding)
├── opus-4-7-features/ Opus 4.7 optimization (xhigh effort, auto mode, caching)
├── phase-workflow/ Phase transition rules and boundaries
├── provider-collaboration-protocol/ Peer review protocol and collaboration rules
├── ralph-loop/ Single-agent autonomous iteration loop
├── review-loop/ Peer review loop with findings and resolutions
├── spec-workflow/ SDD spec gating and artifact management
├── stack-constraints/ Stack version and compatibility guards
├── team-lead-efficiency/ Limit agents, poll frequency, no re-work
└── token-efficiency/ Diff-over-rewrite, context economy
~/.claude/agents/*.md ← Phase + utility agents (9 files)
├── research.md Research phase agent
├── plan.md Plan phase agent
├── build.md Build phase agent
├── review.md Review phase agent
├── code-simplifier.md Simplify recently changed code
├── doc-writer.md Generate and update documentation
├── phase-recap.md Summarize completed phase
├── security-review.md Scan for security vulnerabilities
└── verify-app.md End-to-end project verification
~/.claude/hooks/*.sh ← Deterministic lifecycle hooks (always active, 11 files)
├── verify-on-stop.sh Run test suite when Claude finishes responding
├── verify-after-edit.sh Run type checker after source file edits
├── auto-format.sh Auto-format edited files
├── protect-files.sh Prevent edits to protected files
├── protect-git.sh Guard destructive git commands (push --force, reset --hard)
├── peer-review-on-stop.sh Trigger peer review on phase completion
├── reinject-context.sh Re-inject session context on prompt submit
├── notify.sh Desktop notifications (macOS + Linux)
├── memkernel-recall.sh Recall MemKernel context on session start (self-guarding)
├── memkernel-pre-compact.sh Save checkpoint before compaction (self-guarding)
└── memkernel-post-compact.sh Recover context after compaction (self-guarding)
~/.claude/settings.json ← Hooks wiring and global settings
./CLAUDE.md ← Project-level (overrides global)
./.claude/commands/*.md ← Project slash commands
./CLAUDE.local.md ← Personal overrides (gitignored)
Project-level rules override global rules. More specific always wins.
Four-Phase Workflow
| Phase | Model | Effort | Delegation | What Happens |
|---|---|---|---|---|
| Research | Opus (highest) | High | None | Explore codebase, summarize findings, identify constraints |
| Plan | Opus (highest) | High | None | Design approach, get user approval |
| Build | Sonnet (fast) | Medium | Yes | Team Lead delegates to specialist sub-agents |
| Build (loop) | Sonnet (fast) | Medium | None | Ralph Loop: single agent iterates through stories autonomously |
| Review | Opus (highest) | High | None | Holistic review, run tests, verify consistency |
Each phase has a dedicated agent (~/.claude/agents/) that loads the relevant rules from the rules library. Planning and research must stay in one mind — sub-agents only see fragments and can't reason about the whole system. Delegation only happens during Build. For smaller features, Ralph Loop provides a single-agent alternative: read PRD → implement next failing story → test → commit → repeat.
Supported Tools
All tools share the same rules from shared/skills/. Each adapter formats them for the target tool.
| Tool | Adapter Output | Install Location |
|---|---|---|
| Claude Code | agents, hooks, commands, settings | ~/.claude/ (global) |
| OpenAI Codex | AGENTS.md + 5 skills |
~/.codex/ (global) |
| Cursor | .mdc files with frontmatter |
project/.cursor/rules/ |
| GitHub Copilot | copilot-instructions.md + per-rule instructions |
project/.github/ |
| Windsurf | rules.md |
project/.windsurf/rules/ |
| Aider | CONVENTIONS.md |
project/ |
Repo Structure
code-copilot-team/
├── shared/ ← Single source of truth
│ ├── skills/ 21 skills (SKILL.md format, open Agent Skills spec)
│ ├── docs/ 8 tool-agnostic reference docs
│ ├── templates/ 11 stacks × PROJECT.md + commands/
│ ├── templates/sdd/ 5 SDD templates (spec, plan, tasks, lessons-learned, collaboration)
│ └── templates/provider-profile-template.toml Peer provider profile seed
├── specs/ ← SDD artifacts per feature (versioned)
│ └── <feature-id>/ plan.md, spec.md, tasks.md, lessons-learned.md
├── knowledge/ ← Project knowledge layer (curated wiki + raw notes)
│ ├── README.md Wiki usage guide (read this first)
│ ├── raw/ Unedited candidate material
│ └── wiki/ Curated, cited, agent-maintainable pages
├── benchmarks/ ← Benchmark harness (start at benchmarks/README.md)
│ ├── README.md Harness guide + CLI reference
│ ├── adapters/ aider-polyglot, swe-bench-verified, bigcodebench, stub, …
│ ├── presets/ Curated compare-configs for ./scripts/bench
│ └── calibration/ Judge rubrics + calibration corpora/labels
├── adapters/
│ ├── claude-code/ agents, hooks, commands, settings, setup.sh
│ ├── codex/ AGENTS.md, config.toml, 5 skills, setup.sh
│ ├── cursor/ .cursor/rules/*.mdc, setup.sh
│ ├── github-copilot/ .github/copilot-instructions.md, instructions/, setup.sh
│ ├── windsurf/ .windsurf/rules/rules.md, setup.sh
│ └── aider/ CONVENTIONS.md, setup.sh
├── scripts/
│ ├── generate.sh Builds adapter configs from shared/
│ ├── bench Terse benchmark comparison driver (wraps benchmark)
│ ├── benchmark Benchmark harness CLI (run, compare, judge, calibrate, report)
│ ├── wiki LLM Wiki maintainer CLI (ingest, promote, query, lint, audit-flush)
│ ├── validate-spec.sh SDD spec validator (CI + local)
│ ├── pre-pr-check.sh Pre-PR close-keyword audit gate
│ ├── peer-review-runner.sh Peer review execution engine
│ ├── providers-health.sh Peer provider availability diagnostics
│ └── setup.sh Unified install entry point
├── tests/
│ ├── test-hooks.sh 186 hook tests
│ ├── test-generate.sh 282 generation + adapter tests
│ ├── test-shared-structure.sh 800 structure + content tests
│ ├── test-sync.sh 69 sync + init metadata tests
│ ├── test-peer-review.sh 54 peer-review runner tests
│ └── test-review-loop.sh 31 review loop integration tests
├── claude_code/ Backward-compat wrapper → adapters/claude-code/
├── .github/workflows/sync-check.yml CI: adapter drift + full gate verification
├── README.md
├── CONTRIBUTING.md
└── LICENSE
Rule content is written once in shared/ and adapted per tool via scripts/generate.sh. Generated adapter configs are committed to the repo. CI verifies they never drift.
Documentation
Claude Code specific:
- Setup Cookbook — deep-dive into every configuration option
- Config Guide — templates, agent teams, output styles, and workflow reference
- Hooks Guide — hook installation, customization, and supported stacks
- Sub-Agents Guide — sub-agent configuration and usage
- Agent Traces — locating, reading, and archiving agent transcripts
- Debugging Strategies — /doctor, background tasks, Playwright MCP, trace debugging
- Permissions Guide — per-stack Allow/Deny wildcard patterns for /permissions
- Recommended MCP Servers — Context7, PostgreSQL, Filesystem, and Playwright MCP setup
Shared (all tools):
- Alignment Maintenance Checklist — recurring governance checks to keep framework alignment healthy
- Common Pitfalls — cross-cutting issues and solutions
- Delegation Best Practices — when and how to delegate to agents
- Ralph Loop Guide — Ralph Loop usage and configuration
- Session Management — session commands cheat sheet
- Code Reviewer Assistant Guide — peer review setup, commands, and safety model
- Error Reporting Template — standardized format for bug reports
- Phase Recap Template — end-of-phase handoff checklist
Contributing
See CONTRIBUTING.md. PRs welcome for new templates, rule improvements, and ports to other tools.
Community Standards
- Code of Conduct
- Code Owners
- Security Policy
- Issue Templates
- Pull Request Template
- GitHub Hardening Playbook
Alignment Maintenance
Use the recurring checklist in shared/docs/alignment-maintenance.md to keep this repo aligned as rules, skills, and templates evolve.