
Metacognitive Self Mod
- 88 installs
- 325 repo stars
- Updated August 2, 2026
- athola/claude-night-market
Capture execution traces and feed metacognitive self-modification so your agent improves playbooks from real tool sequences and outcomes—not only aggregate scores.
About
metacognitive-self-mod is a journey-wide agent skill family for solo builders who treat coding agents as systems that should learn from their own runs. The ingested readme centers on Execution Trace Capture: each skill invocation becomes a structured trace (trace_id, skill name, timestamps, outcome, steps with tools like Read/Write and targets) wrapped in the ADR-0011 session-capture envelope so friction signals and traces share one consumer. That gives metacognitive-self-mod concrete sequences for backward-style attribution—aligned with Microsoft Trace metaphors, trajectory-informed memory literature, and ACE-style playbook evolution through generation, reflection, and curation. Use it whenever you want continuous improvement across Idea research spikes, Build implementation, Ship debugging, or Operate incident response, as long as you accept recording overhead. It is meta infrastructure, not a feature shipper: pair with trace-capture (estimated_tokens ~200 in child frontmatter) and avoid when you forbid logging tool paths or lack storage for session artifacts. Intermediate complexity; you need a repo or agent session model willing to persist JSON traces.
- Parent abstract for trace-capture: records tool sequences, decision points, and outcome attribution per skill invocation
- Trace JSON uses shared session-capture envelope (schema_version session-capture/1) for one parser across friction and tr
- Informed by Microsoft Trace-style propagation, trajectory-informed memory (arXiv 2603.10600), and ACE generation/reflect
- Child trace-capture lists metacognitive-self-mod dependency and decision-only capture_mode option
Metacognitive Self Mod by the numbers
- 88 all-time installs (skills.sh)
- Ranked #4,935 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Security screen: LOW risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/athola/claude-night-market --skill metacognitive-self-modAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 88 |
|---|---|
| repo stars | ★ 325 |
| Security audit | 3 / 3 scanners passed |
| Last updated | August 2, 2026 |
| Repository | athola/claude-night-market ↗ |
What it does
Capture execution traces and feed metacognitive self-modification so your agent improves playbooks from real tool sequences and outcomes—not only aggregate scores.
Files
Metacognitive Self-Modification
Analyze the effectiveness of past skill improvements and refine the improvement process itself. This is the core innovation from the Hyperagents paper: not just improving skills, but improving HOW skills are improved.
Context Triggers (auto-invocation)
This skill should be invoked automatically when:
1. Regression detected: The homeostatic monitor finds a skill's evaluation window ended in pending_rollback_review status. The improvement made things worse, and we need to understand why.
2. Low effectiveness rate: When ImprovementMemory.get_effective_strategies() vs get_failed_strategies() shows effectiveness below 50%, the improvement process itself needs refinement.
3. Degradation despite improvements: When PerformanceTracker.get_improvement_trend() returns negative for a skill that was recently improved.
4. Periodic check: After every 10 improvement cycles (tracked via outcome count in ImprovementMemory).
Hook integration
The homeostatic monitor emits "improvement_triggered": true when a skill crosses the flag threshold. At that point, before dispatching the skill-improver, check if metacognitive analysis is warranted:
from abstract.improvement_memory import ImprovementMemory
from pathlib import Path
memory = ImprovementMemory(
Path.home() / ".claude/skills/improvement_memory.json"
)
# Check if metacognitive analysis is warranted
effective = memory.get_effective_strategies()
failed = memory.get_failed_strategies()
total = len(effective) + len(failed)
needs_metacognition = False
# Trigger 1: Low effectiveness rate
if total >= 5 and len(effective) / total < 0.5:
needs_metacognition = True
# Trigger 2: Periodic check (every 10 outcomes)
if total > 0 and total % 10 == 0:
needs_metacognition = True
# Trigger 3: Recent regression
if failed and failed[-1].get("outcome_type") == "failure":
needs_metacognition = True
if needs_metacognition:
# Run metacognitive analysis before next improvement
pass # Skill(abstract:metacognitive-self-mod)When To Use (Manual)
- After a batch of skill improvements to assess what
worked
- When improvement outcomes show regressions
- Periodically (monthly) to refine improvement strategy
- When the skill-improver agent seems ineffective
When NOT To Use
- Routine skill improvements (use skill-improver directly)
- First-time skill creation (use skill-authoring)
Workflow
Step 1: Load improvement data
Read improvement memory and performance tracker data:
# Check for improvement memory
MEMORY_FILE=~/.claude/skills/improvement_memory.json
TRACKER_FILE=~/.claude/skills/performance_history.json
if [ ! -f "$MEMORY_FILE" ]; then
echo "No improvement memory found."
echo "Run skill-improver first to generate improvement data."
exit 0
fiLoad the JSON files using Python:
from abstract.improvement_memory import ImprovementMemory
from abstract.performance_tracker import PerformanceTracker
from pathlib import Path
memory = ImprovementMemory(Path.home() / ".claude/skills/improvement_memory.json")
tracker = PerformanceTracker(Path.home() / ".claude/skills/performance_history.json")Step 2: Classify improvement outcomes
For each improvement outcome in memory, classify:
- Effective:
after_score - before_score >= 0.1 - Neutral:
-0.1 < improvement < 0.1 - Regression:
after_score < before_score
effective = memory.get_effective_strategies()
failed = memory.get_failed_strategies()
# Calculate effectiveness rate
total = len(effective) + len(failed)
if total > 0:
effectiveness_rate = len(effective) / totalStep 3: Extract meta-patterns
Analyze WHAT types of improvements succeed vs fail:
Success patterns to look for:
- Adding error handling (reduces failure rate)
- Adding examples (improves user ratings)
- Adding quiet/verbose modes (reduces friction)
- Simplifying workflow steps (reduces duration)
Failure patterns to look for:
- Over-engineering (adding too many options)
- Breaking existing workflows (regression)
- Adding complexity without validation
- Token budget overflow from verbose additions
For each pattern found, record as a causal hypothesis:
memory.record_insight(
skill_ref="_meta", # Special ref for meta-insights
category="causal_hypothesis",
insight="Error handling improvements have 85% success rate",
evidence=["skill-A v1.1.0: +0.3", "skill-B v2.1.0: +0.15"]
)Step 4: Analyze improvement trends
Use PerformanceTracker to identify:
- Skills with sustained improvement (positive trend)
- Skills with degradation despite improvement attempts
- Domains where improvements are most effective
for skill_ref in tracker.get_all_skill_refs():
trend = tracker.get_improvement_trend(skill_ref)
if trend is not None:
if trend > 0.05:
# Sustained improvement - what's working?
pass
elif trend < -0.05:
# Degrading despite improvements - investigate
passStep 5: Generate strategy recommendations
Based on the meta-analysis, generate recommendations for the skill-improver:
1. Priority formula adjustments: If certain issue types have higher improvement success rates, weight them higher.
2. Approach selection: If "add error handling" has 85% success vs "restructure workflow" at 30%, bias toward error handling.
3. Threshold adjustments: If improvements below priority 3.0 consistently fail, raise the minimum threshold.
4. Avoidance rules: Document anti-patterns to avoid in future improvements.
Step 6: Store meta-insights
Record all findings back into ImprovementMemory under the special _meta skill ref:
# Record strategy recommendation
memory.record_insight(
skill_ref="_meta",
category="strategy_success",
insight="Recommendation: Prioritize error handling and examples over restructuring",
evidence=[f"Success rate: error_handling={eh_rate:.0%}, restructure={rs_rate:.0%}"]
)Step 7: Update skill-improver strategy
If significant meta-insights are found, propose concrete modifications to the skill-improver agent:
- Update priority weights in the priority formula
- Add avoidance rules for known anti-patterns
- Adjust thresholds based on empirical data
- Add new improvement patterns that proved effective
Important: Propose changes, do not auto-apply. The user must approve modifications to the improvement process.
Output
Metacognitive Self-Modification Report
Improvement Data:
Total outcomes analyzed: 15
Effective improvements: 11 (73%)
Regressions: 2 (13%)
Neutral: 2 (13%)
Success Patterns:
1. Error handling additions: 5/6 success (83%)
2. Example additions: 3/3 success (100%)
3. Quiet mode additions: 2/2 success (100%)
Failure Patterns:
1. Workflow restructuring: 1/3 success (33%)
2. Token-heavy additions: 0/1 success (0%)
Performance Trends:
Improving: 8 skills (positive trend)
Stable: 4 skills (no trend)
Degrading: 1 skill (negative trend despite attempts)
Recommendations:
1. Weight error handling improvements 2x in priority
2. Avoid workflow restructuring below priority 8.0
3. Cap additions at 200 tokens to prevent budget overflow
4. Focus next improvement cycle on degrading skill X
Meta-insights stored: 5 new entries in improvement memoryRelated
abstract:skill-improver- The agent this skill analyzes
and proposes modifications for
abstract:skills-eval- Evaluation framework whose
criteria could be refined by meta-insights
abstract:aggregate-logs- Data source for improvement
metrics
Execution Trace Capture
Record execution traces so that metacognitive-self-mod can analyze concrete decision sequences, not just aggregate metrics. Inspired by Microsoft Trace (AutoDiff-like backward propagation through execution traces), Trajectory-Informed Memory Generation (arXiv 2603.10600, decision attribution from trajectories), and the ACE framework (evolving playbooks via generation, reflection, and curation).
Trace Structure
Each trace captures a single skill invocation from start to finish. The trace body sits inside the shared session-capture envelope (ADR-0011) so friction signals and traces can be consumed through one parser:
{
"schema_version": "session-capture/1",
"session_id": "2026-04-14-abc12345",
"timestamp": "2026-04-14T10:30:00Z",
"source": "trace-capture",
"payload": {
"trace_id": "session-{date}-{hash}",
"skill": "attune:project-execution",
"started": "2026-04-14T10:30:00Z",
"completed": "2026-04-14T10:32:15Z",
"outcome": "success",
"capture_mode": "decision-only",
"steps": [
{
"tool": "Read",
"target": "src/main.py",
"purpose": "understand entry point",
"result": "success",
"tokens_used": 1200,
"decision_point": false
},
{
"tool": "Edit",
"target": "src/main.py:45",
"purpose": "add error handling",
"result": "success",
"tokens_used": 800,
"decision_point": true,
"alternatives_considered": [
"try/except",
"result type",
"assertion"
],
"rationale": "try/except matches existing patterns"
}
],
"attribution": {
"success_factors": [
"followed existing patterns",
"tested incrementally"
],
"failure_factors": [],
"key_decisions": [
"chose try/except over result type at step 4"
]
}
}
}Legacy traces written before envelope adoption are read as `session-capture/0 (entire file treated as the payload). See docs/adr/0011-session-capture-envelope.md` for the contract and migration path.
Capture Modes
Not every invocation needs a full trace. Three modes control the recording fidelity.
| Mode | Records | When to use |
|---|---|---|
minimal | Outcome and duration only | High-trust T3 skills |
decision-only | Decision points, outcome, duration | Default for all skills |
full | Every tool call, token counts, all fields | Skills with <85% success rate |
Mode selection logic:
- Default:
decision-only(captures rationale without
flooding storage).
- Enable
fullwith--trace=fullor automatically for
any skill whose rolling success rate falls below 85%.
- Use
minimalfor T3 skills that consistently succeed
and need only aggregate trend data.
What to Capture
- Tool calls (full mode only): tool name, target,
purpose, result, approximate tokens consumed.
- Decision points (all modes except minimal):
alternatives considered (2-5), rationale for the chosen option, whether later revised.
- Trace completion (all modes): overall outcome
(success/failure/partial), wall-clock duration, total tokens consumed.
Attribution Analysis
After a trace completes, run backward attribution to identify which decisions drove the outcome. This follows the Microsoft Trace principle of propagating feedback backward through the execution path.
For successful traces:
1. Identify decisions that aligned with known success patterns (from improvement_memory.json). 2. Flag novel successful patterns as candidate hypotheses. 3. Record success_factors in the trace's attribution block.
For failed traces:
1. Walk backward from the failure point. 2. Identify the earliest decision that diverged from known-good patterns. 3. Record failure_factors and the specific decision. 4. Generate a causal hypothesis for ImprovementMemory:
memory.record_insight(
skill_ref=trace["skill"],
category="causal_hypothesis",
insight="Failure correlated with choosing X over Y",
evidence=[f"trace:{trace['trace_id']}"]
)Cross-trace pattern detection:
When 5 or more traces exist for a skill, scan for recurring decision-outcome correlations:
- Decisions that appear in >70% of successful traces
become "recommended patterns."
- Decisions that appear in >50% of failed traces become
"anti-patterns."
Storage
| Location | Contents | Retention |
|---|---|---|
~/.claude/skills/traces/ | Raw JSON trace files | Rolling 30-day window |
improvement_memory.json | Aggregate patterns and hypotheses | Persistent |
Budget: Maximum 100 trace files, FIFO eviction. Traces linked to active causal hypotheses are protected until the hypothesis is resolved. File naming: {trace_id}.json (one file per trace).
Integration Points
- metacognitive-self-mod (parent): consumes traces
during periodic analysis (Step 3) to inspect specific decisions behind success or failure.
- skill-improver: queries traces for a target skill
before proposing changes. Targets recurring failure points directly.
- friction-detector: cross-references friction signals
with trace data to pinpoint where a workflow broke.
Lightweight by Default
The default decision-only mode records only branching points (typically 3-8 entries per trace vs 20-50 for full mode). Additional storage hygiene:
- Prune full-mode traces older than 7 days down to
decision-only.
- Cap
alternatives_consideredat 5 entries. - Omit
tokens_usedin minimal mode.
Related skills
FAQ
Is Metacognitive Self Mod safe to install?
skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.