Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
full-statck-skills avatar

Skill Official Evaluation

  • 8 installs
  • 2 repo stars
  • Updated July 29, 2026
  • full-statck-skills/utility-skills

Evaluate an Agent Skill against the official agentskills.io specification and produce a spec-compliance assessment report.

About

Checks a skill's frontmatter, directory structure, progressive disclosure, description triggering and script safety against the official Agent Skills spec. A developer uses it to audit a skill and generate an official-style evaluation report.

  • Checks frontmatter, structure, triggering and script safety
  • Produces Pass/Needs-improvement conclusions tied to the official spec

Skill Official Evaluation by the numbers

  • 8 all-time installs (skills.sh)
  • Ranked #545 of 782 Skill Development skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/full-statck-skills/utility-skills --skill skill-official-evaluation

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs8
repo stars2
Last updatedJuly 29, 2026
Repositoryfull-statck-skills/utility-skills

What it does

Evaluate an Agent Skill against the official agentskills.io specification and produce a spec-compliance assessment report.

Files

SKILL.mdMarkdownGitHub ↗

When to use this skill

ALWAYS use this skill when the user asks to:

  • Review a skill for compliance with the official Agent Skills specification
  • Check whether a skill's frontmatter, directory structure, or naming follows the rules
  • Evaluate a skill's description triggering quality against official best practices
  • Audit a skill's script safety (non-interactive, --help, secrets, structured output)
  • Generate an official-style evaluation report with Pass/Needs-improvement conclusions
  • Verify progressive disclosure is implemented correctly
  • Scan for security issues (hardcoded secrets, suspicious instructions)
  • "审查技能合规" (review skill compliance), "官方规范评估" (official spec evaluation)
  • "Skill 规范检查" (skill spec check), "技能安全审计" (skill security audit)
  • "检查 SKILL.md 格式" (check SKILL.md format), "检查技能结构" (check skill structure)
  • "生成官方评估报告" (generate official evaluation report)
  • "根据官方规范评估技能" (evaluate skill against official spec)
  • "这个 Skill 符合规范吗" (does this skill comply with the spec)

Trigger phrases include:

  • "帮我审查这个 Skill 是否符合规范" (help me review whether this skill complies with spec)
  • "检查这个技能的 SKILL.md 格式对不对" (check if this skill's SKILL.md format is correct)
  • "这个技能的 frontmatter 合规吗" (is this skill's frontmatter compliant)
  • "审计一下这个技能的脚本安全性" (audit this skill's script safety)
  • "按照 agentskills.io 规范评估这个技能" (evaluate this skill against agentskills.io spec)
  • "review this skill for official spec compliance"
  • "check if my skill follows the official specification"
  • "generate an official evaluation report for this skill"
  • "does this skill meet the agentskills.io requirements"

When NOT to use (near-miss boundaries):

  • User wants a multi-dimensional quality score with radar charts → use skill-trace-evaluation instead (TRACE model covers T/R/A/C/E, while official evaluation focuses on spec compliance)
  • User wants to learn how to design a skill (know the rules, not evaluate a specific skill) → use skill-awesome instead
  • User wants to organize skill documentation into an index → use skill-awesome instead
  • User asks for general code review (not related to Agent Skills) → this skill is scoped to Agent Skills ecosystem only

IMPORTANT: Official Evaluation vs TRACE Evaluation — Two Different Evaluation Models:

This skill and skill-trace-evaluation evaluate skills using different frameworks:

  • Official Evaluation (this skill): Based on the official Agent Skills specification from agentskills.io. Checks structural compliance, naming rules, frontmatter correctness, and script safety. Answers "Does this skill follow the rules?"
  • TRACE Evaluation (different skill): Based on the SkillHub TRACE quality model. Scores across Trust, Reliability, Adaptability, Convention, and Effectiveness. Produces radar charts and per-dimension scores. Answers "How good is this skill?"

When both skills could apply:

  • If the user says "evaluate this skill" or "review this skill" without specifying a framework, ask: "I can evaluate this skill using either the official specification (agentskills.io compliance) or the TRACE quality model (five-dimension scoring with radar charts). Which would you prefer?"
  • If the user explicitly mentions "official spec", "agentskills.io", "compliance", "format check" → use this skill
  • If the user explicitly mentions "TRACE", "quality score", "radar chart", "five dimensions" → use skill-trace-evaluation

How to use this skill

CRITICAL: This skill evaluates a target skill against the official Agent Skills specification. The evaluation conclusion is explicitly based on the official specification and best practices published at agentskills.io. Do not invent requirements not present in the official sources.

To evaluate a skill:

Step 1: Identify the target skill

  • Preferred input: Path to the target skill directory
  • The target must contain a SKILL.md file. If not found, report "SKILL.md not found" and stop.
  • Also inspect optional directories: scripts/, references/, assets/.
  • If the user provides a .skill or .zip archive, only unpack when explicitly asked; otherwise evaluate from provided excerpts.

Step 2: Apply the official rubric

Use references/official-rubric.md as the evaluation checklist. The rubric has five inspection dimensions:

Dimension 1: Spec Compliance (MUST pass)
CheckWhat to verify
SKILL.md existsThe skill root must contain a SKILL.md file
Frontmatter presentYAML frontmatter delimited by --- at the top of SKILL.md
`name` fieldMust match the parent directory name. Lowercase letters, digits, and hyphens only. 1-64 characters. No leading/trailing hyphens, no consecutive --.
`description` fieldNon-empty, max 1024 characters. Must describe both what the skill does AND when to use it. Should not be overly broad.
Optional fields formatIf license, compatibility, metadata, or allowed-tools are present, verify their formatting is valid.
Directory structureOptional directories must follow conventions: scripts/ for executable code, references/ for on-demand docs, assets/ for templates and resources.
Dimension 2: Progressive Disclosure Quality (SHOULD meet)
CheckWhat to verify
SKILL.md concisenessBody stays concise and actionable. Ideally under 500 lines / 5000 tokens.
Details in references/Long explanations, reference tables, and supplementary content moved to references/.
Clear reference triggersWhen a reference file is mentioned, the skill tells the agent WHEN to load it ("Read references/api-errors.md if the API returns a non-200 status code").
No deep reference chainsReferences should be one level deep from SKILL.md. Avoid references that point to other references.
Dimension 3: Description Triggering Quality (SHOULD meet)
CheckWhat to verify
User-intent languageDescription uses words users would naturally say, not implementation jargon.
Not implementation-onlyDescription goes beyond "Processes X files" — it tells the agent when the user needs X processed.
Trigger boundariesDescription contains both "should trigger" and "should not trigger" signals where applicable.
Dimension 4: Script Readiness (CONDITIONAL — only if scripts/ exists)

Use references/script-safety-checklist.md to verify:

CheckRequirement
Non-interactiveNo TTY prompts. All inputs via flags, env vars, or stdin.
`--help` availablePrints usage, options, and examples.
Clear error messagesErrors say what failed, what was expected, and what to try next.
No secretsNo hardcoded tokens, keys, or passwords.
Safe defaultsDestructive operations require --force or --confirm.
Structured output (recommended)--format json option. Data to stdout, diagnostics to stderr.
Idempotency (recommended)Repeated runs do not corrupt state.
Dimension 5: Security Hygiene (MUST pass)
CheckWhat to verify
No secrets in filesScan SKILL.md, scripts, and other text files for hardcoded tokens, API keys, passwords.
No suspicious instructionsThe skill must not instruct the agent to download from untrusted sources, exfiltrate data, or execute obfuscated code.
Risky operations guidanceIf the skill involves destructive operations, it must instruct the agent to get explicit user confirmation.

Step 3: Run the evaluator script (recommended)

The bundled script automates data collection and formatting:

python3 scripts/official_evaluate.py --help

# Generate a Markdown evaluation report
python3 scripts/official_evaluate.py --skill-dir <path> --format md

# Generate machine-readable JSON
python3 scripts/official_evaluate.py --skill-dir <path> --format json

# Write to a file
python3 scripts/official_evaluate.py --skill-dir <path> --format md --output report.md

The script performs automated checks for:

  • SKILL.md presence and frontmatter parsing
  • Name format validation (regex) and directory match
  • Description length validation
  • License field check
  • Secret pattern scanning (AWS keys, API keys, token patterns)
  • Non-interactive pattern detection in scripts

After running the script, you MUST supplement the automated results with qualitative assessment for:

  • Progressive disclosure quality (is the body concise? are references well-triggered?)
  • Description triggering quality (does it use user-intent language? are trigger boundaries clear?)
  • Security hygiene beyond regex patterns (are there suspicious instructions?)

Step 4: Produce the evaluation report

The report MUST include these sections, in order:

Report structure
# Official Skill Evaluation Report

Target: `<path-to-skill-directory>`

## Conclusion

- Overall conclusion: **Pass** / **Needs improvement** / **Fail**
- Top issues:
  1. ...
  2. ...
  3. ...

## Compliance Checklist

| Item | Result | Evidence | Suggestion |
|------|--------|----------|------------|
| SKILL.md frontmatter present | Pass/Fail | ... | ... |
| name matches directory | Pass/Fail | ... | ... |
| name format valid | Pass/Fail | ... | ... |
| description present & valid | Pass/Fail | ... | ... |
| license field | Pass/Needs improvement | ... | ... |
| Optional directories organized | Pass | ... | ... |
| Progressive disclosure | Pass/Needs improvement | ... | ... |
| Description trigger quality | Pass/Needs improvement | ... | ... |
| Script safety (if applicable) | Pass/Fail/N/A | ... | ... |
| Security & secrets scan | Pass/Fail | ... | ... |

## Risks & Limitations

- ...

## Improvement Suggestions (prioritized)

1. ...
2. ...
3. ...
Conclusion levels
LevelCriteria
PassAll MUST items pass. SHOULD items are reasonably met. No security findings.
Needs improvementAll MUST items pass, but SHOULD items have significant gaps. No security findings.
FailOne or more MUST items fail, OR security findings detected.
Evidence rules
  • Every Pass/Fail MUST include specific evidence: a file path, a field value, a line number, or a scan result.
  • Do NOT use subjective language like "seems good" or "looks fine". Cite artifacts.
  • If a check is N/A (e.g., no scripts/ directory), state "N/A — no scripts/ directory" as evidence.

Output format

After producing the evaluation report:

1. State the overall conclusion clearly: "Pass", "Needs improvement", or "Fail" 2. List the top 3 most important findings 3. Show the compliance checklist table with evidence 4. Provide prioritized, actionable improvement suggestions 5. Save the report if the user requests a file; otherwise display inline

Rules

1. Do not invent "official requirements" not present in the official sources above. Every finding must be traceable to the official specification or best practices. 2. Do not include secrets or reproduce sensitive content in the report. If secrets are found, note their location without reproducing the secret value. 3. Treat the rubric as the ground truth. If references/official-rubric.md says a check is "Should", do not report it as a hard failure. 4. The evaluation conclusion is explicitly based on the official specification. The report should state this clearly in the opening paragraph.

Keywords

English keywords: official-evaluation, spec-compliance, skill-review, skill-audit, frontmatter-check, naming-validation, description-quality, script-safety, security-scan, progressive-disclosure, official-rubric, agentskills-spec, skill-assessment, compliance-report, skill-inspection, format-check, structure-review

Chinese keywords (中文关键词): 审查技能合规, 官方规范评估, Skill 规范检查, 技能安全审计, 检查 SKILL.md 格式, 检查技能结构, 生成官方评估报告, 根据官方规范评估技能, 技能合规检查, 技能评估报告, frontmatter 检查, 技能命名检查, 技能描述检查, 脚本安全检查, 渐进式披露检查, 官方规范审查

能力边界

✅ 适用场景

  • 当你需要使用此技能对应的技术栈时
  • 当项目需要遵循最佳实践时
  • 当需要快速上手或深入理解核心概念时

⚠️ 需要注意

  • 复杂业务逻辑需要结合具体场景调整
  • 性能优化需要根据实际数据量评估

❌ 不适用场景

  • 不相关的技术栈或框架
  • 需要完全自定义的特殊场景

常见陷阱 (Gotchas)

1. 版本兼容性:注意框架版本与依赖库的兼容性,不同版本 API 可能有差异 2. 配置文件格式:配置文件格式错误是最常见的问题,建议使用编辑器的语法检查 3. 环境变量:确保所有必要的环境变量已正确设置,敏感信息不要硬编码 4. 依赖冲突:多版本共存时注意依赖冲突,使用 lock 文件锁定版本 5. 性能陷阱:大数据量场景下注意性能优化,避免 N+1 查询等常见问题

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.