Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
actionbook avatar

M10 Performance

  • 1.5k installs
  • 1.3k repo stars
  • Updated May 24, 2026
  • actionbook/rust-skills

m10-performance is an agent skill for Rust performance optimization using profiling, benchmarks, and measured design choices.

About

The m10-performance skill guides Rust performance optimization with a measure-first mindset using flamegraphs, perf, and criterion benchmarks to find real bottlenecks. It maps goals to design choices such as pre-allocation with with_capacity, contiguous Vec layouts, rayon parallelism, zero-copy Cow references, and smallvec for inline data. The skill asks whether optimization is worth added complexity and prioritizes algorithmic wins over cache or allocation tweaks. Thinking prompts cover measurement, priority ordering from algorithm through cache effects, and tradeoffs between memory, CPU, latency, and throughput. It traces decisions up to domain constraints and down to concrete Rust patterns. Triggers include performance, benchmark, profiling, flamegraph, criterion, SIMD, and allocation keywords. Use when developers profile Rust code and choose optimization strategies grounded in measurement.

  • Measure-first workflow with flamegraph, perf, and criterion before optimizing.
  • Decision table maps goals to pre-allocation, rayon, Cow, and smallvec patterns.
  • Prioritizes algorithmic gains over allocation and cache micro-optimizations.
  • Prompts for complexity versus speed and memory versus CPU tradeoffs.
  • User-invocable false; triggered by performance and benchmark keywords.

M10 Performance by the numbers

  • 1,494 all-time installs (skills.sh)
  • +52 installs in the week ending Jul 28, 2026 (Skillselion tracking)
  • Ranked #7 of 129 Rust skills by installs in the Skillselion catalog
  • Security screen: LOW risk (skills.sh audit)
  • Data as of Jul 28, 2026 (Skillselion catalog sync)
At a glance

m10-performance capabilities & compatibility

Capabilities
profile first bottleneck identification · allocation and cache optimization patterns · rayon parallelism guidance · zero copy cow reference patterns · complexity versus speed tradeoff prompts
Use cases
refactoring · testing · debugging
From the docs

What m10-performance says it does

What's the bottleneck, and is optimization worth it?
SKILL.md
Have you measured? (Don't guess)
SKILL.md
npx skills add https://github.com/actionbook/rust-skills --skill m10-performance

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1.5k
repo stars1.3k
Security audit3 / 3 scanners passed
Last updatedMay 24, 2026
Repositoryactionbook/rust-skills

How do I find and fix Rust performance bottlenecks without guessing or premature micro-optimization?

Optimize Rust performance using profiling, benchmarks, and allocation or cache-aware design choices before micro-optimizing.

Who is it for?

Developers optimizing Rust services or libraries who need benchmark-driven performance decisions.

Skip if: Skip for non-Rust languages or feature development without performance measurement needs.

When should I use this skill?

User asks to optimize Rust performance, run criterion benchmarks, or analyze flamegraphs.

What you get

Profiled hotspots with chosen optimization patterns such as pre-allocation, parallelism, or zero-copy references.

  • Benchmark results
  • Bottleneck analysis
  • Targeted optimization plan

Files

SKILL.mdMarkdownGitHub ↗

Performance Optimization

Layer 2: Design Choices

Core Question

What's the bottleneck, and is optimization worth it?

Before optimizing:

  • Have you measured? (Don't guess)
  • What's the acceptable performance?
  • Will optimization add complexity?

---

Performance Decision → Implementation

GoalDesign ChoiceImplementation
Reduce allocationsPre-allocate, reusewith_capacity, object pools
Improve cacheContiguous dataVec, SmallVec
ParallelizeData parallelismrayon, threads
Avoid copiesZero-copyReferences, Cow<T>
Reduce indirectionInline datasmallvec, arrays

---

Thinking Prompt

Before optimizing:

1. Have you measured?

  • Profile first → flamegraph, perf
  • Benchmark → criterion, cargo bench
  • Identify actual hotspots

2. What's the priority?

  • Algorithm (10x-1000x improvement)
  • Data structure (2x-10x)
  • Allocation (2x-5x)
  • Cache (1.5x-3x)

3. What's the trade-off?

  • Complexity vs speed
  • Memory vs CPU
  • Latency vs throughput

---

Trace Up ↑

To domain constraints (Layer 3):

"How fast does this need to be?"
    ↑ Ask: What's the performance SLA?
    ↑ Check: domain-* (latency requirements)
    ↑ Check: Business requirements (acceptable response time)
QuestionTrace ToAsk
Latency requirementsdomain-*What's acceptable response time?
Throughput needsdomain-*How many requests per second?
Memory constraintsdomain-*What's the memory budget?

---

Trace Down ↓

To implementation (Layer 1):

"Need to reduce allocations"
    ↓ m01-ownership: Use references, avoid clone
    ↓ m02-resource: Pre-allocate with_capacity

"Need to parallelize"
    ↓ m07-concurrency: Choose rayon or threads
    ↓ m07-concurrency: Consider async for I/O-bound

"Need cache efficiency"
    ↓ Data layout: Prefer Vec over HashMap when possible
    ↓ Access patterns: Sequential over random access

---

Quick Reference

ToolPurpose
cargo benchMicro-benchmarks
criterionStatistical benchmarks
perf / flamegraphCPU profiling
heaptrackAllocation tracking
valgrind / cachegrindCache analysis

Optimization Priority

1. Algorithm choice     (10x - 1000x)
2. Data structure       (2x - 10x)
3. Allocation reduction (2x - 5x)
4. Cache optimization   (1.5x - 3x)
5. SIMD/Parallelism     (2x - 8x)

Common Techniques

TechniqueWhenHow
Pre-allocationKnown sizeVec::with_capacity(n)
Avoid cloningHot pathsUse references or Cow<T>
Batch operationsMany small opsCollect then process
SmallVecUsually smallsmallvec::SmallVec<[T; N]>
Inline buffersFixed-size dataArrays over Vec

---

Common Mistakes

MistakeWhy WrongBetter
Optimize without profilingWrong targetProfile first
Benchmark in debug modeMeaninglessAlways --release
Use LinkedListCache unfriendlyVec or VecDeque
Hidden .clone()Unnecessary allocsUse references
Premature optimizationWasted effortMake it work first

---

Anti-Patterns

Anti-PatternWhy BadBetter
Clone to avoid lifetimesPerformance costProper ownership
Box everythingIndirection costStack when possible
HashMap for small setsOverheadVec with linear search
String concat in loopO(n^2)String::with_capacity or format!

---

Related Skills

WhenSee
Reducing clonesm01-ownership
Concurrency optionsm07-concurrency
Smart pointer choicem02-resource
Domain requirementsdomain-*

Related skills

Forks & variants (1)

M10 Performance has 1 known copy in the catalog totaling 911 installs. They canonicalize to this original listing.

How it compares

Pick this over general Rust coding skills when the task is profiling-driven performance decisions rather than language syntax or API design.

FAQ

What does m10-performance recommend first?

Measure with profiling and benchmarks to identify actual bottlenecks before choosing optimization patterns.

When should I use m10-performance?

When optimizing Rust code for speed, allocations, cache behavior, or parallel execution with criterion or flamegraphs.

Is m10-performance safe to install?

Review the Security Audits panel on this page before installing in production.

Rustbackendtesting

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.