Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
mohitmishra786 avatar

Linux Perf

  • 427 installs
  • 155 repo stars
  • Updated June 27, 2026
  • mohitmishra786/low-level-dev-skills

linux-perf is an agent skill that guides developers through Linux perf stat, perf record, perf report, and flamegraph export to locate CPU hotspots, cache misses, and branch mispredictions in native binaries.

About

linux-perf is a low-level-dev-skills agent skill for CPU performance analysis with the Linux perf profiler. It walks through nine workflow sections covering prerequisites, paranoid-level sysctl tuning, compiling with -g and -fno-omit-frame-pointer, perf stat hardware counters (cache-misses, IPC, branch-misses), perf record sampling at configurable frequencies, perf report hotspot review, perf annotate disassembly, off-CPU profiling, and flamegraph handoff. The skill explains how to interpret IPC below 1.0, cache-miss rates above 5%, and branch-miss rates above 5% as optimization signals. Developers reach for linux-perf when profiling C/C++/Rust binaries on Linux, diagnosing [unknown] stack frames, or feeding perf.data into flamegraph generators before shipping latency-sensitive services.

  • perf record/report flamegraphs
  • Hardware counter events
  • Kernel and userspace stacks
  • Off-CPU and syscall analysis
  • Before/after benchmark diffs

Linux Perf by the numbers

  • 427 all-time installs (skills.sh)
  • +25 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #101 of 596 Debugging skills by installs in the Skillselion catalog
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mohitmishra786/low-level-dev-skills --skill linux-perf

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs427
repo stars155
Last updatedJune 27, 2026
Repositorymohitmishra786/low-level-dev-skills

How do you profile CPU hotspots with Linux perf?

Sample CPU cycles, cache misses, and syscall hotspots in production-like Linux workloads to prioritize optimizations before shipping performance-critical native code.

Who is it for?

Systems and native-code developers optimizing C, C++, or Rust binaries on Linux who need sampling profiles and hardware counter interpretation before release.

Skip if: Developers profiling macOS or Windows workloads without Linux perf access should skip this skill because commands and kernel paranoid settings are Linux-specific.

When should I use this skill?

A user asks which function consumes the most CPU, how to measure cache misses or IPC, or how to generate a flamegraph from perf record output.

What you get

perf.data capture, perf stat counter report, annotated hotspot function list, and flamegraph input ready for visualization.

  • perf.data profile
  • perf stat counter summary
  • hotspot function report

By the numbers

  • Documents 9 numbered perf workflow sections from prerequisites through flamegraph handoff

Files

SKILL.mdMarkdownGitHub ↗

Linux perf

Purpose

Guide agents through perf for CPU profiling: sampling, hardware counter measurement, hotspot identification, and integration with flamegraph generation.

Triggers

  • "Which function is consuming the most CPU?"
  • "How do I measure cache misses / IPC?"
  • "How do I use perf to find hotspots?"
  • "How do I generate a flamegraph from perf data?"
  • "perf shows [unknown] or [kernel] frames"

Workflow

1. Prerequisites

# Install
sudo apt install linux-perf    # Debian/Ubuntu (version-matched)
sudo dnf install perf          # Fedora/RHEL

# Check permissions
# By default perf requires root or paranoid level ≤ 1
cat /proc/sys/kernel/perf_event_paranoid
# 2 = only CPU stats (not kernel), 1 = user+kernel, 0 = all, -1 = no restrictions

# Temporarily lower (session only)
sudo sysctl -w kernel.perf_event_paranoid=1

# Persistent
echo 'kernel.perf_event_paranoid=1' | sudo tee /etc/sysctl.d/99-perf.conf
sudo sysctl -p /etc/sysctl.d/99-perf.conf

Compile the target with debug symbols for useful frame data:

gcc -g -O2 -fno-omit-frame-pointer -o prog main.c
# -fno-omit-frame-pointer: essential for frame-pointer-based unwinding
# Alternative: compile with DWARF CFI and use --call-graph=dwarf

2. perf stat — quick counters

# Basic hardware counters
perf stat ./prog

# With specific events
perf stat -e cache-misses,cache-references,instructions,cycles,branch-misses ./prog

# Wall-clock comparison: N runs
perf stat -r 5 ./prog

# Attach to existing process
perf stat -p 12345 sleep 10

Interpret perf stat output:

  • IPC (instructions per cycle) < 1.0: memory-bound or stalled pipeline
  • cache-miss rate > 5%: significant cache pressure
  • branch-miss rate > 5%: branch predictor struggling

3. perf record — sampling

# Default: sample at 1000 Hz (cycles event)
perf record -g ./prog

# Specify frequency
perf record -F 999 -g ./prog

# Specific event
perf record -e cache-misses -g ./prog

# Attach to running process
perf record -F 999 -g -p 12345 sleep 30

# Off-CPU profiling (time spent waiting)
perf record -e sched:sched_switch -ag sleep 10

# DWARF call graphs (better for binaries without frame pointers)
perf record -F 999 --call-graph=dwarf ./prog

# Save to named file
perf record -o myapp.perf.data -g ./prog

4. perf report — interactive analysis

perf report                          # reads perf.data
perf report -i myapp.perf.data
perf report --no-children            # self time only (not cumulative)
perf report --sort comm,dso,sym      # sort by fields
perf report --stdio                  # non-interactive text output

Navigation in TUI:

  • Enter — expand a symbol
  • a — annotate (show assembly with hit counts)
  • s — show source (needs debug info)
  • d — filter by DSO (library)
  • t — filter by thread
  • ? — help

5. perf annotate — hot instructions

# Show assembly with hit percentages
perf annotate sym_name

# From report: press 'a' on a symbol
# Or directly:
perf annotate -i perf.data --symbol=hot_function --stdio

High hit count on a mov or vmovdqa suggests a cache miss at that load.

6. perf top — live profiling

# Live top, like 'top' but for functions
sudo perf top -g

# Filter by process
sudo perf top -p 12345

7. Feed into flamegraphs

# Generate perf script output
perf script > out.perf

# Use Brendan Gregg's FlameGraph tools
git clone https://github.com/brendangregg/FlameGraph
./FlameGraph/stackcollapse-perf.pl out.perf > out.folded
./FlameGraph/flamegraph.pl out.folded > flamegraph.svg

# Open flamegraph.svg in browser

See skills/profilers/flamegraphs for reading flamegraphs and interpreting results.

8. Common issues

ProblemCauseFix
Permission deniedperf_event_paranoid too highLower paranoid level or run with sudo
[unknown] framesMissing frame pointers or debug infoRecompile with -fno-omit-frame-pointer or use --call-graph=dwarf
[kernel] everywhereKernel symbols not visibleUse sudo perf record; install linux-image-$(uname -r)-dbgsym
No kallsymsKernel symbols unavailable`echo 0
Empty report for short programProgram exits too fastUse -F 9999 or instrument longer workload
DWARF unwinding slowLarge DWARF stackLimit with --call-graph dwarf,512

9. Useful events

# List all available events
perf list

# Common hardware events
cycles
instructions
cache-references
cache-misses
branch-instructions
branch-misses
stalled-cycles-frontend
stalled-cycles-backend

# Software events
context-switches
cpu-migrations
page-faults

# Tracepoints (requires root)
sched:sched_switch
syscalls:sys_enter_read

For a counter reference and interpretation guide, see references/events.md.

Related skills

  • Use skills/profilers/flamegraphs for SVG flamegraph generation and reading
  • Use skills/profilers/valgrind for cache simulation and memory profiling
  • Use skills/compilers/gcc or skills/compilers/clang for PGO from perf data (AutoFDO)

Related skills

How it compares

Pick this over generic debugging skills when Linux perf sampling, PMU hardware counters, and flamegraph preparation are the explicit profiling toolchain.

FAQ

What compile flags does linux-perf recommend for useful stack traces?

linux-perf recommends compiling with -g -O2 -fno-omit-frame-pointer so perf can unwind frames reliably. Alternatively, DWARF CFI with --call-graph=dwarf works when frame pointers are omitted.

How does linux-perf interpret low IPC in perf stat output?

linux-perf treats instructions-per-cycle below 1.0 as a memory-bound or pipeline-stall signal. Cache-miss and branch-miss rates above 5% indicate significant optimization pressure worth investigating.

What permissions does perf need on Linux?

linux-perf documents kernel.perf_event_paranoid: level 2 limits kernel visibility, level 1 allows user plus kernel sampling, and level 0 or -1 relaxes restrictions. Root or lowered paranoid is required for full profiles.

Debuggingbackendtesting

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.