Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
firecrawl avatar

Llama Cpp

  • 16 installs
  • 17 repo stars
  • Updated February 6, 2026
  • firecrawl/ai-research-skills

llama-cpp serves GGUF models on CPU and Apple Silicon via llama.cpp.

About

The llama-cpp skill documents llama.cpp pure C++ inference with brew or make builds, Metal CUDA and ROCm flags, GGUF downloads from HuggingFace, llama-cli interactive chat, and llama-server OpenAI-compatible HTTP on port 8080. Quantization table compares Q4_K_M recommended default, Q5_K_M quality, and size tradeoffs for 7B models. Use on CPU-only, M1-M4 Macs, AMD Intel GPUs, Raspberry Pi edge, or when avoiding NVIDIA CUDA stacks. Alternatives note TensorRT-LLM for datacenter NVIDIA throughput and vLLM for Python NVIDIA serving with PagedAttention.

  • Builds llama.cpp with Metal CUDA or ROCm flags.
  • Runs GGUF models via llama-cli and llama-server.
  • Exposes OpenAI-compatible local HTTP API.
  • Documents Q4_K_M and quantization size table.
  • Targets CPU Apple Silicon and non-NVIDIA GPU deploys.

Llama Cpp by the numbers

  • 16 all-time installs (skills.sh)
  • Ranked #1,320 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
At a glance

llama-cpp capabilities & compatibility

Capabilities
when to use llama.cpp criteria · quick start brew make install · quantization formats q4_k_m table
Use cases
api development
Platforms
macOS · Linux · Windows
Runs
Runs locally
From the docs

What llama-cpp says it does

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware
SKILL.md
npx skills add https://github.com/firecrawl/ai-research-skills --skill llama-cpp

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs16
repo stars17
Last updatedFebruary 6, 2026
Repositoryfirecrawl/ai-research-skills

How do I run Llama GGUF on a Mac with Metal?

Run GGUF quantized LLM inference on CPU, Apple Silicon, and non-NVIDIA GPUs.

Who is it for?

Developers deploying models without NVIDIA datacenter GPUs.

Skip if: Skip when already standardized on vLLM CUDA fleet.

When should I use this skill?

User runs llama.cpp, GGUF quantization, or edge inference.

What you get

Local llama-server responding to chat completions curl.

Files

SKILL.mdMarkdownGitHub ↗

llama.cpp

Pure C/C++ LLM inference with minimal dependencies, optimized for CPUs and non-NVIDIA hardware.

When to use llama.cpp

Use llama.cpp when:

  • Running on CPU-only machines
  • Deploying on Apple Silicon (M1/M2/M3/M4)
  • Using AMD or Intel GPUs (no CUDA)
  • Edge deployment (Raspberry Pi, embedded systems)
  • Need simple deployment without Docker/Python

Use TensorRT-LLM instead when:

  • Have NVIDIA GPUs (A100/H100)
  • Need maximum throughput (100K+ tok/s)
  • Running in datacenter with CUDA

Use vLLM instead when:

  • Have NVIDIA GPUs
  • Need Python-first API
  • Want PagedAttention

Quick start

Installation

# macOS/Linux
brew install llama.cpp

# Or build from source
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make

# With Metal (Apple Silicon)
make LLAMA_METAL=1

# With CUDA (NVIDIA)
make LLAMA_CUDA=1

# With ROCm (AMD)
make LLAMA_HIP=1

Download model

# Download from HuggingFace (GGUF format)
huggingface-cli download \
    TheBloke/Llama-2-7B-Chat-GGUF \
    llama-2-7b-chat.Q4_K_M.gguf \
    --local-dir models/

# Or convert from HuggingFace
python convert_hf_to_gguf.py models/llama-2-7b-chat/

Run inference

# Simple chat
./llama-cli \
    -m models/llama-2-7b-chat.Q4_K_M.gguf \
    -p "Explain quantum computing" \
    -n 256  # Max tokens

# Interactive chat
./llama-cli \
    -m models/llama-2-7b-chat.Q4_K_M.gguf \
    --interactive

Server mode

# Start OpenAI-compatible server
./llama-server \
    -m models/llama-2-7b-chat.Q4_K_M.gguf \
    --host 0.0.0.0 \
    --port 8080 \
    -ngl 32  # Offload 32 layers to GPU

# Client request
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama-2-7b-chat",
    "messages": [{"role": "user", "content": "Hello!"}],
    "temperature": 0.7,
    "max_tokens": 100
  }'

Quantization formats

GGUF format overview

FormatBitsSize (7B)SpeedQualityUse Case
Q4_K_M4.54.1 GBFastGoodRecommended default
Q4_K_S4.33.9 GBFasterLowerSpeed critical
Q5_K_M5.54.8 GBMediumBetterQuality critical
Q6_K6.55.5 GBSlowerBestMaximum quality
Q8_08.07.0 GBSlowExcellentMinimal degradation
Q2_K2.52.7 GBFastestPoorTesting only

Choosing quantization

# General use (balanced)
Q4_K_M  # 4-bit, medium quality

# Maximum speed (more degradation)
Q2_K or Q3_K_M

# Maximum quality (slower)
Q6_K or Q8_0

# Very large models (70B, 405B)
Q3_K_M or Q4_K_S  # Lower bits to fit in memory

Hardware acceleration

Apple Silicon (Metal)

# Build with Metal
make LLAMA_METAL=1

# Run with GPU acceleration (automatic)
./llama-cli -m model.gguf -ngl 999  # Offload all layers

# Performance: M3 Max 40-60 tokens/sec (Llama 2-7B Q4_K_M)

NVIDIA GPUs (CUDA)

# Build with CUDA
make LLAMA_CUDA=1

# Offload layers to GPU
./llama-cli -m model.gguf -ngl 35  # Offload 35/40 layers

# Hybrid CPU+GPU for large models
./llama-cli -m llama-70b.Q4_K_M.gguf -ngl 20  # GPU: 20 layers, CPU: rest

AMD GPUs (ROCm)

# Build with ROCm
make LLAMA_HIP=1

# Run with AMD GPU
./llama-cli -m model.gguf -ngl 999

Common patterns

Batch processing

# Process multiple prompts from file
cat prompts.txt | ./llama-cli \
    -m model.gguf \
    --batch-size 512 \
    -n 100

Constrained generation

# JSON output with grammar
./llama-cli \
    -m model.gguf \
    -p "Generate a person: " \
    --grammar-file grammars/json.gbnf

# Outputs valid JSON only

Context size

# Increase context (default 512)
./llama-cli \
    -m model.gguf \
    -c 4096  # 4K context window

# Very long context (if model supports)
./llama-cli -m model.gguf -c 32768  # 32K context

Performance benchmarks

CPU performance (Llama 2-7B Q4_K_M)

CPUThreadsSpeedCost
Apple M3 Max1650 tok/s$0 (local)
AMD Ryzen 9 7950X3235 tok/s$0.50/hour
Intel i9-13900K3230 tok/s$0.40/hour
AWS c7i.16xlarge6440 tok/s$2.88/hour

GPU acceleration (Llama 2-7B Q4_K_M)

GPUSpeedvs CPUCost
NVIDIA RTX 4090120 tok/s3-4×$0 (local)
NVIDIA A1080 tok/s2-3×$1.00/hour
AMD MI25070 tok/s$2.00/hour
Apple M3 Max (Metal)50 tok/s~Same$0 (local)

Supported models

LLaMA family:

  • Llama 2 (7B, 13B, 70B)
  • Llama 3 (8B, 70B, 405B)
  • Code Llama

Mistral family:

  • Mistral 7B
  • Mixtral 8x7B, 8x22B

Other:

  • Falcon, BLOOM, GPT-J
  • Phi-3, Gemma, Qwen
  • LLaVA (vision), Whisper (audio)

Find models: https://huggingface.co/models?library=gguf

References

  • [Quantization Guide](references/quantization.md) - GGUF formats, conversion, quality comparison
  • [Server Deployment](references/server.md) - API endpoints, Docker, monitoring
  • [Optimization](references/optimization.md) - Performance tuning, hybrid CPU+GPU

Resources

  • GitHub: https://github.com/ggerganov/llama.cpp
  • Models: https://huggingface.co/models?library=gguf
  • Discord: https://discord.gg/llama-cpp

Related skills

FAQ

What does llama-cpp do?

llama-cpp serves GGUF models on CPU and Apple Silicon via llama.cpp.

When should I use llama-cpp?

User runs llama.cpp, GGUF quantization, or edge inference.

Is this skill safe to install?

Review the Security Audits panel on this page before installing in production.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.