Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
orchestra-research avatar

Tensorrt Llm

  • 516 installs
  • 11.2k repo stars
  • Updated June 16, 2026
  • orchestra-research/ai-research-skills

tensorrt-llm is an agent skill that configures NVIDIA TensorRT-LLM tensor, pipeline, and expert parallelism plus FP8/INT4 quantization so developers serving large language models on multi-GPU clusters can maximize throug

About

tensorrt-llm is a version 1.0.0 Orchestra Research agent skill that documents production LLM inference with NVIDIA TensorRT-LLM on A100, H100, and GB200 GPUs. It covers pip and Docker installation, the Python LLM API, trtllm-serve with --tp_size and --max_batch_size, and parallelism strategies including tensor parallelism for same-node sharding, pipeline parallelism for 405B-class models, and expert parallelism for MoE architectures like Mixtral. Reference guides span optimization, multi-GPU setup, and serving, with benchmarks citing up to 24,000 tokens/sec on Llama 3-8B and 100× faster inference versus PyTorch in documented H100 tests. Developers reach for tensorrt-llm when a model exceeds single-GPU memory, when NVLink or InfiniBand multi-node serving is required, or when FP8 quantization can halve memory on H100—prefer vLLM or llama.cpp when hardware is non-NVIDIA or setup simplicity matters more.

  • Tensor Parallelism (TP) for single-node low-latency sharding across GPUs
  • Pipeline Parallelism (PP) for very large models across nodes with micro-batching
  • Worked examples (e.g. Llama 3-70B TP=4, Llama 3-405B TP=4 × PP=2 on 8× H100)
  • Guidance on communication overhead, throughput, and when NVLink matters
  • Expert parallelism patterns for MoE-scale deployments

Tensorrt Llm by the numbers

  • 516 all-time installs (skills.sh)
  • +34 installs in the week ending Jul 26, 2026 (Skillselion tracking)
  • Ranked #366 of 1,041 Cloud & Infrastructure skills by installs in the Skillselion catalog
  • Security screen: MEDIUM risk (skills.sh audit)
  • Data as of Jul 28, 2026 (Skillselion catalog sync)
npx skills add https://github.com/orchestra-research/ai-research-skills --skill tensorrt-llm

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs516
repo stars11.2k
Security audit1 / 3 scanners passed
Last updatedJune 16, 2026
Repositoryorchestra-research/ai-research-skills

How do you serve large LLMs on multiple GPUs?

Choose and configure TensorRT-LLM tensor, pipeline, and expert parallelism when serving large models across multiple GPUs or nodes.

Who is it for?

ML platform engineers deploying Llama, Qwen, Mixtral, or DeepSeek models on NVIDIA GPU clusters who need TensorRT-LLM throughput and multi-GPU sharding guidance.

Skip if: Developers on AMD GPUs, CPU-only edge targets, or teams wanting a Python-first PagedAttention setup without TensorRT compilation.

When should I use this skill?

User deploys TensorRT-LLM, configures tensor or pipeline parallelism, tunes trtllm-serve, or optimizes FP8/INT4 LLM inference on NVIDIA hardware.

What you get

TensorRT-LLM LLM or trtllm-serve configuration, parallelism sizing, and multi-GPU deployment reference aligned to model memory and latency targets.

  • Parallelism configuration
  • trtllm-serve command
  • Multi-GPU deployment reference

By the numbers

  • Includes 3 reference guides: optimization, multi-gpu, and serving
  • Documents up to 24,000 tokens/sec for Llama 3-8B on H100
  • Lists 100+ supported HuggingFace models and pip package tensorrt_llm==1.2.0rc3

Files

SKILL.mdMarkdownGitHub ↗

TensorRT-LLM

NVIDIA's open-source library for optimizing LLM inference with state-of-the-art performance on NVIDIA GPUs.

When to use TensorRT-LLM

Use TensorRT-LLM when:

  • Deploying on NVIDIA GPUs (A100, H100, GB200)
  • Need maximum throughput (24,000+ tokens/sec on Llama 3)
  • Require low latency for real-time applications
  • Working with quantized models (FP8, INT4, FP4)
  • Scaling across multiple GPUs or nodes

Use vLLM instead when:

  • Need simpler setup and Python-first API
  • Want PagedAttention without TensorRT compilation
  • Working with AMD GPUs or non-NVIDIA hardware

Use llama.cpp instead when:

  • Deploying on CPU or Apple Silicon
  • Need edge deployment without NVIDIA GPUs
  • Want simpler GGUF quantization format

Quick start

Installation

# Docker (recommended)
docker pull nvidia/tensorrt_llm:latest

# pip install
pip install tensorrt_llm==1.2.0rc3

# Requires CUDA 13.0.0, TensorRT 10.13.2, Python 3.10-3.12

Basic inference

from tensorrt_llm import LLM, SamplingParams

# Initialize model
llm = LLM(model="meta-llama/Meta-Llama-3-8B")

# Configure sampling
sampling_params = SamplingParams(
    max_tokens=100,
    temperature=0.7,
    top_p=0.9
)

# Generate
prompts = ["Explain quantum computing"]
outputs = llm.generate(prompts, sampling_params)

for output in outputs:
    print(output.text)

Serving with trtllm-serve

# Start server (automatic model download and compilation)
trtllm-serve meta-llama/Meta-Llama-3-8B \
    --tp_size 4 \              # Tensor parallelism (4 GPUs)
    --max_batch_size 256 \
    --max_num_tokens 4096

# Client request
curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-8B",
    "messages": [{"role": "user", "content": "Hello!"}],
    "temperature": 0.7,
    "max_tokens": 100
  }'

Key features

Performance optimizations

  • In-flight batching: Dynamic batching during generation
  • Paged KV cache: Efficient memory management
  • Flash Attention: Optimized attention kernels
  • Quantization: FP8, INT4, FP4 for 2-4× faster inference
  • CUDA graphs: Reduced kernel launch overhead

Parallelism

  • Tensor parallelism (TP): Split model across GPUs
  • Pipeline parallelism (PP): Layer-wise distribution
  • Expert parallelism: For Mixture-of-Experts models
  • Multi-node: Scale beyond single machine

Advanced features

  • Speculative decoding: Faster generation with draft models
  • LoRA serving: Efficient multi-adapter deployment
  • Disaggregated serving: Separate prefill and generation

Common patterns

Quantized model (FP8)

from tensorrt_llm import LLM

# Load FP8 quantized model (2× faster, 50% memory)
llm = LLM(
    model="meta-llama/Meta-Llama-3-70B",
    dtype="fp8",
    max_num_tokens=8192
)

# Inference same as before
outputs = llm.generate(["Summarize this article..."])

Multi-GPU deployment

# Tensor parallelism across 8 GPUs
llm = LLM(
    model="meta-llama/Meta-Llama-3-405B",
    tensor_parallel_size=8,
    dtype="fp8"
)

Batch inference

# Process 100 prompts efficiently
prompts = [f"Question {i}: ..." for i in range(100)]

outputs = llm.generate(
    prompts,
    sampling_params=SamplingParams(max_tokens=200)
)

# Automatic in-flight batching for maximum throughput

Performance benchmarks

Meta Llama 3-8B (H100 GPU):

  • Throughput: 24,000 tokens/sec
  • Latency: ~10ms per token
  • vs PyTorch: 100× faster

Llama 3-70B (8× A100 80GB):

  • FP8 quantization: 2× faster than FP16
  • Memory: 50% reduction with FP8

Supported models

  • LLaMA family: Llama 2, Llama 3, CodeLlama
  • GPT family: GPT-2, GPT-J, GPT-NeoX
  • Qwen: Qwen, Qwen2, QwQ
  • DeepSeek: DeepSeek-V2, DeepSeek-V3
  • Mixtral: Mixtral-8x7B, Mixtral-8x22B
  • Vision: LLaVA, Phi-3-vision
  • 100+ models on HuggingFace

References

  • [Optimization Guide](references/optimization.md) - Quantization, batching, KV cache tuning
  • [Multi-GPU Setup](references/multi-gpu.md) - Tensor/pipeline parallelism, multi-node
  • [Serving Guide](references/serving.md) - Production deployment, monitoring, autoscaling

Resources

  • Docs: https://nvidia.github.io/TensorRT-LLM/
  • GitHub: https://github.com/NVIDIA/TensorRT-LLM
  • Models: https://huggingface.co/models?library=tensorrt_llm

Related skills

How it compares

Choose tensorrt-llm for peak NVIDIA GPU serving; prefer vLLM for simpler Python APIs or llama.cpp for CPU and Apple Silicon edge deployment.

FAQ

What parallelism modes does the tensorrt-llm skill cover?

The tensorrt-llm skill documents tensor parallelism for horizontal layer splits, pipeline parallelism for vertical layer distribution on 405B-class models, and expert parallelism for MoE models such as Mixtral-8x22B. It includes decision trees, NVLink guidance, and multi-node Ray

When should developers pick TensorRT-LLM over vLLM?

The tensorrt-llm skill recommends TensorRT-LLM for maximum NVIDIA GPU throughput, FP8/INT4 quantization, and multi-GPU scaling on A100 or H100. It points to vLLM when you want simpler Python-first setup and PagedAttention without TensorRT compilation, or when hardware is not NVID

What performance numbers does tensorrt-llm cite?

The tensorrt-llm skill cites up to 24,000 tokens/sec throughput for Meta Llama 3-8B on H100, roughly 10 ms per token latency, and about 100× faster inference than PyTorch in its documented benchmarks. Llama 3-70B on 4× A100 with FP8 is shown at 10,000–15,000 tokens/sec.

Is Tensorrt Llm safe to install?

skills.sh reports 1 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.