Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
agricidaniel avatar

Claude Video Audio

  • 6 installs
  • 23 repo stars
  • Updated April 6, 2026
  • agricidaniel/claude-video

claude-video-audio is a Claude Code skill that processes video audio with FFmpeg, including loudness normalization and noise reduction.

About

claude-video-audio is a Claude Code skill that processes the audio in video files with FFmpeg. It handles loudness normalization to streaming LUFS standards, noise reduction, silence detection and removal, audio extraction and replacement, equalization, compression, and ducking. A developer uses it to clean up or level a video's audio track for a target platform.

  • FFmpeg audio processing: loudness normalization, noise reduction, silence removal
  • Per-platform LUFS targets for YouTube, Spotify, Podcast, and more
  • Audio extraction, replacement, mixing, EQ, and ducking

Claude Video Audio by the numbers

  • 6 all-time installs (skills.sh)
  • Ranked #1,096 of 1,335 Generative Media skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

claude-video-audio capabilities & compatibility

Free; runs on local FFmpeg with no API keys.

Capabilities
audio processing · loudness normalization · noise reduction · silence removal · audio extraction
Pricing
Free
From the docs

What claude-video-audio says it does

Audio processing for video files using FFmpeg. Loudness normalization to streaming
SKILL.md
Two-pass workflow for accurate normalization:
SKILL.md
npx skills add https://github.com/agricidaniel/claude-video --skill claude-video-audio

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs6
repo stars23
Last updatedApril 6, 2026
Repositoryagricidaniel/claude-video

What it does

Normalize loudness, reduce noise, remove silence, or extract and replace the audio track of a video with FFmpeg.

Who is it for?

Normalizing loudness to platform LUFS targets and reducing noise in video audio

Skip if: AI model-based audio like source separation or speaker diarization, which claude-video-enhance-audio handles

When should I use this skill?

You need to normalize, denoise, extract, replace, or remove silence from a video's audio with FFmpeg

What you get

Produces a video with normalized, cleaned audio ready for the target platform.

  • Loudness-normalized video
  • Denoised audio
  • Extracted audio files

By the numbers

  • Two-pass loudnorm workflow
  • Loudness targets for 6 platforms (YouTube, Spotify, Apple Music, Podcast, Broadcast, TikTok/Reels)

Files

SKILL.mdMarkdownGitHub ↗

claude-video-audio — Audio Processing

Pre-Flight

1. Run bash scripts/preflight.sh "$INPUT" "$OUTPUT" 2. Analyze audio streams: ffprobe -v error -select_streams a -show_entries stream=codec_name,channels,sample_rate,bit_rate -of json "$INPUT"

Loudness Normalization (EBU R128)

Two-pass workflow for accurate normalization:

# Pass 1: Measure current loudness
ffmpeg -i "$INPUT" -af loudnorm=I=-14:TP=-1.5:LRA=11:print_format=json -f null - 2>&1 | tail -12

Extract measured_I, measured_TP, measured_LRA, measured_thresh, offset from the JSON output.

# Pass 2: Apply normalization with measured values
ffmpeg -n -i "$INPUT" -af "loudnorm=I=-14:TP=-1.5:LRA=11:\
measured_I=MEASURED_I:measured_TP=MEASURED_TP:measured_LRA=MEASURED_LRA:\
measured_thresh=MEASURED_THRESH:offset=MEASURED_OFFSET:linear=true" \
  -c:v copy "$OUTPUT"

Target Loudness by Platform

PlatformTarget LUFSTrue PeakNotes
YouTube-14 LUFS-1.0 dBTPYouTube normalizes down, not up
Spotify-14 LUFS-1.0 dBTP"Loud" mode = -11 LUFS
Apple Music-16 LUFS-1.0 dBTPSound Check
Podcast-16 LUFS-1.5 dBTPConversational clarity
Broadcast (EBU R128)-23 LUFS-1.0 dBTPEuropean standard
TikTok/Reels-14 LUFS-1.0 dBTPMatch YouTube

Noise Reduction

FFT-based (good for consistent background noise):

ffmpeg -n -i "$INPUT" -af "afftdn=nf=-25:nt=w" -c:v copy "$OUTPUT"

nf = noise floor in dB (lower = more aggressive). nt=w = white noise model.

Non-local means (better quality, slower):

ffmpeg -n -i "$INPUT" -af "anlmdn=s=7:p=0.002:r=0.002" -c:v copy "$OUTPUT"

Highpass/lowpass to remove rumble and hiss:

ffmpeg -n -i "$INPUT" -af "highpass=f=80,lowpass=f=12000" -c:v copy "$OUTPUT"

Silence Detection and Removal

Detect silence:

ffmpeg -i "$INPUT" -af "silencedetect=noise=-30dB:d=0.5" -f null - 2>&1 | grep silence

Returns silence_start and silence_end timestamps.

Remove silence (using Auto-Editor — much easier):

auto-editor "$INPUT" --margin 0.3s -o "$OUTPUT"

Manual silence removal with FFmpeg:

ffmpeg -n -i "$INPUT" -af "silenceremove=start_periods=1:start_silence=0.5:start_threshold=-30dB:detection=rms" -c:v copy "$OUTPUT"

Audio Extraction

# Extract as WAV (lossless)
ffmpeg -n -i "$INPUT" -vn -c:a pcm_s16le audio.wav

# Extract as AAC
ffmpeg -n -i "$INPUT" -vn -c:a aac -b:a 256k audio.m4a

# Extract as MP3
ffmpeg -n -i "$INPUT" -vn -c:a libmp3lame -b:a 320k audio.mp3

# Extract as Opus (most efficient)
ffmpeg -n -i "$INPUT" -vn -c:a libopus -b:a 128k audio.ogg

# Stream copy (fastest, keeps original codec)
ffmpeg -n -i "$INPUT" -vn -c:a copy audio.m4a

Audio Replacement

# Replace audio track entirely
ffmpeg -n -i video.mp4 -i new_audio.wav -c:v copy -c:a aac -b:a 192k \
  -map 0:v:0 -map 1:a:0 "$OUTPUT"

# Mix original audio with new audio (e.g., background music)
ffmpeg -n -i video.mp4 -i music.mp3 -filter_complex \
  "[0:a]volume=1.0[orig];[1:a]volume=0.3[music];[orig][music]amix=inputs=2:duration=first" \
  -c:v copy "$OUTPUT"

Audio Effects

Volume adjustment:

ffmpeg -n -i "$INPUT" -af "volume=1.5" -c:v copy "$OUTPUT"   # 1.5x louder
ffmpeg -n -i "$INPUT" -af "volume=-6dB" -c:v copy "$OUTPUT"   # 6dB quieter

Equalization:

# Boost bass, cut muddy mids, add presence
ffmpeg -n -i "$INPUT" -af "equalizer=f=80:t=q:w=1:g=3,equalizer=f=300:t=q:w=2:g=-2,equalizer=f=4000:t=q:w=1:g=2" -c:v copy "$OUTPUT"

Compression (reduce dynamic range):

ffmpeg -n -i "$INPUT" -af "acompressor=threshold=-20dB:ratio=4:attack=5:release=50" -c:v copy "$OUTPUT"

Limiter (prevent clipping):

ffmpeg -n -i "$INPUT" -af "alimiter=limit=0.95:level=false" -c:v copy "$OUTPUT"

Fade in/out:

ffmpeg -n -i "$INPUT" -af "afade=t=in:st=0:d=2,afade=t=out:st=OFFSET:d=2" -c:v copy "$OUTPUT"

Bass boost: -af "bass=g=5:f=100:w=0.5" Treble boost: -af "treble=g=3:f=4000:w=0.5"

Audio Format Conversion

CodecFFmpeg EncoderRecommended BitrateUse Case
AACaac128-256kVideo, streaming
MP3libmp3lame192-320kUniversal playback
Opuslibopus96-160kBest efficiency, WebM
FLACflacLosslessArchival, editing
WAVpcm_s16leUncompressedEditing, processing
AC3ac3384-640kSurround sound

Channel Layout

# Stereo to mono
ffmpeg -n -i "$INPUT" -ac 1 -c:v copy "$OUTPUT"

# Mono to stereo (duplicate)
ffmpeg -n -i "$INPUT" -ac 2 -c:v copy "$OUTPUT"

# Extract left channel only
ffmpeg -n -i "$INPUT" -af "pan=mono|c0=FL" -c:v copy "$OUTPUT"

Audio Sync Fix

# Delay audio by 0.5 seconds
ffmpeg -n -i "$INPUT" -itsoffset 0.5 -i "$INPUT" -map 0:v -map 1:a -c copy "$OUTPUT"

# Advance audio by 0.5 seconds
ffmpeg -n -i "$INPUT" -itsoffset -0.5 -i "$INPUT" -map 0:v -map 1:a -c copy "$OUTPUT"

Audio Ducking (Sidechain Compression)

Automatically lower music volume when speech is present.

Basic ducking (voice track ducks music track):

ffmpeg -n -i video_with_voice.mp4 -i music.mp3 -filter_complex \
  "[0:a]aformat=fltp:44100:stereo[voice];\
   [1:a]aformat=fltp:44100:stereo[music];\
   [music][voice]sidechaincompress=threshold=0.015:ratio=6:attack=200:release=1000:level_sc=1[ducked];\
   [voice][ducked]amix=inputs=2:duration=first:weights=1 0.4[out]" \
  -map 0:v -map "[out]" -c:v copy "$OUTPUT"

Parameters guide:

ParameterDefaultPurpose
threshold0.015Voice level that triggers ducking (lower = more sensitive)
ratio6How much to reduce music (higher = more reduction)
attack200msHow quickly music ducks down
release1000msHow quickly music returns after speech stops
weights1 0.4Voice volume : music volume in final mix

Ducking with existing audio streams (video has voice + separate music file):

ffmpeg -n -i "$INPUT" -i music.mp3 -filter_complex \
  "[1:a][0:a]sidechaincompress=threshold=0.02:ratio=8:attack=100:release=800[ducked];\
   [0:a][ducked]amix=inputs=2:duration=first:weights=1 0.3[out]" \
  -map 0:v -map "[out]" -c:v copy "$OUTPUT"

Tips:

  • Lower threshold (e.g., 0.01) for quieter speech
  • Higher ratio (e.g., 10-20) for more aggressive ducking
  • Shorter attack (e.g., 50ms) for spoken word podcasts
  • Longer release (e.g., 2000ms) for smoother music return
  • Adjust weights to control the overall voice-to-music balance

Reference

Load references/audio.md for extended EBU R128 loudness details, advanced filter chains, and format-specific audio settings.

Related skills

FAQ

What loudness target does it use for YouTube?

It targets -14 LUFS with a -1.0 dBTP true peak for YouTube.

How does it normalize loudness accurately?

It uses a two-pass EBU R128 loudnorm workflow, measuring current loudness then applying the measured values.

Generative Mediaagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.