Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
inference-sh avatar

Speech To Text

  • 575 installs
  • 680 repo stars
  • Updated August 3, 2026
  • inference-sh/skills

Speech-to-text is an agent skill that transcribes meetings, podcasts, and voice notes into searchable text or subtitles through inference.sh so developers who need STT can run Whisper and ElevenLabs Scribe models without

About

Speech-to-text is an inference.sh agent skill that transcribes audio via the belt CLI using ElevenLabs Scribe v2, Fast Whisper Large V3, and Whisper V3 Large. It supports transcription, translation, multi-language output, timestamps, speaker diarization, and audio event tagging for meetings, podcasts, subtitles, and voice notes. Developers reach for speech-to-text when they need production STT from the terminal or agent workflows instead of implementing Whisper or ElevenLabs SDK calls directly.

  • Runs ElevenLabs Scribe v2, Fast Whisper Large V3, and Whisper V3 Large via inference.sh app IDs
  • Supports diarization, timestamps, multi-language transcription, translation, and audio event tagging
  • Meeting, podcast, subtitle, and voice-note oriented trigger phrases in SKILL.md
  • Quick start: `belt login` then `belt app run` with an `audio_url` JSON payload
  • Install path documented: `npx skills add belt-sh/cli` for the belt CLI skill

Speech To Text by the numbers

  • 575 all-time installs (skills.sh)
  • +9 installs in the week ending Aug 2, 2026 (Skillselion tracking)
  • Ranked #1,626 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Security screen: MEDIUM risk (skills.sh audit)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/inference-sh/skills --skill speech-to-text

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs575
repo stars680
Security audit1 / 3 scanners passed
Last updatedAugust 3, 2026
Repositoryinference-sh/skills

How do you transcribe audio with Whisper via CLI?

Transcribe meetings, podcasts, and voice notes into searchable text or subtitles through inference.sh without wiring Whisper or ElevenLabs APIs by hand.

Who is it for?

Developers adding meeting transcription, podcast captions, or voice-note search who want inference.sh CLI access to Scribe and Whisper models.

Skip if: Real-time streaming STT in a custom WebRTC client or fine-tuning proprietary acoustic models outside inference.sh.

When should I use this skill?

A developer asks to transcribe a meeting, generate subtitles, or convert voice notes to text with Whisper or ElevenLabs.

What you get

Plain-text or timestamped transcript, optional subtitles, speaker labels, and translated text from source audio.

  • transcript text
  • subtitles
  • timestamped segments

By the numbers

  • ElevenLabs Scribe v2 advertises 98%+ transcription accuracy
  • Supports 3 STT models: Scribe v2, Fast Whisper Large V3, Whisper V3 Large

Files

SKILL.mdMarkdownGitHub ↗
Install the belt CLI skill: npx skills add belt-sh/cli

Speech-to-Text

Transcribe audio to text via inference.sh CLI.

Speech-to-Text
Speech-to-Text

Quick Start

Requires inference.sh CLI (belt). Install instructions
belt login

belt app run infsh/fast-whisper-large-v3 --input '{"audio_url": "https://audio.mp3"}'

Available Models

ModelApp IDBest For
ElevenLabs Scribe v2elevenlabs/stt98%+ accuracy, diarization, 90+ languages
Fast Whisper V3infsh/fast-whisper-large-v3Fast transcription
Whisper V3 Largeinfsh/whisper-v3-largeHighest accuracy

Examples

Basic Transcription

belt app run infsh/fast-whisper-large-v3 --input '{"audio_url": "https://meeting.mp3"}'

With Timestamps

belt app sample infsh/fast-whisper-large-v3 --save input.json

# {
#   "audio_url": "https://podcast.mp3",
#   "timestamps": true
# }

belt app run infsh/fast-whisper-large-v3 --input input.json

Translation (to English)

belt app run infsh/whisper-v3-large --input '{
  "audio_url": "https://french-audio.mp3",
  "task": "translate"
}'

From Video

# Extract audio from video first
belt app run infsh/video-audio-extractor --input '{"video_url": "https://video.mp4"}' > audio.json

# Transcribe the extracted audio
belt app run infsh/fast-whisper-large-v3 --input '{"audio_url": "<audio-url>"}'

Workflow: Video Subtitles

# 1. Transcribe video audio
belt app run infsh/fast-whisper-large-v3 --input '{
  "audio_url": "https://video.mp4",
  "timestamps": true
}' > transcript.json

# 2. Use transcript for captions
belt app run infsh/caption-videos --input '{
  "video_url": "https://video.mp4",
  "captions": "<transcript-from-step-1>"
}'

Supported Languages

Whisper supports 99+ languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, Hindi, Russian, and many more.

Use Cases

  • Meetings: Transcribe recordings
  • Podcasts: Generate transcripts
  • Subtitles: Create captions for videos
  • Voice Notes: Convert to searchable text
  • Interviews: Transcription for research
  • Accessibility: Make audio content accessible

Output Format

Returns JSON with:

  • text: Full transcription
  • segments: Timestamped segments (if requested)
  • language: Detected language

Related Skills

# ElevenLabs STT (98%+ accuracy, diarization)
npx skills add inference-sh/skills@elevenlabs-stt

# ElevenLabs TTS (reverse direction)
npx skills add inference-sh/skills@elevenlabs-tts

# Full platform skill (all 250+ apps)
npx skills add inference-sh/skills@infsh-cli

# Text-to-speech (reverse direction)
npx skills add inference-sh/skills@text-to-speech

# Video generation (add captions)
npx skills add inference-sh/skills@ai-video-generation

# AI avatars (lipsync with transcripts)
npx skills add inference-sh/skills@ai-avatar-video

Browse all audio apps: belt app store --category audio

Documentation

Related skills

Forks & variants (2)

Speech To Text has 2 known copies in the catalog totaling 76 installs. They canonicalize to this original listing.

How it compares

Pick speech-to-text for quick CLI-driven batch transcription via inference.sh; build custom SDK integrations when you need embedded real-time streaming control.

FAQ

Which models does speech-to-text support?

Speech-to-text routes audio through inference.sh to ElevenLabs Scribe v2, Fast Whisper Large V3, and Whisper V3 Large, with Scribe v2 advertising 98%+ accuracy plus diarization.

How do you run speech-to-text?

Speech-to-text uses the inference.sh belt CLI with allowed Bash tooling, letting developers transcribe, translate, timestamp, and diarize audio without wiring Whisper or ElevenLabs APIs manually.

Is Speech To Text safe to install?

skills.sh reports 1 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

AI & Agent Buildingautomationllm

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.