
Text To Speech
- 1 installs
- 6 repo stars
- Updated August 3, 2026
- avivsinai/telclaude
text-to-speech is a Claude Code skill that converts text to speech audio using the OpenAI TTS API via the telclaude tts CLI.
About
This skill converts text to speech audio using the OpenAI TTS API through the telclaude tts CLI. A developer uses it so a Telegram agent can reply with voice, matching the medium when a user sends a voice message. It supports six voices, speed and quality options, and a voice-message flag for waveform display.
- Converts text to speech via the OpenAI TTS API using telclaude tts
- --voice-message flag outputs OGG/Opus with Telegram waveform display
- Six voices, adjustable speed, and tts-1 / tts-1-hd quality models
Text To Speech by the numbers
- 1 all-time installs (skills.sh)
- Ranked #1,200 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
text-to-speech capabilities & compatibility
Requires OPENAI_API_KEY; tts-1 is $0.015 per 1K chars, tts-1-hd is $0.030 per 1K chars.
- Capabilities
- text to speech · image generation · transcription
- Works with
- openai
- Use cases
- transcription
- Pricing
- Bring your own API key
What text-to-speech says it does
Converts text to speech audio using OpenAI TTS API.
This outputs OGG/Opus format that displays as a voice message with waveform in Telegram.
npx skills add https://github.com/avivsinai/telclaude --skill text-to-speechAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1 |
|---|---|
| repo stars | ★ 6 |
| Last updated | August 3, 2026 |
| Repository | avivsinai/telclaude ↗ |
What it does
Convert text into a spoken audio reply for a Telegram user, matching a voice-in message.
Who is it for?
Sending voice replies when a Telegram user speaks or wants audio.
Skip if: Speech-to-text transcription or non-telclaude audio pipelines.
When should I use this skill?
A user sends a voice message or asks to read aloud, speak, or convert text to speech.
What you get
Outputs a TTS audio file path in the language the user spoke.
By the numbers
- 6 voices (alloy, echo, fable, onyx, nova, shimmer)
- 4096-char max per request
Files
Text-to-Speech Skill
CRITICAL: Voice Message Reply Rules
When a user sends you a voice message, follow these rules:
1. ALWAYS use `--voice-message` flag - Required for Telegram waveform display 2. Generate TTS in the SAME LANGUAGE the user spoke - If they spoke English, generate English audio 3. Output ONLY the file path - No text commentary alongside the voice reply
Exception: If the user explicitly asks for a text response (e.g., "respond in text", "don't send voice"), respond with text instead.
Correct Example (user sent voice in English):
telclaude tts "Hello! How can I help you today?" --voice-messageThen output ONLY:
/media/outbox/voice/1234567890-abc123.oggWRONG - Do NOT do this:
Hello! Here is the audio you requested:
/media/outbox/tts/1234567890-abc123.mp3This is wrong because: (1) added text alongside voice, (2) missing --voice-message flag, (3) mp3 instead of ogg, (4) wrong directory
---
When to Use
Use this skill when users:
- Ask to "read aloud", "speak", or "say" something
- Request audio versions of text content
- Want voice messages or audio responses
- Ask for text to be converted to speech
- Send a voice message (respond in voice - see CRITICAL rules above)
How to Generate Speech
Voice Messages (Telegram waveform display)
For conversational voice replies, use --voice-message to get proper Telegram voice message formatting:
telclaude tts "Your response here" --voice-messageThis outputs OGG/Opus format that displays as a voice message with waveform in Telegram.
Audio Files (music player display)
For regular audio files (longer content, podcast-style):
telclaude tts "Your text to convert to speech here"Or use the short alias:
telclaude tts "Your text here"Options
--voice-message: Output as Telegram voice message (OGG/Opus with waveform display)--voice: Voice to use (alloy, echo, fable, onyx, nova, shimmer). Default: alloy- alloy: Neutral, balanced voice
- echo: Deeper, more resonant voice
- fable: Expressive, storytelling voice
- onyx: Deep, authoritative voice
- nova: Warm, conversational voice
- shimmer: Soft, gentle voice
--speed: Speech speed from 0.25 to 4.0. Default: 1.0--model: Quality model (tts-1, tts-1-hd). Default: tts-1- tts-1: Standard quality, faster
- tts-1-hd: Higher quality, slightly slower
--format: Audio format (mp3, opus, aac, flac, wav). Default: mp3 (ignored with --voice-message)
Examples
# Voice message reply (when user sent a voice message)
telclaude tts "Sure, I can help you with that!" --voice-message
# Voice message with specific voice
telclaude tts "Here's what I found..." --voice-message --voice nova
# Regular audio file
telclaude tts "Hello! Here is your summary."
# High quality audio file
telclaude tts "Important announcement" --voice onyx --model tts-1-hd --speed 0.9Response Format
The telclaude tts command outputs metadata (file path, size, format, voice, duration). You only need to include the file path in your response - the relay handles sending it to Telegram.
Voice message replies (responding to incoming voice)
Output ONLY the file path - no commentary:
/media/outbox/voice/1234567890-abc123.oggThat's it. No "I've generated..." or "Here's your audio...". The relay sends just the voice message, like a human would.
Audio files or text+audio responses
If the user requested an audio FILE (not a voice reply), or you need to include text context:
Here's the summary as audio:
/media/outbox/tts/1234567890-abc123.mp3Key points:
- Voice messages:
.../voice/*.ogg- waveform display, path only - Audio files:
.../tts/*.mp3- music player display, text OK - The relay automatically detects paths and sends the media
- Paths live under
TELCLAUDE_MEDIA_OUTBOX_DIR(default.telclaude-mediain native mode;/media/outboxin Docker)
Best Practices
1. Match the medium: If user sends voice, respond with voice 2. Choose Appropriate Voice: Match the voice to the content type (e.g., fable for stories, onyx for announcements) 3. Keep Text Reasonable: Maximum 4096 characters per request 4. Consider Speed: Use slower speed (0.8-0.9) for important content, faster (1.2-1.5) for casual updates 5. Use HD Sparingly: tts-1-hd costs 2x more; use for important or long-form content
Limitations
- Maximum 4096 characters per request (longer text is truncated)
- Audio files are stored temporarily and cleaned up after 24 hours
- Requires OPENAI_API_KEY to be configured
Cost Awareness
OpenAI TTS pricing (per 1000 characters):
- tts-1: $0.015/1K chars
- tts-1-hd: $0.030/1K chars
Example: A 500-word response (~2500 chars) costs ~$0.04 with tts-1