Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
cinience avatar

Alicloud Ai Audio Tts Realtime

  • 272 installs
  • 396 repo stars
  • Updated July 18, 2026
  • cinience/alicloud-skills

alicloud-ai-audio-tts-realtime is a version 1.0.0 agent skill that wires Alibaba Cloud Model Studio Qwen realtime TTS models into voice agents for developers who need low-latency streaming speech synthesis with instructi

About

alicloud-ai-audio-tts-realtime is a version 1.0.0 Cinience agent skill for Alibaba Cloud Model Studio Qwen realtime text-to-speech. It supports five exact model strings—including qwen3-tts-flash-realtime and instruction-controlled variants—via the dashscope SDK in a Python virtual environment with DASHSCOPE_API_KEY authentication. The normalized tts.realtime interface accepts text, voice, optional instruction and sample_rate, returning audio_base64_pcm_chunks over websocket streaming. Developers reach for this skill when building voice agents, live assistants, or interactive apps needing immediate spoken responses rather than batch TTS. A bundled realtime_tts_demo.py script probes SDK compatibility, supports --strict CI gating, and can fallback to non-realtime models. Output lands in output/ai-audio-tts-realtime/audio/ with py_compile validation scripts. Operational guidance keeps utterances short for lower latency and instructions concise on instruct models.

  • Streaming TTS session management
  • Low-latency audio chunk delivery
  • Live voice agent integration
  • WebSocket or stream API wiring
  • Interactive assistant speech output

Alicloud Ai Audio Tts Realtime by the numbers

  • 272 all-time installs (skills.sh)
  • Ranked #534 of 1,335 Generative Media skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/cinience/alicloud-skills --skill alicloud-ai-audio-tts-realtime

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs272
repo stars396
Last updatedJuly 18, 2026
Repositorycinience/alicloud-skills

How do you add realtime TTS to a voice agent?

Wire low-latency streaming TTS into voice agents, live assistants, or interactive apps requiring immediate spoken responses from Alibaba Cloud.

Who is it for?

Developers building voice agents or live assistants on Alibaba Cloud who need streaming Qwen TTS with dashscope SDK integration.

Skip if: Batch offline TTS or non-Alibaba speech providers should use aliyun-qwen-tts or other cloud audio skills instead.

When should I use this skill?

User needs low-latency streaming speech synthesis with Alibaba Cloud Qwen realtime TTS models.

What you get

Streaming PCM audio chunks, sample_rate metadata, probe-script WAV output, and validation artifacts under output/aliyun-qwen-tts-realtime/.

  • Streaming PCM audio output
  • realtime_tts_demo.py probe results
  • Validation artifacts in output/aliyun-qwen-tts-realtime/

By the numbers

  • Version 1.0.0 skill with 5 exact Qwen realtime TTS model strings
  • Includes realtime_tts_demo.py probe script with --strict CI mode

Files

SKILL.mdMarkdownGitHub ↗

Category: provider

Model Studio Qwen TTS Realtime

Use realtime TTS models for low-latency streaming speech output.

Critical model names

Use one of these exact model strings:

  • qwen3-tts-flash-realtime
  • qwen3-tts-instruct-flash-realtime
  • qwen3-tts-instruct-flash-realtime-2026-01-22
  • qwen3-tts-vd-realtime-2026-01-15
  • qwen3-tts-vc-realtime-2026-01-15

Prerequisites

  • Install SDK in a virtual environment:
python3 -m venv .venv
. .venv/bin/activate
python -m pip install dashscope
  • Set DASHSCOPE_API_KEY in your environment, or add dashscope_api_key to ~/.alibabacloud/credentials.

Normalized interface (tts.realtime)

Request

  • text (string, required)
  • voice (string, required)
  • instruction (string, optional)
  • sample_rate (int, optional)

Response

  • audio_base64_pcm_chunks (array<string>)
  • sample_rate (int)
  • finish_reason (string)

Operational guidance

  • Use websocket or streaming endpoint for realtime mode.
  • Keep each utterance short for lower latency.
  • For instruction models, keep instruction explicit and concise.
  • Some SDK/runtime combinations may reject realtime model calls over MultiModalConversation; use the probe script below to verify compatibility.

Local demo script

Use the probe script to verify realtime compatibility in your current SDK/runtime, and optionally fallback to a non-realtime model for immediate output:

.venv/bin/python skills/ai/audio/alicloud-ai-audio-tts-realtime/scripts/realtime_tts_demo.py \
  --text "This is a realtime speech demo." \
  --fallback \
  --output output/ai-audio-tts-realtime/audio/fallback-demo.wav

Strict mode (for CI / gating):

.venv/bin/python skills/ai/audio/alicloud-ai-audio-tts-realtime/scripts/realtime_tts_demo.py \
  --text "realtime health check" \
  --strict

Output location

  • Default output: output/ai-audio-tts-realtime/audio/
  • Override base dir with OUTPUT_DIR.

Validation

mkdir -p output/alicloud-ai-audio-tts-realtime
for f in skills/ai/audio/alicloud-ai-audio-tts-realtime/scripts/*.py; do
  python3 -m py_compile "$f"
done
echo "py_compile_ok" > output/alicloud-ai-audio-tts-realtime/validate.txt

Pass criteria: command exits 0 and output/alicloud-ai-audio-tts-realtime/validate.txt is generated.

Output And Evidence

  • Save artifacts, command outputs, and API response summaries under output/alicloud-ai-audio-tts-realtime/.
  • Include key parameters (region/resource id/time range) in evidence files for reproducibility.

Workflow

1) Confirm user intent, region, identifiers, and whether the operation is read-only or mutating. 2) Run one minimal read-only query first to verify connectivity and permissions. 3) Execute the target operation with explicit parameters and bounded scope. 4) Verify results and save output/evidence files.

References

  • references/sources.md

Related skills

How it compares

Pick alicloud-ai-audio-tts-realtime for websocket streaming voice agents; use aliyun-qwen-tts for batch non-realtime speech generation.

FAQ

Which Qwen models does alicloud-ai-audio-tts-realtime support?

alicloud-ai-audio-tts-realtime lists five exact model strings: qwen3-tts-flash-realtime, qwen3-tts-instruct-flash-realtime, qwen3-tts-instruct-flash-realtime-2026-01-22, qwen3-tts-vd-realtime-2026-01-15, and qwen3-tts-vc-realtime-2026-01-15.

How do you authenticate alicloud-ai-audio-tts-realtime calls?

alicloud-ai-audio-tts-realtime requires the dashscope Python SDK with DASHSCOPE_API_KEY set as an environment variable or dashscope_api_key in ~/.alibabacloud/credentials.

How do you validate alicloud-ai-audio-tts-realtime locally?

alicloud-ai-audio-tts-realtime provides realtime_tts_demo.py to probe websocket streaming compatibility, with --strict for CI gating and py_compile validation of bundled scripts.

Generative Mediallmagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.