
Aliyun Pixverse Generation
- 53 installs
- 396 repo stars
- Updated July 18, 2026
- cinience/alicloud-skills
Generate videos with Alibaba Cloud Model Studio PixVerse v5.6 models for text-to-video, first-frame, keyframe, and reference-to-video workflows.
About
Uses Model Studio PixVerse v5.6 models to build non-Wan text-to-video, image-to-video, keyframe, and reference-to-video generation. A developer uses it when they explicitly want the PixVerse family for video generation.
- Four PixVerse v5.6 model variants (t2v/it2v/kf2v/r2v)
- Records resolution, duration, and audio-generation flags
Aliyun Pixverse Generation by the numbers
- 53 all-time installs (skills.sh)
- Ranked #879 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/cinience/alicloud-skills --skill aliyun-pixverse-generationAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 53 |
|---|---|
| repo stars | ★ 396 |
| Last updated | July 18, 2026 |
| Repository | cinience/alicloud-skills ↗ |
What it does
Generate videos with Alibaba Cloud Model Studio PixVerse v5.6 models for text-to-video, first-frame, keyframe, and reference-to-video workflows.
Files
Category: provider
Model Studio Aishi Video Generation
Validation
mkdir -p output/aliyun-pixverse-generation
python -m py_compile skills/ai/video/aliyun-pixverse-generation/scripts/prepare_aishi_request.py && echo "py_compile_ok" > output/aliyun-pixverse-generation/validate.txtPass criteria: command exits 0 and output/aliyun-pixverse-generation/validate.txt is generated.
Output And Evidence
- Save normalized request payloads, chosen model variant, and task polling snapshots under
output/aliyun-pixverse-generation/. - Record region, resolution/size, duration, and whether audio generation was enabled.
Use Aishi when the user explicitly wants the non-Wan PixVerse family for video generation.
Critical model names
Use one of these exact model strings:
pixverse/pixverse-v5.6-t2vpixverse/pixverse-v5.6-it2vpixverse/pixverse-v5.6-kf2vpixverse/pixverse-v5.6-r2v
Selection guidance:
- Use
pixverse/pixverse-v5.6-t2vfor text-only generation. - Use
pixverse/pixverse-v5.6-it2vfor first-frame image-to-video. - Use
pixverse/pixverse-v5.6-kf2vfor first-frame + last-frame transitions. - Use
pixverse/pixverse-v5.6-r2vfor multi-image character/style consistency.
Prerequisites
- This family currently only supports China mainland (Beijing).
- Install SDK or call HTTP directly:
python3 -m venv .venv
. .venv/bin/activate
python -m pip install dashscope- Set
DASHSCOPE_API_KEYin your environment, or adddashscope_api_keyto~/.alibabacloud/credentials.
Normalized interface (video.generate)
Request
model(string, required)prompt(string, optional forit2v, required for other variants)media(array<object>, optional)size(string, optional): direct pixel size such as1280*720, used byt2vandr2vresolution(string, optional):360P/540P/720P/1080P, used byit2vandkf2vduration(int, required):5/8/10, except 1080P only supports5/8audio(bool, optional)watermark(bool, optional)seed(int, optional)
Response
task_id(string)task_status(string)video_url(string, when finished)
Endpoint and execution model
- Submit task:
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/video-generation/video-synthesis - Poll task:
GET https://dashscope.aliyuncs.com/api/v1/tasks/{task_id} - HTTP calls are async only and must set header
X-DashScope-Async: enable.
Quick start
Text-to-video:
python skills/ai/video/aliyun-pixverse-generation/scripts/prepare_aishi_request.py \
--model pixverse/pixverse-v5.6-t2v \
--prompt "A compact robot walks through a rainy neon alley." \
--size 1280*720 \
--duration 5Image-to-video:
python skills/ai/video/aliyun-pixverse-generation/scripts/prepare_aishi_request.py \
--model pixverse/pixverse-v5.6-it2v \
--prompt "The turtle swims slowly as the camera rises." \
--media image_url=https://example.com/turtle.webp \
--resolution 720P \
--duration 5Operational guidance
t2vandr2vusesize;it2vandkf2vuseresolution.- For
kf2v, provide exactly onefirst_frameand onelast_frame. - For
r2v, you can pass up to 7 reference images. - Aishi returns task IDs first; do not treat the initial response as the final video result.
Output location
- Default output:
output/aliyun-pixverse-generation/request.json - Override base dir with
OUTPUT_DIR.
References
references/sources.md
interface:
display_name: "Alibaba Cloud AI Video Aishi Generation"
short_description: "Non-Wan video generation with PixVerse models"
default_prompt: "Use $aliyun-pixverse-generation to complete this ai/video PixVerse task on Alibaba Cloud."
- 视频生成总览(爱诗条目): https://help.aliyun.com/zh/model-studio/use-video-generation
- 爱诗文生视频 API 参考: https://help.aliyun.com/zh/model-studio/pixverse-text-to-video-api-reference
- 爱诗图生视频-基于首帧 API 参考: https://help.aliyun.com/zh/model-studio/pixverse-image-to-video-api-reference
- 爱诗图生视频-基于首尾帧 API 参考: https://help.aliyun.com/zh/model-studio/pixverse-keyframe-to-video-api-reference
- 爱诗参考生视频 API 参考: https://help.aliyun.com/zh/model-studio/pixverse-reference-to-video-api-reference
#!/usr/bin/env python3
"""Prepare a normalized request for Model Studio Aishi (PixVerse) video generation."""
from __future__ import annotations
import argparse
import json
from pathlib import Path
def parse_media(values: list[str]) -> list[dict[str, str]]:
items: list[dict[str, str]] = []
for value in values:
if "=" not in value:
raise ValueError(f"Invalid media value: {value}. Expected type=url format.")
media_type, url = value.split("=", 1)
items.append({"type": media_type, "url": url})
return items
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("--model", required=True)
parser.add_argument("--prompt")
parser.add_argument("--media", action="append", default=[])
parser.add_argument("--size")
parser.add_argument("--resolution")
parser.add_argument("--duration", type=int, required=True)
parser.add_argument("--audio", action="store_true")
parser.add_argument("--watermark", action="store_true")
parser.add_argument("--seed", type=int)
parser.add_argument("--output", default="output/aliyun-pixverse-generation/request.json")
args = parser.parse_args()
payload: dict[str, object] = {
"model": args.model,
"input": {},
"parameters": {
"duration": args.duration,
"audio": args.audio,
"watermark": args.watermark,
},
}
if args.prompt:
payload["input"]["prompt"] = args.prompt
media = parse_media(args.media)
if media:
payload["input"]["media"] = media
if args.size:
payload["parameters"]["size"] = args.size
if args.resolution:
payload["parameters"]["resolution"] = args.resolution
if args.seed is not None:
payload["parameters"]["seed"] = args.seed
output = Path(args.output)
output.parent.mkdir(parents=True, exist_ok=True)
output.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
print(json.dumps({"ok": True, "request_path": str(output)}, ensure_ascii=False))
if __name__ == "__main__":
main()