Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
huaweicloud avatar

Huawei Cloud Msmodelslim Model Adapt

  • 50 installs
  • 19 repo stars
  • Updated July 31, 2026
  • huaweicloud/huaweicloud-skills

Create msModelSlim adapters for Transformers models to run W8A8/W4A16 quantization, with a four-step generate-quantize-verify workflow.

About

Guides building basic Transformers model adapters for msModelSlim to enable W8A8/W4A16 quantization, via a four-step generate, fallback-quantize, weight-verify, and validation workflow. A developer uses it to make new decoder-only LLMs or VLM text backbones quantizable.

  • Implements required interfaces for new-model quantization
  • Four-step verification: generate to quantize to verify to validate

Huawei Cloud Msmodelslim Model Adapt by the numbers

  • 50 all-time installs (skills.sh)
  • +4 installs in the week ending Aug 2, 2026 (Skillselion tracking)
  • Ranked #930 of 2,101 Data Science & ML skills by installs in the Skillselion catalog
  • Data as of Aug 2, 2026 (Skillselion catalog sync)
npx skills add https://github.com/huaweicloud/huaweicloud-skills --skill huawei-cloud-msmodelslim-model-adapt

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs50
repo stars19
Last updatedJuly 31, 2026
Repositoryhuaweicloud/huaweicloud-skills

What it does

Create msModelSlim adapters for Transformers models to run W8A8/W4A16 quantization, with a four-step generate-quantize-verify workflow.

Files

SKILL.mdMarkdownGitHub ↗

Huawei Cloud msModelSlim Model Adapter

Overview

This skill guides how to create basic adapters for new models to run W8A8/W4A16 quantization workflows in msModelSlim.

Architecture: Model Analysis -> Adapter Creation -> Registration -> Verification (4 Steps)

Related Skills:

  • huawei-cloud-msmodelslim-model-analysis - Model structure analysis

before adapter implementation

  • huawei-cloud-ascend-profiler-db-explorer - Optional: Performance

analysis after deployment

Scope

Supported:

  • Decoder-only LLM
  • Understanding VLM (text/LLM backbone only)

Not supported:

  • Multimodal generation (Stable Diffusion/Flux/Wan)
  • Encoder-only models
  • Non-Transformers architectures

Architecture

┌─────────────────────────────────────────────────────────────┐
│              msModelSlim Model Adapter Skill                │
├─────────────────────────────────────────────────────────────┤
│  ┌──────────────────┐    ┌──────────────────────────────┐  │
│  │  Model Analysis │───▶│    Adapter Creation          │  │
│  │  - config.json  │    │    - LLM Adapter Template     │  │
│  │  - modeling_*.py│    │    - VLM Adapter Template     │  │
│  └──────────────────┘    │    - Required Interfaces     │  │
│                          └──────────────────────────────┘  │
│                                    │                       │
│                                    ▼                       │
│                          ┌──────────────────┐             │
│                          │  Registration    │             │
│                          │  & Installation  │             │
│                          └──────────────────┘             │
│                                    │                       │
│                                    ▼                       │
│  ┌──────────────────────────────────────────────────────┐ │
│  │                   Verification (4 Steps)              │ │
│  │  1. Generate Test Model → 2. Full Fallback Quant     │ │
│  │  3. Weight Verification → 4. Quant Description     │ │
│  └──────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘

Architecture Components

This skill involves the following cloud services and components:

  • msModelSlim: Huawei Cloud's model quantization framework for

efficient model compression

  • Transformers Library: Hugging Face Transformers for model loading

and processing

  • ModelScope: Model download and management platform
  • Ascend NPU: Target hardware for quantized model deployment

Use Cases

Typical Problem Scenarios:

  • Need to deploy LLM models with reduced memory footprint on Ascend NPU
  • Want to optimize inference speed without significant accuracy loss
  • Migrating models that don't have built-in msModelSlim support
  • Need W8A8/W4A16 quantization for decoder-only LLM or VLM text backbones

Typical User Phrases:

  • "How to quantize my custom LLM model for Ascend?"
  • "Create msModelSlim adapter for Qwen model"
  • "Implement W4A16 quantization workflow"
  • "Adapt my VLM text backbone for quantization"
  • "How to add quantization support for new models?"

Core Workflow

1. Preparation

  • Download Model: Recommended to use modelscope download for

non-weight files.

  • Example: `modelscope download --model <org>/<model> --local_dir

./models/<name> --exclude '*.safetensors'`

  • Analyze Model: Read config.json and modeling_*.py to confirm

structure and implementation.

  • See: Model Analysis Guide

2. Create Adapter

  • Use Templates:
  • LLM: assets/model_adapter_template.py
  • VLM: assets/vlm_model_adapter_template.py
  • Implement Interfaces: Implement handle_dataset, init_model,

generate_model_visit, generate_model_forward, enable_kv_cache.

  • Key Principles:
  • visit and forward must be strictly consistent.
  • MoE models recommended to unpack to pure linear layers.
  • See: Implementation Guide

3. Registration & Installation

  • Register model and entry in config/config.ini, then execute

bash install.sh.

  • See: Registration Guide

4. Verify Adapter (Required)

  • Must execute four-step verification: Generate test model -> Full

fallback quantization -> Verify full fallback model matches float weights exactly and can load/save completely -> Verify actual quantization workflow works (including description file rule validation).

  • See: Verification Guide

Common Scripts

Scripts located in scripts/ directory:

  • scripts/step1_generate_test_model.py
  • scripts/step2_run_quantization.py
  • scripts/step3_verify_weights.py
  • scripts/step4_verify_quant_description.py

Prerequisites

System Requirements

  • Python 3.8+
  • transformers >= 4.40.0
  • msmodelslim >= 1.0.0

Environment Check

Prerequisite check: Python3 + transformers + msmodelslim required

>

```bash
python3 --version # Python3 >= 3.8
python3 -c "import transformers; print('OK')" # Transformers library
python3 -c "import msmodelslim; print('OK')" # msModelSlim library
```

>

If not installed: pip3 install --user transformers msmodelslim

Reference Documents

DocumentDescription
Model Analysis GuideModel structure analysis guide
Implementation GuideAdapter implementation instructions
Registration GuideRegistration and installation guide
Verification GuideFour-step verification workflow
Interface ChecklistRequired interface implementation checklist
Core WorkflowCore workflow documentation
Acceptance CriteriaFunctional acceptance criteria
TroubleshootingCommon issues and solutions

Requirements

  • transformers >= 4.40.0 installed
  • msmodelslim >= 1.0.0 installed
  • Transformers model to be adapted
  • Understanding of target quantization scheme (W8A8/W4A16)

Core Commands

# Create model adapter
python3 scripts/create_adapter.py \
  --model Qwen2-7B \
  --quantization W8A8

# Run four-step verification
python3 scripts/verify_adapter.py --adapter ./adapter.py

Parameter Confirmation

ParameterDescriptionRequired
modelModel name or pathYes
quantizationQuantization scheme (W8A8/W4A16)Yes
outputAdapter output pathNo

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.