
Deepseek Vision
- 2 installs
- 58 repo stars
- Updated July 31, 2026
- agents365-ai/dsclaude
deepseek-vision is a Claude Code skill that calls a DashScope vision model to return a text description of an image so an agent can reason over it.
About
deepseek-vision is a Claude Code skill that calls a vision model through DashScope and returns a text description of an image the agent can then reason over. A developer uses it when the user references a screenshot, photo, diagram, or error dialog by file path or URL and the agent needs to know what is in it. It is especially useful on text-only backends, and also serves as a dedicated OCR or detail extractor.
- Sends a local image path or http/https URL to a vision model and returns a text description to reason over
- Adds vision/OCR to text-only backends like DeepSeek V4 via DashScope (Qwen vision by default)
- Cannot read images pasted inline in chat; needs a filesystem path or URL
Deepseek Vision by the numbers
- 2 all-time installs (skills.sh)
- Ranked #13,958 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
deepseek-vision capabilities & compatibility
Requires a DASHSCOPE_API_KEY (Alibaba Cloud Bailian); usage billed by DashScope.
- Capabilities
- image analysis · ocr · vision model call
- Works with
- openai
- Use cases
- image generation
- Pricing
- Bring your own API key
What deepseek-vision says it does
The tool sends the image to a vision model and returns a text description.
export DASHSCOPE_API_KEY=sk-xxxxxxxxxxxxxxxxxx # add to ~/.zshrc
npx skills add https://github.com/agents365-ai/dsclaude --skill deepseek-visionAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 2 |
|---|---|
| repo stars | ★ 58 |
| Last updated | July 31, 2026 |
| Repository | agents365-ai/dsclaude ↗ |
What it does
Turn an image at a file path or URL into a text description via a DashScope vision model so a text-only agent can reason over it.
Who is it for?
Describing or OCR-ing an image at a local file path or http/https URL, especially on text-only backends.
Skip if: Reading images dropped, pasted, or attached inline in chat with no filesystem path.
When should I use this skill?
The user references an image by local file path or http/https URL and you need to know what is in it to answer or act.
What you get
A plaintext description of the image the agent can read as if it saw it.
- A plaintext description of the supplied image
By the numbers
- 3 configurable vision models (Qwen3.6-Flash/Plus, Qwen3-VL-Plus)
Files
Deepseek Vision Helper
When the user shares an image — file path or URL — and you need to understand its content, call this skill's analyze-image tool instead of trying to "see" it directly. The tool sends the image to a vision model and returns a text description.
How to use
./skills/deepseek-vision/analyze-image <image-path-or-url> [focus prompt]Accepts:
- A local file path (
/Users/me/screenshot.png,~/Desktop/diagram.jpg, relative paths) - An http/https URL — passed through directly to the API, no download needed
The script prints plaintext. Read it as if you saw the image yourself, then continue your reasoning. On error it prints to stderr and exits non-zero.
What this skill cannot do
Read inline images dropped, pasted, or attached into the chat. When the user drops or pastes an image, or uses Claude Desktop's "+ → Add files or photos" attachment, the image becomes an image_url content block embedded in the message — there is no filesystem path the bash script can reach, and on a text-only backend it usually surfaces as [Unsupported Image].
For those cases, use the `dsvision-mcp` server (sibling tool in this repo) instead — it runs outside Cowork's sandbox and auto-finds Claude Code's ~/.claude/image-cache/ directory.
If only this skill is available and the user shares an image inline with no path:
1. Save the image to disk and re-share with the file path, or 2. Paste an http/https URL to the image
Do not try to invent a path or guess at the image content — say what you can't see, then propose one of the two options above.
Setup (one-time)
export DASHSCOPE_API_KEY=sk-xxxxxxxxxxxxxxxxxx # add to ~/.zshrcGet a key at https://bailian.console.aliyun.com/.
Optional env overrides
| Variable | Default | Purpose |
|---|---|---|
DSVISION_MODEL | qwen3.6-flash | swap to qwen3.6-plus for higher quality, qwen3-vl-plus for cheaper |
DSVISION_BASE_URL | https://dashscope.aliyuncs.com/compatible-mode/v1 | self-hosted or alternative provider |
Example
User: "What's the error in /Users/me/screenshot.png?"
You:
./skills/deepseek-vision/analyze-image /Users/me/screenshot.png "What error message is shown?"Tool output: TypeError: cannot read property 'foo' of undefined at line 42 in app.js
You: The screenshot shows a TypeError on line 42 — foo is being accessed on an undefined value. Looking at app.js:42...
#!/usr/bin/env bash
# analyze-image — describe an image via Qwen3.6-Flash for text-only models.
# Usage: analyze-image <path-or-url> [focus prompt]
# <path-or-url> is either a local file path or an http(s) URL.
# Requires: DASHSCOPE_API_KEY in env or shell rc. Returns plaintext on stdout.
#
# For inline-attached / pasted / clipboard images in Cowork or Claude Desktop,
# use dsvision-mcp instead — it auto-finds Claude Code's image-cache and
# bypasses Cowork's egress sandbox.
set -euo pipefail
INPUT="${1:-}"
PROMPT="${2:-Describe this image in detail. Focus on text content, UI elements, code snippets, charts, error messages — anything a coding/agent task might care about.}"
[[ -n "$INPUT" ]] || { echo "analyze-image: missing image path or URL" >&2; exit 1; }
# Resolve DASHSCOPE_API_KEY: env first, then shell rc files. Claude Code's
# subprocesses don't always inherit interactive-shell exports, so we look in
# the rc files directly as a fallback.
if [[ -z "${DASHSCOPE_API_KEY:-}" ]]; then
for rc in "$HOME/.zshrc" "$HOME/.bashrc" "$HOME/.bash_profile" "$HOME/.profile"; do
[[ -r "$rc" ]] || continue
found=$(grep -E '^[[:space:]]*export[[:space:]]+DASHSCOPE_API_KEY=' "$rc" 2>/dev/null \
| tail -1 \
| sed -E 's/^[^=]*=//; s/^"(.*)"$/\1/; s/^'\''(.*)'\''$/\1/' \
|| true)
[[ -n "$found" ]] && { export DASHSCOPE_API_KEY="$found"; break; }
done
fi
[[ -n "${DASHSCOPE_API_KEY:-}" ]] || {
echo "analyze-image: DASHSCOPE_API_KEY not set in env or shell rc files (~/.zshrc, ~/.bashrc, ...). Add 'export DASHSCOPE_API_KEY=sk-...' and retry." >&2
exit 1
}
# URL input → pass through directly (DashScope accepts public http/https URLs).
if [[ "$INPUT" =~ ^https?:// ]]; then
IMAGE_URL="$INPUT"
else
[[ -f "$INPUT" ]] || { echo "analyze-image: image not found: $INPUT" >&2; exit 1; }
# DashScope rejects payloads beyond ~10MB; fail early with a clear message.
SIZE=$(wc -c < "$INPUT" | tr -d ' ')
if (( SIZE > 10485760 )); then
echo "analyze-image: image too large ($SIZE bytes; limit 10485760). Resize and retry." >&2
exit 1
fi
case "$(echo "${INPUT##*.}" | tr 'A-Z' 'a-z')" in
jpg|jpeg) MIME=image/jpeg ;;
png) MIME=image/png ;;
webp) MIME=image/webp ;;
gif) MIME=image/gif ;;
*) MIME=$(file --mime-type -b "$INPUT") ;;
esac
B64=$(base64 < "$INPUT" | tr -d '\n')
IMAGE_URL="data:$MIME;base64,$B64"
fi
PAYLOAD=$(jq -n \
--arg model "${DSVISION_MODEL:-qwen3.6-flash}" \
--arg url "$IMAGE_URL" \
--arg prompt "$PROMPT" \
'{
model: $model,
messages: [{
role: "user",
content: [
{ type: "image_url", image_url: { url: $url } },
{ type: "text", text: $prompt }
]
}]
}')
BASE_URL="${DSVISION_BASE_URL:-https://dashscope.aliyuncs.com/compatible-mode/v1}"
RESP=$(curl -sS -X POST "$BASE_URL/chat/completions" \
--connect-timeout 10 \
--max-time 60 \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
--data-raw "$PAYLOAD")
if echo "$RESP" | jq -e '.error' >/dev/null 2>&1; then
echo "analyze-image: API error" >&2
echo "$RESP" | jq '.error' >&2
exit 1
fi
CONTENT=$(echo "$RESP" | jq -r '.choices[0].message.content // empty')
if [[ -z "$CONTENT" ]]; then
echo "analyze-image: empty response from model" >&2
echo "$RESP" | jq '.' >&2
exit 1
fi
echo "$CONTENT"
Related skills
FAQ
Can deepseek-vision read images pasted into chat?
No. It needs a local file path or an http/https URL; inline-pasted images have no filesystem path the script can reach.
What model does it use?
A vision model via DashScope, Qwen3.6-Flash by default, overridable with DSVISION_MODEL.