Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
agents365-ai avatar

Deepseek Vision

  • 2 installs
  • 58 repo stars
  • Updated July 31, 2026
  • agents365-ai/dsclaude

deepseek-vision is a Claude Code skill that calls a DashScope vision model to return a text description of an image so an agent can reason over it.

About

deepseek-vision is a Claude Code skill that calls a vision model through DashScope and returns a text description of an image the agent can then reason over. A developer uses it when the user references a screenshot, photo, diagram, or error dialog by file path or URL and the agent needs to know what is in it. It is especially useful on text-only backends, and also serves as a dedicated OCR or detail extractor.

  • Sends a local image path or http/https URL to a vision model and returns a text description to reason over
  • Adds vision/OCR to text-only backends like DeepSeek V4 via DashScope (Qwen vision by default)
  • Cannot read images pasted inline in chat; needs a filesystem path or URL

Deepseek Vision by the numbers

  • 2 all-time installs (skills.sh)
  • Ranked #13,958 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
At a glance

deepseek-vision capabilities & compatibility

Requires a DASHSCOPE_API_KEY (Alibaba Cloud Bailian); usage billed by DashScope.

Capabilities
image analysis · ocr · vision model call
Works with
openai
Use cases
image generation
Pricing
Bring your own API key
From the docs

What deepseek-vision says it does

The tool sends the image to a vision model and returns a text description.
SKILL.md
export DASHSCOPE_API_KEY=sk-xxxxxxxxxxxxxxxxxx # add to ~/.zshrc
SKILL.md
npx skills add https://github.com/agents365-ai/dsclaude --skill deepseek-vision

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs2
repo stars58
Last updatedJuly 31, 2026
Repositoryagents365-ai/dsclaude

What it does

Turn an image at a file path or URL into a text description via a DashScope vision model so a text-only agent can reason over it.

Who is it for?

Describing or OCR-ing an image at a local file path or http/https URL, especially on text-only backends.

Skip if: Reading images dropped, pasted, or attached inline in chat with no filesystem path.

When should I use this skill?

The user references an image by local file path or http/https URL and you need to know what is in it to answer or act.

What you get

A plaintext description of the image the agent can read as if it saw it.

  • A plaintext description of the supplied image

By the numbers

  • 3 configurable vision models (Qwen3.6-Flash/Plus, Qwen3-VL-Plus)

Files

SKILL.mdMarkdownGitHub ↗

Deepseek Vision Helper

When the user shares an image — file path or URL — and you need to understand its content, call this skill's analyze-image tool instead of trying to "see" it directly. The tool sends the image to a vision model and returns a text description.

How to use

./skills/deepseek-vision/analyze-image <image-path-or-url> [focus prompt]

Accepts:

  • A local file path (/Users/me/screenshot.png, ~/Desktop/diagram.jpg, relative paths)
  • An http/https URL — passed through directly to the API, no download needed

The script prints plaintext. Read it as if you saw the image yourself, then continue your reasoning. On error it prints to stderr and exits non-zero.

What this skill cannot do

Read inline images dropped, pasted, or attached into the chat. When the user drops or pastes an image, or uses Claude Desktop's "+ → Add files or photos" attachment, the image becomes an image_url content block embedded in the message — there is no filesystem path the bash script can reach, and on a text-only backend it usually surfaces as [Unsupported Image].

For those cases, use the `dsvision-mcp` server (sibling tool in this repo) instead — it runs outside Cowork's sandbox and auto-finds Claude Code's ~/.claude/image-cache/ directory.

If only this skill is available and the user shares an image inline with no path:

1. Save the image to disk and re-share with the file path, or 2. Paste an http/https URL to the image

Do not try to invent a path or guess at the image content — say what you can't see, then propose one of the two options above.

Setup (one-time)

export DASHSCOPE_API_KEY=sk-xxxxxxxxxxxxxxxxxx   # add to ~/.zshrc

Get a key at https://bailian.console.aliyun.com/.

Optional env overrides

VariableDefaultPurpose
DSVISION_MODELqwen3.6-flashswap to qwen3.6-plus for higher quality, qwen3-vl-plus for cheaper
DSVISION_BASE_URLhttps://dashscope.aliyuncs.com/compatible-mode/v1self-hosted or alternative provider

Example

User: "What's the error in /Users/me/screenshot.png?"

You:

./skills/deepseek-vision/analyze-image /Users/me/screenshot.png "What error message is shown?"

Tool output: TypeError: cannot read property 'foo' of undefined at line 42 in app.js

You: The screenshot shows a TypeError on line 42 — foo is being accessed on an undefined value. Looking at app.js:42...

Related skills

FAQ

Can deepseek-vision read images pasted into chat?

No. It needs a local file path or an http/https URL; inline-pasted images have no filesystem path the script can reach.

What model does it use?

A vision model via DashScope, Qwen3.6-Flash by default, overridable with DSVISION_MODEL.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.