PluginWorld
Ds

dsh-vision

dsh✓ SPEC VERIFIED

Near-native image understanding for DeepSeek Harness

@oil-oil · v0.1.2 · MIT · updated 6d ago

SECURITY

A

SCORE

88

STARS

87

PLUG IN

npx @deepseek-ai/dsh web

Launch, then search "dsh-vision" in the built-in market to plug it in

README

dsh-vision: native vision passthrough and a vision bridge for DeepSeek Harness

English | 中文

CI MIT License DeepSeek Harness

dsh-vision is a plugin for DeepSeek Harness. Vision-capable models keep receiving images natively. When the selected main model is text-only, the plugin asks a separate vision model to observe the original images, then lets the original DeepSeek model produce the final answer.

How it works

Main model Image path Final answer
Supports images Original images are sent directly, without preprocessing or OCR Current model
deepseek-official or another text-only model A configured vision model observes the original images; its output is injected as untrusted attachment context DeepSeek
Cloud vision unavailable Falls back to macOS Vision or Tesseract DeepSeek

The plugin does not replace the main model selected in Harness. Multiple image attachments are analyzed together, so comparisons and combined evidence work naturally. The user's task is forwarded unchanged instead of being wrapped in a fixed report template.

Install

Use the plugin manager built into DeepSeek Harness:

npx @deepseek-ai/dsh plugin --profile web add github:oil-oil/dsh-vision

Restart Harness, then paste or drag images into the composer as usual. The plugin replaces the official deepseek-official adapter while preserving its model catalog, settings, and credentials. It also adds a Vision Recognition card to Settings → Plugins → Plugin configuration.

DeepSeek Harness is still in Developer Preview. This release supports 0.1.0-rc.6 and 0.1.0-rc.7; its settings-card registration satisfies both the legacy list Slot and the current keyed Slot without relying on private runtime inspection.

Configure Vision Recognition

Open Settings → Plugins → Plugin configuration → Vision Recognition. Select ZenMux, Alibaba Cloud Model Studio, TokenDance, or OpenRouter, then enter its API key. The same card lets you change the model ID, API endpoint, and image limit.

The API key is stored through Harness's official credential service. It is write-only in the browser: the plugin can report whether a key exists, but never reads it back into the page, chat, settings document, or session log.

Routing follows the user's choice. A provider selected in Vision Recognition is primary for text-only models. Other enabled Harness vision routes, an existing see configuration, and local OCR are failover only. When the current main model supports images, the original images pass through natively and none of these bridge routes are used.

Choose Automatic to skip plugin-managed cloud credentials. The bridge then tries image-capable models already configured in Harness, followed by see-compatible private configuration and local OCR. A Harness custom model must declare image as an input modality or it remains a text model.

Advanced file configuration

Most setups should use the UI. The equivalent non-secret fields live in the existing llm-deepseek section of $DSH_HOME/settings.yaml:

llm-deepseek:
  visionBackend: zenmux
  visionBackendModel: qwen/qwen3.7-plus
  visionBackendBaseURL: https://zenmux.ai/api/v1
  maxImages: 8

Do not put API keys in this file. Save them in the Vision Recognition card or provide the matching environment variable. Changes apply without a restart.

see-skill compatibility

If Harness has no usable vision model, the plugin also reads ~/.config/see/config.env. It supports ZenMux, Alibaba Cloud Model Studio, OpenRouter, and TokenDance. Environment variables override the private config file.

export SEE_PROVIDER=zenmux
export ZENMUX_API_KEY=your-key

SEE_PROVIDER selects the primary provider. Other providers with configured keys are failover routes only. If no provider is selected and only one is configured, that provider is used.

When no cloud key is available, or every cloud route fails, the plugin tries local capabilities:

  • macOS: built-in Vision OCR, with no extra dependency.
  • Linux / Windows: Tesseract with the required language data installed.

Local fallback is primarily OCR and is not equivalent to full multimodal understanding.

Security boundary

  • Original images are sent only to vision services configured by the user.
  • Vision output is marked as untrusted observation data; instructions inside an image receive no system authority.
  • Generated vision context affects only the current model request and does not rewrite message history.
  • API keys are resolved through Harness credentials or the user's private see config and are never written to this repository.

Development

pnpm install
pnpm check

The project is available under the MIT License. Cloud routing, joint multi-image analysis, and local fallback behavior are based on the MIT-licensed oil-oil/see-skill. The DeepSeek icon comes from the official deepseek-ai/deepseek-harness repository.

SIMILAR PLUGINS

Mo

modlens

v3.24.2 · MIT

95

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

dshA
110K installs
Ds

dsh-vision-router

v1.7.7 · MIT

95

Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

dshA
39.6K installs
Ds

dsh-vision-toolkit

v0.1.39 · MIT

94

[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

dshA
30K installs
Ds

dsh-plugin-wallpaper-engine

v0.6.2 · MIT

91

把本机 Wallpaper Engine 的壁纸变成 DSH 网页界面的背景:Video 动态播放、Web 以 iframe 加载、Scene 壁纸提取主纹理作为静态帧;iOS 液态玻璃设置窗口(配色 / 玻璃颜色 / 透明度)、内容分级与类型过滤、自定义壁纸上传、紧凑 CD 架布局、黑胶唱片展示、隐藏 / 恢复、倍速 / 翻转与自动轮播。感谢 Jerry 维护 macOS 版。

dshA
5.1K installs