— · MIT
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
v3.24.2 · MIT
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
v0.1.39 · MIT
[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
v0.51.57 · MIT
Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in natural language on ANY LLM (Claude, ChatGPT, Gemini, offline Ollama, or any hosted model). 178 tools, 36 AI skills, 55 installer packs. Local, LAN, VPS, or Comfy Cloud.
v0.1.0-rc.6 · MIT
DeepSeek Harness (DSH) plugin: dispatch work to DSH agents from Claude Code / Codex — native subagent progress, in-host worker sessions with per-tier presets, and a multimodal bridge that lends the text-only harness vision and image generation.
— · MIT
A Swift library for macOS automation — mouse, keyboard, screenshots, image recognition, and AI-powered agents.
v0.0.0 · MIT
Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file.
— · Apache-2.0
The fastest way to put Volcengine Ark in your terminal and your AI agent — go from prompt to generated media, multimodal answer, or deployed endpoint in a single command, no API glue code.
v0.4.0 · MIT License
Enables MCP clients to run project-scoped video editing workflows: propose and approve editing strategies, apply validated plans, review immutable versions, and export final renders via FFmpeg, with durable persistence and approval gates.
— · no license
Atlas Cloud skills for Claude Code, Codex & Gemini CLI — generate images/videos and call 300+ AI models from your coding agent.
— · no license
A collection of open-source Agent Skills for content creation — images, audio, and video.
v0.1.1 · no license
Turn designs into shipped parts: quote 3D printing, CNC, and decals, then check out.
v0.3.0 · MIT
Open source implementation and extension of Google Research’s PaperBanana for automated academic figures, diagrams, and research visuals, expanded to new domains like slide generation.
v0.4.0 · MIT
Clip any video into a narration recap with claude code skill|用 claude code skill 把任何视频剪辑成中文解说视频,支持剪映导出
v0.4.17 · MIT
你的随身 AI 剪辑搭档 | Pocket AI co-editor for video montage — AI video editing plugin & MCP server for Claude Code, Codex, Hermes & OpenCode
v1.0.0 · MIT
Recreate any web animation from a screen recording. A motion-design plugin: analyzes motion (timing, easing, stagger, loops) and rebuilds it as GSAP / CSS / Framer Motion / Lottie code.
v0.1.7 · MIT
Generate images directly in DeepSeek Harness chats
v1.0.0 · MIT
Prompt as Code | GPT-Image2 工业级提示词引擎与模板库,530+ 个案例逆向工程,20+ 套工业级模板,并提炼出Skills,持续更新中
v0.13.0 · MIT
MCP server for AI image generation
v0.0.10 · AGPL-3.0
Streaming music player that finds free music for you
v3.24.2 · MIT
Plug-in vision for text-only LLMs, powered by the free Antigravity CLI
v0.3.4 · AGPL-3.0
TongFlow — Multimodal GenAI Studio
v0.1.142 · MIT
AI ad studio and marketing MCP server with 681 tools. Research the ads already running in any market, generate finished image, video and UGC avatar ads, publish and schedule them to your own channels, build and manage the ad campaigns behind them, and rea
v0.4.0 · MIT
Give text-only models eyes: an analyze_image tool for DeepSeek Harness, backed by free Chinese vision APIs (GLM-4V-Flash / Qwen-VL) or any OpenAI-compatible vision endpoint. 给纯文本模型装上眼睛。