— · MIT
An MCP server providing tools for image processing operations
v0.13.0 · MIT
MCP server for AI image generation
v0.51.57 · MIT
Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in natural language on ANY LLM (Claude, ChatGPT, Gemini, offline Ollama, or any hosted model). 178 tools, 36 AI skills, 55 installer packs. Local, LAN, VPS, or Comfy Cloud.
v1.4.0 · MIT
Supports GPT Image 2, Seedance & ComfyUI, with a 1,400+ prompt library, carefully crafted hooks and a multi-task orchestration system
v2.1.4 · MIT
The measurable image + document toolkit for JavaScript. On flat art -- logos, icons, UI, screenshots, pixel art -- the SVG is bit-exact: SSIM 1.0000, zero differing pixels, verified by rendering it back. On photographs it leads potrace, imagetracerjs and
— · MIT
Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.
v0.3.0 · MIT
Open source implementation and extension of Google Research’s PaperBanana for automated academic figures, diagrams, and research visuals, expanded to new domains like slide generation.
v0.0.0 · Apache-2.0
Official Model Studio CLI(阿里云百炼 CLI)built for AI Agent frameworks, exposing models, search, multimodal, and workflow capabilities as structured tool calls.
v0.2.43 · MIT
The canonical 12ui CLI, SDK, MCP server, and visual interface design skills
— · Apache-2.0
lightweight Python-based MCP (Model Context Protocol) server for local ComfyUI
v1.25.0 · MIT
MCP server giving AI assistants 17 weather tools with zero API keys: global forecasts, current conditions, alerts, air quality, marine conditions, lightning detection, radar imagery, river levels, wildfire tracking, and historical weather back to 1940. Bu
v0.1.142 · MIT
AI ad studio and marketing MCP server with 681 tools. Research the ads already running in any market, generate finished image, video and UGC avatar ads, publish and schedule them to your own channels, build and manage the ad campaigns behind them, and rea
v1.0.0 · MIT License
Generate AI images and videos (Flux, Nano Banana, Kling, Veo, 48+ models), refunds on failure.
v3.9.0 · AGPL-3.0
Self-hosted URL- and file-to-Markdown service for humans and AI agents - web pages, documents, images, audio, YouTube. PWA + REST + MCP + Claude Code skill, Reddit-aware, refreshable share links.
— · MIT
Android Full-Stack Device Control Platform: WebRTC/H.264 remote desktop, UI/OCR/image-matching automation, one-click MITM, built-in Frida, proxy/VPN/frp/P2P networking, MCP/Agent, 160+ APIs, designed for multi-device clusters and engineered deployments.
v1.0.0 · no license
AI-powered ECG analysis: submit ECG images and receive diagnosis reports from ARPI's ECG AI.
v3.0.1 · MIT
Agent-first, provider-neutral multimodal OCR CLI for images, PDFs, URLs, JSON schemas, and agentic extraction with Gemini, Kimi, Muse, and OpenRouter.
v0.2.0 · MIT
Enables AI assistants to control Adobe InDesign on Windows via the Model Context Protocol, allowing document inspection, text and style editing, and exports to formats like PDF, images, and EPUB.
— · MIT
Enables AI agents to generate images on a local Stable Diffusion Forge Neo instance, automatically inferring prompt style and sampling parameters from the user's setup and past generations, and providing tools for LoRA search, model management, and module checks.
— · MIT
Unreal Engine plugin for LLM/GenAI models & MCP UE5 server. OpenAI GPT-5, Deepseek R1, Claude Opus/Sonnet, Gemini 3, Grok 4, Alibaba Qwen, Kimi, ElevenLabs TTS, Inworld, OpenRouter, Groq, GLM, Ollama, Local, Meshy, Tripo, Hunyuan3D, Rodin, fal, Dashscope, Seedream. NPC AI, agentic, chat, 3D gen, TTS, multimodal, image gen. UnrealMCP/UnrealClaude
— · MIT License
Enables agents to generate and edit images, sprites, icons, and animated sprite sheets via Gemini's Nano Banana model. It uses the existing Gemini plan's quota without per-image costs.
v1.0.0 · no license
Generate images, videos, voiceovers, and captions from a chat prompt.
v1.0.0 · no license
Generate vector art, vectorize images, and return SVG, PNG, and logo kits to AI agents.
v1.0.1 · no license
Extract text from documents, manipulate PDFs, and perform OCR on images.