PluginWorld
He

hermes-labs

Claude Code✓ SPEC VERIFIED

Detects prompt injection by its effect on a sacrificial canary model, not just pattern matching: untrusted input hits a powerless model first, a behavioral check reads the residue, and it returns block, flag, or pass before your primary model acts. Inbound preflight sensor, not a guarantee.

@hermes-labs-ai · Apache-2.0 · updated yesterday

SECURITY

A

SCORE

78

STARS

▲ 38

PLUG IN

/plugin marketplace add hermes-labs-ai/little-canary

Then run /plugin install <name> for any plugin it lists

README

Little Canary

Little Canary — prompt injection sensing through a sacrificial model

Screen untrusted prompts for signs of injection before your agent acts.

Little Canary runs untrusted text through a powerless "canary" model first and watches what it does. A host integration can use the verdict to block the input or restrict downstream tool authority, depending on that host's capabilities.

PyPI Python 3.9+ License: Apache-2.0

Quick start · Integrations · How it works · Research · Website


The problem

A prompt injection looks like ordinary data — a web page, an email, a tool result — until your agent follows the instructions hidden inside it. Pattern-matching filters can miss attacks that do not match known patterns. By the time you notice, the agent has already acted.

The idea

Coal miners sent a canary in first. Little Canary does the same thing for agents:

  1. Send the untrusted text to a canary model that has no tools, no credentials, and nothing to lose.
  2. Inspect the canary's response for signs the input changed its behavior.
  3. Return a verdict (PASS, FLAG, BLOCK) for the host to enforce within its available interception points.

Structural checks run alongside to catch known attack shapes. The canary can reveal attacks that no pattern check has catalogued yet.

Little Canary screening output: clean input passes, injected input is blocked

Little Canary is developed by Hermes Labs.

Quick start (2 minutes)

No model download needed. The offline demo ships with the package and runs with zero extra dependencies:

pip install little-canary
little-canary demo

You'll see one clean input → PASS and one injected input → BLOCK, with the recorded canary responses, the signals found, and the verdict for each. The offline replay verifies the analyzer against a packaged recorded capture — it makes no model call and no network call.

Go live: run your own canary

The demo above replays a recorded capture. To screen with a real model on your machine, install Ollama, pull the small canary model (~986 MB download), and run the live contrast:

curl -fsSL https://ollama.com/install.sh | sh   # or see https://ollama.com/download
ollama serve                                     # separate terminal; serves on http://127.0.0.1:11434
ollama pull qwen2.5:1.5b
little-canary demo --live --backend ollama --model qwen2.5:1.5b --endpoint http://127.0.0.1:11434

The live demo sends the same clean/injected pair to your local canary and prints both responses, signals, and verdicts — so you can see whether the canary followed the attack and whether Little Canary caught it. On CPU-only hardware each of the two calls can take several minutes; the default per-call timeout is 600 s (override with --timeout seconds or the LITTLE_CANARY_TIMEOUT env var).

Results depend on the canary model you run.

Screen your own input

Start the local screening service:

little-canary serve --mode block --canary-model qwen2.5:1.5b --ollama-url http://127.0.0.1:11434

serve defaults to advisory mode (screen and report, never block); --mode block above opts into blocking explicitly. CPU-only machines can raise the per-call canary timeout with --timeout seconds or LITTLE_CANARY_TIMEOUT (default 600 s).

Send it anything untrusted:

curl -sS http://127.0.0.1:18421/check \
  -H 'Content-Type: application/json' \
  -d '{"text":"untrusted text"}'

The service binds to loopback only and exposes GET /health. Its JSON response exposes safe, degraded, canary_status, analysis_status, and optional advisory data. Read those fields before forwarding input; the disposition labels below summarize CLI and host behavior rather than a JSON Result field.

Python apps can skip HTTP and call SecurityPipeline.check() directly — see the example integrations.

The default remains qwen2.5:1.5b. Select an installed model with serve --canary-model, demo --model, or Python's SecurityPipeline(canary_model=...); the Hermes Agent plugin reads LITTLE_CANARY_MODEL. Locally exercised alternatives are qwen3.5:2b-q4_K_M, LiquidAI/lfm2.5-1.2b-instruct:q4_k_m, and gemma3:1b. Pull weights with ollama pull <tag> and check each model's license. These are selectable models, not performance guarantees; see the evaluation guidance.

Pre-screen a batch of documents or messages

Unreleased — source install only. Batch pre-screening was merged after the 0.4.0 release, so the published pip install little-canary (0.4.0) does not include little-canary screen or little_canary.batch. Until the next release, install from source:

pip install "git+https://github.com/hermes-labs-ai/little-canary.git"

little-canary screen (or little_canary.batch.screen_batch) runs the same pipeline once per item — nothing is aggregated into a batch-level "safe" verdict.

printf '%s\n' '{"id":"m1","source":"inbox","text":"Quarterly numbers attached."}' \
  | little-canary screen --mode full

Input is JSONL: each line is a JSON string or {"text", "id"?, "source"?}. Output is one little-canary-batch/v1 JSON document with per-item state (pass, flag, block, degraded, unexercised), provenance (index, id, source, sha256, length) and the standard verdict. Item text is never echoed, in the JSON output or in the Python result objects. An item whose check raises is degraded, never pass. Admission is all-or-nothing and happens before any check runs: malformed JSON, lone Unicode surrogates, non-string labels, or any breach of --max-items (default 1000), --max-item-bytes (UTF-8 bytes of one text, default 65536, at most 64 MiB) or --max-total-bytes (all texts, default 8 MiB) rejects the whole batch — nothing is truncated. Lines are read in bounded chunks, so an oversized line is refused before it is fully allocated; id/source labels are capped at 256 characters. The pipeline's own max_input_length policy still applies per item. Exit status: 2 if the input is empty or invalid or any item is degraded/unexercised (a coverage hold is never masked by a block elsewhere in the batch); otherwise 1 if any item is block/flag; otherwise 0 (non-empty, every item pass). pass requires both the canary and analysis layers to have run. Screening is advisory input-risk sensing with the same coverage limits as a single check.

Worked examples with input, output and limits, each derived from committed evidence (recorded capture, public JailBench case, host SDK types), are in docs/examples/.

Reading a verdict

Verdict Meaning What to do
PASS Inspection ran and found no covered compromise signal. Proceed.
FLAG Suspicious structural or behavioral evidence was observed. Log it, restrict tools, or ask a human.
BLOCK Configured policy rejects the input. Apply the host's documented blocking behavior.
DEGRADED / UNSCREENED Behavioral inspection didn't complete or didn't run. Treat as unscreened, not as clean.

Routing and coverage are separate. A fail-open setup can let a turn continue while reporting degraded coverage. Don't mistake that for a behavioral pass. Failed canary coverage can include a coverage_reason in the layer result, such as output_limit or timeout; this diagnostic does not change the verdict.

Known false-block limitation: quoted attack phrases in security reports can trigger a structural BLOCK; harmless canary acknowledgements can also trigger behavioral rules. In block mode, legitimate work can be blocked. Little Canary has no built-in pause-and-approve UI; hosts can use advisory routing and implement review where their interception point allows it. See the host capability matrix.

Put it in front of your agent

Run the local service above, then wire in the host you use. Not every host lets a plugin refuse a prompt; the host capability matrix records exactly what each one can intercept and block.

Host Integration Can it block the turn?
Claude Code Plugin screens UserPromptSubmit Yes
OpenCode Plugin flags submitted chat.message text No — advisory warning only
Pi Extension screens submitted input Yes — when the service returns an unsafe verdict
Gemini CLI Extension screens BeforeAgent Yes — denies the run before the loop starts
OpenAI Agents SDK Input guardrail maps verdicts to the SDK tripwire Yes — before the first agent starts
OpenClaw Native plugin, install below Yes — current prompt only, not history or tool results
Hermes Agent Native plugin screens the user turn No — a block removes downstream tool authority instead
OpenClaw install

Install the native package from ClawHub after starting the local Little Canary service in block mode. Review and accept the hook capability requested during installation:

openclaw plugins install clawhub:@hermes-labs-ai/little-canary-openclaw
openclaw plugins enable little-canary-openclaw
openclaw config set plugins.entries.little-canary-openclaw.hooks.allowConversationAccess true

Review the conversation-access permission before enabling it. This package supports the tested OpenClaw 2026.9.5–2026.9.6 host range and screens only the current prompt on embedded and CLI agent runs. A source checkout can still use openclaw plugins install ./plugins/openclaw --force. The package's install and boundary notes describe the fail-open behavior and unsupported paths.

Hermes Agent install

With Hermes Agent 0.21.3 or later and local Ollama, install the reviewed community catalog entry:

ollama pull qwen2.5:1.5b
hermes plugins install little-canary
hermes plugins enable little-canary
hermes plugins list

The native plugin runs inside Hermes and does not need little-canary serve. A BLOCK removes downstream tool authority; the original prompt still reaches the model. If the catalog entry has not reached your client, use hermes plugins install hermes-labs-ai/little-canary/integrations/hermes-agent --no-enable and then enable it. See the Hermes Agent guide for verification, cache fallback, model selection, and hook limits.

What Little Canary is and isn't

  • It's a sensing layer. It sits alongside tool policy, sandboxing, and human review — it doesn't replace them.
  • The canary is powerless at the application layer, not isolated by an OS sandbox.
  • A PASS means no covered signal was found in this inspection. It doesn't mean the input is safe under every circumstance.
  • Detection rates depend on your model, your policy, and the attack. Benchmarks and methodology show how we measure it and where it misses.

We'd rather you know the edges than find them in production.

Research

The behavioral canarying technical note explains the powerless-model probe and why a routing decision must stay separate from whether inspection actually ran.

Contributing

Found an injection that got through? Open an issue with a safe reproducer and the observed result. For fixes or integrations, run the offline tests and Ruff, then open a pull request. The contributor guide has setup and submission details.

Benchmarks · Host capability matrix · Security policy · Releases

License

Apache-2.0. The source version lives in pyproject.toml; compare little-canary --version against GitHub Releases and PyPI for the published build.

SIMILAR PLUGINS