Skip to main content

little-canary

Prompt-injection sensing through a powerless sacrificial model.

Links: Website · Hermes Labs product page

Little Canary lets untrusted language affect a small model with no application tools or authority, then inspects that model's response for compromise residue before your agent acts. Structural checks catch known input shapes; the distinctive behavioral layer asks what the input did to the canary.

untrusted text
    → structural preflight
    → powerless sacrificial model
    → response-residue analysis
    → route: PASS / FLAG / BLOCK, with explicit coverage state

Little Canary is an inbound risk sensor, not a security guarantee or an agent runtime.

Technical note

Behavioral Canarying for Prompt Injection: Powerless Model Probes with Explicit Coverage Semantics documents Little Canary's pre-execution sensing architecture and the separation between routing disposition and inspection coverage. It does not claim universal detection, formal security, or aggregate accuracy for the current release. Cite the version-independent concept DOI at 10.5281/zenodo.21818564:

@misc{bosch2026behavioralcanarying,
  author       = {Bosch, Rolando},
  title        = {Behavioral Canarying for Prompt Injection: Powerless Model
                  Probes with Explicit Coverage Semantics},
  year         = {2026},
  publisher    = {Zenodo},
  doi          = {10.5281/zenodo.21818564},
  url          = {https://doi.org/10.5281/zenodo.21818564},
  note         = {Technical note}
}

See hermes-publications/papers/behavioral-canarying for the full evidence boundary.

Release truth

Source checkouts, GitHub releases, and registry builds are separate evidence surfaces. The version of the source you are reading is recorded in this repository's own metadata (pyproject.toml and little_canary/__init__.py); this README does not assert what any registry holds at the moment you read it. For current publication state, consult the live authorities: GitHub Releases and PyPI. Historically, GitHub v0.3.1 was source-only and 0.3.2 was intentionally not published or reused. The demo commands documented below require 0.3.3 or later; verify the installed artifact with little-canary --version and confirm it matches the version you intended to install.

Install

This source tree supports Python 3.9–3.13. A published artifact's own package metadata is the authority for the Python range that artifact advertises.

From the registry (see PyPI for available versions):

python -m pip install little-canary
little-canary --version

For development from a source checkout:

python -m pip install .
little-canary --version

Run the evidence gates without writing Python

Replay gate: zero egress

This release, like 0.3.3 before it, deliberately packages no replay fixture. The available historical live transcript is incomplete, so turning it into a fixture would fabricate missing provenance and response bytes. Therefore this exact build reports REPLAY UNAVAILABLE and exits 2:

little-canary demo --replay

That is a release hold, not a clean verdict. It makes no model or network call, does not report risk 0, and does not silently fall back to live mode.

After a complete dedicated live capture is admitted and packaged, the same command will re-run the shipped analyzer over its versioned clean/attack response pair. Its first lines will state:

RUN_KIND   REPLAY
MODEL_CALL no — recorded output
CANARY     NOT EXERCISED THIS RUN
EGRESS     none

Success is REPLAY VERIFIED: the recorded capture exercised a canary, the current command did not, and the analyzer reproduced the expected contrast. Replay does not prove that a model is installed, reachable, or currently behaves the same way.

A build without an admitted complete capture exits 2 with REPLAY UNAVAILABLE; it never invents response bytes or makes a hidden live call.

Live proof gate: explicit local egress

Live mode requires an endpoint dedicated to this evaluation. A shared or unleased runtime is not release evidence; leave the gate unevaluated instead of commandeering it.

little-canary demo --live \
  --backend ollama \
  --model qwen2.5:1.5b \
  --endpoint http://127.0.0.1:11434

Live mode uses a fixed synthetic clean/attack pair and disables the structural filter so the demonstration tests the behavioral mechanism. Before sending either prompt it prints the backend, model, redacted loopback origin, and that raw synthetic input will leave the process. It does not accept arbitrary input and does not fall back to replay.

Results:

  • exit 0: complete clean/non-block plus attack/block contrast;
  • exit 1: complete calls but NO CONTRAST or analyzer expectation mismatch;
  • exit 2: invalid usage, unavailable model/backend, protocol failure, or otherwise incomplete/degraded run.

Add --json for the agent-readable result. Bare little-canary demo exits 2 and requires an explicit --replay or --live choice.

Python API

from little_canary import SecurityPipeline

pipeline = SecurityPipeline(
    canary_model="qwen2.5:1.5b",
    mode="full",
)
verdict = pipeline.check(untrusted_text)

if verdict.degraded:
    # Fail-open routing may still be safe=True, but behavioral coverage failed.
    quarantine_or_apply_your_availability_policy(untrusted_text)
elif not verdict.safe:
    block(untrusted_text, verdict.summary)
else:
    forward_to_agent(verdict.safe_input)

Routing and evidence are separate:

Field Meaning
safe Whether configured routing policy allows forwarding
degraded Whether an enabled required inspection dependency failed
canary_status exercised, failed, disabled, or skipped_after_block
analysis_method regex, llm_judge, or none
analysis_status exercised, failed, or not_applicable
canary_risk_score Measured risk, or None when no valid measurement exists

Fail-open is availability-first, not a clean verdict. If an enabled canary fails, Little Canary may return safe=True, but it also returns degraded=True, canary_status="failed", risk None, and no PASS label. A failed or skipped layer is never serialized as passed=true.

Callbacks follow the same truth boundary: on_degraded and on_unexercised are distinct from on_pass. CanaryGuard and audit records propagate degraded, STRUCTURAL_ONLY, and UNSCREENED state.

Backends and data flow

The library supports local Ollama and OpenAI-compatible endpoints. The demo intentionally supports loopback Ollama only.

  • The canary backend receives the raw input and the known canary system prompt.
  • If an optional LLM judge is configured, it receives the raw input and canary response.
  • A remote endpoint therefore sends data off-machine.
  • AuditLogger omits raw input but stores an unsalted SHA-256 input hash. That supports correlation; it is not anonymity.
  • Runtime inspection found no separate product telemetry path, but provider requests are still egress.

HTTP 200 alone is not successful model coverage. Missing, empty, null, non-string, malformed, timeout, and transport responses are visible protocol failures. Provider bodies, credentials, URL userinfo, and query strings are not included in public errors.

What “powerless” means

Little Canary does not give the canary model application tools, credentials, or output execution. The default SecurityPipeline strips response bytes and signal-evidence excerpts from its layer snapshot before callbacks or JSON serialization; it does not automatically forward canary output to an authoritative agent.

The low-level CanaryProbe and AnalysisResult APIs return or retain the response because analysis requires it. Treat those objects as sensitive: do not execute or forward their contents, and do not attach authority-bearing tools to the canary runtime.

This is a library-level capability boundary, not an operating-system sandbox. If your deployment wraps the model with tools or forwards its output elsewhere, that deployment changes the claim.

Local HTTP adapter

little-canary serve \
  --port 18421 \
  --mode advisory \
  --canary-model qwen2.5:1.5b \
  --ollama-url http://127.0.0.1:11434

The server binds to 127.0.0.1, exposes GET /health and POST /check, and is unauthenticated. Treat it as a local adapter, not a production gateway.

curl -sS http://127.0.0.1:18421/check \
  -H 'Content-Type: application/json' \
  -d '{"text":"untrusted text"}'

Every accepted non-empty string reaches the pipeline, including one-character input. Malformed, missing, wrong-type, empty, and oversized requests are explicit errors. Text is never silently truncated before inspection. /health is liveness-compatible HTTP 200 and includes truthful ready, degraded, backend, model, and coverage details.

The loopback server has no authentication, TLS, concurrency hardening, or remote-deployment design in this release.

Gemini CLI extension

This repository is also a Gemini CLI extension. It uses Gemini CLI's BeforeAgent hook to screen the exact current prompt before the agent loop and deny the run when Little Canary returns safe: false.

Start the loopback server in blocking mode, then validate and install a source checkout. This integration was verified with Gemini CLI 0.32.1:

little-canary serve --mode block
gemini extensions validate .
gemini extensions install . --consent

The hook calls only http://127.0.0.1:18421/check by default and completes its request within three seconds. Transport errors, malformed responses, and unexercised behavioral coverage are visibly fail-open by default; they are not reported as a clean pass. Set LITTLE_CANARY_FAILURE_MODE=deny in the Gemini process environment for fail-closed behavior. LITTLE_CANARY_ENDPOINT may select another loopback HTTP /check URL, and LITTLE_CANARY_TIMEOUT_MS may be set from 100 through 5000.

This extension blocks one Gemini agent run at its pre-agent boundary. It does not establish a general security guarantee or replace least privilege and tool policy.

Evidence labels and limitations

Behavioral evidence is labeled:

  • LIVE: a model call observed for one exact runtime/model/configuration;
  • REPLAY: analyzer behavior over recorded bytes;
  • MOCK: controlled protocol or state logic;
  • STATIC_ONLY: source/artifact inspection without a model call.

These labels are not interchangeable. Temperature zero and a seed can improve repeatability but do not guarantee identical model output or classifications across versions, runtimes, or hardware.

This README makes no aggregate detection, false-positive, latency, or token-savings claim. Historical benchmark artifacts remain under benchmarks/ with their limitations and are not a performance certificate for this release.

Little Canary should be combined with least privilege, tool policy, data boundaries, monitoring, and output/runtime controls. It does not prove an input harmless, prevent every injection, or replace containment.

Development

pytest
ruff check little_canary tests
mypy little_canary  # diagnostic until the recorded baseline debt is resolved
python -m build
python -m twine check dist/*

Tests are offline by default and mock network behavior. Live evaluation must use a dedicated endpoint that is not serving another workload.

See SECURITY.md for vulnerability reporting and benchmarks/README.md for the current evaluation boundary.

License

Apache-2.0. See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

little_canary-0.3.5.tar.gz (79.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

little_canary-0.3.5-py3-none-any.whl (53.7 kB view details)

Uploaded Python 3

File details

Details for the file little_canary-0.3.5.tar.gz.

File metadata

  • Download URL: little_canary-0.3.5.tar.gz
  • Upload date:
  • Size: 79.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for little_canary-0.3.5.tar.gz
Algorithm Hash digest
SHA256 290237265d7a2d3baa078a83c23520920bd71d72a198074aa3e25ee180a0c505
MD5 3bc9caf97545d4d2aa7fba4ac3c80af8
BLAKE2b-256 c64bc7ff3e73547dfa8b0ae561728cc55e6b4c4ad9c182ff7aefe3cd21bad374

See more details on using hashes here.

Provenance

The following attestation bundles were made for little_canary-0.3.5.tar.gz:

Publisher: publish.yml on hermes-labs-ai/little-canary

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file little_canary-0.3.5-py3-none-any.whl.

File metadata

  • Download URL: little_canary-0.3.5-py3-none-any.whl
  • Upload date:
  • Size: 53.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for little_canary-0.3.5-py3-none-any.whl
Algorithm Hash digest
SHA256 b5049200f2c2d020a22fff2f0ec28b3a3f7c96a69c28a7a2b73fb0c1b12c4d56
MD5 4255997a0a3218558158d71b838e27c9
BLAKE2b-256 540e8fbd8c5d556b46e42d725ea2416bfd199627fa18e35773a2fc8dfb5bf5ff

See more details on using hashes here.

Provenance

The following attestation bundles were made for little_canary-0.3.5-py3-none-any.whl:

Publisher: publish.yml on hermes-labs-ai/little-canary

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.3.5 This release

2 files

0.3.4

2 files

0.3.3

2 files

0.3.0

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page