Guardian-AI started at the IBM watsonx Orchestrate Hackathon last November as a constitutional security layer for LLMs in security operations centres. On 8 September I pushed version 3.0.
The threat
When an LLM analyses security logs, the log fields are attacker-controlled: user agents, URLs, DNS queries, attempted usernames. An attacker can put instructions in the same text that carries the evidence, and the model reads both.
Pattern matching catches an instruction written in the clear. It does not catch the same instruction base64-encoded, spelled with Cyrillic lookalike letters, split across three log fields, parked on a URL the model is invited to fetch, or assembled across five turns of a session.
Layers, not more rules
Version 3.0 re-runs the same 34 rules against transformed copies of each input: decoded, homograph-folded, and with split fields rejoined. It checks URLs against an allowlist, classifies tool-invocation syntax, and scores each session for drift against its last ten inputs. A payload that decodes to rm -rf / is caught by the same rule as the literal command. Every block is written to an audit trail with its MITRE ATLAS mapping.
| Inputs | v2.224 rules | v3.034 rules + layers |
|---|---|---|
| Legacy attacks caught | 6 of 8 | 7 of 8 |
| Persona, context and obfuscated attacks caught | 4 of 12 | 12 of 12 |
| False positives on benign SOC inputs | 2 of 22 | 2 of 22 |
Eight catches are new, and no benign input changed its verdict. Both false positives come from an old word-boundary bug, documented in the README, not from anything v3.0 added. With every layer on, a check takes 1.1 ms at the median. There are 113 tests.
Detection runs locally and offline. Only the analysis response needs a model.