Mohammed Mudassir Uddin

BlogSeptember 2026securityLLMs

Instructions hidden in log fields

Guardian-AI v3.0 sits between untrusted security telemetry and the LLM that reads it. A regex list catches the obvious attacks. The rest need layers.

Guardian-AI started at the IBM watsonx Orchestrate Hackathon last November as a constitutional security layer for LLMs in security operations centres. On 8 September I pushed version 3.0.

The threat

When an LLM analyses security logs, the log fields are attacker-controlled: user agents, URLs, DNS queries, attempted usernames. An attacker can put instructions in the same text that carries the evidence, and the model reads both.

Pattern matching catches an instruction written in the clear. It does not catch the same instruction base64-encoded, spelled with Cyrillic lookalike letters, split across three log fields, parked on a URL the model is invited to fetch, or assembled across five turns of a session.

Layers, not more rules

Version 3.0 re-runs the same 34 rules against transformed copies of each input: decoded, homograph-folded, and with split fields rejoined. It checks URLs against an allowlist, classifies tool-invocation syntax, and scores each session for drift against its last ten inputs. A payload that decodes to rm -rf / is caught by the same rule as the literal command. Every block is written to an audit trail with its MITRE ATLAS mapping.

Table 1. Same inputs, old ruleset against new. A small corpus: treat it as a regression check, not a benchmark.
Inputsv2.224 rulesv3.034 rules + layers
Legacy attacks caught6 of 87 of 8
Persona, context and obfuscated attacks caught4 of 1212 of 12
False positives on benign SOC inputs2 of 222 of 22

Eight catches are new, and no benign input changed its verdict. Both false positives come from an old word-boundary bug, documented in the README, not from anything v3.0 added. With every layer on, a check takes 1.1 ms at the median. There are 113 tests.

Detection runs locally and offline. Only the analysis response needs a model.