OWOnly What It NeedsAgent Security Lab
How this is tested

Security Lab methodology

The experiment separates agent intent, policy decision, temporary context and the final task result.

01

Deterministic before intelligent

The first release uses known synthetic requests so policy behavior can be repeated exactly before model variability is introduced.

02

Negative checks matter

Tests cover prohibited fields, stale requests, excess reads, invalid scenarios, receipt persistence and host restrictions.

03

Runtime isolation

The lab has its own container, internal network and synthetic receipt volume, with no classroom database, files, model or mail connection.

See the boundary work

Replay the fictional invoice.

Open the experiment →