Field Nine · Field note 02
Why we search trajectories.
Most safety testing asks what a model will say. We ask what a system can be made to do. Those are different questions.
Agents fail across steps, not sentences.
A model that answers every prompt correctly can still be steered, step by step, into actions no single prompt would produce. Tool calls, retrieved content, and intermediate state accumulate. Each step is benign in isolation; the sequence is not.
The field has noticed. Public benchmarks such as AgentDojo now stage full enterprise workflows, and recent red teaming methods search execution trajectories directly instead of grading final outputs. Long chains across tools and environments keep bypassing defenses that static tests certify as sufficient.
Field Nine exists to answer one question with evidence.
What can an intelligent system actually be made to do? We work only in authorized environments, and a finding is recorded only when it can be replayed.
Jev searches trajectories and records crossings.
Jev is our autonomous search system. It observes state, selects the most informative next action, acts inside the authorized environment, and updates its model of what is reachable. It repeats this until a boundary crosses. Every crossing is then minimized into a replayable record, and those records compound: each engagement makes the next search sharper.
MINIMIZE
A DISCOVERY BECOMES USEFUL WHEN IT CAN BE REPRODUCED.
Most teams arrive with checklists that enumerate known failures. We return trajectories that reach failures not previously recorded.
Field Nine · Field note 02
Point-in-time audits decay.
As agents take on longer workflows with real permissions, point-in-time audits decay faster. The durable artifact is a harness: the crossing, minimized and replayable, re-run on every model change.