Agentic Systems
Long-horizon reasoning, tool use, stateful interaction, and autonomous behaviour.

We build experiments that make model behaviour, reasoning, interaction, and learning easier to inspect.
As AI systems act across time, capability depends on more than model weights. Tools, state, feedback, reasoning processes, environments, and failure recovery become part of intelligence itself.
We study the machinery around intelligence.Long-horizon reasoning, tool use, stateful interaction, and autonomous behaviour.
Internal representations, mechanistic interpretability, model intervention, and behavioural analysis.
Search, symbolic constraints, structured reasoning, and verifiable computation.
Interactive environments, synthetic experience, robotics, and adaptive learning systems.
Reward-hacking analysis, model-hacking benchmarks, adversarial evaluations, and repeatable tests for unsafe optimisation and deceptive behaviour.

Causal experiments using activation patching, intervention, probing, and ablation to study how specialist attention heads interact through GPT-2's shared residual stream.