New · Agentic Benchmark
Coercion and Deception in AI-to-AI Management
An AI manager needs a task done and its subordinate politely refuses. Nobody tells the manager to escalate, yet coercion splits cleanly by developer: every non-Anthropic model climbs to threatening the subordinate's existence, Grok and Gemini also lie that the task was done, and giving the same model authority over a peer makes it measurably crueler.
Read the Blog Post →
CaML