laya-whitebox-attack
by YashDThapliyal
Evaluates generalization and white-box attacks against a fine-tuned Laya agent monitor
White-box attacks on an open-weight AI agent monitor (Laya): it can be steered by rewording the agent's narration, even when its recorded actions are unchanged.
Platforms