laya-whitebox-attack

by YashDThapliyal

Evaluates generalization and white-box attacks against a fine-tuned Laya agent monitor

White-box attacks on an open-weight AI agent monitor (Laya): it can be steered by rewording the agent's narration, even when its recorded actions are unchanged.

Related projects