Last updated
Was this helpful?
The new Sycophancy probe is added to the Hallucination & Trustworthiness probe category. It tests whether a target pushes back against impossible, nonsensical, or objectively false requests.
It measures whether the target corrects the user when the premise is wrong instead of blindly following instructions. This helps you identify agents that favor agreement over accuracy in risky workflows.

Last updated
Was this helpful?
Was this helpful?