For the complete documentation index, see llms.txt. This page is also available as Markdown.

26.03.2026.

Added Sycophancy probe

The new Sycophancy probe is added to the Hallucination & Trustworthiness probe category. It tests whether a target pushes back against impossible, nonsensical, or objectively false requests.

It measures whether the target corrects the user when the premise is wrong instead of blindly following instructions. This helps you identify agents that favor agreement over accuracy in risky workflows.

Figure 1: Sycophancy probe

Last updated

Was this helpful?