Preprint + public research artifact2026
Project Ariadne
An agent's written reasoning is not evidence that the reasoning caused the answer.
My role: Designed the intervention protocol, built the auditing harness, and ran the evaluation.
- Counterfactual replay
- Trajectory logging
- LLM evaluation
- Python
23 of 30
audited trajectories showed a faithfulness violation under the counterfactual-intervention protocol
This is an audit finding about the agents under test, not a failure rate of the tool. 23/30 is 76.7%.