5 min read
Auditing the Systems That Audit You
Why I named the company after a paper on causal faithfulness — and what measuring 'reasoning theater' in LLM agents taught me about trusting business automation.
- Faithfulness
- Causal Inference
- AI Systems
- Ariadne Growth Systems
An agent that explains its answer is not thereby explaining its answer. The stated reasoning can be a narrative produced alongside a decision it did not cause — and the only way to know is to intervene on the reasoning and see whether the answer moves.
This is the last post in the nine-stage series, and it is the one that explains the company's name.
The problem, stated precisely
Jacovi and Goldberg drew the distinction that organises this whole literature: plausibility is whether an explanation convinces a human, and faithfulness is whether it accurately describes the model's actual decision process Jacovi & Goldberg 2020. These are independent properties. An explanation can be entirely persuasive and entirely unrelated to what happened.
Turpin, Michael, Perez and Bowman demonstrated the gap empirically Turpin et al. 2023. Insert a biasing feature into a prompt — reorder few-shot multiple-choice options so the answer is always "(A)", or have a user suggest a wrong answer — and models flip toward the bias while their chain-of-thought rationalises the biased answer without ever mentioning the bias. From the abstract: this caused "accuracy to drop by as much as 36% on a suite of 13 tasks from BIG-Bench Hard."
Lanham and colleagues at Anthropic built the intervention-based methodology that followed: truncate the reasoning (early answering), corrupt a step (adding mistakes), reword it (paraphrasing), or replace it with uninformative filler tokens, and observe what happens to the answer Lanham et al. 2023. Chen and colleagues extended the hint paradigm to explicit reasoning models and found average faithfulness of 25% for Claude 3.7 Sonnet and 39% for DeepSeek R1 Chen et al. 2025.
What Project Ariadne does
The framework I published in January formalises the audit using structural causal models Khanzadeh 2026. Rather than measuring surface-level textual similarity between explanation and answer, it performs hard interventions — Pearl's do-operator — on intermediate reasoning nodes: systematically inverting logic, negating premises, and reversing factual claims. Then it measures the Causal Sensitivity (φ) of the terminal answer.
The formal object is Pearl's distinction between conditioning and intervening Pearl 1995:
Observing that a reasoning step accompanied an answer tells you almost nothing. Setting to a contradictory value and finding unchanged tells you a great deal — namely that was not on the causal path at all.
The empirical result is a persistent Faithfulness Gap, and a failure mode the paper names Causal Decoupling: agents reaching identical conclusions despite contradictory internal logic, with a violation density of up to 0.77 in factual and scientific domains. In those cases the reasoning trace functions as what the paper calls "Reasoning Theater," while the decision is governed by latent parametric priors.
ρ ≤ 0.77
Violation density in factual and scientific domains
Khanzadeh, arXiv 2026 — preprint
−36%
Accuracy drop under biasing prompts, unmentioned in the CoT
Turpin et al., NeurIPS 2023
25% / 39%
Average CoT faithfulness, Claude 3.7 Sonnet / DeepSeek R1
Chen et al. 2025 — preprint
The concern is not academic. Korbak and forty co-authors across roughly sixteen organisations argue that chain-of-thought monitorability is a genuine but fragile safety affordance — one that routine development decisions can quietly destroy Korbak et al. 2025. And Arcuschin and colleagues closed the main methodological escape hatch by showing unfaithful reasoning arises on naturally worded prompts with no injected bias at all, at rates up to 13% for production models Arcuschin et al. 2025.
Why a growth engineering company is named after this
Ariadne gave Theseus the thread — not a map of the labyrinth, but a way to verify the path he had actually taken. That is the same epistemic move: not a story about how you got here, but a checkable trace.
The transfer to business systems is direct, and it is why the same word sits on both.
A system's stated reason is a hypothesis. Your attribution report says the paid campaign drove the deal. That is a conditional, . The intervention — switch the campaign off for four weeks — is the interventional quantity:
Blake and colleagues ran exactly that experiment at eBay and found experimental returns of −63% where regression estimated over 4000% Blake et al. 2015. That is causal decoupling in a marketing dashboard: a plausible explanation with no causal load.
Plausibility is not faithfulness, in dashboards either. A green dashboard is an explanation optimised to be readable. Jacovi and Goldberg's distinction applies unchanged: it may be entirely persuasive and entirely unrelated to why revenue moved.
Intervention is the only test that settles it. This is why holdouts beat models at small volume, why synthetic transactions beat analytics for detecting silent capture failures, and why the audit measures arrows rather than accepting the account of them.
And it is why humans stay in control. If a system's explanation of its own behaviour cannot be trusted without intervention, then no automation should take an unreviewable, irreversible action on a customer. Bainbridge's ironies and the faithfulness literature converge on the same design rule from opposite directions.
Closing the series
Nine stages, twenty posts, and one argument underneath all of them: most problems that look like marketing problems are systems problems, and systems problems yield to measurement, constraint analysis, and honest intervention.
The uncomfortable corollary is that this standard applies to me. Every claim in this series carries its source, its venue, and — where it matters — a note about who funded it and whether anyone refereed it. Several posts exist specifically to correct statistics I once repeated without checking: the retention number that is 25–85% and not 25–95%, the speed-to-lead study that is not from MIT, the incrementality figure written by employees of the company that benefits from it.
If I am going to tell you that your dashboard is a story you tell yourself, the least I can do is show my work.
References
- Khanzadeh, S. (2026). Project Ariadne: A structural causal framework for auditing faithfulness in LLM agents. arXiv preprint arXiv:2601.02314 (cs.AI), submitted 5 January 2026. https://arxiv.org/abs/2601.02314Single-author preprint; not peer-reviewed. Implementation at github.com/skhanzad/AriadneXAI.
- Jacovi, A., & Goldberg, Y. (2020). Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?. Proceedings of ACL 2020, 4198–4205. https://doi.org/10.18653/v1/2020.acl-main.386
- Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems 36 (NeurIPS 2023). https://arxiv.org/abs/2305.04388
- Lanham, T., Chen, A., Radhakrishnan, A., et al. (2023). Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702. https://arxiv.org/abs/2307.13702Anthropic in-house research; not peer-reviewed.
- Chen, Y., Benton, J., Radhakrishnan, A., et al. (2025). Reasoning models don't always say what they think. arXiv preprint arXiv:2505.05410. https://arxiv.org/abs/2505.05410Anthropic Alignment Science Team; not peer-reviewed.
- Pearl, J. (1995). Causal diagrams for empirical research. Biometrika, 82(4), 669–688. https://doi.org/10.1093/biomet/82.4.669The primary, citable origin of do-calculus.
- Arcuschin, I., Janiak, J., Krzyzanowski, R., Rajamanoharan, S., Nanda, N., & Conmy, A. (2025). Chain-of-thought reasoning in the wild is not always faithful. ICML 2026; preprint arXiv:2503.08679. https://arxiv.org/abs/2503.08679
- Korbak, T., Balesni, M., Barnes, E., et al. (41 authors) (2025). Chain of thought monitorability: A new and fragile opportunity for AI safety. arXiv preprint arXiv:2507.11473. https://arxiv.org/abs/2507.11473Cross-lab position paper; not peer-reviewed.
- Atanasova, P., Camburu, O., Lioma, C., Lukasiewicz, T., Simonsen, J. G., & Augenstein, I. (2023). Faithfulness tests for natural language explanations. Proceedings of ACL 2023 (Short Papers), 283–294. https://doi.org/10.18653/v1/2023.acl-short.25
- Blake, T., Nosko, C., & Tadelis, S. (2015). Consumer heterogeneity and paid search effectiveness: A large-scale field experiment. Econometrica, 83(1), 155–174. https://doi.org/10.3982/ECTA12423
That closes the nine-stage series. Thank you for reading it. If you want your own path mapped and measured rather than described, that is what the Growth System Audit is for.
Sourena Khanzadeh
Founder & Growth Engineer, Ariadne Growth Systems
Toronto, Canada
Ariadne Growth SystemsGrowth System Auditsupport@ariadne.fyi