A study published in The International Tax Journal tests, with data rather than assertion, a claim commonly made for AI-driven transfer pricing benchmarking tools: that autonomous systems can produce faster analysis without a corresponding loss of rigour. The findings deserve closer attention from Indian practitioners than they are likely to receive.
The authors ran agentic AI, systems that execute a multi-step workflow such as a benchmarking search or functional analysis without step-by-step human prompting, against 20 simulated MNE cases across three sectors. They compared the outputs against a model of how tax administrations in 11 countries are deploying similar technology. The efficiency gain was substantial: agentic AI cut processing time by roughly 98% while increasing the depth of the functional analyses produced.
The same study reports three failure rates that merit attention before any TP head signs off on an AI-assisted benchmarking pipeline: a 15% error rate in complex functional characterisations, a 22% irrelevance rate in comparable selection, and a 15% hallucination rate in legal citations. The authors propose a governance framework they call "Tracer-Wire," under which every AI-generated conclusion must carry a visible, auditable path back to its source data, with a mandatory human checkpoint before any output is finalised.
Explainability-by-design and human-in-the-loop review are now standard features of AI governance proposals, so the framework itself is not the notable part of the paper. What is notable is the coincidence between the study's error rates and recent Indian tribunal outcomes.
This month alone, three decisions have turned on the same issue the study measures. The Delhi bench of the ITAT excluded two comparables from Dixon Technologies' set for functional dissimilarity, even though the taxpayer had itself flagged the issue years earlier. The Chennai bench devoted an entire order to whether a single internal comparable can still claim the statutory tolerance band. The Karnataka High Court's SAP Labs line of rulings has produced a further set of follow-on decisions this month on whether a TPO may discard a taxpayer's comparables in favour of a "standard set."
Each of these disputes falls within the 22% irrelevance-rate failure mode the study measured. If agentic AI misjudges comparable relevance roughly one time in five even in a controlled simulation, and Indian tribunals are already spending full orders correcting comparable-selection errors made by humans, the practical question for TP practice in 2026 is concrete rather than conceptual: whose signature appears on the local file when an AI-selected comparable turns out, on review by a TPO or an ITAT bench two years later, to be a functionally dissimilar entity such as a plastics manufacturer.
Coverage of AI in transfer pricing tends to avoid a question that is uncomfortable for vendors and practitioners alike: whether an audit trail reduces liability or merely relocates it. A Tracer-Wire log showing that an AI system considered and rejected a comparable for a documented reason does not make that rejection correct. It makes the error more visible, and it arguably shifts accountability toward whoever approved the workflow rather than toward the tool itself.
It is worth asking whether India's Master File and Form 3CEB documentation requirements are structured to capture this kind of AI-decision provenance at all. If they are not, practitioners may be looking at a new category of documentation gap in an area the OECD's Chapter V framework was never designed to address. Firms adopting agentic AI for benchmarking would do well to build a sign-off protocol now, one that fixes responsibility for each accepted comparable before a tribunal does it for them.