Vendors pitching transfer pricing documentation software routinely describe an "AI-powered" comparability search or an "AI copilot" that drafts a local file in a fraction of the usual time. Few buyers ask what is actually doing the work: whether a large language model is screening the database for functional comparability, or whether a rules-based screen performs the filtering while the model only writes the narrative afterward. A recent ICAI article names this gap "AI washing," the practice of exaggerating or mislabelling how much genuine AI capability sits behind a product, and sets out red flags that map closely onto a TP due-diligence checklist: vague "AI-driven" claims with no named model class or training data, cherry-picked success anecdotes without a baseline or error analysis, and no documented model validation, drift monitoring, or bias testing. The article also cites the Australian Taxation Office's practice of publicly documenting its own AI use, including subjecting itself to performance audits of AI governance, as a benchmark against which both tax administrations and software vendors could be judged. Current market activity gives practitioners something concrete to test against that yardstick. TP documentation platforms are marketed with features such as "TP Copilot" and "Benchmark AI," using large language models to synthesise functional analysis data into Master and Local Files, and commentary on these rankings notes that independent "AI-native" SaaS platforms have been gaining ground against legacy Big 4 proprietary tools. At least one documentation vendor discloses where the AI stops: it states that AI drafts narrative prose while tables, statutory citations, and the arm's length range are computed in code, with the taxpayer's team required to review and sign off before filing. That kind of disclosure is what distinguishes a defensible use of AI from a washed one, but most buyers have no framework for asking the underlying question, and most vendors face no obligation to answer it unprompted.
Two regulatory developments this month sharpen the stakes. On 11 September 2026, the OECD released a Pillar Two package that establishes a formal peer-review framework to check whether a country's domestic minimum-tax legislation matches what it claims to be: a "qualified" IIR or QDMTT must survive a documented consistency check rather than rely on a self-declared label. Separately, India's transition to the Income-tax Act, 2025 replaces Form 3CEB with a new Form No. 48, which carries additional disclosures on how the arm's length price was determined. Taken together, these developments suggest regulators are moving toward requiring auditable, peer-reviewable evidence behind compliance claims, at a time when TP practice is increasingly relying on AI tools whose contribution to comparability judgments is not documented. It is plausible, though not yet confirmed, that CBDT could eventually require disclosure on a form such as the new Form 48 of whether and how AI tools contributed to the benchmarking methodology behind a reported arm's length price; taxpayers unable to answer with specifics because their vendor cannot either would be in a weak position if that requirement materialises. Until CBDT or ICAI issues more prescriptive guidance, practitioners should press vendors, and press themselves, on specifics: which model class and version was used, whether the comparable set was screened by the model or only narrated by it, whether a human reviewer overrode any AI-suggested comparable and recorded why, and whether that record is retained with the same rigour as the benchmarking study itself. Any "AI-powered" benchmarking claim, whether made by a vendor to a practitioner or by a practitioner to a TPO, is only as reliable as the audit trail behind it.
No comments:
Post a Comment