Thursday, September 17, 2026

AI-Performed DEMPE Functions and the Limits of Transfer Pricing's Intangibles Framework

The OECD's DEMPE framework was designed to stop groups from parking valuable intangibles in a low-tax entity that holds legal title but performs none of the underlying work. The test traces the relevant functions, development, enhancement, maintenance, protection and exploitation, back to the people who exercise judgment over them: who decided to pursue a research direction, who approved a patent filing, who manages commercialisation risk. For two decades this test has worked reasonably well because those functions were, in fact, performed by identifiable people.

That assumption is being tested where AI systems perform DEMPE functions. Commentary in Australia has raised this issue in relation to the ATO's Practical Compliance Guideline 2024/1 on intangibles migration, which requires multinationals to self-assess and disclose where DEMPE activities for offshore-held intangibles actually occur, using a risk-zone framework that determines how closely the ATO will scrutinise a taxpayer's position. The commentary notes that the PCG's evidence requirements assume human decision-makers, leaving a gap where AI systems (training models, tuning algorithms, monitoring outputs, iterating on protection measures) perform these functions without a person who can be identified as the locus of judgment.

A companion piece extends this analysis to M&A due diligence. It argues that AI systems, including the algorithms, training data and infrastructure, are themselves valuable intangibles under transfer pricing principles, and that a target lacking DEMPE documentation for its AI footprint may carry dormant tax risk that an acquirer inherits on completion.

This is not confined to Australia, and it is not a distant issue for India. India's global capability centre sector has moved well beyond routine coding and support work. Commentary directed at India's GCC market has told clients that where a captive centre designs a core AI algorithm used across the global group, arm's-length pricing cannot be justified by comparing developer hourly rates; the pricing must reflect that creative, value-creating contribution. That is a DEMPE argument, framed in language close to the ATO's, applied to the kind of AI-first GCC that has become common in Bengaluru, Hyderabad and Pune over the last two years.

The timing adds to the difficulty. The Finance Act 2026 safe-harbour rationalisation consolidated software, ITeS, KPO and contract R&D into a single IT-services category, set a flat 15.5% margin, and raised the eligible transaction threshold to ₹2,000 crore. The change was presented as compliance simplification for routine captive units. A captive unit whose engineers are training or fine-tuning a model whose commercial exploitation happens offshore may not be 'routine' in the DEMPE sense, even if it fits comfortably within the new safe-harbour band. A group that elects into the safe harbour for five years on the strength of a routine cost-plus characterisation may be locking in a margin that a DEMPE-consistent functional analysis would not support. That exposure may surface only when the safe-harbour election lapses and a TPO examines what the unit was actually doing.

This raises a harder question about the DEMPE test itself. DEMPE and the 'significant people functions' concept exist to identify who exercises judgment over value creation. Where that judgment is distributed across a model's training run, a human supervisor who approves outputs in bulk, and an infrastructure team in a different jurisdiction, none of these resembles the risk-bearing decision-maker the framework was written for. Tribunals and tax authorities have decades of practice tracing DEMPE to people, but little developed practice for tracing it to a system.

Whether CBDT, the ITAT or the OECD develops a workable answer first, and whether India follows the ATO's disclosure-and-self-assessment approach or adopts something more prescriptive, remains unclear. For now, practitioners advising GCCs and AI-heavy captives should treat DEMPE documentation for AI-performed functions as a present compliance issue, and should examine whether a safe-harbour election understates the functional profile of a unit that is doing more than routine service delivery.

Quantifying AI's Error Rate in Transfer Pricing Benchmarking

A study in The International Tax Journal tested multi-agent AI systems against simulated data covering twenty MNE cases across three sectors, comparing deployment models across eleven countries. Agentic AI cut processing time by 98% and performed deeper functional analysis than a human team in the same window. The study also recorded a 15% error rate in complex functional characterizations, a 22% irrelevance rate in comparable selection, and a 15% hallucination rate in legal citations.

These figures matter for Indian practice because the comparable set is the most contested element in most disputes reaching the ITAT, on turnover filters, FAR mismatches, and functional dissimilarity. A 22% irrelevance rate applied to a twelve-company set implies two or three comparables may not withstand scrutiny. A 15% citation hallucination rate is more serious, since courts elsewhere have sanctioned practitioners for AI-fabricated citations. An ABA Tax Section panel on Section 482 similarly concluded that practitioners remain responsible for method selection and defensibility despite AI assistance.

Until verification protocols match these error rates, AI-assisted local files and benchmarking studies warrant the same partner-level scrutiny as any junior draft, if not more.pemesan

Ingram Micro: TPO's PE Finding Falls Outside a Section 92CA Reference

A Section 92CA reference is meant to test the arm's length price of a specifically identified international transaction. The Assessing Officer's reference defines that scope, and a TPO cannot expand it.

In Ingram Micro (India) Exports Pte. Ltd. v. DCIT, the AO referred the matter to the TPO based on search statements that the Indian affiliate was "carrying on the actual business" of the Singapore entity, a general assertion rather than an identified transaction. The TPO went further, holding that the assessee had a Permanent Establishment in India and raising an adjustment of about Rs. 61.32 crore, feeding a proposed assessment of roughly Rs. 72.39 crore.

The Mumbai ITAT set aside the order. It held that a TPO's jurisdiction under Section 92CA is confined to the arm's length price of a specifically identified transaction. PE existence under Article 5 and taxability of profits under Article 7 remain matters for the AO, not the TPO.

Practitioners reviewing TPO orders built on broadly worded references should examine whether the reference itself identifies an actual transaction before accepting findings made on the strength of it.

AI Washing and Transfer Pricing Benchmarking: Closing the Due Diligence Gap

Vendors pitching transfer pricing documentation software routinely describe an "AI-powered" comparability search or an "AI copilot" that drafts a local file in a fraction of the usual time. Few buyers ask what is actually doing the work: whether a large language model is screening the database for functional comparability, or whether a rules-based screen performs the filtering while the model only writes the narrative afterward. A recent ICAI article names this gap "AI washing," the practice of exaggerating or mislabelling how much genuine AI capability sits behind a product, and sets out red flags that map closely onto a TP due-diligence checklist: vague "AI-driven" claims with no named model class or training data, cherry-picked success anecdotes without a baseline or error analysis, and no documented model validation, drift monitoring, or bias testing. The article also cites the Australian Taxation Office's practice of publicly documenting its own AI use, including subjecting itself to performance audits of AI governance, as a benchmark against which both tax administrations and software vendors could be judged. Current market activity gives practitioners something concrete to test against that yardstick. TP documentation platforms are marketed with features such as "TP Copilot" and "Benchmark AI," using large language models to synthesise functional analysis data into Master and Local Files, and commentary on these rankings notes that independent "AI-native" SaaS platforms have been gaining ground against legacy Big 4 proprietary tools. At least one documentation vendor discloses where the AI stops: it states that AI drafts narrative prose while tables, statutory citations, and the arm's length range are computed in code, with the taxpayer's team required to review and sign off before filing. That kind of disclosure is what distinguishes a defensible use of AI from a washed one, but most buyers have no framework for asking the underlying question, and most vendors face no obligation to answer it unprompted.

Two regulatory developments this month sharpen the stakes. On 11 September 2026, the OECD released a Pillar Two package that establishes a formal peer-review framework to check whether a country's domestic minimum-tax legislation matches what it claims to be: a "qualified" IIR or QDMTT must survive a documented consistency check rather than rely on a self-declared label. Separately, India's transition to the Income-tax Act, 2025 replaces Form 3CEB with a new Form No. 48, which carries additional disclosures on how the arm's length price was determined. Taken together, these developments suggest regulators are moving toward requiring auditable, peer-reviewable evidence behind compliance claims, at a time when TP practice is increasingly relying on AI tools whose contribution to comparability judgments is not documented. It is plausible, though not yet confirmed, that CBDT could eventually require disclosure on a form such as the new Form 48 of whether and how AI tools contributed to the benchmarking methodology behind a reported arm's length price; taxpayers unable to answer with specifics because their vendor cannot either would be in a weak position if that requirement materialises. Until CBDT or ICAI issues more prescriptive guidance, practitioners should press vendors, and press themselves, on specifics: which model class and version was used, whether the comparable set was screened by the model or only narrated by it, whether a human reviewer overrode any AI-suggested comparable and recorded why, and whether that record is retained with the same rigour as the benchmarking study itself. Any "AI-powered" benchmarking claim, whether made by a vendor to a practitioner or by a practitioner to a TPO, is only as reliable as the audit trail behind it.

 

Saturday, September 12, 2026

India Is Negotiating the Future of AI Tax Sourcing at the UN

The arm's length principle, and the DEMPE test that operationalizes it for intangibles, was built on an assumption so basic that nobody bothered to write it down: that a human being somewhere is developing, enhancing, maintaining, protecting or exploiting the asset, and that you can find that human, interview them, and attribute value to their employer based on what they actually did. A recent industry analysis of AI's effect on TP disputes puts the resulting question bluntly - when risk is managed through algorithms and data centres spread across jurisdictions while a smaller, dispersed human team just sets parameters, the practical question becomes where exactly control over risk is anchored. Once AI genuinely performs the function rather than merely assisting a human who performs it, the entire evidentiary architecture of functional analysis has nothing to point to.

This isn't an abstract worry anymore; it's showing up in real guidance in at least two places practitioners should be tracking together. Australia's ATO has PCG 2024/1, its compliance guideline on intangibles migration, and a MinterEllison alert has flagged that its DEMPE evidence requirements assume human decision-makers : leaving a gap precisely where AI performs those functions, with the firm advising multinationals to map AI-related DEMPE functions and reassess their transfer pricing positions now. Separately, and on a longer institutional timeline, the UN Tax Committee's 32nd session mandated two things worth pairing: a subcommittee tasked with a Practical Guide to Implementing Artificial Intelligence for Tax Administrations (due no later than October 2027), and an update to the 2021 UN Transfer Pricing Manual's chapters on intragroup services, intangibles and financial transactions specifically to account for new business models. Meanwhile, at the treaty level, the Fifth Session of the UN Framework Convention on International Tax Cooperation negotiations, held in New York from August 3–13, 2026, saw nations explicitly debate whether AI should be covered under the draft protocol on taxing cross-border services - with live tracking of the session showing India engaged on a closely related scope question, backing further work on Article 2 while questioning how it was drafted.

India is currently playing both sides of this problem without anyone connecting the dots publicly. At the UN negotiating table, India is a rule-maker, shaping how cross-border AI-driven services income gets sourced and taxed under a protocol that will matter enormously given India's position as the world's largest hub for Global Capability Centres and IT-enabled services exports. At home, India is simultaneously a rule-enforcer, and the domestic playbook hasn't caught up. Mainstream GCC audit-defense advisory is already telling clients that if an Indian GCC is designing a core AI algorithm used globally, the arm's length price must reflect that creative contribution, and that authorities are using data-mining tools to compare GCC margins across the industry : in other words, Indian TPOs are already pricing AI-DEMPE contributions using the old human-centric functional analysis framework, years before the UN's own Practical Guide or protocol language is finalized. There's also a second, quieter version of this same problem sitting inside the OECD's plumbing: the OECD has just published the public comments on its own revision of Chapter VII (intra-group services guidance), with a consultation meeting set for November 2026 - the very guidance that ITAT benches are currently applying to decide intra-group services benefit-test disputes, using a chapter the OECD itself has decided needs a structural rewrite.

The open question worth raising is whether India should be trying to actively harmonize its emerging domestic AI-DEMPE case law with the position it's taking at the UN negotiating table - or whether it is, without quite intending to, building two separate and eventually incompatible bodies of doctrine: one forged in ITAT orders and CBDT safe harbour notifications that assume a human designed the algorithm, and another being drafted in a UN conference room that may define AI-driven service income sourcing on entirely different terms.

Wednesday, September 9, 2026

The UN Just Put AI on the Transfer Pricing Table

Start with the problem most transfer pricing practitioners are tracking the wrong forum for. When people think about how AI-delivered cross-border services will be taxed, the reflex is to look at the OECD - Pillar One, Amount B, the endless refinements to the arm's-length principle for baseline distribution and marketing functions. That's not wrong, but it's incomplete. There is a second, quieter negotiation running in parallel at the United Nations that could end up mattering just as much, and it is the one actually putting AI's name on paper right now.

Here's what changed. The UN Framework Convention on International Tax Cooperation has been negotiating, since early 2025, a framework treaty plus two 'early protocols' - one on dispute resolution, and one specifically on taxing cross-border services. The Co-Leads published a draft text of that services protocol on 20 July 2026, and during the Fifth Session of negotiations, held at UN Headquarters from 3 to 13 August 2026, multiple member states pushed to have artificial intelligence explicitly captured within its scope. That same session saw other states raise concerns about overlapping nexus claims for service fees, and a separate bloc pushing for optionality in how the protocol's provisions get applied. In other words: the machinery for deciding which country gets to tax AI-delivered services is being built right now, in a forum most Indian TP practitioners aren't reading transcripts of.

Why this matters more than it looks: India occupies an unusually exposed - and unusually powerful - position in this specific fight. On one hand, India has historically been among the loudest voices for expanding source-country taxing rights over digital and cross-border service income; the UN process exists substantially because countries like India argued the OECD's two-pillar solution didn't go far enough for market/source jurisdictions. On the other hand, India is also the world's largest base for AI-enabled global capability centres and IT/ITeS delivery - the exact category of cross-border service flow this protocol is trying to pin down. If the AI carve-out in the services protocol ends up defining nexus or taxing rights in a way that diverges from how OECD Pillar One and Amount B treat the same AI-delivered functions, Indian-headquartered groups and Indian subsidiaries of multinationals could find themselves benchmarking the same intercompany service flow against two different rulebooks depending on which counterparty jurisdiction is involved. That is not a hypothetical compliance headache; it is a structural one, because transfer pricing documentation is built around a single delineated transaction and a single most-appropriate-method choice, not a dual-track nexus test.

There is also a functional-analysis problem lurking underneath the treaty politics. Cross-border services protocols, going back to earlier UN work like Article 12B on automated digital services, tend to draw bright lines based on where a service is 'delivered' or 'consumed.' AI complicates that in the same way it complicates DEMPE analysis for intangibles: an AI system trained in one jurisdiction, fine-tuned or RAG-augmented with client data in a second, and delivering inference-based output to end customers in a third doesn't map cleanly onto any single-jurisdiction nexus concept the drafters are likely working from. If negotiators write a definition of 'AI services' into the protocol without engaging with how multi-jurisdictional AI pipelines actually function, they risk creating a nexus rule that transfer pricing professionals will spend the next decade trying to reconcile with functional reality - much the way 'significant people functions' language under the OECD's authorized approach took years of practice to operationalize.

The open question, and the one worth writing toward rather than around: does India's negotiating position on this protocol actually reflect an analysis of how Indian GCCs and IT exporters would be affected if AI services get a distinct, UN-defined nexus test that differs from OECD treatment - or is India's source-country advocacy here running on inertia from an earlier era of BPO and call-centre economics, before AI-native delivery models existed? That's not a rhetorical question so much as a genuine gap in the public record; the UN's own tracking shows the draft protocol text was only published in July 2026 and is still being contested clause by clause. A practitioner with both technical AI fluency and TP grounding is unusually well positioned to make that case publicly before the text hardens - which may be the real opportunity here, separate from whatever the final protocol says.

 

Who Performed the Function?

A multinational files a patent for a new compound. Tax authorities in two countries want to know which entity in the group is entitled to the royalty stream the patent will eventually generate. The test they apply is old and, until recently, reasonably reliable: find out who performed the Development, Enhancement, Maintenance, Protection and Exploitation functions for the intangible - DEMPE, in transfer pricing shorthand - because whoever bears the risk and does the deciding gets the residual profit. For decades this worked because you could point to a building: a lab, a notebook, a committee that signed off on which candidate molecule to pursue. Increasingly, the person you would point to typed a prompt, set an agent running overnight, and reviewed a shortlist of candidates over coffee the next morning. The system generated the options, ran the simulations, in some cases proposed the formulation that went into the filing. Who, exactly, performed the function? 

This is not a hypothetical irritating a handful of transfer pricing specialists. An Australian tax advisory has recently flagged what it calls a DEMPE-shaped hole - cases where AI itself is doing the value-creating work and the existing test simply has no place to put that. A parallel problem is surfacing in patent law, where AI-generated inventions are quietly breaking the doctrine of inventorship the same way they are breaking the tax doctrine of significant people functions. Two different bodies of law, built for different purposes, are running into the identical wall: both assume that value is created by an identifiable human making an identifiable decision, and both are discovering that the human's contribution has thinned to something closer to selection and approval. 

 The instinct is to treat this as a compliance problem for corporate tax departments. It is that. It is also a preview of a question that is going to reach every individual who has taken the advice - now common, and not wrong - to stop being a mere user of AI and start being a builder with it. The advice is sound as far as it goes. But it assumes the world has a stable way of recognising, after the fact, that a given piece of work was built rather than merely retrieved. Increasingly, it does not. If the design decision, the judgment call, the actual creative leap happened inside the model, and the human's visible contribution is a prompt and a sign-off, then whatever institution eventually has to decide who gets the credit - a patent office, a performance review, a promotion committee, a funding panel, or a tax auditor deciding which entity in a group actually earned a royalty - is going to ask the same question the DEMPE test is asking of multinationals right now. Where, precisely, did you sit in this? 

It gets sharper still. Some of the institutions asking that question are, at the same moment, handing the question itself to AI. A peer-reviewed study on agentic transfer pricing work found it running nearly a hundred times faster than a human analyst, and wrong in roughly one case out of five or six. That is the position builders are actually in: producing work at a pace where the artifact looks finished, inside a system that is simultaneously trying to verify whether the artifact - or the reasoning behind it - was genuinely someone's, at a speed that does not allow careful verification either. Credit and reliability are being tested at the same moment, by the same overstretched apparatus, and neither test is mature. 

None of this means the advice to build rather than merely search was wrong. It means the second half of that advice has arrived earlier than expected: it is not enough to have used AI to produce something. What will be asked, by whoever is deciding whether the work was yours, is narrower and harder - what did you decide that the model could not have decided on its own; what judgment did you make that the output alone does not show; what risk did you actually carry if the decision turned out to be wrong. That residue is what DEMPE is trying, clumsily, to locate inside multinational R&D right now, and it is what will eventually be asked of anyone whose work passed through an agent before it reached a reader, a patent examiner, or a manager. 

The practical instruction for builders is not to use AI less. It is to keep a different kind of record than most people currently bother with - not a log of what the model produced, but a trace of the decision points where a different judgment would have produced a different outcome, and where that judgment was demonstrably yours. Tax authorities are already learning, case by case, that the old test - find the person who performed the function - cannot be answered by pointing at a finished patent or a tidy set of transfer pricing documentation. It has to be answered by reconstructing a decision trail. Builders who have not been keeping one are going to discover, the way multinationals currently are, that the absence of a trail is itself the answer to the question of who performed the function - and it is rarely the answer they wanted.

Post-SAP Labs Rulings Tighten Scrutiny of Comparability Filters in TP Benchmarking

The Karnataka High Court has disposed of a batch of transfer pricing appeals that had been pending since the Supreme Court's 2023 judgme...