Showing posts with label Tax Technology. Show all posts
Showing posts with label Tax Technology. Show all posts

Tuesday, October 6, 2026

The Facebook Royalty Ruling and the Limits of AI in Transfer Pricing

AI vendors frequently market tools that compress transfer pricing benchmarking into minutes. For certain tasks the claim holds: screening a database of five hundred potential comparables down to a defensible shortlist, or drafting a first-cut local file, is a repetitive, pattern-matching exercise that large language models and agentic workflows can handle. These tasks, however, are not where transfer pricing disputes carry the highest financial stakes. Intangible valuation, royalty characterisation and cost-sharing disputes remain the hardest interpretive questions in the field, as a recent Facebook/Meta ruling from the US Tax Court illustrates.

On September 29, 2026, the Tax Court issued a supplemental opinion closing out the long-running dispute over how Facebook Ireland should compensate Facebook US for the platform technology, user-community rights and marketing intangibles it received under a 2010 cost-sharing arrangement. The court had earlier determined the present value of these contributions at roughly $7.8 billion. The remaining question was mechanical but consequential: how to convert a single lump-sum valuation into a stream of annual royalty payments over time.

The IRS sought one aggregate flat-rate royalty on rest-of-world revenue, spread evenly over 6.29 years. The court rejected this and allowed Facebook to rely on its own contemporaneous cost-sharing documentation, which split the payment into three separate streams with different bases and durations. The regulations, the court held, require the form of payment to be specified upfront, not the precise royalty base or period. The dispute turned on reading a specific set of facts against specific regulatory text, not on matching the company against a database of comparables.

This ruling is relevant to the debate on AI adoption in transfer pricing because it marks a split in the kind of work routine automation can handle versus disputes that require sustained interpretive judgment. Routine compliance work, safe-harbour eligibility checks, comparable-set generation, local file drafting, is being automated quickly, and tax authorities are moving in the same direction. India's 2026 TP rule changes already feature rule-based automated safe-harbour processes and data-analytics-driven case referral.

At the high-value end of TP, cost sharing, platform contribution valuation, royalty mechanics, documentation form-over-substance arguments, the Facebook dispute shows that resolution depends on years of adversarial expert reconstruction. Competing discount-rate models, arguments over which revenue line is too "aspirational" to count, and disagreements over whether a specific regulatory provision requires the royalty base to be stated in the agreement itself or merely in supporting documentation, are the stuff of this kind of litigation.

Current AI tools do not attempt this kind of analysis. A recently published academic stress test gives reason for caution even where they do attempt related tasks: in a controlled simulation across 20 MNE cases, agentic AI systems performing TP functional analysis and comparable selection cut processing time by 98%, but also produced a 15% error rate in complex functional characterisations, a 22% irrelevance rate in comparable selection, and a 15% hallucination rate in legal citations.

These two developments together carry a practical implication. The parts of TP work being automated fastest are the parts where errors are least costly: an imperfect comparable set in a routine ITeS benchmarking study rarely puts significant tax at stake. The parts of TP work where errors are most costly, intangible valuation, royalty structuring, the kind of judgment call at the centre of the Facebook ruling, are precisely where current AI tools, per the error rates above, are least reliable. Courts in such disputes examine competing expert methodologies line by line rather than defer to any formulaic shortcut.

For Indian GCCs, IP-holding structures, and groups running long-duration cost-sharing or licensing arrangements with affiliates, AI-assisted documentation should be treated as a first draft rather than a substitute for the granular factual and methodological rigour this case rewarded. If tax administrations deploy agentic AI for risk-scoring or functional analysis in audit selection, the same 15-22% error and hallucination rates apply on their side as well. This raises governance questions, including traceability, human-in-the-loop sign-off, and explainability, that Indian TPOs and the CBDT have not yet had to address publicly but may need to as data-analytics-driven case selection expands.

Monday, October 5, 2026

Royalty or Profit-Shifting: Australia's Risk-Zone Test and Its Implications for Indian TP Characterisation Disputes

Cross-border payments for software, brand use, or know-how raise a recurring transfer pricing question: how much of the payment compensates the foreign parent for genuine IP use, and how much constitutes profit shifted out of the local entity. India has addressed this question primarily through tribunal litigation. Samsung, Sony, Vodafone Idea and other taxpayers have contested the same AMP/royalty issue across multiple assessment years, with outcomes turning on functional characterisation: whether the Indian entity operated as a full-risk licensed manufacturer or a disguised contract manufacturer, and whether the royalty comparable used by the TPO was genuinely similar. In its August 2026 ruling on Samsung India, the Delhi ITAT found that the TPO had benchmarked a consumer electronics brand licence against royalty agreements from the agricultural and biotech sectors, a comparison the tribunal rejected along with roughly ₹7,800 crore in other adjustments.

Australia has taken a structurally different approach. On 4 September 2026, the ATO finalised Taxation Ruling TR 2026/2, addressing when cross-border software and IP payments constitute royalties, and released draft Practical Compliance Guideline PCG 2026/D4 alongside it. The PCG sets out a five-zone, colour-coded risk framework, from white to red, that software distributors use to self-assess their royalty withholding exposure.

The zones rest substantially on a quantitative trigger: an Australian distributor's operating margin relative to its global group's margin, with a 10-percentage-point band used as a proxy for whether an embedded royalty sits inside an otherwise 'royalty-free' distribution arrangement. Consultation on the draft PCG closed on 2 October 2026. The same package included a Decision Impact Statement on the Oracle case, confirming that MAP and treaty arbitration remain available even where domestic litigation on the same royalty question is still running in parallel.

The two approaches differ in method and cost. India's approach is precedent-driven and functionally granular: each case requires its own FAR analysis, its own comparable search, and its own tribunal hearing, sometimes across a decade of appeals before the characterisation question is settled for that taxpayer and that year. Australia's approach is administrative and mechanical: a margin threshold either trips a risk flag or it does not, without relitigating functional characterisation in every case.

Each method has a cost. India's approach protects taxpayers from being classified on crude proxies, but generates litigation overhead that both the department and taxpayers have raised concerns about for years. Australia's approach is more scalable and predictable, but practitioners have already flagged that a single margin band applied uniformly across SaaS, cloud and subscription models may catch arrangements unrelated to embedded royalties.

The comparison is relevant to Indian practice independent of the Australian context. CBDT has, over the last two budget cycles, moved India's TP regime toward more systemised and less discretionary processes: block assessments that carry a TPO's margin determination forward two years, a rationalised safe harbour regime with fixed bands by transaction category, and a stated intent to limit automated-analytics flags from triggering full scrutiny unless paired with other evidence of evasion. This is consistent with a shift toward rules-based risk-zoning, though the department has not framed it in those terms.

If CBDT, or an AI-assisted TPO, were to adopt a quantitative first pass for royalty and FTS characterisation along the ATO's lines, drawing on CbCR and segmental data the government already collects, such a system could replace inconsistent manual judgment with predictable, published zones. It could equally reproduce the same reliance on a mismatched comparable that the Samsung tribunal had to correct, embedded instead in a model that taxpayers outside the department cannot audit.

The ATO's framework was put out for public comment; an Indian equivalent, if one emerges, may not follow the same consultative process. For TP practitioners, that difference in process, more than the underlying royalty question, is worth tracking as India's own risk-assessment tools develop.

Sunday, October 4, 2026

AI-Based Comparable Screening Does Not Address India's Core Transfer Pricing Disputes

At the WU-TA Advanced Transfer Pricing Programme in Singapore this week, PwC's transfer pricing partners presented AI-based comparable screening as a significant efficiency gain for transfer pricing teams, and also set out its limits. One partner compared AI to a GPS: given the wrong address, it will still get you there efficiently, but a human has to recognise that the destination itself is wrong. A tax head from P&G compared AI's output to that of a junior associate: useful for a first draft, but requiring supervision, risk assessment and quality control before anyone relies on it.

Comparable screening once meant examining databases or company websites one company at a time. AI tools can now gather that information simultaneously, cutting the search stage substantially. Participants at the conference did not dispute that this is a genuine efficiency gain. The question from an Indian perspective is what that gain addresses within the broader TP dispute process.

Indian transfer pricing litigation in 2026 continues to turn on characterization issues rather than comparable selection. The Delhi ITAT's ruling this week in Discovery Communication India illustrates this. The taxpayer had adopted TNMM, its margin came in above the comparables' average, and the TPO did not challenge the comparable set.

The dispute concerned whether the TPO could carve out AMP expenditure already included within the accepted operating cost base, treat it as a separate international transaction, and re-price it using the Bright Line Test. The Tribunal rejected this approach, following the line of Delhi High Court and ITAT precedent in Sony India, Casio, and the Special Bench ruling in LG Electronics. The reasoning turned on the terms of the intercompany agreement, whether a genuine cost-sharing arrangement existed, and whether the expenditure benefited the taxpayer or the foreign AE. Comparable benchmarking played no part in this analysis.

Much of Indian TP litigation that consumes years and significant tax amounts follows this pattern: captive power CUP arguments, GCC or FAR mischaracterization, receivables treated as separate transactions, and tested-party selection disputes. These are disputes about judgment on facts, not about identifying additional comparable companies.

AI-based comparable screening does not reach this category of dispute. If the contested, revenue-significant disputes in India are predominantly characterization disputes, concerning whether a transaction exists at all, whether an entity is genuinely a captive, or whether expenditure qualifies as AMP, a tool that shortens the comparable search by some weeks does not touch the stage of the process where litigation risk is concentrated.

Such tools make the uncontested, mechanical part of TP documentation cheaper and faster, which has value. That is a different claim from suggesting AI will reduce TP disputes or TP risk, a claim made in vendor marketing and at conference panels, including at the Singapore event. The P&G comparison of AI to a junior associate acknowledges this limitation.

The GPS comparison makes the same point more directly. It concedes that AI's efficiency is irrelevant, and potentially unhelpful, if the underlying functional characterization (the "address") is wrong, and that getting the characterization right remains a human judgment task.

The implication for Indian tax administration concerns where AI-based risk scoring may next be deployed. Tax administrations have tended to adopt AI tools some time after industry does, and CBDT or TPOs could deploy similar AI benchmarking tools to accelerate comparable searches and sharpen risk-based case selection.

If AI's comparative advantage stops where functional characterization begins, as the Singapore discussion suggested, AI-assisted TPOs may build adjustment cases faster on the same mechanical grounds that are already failing in the Tribunals, without improving their ability to predict which cases will survive judicial scrutiny on characterization. Whether any firm or tax administration is developing an AI model trained specifically on characterization outcomes, such as AMP treatment, captive recharacterization, or DEMPE allocation, rather than on comparable company financials, was not evident from the Singapore discussion.

For TP practitioners, the practical implication is to assess AI tools by the stage of the process they improve. Faster comparable searches do not reduce exposure on characterization issues, which continue to require substantive factual and legal analysis, and documentation strategy should treat that distinction as material.

Saturday, October 3, 2026

Flyjac Logistics: Bombay High Court Holds That a System-Generated Notice Is Not Proof of Service in Transfer Pricing Proceedings

Transfer pricing law requires a Transfer Pricing Officer (TPO) to issue a show-cause notice before determining an arm's length price adjustment, giving the taxpayer a genuine opportunity to respond. Courts have generally dealt with breaches of this requirement where a notice was lost in transit, never issued, or deliberately bypassed. The Bombay High Court's recent ruling in Flyjac Logistics' case addresses a different failure mode: the notice was generated and logged by the department's own system, yet never reached the taxpayer because the system malfunctioned.

The TPO proposed a ₹20.16 crore adjustment to Flyjac's arm's length price and recorded that a show-cause notice had been issued and had gone unanswered. Flyjac stated that it never received any such notice and learned of its existence only because the TPO's order referred to it. In its affidavit, the Department attributed this to a technical failure rather than any deliberate lapse: the email carrying the notice had bounced, an unexplained glitch in the ITBA backend kept the notice from appearing on Flyjac's e-filing portal, and no SMS alert was triggered either.

The Bombay High Court held that the statutory show-cause requirement under Section 92C(3) is distinct from the information-seeking notices issued earlier under Section 92CA(2), and that the order could not stand without actual service on the taxpayer. A system log recording that a notice had been issued did not, in the court's view, satisfy that requirement.

The ruling can be read alongside other developments this year that touch on transparency and procedural rigour in tax administration, though these do not point to a single established trend. Chile's SII has published, in some detail, the risk criteria and audit yield behind its enforcement program. HMRC has tightened the link between its TP compliance guidance and its formal error-correction facility, creating a more structured route from self-identified error to resolution. India's Supreme Court, in the Samsung India matter, dismissed the Revenue's appeals because of unexplained delays of several hundred days in filing, without reaching the underlying transfer pricing questions. Taken together, these developments suggest that as tax administrations rely more on systems, logs and published criteria to manage scale, courts and taxpayers are likely to scrutinise those systems more closely rather than less.

CBDT has publicly stated its intention to expand automated risk-flagging through initiatives such as Project Insight and NUDGE, and the broader direction of travel is toward AI-assisted scrutiny. Tax-technology vendors, including infer360 (now allied with Nexdigm) and comparable platforms, market tools for machine-assisted TP monitoring and documentation. Flyjac illustrates a specific risk that arises as these systems take on a larger role: a system can generate a notice, log that it generated the notice, and still fail to deliver it, and such a failure may not surface until the taxpayer is already contesting the resulting order in court. As automated systems take on more of the notice-generation and routing function in TP enforcement, similar service failures are likely to recur as adoption increases.

For practitioners, the ruling raises a practical evidentiary question: whether a system log showing that a notice was issued will be treated as sufficient proof of service, or whether courts will continue to require affirmative proof of actual receipt, as the Bombay High Court did in this case. Until that question is settled through further litigation, TP compliance functions may find it prudent to maintain independent records of portal access, email delivery and SMS alerts, so that service failures originating in the department's own systems can be identified and raised at the earliest stage of any proceeding.

Friday, October 2, 2026

ITAT Mumbai's Captive Power Ruling and the Limits of AI-Driven TP Benchmarking

A recent Mumbai ITAT ruling in a captive power transfer pricing dispute turned on a question that AI benchmarking tools are not built to ask: whether the comparable a taxpayer needed was already sitting inside its own books.

The case involved a cement manufacturer whose captive power plants supplied electricity to its own cement units. The dispute arose under the 2013 amendment that brought specified domestic transactions within the arm's length pricing framework. The Revenue argued that the captive generator had to be benchmarked against the rate at which other generators sell to state distribution licensees. The taxpayer argued that its own cement units paid documented market rates to state power distribution companies for electricity purchased on the open market, and that this internal rate was a valid comparable for what the captive plant charged them. The tribunal agreed. It held that the amendment does not require a generator to be benchmarked only against another generator, or that a distribution licensee's procurement rate must invariably be adopted, and that CUP selection remains governed by the transaction-specific requirements of section 92C and Rule 10B. Revenue's appeals against adjustments of roughly Rs. 64 crore and Rs. 43 crore across two years were dismissed.

The outcome is less significant than the source of the winning comparable. It was not retrieved from a database of unrelated third-party companies, the kind of external, industry-classified, margin-screened universe that commercial TP benchmarking tools are built to search. It was already in the taxpayer's own transaction records: the price its own units paid the grid for the same commodity, in the same period, in the same geography. An external comparables search, however well it ranks functional similarity, would not have surfaced this fact pattern, because the relevant exercise was not searching further afield but looking at what the group itself was already paying for the identical input elsewhere in its operations.

This points to a gap in how AI benchmarking tools are currently designed. Internal CUPs sit at the top of the comparability hierarchy because they avoid much of the functional-comparability subjectivity that external searches have to approximate statistically. Yet most AI tooling marketed to TP teams is built to screen external databases faster, not to mine a group's own intercompany and third-party transaction data for an internal benchmark. These tools automate work that used to be slow, external database screening, but they are not generally designed to surface the kind of internal comparable a case like this one turned on.

A separate observation from this filing season's TP engagements reinforces the point. Some practitioners report that adjustments are more often lost on evidence discipline than on legal interpretation: the comparable search cannot be reproduced, or the Local File and the counterparty's file tell different functional stories. Reproducibility of an external search is a baseline requirement. It says nothing about whether a better internal comparable was considered at all.

For TP practitioners evaluating AI benchmarking tools, this ruling is a useful prompt to ask a specific question: does the tool query the group's own ERP and intercompany transaction history for internal comparables before it touches an external database, or does it only speed up external screening under a new label. Before commissioning an external search, it is worth checking whether the answer was already in the ledger.

Thursday, October 1, 2026

The Asymmetry in AI Use Between Tax Authorities and Taxpayers in Transfer Pricing Defense

A recent Tax Notes article on a Turkish transfer pricing case traces the dispute through to its downstream effect on a U.S. foreign tax credit claim. In doing so, it sets out an asymmetry that deserves more attention in transfer pricing practice: Turkey regulates how professional advisors may use generative AI in giving tax advice, but places no comparable constraint on its own tax administration's use of machine analysis for risk-scoring, flagging related-party anomalies, or selecting cases for audit. The rules on the taxpayer side are specific and restrictive. The rules on the authority side are largely unwritten.

The same week the Turkey piece appeared, a Manila-based practitioner column described how the Philippines' Bureau of Internal Revenue has operationalised a system-assisted, risk-based audit selection framework under Revenue Memorandum Order 1-2026. The indicators are aimed squarely at related-party transactions: persistent losses against strong revenue, tax-to-sales ratios that look too low, and heavy reliance on a single related counterparty. Under this framework, transfer pricing risk is pre-flagged by the system before an examiner opens the file. Read together with Turkey's dual posture, restrictive on taxpayer-side AI reliance and expansive on authority-side machine analysis, the two examples suggest the asymmetry is not confined to one jurisdiction.

The same question arises for anyone advising Indian multinationals or GCCs. India's safe harbour election process is moving toward an automated, rules-based model, CBDT's compliance apparatus already uses AI-assisted risk profiling, and the APA and audit infrastructure is becoming steadily more analytics-driven. All of this sits on the authority side of the same asymmetry described in the Turkish and Philippine examples. Several Big 4 and boutique TP technology vendors have published material this year describing increased use of generative AI in preparing local files, benchmarking memoranda, and functional analyses. Whether that assistance will be treated on the same footing as a signed opinion from a human expert, when the question becomes whether a position was taken in good faith or whether a reasonable-cause defence survives scrutiny, is not yet settled by any rule or ruling in India.

The Turkish case does not resolve this question. Its value lies in naming the gap explicitly, rather than treating AI in tax administration as a single undifferentiated trend of efficiency gains for everyone. Whether the standards governing reliance, documentation, and reasonable cause will develop in step for both tax authorities and taxpayers, or whether taxpayers will end up defending AI-informed positions against AI-generated risk scores under rules drafted for one-sided use, remains open.

For Indian practitioners, the practical implication is to document how AI tools are used in preparing local files, benchmarking analyses, and functional analyses now, before the question is tested in audit or litigation, so that the basis for any position can be explained and defended independently of the tool used to generate it.

Wednesday, September 30, 2026

India's Automated Safe Harbour Approval and the Administrative-Law Question Canada Is Now Asking About AI-Driven Audit Selection

Budget 2026 moved Safe Harbour approval for IT services onto an automated, rule-driven framework and removed the requirement for an officer to examine the application. Applicants who meet the prescribed margin thresholds receive approval without human review, and the outcome is locked in for up to five years.

A recent Canadian Tax Journal paper by Theertha Narayanan and Pramod Kumar Siva examines a related but distinct question: whether the Canada Revenue Agency's use of AI-based risk scoring to select transfer-pricing files for audit must satisfy the same procedural-fairness standard that Canadian courts apply to other administrative decisions. The paper argues that it should, particularly after Canada's 2025 federal budget introduced stricter TP methodologies, expanded recharacterization powers, higher penalty thresholds and shorter documentation deadlines, which raised the consequences of being selected for audit. On the paper's argument, the selection decision itself should be reviewable under the reasonableness standard set out in the Supreme Court of Canada's Vavilov framework: justified, transparent and intelligible.

The Canadian debate concerns whether an algorithm can lawfully flag a file for human review. India's Budget 2026 reform goes further: it has let an algorithm replace the human reviewer for a determination that carries up to five years of certainty on arm's length pricing. Indian commentary on the change does not appear to have framed it as raising a comparable administrative-law question.

The Canadian paper's argument rests on specific case law rather than general concerns about AI. It draws on Vavilov's requirement that administrative reasoning be internally coherent and "justified in relation to the facts and law that constrain the decision maker," and on Dow Chemical's confirmation that discretionary CRA decisions under the Income Tax Act are reviewable on that same reasonableness standard. The paper's contribution is to apply these doctrines to a system that produces a risk score rather than a chain of reasoning, and to ask whether a score alone can meet a standard that requires justification.

The Indian changes are comparably specific. Budget 2026 pushed Safe Harbour approval into what industry commentary has called an "auto-pilot mode": applications are processed through a fully automated framework without officer discretion, in exchange for locking in outcomes for five consecutive years. The stated rationale, reducing subjective review, reducing litigation and increasing predictability, is consistent with the reasoning most jurisdictions offer for automating tax administration. Automating an approval, however, is a different act from automating a flag for human review, and that difference is what the Canadian paper is built to examine.

The practical stakes for Indian practice are concrete. GCCs are the primary beneficiaries of the new Safe Harbour automation, and the pitch to them has centred on certainty: file the declaration, receive automatic approval, and bypass officer review. An approval granted without examination is, in substance, an algorithmic determination of an arm's length outcome. It is made without safeguards that the international literature on automated decision-making treats as significant, including explainability, an audit trail showing why a given margin was accepted, and a documented basis for treating a taxpayer's declared facts as sufficient without human verification.

India has no published equivalent of the Vavilov standard for testing whether a rule-driven tax decision is reasonable. Nor is there public discussion, so far, of what a taxpayer can do if an automated approval later turns out to rest on a misclassification the system had no way to detect. On most commentary on the reform, officers who did not examine the application at the front end can still reopen the matter on audit later. That leaves taxpayers with a certainty that appears binding on its face but may not bind the department in substance, because no human made a reviewable decision at the outset.

Whether Indian tax administrative law needs an equivalent of the reasoned-decision requirement that Canadian courts are now applying to algorithmic audit selection remains open. If Budget 2026 signals a broader shift toward rule-driven, officer-free approval extending to APA processing or audit selection, the question will carry more weight, though the current material does not establish how far that shift will go. There is also a defensible counter-view, that Safe Harbour was always meant to operate as a self-assessment regime in which the absence of officer discretion is a deliberate feature rather than a due-process gap. Practitioners advising on Safe Harbour elections should flag to clients that automated approval does not necessarily foreclose later audit scrutiny, and should treat the administrative-law framing, not merely the compliance mechanics, as part of the risk assessment.

Tuesday, September 29, 2026

From Annual Documentation to Continuous Monitoring: What the Nexdigm-infer360 Alliance Signals for TP Practice

Transfer pricing documentation in India is built after the fact. Benchmarking studies are prepared once a year, local files are frozen at a point in time, and the arm's length position a taxpayer defends before a TPO reflects a snapshot taken months or years after the transactions occurred. The analysis need not be wrong, but it is backward-looking by design: assembled once the business has already priced the transaction, to justify a decision already taken.

A new generation of AI-native platforms is pitching something different: continuous, always-on monitoring of intercompany pricing through the year, rather than faster annual documentation. Nexdigm, the Mumbai-headquartered advisory firm, has recently announced an alliance with infer360, a Singapore-based AI platform built by former Big Four transfer pricing partners. The stated aim is to help clients move from periodic compliance to continuous transfer pricing management: automating documentation, running ongoing risk monitoring, and maintaining an audit-ready trail through the year rather than reconstructing one afterward.

The marketing language is unremarkable; most vendors in this space now describe themselves as AI-native. What is worth noting is that a serious Indian-origin advisory practice with established APA, controversy and operational TP credentials is putting its own resources behind the view that continuous monitoring, rather than faster annual studies, is where the market is heading.

This sits alongside a parallel move on the US side of the industry, where an established transfer pricing technology vendor's benchmarking and documentation AI suite continues to draw trade-press attention as a template for how mid-market and boutique practices are narrowing the technology gap with the Big Four.

The relevance for practitioners lies in the kind of evidentiary record these platforms are built to produce: a clean, time-stamped account of the methodology used and how it performed against comparables through the year. That is the kind of record Indian tribunals have shown they will act on.

In a Delhi ITAT ruling this month in Honda R&D (India)'s case, the taxpayer faced a ₹50.20-lakh adjustment for AY 2020-21 and pointed the Tribunal to a TPO order for the following assessment year, AY 2021-22, in which the department had accepted the identical benchmarking position it was disputing for the earlier year. The Tribunal directed relief consistent with the later year's accepted treatment.

This is not an isolated result. Indian tribunals have repeatedly leaned on a consistency principle when a TPO's own later-year acceptance undercuts an earlier-year adjustment. A continuous-monitoring platform, by construction, produces the kind of multi-year, methodology-stable record that makes this argument easier to run: the more granular and continuous a taxpayer's own TP data trail becomes, the stronger its consistency-principle arguments are likely to be.

This also raises a question about asymmetry between well-resourced multinationals and the tax administration. If sophisticated taxpayers build continuous, audit-ready monitoring systems while the department still audits largely on an annual, backward-looking cycle anchored to Form 3CEB-style filings, an information and preparedness gap could open up between the two sides of the table.

India's income tax apparatus already runs AI-based risk assessment for scrutiny selection at the return level, and the CBDT's APA programme leans heavily on post-agreement monitoring of critical assumptions for bilateral agreements. From there, it is a reasonable question whether the department should build its own continuous-monitoring capability specifically for TP, to match rather than only react to the tooling that advisory firms are now selling to taxpayers.

There is also a discovery-related question worth flagging. As continuous monitoring platforms generate richer contemporaneous records than the old annual documentation model, those records could prove double-edged: they may help taxpayers win consistency arguments in years like this one, but they could also give revenue authorities a more granular trail to probe when a deviation does appear.

Neither the vendors marketing these platforms nor the tribunals applying the consistency principle have had occasion to address this tension yet. As continuous monitoring tools become more common, practitioners advising on TP documentation strategy will need to weigh the benefit of a stronger contemporaneous record against the risk that the same record gives the department more material to examine when a taxpayer's position shifts from one year to the next.

Monday, September 28, 2026

Karnataka HC's SAP Labs Remand Ruling and the Case for Reasoned Comparability Filters in AI Benchmarking

The Karnataka High Court's 2018 decision in Softbrands treated the selection and rejection of comparables as a pure question of fact, which left the ITAT as the final word on comparability disputes; High Courts could not reweigh turnover filters, RPT thresholds, or FAR analyses under Section 260A. The Supreme Court's ruling in SAP Labs India v. ITO rejected the idea that every ALP determination is immune from scrutiny and sent a batch of pending appeals back to Karnataka for fresh consideration on the merits. How the High Court would apply that reopened scrutiny remained unclear until its recent consolidated ruling.

In its 28 August 2026 consolidated ruling, the Karnataka High Court restated the SAP Labs test and attached specific numerical thresholds to it. A ₹200-crore upper turnover filter was held to be rational and logical, not because the statute prescribes it, but because size differences plausibly correlate with brand value, bargaining power, and economies of scale. A 15% RPT (related-party-transaction) filter was endorsed as the ordinary default, with a departure to 20% or 25% permitted only where the authority records a specific, reasoned finding that comparables meeting the lower threshold are scarce. The Court also held that the ±5% tolerance range under Section 92C is a threshold, not a standing deduction that a taxpayer can claim once the transaction price exceeds it. Within days, at least two more Karnataka HC benches cited this ruling to dispose of pending Revenue appeals, including one that upheld the exclusion of Bodhtree Consulting as a comparable by applying the same 200-crore ceiling.

Barely two weeks after the Karnataka HC ruling, ITAT Delhi reached a different outcome in GE India Industrial v. DCIT, rejecting the Revenue's use of a rigid 50% turnover band to strike out lower-end comparables and criticising the DRP for being "guided by a general opinion" rather than the statutory FAR-based comparability test under Rule 10B. Read together, the two rulings point to a common test: a filter survives scrutiny only if the tribunal or officer applying it shows a specific, reasoned link between the chosen threshold and functional comparability on the facts of the case. A 200-crore ceiling with a size-based rationale passes. A blanket 50% band applied mechanically, without engaging the functional analysis, fails.

These rulings have implications for how AI tools are used in TP documentation. Benchmarking platforms, whether Big 4 proprietary tools or newer AI-native SaaS products, increasingly automate comparable searches using statistical filters such as turnover bands, RPT thresholds, and quartile screens. A vendor could hard-code the thresholds recently upheld, such as 200 crore and 15%, as default parameters on the assumption that a threshold accepted once will be accepted again.

That approach risks the failure mode both rulings identify from opposite directions: a threshold applied without a documented, case-specific rationale is what gets struck down, whether the Revenue or the taxpayer relies on it. An AI benchmarking tool that outputs a comparables set with a turnover filter applied, but no accompanying explanation tying that filter to the tested party's specific brand value, IP ownership, or scale economics, is building a study that looks defensible today and may not survive the next remand.

TP teams, and the vendors building the AI layer underneath them, may need to consider whether benchmarking software should generate a reasoned-finding log alongside every filter it applies: not merely that a comparable was excluded because turnover exceeded 200 crore, but a documented basis for why that threshold was chosen for the specific tested party, FAR profile, and data set. The Karnataka High Court's ruling signals that courts are now willing to test that reasoning on the merits rather than treat comparable selection as an unreviewable factual finding.

For AI-assisted TP work, the practical advantage may lie less in the speed of comparable identification than in the tool's capacity to produce explainable, fact-specific justification that a Division Bench applying the SAP Labs test would accept.

Sunday, September 27, 2026

ITAT Hyderabad Extends BAPA Margin to Non-Covered AE Transactions on FAR-Identity Grounds

A Bilateral Advance Pricing Agreement negotiated with the CBDT and a foreign competent authority, usually the US given where most AE relationships sit, can cover more than 90 percent of a captive service provider's international transactions. The remainder, revenue earned from AEs in other jurisdictions such as the UK or Singapore, falls outside the agreement's scope. This residual portion must be benchmarked afresh each year and remains open to TPO scrutiny, since the certainty negotiated under the BAPA does not formally extend to it.

The ITAT Hyderabad Bench addressed this fact pattern in Synchrony International Services Private Limited v. ACIT, an order pronounced on 30 March 2026. Synchrony had a BAPA with the US covering roughly 95.75 percent of its revenue as a captive ITeS provider. The remaining 4.25 percent came from non-US AEs and sat outside the agreement. The TPO did not conduct a separate benchmarking exercise for this residual portion and proposed an adjustment on it. The Tribunal held that a BAPA margin negotiated for one country's AEs cannot be restricted to that country alone where the functions, assets and risk (FAR) profile of the non-covered AE transactions is identical, and directed the TPO to apply the BAPA rate across all three assessment years under appeal.

The outcome favours taxpayers: if a captive performs identical back-office work for a UK entity and a US entity, the pricing need not differ merely because only one entity's AE relationship is covered by the signed APA. But the ruling shifts the locus of the next dispute rather than removing it. The burden of proof on non-covered transactions is relocated, not eliminated.

Instead of a fresh comparables search each year, the taxpayer must now build and defend a case that the FAR profile across AEs is genuinely identical, covering service descriptions, decision rights, risk allocation, contractual terms, and reporting lines. This is a different exercise from a standard benchmarking study, and in some respects a harder one, because there is no external database of third-party comparables to draw on. The comparison is intra-group, AE to AE, and the evidence must come from internal documentation: service agreements, organisation charts, SLAs, cost allocation keys, and correspondence showing who directed the work.

Current AI-enabled TP tools, including benchmarking platforms that automate comparable searches and NLP tools that flag inconsistent documentation, are built mainly to address a different problem: finding and screening third-party comparables faster, or checking a local file against a jurisdiction's formatting requirements. Few, if any, are designed to assess whether AE-A's functional profile is identical to AE-B's functional profile within a single multinational group. That is a more bespoke comparability question, closer to internal audit than to database screening, and it now sits at the centre of the scope this ruling has opened.

India's APA programme has crossed 1,034 agreements since inception, with 284 bilateral, a population of taxpayers who may now have grounds to extend negotiated certainty to residual AE transactions. Doing so will require a FAR-identity case capable of withstanding scrutiny on points such as a contractual clause, headcount difference, or decision-rights nuance that could break the claim of identity.

CBDT has not yet addressed BAPA scope in this context. The Board has previously issued administrative clarifications where APA and Safe Harbour regimes interact awkwardly; the March 2026 Office Memorandum permitting taxpayers with UAPAs spanning the Safe Harbour transition to opt into the new regime for later years is a precedent for this kind of housekeeping.

A similar clarification on BAPA scope, an administrative mechanism to formally extend or fast-track non-covered AE transactions where FAR identity is not seriously disputed, could save taxpayers from re-litigating this question bench by bench, year by year, across every captive with a partial BAPA. Absent such clarification, Synchrony is likely to become a citation that captives with partial-scope BAPAs raise routinely, and one that TPOs will need to engage with on the merits rather than dismiss at the threshold.

For practitioners advising captives with partial-scope BAPAs, the practical task is to assemble FAR-identity documentation now, before the TPO raises the issue, rather than treat the Tribunal's reasoning as self-executing.

Saturday, September 26, 2026

TR 2026/2: Australia's Software Royalty Ruling and Its Implications for Groups with Indian Operations

On 4 September 2026, the Australian Taxation Office finalised Taxation Ruling TR 2026/2, along with a draft Practical Compliance Guideline, setting out its position on when cross-border payments for software, SaaS access, and related IP arrangements amount to 'royalties' subject to withholding tax. The ruling took five years to finalise, replacing a position that traces back to the withdrawal of an older ruling in 2021. It follows the Australian High Court's decision in PepsiCo, which the ATO has relied on to justify a substance-over-form, purpose-focused test for what counts as a royalty.

The fact pattern is close to one India's Supreme Court resolved in 2021, in Engineering Analysis Centre of Excellence v CIT. That decision ended nearly two decades of litigation by holding that payments to non-resident software suppliers for resale or use under distribution agreements and end-user licences are not royalty payments, because what is transferred is a copyrighted article, not an interest in the copyright itself. India adopted the narrow, taxpayer-favourable reading. TR 2026/2 goes the other way, and further: it takes the position that where IP rights are practically inseparable from the other commercial rights in a software intermediation arrangement, which the ATO treats as the common case, the entire payment is characterised as a royalty, with no workable apportionment pathway in the final text. The draft compliance guideline layered on top classifies most royalty-free distribution structures as medium-to-high risk unless the local margin clears a threshold meaningfully higher than the ATO's own prior benchmark for low-risk distributor margins.

The commercial reality of how software gets distributed cross-border has not changed. What has changed is the doctrinal lens two major tax administrations apply to that reality, and the two now point in opposite directions. This matters because many multinational software distribution structures are built on a common template rolled out across several jurisdictions at once. A group that priced its Indian and Australian distribution entities on the same functional analysis and the same intercompany agreement now faces two different downstream outcomes. In India, the royalty question was settled years ago, so the TP analysis stands on its own. In Australia, the royalty characterisation question sits ahead of the TP question and can override it: if the whole payment is a royalty, the arm's length distribution margin analysis becomes secondary to a withholding tax exposure that was never priced into the original structure. MinterEllison's analysis of the final ruling described this as a 'full-royalty rather than apportioned outcome' as the default assumption, a marked hardening from where the draft guidance started.

Whether this is an Australian idiosyncrasy or an early sign of a broader trend is not yet clear. Some revenue authorities, facing fiscal pressure, have shown greater willingness to use substance-based, purpose-focused tests to recharacterise payments, an interpretive approach not unlike the one behind India's GAAR and its expanded commercial substance doctrines. If more jurisdictions begin reading embedded software royalties this expansively, Indian multinationals licensing software into distribution networks abroad, and Indian subsidiaries of global software groups receiving payments from Indian customers, may need a characterisation review ahead of every TP benchmarking exercise, rather than a periodic refresh of comparables alone. For now, groups with common distribution documentation across India and Australia should map where the two jurisdictions' positions on the same contract diverge, rather than assume that a settled Indian royalty position carries over wherever the same paperwork is used.

Friday, September 25, 2026

When the Service Provider Is an Algorithm: What Chapter VII's AI Problem Means for Indian Captives

A captive Indian entity, such as a GCC, a KPO, or a testing centre like the one at the heart of the Honda R&D India ruling this month, performs a defined function for its overseas parent. The transfer pricing analysis asks three questions: whether a service was rendered, whether the recipient derived a benefit, and what cost-plus markup reflects an arm's length charge for that service.

This is the architecture of Chapter VII of the OECD Transfer Pricing Guidelines, and it underlies most of India's Safe Harbour Rules, most APAs for ITES/BPO structures, and a large share of the ITAT docket. It works cleanly as long as the item being delivered is recognisably a service: bounded, routine, and separable from any underlying intangible.

In June 2026, the OECD's Working Party 6 released a discussion draft proposing a substantive rewrite of Chapter VII, including a new accurate-delineation analysis, an expanded benefit test, a sharper shareholder/stewardship distinction, and new guidance on the boundary between a service and an intangible. Comments closed on 22 July, and the OECD published the full set of responses on 24 August: more than one hundred submissions from businesses, industry bodies, and advisory firms.

A recurring theme across those submissions was a call for clearer distinctions between intra-group services and intangible transfers, particularly for AI-enabled and digital service models. This suggests that practitioners find the existing framework difficult to apply to AI-delivered outputs, a chapter otherwise treated as settled doctrine since the 1990s.

The captive/GCC model is central to India's transfer pricing practice, not a peripheral segment of it. Thousands of Indian entities are remunerated on a cost-plus basis for the kind of routine, human-performed functions that Chapter VII was built to price: testing, market research, back-office processing, and KPO analytics. Many of these entities are, in 2026, incorporating agentic AI into the delivery of these functions, which changes what sits inside the service line item rather than replacing it.

If a GCC's cost-plus-remunerated data analytics service is increasingly produced by an AI agent trained on group data, proprietary models, or a parent's algorithms, the question is whether the Indian entity is still delivering a 'service' in the Chapter VII sense, or something closer to an output derived from an intangible it does not own. That distinction could imply a different pricing method, a different DEMPE analysis, and possibly disqualification from the Safe Harbour margins (now consolidated into the unified 15.5% IT Services band) that assume a routine, low-risk service function.

Two other developments this week bear on the same issue. Turkey's Tax Inspection Board launched an AI system in August built specifically to flag transfer pricing risk in related-party transactions, which suggests some revenue authorities are building AI capability aimed at this kind of transaction. On the advisory side, Nexdigm's alliance with the AI-native TP platform infer360 shows Indian mid-tier firms building AI-TP capability of their own, partly in anticipation of this characterisation issue becoming a live audit question rather than an academic one.

Whether the CBDT should clarify, before the 2027 filing season, whether AI-enabled service delivery inside a GCC/ITES structure remains within the Safe Harbour and cost-plus framework, or whether it needs a separate carve-out, is a question worth raising now. The Chapter VII revision will not be finalised until after the November 2026 Paris consultation, so India has a window to shape its own domestic position rather than import whatever the OECD eventually adopts.

Given how much of India's outbound service economy sits on this fault line, practitioners advising GCC and ITES clients would do well to flag the characterisation risk in current TP documentation, ahead of any assessment order that forces the issue.

Thursday, September 24, 2026

Post-SAP Labs Rulings Tighten Scrutiny of Comparability Filters in TP Benchmarking

The Karnataka High Court has disposed of a batch of transfer pricing appeals that had been pending since the Supreme Court's 2023 judgment in SAP Labs sent them back for fresh consideration. The appeals turned on a familiar but consequential question: how much latitude a High Court has to examine the comparability filters, such as turnover bands, related-party-transaction (RPT) thresholds and export-ratio cutoffs, that determine which companies are shortlisted for TNMM or CPM benchmarking.

Under the Karnataka High Court's earlier Softbrands precedent, comparability disputes, including which companies get excluded and which filters apply, were treated as pure questions of fact on which a High Court would not ordinarily interfere. SAP Labs removed that immunity. In this month's ruling, the Court held that it can examine whether the selection of comparables and the choice of filters was made judiciously and on the basis of relevant material, though it will continue to defer to the Tribunal unless the process is shown to violate Section 92C or Rule 10B, or is otherwise perverse. On the facts, the Court upheld a Rs 200 crore upper turnover filter as rational given the size, brand value and economies of scale of the tested party, and held that a 15% RPT filter is ordinarily appropriate. A higher threshold of 20% or 25% is permissible only where the TPO records a specific finding explaining why comparables meeting the lower threshold were unavailable. The Court also held that the tolerance band of plus or minus 5% under Section 92C is not an automatic standard deduction from the arithmetic mean; it is a threshold that determines when an adjustment is required, not a discount to be applied as a matter of course.

In the same week, the Delhi ITAT reached a comparable conclusion in GE India Industrial's case. The Tribunal rejected the TPO's rigid 50% turnover filter and held that functionally comparable companies cannot be excluded on size grounds alone; comparability must be tested against the functions, assets and risks framework under Rule 10B, not against an arbitrary turnover band. Read together, the two rulings direct TPOs and taxpayers to support each filter with a documented, fact-specific rationale tied to the taxpayer's actual FAR profile, rather than applying it as a default setting.

This has a direct bearing on how benchmarking is now conducted. Commercial and in-house benchmarking tools typically ship with default screening logic: an RPT filter set at 15% or 25%, a turnover band expressed as a multiple of the tested party's revenue, an export-ratio cutoff for captive units. Agentic benchmarking tools, which some practices have begun piloting, go further: they select the filter to apply, often based on the vendor's training data or built-in heuristics, rather than on a documented judgment call by the practitioner running the analysis.

Read against the Karnataka High Court's ruling, the choice of filter thresholds, not merely the choice of method, is now a part of the TP file that could attract closer scrutiny. A TPO or a taxpayer's advisor who cannot explain why a tool applied a 25% RPT filter instead of 15%, beyond the fact that this is the platform's default, may find that position harder to sustain in a perversity challenge than it would have been before SAP Labs.

This raises a practical question for practitioners advising on AI adoption in benchmarking. Building an auditable, tool-agnostic justification for each filter choice adds documentation overhead, but it may also be the record that a TPO or Tribunal now expects to see. Conversely, the push toward faster, high-volume automated benchmarking runs could encourage greater reliance on unexamined defaults at the same time as the case law raises the cost of doing so. Industry commentary describing 2026 as the year in which touchless compliance becomes a practical necessity has largely framed that shift around GloBE Information Return and Form 6765 deadlines rather than TP substance, which suggests that current AI-in-tax roadmaps are built primarily for data processing rather than for the judgment-intensive task of justifying filter choices that these rulings have placed back on the table.

For practitioners, the immediate implication is procedural rather than strategic: any benchmarking filter applied by a tool, whether commercial or agentic, should be accompanied by a documented rationale linking the threshold to the taxpayer's FAR profile, so that the choice can be defended as a reasoned exercise rather than a software default if it is questioned.

Wednesday, September 23, 2026

Agentic AI in Transfer Pricing Benchmarking: What a New Study's Error Rates Mean for Indian Practice

A study published in The International Tax Journal tests, with data rather than assertion, a claim commonly made for AI-driven transfer pricing benchmarking tools: that autonomous systems can produce faster analysis without a corresponding loss of rigour. The findings deserve closer attention from Indian practitioners than they are likely to receive.

The authors ran agentic AI, systems that execute a multi-step workflow such as a benchmarking search or functional analysis without step-by-step human prompting, against 20 simulated MNE cases across three sectors. They compared the outputs against a model of how tax administrations in 11 countries are deploying similar technology. The efficiency gain was substantial: agentic AI cut processing time by roughly 98% while increasing the depth of the functional analyses produced.

The same study reports three failure rates that merit attention before any TP head signs off on an AI-assisted benchmarking pipeline: a 15% error rate in complex functional characterisations, a 22% irrelevance rate in comparable selection, and a 15% hallucination rate in legal citations. The authors propose a governance framework they call "Tracer-Wire," under which every AI-generated conclusion must carry a visible, auditable path back to its source data, with a mandatory human checkpoint before any output is finalised.

Explainability-by-design and human-in-the-loop review are now standard features of AI governance proposals, so the framework itself is not the notable part of the paper. What is notable is the coincidence between the study's error rates and recent Indian tribunal outcomes.

This month alone, three decisions have turned on the same issue the study measures. The Delhi bench of the ITAT excluded two comparables from Dixon Technologies' set for functional dissimilarity, even though the taxpayer had itself flagged the issue years earlier. The Chennai bench devoted an entire order to whether a single internal comparable can still claim the statutory tolerance band. The Karnataka High Court's SAP Labs line of rulings has produced a further set of follow-on decisions this month on whether a TPO may discard a taxpayer's comparables in favour of a "standard set."

Each of these disputes falls within the 22% irrelevance-rate failure mode the study measured. If agentic AI misjudges comparable relevance roughly one time in five even in a controlled simulation, and Indian tribunals are already spending full orders correcting comparable-selection errors made by humans, the practical question for TP practice in 2026 is concrete rather than conceptual: whose signature appears on the local file when an AI-selected comparable turns out, on review by a TPO or an ITAT bench two years later, to be a functionally dissimilar entity such as a plastics manufacturer.

Coverage of AI in transfer pricing tends to avoid a question that is uncomfortable for vendors and practitioners alike: whether an audit trail reduces liability or merely relocates it. A Tracer-Wire log showing that an AI system considered and rejected a comparable for a documented reason does not make that rejection correct. It makes the error more visible, and it arguably shifts accountability toward whoever approved the workflow rather than toward the tool itself.

It is worth asking whether India's Master File and Form 3CEB documentation requirements are structured to capture this kind of AI-decision provenance at all. If they are not, practitioners may be looking at a new category of documentation gap in an area the OECD's Chapter V framework was never designed to address. Firms adopting agentic AI for benchmarking would do well to build a sign-off protocol now, one that fixes responsibility for each accepted comparable before a tribunal does it for them.

Tuesday, September 22, 2026

DEMPE's Human-Centric Assumptions Face a Test as AI Performs Development and Enhancement Functions

The DEMPE framework determines which group entity is entitled to the return on an intangible by examining which entity performed the Development, Enhancement, Maintenance, Protection and Exploitation functions, made the key decisions, bore the risk, and had the people and capability to control what was happening. The analytical apparatus, functional interviews, organisation charts, decision logs, and the contemporaneous email trail showing an engineer approving a design change, rests on an assumption that a human being performed the function. A recent piece of Australian tax commentary notes that this assumption is becoming harder to sustain. The Australian Taxation Office's guidance on intangibles migration, PCG 2024/1, sets out the evidence it expects a multinational to produce when documenting its DEMPE functions. That evidence regime, the commentary observes, presumes human decision-making and physical performance. It does not address what happens when the decision is made by a model.

PCG 2024/1 has been in force since January 2024. The law itself has not changed; what has changed is a growing recognition that the gap between the rule and the operating reality is becoming practically significant. AI systems are themselves valuable intangible assets under ordinary transfer pricing principles: the algorithms, the training data, and the fine-tuned models sitting inside a captive centre's codebase. A separate issue arises one level up. When an AI system performs the actual development and enhancement work, writing and testing code, running design iterations, flagging defects, optimising a process, the question is who counts as the DEMPE contributor. It could be the engineers who built and supervise the model, the entity hosting the compute, the parent that trained the underlying model elsewhere, or no clearly identifiable party if performance becomes diffuse. The commentary frames this as a mismatch between a human-centric international tax architecture and an operating reality that is automating functions the architecture was built to track.

The implications are likely to be more pressing for India than for Australia. India's transfer pricing landscape relies heavily on captive centres and GCCs performing the kind of granular, iterative technical work, software development, testing, engineering support, and increasingly research, that is most amenable to AI augmentation and, potentially, AI substitution. For two decades, the Indian TP dispute over these centres has centred on characterisation: whether a centre is a routine cost-plus service provider or a genuine value-creating R&D contributor, and which comparables support either position. That analysis has assumed the underlying question, who did the work, could be answered by examining headcount, job descriptions, and reporting lines. If an increasing share of the development and enhancement work inside these centres is performed by AI tools that the centre operates but did not necessarily build, the functional analysis becomes harder to write convincingly in either direction. A TPO could argue that the Indian entity performs less genuine DEMPE work than its cost base suggests, because the model does the underlying processing. A taxpayer could equally argue that the Indian entity deserves more than a routine cost-plus return, because supervising, curating, and directing an AI system that performs high-value technical work is itself a sophisticated function that current benchmarking studies are not equipped to price.

The functional analysis may need a distinct category, one that asks not only who performed the function but who exercised meaningful control and judgment over the system performing it. Whether existing documentation templates, functional interviews, organisation charts, decision logs, can capture that distinction without substantial revision is unclear.

A related institutional question follows. If DEMPE evidentiary standards assume human decision-makers, and the CBDT's own scrutiny and risk-selection processes are moving toward AI-driven flagging at the same time, the tax administration sits on both sides of the same conceptual gap: applying a human-centric framework through tools that are themselves not human-centric. For practitioners, the practical implication is to begin documenting, now, the extent of human supervision, curation, and judgment exercised over AI tools used in development and enhancement work inside Indian captive centres, since existing DEMPE templates may not capture that distinction without adaptation.

Monday, September 21, 2026

Aggregate or Segregate: Why the Choice Cannot Be a Default Setting in AI Benchmarking Tools

Transfer pricing benchmarking requires an early decision on whether to test a transaction on a standalone basis or aggregate it with other international transactions of the entity. Rule 10A's language on aggregation can support either approach depending on the facts, and taxpayers and tax authorities routinely take opposing positions on which approach the facts justify. A recent ITAT Mumbai ruling in the case of NTT India (formerly Dimension Data India) addressed this question directly, and its reasoning has implications for how AI-assisted benchmarking tools are being designed and used by Indian tax teams.

NTT India had benchmarked its management-fee payment to its Asian regional AE as part of an aggregate, entity-level TNMM analysis, arguing that the overall margin was at arm's length once all international transactions were considered together. The TPO disagreed, extracted the management-fee transaction from that aggregate analysis, applied the CUP method, found no comparable uncontrolled data, and valued the entire service at nil, resulting in an adjustment of nearly ₹93.23 crore. The ITAT deleted the adjustment. In November 2025, a separate ITAT Mumbai bench reached a similar conclusion in an unrelated case, holding that a TPO cannot accept TNMM for a taxpayer's transactions in aggregate and then isolate a single line item for independent nil valuation. The two rulings, from different benches and about ten months apart, apply the same underlying principle.

Rule 10A's aggregation language has not changed. What appears to have shifted is how often tribunals are being asked to police where the aggregation boundary sits, and the consistency with which they are ruling against the department on this point. For a TP practitioner, this strengthens the argument that once an aggregate TNMM position has been accepted, or at least not affirmatively rejected, for an assessee's transactions as a whole, the TPO's room to isolate and independently value a single line item is narrower than it may once have appeared.

This reasoning is also relevant to AI-assisted benchmarking and documentation tools, including agentic platforms and GenAI-based comparable-search products, that are being marketed to Indian tax teams and Big Four practices. Most of these tools make an implicit aggregation-or-segregation choice somewhere in their workflow. Some default to pulling entity-level financials and running margin comparisons across the whole profit and loss account. Others are built to isolate and test each intercompany transaction separately because that is easier to automate and audit.

Neither default is safe on its own. A tool that always aggregates risks reproducing the outcome favourable to the taxpayer in NTT India even in fact patterns where aggregation is not actually justified, inviting a TPO challenge on the opposite theory. A tool that always segregates transactions for cleaner, auditable output risks reproducing the same TPO error that was overturned twice within about a year, testing a management fee, a cost-contribution arrangement, or an IT service fee in isolation when it was never meant to be tested that way. The tribunals' reasoning indicates that this decision has to rest on how closely the transactions are linked on the specific facts, not on a default setting built into a product.

This has a direct implication for how TP teams evaluate any AI benchmarking tool they consider buying or building. Vendors are likely to emphasise comparable-search speed and documentation drafting, but the more relevant question for audit defensibility is narrower: does the tool make its aggregation-or-segregation choice explicit, does it require a person to record the specific factual basis for that choice, and would that basis survive a TPO challenge along the lines the TPO raised in NTT India.

As more Indian captives and GCCs adopt AI-assisted TP documentation workflows, practitioners should treat the aggregation-or-segregation call as a documented, fact-specific judgment that sits with a person on the team, not a default a vendor sets. The NTT India line of rulings gives that judgment more weight than it may have carried before, and it is a reasonable basis for reviewing any AI-generated benchmarking file before it is relied upon in a submission to the TPO or the DRP.

Sunday, September 20, 2026

SAP Labs Ruling on Standard Sets and Its Implications for AI-Assisted Benchmarking

Transfer pricing practitioners frequently encounter this pattern: a taxpayer builds a comparable set, the TPO rejects most of it, and the replacement set resembles the set the TPO has used in assessments of similar taxpayers. The Karnataka High Court's ruling in the SAP Labs India appeals addressed this practice directly. The Court held that a TPO cannot reject a taxpayer's comparables merely to substitute a standard set of comparables routinely used by the Income Tax Department, and that the comparability exercise must be tied to the particular international transaction and conform strictly to Rule 10B. The Court also closed off another common approach, holding that the plus or minus 5% tolerance range under Section 92C is a threshold for determining when no adjustment is required, not an automatic deduction from the arithmetic mean once a transaction falls outside it.

Considered solely as a case about TPO conduct, this ruling restates a familiar principle: tribunals have long required that FAR analysis be conducted properly. Its significance is heightened by timing. A small but growing set of commercial platforms, including Tessera, ArmsLength AI, TPGenie and others, now offer a workflow similar to the one the Court found impermissible, without the government's involvement. These vendors market consistency: a fixed, repeatable accept or reject logic applied to a comparable-set export, uniformly across each row, with human review limited to exceptions. Vendors advertise measurable gains, including claims of freeing up preparation time substantially and achieving high accuracy rates on automated accept or reject decisions. The feature underlying this pitch, a single decision logic applied consistently across every candidate company, closely resembles the feature the Karnataka High Court held a TPO cannot rely on when substituting a standard set for taxpayer-specific analysis.

This ruling does not hold that AI-driven benchmarking is impermissible. Nothing in it addresses AI, and no Indian tribunal has yet evaluated an AI-generated comparable set on its own terms. Its relevance lies in the doctrinal language now available for testing a benchmarking study whose selection logic was a repeatable template rather than transaction-specific FAR judgment. A taxpayer whose TP study relied substantially on an automated accept or reject pass, and whose audit trail records only which template rule fired for which company, may find that this traceability does not, on its own, demonstrate that the comparable search was tailored to the taxpayer's controlled transaction. The same risk applies to the Department: if a TPO's office uses AI-assisted searches that effectively reconstitute the Department's familiar standard set under a different label, taxpayer's counsel could cite SAP Labs against that approach.

Practitioners advising on AI-assisted benchmarking need to consider what a defensible workflow looks like under this standard. One possibility is that a traceable, overridable accept or reject log is sufficient, provided a human reviewer documents transaction-specific reasoning for the final set. Another possibility is that the underlying decision logic itself must be shown to respond to the specific FAR profile of the tested party, rather than being applied uniformly across an unrelated population of candidate comparables. No case has tested this question yet, and vendors are unlikely to raise it themselves. Firms advising clients on adopting AI benchmarking tools, or defending a study built using one, would be well advised to build in an explicit, documented step where a reviewer records why the FAR profile of the tested party justified each inclusion or exclusion, independent of what the algorithm flagged, and to treat AI output as a first-pass screen rather than as the analysis itself.

Although the SAP Labs ruling does not mention AI, its reasoning on standard sets and transaction-specific comparability is likely to inform how AI-assisted benchmarking studies are tested going forward. Practitioners should document the human judgment behind each comparable decision accordingly, rather than relying on the traceability of an automated log as a substitute for that judgment.

Saturday, September 19, 2026

Agentic AI in Transfer Pricing: The Practical Problem Is the Handoff Between Agents

Discussion of AI and transfer pricing has largely centred on whether a single autonomous agent could take a set of intercompany agreements, run a functional analysis, select comparables and produce a defensible benchmarking range with minimal human involvement. Vendors market toward that capability, and practitioner panels debate whether such an agent could meet the reliability standards implicit in Section 92C or Section 482. A hackathon held in Vienna earlier this year, organised with the WU Tax Law Technology Center, Microsoft and TPA Global and reported only this week, points to a different pattern taking shape in practice. Rather than building one model to perform the entire task, participating teams chained together several narrower agents, each handling a bounded function.

The case studies covered intra-group financing and intercompany services: arm's length interest rates, creditworthiness assessment, the benefit test, cost allocation, method selection and documentation. Teams built separate agents for data extraction, service classification, benefit testing, cost allocation, compliance monitoring, documentation and audit readiness, and linked them into a workflow. The organisers were explicit that the intent is not to replace professional judgment: outputs are meant to remain traceable to source, reviewed before use, and subject to human oversight at each step. This combination of decomposition and human-in-the-loop review appears to be the practical model emerging from the exercise, even as vendor marketing continues to emphasise single-agent capability.

This distinction matters for Indian TP practice because the architecture of contemporaneous documentation, under Rule 10D, the erstwhile Form 3CEB and now Form 48 under the 2025 Act, assumes a single preparer's judgment trail. A TPO can ask why a particular comparable was included and expect an answer from one analyst or one firm. A pipeline of five narrow agents does not fit that assumption. If a data-extraction agent misclassifies a transaction, a downstream benefit-test agent may inherit that error and proceed regardless, since it is not designed to question upstream inputs, only to execute its own task. The more likely failure mode in a multi-agent TP workflow is not a single agent producing a wrong answer, but an error propagating silently across a handoff that no one is specifically assigned to audit. India's TP documentation requirements, safe harbour disclosures and APA application forms do not currently address this scenario; they assume one preparer whose competence and good faith can be tested under cross-examination or TPO scrutiny.

Practitioners, and possibly CBDT, will need to consider what a chain-of-custody requirement for multi-agent TP work product should look like before a dispute forces the issue. Knowing that a benchmarking output traces back to a database source, as most current vendor claims are framed, is not sufficient. It would also require knowing which agent touched the data at each stage, what it changed or flagged, and whether a human actually reviewed the boundary between two agents' work rather than only the final output. India is already working through related questions, such as how DEMPE functions performed by AI systems fit within the intangibles framework, and whether GCC functional segmentation holds up against agentic AI restructuring inside captive centres. The handoff-audit issue sits a level below those debates: it concerns not whether AI can perform a TP function defensibly, but whether anyone can reconstruct, after the fact, which of several AI agents was responsible when something went wrong. Given how current documentation standards are framed, that question is more likely to surface first in a TPO's show-cause notice than in a policy paper, and practitioners relying on multi-agent tools would do well to build their own audit trail across agent handoffs before that happens.

Friday, September 18, 2026

India's New GCC Benchmarking Advice Meets an Agentic AI Problem It Has Not Addressed

For several years, the standard defensive approach for a captive Global Capability Centre (GCC) facing an Indian Transfer Pricing Officer (TPO) has been fairly mechanical: apply TNMM on an operating cost base, benchmark against routine service-provider comparables, and settle within the accepted cost-plus range. A jurisdiction briefing circulated this week for International Tax Review describes FY2025-26 as a year in which TPO scrutiny of GCC margins and intra-group services intensified, with officers moving toward fewer but deeper adjustments. The advisory response is to segment the GCC into three separately tested activities, namely support, delivery, and decision-rights, with each priced on its own terms rather than blended into a single entity-level margin.

This approach aligns with the OECD's current work at the multilateral level. The OECD has published the full set of public comments on its proposed rewrite of Chapter VII, which governs intra-group services, ahead of a consultation meeting scheduled for November in Paris. Practitioner submissions describe the draft as a substantive rewrite rather than a tidy-up: it requires accurate delineation of what was actually done, by whom, and under what conduct, as the necessary first step before any pricing method is chosen, and it expands the benefit test that has long been the fault line in service fee disputes. Read together, the Indian advisory and the OECD draft point in the same direction: MNEs are being asked to stop pricing services as a single blended category and instead demonstrate, activity by activity, that a real economic function occurred and that an independent party would have paid for it.

This advice may be harder to execute than it appears, not because of documentation gaps in the traditional sense but because of a shift underway in how GCCs operate. India's GCCs are moving away from task-by-task human execution toward agentic AI systems that plan and execute multi-step workflows autonomously. Industry commentary has described Google's move from single-task AI assistants to an agent platform that can be deployed, supervised, and audited across a company's systems as a development that could affect the labour-arbitrage economics on which the GCC sector was built. Separately, EY's GCC survey work finds a large majority of centres already testing agentic technology, with over half piloting agent-based systems specifically. Inside a GCC, this means a workflow once visibly split between a junior analyst performing support work and a senior lead exercising decision-rights judgment can now be executed end-to-end by a single agent, or by a human-agent pair, in a manner not observable from outside the system the way an organisation chart or job description made it observable five years ago.

The Indian TP advisory recommends segmenting the GCC into three tested activities before benchmarking. The OECD requires accurate delineation of the transaction before pricing it. Both instructions assume a documentable boundary between routine support and higher-value decision-making that a TPO or comparability analyst can observe and test. Agentic AI does not necessarily respect that boundary. If an exception-handling agent inside a GCC performs functions that combine what used to be tier-1 support and tier-2 judgment, the functional analysis section of the Local File, which describes who does what, would either need to become considerably more granular about which agent or human made a given call, or risk reverting to the blended, entity-level treatment that both the OECD and Indian practice are trying to move away from.

The practical question is not whether AI adoption in GCCs is occurring; the survey data and industry commentary indicate that it is. It is whether the functional-segmentation and accurate-delineation frameworks currently being refined by the OECD and recommended by Indian advisors were designed for a labour model that may be changing faster than the guidance can be finalised. If accurate delineation depends on establishing what was actually done and by whom, and 'whom' increasingly means a shifting combination of humans and autonomous agents operating across functional boundaries that were previously organisationally distinct, this raises a documentation question that neither the OECD's November consultation agenda nor the Indian safe harbour and Local File templates currently address in detail. For practitioners, the practical implication is to start building functional analyses capable of tracking agent-level activity now, rather than waiting for the guidance to catch up.

Thursday, September 17, 2026

AI-Performed DEMPE Functions and the Limits of Transfer Pricing's Intangibles Framework

The OECD's DEMPE framework was designed to stop groups from parking valuable intangibles in a low-tax entity that holds legal title but performs none of the underlying work. The test traces the relevant functions, development, enhancement, maintenance, protection and exploitation, back to the people who exercise judgment over them: who decided to pursue a research direction, who approved a patent filing, who manages commercialisation risk. For two decades this test has worked reasonably well because those functions were, in fact, performed by identifiable people.

That assumption is being tested where AI systems perform DEMPE functions. Commentary in Australia has raised this issue in relation to the ATO's Practical Compliance Guideline 2024/1 on intangibles migration, which requires multinationals to self-assess and disclose where DEMPE activities for offshore-held intangibles actually occur, using a risk-zone framework that determines how closely the ATO will scrutinise a taxpayer's position. The commentary notes that the PCG's evidence requirements assume human decision-makers, leaving a gap where AI systems (training models, tuning algorithms, monitoring outputs, iterating on protection measures) perform these functions without a person who can be identified as the locus of judgment.

A companion piece extends this analysis to M&A due diligence. It argues that AI systems, including the algorithms, training data and infrastructure, are themselves valuable intangibles under transfer pricing principles, and that a target lacking DEMPE documentation for its AI footprint may carry dormant tax risk that an acquirer inherits on completion.

This is not confined to Australia, and it is not a distant issue for India. India's global capability centre sector has moved well beyond routine coding and support work. Commentary directed at India's GCC market has told clients that where a captive centre designs a core AI algorithm used across the global group, arm's-length pricing cannot be justified by comparing developer hourly rates; the pricing must reflect that creative, value-creating contribution. That is a DEMPE argument, framed in language close to the ATO's, applied to the kind of AI-first GCC that has become common in Bengaluru, Hyderabad and Pune over the last two years.

The timing adds to the difficulty. The Finance Act 2026 safe-harbour rationalisation consolidated software, ITeS, KPO and contract R&D into a single IT-services category, set a flat 15.5% margin, and raised the eligible transaction threshold to ₹2,000 crore. The change was presented as compliance simplification for routine captive units. A captive unit whose engineers are training or fine-tuning a model whose commercial exploitation happens offshore may not be 'routine' in the DEMPE sense, even if it fits comfortably within the new safe-harbour band. A group that elects into the safe harbour for five years on the strength of a routine cost-plus characterisation may be locking in a margin that a DEMPE-consistent functional analysis would not support. That exposure may surface only when the safe-harbour election lapses and a TPO examines what the unit was actually doing.

This raises a harder question about the DEMPE test itself. DEMPE and the 'significant people functions' concept exist to identify who exercises judgment over value creation. Where that judgment is distributed across a model's training run, a human supervisor who approves outputs in bulk, and an infrastructure team in a different jurisdiction, none of these resembles the risk-bearing decision-maker the framework was written for. Tribunals and tax authorities have decades of practice tracing DEMPE to people, but little developed practice for tracing it to a system.

Whether CBDT, the ITAT or the OECD develops a workable answer first, and whether India follows the ATO's disclosure-and-self-assessment approach or adopts something more prescriptive, remains unclear. For now, practitioners advising GCCs and AI-heavy captives should treat DEMPE documentation for AI-performed functions as a present compliance issue, and should examine whether a safe-harbour election understates the functional profile of a unit that is doing more than routine service delivery.

The Facebook Royalty Ruling and the Limits of AI in Transfer Pricing

AI vendors frequently market tools that compress transfer pricing benchmarking into minutes. For certain tasks the claim holds: screening a ...