Sunday, October 4, 2026

AI-Based Comparable Screening Does Not Address India's Core Transfer Pricing Disputes

At the WU-TA Advanced Transfer Pricing Programme in Singapore this week, PwC's transfer pricing partners presented AI-based comparable screening as a significant efficiency gain for transfer pricing teams, and also set out its limits. One partner compared AI to a GPS: given the wrong address, it will still get you there efficiently, but a human has to recognise that the destination itself is wrong. A tax head from P&G compared AI's output to that of a junior associate: useful for a first draft, but requiring supervision, risk assessment and quality control before anyone relies on it.

Comparable screening once meant examining databases or company websites one company at a time. AI tools can now gather that information simultaneously, cutting the search stage substantially. Participants at the conference did not dispute that this is a genuine efficiency gain. The question from an Indian perspective is what that gain addresses within the broader TP dispute process.

Indian transfer pricing litigation in 2026 continues to turn on characterization issues rather than comparable selection. The Delhi ITAT's ruling this week in Discovery Communication India illustrates this. The taxpayer had adopted TNMM, its margin came in above the comparables' average, and the TPO did not challenge the comparable set.

The dispute concerned whether the TPO could carve out AMP expenditure already included within the accepted operating cost base, treat it as a separate international transaction, and re-price it using the Bright Line Test. The Tribunal rejected this approach, following the line of Delhi High Court and ITAT precedent in Sony India, Casio, and the Special Bench ruling in LG Electronics. The reasoning turned on the terms of the intercompany agreement, whether a genuine cost-sharing arrangement existed, and whether the expenditure benefited the taxpayer or the foreign AE. Comparable benchmarking played no part in this analysis.

Much of Indian TP litigation that consumes years and significant tax amounts follows this pattern: captive power CUP arguments, GCC or FAR mischaracterization, receivables treated as separate transactions, and tested-party selection disputes. These are disputes about judgment on facts, not about identifying additional comparable companies.

AI-based comparable screening does not reach this category of dispute. If the contested, revenue-significant disputes in India are predominantly characterization disputes, concerning whether a transaction exists at all, whether an entity is genuinely a captive, or whether expenditure qualifies as AMP, a tool that shortens the comparable search by some weeks does not touch the stage of the process where litigation risk is concentrated.

Such tools make the uncontested, mechanical part of TP documentation cheaper and faster, which has value. That is a different claim from suggesting AI will reduce TP disputes or TP risk, a claim made in vendor marketing and at conference panels, including at the Singapore event. The P&G comparison of AI to a junior associate acknowledges this limitation.

The GPS comparison makes the same point more directly. It concedes that AI's efficiency is irrelevant, and potentially unhelpful, if the underlying functional characterization (the "address") is wrong, and that getting the characterization right remains a human judgment task.

The implication for Indian tax administration concerns where AI-based risk scoring may next be deployed. Tax administrations have tended to adopt AI tools some time after industry does, and CBDT or TPOs could deploy similar AI benchmarking tools to accelerate comparable searches and sharpen risk-based case selection.

If AI's comparative advantage stops where functional characterization begins, as the Singapore discussion suggested, AI-assisted TPOs may build adjustment cases faster on the same mechanical grounds that are already failing in the Tribunals, without improving their ability to predict which cases will survive judicial scrutiny on characterization. Whether any firm or tax administration is developing an AI model trained specifically on characterization outcomes, such as AMP treatment, captive recharacterization, or DEMPE allocation, rather than on comparable company financials, was not evident from the Singapore discussion.

For TP practitioners, the practical implication is to assess AI tools by the stage of the process they improve. Faster comparable searches do not reduce exposure on characterization issues, which continue to require substantive factual and legal analysis, and documentation strategy should treat that distinction as material.

Saturday, October 3, 2026

Flyjac Logistics: Bombay High Court Holds That a System-Generated Notice Is Not Proof of Service in Transfer Pricing Proceedings

Transfer pricing law requires a Transfer Pricing Officer (TPO) to issue a show-cause notice before determining an arm's length price adjustment, giving the taxpayer a genuine opportunity to respond. Courts have generally dealt with breaches of this requirement where a notice was lost in transit, never issued, or deliberately bypassed. The Bombay High Court's recent ruling in Flyjac Logistics' case addresses a different failure mode: the notice was generated and logged by the department's own system, yet never reached the taxpayer because the system malfunctioned.

The TPO proposed a ₹20.16 crore adjustment to Flyjac's arm's length price and recorded that a show-cause notice had been issued and had gone unanswered. Flyjac stated that it never received any such notice and learned of its existence only because the TPO's order referred to it. In its affidavit, the Department attributed this to a technical failure rather than any deliberate lapse: the email carrying the notice had bounced, an unexplained glitch in the ITBA backend kept the notice from appearing on Flyjac's e-filing portal, and no SMS alert was triggered either.

The Bombay High Court held that the statutory show-cause requirement under Section 92C(3) is distinct from the information-seeking notices issued earlier under Section 92CA(2), and that the order could not stand without actual service on the taxpayer. A system log recording that a notice had been issued did not, in the court's view, satisfy that requirement.

The ruling can be read alongside other developments this year that touch on transparency and procedural rigour in tax administration, though these do not point to a single established trend. Chile's SII has published, in some detail, the risk criteria and audit yield behind its enforcement program. HMRC has tightened the link between its TP compliance guidance and its formal error-correction facility, creating a more structured route from self-identified error to resolution. India's Supreme Court, in the Samsung India matter, dismissed the Revenue's appeals because of unexplained delays of several hundred days in filing, without reaching the underlying transfer pricing questions. Taken together, these developments suggest that as tax administrations rely more on systems, logs and published criteria to manage scale, courts and taxpayers are likely to scrutinise those systems more closely rather than less.

CBDT has publicly stated its intention to expand automated risk-flagging through initiatives such as Project Insight and NUDGE, and the broader direction of travel is toward AI-assisted scrutiny. Tax-technology vendors, including infer360 (now allied with Nexdigm) and comparable platforms, market tools for machine-assisted TP monitoring and documentation. Flyjac illustrates a specific risk that arises as these systems take on a larger role: a system can generate a notice, log that it generated the notice, and still fail to deliver it, and such a failure may not surface until the taxpayer is already contesting the resulting order in court. As automated systems take on more of the notice-generation and routing function in TP enforcement, similar service failures are likely to recur as adoption increases.

For practitioners, the ruling raises a practical evidentiary question: whether a system log showing that a notice was issued will be treated as sufficient proof of service, or whether courts will continue to require affirmative proof of actual receipt, as the Bombay High Court did in this case. Until that question is settled through further litigation, TP compliance functions may find it prudent to maintain independent records of portal access, email delivery and SMS alerts, so that service failures originating in the department's own systems can be identified and raised at the earliest stage of any proceeding.

Friday, October 2, 2026

ITAT Mumbai's Captive Power Ruling and the Limits of AI-Driven TP Benchmarking

A recent Mumbai ITAT ruling in a captive power transfer pricing dispute turned on a question that AI benchmarking tools are not built to ask: whether the comparable a taxpayer needed was already sitting inside its own books.

The case involved a cement manufacturer whose captive power plants supplied electricity to its own cement units. The dispute arose under the 2013 amendment that brought specified domestic transactions within the arm's length pricing framework. The Revenue argued that the captive generator had to be benchmarked against the rate at which other generators sell to state distribution licensees. The taxpayer argued that its own cement units paid documented market rates to state power distribution companies for electricity purchased on the open market, and that this internal rate was a valid comparable for what the captive plant charged them. The tribunal agreed. It held that the amendment does not require a generator to be benchmarked only against another generator, or that a distribution licensee's procurement rate must invariably be adopted, and that CUP selection remains governed by the transaction-specific requirements of section 92C and Rule 10B. Revenue's appeals against adjustments of roughly Rs. 64 crore and Rs. 43 crore across two years were dismissed.

The outcome is less significant than the source of the winning comparable. It was not retrieved from a database of unrelated third-party companies, the kind of external, industry-classified, margin-screened universe that commercial TP benchmarking tools are built to search. It was already in the taxpayer's own transaction records: the price its own units paid the grid for the same commodity, in the same period, in the same geography. An external comparables search, however well it ranks functional similarity, would not have surfaced this fact pattern, because the relevant exercise was not searching further afield but looking at what the group itself was already paying for the identical input elsewhere in its operations.

This points to a gap in how AI benchmarking tools are currently designed. Internal CUPs sit at the top of the comparability hierarchy because they avoid much of the functional-comparability subjectivity that external searches have to approximate statistically. Yet most AI tooling marketed to TP teams is built to screen external databases faster, not to mine a group's own intercompany and third-party transaction data for an internal benchmark. These tools automate work that used to be slow, external database screening, but they are not generally designed to surface the kind of internal comparable a case like this one turned on.

A separate observation from this filing season's TP engagements reinforces the point. Some practitioners report that adjustments are more often lost on evidence discipline than on legal interpretation: the comparable search cannot be reproduced, or the Local File and the counterparty's file tell different functional stories. Reproducibility of an external search is a baseline requirement. It says nothing about whether a better internal comparable was considered at all.

For TP practitioners evaluating AI benchmarking tools, this ruling is a useful prompt to ask a specific question: does the tool query the group's own ERP and intercompany transaction history for internal comparables before it touches an external database, or does it only speed up external screening under a new label. Before commissioning an external search, it is worth checking whether the answer was already in the ledger.

Thursday, October 1, 2026

The Asymmetry in AI Use Between Tax Authorities and Taxpayers in Transfer Pricing Defense

A recent Tax Notes article on a Turkish transfer pricing case traces the dispute through to its downstream effect on a U.S. foreign tax credit claim. In doing so, it sets out an asymmetry that deserves more attention in transfer pricing practice: Turkey regulates how professional advisors may use generative AI in giving tax advice, but places no comparable constraint on its own tax administration's use of machine analysis for risk-scoring, flagging related-party anomalies, or selecting cases for audit. The rules on the taxpayer side are specific and restrictive. The rules on the authority side are largely unwritten.

The same week the Turkey piece appeared, a Manila-based practitioner column described how the Philippines' Bureau of Internal Revenue has operationalised a system-assisted, risk-based audit selection framework under Revenue Memorandum Order 1-2026. The indicators are aimed squarely at related-party transactions: persistent losses against strong revenue, tax-to-sales ratios that look too low, and heavy reliance on a single related counterparty. Under this framework, transfer pricing risk is pre-flagged by the system before an examiner opens the file. Read together with Turkey's dual posture, restrictive on taxpayer-side AI reliance and expansive on authority-side machine analysis, the two examples suggest the asymmetry is not confined to one jurisdiction.

The same question arises for anyone advising Indian multinationals or GCCs. India's safe harbour election process is moving toward an automated, rules-based model, CBDT's compliance apparatus already uses AI-assisted risk profiling, and the APA and audit infrastructure is becoming steadily more analytics-driven. All of this sits on the authority side of the same asymmetry described in the Turkish and Philippine examples. Several Big 4 and boutique TP technology vendors have published material this year describing increased use of generative AI in preparing local files, benchmarking memoranda, and functional analyses. Whether that assistance will be treated on the same footing as a signed opinion from a human expert, when the question becomes whether a position was taken in good faith or whether a reasonable-cause defence survives scrutiny, is not yet settled by any rule or ruling in India.

The Turkish case does not resolve this question. Its value lies in naming the gap explicitly, rather than treating AI in tax administration as a single undifferentiated trend of efficiency gains for everyone. Whether the standards governing reliance, documentation, and reasonable cause will develop in step for both tax authorities and taxpayers, or whether taxpayers will end up defending AI-informed positions against AI-generated risk scores under rules drafted for one-sided use, remains open.

For Indian practitioners, the practical implication is to document how AI tools are used in preparing local files, benchmarking analyses, and functional analyses now, before the question is tested in audit or litigation, so that the basis for any position can be explained and defended independently of the tool used to generate it.

Wednesday, September 30, 2026

India's Automated Safe Harbour Approval and the Administrative-Law Question Canada Is Now Asking About AI-Driven Audit Selection

Budget 2026 moved Safe Harbour approval for IT services onto an automated, rule-driven framework and removed the requirement for an officer to examine the application. Applicants who meet the prescribed margin thresholds receive approval without human review, and the outcome is locked in for up to five years.

A recent Canadian Tax Journal paper by Theertha Narayanan and Pramod Kumar Siva examines a related but distinct question: whether the Canada Revenue Agency's use of AI-based risk scoring to select transfer-pricing files for audit must satisfy the same procedural-fairness standard that Canadian courts apply to other administrative decisions. The paper argues that it should, particularly after Canada's 2025 federal budget introduced stricter TP methodologies, expanded recharacterization powers, higher penalty thresholds and shorter documentation deadlines, which raised the consequences of being selected for audit. On the paper's argument, the selection decision itself should be reviewable under the reasonableness standard set out in the Supreme Court of Canada's Vavilov framework: justified, transparent and intelligible.

The Canadian debate concerns whether an algorithm can lawfully flag a file for human review. India's Budget 2026 reform goes further: it has let an algorithm replace the human reviewer for a determination that carries up to five years of certainty on arm's length pricing. Indian commentary on the change does not appear to have framed it as raising a comparable administrative-law question.

The Canadian paper's argument rests on specific case law rather than general concerns about AI. It draws on Vavilov's requirement that administrative reasoning be internally coherent and "justified in relation to the facts and law that constrain the decision maker," and on Dow Chemical's confirmation that discretionary CRA decisions under the Income Tax Act are reviewable on that same reasonableness standard. The paper's contribution is to apply these doctrines to a system that produces a risk score rather than a chain of reasoning, and to ask whether a score alone can meet a standard that requires justification.

The Indian changes are comparably specific. Budget 2026 pushed Safe Harbour approval into what industry commentary has called an "auto-pilot mode": applications are processed through a fully automated framework without officer discretion, in exchange for locking in outcomes for five consecutive years. The stated rationale, reducing subjective review, reducing litigation and increasing predictability, is consistent with the reasoning most jurisdictions offer for automating tax administration. Automating an approval, however, is a different act from automating a flag for human review, and that difference is what the Canadian paper is built to examine.

The practical stakes for Indian practice are concrete. GCCs are the primary beneficiaries of the new Safe Harbour automation, and the pitch to them has centred on certainty: file the declaration, receive automatic approval, and bypass officer review. An approval granted without examination is, in substance, an algorithmic determination of an arm's length outcome. It is made without safeguards that the international literature on automated decision-making treats as significant, including explainability, an audit trail showing why a given margin was accepted, and a documented basis for treating a taxpayer's declared facts as sufficient without human verification.

India has no published equivalent of the Vavilov standard for testing whether a rule-driven tax decision is reasonable. Nor is there public discussion, so far, of what a taxpayer can do if an automated approval later turns out to rest on a misclassification the system had no way to detect. On most commentary on the reform, officers who did not examine the application at the front end can still reopen the matter on audit later. That leaves taxpayers with a certainty that appears binding on its face but may not bind the department in substance, because no human made a reviewable decision at the outset.

Whether Indian tax administrative law needs an equivalent of the reasoned-decision requirement that Canadian courts are now applying to algorithmic audit selection remains open. If Budget 2026 signals a broader shift toward rule-driven, officer-free approval extending to APA processing or audit selection, the question will carry more weight, though the current material does not establish how far that shift will go. There is also a defensible counter-view, that Safe Harbour was always meant to operate as a self-assessment regime in which the absence of officer discretion is a deliberate feature rather than a due-process gap. Practitioners advising on Safe Harbour elections should flag to clients that automated approval does not necessarily foreclose later audit scrutiny, and should treat the administrative-law framing, not merely the compliance mechanics, as part of the risk assessment.

Tuesday, September 29, 2026

From Annual Documentation to Continuous Monitoring: What the Nexdigm-infer360 Alliance Signals for TP Practice

Transfer pricing documentation in India is built after the fact. Benchmarking studies are prepared once a year, local files are frozen at a point in time, and the arm's length position a taxpayer defends before a TPO reflects a snapshot taken months or years after the transactions occurred. The analysis need not be wrong, but it is backward-looking by design: assembled once the business has already priced the transaction, to justify a decision already taken.

A new generation of AI-native platforms is pitching something different: continuous, always-on monitoring of intercompany pricing through the year, rather than faster annual documentation. Nexdigm, the Mumbai-headquartered advisory firm, has recently announced an alliance with infer360, a Singapore-based AI platform built by former Big Four transfer pricing partners. The stated aim is to help clients move from periodic compliance to continuous transfer pricing management: automating documentation, running ongoing risk monitoring, and maintaining an audit-ready trail through the year rather than reconstructing one afterward.

The marketing language is unremarkable; most vendors in this space now describe themselves as AI-native. What is worth noting is that a serious Indian-origin advisory practice with established APA, controversy and operational TP credentials is putting its own resources behind the view that continuous monitoring, rather than faster annual studies, is where the market is heading.

This sits alongside a parallel move on the US side of the industry, where an established transfer pricing technology vendor's benchmarking and documentation AI suite continues to draw trade-press attention as a template for how mid-market and boutique practices are narrowing the technology gap with the Big Four.

The relevance for practitioners lies in the kind of evidentiary record these platforms are built to produce: a clean, time-stamped account of the methodology used and how it performed against comparables through the year. That is the kind of record Indian tribunals have shown they will act on.

In a Delhi ITAT ruling this month in Honda R&D (India)'s case, the taxpayer faced a ₹50.20-lakh adjustment for AY 2020-21 and pointed the Tribunal to a TPO order for the following assessment year, AY 2021-22, in which the department had accepted the identical benchmarking position it was disputing for the earlier year. The Tribunal directed relief consistent with the later year's accepted treatment.

This is not an isolated result. Indian tribunals have repeatedly leaned on a consistency principle when a TPO's own later-year acceptance undercuts an earlier-year adjustment. A continuous-monitoring platform, by construction, produces the kind of multi-year, methodology-stable record that makes this argument easier to run: the more granular and continuous a taxpayer's own TP data trail becomes, the stronger its consistency-principle arguments are likely to be.

This also raises a question about asymmetry between well-resourced multinationals and the tax administration. If sophisticated taxpayers build continuous, audit-ready monitoring systems while the department still audits largely on an annual, backward-looking cycle anchored to Form 3CEB-style filings, an information and preparedness gap could open up between the two sides of the table.

India's income tax apparatus already runs AI-based risk assessment for scrutiny selection at the return level, and the CBDT's APA programme leans heavily on post-agreement monitoring of critical assumptions for bilateral agreements. From there, it is a reasonable question whether the department should build its own continuous-monitoring capability specifically for TP, to match rather than only react to the tooling that advisory firms are now selling to taxpayers.

There is also a discovery-related question worth flagging. As continuous monitoring platforms generate richer contemporaneous records than the old annual documentation model, those records could prove double-edged: they may help taxpayers win consistency arguments in years like this one, but they could also give revenue authorities a more granular trail to probe when a deviation does appear.

Neither the vendors marketing these platforms nor the tribunals applying the consistency principle have had occasion to address this tension yet. As continuous monitoring tools become more common, practitioners advising on TP documentation strategy will need to weigh the benefit of a stronger contemporaneous record against the risk that the same record gives the department more material to examine when a taxpayer's position shifts from one year to the next.

Monday, September 28, 2026

Karnataka HC's SAP Labs Remand Ruling and the Case for Reasoned Comparability Filters in AI Benchmarking

The Karnataka High Court's 2018 decision in Softbrands treated the selection and rejection of comparables as a pure question of fact, which left the ITAT as the final word on comparability disputes; High Courts could not reweigh turnover filters, RPT thresholds, or FAR analyses under Section 260A. The Supreme Court's ruling in SAP Labs India v. ITO rejected the idea that every ALP determination is immune from scrutiny and sent a batch of pending appeals back to Karnataka for fresh consideration on the merits. How the High Court would apply that reopened scrutiny remained unclear until its recent consolidated ruling.

In its 28 August 2026 consolidated ruling, the Karnataka High Court restated the SAP Labs test and attached specific numerical thresholds to it. A ₹200-crore upper turnover filter was held to be rational and logical, not because the statute prescribes it, but because size differences plausibly correlate with brand value, bargaining power, and economies of scale. A 15% RPT (related-party-transaction) filter was endorsed as the ordinary default, with a departure to 20% or 25% permitted only where the authority records a specific, reasoned finding that comparables meeting the lower threshold are scarce. The Court also held that the ±5% tolerance range under Section 92C is a threshold, not a standing deduction that a taxpayer can claim once the transaction price exceeds it. Within days, at least two more Karnataka HC benches cited this ruling to dispose of pending Revenue appeals, including one that upheld the exclusion of Bodhtree Consulting as a comparable by applying the same 200-crore ceiling.

Barely two weeks after the Karnataka HC ruling, ITAT Delhi reached a different outcome in GE India Industrial v. DCIT, rejecting the Revenue's use of a rigid 50% turnover band to strike out lower-end comparables and criticising the DRP for being "guided by a general opinion" rather than the statutory FAR-based comparability test under Rule 10B. Read together, the two rulings point to a common test: a filter survives scrutiny only if the tribunal or officer applying it shows a specific, reasoned link between the chosen threshold and functional comparability on the facts of the case. A 200-crore ceiling with a size-based rationale passes. A blanket 50% band applied mechanically, without engaging the functional analysis, fails.

These rulings have implications for how AI tools are used in TP documentation. Benchmarking platforms, whether Big 4 proprietary tools or newer AI-native SaaS products, increasingly automate comparable searches using statistical filters such as turnover bands, RPT thresholds, and quartile screens. A vendor could hard-code the thresholds recently upheld, such as 200 crore and 15%, as default parameters on the assumption that a threshold accepted once will be accepted again.

That approach risks the failure mode both rulings identify from opposite directions: a threshold applied without a documented, case-specific rationale is what gets struck down, whether the Revenue or the taxpayer relies on it. An AI benchmarking tool that outputs a comparables set with a turnover filter applied, but no accompanying explanation tying that filter to the tested party's specific brand value, IP ownership, or scale economics, is building a study that looks defensible today and may not survive the next remand.

TP teams, and the vendors building the AI layer underneath them, may need to consider whether benchmarking software should generate a reasoned-finding log alongside every filter it applies: not merely that a comparable was excluded because turnover exceeded 200 crore, but a documented basis for why that threshold was chosen for the specific tested party, FAR profile, and data set. The Karnataka High Court's ruling signals that courts are now willing to test that reasoning on the merits rather than treat comparable selection as an unreviewable factual finding.

For AI-assisted TP work, the practical advantage may lie less in the speed of comparable identification than in the tool's capacity to produce explainable, fact-specific justification that a Division Bench applying the SAP Labs test would accept.

AI-Based Comparable Screening Does Not Address India's Core Transfer Pricing Disputes

At the WU-TA Advanced Transfer Pricing Programme in Singapore this week, PwC's transfer pricing partners presented AI-based comparable s...