Thursday, September 24, 2026

Post-SAP Labs Rulings Tighten Scrutiny of Comparability Filters in TP Benchmarking

The Karnataka High Court has disposed of a batch of transfer pricing appeals that had been pending since the Supreme Court's 2023 judgment in SAP Labs sent them back for fresh consideration. The appeals turned on a familiar but consequential question: how much latitude a High Court has to examine the comparability filters, such as turnover bands, related-party-transaction (RPT) thresholds and export-ratio cutoffs, that determine which companies are shortlisted for TNMM or CPM benchmarking.

Under the Karnataka High Court's earlier Softbrands precedent, comparability disputes, including which companies get excluded and which filters apply, were treated as pure questions of fact on which a High Court would not ordinarily interfere. SAP Labs removed that immunity. In this month's ruling, the Court held that it can examine whether the selection of comparables and the choice of filters was made judiciously and on the basis of relevant material, though it will continue to defer to the Tribunal unless the process is shown to violate Section 92C or Rule 10B, or is otherwise perverse. On the facts, the Court upheld a Rs 200 crore upper turnover filter as rational given the size, brand value and economies of scale of the tested party, and held that a 15% RPT filter is ordinarily appropriate. A higher threshold of 20% or 25% is permissible only where the TPO records a specific finding explaining why comparables meeting the lower threshold were unavailable. The Court also held that the tolerance band of plus or minus 5% under Section 92C is not an automatic standard deduction from the arithmetic mean; it is a threshold that determines when an adjustment is required, not a discount to be applied as a matter of course.

In the same week, the Delhi ITAT reached a comparable conclusion in GE India Industrial's case. The Tribunal rejected the TPO's rigid 50% turnover filter and held that functionally comparable companies cannot be excluded on size grounds alone; comparability must be tested against the functions, assets and risks framework under Rule 10B, not against an arbitrary turnover band. Read together, the two rulings direct TPOs and taxpayers to support each filter with a documented, fact-specific rationale tied to the taxpayer's actual FAR profile, rather than applying it as a default setting.

This has a direct bearing on how benchmarking is now conducted. Commercial and in-house benchmarking tools typically ship with default screening logic: an RPT filter set at 15% or 25%, a turnover band expressed as a multiple of the tested party's revenue, an export-ratio cutoff for captive units. Agentic benchmarking tools, which some practices have begun piloting, go further: they select the filter to apply, often based on the vendor's training data or built-in heuristics, rather than on a documented judgment call by the practitioner running the analysis.

Read against the Karnataka High Court's ruling, the choice of filter thresholds, not merely the choice of method, is now a part of the TP file that could attract closer scrutiny. A TPO or a taxpayer's advisor who cannot explain why a tool applied a 25% RPT filter instead of 15%, beyond the fact that this is the platform's default, may find that position harder to sustain in a perversity challenge than it would have been before SAP Labs.

This raises a practical question for practitioners advising on AI adoption in benchmarking. Building an auditable, tool-agnostic justification for each filter choice adds documentation overhead, but it may also be the record that a TPO or Tribunal now expects to see. Conversely, the push toward faster, high-volume automated benchmarking runs could encourage greater reliance on unexamined defaults at the same time as the case law raises the cost of doing so. Industry commentary describing 2026 as the year in which touchless compliance becomes a practical necessity has largely framed that shift around GloBE Information Return and Form 6765 deadlines rather than TP substance, which suggests that current AI-in-tax roadmaps are built primarily for data processing rather than for the judgment-intensive task of justifying filter choices that these rulings have placed back on the table.

For practitioners, the immediate implication is procedural rather than strategic: any benchmarking filter applied by a tool, whether commercial or agentic, should be accompanied by a documented rationale linking the threshold to the taxpayer's FAR profile, so that the choice can be defended as a reasoned exercise rather than a software default if it is questioned.

Wednesday, September 23, 2026

Agentic AI in Transfer Pricing Benchmarking: What a New Study's Error Rates Mean for Indian Practice

A study published in The International Tax Journal tests, with data rather than assertion, a claim commonly made for AI-driven transfer pricing benchmarking tools: that autonomous systems can produce faster analysis without a corresponding loss of rigour. The findings deserve closer attention from Indian practitioners than they are likely to receive.

The authors ran agentic AI, systems that execute a multi-step workflow such as a benchmarking search or functional analysis without step-by-step human prompting, against 20 simulated MNE cases across three sectors. They compared the outputs against a model of how tax administrations in 11 countries are deploying similar technology. The efficiency gain was substantial: agentic AI cut processing time by roughly 98% while increasing the depth of the functional analyses produced.

The same study reports three failure rates that merit attention before any TP head signs off on an AI-assisted benchmarking pipeline: a 15% error rate in complex functional characterisations, a 22% irrelevance rate in comparable selection, and a 15% hallucination rate in legal citations. The authors propose a governance framework they call "Tracer-Wire," under which every AI-generated conclusion must carry a visible, auditable path back to its source data, with a mandatory human checkpoint before any output is finalised.

Explainability-by-design and human-in-the-loop review are now standard features of AI governance proposals, so the framework itself is not the notable part of the paper. What is notable is the coincidence between the study's error rates and recent Indian tribunal outcomes.

This month alone, three decisions have turned on the same issue the study measures. The Delhi bench of the ITAT excluded two comparables from Dixon Technologies' set for functional dissimilarity, even though the taxpayer had itself flagged the issue years earlier. The Chennai bench devoted an entire order to whether a single internal comparable can still claim the statutory tolerance band. The Karnataka High Court's SAP Labs line of rulings has produced a further set of follow-on decisions this month on whether a TPO may discard a taxpayer's comparables in favour of a "standard set."

Each of these disputes falls within the 22% irrelevance-rate failure mode the study measured. If agentic AI misjudges comparable relevance roughly one time in five even in a controlled simulation, and Indian tribunals are already spending full orders correcting comparable-selection errors made by humans, the practical question for TP practice in 2026 is concrete rather than conceptual: whose signature appears on the local file when an AI-selected comparable turns out, on review by a TPO or an ITAT bench two years later, to be a functionally dissimilar entity such as a plastics manufacturer.

Coverage of AI in transfer pricing tends to avoid a question that is uncomfortable for vendors and practitioners alike: whether an audit trail reduces liability or merely relocates it. A Tracer-Wire log showing that an AI system considered and rejected a comparable for a documented reason does not make that rejection correct. It makes the error more visible, and it arguably shifts accountability toward whoever approved the workflow rather than toward the tool itself.

It is worth asking whether India's Master File and Form 3CEB documentation requirements are structured to capture this kind of AI-decision provenance at all. If they are not, practitioners may be looking at a new category of documentation gap in an area the OECD's Chapter V framework was never designed to address. Firms adopting agentic AI for benchmarking would do well to build a sign-off protocol now, one that fixes responsibility for each accepted comparable before a tribunal does it for them.

Tuesday, September 22, 2026

DEMPE's Human-Centric Assumptions Face a Test as AI Performs Development and Enhancement Functions

The DEMPE framework determines which group entity is entitled to the return on an intangible by examining which entity performed the Development, Enhancement, Maintenance, Protection and Exploitation functions, made the key decisions, bore the risk, and had the people and capability to control what was happening. The analytical apparatus, functional interviews, organisation charts, decision logs, and the contemporaneous email trail showing an engineer approving a design change, rests on an assumption that a human being performed the function. A recent piece of Australian tax commentary notes that this assumption is becoming harder to sustain. The Australian Taxation Office's guidance on intangibles migration, PCG 2024/1, sets out the evidence it expects a multinational to produce when documenting its DEMPE functions. That evidence regime, the commentary observes, presumes human decision-making and physical performance. It does not address what happens when the decision is made by a model.

PCG 2024/1 has been in force since January 2024. The law itself has not changed; what has changed is a growing recognition that the gap between the rule and the operating reality is becoming practically significant. AI systems are themselves valuable intangible assets under ordinary transfer pricing principles: the algorithms, the training data, and the fine-tuned models sitting inside a captive centre's codebase. A separate issue arises one level up. When an AI system performs the actual development and enhancement work, writing and testing code, running design iterations, flagging defects, optimising a process, the question is who counts as the DEMPE contributor. It could be the engineers who built and supervise the model, the entity hosting the compute, the parent that trained the underlying model elsewhere, or no clearly identifiable party if performance becomes diffuse. The commentary frames this as a mismatch between a human-centric international tax architecture and an operating reality that is automating functions the architecture was built to track.

The implications are likely to be more pressing for India than for Australia. India's transfer pricing landscape relies heavily on captive centres and GCCs performing the kind of granular, iterative technical work, software development, testing, engineering support, and increasingly research, that is most amenable to AI augmentation and, potentially, AI substitution. For two decades, the Indian TP dispute over these centres has centred on characterisation: whether a centre is a routine cost-plus service provider or a genuine value-creating R&D contributor, and which comparables support either position. That analysis has assumed the underlying question, who did the work, could be answered by examining headcount, job descriptions, and reporting lines. If an increasing share of the development and enhancement work inside these centres is performed by AI tools that the centre operates but did not necessarily build, the functional analysis becomes harder to write convincingly in either direction. A TPO could argue that the Indian entity performs less genuine DEMPE work than its cost base suggests, because the model does the underlying processing. A taxpayer could equally argue that the Indian entity deserves more than a routine cost-plus return, because supervising, curating, and directing an AI system that performs high-value technical work is itself a sophisticated function that current benchmarking studies are not equipped to price.

The functional analysis may need a distinct category, one that asks not only who performed the function but who exercised meaningful control and judgment over the system performing it. Whether existing documentation templates, functional interviews, organisation charts, decision logs, can capture that distinction without substantial revision is unclear.

A related institutional question follows. If DEMPE evidentiary standards assume human decision-makers, and the CBDT's own scrutiny and risk-selection processes are moving toward AI-driven flagging at the same time, the tax administration sits on both sides of the same conceptual gap: applying a human-centric framework through tools that are themselves not human-centric. For practitioners, the practical implication is to begin documenting, now, the extent of human supervision, curation, and judgment exercised over AI tools used in development and enhancement work inside Indian captive centres, since existing DEMPE templates may not capture that distinction without adaptation.

Monday, September 21, 2026

Aggregate or Segregate: Why the Choice Cannot Be a Default Setting in AI Benchmarking Tools

Transfer pricing benchmarking requires an early decision on whether to test a transaction on a standalone basis or aggregate it with other international transactions of the entity. Rule 10A's language on aggregation can support either approach depending on the facts, and taxpayers and tax authorities routinely take opposing positions on which approach the facts justify. A recent ITAT Mumbai ruling in the case of NTT India (formerly Dimension Data India) addressed this question directly, and its reasoning has implications for how AI-assisted benchmarking tools are being designed and used by Indian tax teams.

NTT India had benchmarked its management-fee payment to its Asian regional AE as part of an aggregate, entity-level TNMM analysis, arguing that the overall margin was at arm's length once all international transactions were considered together. The TPO disagreed, extracted the management-fee transaction from that aggregate analysis, applied the CUP method, found no comparable uncontrolled data, and valued the entire service at nil, resulting in an adjustment of nearly ₹93.23 crore. The ITAT deleted the adjustment. In November 2025, a separate ITAT Mumbai bench reached a similar conclusion in an unrelated case, holding that a TPO cannot accept TNMM for a taxpayer's transactions in aggregate and then isolate a single line item for independent nil valuation. The two rulings, from different benches and about ten months apart, apply the same underlying principle.

Rule 10A's aggregation language has not changed. What appears to have shifted is how often tribunals are being asked to police where the aggregation boundary sits, and the consistency with which they are ruling against the department on this point. For a TP practitioner, this strengthens the argument that once an aggregate TNMM position has been accepted, or at least not affirmatively rejected, for an assessee's transactions as a whole, the TPO's room to isolate and independently value a single line item is narrower than it may once have appeared.

This reasoning is also relevant to AI-assisted benchmarking and documentation tools, including agentic platforms and GenAI-based comparable-search products, that are being marketed to Indian tax teams and Big Four practices. Most of these tools make an implicit aggregation-or-segregation choice somewhere in their workflow. Some default to pulling entity-level financials and running margin comparisons across the whole profit and loss account. Others are built to isolate and test each intercompany transaction separately because that is easier to automate and audit.

Neither default is safe on its own. A tool that always aggregates risks reproducing the outcome favourable to the taxpayer in NTT India even in fact patterns where aggregation is not actually justified, inviting a TPO challenge on the opposite theory. A tool that always segregates transactions for cleaner, auditable output risks reproducing the same TPO error that was overturned twice within about a year, testing a management fee, a cost-contribution arrangement, or an IT service fee in isolation when it was never meant to be tested that way. The tribunals' reasoning indicates that this decision has to rest on how closely the transactions are linked on the specific facts, not on a default setting built into a product.

This has a direct implication for how TP teams evaluate any AI benchmarking tool they consider buying or building. Vendors are likely to emphasise comparable-search speed and documentation drafting, but the more relevant question for audit defensibility is narrower: does the tool make its aggregation-or-segregation choice explicit, does it require a person to record the specific factual basis for that choice, and would that basis survive a TPO challenge along the lines the TPO raised in NTT India.

As more Indian captives and GCCs adopt AI-assisted TP documentation workflows, practitioners should treat the aggregation-or-segregation call as a documented, fact-specific judgment that sits with a person on the team, not a default a vendor sets. The NTT India line of rulings gives that judgment more weight than it may have carried before, and it is a reasonable basis for reviewing any AI-generated benchmarking file before it is relied upon in a submission to the TPO or the DRP.

Sunday, September 20, 2026

SAP Labs Ruling on Standard Sets and Its Implications for AI-Assisted Benchmarking

Transfer pricing practitioners frequently encounter this pattern: a taxpayer builds a comparable set, the TPO rejects most of it, and the replacement set resembles the set the TPO has used in assessments of similar taxpayers. The Karnataka High Court's ruling in the SAP Labs India appeals addressed this practice directly. The Court held that a TPO cannot reject a taxpayer's comparables merely to substitute a standard set of comparables routinely used by the Income Tax Department, and that the comparability exercise must be tied to the particular international transaction and conform strictly to Rule 10B. The Court also closed off another common approach, holding that the plus or minus 5% tolerance range under Section 92C is a threshold for determining when no adjustment is required, not an automatic deduction from the arithmetic mean once a transaction falls outside it.

Considered solely as a case about TPO conduct, this ruling restates a familiar principle: tribunals have long required that FAR analysis be conducted properly. Its significance is heightened by timing. A small but growing set of commercial platforms, including Tessera, ArmsLength AI, TPGenie and others, now offer a workflow similar to the one the Court found impermissible, without the government's involvement. These vendors market consistency: a fixed, repeatable accept or reject logic applied to a comparable-set export, uniformly across each row, with human review limited to exceptions. Vendors advertise measurable gains, including claims of freeing up preparation time substantially and achieving high accuracy rates on automated accept or reject decisions. The feature underlying this pitch, a single decision logic applied consistently across every candidate company, closely resembles the feature the Karnataka High Court held a TPO cannot rely on when substituting a standard set for taxpayer-specific analysis.

This ruling does not hold that AI-driven benchmarking is impermissible. Nothing in it addresses AI, and no Indian tribunal has yet evaluated an AI-generated comparable set on its own terms. Its relevance lies in the doctrinal language now available for testing a benchmarking study whose selection logic was a repeatable template rather than transaction-specific FAR judgment. A taxpayer whose TP study relied substantially on an automated accept or reject pass, and whose audit trail records only which template rule fired for which company, may find that this traceability does not, on its own, demonstrate that the comparable search was tailored to the taxpayer's controlled transaction. The same risk applies to the Department: if a TPO's office uses AI-assisted searches that effectively reconstitute the Department's familiar standard set under a different label, taxpayer's counsel could cite SAP Labs against that approach.

Practitioners advising on AI-assisted benchmarking need to consider what a defensible workflow looks like under this standard. One possibility is that a traceable, overridable accept or reject log is sufficient, provided a human reviewer documents transaction-specific reasoning for the final set. Another possibility is that the underlying decision logic itself must be shown to respond to the specific FAR profile of the tested party, rather than being applied uniformly across an unrelated population of candidate comparables. No case has tested this question yet, and vendors are unlikely to raise it themselves. Firms advising clients on adopting AI benchmarking tools, or defending a study built using one, would be well advised to build in an explicit, documented step where a reviewer records why the FAR profile of the tested party justified each inclusion or exclusion, independent of what the algorithm flagged, and to treat AI output as a first-pass screen rather than as the analysis itself.

Although the SAP Labs ruling does not mention AI, its reasoning on standard sets and transaction-specific comparability is likely to inform how AI-assisted benchmarking studies are tested going forward. Practitioners should document the human judgment behind each comparable decision accordingly, rather than relying on the traceability of an automated log as a substitute for that judgment.

Saturday, September 19, 2026

Agentic AI in Transfer Pricing: The Practical Problem Is the Handoff Between Agents

Discussion of AI and transfer pricing has largely centred on whether a single autonomous agent could take a set of intercompany agreements, run a functional analysis, select comparables and produce a defensible benchmarking range with minimal human involvement. Vendors market toward that capability, and practitioner panels debate whether such an agent could meet the reliability standards implicit in Section 92C or Section 482. A hackathon held in Vienna earlier this year, organised with the WU Tax Law Technology Center, Microsoft and TPA Global and reported only this week, points to a different pattern taking shape in practice. Rather than building one model to perform the entire task, participating teams chained together several narrower agents, each handling a bounded function.

The case studies covered intra-group financing and intercompany services: arm's length interest rates, creditworthiness assessment, the benefit test, cost allocation, method selection and documentation. Teams built separate agents for data extraction, service classification, benefit testing, cost allocation, compliance monitoring, documentation and audit readiness, and linked them into a workflow. The organisers were explicit that the intent is not to replace professional judgment: outputs are meant to remain traceable to source, reviewed before use, and subject to human oversight at each step. This combination of decomposition and human-in-the-loop review appears to be the practical model emerging from the exercise, even as vendor marketing continues to emphasise single-agent capability.

This distinction matters for Indian TP practice because the architecture of contemporaneous documentation, under Rule 10D, the erstwhile Form 3CEB and now Form 48 under the 2025 Act, assumes a single preparer's judgment trail. A TPO can ask why a particular comparable was included and expect an answer from one analyst or one firm. A pipeline of five narrow agents does not fit that assumption. If a data-extraction agent misclassifies a transaction, a downstream benefit-test agent may inherit that error and proceed regardless, since it is not designed to question upstream inputs, only to execute its own task. The more likely failure mode in a multi-agent TP workflow is not a single agent producing a wrong answer, but an error propagating silently across a handoff that no one is specifically assigned to audit. India's TP documentation requirements, safe harbour disclosures and APA application forms do not currently address this scenario; they assume one preparer whose competence and good faith can be tested under cross-examination or TPO scrutiny.

Practitioners, and possibly CBDT, will need to consider what a chain-of-custody requirement for multi-agent TP work product should look like before a dispute forces the issue. Knowing that a benchmarking output traces back to a database source, as most current vendor claims are framed, is not sufficient. It would also require knowing which agent touched the data at each stage, what it changed or flagged, and whether a human actually reviewed the boundary between two agents' work rather than only the final output. India is already working through related questions, such as how DEMPE functions performed by AI systems fit within the intangibles framework, and whether GCC functional segmentation holds up against agentic AI restructuring inside captive centres. The handoff-audit issue sits a level below those debates: it concerns not whether AI can perform a TP function defensibly, but whether anyone can reconstruct, after the fact, which of several AI agents was responsible when something went wrong. Given how current documentation standards are framed, that question is more likely to surface first in a TPO's show-cause notice than in a policy paper, and practitioners relying on multi-agent tools would do well to build their own audit trail across agent handoffs before that happens.

Friday, September 18, 2026

India's New GCC Benchmarking Advice Meets an Agentic AI Problem It Has Not Addressed

For several years, the standard defensive approach for a captive Global Capability Centre (GCC) facing an Indian Transfer Pricing Officer (TPO) has been fairly mechanical: apply TNMM on an operating cost base, benchmark against routine service-provider comparables, and settle within the accepted cost-plus range. A jurisdiction briefing circulated this week for International Tax Review describes FY2025-26 as a year in which TPO scrutiny of GCC margins and intra-group services intensified, with officers moving toward fewer but deeper adjustments. The advisory response is to segment the GCC into three separately tested activities, namely support, delivery, and decision-rights, with each priced on its own terms rather than blended into a single entity-level margin.

This approach aligns with the OECD's current work at the multilateral level. The OECD has published the full set of public comments on its proposed rewrite of Chapter VII, which governs intra-group services, ahead of a consultation meeting scheduled for November in Paris. Practitioner submissions describe the draft as a substantive rewrite rather than a tidy-up: it requires accurate delineation of what was actually done, by whom, and under what conduct, as the necessary first step before any pricing method is chosen, and it expands the benefit test that has long been the fault line in service fee disputes. Read together, the Indian advisory and the OECD draft point in the same direction: MNEs are being asked to stop pricing services as a single blended category and instead demonstrate, activity by activity, that a real economic function occurred and that an independent party would have paid for it.

This advice may be harder to execute than it appears, not because of documentation gaps in the traditional sense but because of a shift underway in how GCCs operate. India's GCCs are moving away from task-by-task human execution toward agentic AI systems that plan and execute multi-step workflows autonomously. Industry commentary has described Google's move from single-task AI assistants to an agent platform that can be deployed, supervised, and audited across a company's systems as a development that could affect the labour-arbitrage economics on which the GCC sector was built. Separately, EY's GCC survey work finds a large majority of centres already testing agentic technology, with over half piloting agent-based systems specifically. Inside a GCC, this means a workflow once visibly split between a junior analyst performing support work and a senior lead exercising decision-rights judgment can now be executed end-to-end by a single agent, or by a human-agent pair, in a manner not observable from outside the system the way an organisation chart or job description made it observable five years ago.

The Indian TP advisory recommends segmenting the GCC into three tested activities before benchmarking. The OECD requires accurate delineation of the transaction before pricing it. Both instructions assume a documentable boundary between routine support and higher-value decision-making that a TPO or comparability analyst can observe and test. Agentic AI does not necessarily respect that boundary. If an exception-handling agent inside a GCC performs functions that combine what used to be tier-1 support and tier-2 judgment, the functional analysis section of the Local File, which describes who does what, would either need to become considerably more granular about which agent or human made a given call, or risk reverting to the blended, entity-level treatment that both the OECD and Indian practice are trying to move away from.

The practical question is not whether AI adoption in GCCs is occurring; the survey data and industry commentary indicate that it is. It is whether the functional-segmentation and accurate-delineation frameworks currently being refined by the OECD and recommended by Indian advisors were designed for a labour model that may be changing faster than the guidance can be finalised. If accurate delineation depends on establishing what was actually done and by whom, and 'whom' increasingly means a shifting combination of humans and autonomous agents operating across functional boundaries that were previously organisationally distinct, this raises a documentation question that neither the OECD's November consultation agenda nor the Indian safe harbour and Local File templates currently address in detail. For practitioners, the practical implication is to start building functional analyses capable of tracking agent-level activity now, rather than waiting for the guidance to catch up.

Thursday, September 17, 2026

AI-Performed DEMPE Functions and the Limits of Transfer Pricing's Intangibles Framework

The OECD's DEMPE framework was designed to stop groups from parking valuable intangibles in a low-tax entity that holds legal title but performs none of the underlying work. The test traces the relevant functions, development, enhancement, maintenance, protection and exploitation, back to the people who exercise judgment over them: who decided to pursue a research direction, who approved a patent filing, who manages commercialisation risk. For two decades this test has worked reasonably well because those functions were, in fact, performed by identifiable people.

That assumption is being tested where AI systems perform DEMPE functions. Commentary in Australia has raised this issue in relation to the ATO's Practical Compliance Guideline 2024/1 on intangibles migration, which requires multinationals to self-assess and disclose where DEMPE activities for offshore-held intangibles actually occur, using a risk-zone framework that determines how closely the ATO will scrutinise a taxpayer's position. The commentary notes that the PCG's evidence requirements assume human decision-makers, leaving a gap where AI systems (training models, tuning algorithms, monitoring outputs, iterating on protection measures) perform these functions without a person who can be identified as the locus of judgment.

A companion piece extends this analysis to M&A due diligence. It argues that AI systems, including the algorithms, training data and infrastructure, are themselves valuable intangibles under transfer pricing principles, and that a target lacking DEMPE documentation for its AI footprint may carry dormant tax risk that an acquirer inherits on completion.

This is not confined to Australia, and it is not a distant issue for India. India's global capability centre sector has moved well beyond routine coding and support work. Commentary directed at India's GCC market has told clients that where a captive centre designs a core AI algorithm used across the global group, arm's-length pricing cannot be justified by comparing developer hourly rates; the pricing must reflect that creative, value-creating contribution. That is a DEMPE argument, framed in language close to the ATO's, applied to the kind of AI-first GCC that has become common in Bengaluru, Hyderabad and Pune over the last two years.

The timing adds to the difficulty. The Finance Act 2026 safe-harbour rationalisation consolidated software, ITeS, KPO and contract R&D into a single IT-services category, set a flat 15.5% margin, and raised the eligible transaction threshold to ₹2,000 crore. The change was presented as compliance simplification for routine captive units. A captive unit whose engineers are training or fine-tuning a model whose commercial exploitation happens offshore may not be 'routine' in the DEMPE sense, even if it fits comfortably within the new safe-harbour band. A group that elects into the safe harbour for five years on the strength of a routine cost-plus characterisation may be locking in a margin that a DEMPE-consistent functional analysis would not support. That exposure may surface only when the safe-harbour election lapses and a TPO examines what the unit was actually doing.

This raises a harder question about the DEMPE test itself. DEMPE and the 'significant people functions' concept exist to identify who exercises judgment over value creation. Where that judgment is distributed across a model's training run, a human supervisor who approves outputs in bulk, and an infrastructure team in a different jurisdiction, none of these resembles the risk-bearing decision-maker the framework was written for. Tribunals and tax authorities have decades of practice tracing DEMPE to people, but little developed practice for tracing it to a system.

Whether CBDT, the ITAT or the OECD develops a workable answer first, and whether India follows the ATO's disclosure-and-self-assessment approach or adopts something more prescriptive, remains unclear. For now, practitioners advising GCCs and AI-heavy captives should treat DEMPE documentation for AI-performed functions as a present compliance issue, and should examine whether a safe-harbour election understates the functional profile of a unit that is doing more than routine service delivery.

Quantifying AI's Error Rate in Transfer Pricing Benchmarking

A study in The International Tax Journal tested multi-agent AI systems against simulated data covering twenty MNE cases across three sectors, comparing deployment models across eleven countries. Agentic AI cut processing time by 98% and performed deeper functional analysis than a human team in the same window. The study also recorded a 15% error rate in complex functional characterizations, a 22% irrelevance rate in comparable selection, and a 15% hallucination rate in legal citations.

These figures matter for Indian practice because the comparable set is the most contested element in most disputes reaching the ITAT, on turnover filters, FAR mismatches, and functional dissimilarity. A 22% irrelevance rate applied to a twelve-company set implies two or three comparables may not withstand scrutiny. A 15% citation hallucination rate is more serious, since courts elsewhere have sanctioned practitioners for AI-fabricated citations. An ABA Tax Section panel on Section 482 similarly concluded that practitioners remain responsible for method selection and defensibility despite AI assistance.

Until verification protocols match these error rates, AI-assisted local files and benchmarking studies warrant the same partner-level scrutiny as any junior draft, if not more.pemesan

Ingram Micro: TPO's PE Finding Falls Outside a Section 92CA Reference

A Section 92CA reference is meant to test the arm's length price of a specifically identified international transaction. The Assessing Officer's reference defines that scope, and a TPO cannot expand it.

In Ingram Micro (India) Exports Pte. Ltd. v. DCIT, the AO referred the matter to the TPO based on search statements that the Indian affiliate was "carrying on the actual business" of the Singapore entity, a general assertion rather than an identified transaction. The TPO went further, holding that the assessee had a Permanent Establishment in India and raising an adjustment of about Rs. 61.32 crore, feeding a proposed assessment of roughly Rs. 72.39 crore.

The Mumbai ITAT set aside the order. It held that a TPO's jurisdiction under Section 92CA is confined to the arm's length price of a specifically identified transaction. PE existence under Article 5 and taxability of profits under Article 7 remain matters for the AO, not the TPO.

Practitioners reviewing TPO orders built on broadly worded references should examine whether the reference itself identifies an actual transaction before accepting findings made on the strength of it.

AI Washing and Transfer Pricing Benchmarking: Closing the Due Diligence Gap

Vendors pitching transfer pricing documentation software routinely describe an "AI-powered" comparability search or an "AI copilot" that drafts a local file in a fraction of the usual time. Few buyers ask what is actually doing the work: whether a large language model is screening the database for functional comparability, or whether a rules-based screen performs the filtering while the model only writes the narrative afterward. A recent ICAI article names this gap "AI washing," the practice of exaggerating or mislabelling how much genuine AI capability sits behind a product, and sets out red flags that map closely onto a TP due-diligence checklist: vague "AI-driven" claims with no named model class or training data, cherry-picked success anecdotes without a baseline or error analysis, and no documented model validation, drift monitoring, or bias testing. The article also cites the Australian Taxation Office's practice of publicly documenting its own AI use, including subjecting itself to performance audits of AI governance, as a benchmark against which both tax administrations and software vendors could be judged. Current market activity gives practitioners something concrete to test against that yardstick. TP documentation platforms are marketed with features such as "TP Copilot" and "Benchmark AI," using large language models to synthesise functional analysis data into Master and Local Files, and commentary on these rankings notes that independent "AI-native" SaaS platforms have been gaining ground against legacy Big 4 proprietary tools. At least one documentation vendor discloses where the AI stops: it states that AI drafts narrative prose while tables, statutory citations, and the arm's length range are computed in code, with the taxpayer's team required to review and sign off before filing. That kind of disclosure is what distinguishes a defensible use of AI from a washed one, but most buyers have no framework for asking the underlying question, and most vendors face no obligation to answer it unprompted.

Two regulatory developments this month sharpen the stakes. On 11 September 2026, the OECD released a Pillar Two package that establishes a formal peer-review framework to check whether a country's domestic minimum-tax legislation matches what it claims to be: a "qualified" IIR or QDMTT must survive a documented consistency check rather than rely on a self-declared label. Separately, India's transition to the Income-tax Act, 2025 replaces Form 3CEB with a new Form No. 48, which carries additional disclosures on how the arm's length price was determined. Taken together, these developments suggest regulators are moving toward requiring auditable, peer-reviewable evidence behind compliance claims, at a time when TP practice is increasingly relying on AI tools whose contribution to comparability judgments is not documented. It is plausible, though not yet confirmed, that CBDT could eventually require disclosure on a form such as the new Form 48 of whether and how AI tools contributed to the benchmarking methodology behind a reported arm's length price; taxpayers unable to answer with specifics because their vendor cannot either would be in a weak position if that requirement materialises. Until CBDT or ICAI issues more prescriptive guidance, practitioners should press vendors, and press themselves, on specifics: which model class and version was used, whether the comparable set was screened by the model or only narrated by it, whether a human reviewer overrode any AI-suggested comparable and recorded why, and whether that record is retained with the same rigour as the benchmarking study itself. Any "AI-powered" benchmarking claim, whether made by a vendor to a practitioner or by a practitioner to a TPO, is only as reliable as the audit trail behind it.

 

Saturday, September 12, 2026

India Is Negotiating the Future of AI Tax Sourcing at the UN

The arm's length principle, and the DEMPE test that operationalizes it for intangibles, was built on an assumption so basic that nobody bothered to write it down: that a human being somewhere is developing, enhancing, maintaining, protecting or exploiting the asset, and that you can find that human, interview them, and attribute value to their employer based on what they actually did. A recent industry analysis of AI's effect on TP disputes puts the resulting question bluntly - when risk is managed through algorithms and data centres spread across jurisdictions while a smaller, dispersed human team just sets parameters, the practical question becomes where exactly control over risk is anchored. Once AI genuinely performs the function rather than merely assisting a human who performs it, the entire evidentiary architecture of functional analysis has nothing to point to.

This isn't an abstract worry anymore; it's showing up in real guidance in at least two places practitioners should be tracking together. Australia's ATO has PCG 2024/1, its compliance guideline on intangibles migration, and a MinterEllison alert has flagged that its DEMPE evidence requirements assume human decision-makers : leaving a gap precisely where AI performs those functions, with the firm advising multinationals to map AI-related DEMPE functions and reassess their transfer pricing positions now. Separately, and on a longer institutional timeline, the UN Tax Committee's 32nd session mandated two things worth pairing: a subcommittee tasked with a Practical Guide to Implementing Artificial Intelligence for Tax Administrations (due no later than October 2027), and an update to the 2021 UN Transfer Pricing Manual's chapters on intragroup services, intangibles and financial transactions specifically to account for new business models. Meanwhile, at the treaty level, the Fifth Session of the UN Framework Convention on International Tax Cooperation negotiations, held in New York from August 3–13, 2026, saw nations explicitly debate whether AI should be covered under the draft protocol on taxing cross-border services - with live tracking of the session showing India engaged on a closely related scope question, backing further work on Article 2 while questioning how it was drafted.

India is currently playing both sides of this problem without anyone connecting the dots publicly. At the UN negotiating table, India is a rule-maker, shaping how cross-border AI-driven services income gets sourced and taxed under a protocol that will matter enormously given India's position as the world's largest hub for Global Capability Centres and IT-enabled services exports. At home, India is simultaneously a rule-enforcer, and the domestic playbook hasn't caught up. Mainstream GCC audit-defense advisory is already telling clients that if an Indian GCC is designing a core AI algorithm used globally, the arm's length price must reflect that creative contribution, and that authorities are using data-mining tools to compare GCC margins across the industry : in other words, Indian TPOs are already pricing AI-DEMPE contributions using the old human-centric functional analysis framework, years before the UN's own Practical Guide or protocol language is finalized. There's also a second, quieter version of this same problem sitting inside the OECD's plumbing: the OECD has just published the public comments on its own revision of Chapter VII (intra-group services guidance), with a consultation meeting set for November 2026 - the very guidance that ITAT benches are currently applying to decide intra-group services benefit-test disputes, using a chapter the OECD itself has decided needs a structural rewrite.

The open question worth raising is whether India should be trying to actively harmonize its emerging domestic AI-DEMPE case law with the position it's taking at the UN negotiating table - or whether it is, without quite intending to, building two separate and eventually incompatible bodies of doctrine: one forged in ITAT orders and CBDT safe harbour notifications that assume a human designed the algorithm, and another being drafted in a UN conference room that may define AI-driven service income sourcing on entirely different terms.

Wednesday, September 9, 2026

The UN Just Put AI on the Transfer Pricing Table

Start with the problem most transfer pricing practitioners are tracking the wrong forum for. When people think about how AI-delivered cross-border services will be taxed, the reflex is to look at the OECD - Pillar One, Amount B, the endless refinements to the arm's-length principle for baseline distribution and marketing functions. That's not wrong, but it's incomplete. There is a second, quieter negotiation running in parallel at the United Nations that could end up mattering just as much, and it is the one actually putting AI's name on paper right now.

Here's what changed. The UN Framework Convention on International Tax Cooperation has been negotiating, since early 2025, a framework treaty plus two 'early protocols' - one on dispute resolution, and one specifically on taxing cross-border services. The Co-Leads published a draft text of that services protocol on 20 July 2026, and during the Fifth Session of negotiations, held at UN Headquarters from 3 to 13 August 2026, multiple member states pushed to have artificial intelligence explicitly captured within its scope. That same session saw other states raise concerns about overlapping nexus claims for service fees, and a separate bloc pushing for optionality in how the protocol's provisions get applied. In other words: the machinery for deciding which country gets to tax AI-delivered services is being built right now, in a forum most Indian TP practitioners aren't reading transcripts of.

Why this matters more than it looks: India occupies an unusually exposed - and unusually powerful - position in this specific fight. On one hand, India has historically been among the loudest voices for expanding source-country taxing rights over digital and cross-border service income; the UN process exists substantially because countries like India argued the OECD's two-pillar solution didn't go far enough for market/source jurisdictions. On the other hand, India is also the world's largest base for AI-enabled global capability centres and IT/ITeS delivery - the exact category of cross-border service flow this protocol is trying to pin down. If the AI carve-out in the services protocol ends up defining nexus or taxing rights in a way that diverges from how OECD Pillar One and Amount B treat the same AI-delivered functions, Indian-headquartered groups and Indian subsidiaries of multinationals could find themselves benchmarking the same intercompany service flow against two different rulebooks depending on which counterparty jurisdiction is involved. That is not a hypothetical compliance headache; it is a structural one, because transfer pricing documentation is built around a single delineated transaction and a single most-appropriate-method choice, not a dual-track nexus test.

There is also a functional-analysis problem lurking underneath the treaty politics. Cross-border services protocols, going back to earlier UN work like Article 12B on automated digital services, tend to draw bright lines based on where a service is 'delivered' or 'consumed.' AI complicates that in the same way it complicates DEMPE analysis for intangibles: an AI system trained in one jurisdiction, fine-tuned or RAG-augmented with client data in a second, and delivering inference-based output to end customers in a third doesn't map cleanly onto any single-jurisdiction nexus concept the drafters are likely working from. If negotiators write a definition of 'AI services' into the protocol without engaging with how multi-jurisdictional AI pipelines actually function, they risk creating a nexus rule that transfer pricing professionals will spend the next decade trying to reconcile with functional reality - much the way 'significant people functions' language under the OECD's authorized approach took years of practice to operationalize.

The open question, and the one worth writing toward rather than around: does India's negotiating position on this protocol actually reflect an analysis of how Indian GCCs and IT exporters would be affected if AI services get a distinct, UN-defined nexus test that differs from OECD treatment - or is India's source-country advocacy here running on inertia from an earlier era of BPO and call-centre economics, before AI-native delivery models existed? That's not a rhetorical question so much as a genuine gap in the public record; the UN's own tracking shows the draft protocol text was only published in July 2026 and is still being contested clause by clause. A practitioner with both technical AI fluency and TP grounding is unusually well positioned to make that case publicly before the text hardens - which may be the real opportunity here, separate from whatever the final protocol says.

 

Who Performed the Function?

A multinational files a patent for a new compound. Tax authorities in two countries want to know which entity in the group is entitled to the royalty stream the patent will eventually generate. The test they apply is old and, until recently, reasonably reliable: find out who performed the Development, Enhancement, Maintenance, Protection and Exploitation functions for the intangible - DEMPE, in transfer pricing shorthand - because whoever bears the risk and does the deciding gets the residual profit. For decades this worked because you could point to a building: a lab, a notebook, a committee that signed off on which candidate molecule to pursue. Increasingly, the person you would point to typed a prompt, set an agent running overnight, and reviewed a shortlist of candidates over coffee the next morning. The system generated the options, ran the simulations, in some cases proposed the formulation that went into the filing. Who, exactly, performed the function? 

This is not a hypothetical irritating a handful of transfer pricing specialists. An Australian tax advisory has recently flagged what it calls a DEMPE-shaped hole - cases where AI itself is doing the value-creating work and the existing test simply has no place to put that. A parallel problem is surfacing in patent law, where AI-generated inventions are quietly breaking the doctrine of inventorship the same way they are breaking the tax doctrine of significant people functions. Two different bodies of law, built for different purposes, are running into the identical wall: both assume that value is created by an identifiable human making an identifiable decision, and both are discovering that the human's contribution has thinned to something closer to selection and approval. 

 The instinct is to treat this as a compliance problem for corporate tax departments. It is that. It is also a preview of a question that is going to reach every individual who has taken the advice - now common, and not wrong - to stop being a mere user of AI and start being a builder with it. The advice is sound as far as it goes. But it assumes the world has a stable way of recognising, after the fact, that a given piece of work was built rather than merely retrieved. Increasingly, it does not. If the design decision, the judgment call, the actual creative leap happened inside the model, and the human's visible contribution is a prompt and a sign-off, then whatever institution eventually has to decide who gets the credit - a patent office, a performance review, a promotion committee, a funding panel, or a tax auditor deciding which entity in a group actually earned a royalty - is going to ask the same question the DEMPE test is asking of multinationals right now. Where, precisely, did you sit in this? 

It gets sharper still. Some of the institutions asking that question are, at the same moment, handing the question itself to AI. A peer-reviewed study on agentic transfer pricing work found it running nearly a hundred times faster than a human analyst, and wrong in roughly one case out of five or six. That is the position builders are actually in: producing work at a pace where the artifact looks finished, inside a system that is simultaneously trying to verify whether the artifact - or the reasoning behind it - was genuinely someone's, at a speed that does not allow careful verification either. Credit and reliability are being tested at the same moment, by the same overstretched apparatus, and neither test is mature. 

None of this means the advice to build rather than merely search was wrong. It means the second half of that advice has arrived earlier than expected: it is not enough to have used AI to produce something. What will be asked, by whoever is deciding whether the work was yours, is narrower and harder - what did you decide that the model could not have decided on its own; what judgment did you make that the output alone does not show; what risk did you actually carry if the decision turned out to be wrong. That residue is what DEMPE is trying, clumsily, to locate inside multinational R&D right now, and it is what will eventually be asked of anyone whose work passed through an agent before it reached a reader, a patent examiner, or a manager. 

The practical instruction for builders is not to use AI less. It is to keep a different kind of record than most people currently bother with - not a log of what the model produced, but a trace of the decision points where a different judgment would have produced a different outcome, and where that judgment was demonstrably yours. Tax authorities are already learning, case by case, that the old test - find the person who performed the function - cannot be answered by pointing at a finished patent or a tidy set of transfer pricing documentation. It has to be answered by reconstructing a decision trail. Builders who have not been keeping one are going to discover, the way multinationals currently are, that the absence of a trail is itself the answer to the question of who performed the function - and it is rarely the answer they wanted.

Tuesday, September 8, 2026

Three Days to Prove Arm's Length

Start with the fact pattern, because it's more revealing than any policy paper. A Lloyd's India-regulated insurance entity received a TPO show-cause notice giving it three days to produce documentary support for its transfer pricing position. It didn't manage it in time - not because the evidence didn't exist, but because its compliance bandwidth had gone into IRDAI regulatory requirements rather than TP file-building. By the time of the appeal, the taxpayer had assembled contemporaneous emails, contracts and invoices that plausibly supported its position. The ITAT Mumbai admitted this evidence under Rule 29 and sent the matter back for fresh adjudication. On the surface this is a routine remand order. Read against the backdrop of where Indian tax administration is heading, it's a warning shot.

What's changed is the operating environment around these notices. CBDT has been public about using AI-driven risk analytics - NUDGE, Project Insight, INTRAC - to flag taxpayers and compress the distance between detection and action; the department has described the next phase of AI deployment as 'more intense.' Separately, the Income-tax Act, 2025 and the Income-tax Rules, 2026 were framed explicitly as ushering in an algorithm-based, technology-driven assessment ecosystem, with the CBDT Chairman describing simplified statutory language as something that will 'enable algorithm-based implementation.' None of this is inherently bad - faster detection and cleaner drafting are good things. But faster detection paired with unchanged (or shortened) response windows means the gap between 'flagged by the algorithm' and 'must produce five years of DEMPE-level documentation' keeps narrowing, while the underlying task - assembling functional, contractual and economic evidence for related-party transactions - hasn't gotten any easier or faster to do properly.

Here is the piece a generic Taxmann write-up on this ruling will miss: India has, in parallel, been building an entire vocabulary for governing AI systems responsibly. The Finance Ministry has laid out how RBI's FREE-AI framework and MeitY's India AI Governance Guidelines are meant to apply to financial-sector AI, organised around principles like 'Accountability,' 'Understandable by Design,' and 'Trust is the Foundation.' These are aimed at banks and NBFCs deploying AI in lending and risk models. But the tax department's own algorithmic risk-scoring - the system that generates the show-cause notices and compressed timelines that produced this ITAT remand - sits outside that governance conversation entirely. Nobody is asking whether NUDGE or INTRAC's flagging logic is 'understandable by design' to the taxpayer who receives a three-day notice, or who is 'accountable' when an automated flag compresses due process past the point where a genuine, good-faith taxpayer can respond. The state has written AI governance principles for everyone except itself.

The CBDT's own APA numbers hint at how sophisticated taxpayers are already responding to this asymmetry: a record 220 APAs signed in FY2025-26, cumulative signings past 1,035, and a Finance Act 2026 that consolidated safe harbour categories and streamlined APA administration. APAs and safe harbours are, in effect, taxpayers pre-negotiating their way out of exactly the compressed, algorithm-driven scrutiny process that produced the insurance company's evidence crunch. That's a rational choice if you can afford the APA application fee and multi-year process. It's not a choice available to a mid-sized captive or a regulated entity whose compliance bandwidth is already stretched across sectoral requirements, as this case shows. The open question for practitioners to sit with: as India's tax administration becomes more algorithmic, should the same accountability and explainability standards the government is asking of private-sector AI deployers - logged reasoning, contestability, human-in-the-loop review before adverse action - apply symmetrically to the department's own risk-scoring systems, especially once (not if) that scoring logic extends into TP-specific audit selection and comparable-set flagging? Right now, that symmetry doesn't exist, and the insurance company's three-day notice is what the gap looks like in practice.

Monday, August 31, 2026

Tax Deadlines Need Grid-Style Planning

Monday is deadline day for the roughly two crore taxpayers filing ITR-3 and ITR-4 for assessment year 2026-27, the freelancers, small traders and professionals who got the extra month that salaried filers did not. And true to form, the run-up has produced the same story administrators watch unfold every year: login failures, slow challan updates, and a Karnataka taxpayers' association writing to the CBDT about credentials that work on the third attempt but not the first. The department's answer, reported this weekend, was to keep the helpdesk open through the night rather than move the date.

the Income Tax Department has decided to keep its helpdesk operational 24x7 until 23:59 on August 31st

Read that sentence carefully and it tells you something about how large public systems learn, or fail to. The department already knows, to the day, when the surge will hit; it built the staggered calendar precisely to spread that load, moving business filers to August so July would carry only salaried returns, as the report in BusinessToday on this deadline makes clear. That is real design thinking, and it worked, mostly. But the response to the predictable residual crunch is still human overtime: longer helpline hours, staff on call, a circular reminding everyone to save drafts. What is missing is the habit of treating the last three days of a filing window as a known demand spike to be engineered for in advance, the way a power utility plans for peak load on the hottest afternoon of the year, and not as an emergency to be staffed through after it begins.

I have watched this pattern from inside a tax administration for years: the software teams are excellent, the front line staff heroic on deadline week, and yet the institutional memory resets every season. Nobody owns the question of what tested, reserved capacity the portal needs on day minus one. A public organisation that logs this exact data every single year, hourly login volumes, exact failure points, has no real excuse for surprise. The fix is not more goodwill from the helpdesk. It is a standing peak load protocol, reviewed and stress tested months before the date, treated with the same seriousness a grid operator gives a heatwave.

#IncomeTax #ITRFiling #DigitalGovernance #TaxAdministration #PublicSector #CBDT #GovTech

Wednesday, July 15, 2026

A Payroll Wedge Just Vanished

India's Comprehensive Economic and Trade Agreement with the United Kingdom enters into force today, and most of the coverage will lead with the 99 percent duty-free access on tariff lines. The more interesting instrument arrived beside it. The Double Contribution Convention exempts Indian professionals on temporary UK assignments (five years, extended from three) from paying into Britain's National Insurance while they continue to contribute at home. The report in India Briefing pegs the annual saving at roughly USD 500 million, covering 90 to 95 percent of Indian professionals sent through Indian employers. That number matters less than the frame. India's largest single export is not steel or leather; it is people-hours priced in foreign currency. Tariff talk fixates on goods, but the real friction in services trade has always been a payroll wedge: mandatory contributions the visiting worker never gets to draw down. Neutralising that wedge inside a trade treaty is a subtler innovation than any tariff schedule, and probably a more durable one. The next FTA worth watching is the one that repeats this move.

#CETA #IndiaUKTrade #DCC #ServicesTrade #FreeTradeAgreement #GlobalMobility #SocialSecurity

Tuesday, July 14, 2026

The Line Above Four

India's June CPI print landed at 4.38%. It sounds unremarkable. It is not. That is the first time in seventeen months the number has crossed the RBI's four percent target, as recorded in a Bloomberg report on Monday. The reading stayed inside the tolerance band. But a line was quietly crossed, and lines matter in monetary policy the way statute matters in tax administration: not because breaking them is catastrophic, but because they change the standard of proof.

The RBI cut 125 basis points last year and pushed liquidity into the banking system on a scale that would have been unusual in a tighter era. The private capital cycle was supposed to receive that liquidity and put it to work. Then food happened. Nearly forty percent of the consumer basket in India is food, and this year's monsoon has run below normal. The government's own Monthly Economic Review argues the economy is now less exposed to rainfall than before, and that is broadly true in the aggregate. But the price index does not measure the aggregate. It measures households in their kitchens.

From inside a national tax administration, one learns that macro forecasts and taxpayer reality diverge more than models suggest. TDS collections track nominal transactions. Nominal transactions carry inflation inside them. Every basis point of CPI drift shows up months later in the shape of the tax base, in the composition of refunds, and in the arguments taxpayers make about real versus nominal incomes. Inflation is not just a monetary story. With a lag, it becomes a tax administration story too.

The immediate temptation, when a target line is crossed, is to over-read the print. One month is a data point, not a trend. But the WPI figure released the same week printed close to its three-year high. Two indices, drifting apart, telling different stories about the same economy. Wholesale is a producer's world; retail is a consumer's. When they diverge this widely, the middle layer of small firms, informal wage earners and first-year borrowers absorbs the friction.

The takeaway is not for the RBI. They will do what the numbers oblige them to do. The takeaway is for anyone building fiscal, tax or investment plans for the second half of this year. Assume a firmer floor under prices, and design for it now, quietly, before it becomes fashionable to do so.

#IndiaEconomy #Inflation #RBI #MonetaryPolicy #CPI #WPI #TaxPolicy #FiscalPolicy

Monday, July 13, 2026

One Screen, Two Statutes

A small notice on the e-Filing portal this week is worth pausing on. The department has quietly rolled out an integrated payment module that lets taxpayers pay dues under the Income-tax Act, 1961 for periods up to FY 2025-26 and under the Income-tax Act, 2025 for Tax Year 2026-27 onwards, all from a single interface. The notice on the Income Tax Department portal puts it plainly:

Seamless payments now enabled across both the Income-tax Act, 1961 and the Income-tax Act, 2025.

That word, seamless, does a lot of work. India is running two direct tax statutes in parallel, one for closing out old years and one for opening new ones. Every legal transition of this scale creates a temptation to make the citizen learn the transition too, to force them to pick which Act their payment belongs to, to route them through different portals, different challan formats, different mental models. The interesting bit of engineering here is the opposite instinct: absorb the complexity inside the system so that a person paying a demand from AY 2022-23 and a person paying advance tax for TY 2026-27 use the same three clicks. The citizen does not need to know which statute is doing the arithmetic.

There is a wider lesson for public administration in this. The measure of a good transition is how quickly it becomes invisible to the person on the other side of the counter. Officers see two Acts, two sets of rules, two saving clauses, two mental frameworks running simultaneously. The taxpayer should ideally see one screen. That gap, between the complexity we carry internally and the simplicity we present externally, is where administrative craft lives. It is worth building for. And it is worth defending against every impulse to expose the plumbing.

#IncomeTax #IncomeTaxAct2025 #CBDT #TaxAdministration #DigitalIndia #eFiling #TaxReform

Sunday, July 12, 2026

UPI Prepares For Machine Users

Between the noise of frontier model releases and the ITR filing rush, a quieter Indian move this week deserves more attention. According to a report in Business Standard, NPCI is working on a Unified Agent Protocol so that AI agents can, with the user's permission, make payments over UPI. The plumbing is being drawn now, in consultation with the industry, and it will need RBI clearance before it goes live.

This is not a small feature. It changes who the "user" of a national payments rail is. For eighteen years of digital India, the working assumption has been one human, one device, one intention per transaction. UAP contemplates a different assumption: a piece of software acting on your behalf, initiating value transfers between banks, at speed, at scale, and often when you are asleep.

The rail, not the app, is the state's job

The instinct in newsroom coverage is to ask which AI assistant will pay first, which quick-commerce platform will move first, which bank will onboard first. That framing misses the point. The private sector will produce the agents. But agents cannot talk to a payments system unless someone builds a shared, trusted way of introducing them, verifying them, and revoking them when they misbehave. That is a public infrastructure question, not a product question.

The same reporting notes that Visa is building a Trusted Agent Protocol, Google has an Agent Payments Protocol, OpenAI has an Agentic Commerce Protocol, and Pine Labs has P3P. Each of these is a private schema competing for adoption. India's answer is different in shape: a common protocol, sitting above private agents but below the bank rails, run by a body that already coordinates the industry. That shape difference matters more than the technical details.

Why a state adjacent register is the right answer

An AI agent that moves money on your behalf is, in effect, a narrow private power of attorney executed at machine speed. The three questions the law has always asked about such powers are unchanged. Is the delegate real? What is the scope? What happens when scope is exceeded? UAP appears to want answers to all three: a registry, an authorisation envelope with spending limits, and audit logs that allow a payment to be reconstructed after the fact. None of these can be credibly provided by any one of the competing private schemas, because their commercial incentives run in the other direction.

What this means for tax and compliance

From inside a large tax administration, the second order effects here are more interesting than the first order ones. Consider three.

Attribution of transactions

Every agentic payment is a transaction the user initiated in intention but not in execution. Tax law has thin machinery for that distinction. When your agent buys a cross border subscription, or repeatedly tops up a wallet, the audit trail must show that the human principal, not the agent, is the taxable person. A registry of agents, plus a log of who authorised which agent to spend how much within what window, is exactly the primary evidence a revenue officer will one day want. UAP, incidentally, is building that evidence layer.

Fraud patterns will migrate

Every payments innovation is followed by a fraud innovation. UPI's own history is proof. When agents start executing recurring low value purchases, the fraud will not be brute impersonation of the human. It will be quiet capture of the agent, either by compromising the credential or by prompt level manipulation of what the agent decides to buy. Rules that ask only whether the payment was authorised will not catch this. The real question becomes whether the decision the agent took was one the user would have taken. That is a new class of dispute, and the consumer protection frameworks around UPI will have to grow into it.

The invoice, the payment log, the return

If a routine grocery order is executed by an agent that also holds the user's GSTIN and preferences, the natural next step is that the invoice, the payment log, and the return pre-fill start speaking to each other automatically. Tax administrations everywhere have been talking about pre-filled returns for a decade. Agentic commerce is what finally forces pre-fill to become the default rather than the exception.

What a tax department should be doing this quarter

Two things, quietly.

First, seat someone in the room during UAP design. Not to slow it down, but to make sure the log schema captures the fields a revenue authority will one day need: agent identity, principal identity, spend envelope, timestamp, merchant category, and the human confirmation trace. Retrofitting those fields later is always more expensive than agreeing them now.

Second, begin scenario work on what pre-filled returns look like when the underlying spend is agent initiated. Category assignment, personal versus business use, and the taxpayer's ability to challenge an entry generated by software the taxpayer barely understands. These are not futuristic problems. If UAP moves at UPI speed, they arrive within three assessment cycles.

India tends to build payment rails first and think about their tax and legal downstream later. UPI is itself the case study. There is a narrow, useful window here to do it the other way round.

#UPI #NPCI #AgenticAI #DigitalPayments #IndiaAI #PublicInfrastructure #TaxAdmin #UAP

Post-SAP Labs Rulings Tighten Scrutiny of Comparability Filters in TP Benchmarking

The Karnataka High Court has disposed of a batch of transfer pricing appeals that had been pending since the Supreme Court's 2023 judgme...