Transfer pricing practitioners frequently encounter this pattern: a taxpayer builds a comparable set, the TPO rejects most of it, and the replacement set resembles the set the TPO has used in assessments of similar taxpayers. The Karnataka High Court's ruling in the SAP Labs India appeals addressed this practice directly. The Court held that a TPO cannot reject a taxpayer's comparables merely to substitute a standard set of comparables routinely used by the Income Tax Department, and that the comparability exercise must be tied to the particular international transaction and conform strictly to Rule 10B. The Court also closed off another common approach, holding that the plus or minus 5% tolerance range under Section 92C is a threshold for determining when no adjustment is required, not an automatic deduction from the arithmetic mean once a transaction falls outside it.
Considered solely as a case about TPO conduct, this ruling restates a familiar principle: tribunals have long required that FAR analysis be conducted properly. Its significance is heightened by timing. A small but growing set of commercial platforms, including Tessera, ArmsLength AI, TPGenie and others, now offer a workflow similar to the one the Court found impermissible, without the government's involvement. These vendors market consistency: a fixed, repeatable accept or reject logic applied to a comparable-set export, uniformly across each row, with human review limited to exceptions. Vendors advertise measurable gains, including claims of freeing up preparation time substantially and achieving high accuracy rates on automated accept or reject decisions. The feature underlying this pitch, a single decision logic applied consistently across every candidate company, closely resembles the feature the Karnataka High Court held a TPO cannot rely on when substituting a standard set for taxpayer-specific analysis.
This ruling does not hold that AI-driven benchmarking is impermissible. Nothing in it addresses AI, and no Indian tribunal has yet evaluated an AI-generated comparable set on its own terms. Its relevance lies in the doctrinal language now available for testing a benchmarking study whose selection logic was a repeatable template rather than transaction-specific FAR judgment. A taxpayer whose TP study relied substantially on an automated accept or reject pass, and whose audit trail records only which template rule fired for which company, may find that this traceability does not, on its own, demonstrate that the comparable search was tailored to the taxpayer's controlled transaction. The same risk applies to the Department: if a TPO's office uses AI-assisted searches that effectively reconstitute the Department's familiar standard set under a different label, taxpayer's counsel could cite SAP Labs against that approach.
Practitioners advising on AI-assisted benchmarking need to consider what a defensible workflow looks like under this standard. One possibility is that a traceable, overridable accept or reject log is sufficient, provided a human reviewer documents transaction-specific reasoning for the final set. Another possibility is that the underlying decision logic itself must be shown to respond to the specific FAR profile of the tested party, rather than being applied uniformly across an unrelated population of candidate comparables. No case has tested this question yet, and vendors are unlikely to raise it themselves. Firms advising clients on adopting AI benchmarking tools, or defending a study built using one, would be well advised to build in an explicit, documented step where a reviewer records why the FAR profile of the tested party justified each inclusion or exclusion, independent of what the algorithm flagged, and to treat AI output as a first-pass screen rather than as the analysis itself.
Although the SAP Labs ruling does not mention AI, its reasoning on standard sets and transaction-specific comparability is likely to inform how AI-assisted benchmarking studies are tested going forward. Practitioners should document the human judgment behind each comparable decision accordingly, rather than relying on the traceability of an automated log as a substitute for that judgment.