LEX.TXT Essay 002
Essay 002
An Indian Opt-Out for AI Retrieval?
Why the copy made to answer a question deserves its own inquiry under section 52.
Cite as Manraj Singh Chandpuri, 'An Indian Opt-Out for AI Retrieval?' (LEX.TXT, 24 September 2026) <https://manrajchandpuri.com/lex/essays/indias-only-statutory-opt-out>.
In this essay
Abstract
I argue that section 52(1)(c) deserves examination as a possible defence for particular copies made during retrieval, because its reference to an express prohibition gives the rightsholder’s instruction a role that cannot simply be imported into fair dealing. The argument depends on the purpose and function of the storage, rather than on calling a service retrieval-augmented generation, and I test it against technical transmission under clause (b), research under clause (a) and the possibility of durable retention. My conclusion is a limited route for analysing a particular copy, rather than a general power to prohibit an answer or displace every available exception.
I am interested in the copy made between a reader’s question and a system’s answer, because that copy can disappear from a debate organised around training on one side and infringing outputs on the other. A service may fetch a news article to answer a particular question, hold some or all of it while composing a response and then discard it, while another service may answer from a stored collection without contacting the publisher at that moment, so neither the word retrieval nor the presence of a source link tells me enough about the storage to classify it legally.
The distinction becomes important after ANI Media Pvt. Ltd. v. Open AI OpCo LLC ('ANI'), where the Delhi High Court ("DHC") accepted a prima facie fair-dealing argument for training under section 52(1)(a) of the Copyright Act, 1957 ("Copyright Act") [¶256]. I do not read that conclusion as resolving every copy made by an artificial intelligence ("AI") service, because the purpose, retention and function of a query-related copy can differ from those examined in the training claim.
The limits of that interim order matter to my argument because the court recorded that the retrieval-based claim had not been pleaded, although its conclusion addressed the particular outputs before it [¶84], [¶271]. Its observations were confined to the interim application [¶274], while the Division Bench’s notice order of 15 September 2026 directed listing on 8 December 2026 without deciding the appeal’s merits [¶4], so I treat the question developed here as an argument that still needs its own factual and legal examination.
From the judgment · ANI Media v OpenAI
The paragraphs ESSAY 002 relies on
The essay argues that the copy made to answer a question deserves its own inquiry. The paragraphs below show what the Court decided about training, what it recorded about retrieval without deciding it, and the limit it placed on everything it said.
- ¶26 How LLMs work · ¶12–27
- ¶84 Issue 2 · Outputs · ¶56–125
- ¶196 Issues 1 and 3 · Storage for training and fair dealing · ¶126–256
- ¶211 Issues 1 and 3 · Storage for training and fair dealing · ¶126–256
- ¶214 Issues 1 and 3 · Storage for training and fair dealing · ¶126–256
- ¶256 Issues 1 and 3 · Storage for training and fair dealing · ¶126–256
- ¶271 Conclusion and order · ¶271–275
- ¶274 Conclusion and order · ¶271–275
In the High Court of Delhi at New Delhi
ANI Media Pvt Ltd v Open AI OpCo LLC
How retrieval works“Retrieval-Augmented Generation (‘RAG’) is an innovative feature of LLMs that optimises their output. This feature references an authoritative knowledge base outside of its training data sources before generating a response. This feature does not predominantly rely on training data to generate output. Instead, it uses an information retrieval component that utilises the user prompt to pull information from external data storage.”
Retrieval was not pleaded“… this Court is of the prima facie view that the instances given in the plaint alleging infringement are not a result of memorisation, rather they are in the nature of live links, perhaps reflecting RAG technique. Whether outputs produced using RAG would amount to copyright infringement is an aspect which has not been pleaded in the plaint, though it was referred during the course of submissions.”
Clause (c) as a contrast“In Sections 52(1)(aa), 52(1)(ab) and 52(1)(ad), the words “lawful”/ “legally obtained copy” has been used only in respect of computer programmes. Therefore, it cannot be said that “non-infringing copy” would apply to storage of other works unless it has been specifically mentioned. For instance, in Section 52(1)(c) it has been specifically mentioned that for the purposes of providing electronic links, the person responsible must have reasonable grounds to believe that an infringing copy is not being stored.”
A closed space“In the present case, Open AI stores the literary works in a closed space without access to the public. The data obtained by the LLMs for training purposes is used for private purposes. The said data is accessible only to the LLM models themselves. The said data is not publicly available to any human entity either for access or for download. Therefore, in my opinion, the use amounts to being purely private.”
Research as a closed process“The process of “research” is generally an intermediary process in all cases. It is undertaken before an output is generated. … It is normally a closed activity and is not disclosed to the general public. It is always the output of the research that is communicated to the public.”
The copyright analysis closes“In light of the discussion above, both the purpose test as well as the fairness test under Section 52(1)(a) stand fulfilled. Hence, in my prima facie view, Open AI’s acts of storage of the literary works of ANI for the training of its LLMs would fall under Section 52(1)(a) of the Copyright Act and hence, would not amount to infringement.”
The conclusion on outputs“… I am also of the prima facie view that the outputs generated by ChatGPT using RAG technique does not amount to infringement under Section 51 of the Copyright Act since the outputs generated by Open AI were not substantially similar to ANI’s original literary works.”
The Court’s own limit“Needless to say, any observations made herein are only for the purpose of adjudication of the aforesaid application and would have no bearing on the final outcome of the suit.”
Quoted from the signed copy · ellipses mark omissions · footnote markers omitted
- ¶26
Records
The Court’s account of retrieval-augmented generation. The essay notes that the external source need not be a live website fetched afresh for every answer.
- ¶84
Records
The retrieval-based claim had not been pleaded. The essay treats the lawfulness of retrieval copies as a question the decision leaves open.
- ¶196
Decides
The Court mentions clause (c) only as a contrast, to show that Parliament names infringing copies when it means to. It did not decide that clause (c) governs retrieval.
- ¶211
Decides
Training was private because the stored works sat in a closed space. The essay asks whether the same can be said of a fetch made for a particular user’s question.
- ¶214
Decides
Research as a closed process whose output is published. The essay warns against assuming that opening a service to users excludes every research argument.
- ¶256
Decides
The prima facie conclusion on training, which the essay does not read as resolving every copy an AI service makes.
- ¶271
Decides
The conclusion about the particular outputs before the Court, which the essay keeps tied to that record.
- ¶274
Limits
The observations are confined to the interim application, which is why the essay presents its argument as one still needing examination.
The statutory route I would test
Section 52(1)(c) concerns transient or incidental storage for providing electronic links, access or integration, and its conditions include the absence of an express prohibition by the right holder and a qualification concerning awareness or reasonable grounds for believing that the storage is of an infringing copy. I regard that wording as a reason to examine whether a particular retrieval copy falls within the provision, without assuming that a publisher’s objection has the same effect under every other exception.
The court in ANI referred to clause (c) while contrasting its express treatment of infringing copies with the Explanation to clause (a), rather than deciding that clause (c) governed the retrieval arrangements before it [¶196]. My proposed application therefore remains an interpretation to be tested against the architecture and statutory purpose, rather than a rule already established by that judgment.
The strongest case for the interpretation is a fetched copy retained only as an intermediate step in supplying access to, or integrating, the source in response to a question. The strongest objection is that a system composing a new answer may be using the source as material for its own product, rather than storing it for the linking or access function contemplated by the clause, and the appearance of a citation cannot settle that objection because a link may accompany many different kinds of copying.
One answer can depend on several different copies
I begin with evidence of the acquisition copy, any retained source text used in an index, the material supplied to the model for the particular query and the text eventually shown to the user. Those stages need not all occur in every system, while the legal relevance of a numerical representation or extracted passage also depends on what it retains, so the inquiry must follow the actual implementation rather than a diagram treated as universal.
| Part of an illustrative system | The evidence I would seek | The question that evidence helps answer |
|---|---|---|
| Fetch from the publisher | Requested resource, response and purpose of acquisition | Whether the copy was made for this query, indexing or another use. |
| Retained collection or index | Stored content, retention arrangements and reuse | Whether storage is a central continuing function or subordinate to another function. |
| Query-related working copy | Material supplied to the model and its lifetime | Whether the claimed linking, access or integration function explains this storage. |
| Answer delivered to the user | Expression reproduced and relationship to the source | Whether the output presents a separate infringement issue. |
Retrieval-augmented generation ("RAG") uses an external information source in generating a response, but that source need not be a live website fetched afresh for every answer, which is why the DHC’s account of external retrieval should not be expanded into an assumption about every request [¶26]. OpenAI’s crawler documentation usefully distinguishes GPTBot, OAI-SearchBot and ChatGPT-User, while its statement that robots.txt may not apply to user-initiated visits concerns described crawler behaviour and does not itself decide the effect of a statutory prohibition.
In motion · scroll or step through
One answer can depend on four different copies
Follow a reader’s question through an illustrative retrieval system, and see why each copy along the way may call for a different clause of section 52.
Step 1 of 8 · A question
A reader asks about today’s news
A reader asks an AI service about something reported this morning. The model’s training ended long before the article was written, so the service must obtain the article from somewhere else before it can answer.
Step 2 of 8 · Three crawlers
Three crawlers with three different jobs
OpenAI’s crawler documentation distinguishes GPTBot, which collects material for training, OAI-SearchBot, which builds a search index, and ChatGPT-User, which visits pages when a user’s request calls for it. The documentation says robots.txt may not apply to user-initiated visits, which describes crawler behaviour without deciding the effect of any statutory prohibition.
Step 3 of 8 · Copy 1
The fetch copies the article from the publisher
The first copy is made when the article is fetched from the publisher’s server, drawn here as the page the crawler carries away on its back. The evidence that matters is the requested resource, the response and the purpose of the acquisition, because a copy made for this query differs from one made to build an index or to train a model.
Step 4 of 8 · Copy 2
A retained collection may keep the text
Some systems keep fetched text in an index that later questions can reuse. The question is whether that storage is a central, continuing function of its own or remains subordinate to providing access to the source.
Step 5 of 8 · Copy 3
A working copy is handed to the model
For the particular question, the relevant passage is supplied to the model as a working copy. Section 52(1)(c) of the Copyright Act, 1957 concerns transient or incidental storage for providing electronic links, access or integration, where these have not been expressly prohibited by the right holder, and this is the copy the essay asks the clause to examine.
Step 6 of 8 · Function
How long a copy lasts does not decide what it is for
A copy held for two seconds is not incidental merely because it is brief, and a copy held for thirty days is not excluded merely because it lasts. Drawing on the discussion in My Space v Super Cassettes, the essay asks whether the storage serves an identifiable linking, access or integration activity or instead forms a reusable content resource in its own right.
Step 7 of 8 · Other routes
Clauses (b) and (a) offer different doors
Section 52(1)(b) concerns transient or incidental storage in the technical process of electronic transmission, without clause (c)’s condition about express prohibition, and section 52(1)(a) concerns fair dealing for private or personal use, including research. A developer may argue through either door, so a publisher’s prohibition under clause (c) does not automatically defeat every defence.
Step 8 of 8 · Copy 4
The answer shown to the reader is a separate question
The last copy is whatever expression reaches the reader in the answer. Protection for an intermediate copy would not by itself make reproduced expression in the output lawful, and the Delhi High Court’s conclusion about the particular outputs before it remains tied to that record [¶271].
Why duration cannot carry the argument alone
In My Space Inc. v. Super Cassettes Industries Ltd. ('My Space'), the Division Bench discussed transient storage as temporary and incidental storage as subordinate to a principal function, while distinguishing automatically generated storage from permanent hosting [¶59]. I use that discussion with its express qualification that the 2012 amendment did not apply to the circumstances before the court, rather than presenting it as a holding about contemporary retrieval systems [¶58].
The disjunctive wording matters because storage does not necessarily cease to be incidental merely because it lasts longer than one answer, just as deleting a copy quickly does not establish that it was made for a qualifying purpose. I therefore examine both retention and function, asking whether the copy serves an identifiable linking, access or integration activity or instead forms a substantial, reusable content resource in its own right.
That approach creates a narrower argument than a distinction between temporary retrieval and permanent indexing, but it is also more defensible because it does not allow a retention timer to determine a legal category without examining what the stored material does. A developer asserting that a persistent copy remains incidental would need to explain the principal function to which it is subordinate, while a publisher contesting the explanation would need evidence beyond the bare fact that an index exists.
The alternative route through technical transmission
Section 52(1)(b) separately concerns transient or incidental storage in the technical process of electronic transmission or communication to the public, without reproducing clause (c)’s express-prohibition condition. I therefore test that route before suggesting that a refusal under clause (c) defeats the defence for a copy, since the neighbouring provisions must be distinguished through their purposes and wording rather than by treating the more restrictive clause as automatically controlling every overlap.
A developer might argue that a working copy is only part of the technical process through which requested information reaches a user, whereas the publisher might answer that selecting and integrating source material to compose an answer goes beyond that description. I do not think either characterisation can prevail simply by naming the service, because the inquiry concerns the function of the particular storage and whether the claimed technical role explains the copying in question.
My argument survives this objection only as a conditional one, because an effective defence under clause (b) would require its own analysis and would not disappear merely because the publisher had made a statement relevant to clause (c). Conversely, a service cannot establish that defence simply by observing that computers necessarily make temporary copies, since that observation leaves the statutory purpose requirement unanswered.
The alternative route through research
A developer could also invoke section 52(1)(a), which makes the characterisation of the particular dealing and its fairness central to the inquiry, and I do not assume that opening a service to users necessarily excludes every research argument. The DHC’s treatment of closed inputs and publicly available outputs in ANI makes such an automatic distinction especially unsafe [¶211], [¶214].
I nevertheless think that a fetch made to satisfy a particular customer’s information request calls for an explanation different from the model-development process examined in the training claim. The developer would need to identify the relevant dealing and show why its purpose and fairness satisfy clause (a), rather than allowing an earlier finding about training to travel unexamined to every subsequent interaction with a source.
The output remains a further question because protection for an intermediate copy would not by itself establish the lawfulness of reproduced expression delivered to a user, and the conclusion concerning the outputs before the DHC should remain tied to that record [¶271]. I therefore keep access to the source, storage during retrieval and the resulting answer separate without suggesting that the boundaries between them eliminate their factual connections.
What I would require of the prohibition
If clause (c) is the relevant defence, the inquiry requires a statement attributable to the right holder that expressly concerns the links, access or integration at issue, rather than treating an unexplained 403 or an objection to training as necessarily sufficient. ESSAY 001 examines attribution in greater depth, while PROTOCOL 001 explains why even a clearly parsed signal may leave its legal scope unresolved.
I would encourage a publisher to describe the relevant use in clear prose and maintain a consistent structured expression where available, but I reject importing the European machine-readability requirement into a clause that does not state it. The allocation of interpretive responsibility proposed in ESSAY 003 is a practical recommendation whose relationship to legal sufficiency must be assessed under the applicable provision.
The complaint procedure in the proviso also needs independent attention because a published objection is not a completed statutory complaint, and rule 75 of the Copyright Rules, 2013 requires particulars concerning the work, the infringing copy and the grounds of complaint, together with the undertaking to seek the necessary court order. Posting a signal cannot automatically activate the temporary restraint procedure or establish infringement, since those conclusions depend on additional conditions and any competing defence.
I therefore propose a question about a defined copy rather than a general veto over a technology, asking whether its storage falls within clause (c), whether an attributable and appropriately scoped prohibition affects that defence, and whether another route independently protects the dealing. That inquiry gives a publisher’s instruction legal work to do where the statute permits it, while leaving the developer the opportunity to establish the defence its actual conduct supports.