LEX.TXT Note 001
Note 001
The File I Could Read but My Crawler Could Not
A research note on aninews.in, an unread instruction and the difference between a default and permission.
Cite as Manraj Singh Chandpuri, 'The File I Could Read but My Crawler Could Not' (LEX.TXT, 24 September 2026) <https://manrajchandpuri.com/lex/notes/is-a-403-a-reservation>.
Abstract
I examine a discrepancy in my RightSignal 1.1 audit of aninews.in, where an automated request for robots.txt received a 403 response although I could retrieve the file manually, and where the instrument nevertheless produced silence-based ALLOW results. I argue that the legally significant failure occurs when an unavailable instruction is treated as an absent instruction, because that substitution turns uncertainty about the evidence into apparent certainty about permission. By separating the request, the protocol fallback and the research classification, I identify the limited finding the record supports and the additional evidence needed before attributing a refusal or consent to the publisher.
On 25 July 2026, the research instrument I developed, RightSignal, could not read the robots.txt file at aninews.in, although I could retrieve that file manually, and the discrepancy matters because the instrument subsequently displayed ALLOW for each of the four uses it assessed. I begin this edition with that result because it exposes a problem that a general discussion of opt-outs can easily miss, which is that an affirmative-looking answer may describe the instrument’s fallback rather than anything a publisher actually said.
My archived empirical report records the version 1.1 audit, while the underlying export identifies the unsuccessful request on page 473. I treat the report’s deposit as a way to make my research inspectable, without suggesting that uploading it caused the Internet Archive independently to witness or certify the requests described in it.
In motion · scroll or step through
Two encounters with one instruction file
Follow the automated request and the manual retrieval separately, and watch three different operations turn one refusal into an apparently permissive answer.
Step 1 of 8 · The file
Every site can leave instructions at one address
A website can publish instructions for automated visitors in a plain-text file kept at
/robots.txt. A well-behaved crawler, drawn here as a small mechanical spider that walks the web, asks for that file before it asks for anything else, because the file tells it which paths it may visit, and it carries a clipboard on which to note what the file says.Step 2 of 8 · The request
The instrument walks up and asks for aninews.in/robots.txt
On 25 July 2026 my research instrument, RightSignal 1.1, sent an automated request for the instruction file at
aninews.in. The request is the first thing the record must preserve, together with the client that sent it and the conditions under which it was sent.Step 3 of 8 · The refusal
The server answers 403 and the clipboard stays empty
The server returned Hypertext Transfer Protocol ("HTTP") 403, which means that it understood the request and refused to fulfil it. The crawler therefore received no instructions at all, and the response by itself cannot tell us whether a security filter, a rule about this client or a publisher’s choice produced the refusal.
Step 4 of 8 · By hand
A separate, manual retrieval succeeds
When I opened the same address manually in a browser, the file loaded. It contained four disallowed paths and two sitemap references, including a Google News sitemap, and it named no artificial intelligence ("AI") crawler.
Step 5 of 8 · Kept apart
The manual file cannot be handed to the automated request
The two encounters happened through different clients under different conditions, so the manually retrieved text cannot be passed back to the refused request and described as what the instrument received. The record keeps them in separate lanes, and the attempted hand-over is refused.
Step 6 of 8 · The fallback
The protocol lets a crawler proceed after a client error
The Robots Exclusion Protocol treats robots.txt as unavailable when the request receives a client error in the 400 to 499 range, and it allows a crawler to proceed in that case. That rule describes crawler behaviour, and the same standard states that its rules are not a form of access authorisation.
Step 7 of 8 · The default
The instrument turns silence into four ALLOW results
RightSignal then applied its own silence default, which translates the absence of a readable, qualifying signal into a result for each audited purpose. Search indexing, AI training, generative AI training and commercial text and data mining ("TDM") all flipped to ALLOW, although no one had granted anything.
Step 8 of 8 · The assessment
The research record keeps every layer, and stays indeterminate
The report assesses these observations as indeterminate, because an unread instruction is not an absent instruction. The three layers stay stacked and separately visible, namely what was observed, what the protocol and the instrument inferred, and what the research can actually support.
What I observed and what the instrument inferred
The automated request to https://aninews.in/robots.txt received Hypertext Transfer Protocol ("HTTP") 403, which indicates that a server understood a request and refused to fulfil it under section 15.5.4 of the HTTP semantics standard. That observation establishes a refusal encountered by this client, while leaving open whether its cause was a security filter, a restriction on the requesting client, a rule selected by the publisher or some other condition affecting the request.
The file I retrieved manually contained four disallowed paths and two sitemap references, including a Google News sitemap, without naming an artificial intelligence ("AI") crawler. I reproduce the retained figure below because the actual content matters more than a description of the site as either open or closed, although the figure does not independently establish the time, headers or network conditions of the manual retrieval.
The sitemap references suggest that search discovery was contemplated, but I cannot turn that observation into an express permission for training, just as the absence of a named training crawler does not erase the rules addressed to the wildcard user agent. More importantly, I cannot silently supply the manually obtained text to the failed automated request and describe the combined result as what the instrument received, because that would replace an observed limitation with evidence acquired through a different encounter.
| Part of the record | What I can establish | What I cannot establish from it |
|---|---|---|
| Automated request | The recorded request for robots.txt received 403. | Which rule caused the refusal or who authorised it. |
| Manual retrieval | The retained figure shows four disallowed paths and two sitemap references. | That the automated client received those instructions under the same conditions. |
| Version 1.1 output | The resolver returned ALLOW for four purposes through its silence default, without a winning permission signal. | That the publisher affirmatively permitted those purposes. |
| Research assessment | The report treats the observations as indeterminate. | What OpenAI encountered during any earlier collection. |
The four audited purposes were search indexing, AI training, generative AI training and commercial text and data mining ("TDM"), and the report’s methodology and findings explain why the absence of a readable, qualifying instruction did not justify treating the resulting classifications as affirmative permission.
Three meanings of an apparently permissive result
The location of the failed request is critical because it sought the instruction file, rather than a news article, and the Robots Exclusion Protocol allows a crawler to proceed when robots.txt is unavailable through a client-error response in the 400–499 range. Its separate treatment of server errors and network failures, together with its express statement that robots.txt rules are not access authorisation, shows why a fallback must be read within the technical system that defines it.
RightSignal’s silence default was a second operation, which translated the absence of a qualifying signal in the material it could read into a result for each audited purpose. Even if that operation followed the instrument’s own rules consistently, consistency would establish only that the rules were applied, while leaving me responsible for explaining whether the result was an adequate description of the evidence.
A legal conclusion would require a third inquiry into the relevant act, the applicable law and the authority and scope of any instruction, so I would not allow either of the first two operations to supply that conclusion by changing the label on its output. My criticism therefore reaches my own instrument’s presentation, because an ALLOW result should disclose when it rests on an unread carrier and should not make that uncertainty disappear from the reader’s view.
Why the distinction matters after the judgment
In ANI Media Pvt. Ltd. v. Open AI OpCo LLC ('ANI'), the Delhi High Court ("DHC") recorded the ability to block crawlers while considering balance of convenience, after concluding its prima facie copyright analysis [¶256], [¶262]. I read the location of those observations as a reason to resist turning non-blocking into a general rule of permission, while recognising why the practical availability and operation of a block deserve careful investigation.
From the judgment · ANI Media v OpenAI
The two paragraphs NOTE 001 relies on
The note reads the Court’s observation about blocking crawlers together with its location in the decision. The map shows that paragraph 262 sits in the part about interim relief, after the copyright analysis has already closed at paragraph 256.
In the High Court of Delhi at New Delhi
ANI Media Pvt Ltd v Open AI OpCo LLC
The copyright analysis closes“In light of the discussion above, both the purpose test as well as the fairness test under Section 52(1)(a) stand fulfilled. Hence, in my prima facie view, Open AI’s acts of storage of the literary works of ANI for the training of its LLMs would fall under Section 52(1)(a) of the Copyright Act and hence, would not amount to infringement.”
The opting-out option“It is an admitted position that ANI has the ability to block its website vis-à-vis any third-party including Open AI. The opting-out option is available to ANI for blocking the third-party web crawlers from copying their data as well as from scraping their website for the search function/RAG. Despite having an option of opt-out, evidently ANI has not exercised the same. In fact, it has been stated on behalf of Open AI that it has internally blocked ANI’s website from its web crawlers or bots for the purposes of scraping of data. During the course of oral submissions Open AI has also submitted that Open AI itself has blocked ANI’s website from ‘ChatGPT search function/RAG’.”
Quoted from the signed copy · ellipses mark omissions · footnote markers omitted
- ¶256
Decides
This is where the Court finishes deciding whether storage for training infringes copyright, on a prima facie view. Nothing after this paragraph adds to that analysis.
- ¶262
Records
This is where the Court records that ANI could have blocked crawlers. Because it appears under balance of convenience, the note resists reading non-blocking as a general rule of permission.
My audit supplies no evidence of the requests OpenAI made when collecting material earlier, and it cannot retrospectively settle the issue in ANI, even though it illustrates why a later investigator should retain the request conditions before drawing an inference from a site’s apparent openness. The same restraint applies to the number of failures, because the report’s eleven unreadable probes depended on five underlying requests, rather than eleven independent refusals (section 4, table 5).
I would describe the finding as an unresolved difference between two encounters with an instruction file, accompanied by a default whose legal meaning must not be overstated, because that description preserves the evidence without pretending to know the missing cause. ESSAY 001 examines the authority behind a refusal in greater depth, while NOTE 002 develops the dated record I would preserve before either the response or its interpretation changes.