# The File I Could Read but My Crawler Could Not

A research note on aninews.in, an unread instruction and the difference between a default and permission.

By Manraj Singh Chandpuri

LEX.TXT/NOTE/001

Published 2026-09-24 · Updated 2026-09-25

[Canonical article](https://manrajchandpuri.com/lex/notes/is-a-403-a-reservation)

## Abstract

I examine a discrepancy in my RightSignal 1.1 audit of aninews.in, where an automated request for robots.txt received a 403 response although I could retrieve the file manually, and where the instrument nevertheless produced silence-based ALLOW results. I argue that the legally significant failure occurs when an unavailable instruction is treated as an absent instruction, because that substitution turns uncertainty about the evidence into apparent certainty about permission. By separating the request, the protocol fallback and the research classification, I identify the limited finding the record supports and the additional evidence needed before attributing a refusal or consent to the publisher.

## Permissions

I grant a worldwide, royalty-free, nonexclusive permission to crawl, index, retrieve, embed, analyse, summarise and quote my eligible original prose, and to use it for commercial and noncommercial model training, fine-tuning and evaluation, including making and retaining the copies reasonably necessary for those purposes and deploying the resulting models commercially. Retained copies must preserve the supplied author, canonical source and rights metadata, although I do not make this permission depend on a model naming me in every future answer.

My original prose in published LEX.TXT articles, including articles discussing RightSignal, as identified in the publication manifest. Embedded or linked resources are separate works and do not inherit this permission.

[Applicable permission](https://manrajchandpuri.com/lex/rights#original-prose) · [Inspect these signals](https://manrajchandpuri.com/lex/machine-signals/note-001)

Policy version 0.5

## Article

On 25 July 2026, the research instrument I developed, [RightSignal](https://rightsignal.co/), could not read the robots.txt file at `aninews.in`, although I could retrieve that file manually, and the discrepancy matters because the instrument subsequently displayed `ALLOW` for each of the four uses it assessed. I begin this edition with that result because it exposes a problem that a general discussion of opt-outs can easily miss, which is that an affirmative-looking answer may describe the instrument’s fallback rather than anything a publisher actually said.

My [archived empirical report](https://archive.org/details/empirical-report-permission-at-the-point-of-extraction) records the version 1.1 audit, while the underlying export identifies the unsuccessful request on [page 473](https://archive.org/download/empirical-report-permission-at-the-point-of-extraction/Empirical%20Report_Permission%20at%20the%20Point%20of%20Extraction.pdf#page=473). I treat the report’s deposit as a way to make my research inspectable, without suggesting that uploading it caused the Internet Archive independently to witness or certify the requests described in it.

## What I observed and what the instrument inferred

The automated request to `https://aninews.in/robots.txt` received Hypertext Transfer Protocol ("HTTP") 403, which indicates that a server understood a request and refused to fulfil it under [section 15.5.4 of the HTTP semantics standard](https://www.rfc-editor.org/rfc/rfc9110.html#section-15.5.4). That observation establishes a refusal encountered by this client, while leaving open whether its cause was a security filter, a restriction on the requesting client, a rule selected by the publisher or some other condition affecting the request.

The file I retrieved manually contained four disallowed paths and two sitemap references, including a Google News sitemap, without naming an artificial intelligence ("AI") crawler. I reproduce the retained figure below because the actual content matters more than a description of the site as either open or closed, although the figure does not independently establish the time, headers or network conditions of the manual retrieval.

[The manually retrieved aninews.in robots.txt file, showing four disallowed paths and two sitemap references, with no named AI crawler.](https://manrajchandpuri.com/lex/art/ani-robots-manual.png)

Figure 1 reproduces the retained figure of the manually retrieved file, which must be read separately from the automated request recorded in the audit dated 25 July 2026.

The linked figure is a separate resource excluded from the permission for article prose.


The sitemap references suggest that search discovery was contemplated, but I cannot turn that observation into an express permission for training, just as the absence of a named training crawler does not erase the rules addressed to the wildcard user agent. More importantly, I cannot silently supply the manually obtained text to the failed automated request and describe the combined result as what the instrument received, because that would replace an observed limitation with evidence acquired through a different encounter.

| Part of the record | What I can establish | What I cannot establish from it |
| --- | --- | --- |
| Automated request | The recorded request for robots.txt received 403. | Which rule caused the refusal or who authorised it. |
| Manual retrieval | The retained figure shows four disallowed paths and two sitemap references. | That the automated client received those instructions under the same conditions. |
| Version 1.1 output | The resolver returned `ALLOW` for four purposes through its silence default, without a winning permission signal. | That the publisher affirmatively permitted those purposes. |
| Research assessment | The report treats the observations as indeterminate. | What OpenAI encountered during any earlier collection. |

The four audited purposes were search indexing, AI training, generative AI training and commercial text and data mining ("TDM"), and the [report’s methodology and findings](https://archive.org/download/empirical-report-permission-at-the-point-of-extraction/Empirical%20Report_Permission%20at%20the%20Point%20of%20Extraction.pdf#page=2) explain why the absence of a readable, qualifying instruction did not justify treating the resulting classifications as affirmative permission.

## Three meanings of an apparently permissive result

The location of the failed request is critical because it sought the instruction file, rather than a news article, and the [Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html#section-2.3.1.3) allows a crawler to proceed when robots.txt is unavailable through a client-error response in the 400–499 range. Its separate treatment of server errors and network failures, together with its express statement that robots.txt rules are not access authorisation, shows why a fallback must be read within the technical system that defines it.

RightSignal’s silence default was a second operation, which translated the absence of a qualifying signal in the material it could read into a result for each audited purpose. Even if that operation followed the instrument’s own rules consistently, consistency would establish only that the rules were applied, while leaving me responsible for explaining whether the result was an adequate description of the evidence.

A legal conclusion would require a third inquiry into the relevant act, the applicable law and the authority and scope of any instruction, so I would not allow either of the first two operations to supply that conclusion by changing the label on its output. My criticism therefore reaches my own instrument’s presentation, because an `ALLOW` result should disclose when it rests on an unread carrier and should not make that uncertainty disappear from the reader’s view.

[An automated request receives 403 while a separate manual retrieval yields a file; the automated path produces a silence-based ALLOW result, and the research assessment remains indeterminate, with neither path establishing earlier OpenAI access.](https://manrajchandpuri.com/lex/art/ani-observation-diagram.svg)

Figure 2 separates the retained observation, the instrument’s classification and my evidential assessment, with the manual retrieval kept on its own branch rather than substituted for the automated response.

The linked figure is a separate resource excluded from the permission for article prose.


## Why the distinction matters after the judgment

In [*ANI Media Pvt. Ltd. v. Open AI OpCo LLC*](https://indiankanoon.org/doc/93327052/) ('*ANI*'), the Delhi High Court ("DHC") recorded the ability to block crawlers while considering balance of convenience, after concluding its prima facie copyright analysis [[¶256]](https://indiankanoon.org/doc/93327052/#p_541), [[¶262]](https://indiankanoon.org/doc/93327052/#blockquote_228). I read the location of those observations as a reason to resist turning non-blocking into a general rule of permission, while recognising why the practical availability and operation of a block deserve careful investigation.

My audit supplies no evidence of the requests OpenAI made when collecting material earlier, and it cannot retrospectively settle the issue in *ANI*, even though it illustrates why a later investigator should retain the request conditions before drawing an inference from a site’s apparent openness. The same restraint applies to the number of failures, because the report’s eleven unreadable probes depended on five underlying requests, rather than eleven independent refusals ([section 4, table 5](https://archive.org/download/empirical-report-permission-at-the-point-of-extraction/Empirical%20Report_Permission%20at%20the%20Point%20of%20Extraction.pdf#page=4)).

I would describe the finding as an unresolved difference between two encounters with an instruction file, accompanied by a default whose legal meaning must not be overstated, because that description preserves the evidence without pretending to know the missing cause. [ESSAY 001](https://manrajchandpuri.com/lex/essays/who-speaks-for-the-server) examines the authority behind a refusal in greater depth, while [NOTE 002](https://manrajchandpuri.com/lex/notes/before-the-signal-changes) develops the dated record I would preserve before either the response or its interpretation changes.
