LEX.TXT Note 002
Note 002
A Wayback Machine for Content-Use Signals
A proposal for preserving the request, the response and the interpretation without replacing one with another.
Cite as Manraj Singh Chandpuri, 'A Wayback Machine for Content-Use Signals' (LEX.TXT, 24 September 2026) <https://manrajchandpuri.com/lex/notes/before-the-signal-changes>.
Abstract
I propose a dated archive of content-use signals that preserves the conditions of each observation alongside the raw response and successive interpretations, because a later successful request cannot reconstruct an earlier refusal. Using my aninews.in audit, I show how a record can distinguish the automated 403, the separately retrieved file and the report’s indeterminate classification without claiming to establish what another collector encountered. The proposal builds on existing archival formats and treats hashes, timestamps and admissibility requirements as complementary safeguards, while refusing to equate the preservation of bytes with proof of publication, authority or receipt.
A dispute about what a website permitted can outlive the configuration that supplied the answer, which is why I want a history of content-use signals that remains useful after the live page changes. My starting example is the aninews.in discrepancy examined in NOTE 001, where the RightSignal 1.1 audit dated 25 July 2026 recorded an automated Hypertext Transfer Protocol ("HTTP") 403 response while I could retrieve the instruction file manually, leaving two encounters that a later successful fetch should not collapse into one.
The Internet Archive’s Wayback Machine offers a useful model for looking backwards, and its capture calendar already distinguishes successful responses, redirects and errors, although saving a page does not capture an entire site or arrange continuing preservation. I would build the proposed history around the instruction locations and request conditions relevant to a particular observation, allowing a reader to inspect both what was collected and what the collection failed to reach.
A specimen entry for the ANI observation
The table below illustrates how I would organise the retained research, rather than claiming that a new archive has already captured all the fields I propose. The empirical report and underlying export remain the sources for the automated observation and classification, while the figures reproduced in NOTE 001 keep the manual retrieval as a separate item with its own provenance limits.
| Associated record | Material I would retain | Limit that must remain visible |
|---|---|---|
| Automated encounter | The request for aninews.in/robots.txt, its recorded date and 403 result, with the export location. | Unrecorded client or network details cannot be reconstructed by assumption. |
| Manual encounter | The retained figure of the retrieved instruction file, associated with the author’s account. | It cannot be substituted for the automated response or represented as a simultaneous controlled comparison. |
| Instrument interpretation | The version 1.1 silence-based ALLOW outputs for the four audited uses. | The output identifies a resolver result without establishing permission. |
| Research interpretation | The report’s indeterminate classification, linked to the observations it assesses. | The classification does not establish what OpenAI received during earlier collection. |
The report’s distinction between eleven unreadable probes and five underlying requests also belongs in this record, because several checks that depend on the same failed response do not become independent observations simply because they occupy separate rows (section 4, table 5). I would connect each interpretation to its underlying request so that a reader can assess the actual evidential weight rather than count repeated dependencies as corroboration.
In motion · scroll or step through
An archive that keeps every encounter
Watch observations and interpretations join the record as separate, dated cards, and see why a later success is added beside an earlier refusal rather than written over it.
Step 1 of 7 · The live page
Without an archive, yesterday’s answer is lost
A website can refuse a request today and serve it tomorrow. If only the live page is consulted, the earlier refusal disappears, and a dispute about what the site permitted at a particular moment loses its most important evidence.
Step 2 of 7 · Capture 1
The automated encounter is captured with its conditions
The crawler brings back what it met and files it as the first card, which records the request for
aninews.in/robots.txt, its date of 25 July 2026, its 403 result and the location of the export. Details that were never recorded, such as the network conditions, remain marked as unknown rather than reconstructed by assumption.Step 3 of 7 · Capture 2
The manual encounter sits on its own card
The manually retrieved file is kept as a separate card with its own provenance limits. It cannot be substituted for the automated response or presented as a simultaneous, controlled comparison.
Step 4 of 7 · Readings
Interpretations are separate cards, tied to what they read
The instrument’s version 1.1 result of ALLOW for four uses and the report’s indeterminate assessment are recorded as interpretations, each linked by a thread to the observation it assesses. A corrected resolver would add a new, dated interpretation rather than replacing the old one.
Step 5 of 7 · Dependencies
Eleven probes rest on five requests
The report’s eleven unreadable probes depended on five underlying requests. Connecting each interpretation to its request lets a reader see the real evidential weight, instead of counting repeated dependencies as independent corroboration (section 4, table 5).
Step 6 of 7 · Later
A later success is added, never written over
If a later request succeeds, it receives its own card further along the shelf. Replacing yesterday’s refusal with today’s success would erase the very difference the archive exists to preserve.
Step 7 of 7 · Seal and clock
A digest and a timestamp prove less than they seem
A digest and an independently issued timestamp can show that the retained bytes existed and have not changed since, but the time-stamping standard does not show that a publisher served them, authorised them or sent them to a particular recipient. In India, section 63 of the Bharatiya Sakshya Adhiniyam, 2023 ("BSA") governs how such electronic evidence is produced and certified.
What a future capture should preserve
For a newly designed capture, I would retain the requested address, method, relevant request headers and client configuration, start and completion times in Coordinated Universal Time ("UTC"), redirects, response status, response headers and returned bytes, while marking a timeout or missing body explicitly. The instruction file, linked policy and article-level response should remain separate associated records within a disclosed collection window, because successive requests are not a perfectly simultaneous account of a website and may encounter different configurations.
The Web ARChive ("WARC") format already accommodates requests, responses and associated metadata, so my proposal concerns the organisation and interpretation of evidence rather than the invention of another container. I would add an index that connects signal locations, capture conditions and instrument versions, allowing a later reader to distinguish a change in the website from a change in the rules used to interpret the same retained material.
If a corrected resolver reaches a different result, I would publish that as a new interpretation with its own date and explanation, leaving the original response and earlier result available for inspection. A later capture should likewise receive a separate reference, because replacing yesterday’s refusal with today’s success would remove the very difference the archive was meant to preserve.
What preservation can and cannot prove
A digest and an independently issued timestamp would help detect subsequent alteration and support the prior existence of retained material, but the time-stamping standard does not establish that a publisher served the bytes, authorised their contents or sent them to a particular recipient. I would distinguish the capture clock from any server-supplied time and independent timestamp, while recording their sources rather than treating several differently generated times as mutually confirming by default.
The same distinction matters to the litigation that prompted my interest, because ANI Media Pvt. Ltd. v. Open AI OpCo LLC ('ANI') records the ability to block crawlers within its assessment of interim relief [¶262]. A dated record may help investigate what a particular collector encountered, but two captures on either side of a disputed date do not prove that the same instruction remained continuously in force between them, and my later audit cannot settle the historical facts in ANI.
From the judgment · ANI Media v OpenAI
The paragraph NOTE 002 relies on
The note cites the Court’s record of ANI’s ability to block crawlers to show why dated evidence of what a collector encountered matters, and why a later audit cannot settle what happened earlier.
- ¶262 Balance of convenience and irreparable injury · ¶257–270
In the High Court of Delhi at New Delhi
ANI Media Pvt Ltd v Open AI OpCo LLC
The opting-out option“It is an admitted position that ANI has the ability to block its website vis-à-vis any third-party including Open AI. The opting-out option is available to ANI for blocking the third-party web crawlers from copying their data as well as from scraping their website for the search function/RAG. Despite having an option of opt-out, evidently ANI has not exercised the same. In fact, it has been stated on behalf of Open AI that it has internally blocked ANI’s website from its web crawlers or bots for the purposes of scraping of data. During the course of oral submissions Open AI has also submitted that Open AI itself has blocked ANI’s website from ‘ChatGPT search function/RAG’.”
Quoted from the signed copy · ellipses mark omissions · footnote markers omitted
- ¶262
Records
The Court’s observation concerns ANI’s ability to block at the time of the dispute. A dated archive could help investigate such claims, but captures on either side of a date do not prove that an instruction stayed in force between them.
For an Indian proceeding, section 63 and the Schedule to the Bharatiya Sakshya Adhiniyam, 2023 ("BSA") require attention to the production and certification of electronic evidence, which makes provenance and the people able to explain the relevant systems important alongside the stored objects. I would preserve that information when collecting the record, without describing a hash, a screenshot or an Internet Archive deposit as a substitute for the applicable admissibility requirements.
A public version could contain documented redactions of credentials or personal information while controlled originals remain preserved, with the relationship between the two versions explained rather than concealed. The archive would then support the separate inquiries into authority in ESSAY 001 and interpretation in ESSAY 003, giving each an inspectable history without claiming to decide either question merely by storing it.