# A Wayback Machine for Content-Use Signals

A proposal for preserving the request, the response and the interpretation without replacing one with another.

By Manraj Singh Chandpuri

LEX.TXT/NOTE/002

Published 2026-09-24 · Updated 2026-09-25

[Canonical article](https://manrajchandpuri.com/lex/notes/before-the-signal-changes)

## Abstract

I propose a dated archive of content-use signals that preserves the conditions of each observation alongside the raw response and successive interpretations, because a later successful request cannot reconstruct an earlier refusal. Using my aninews.in audit, I show how a record can distinguish the automated 403, the separately retrieved file and the report’s indeterminate classification without claiming to establish what another collector encountered. The proposal builds on existing archival formats and treats hashes, timestamps and admissibility requirements as complementary safeguards, while refusing to equate the preservation of bytes with proof of publication, authority or receipt.

## Permissions

I grant a worldwide, royalty-free, nonexclusive permission to crawl, index, retrieve, embed, analyse, summarise and quote my eligible original prose, and to use it for commercial and noncommercial model training, fine-tuning and evaluation, including making and retaining the copies reasonably necessary for those purposes and deploying the resulting models commercially. Retained copies must preserve the supplied author, canonical source and rights metadata, although I do not make this permission depend on a model naming me in every future answer.

My original prose in published LEX.TXT articles, including articles discussing RightSignal, as identified in the publication manifest. Embedded or linked resources are separate works and do not inherit this permission.

[Applicable permission](https://manrajchandpuri.com/lex/rights#original-prose) · [Inspect these signals](https://manrajchandpuri.com/lex/machine-signals/note-002)

Policy version 0.5

## Article

A dispute about what a website permitted can outlive the configuration that supplied the answer, which is why I want a history of content-use signals that remains useful after the live page changes. My starting example is the aninews.in discrepancy examined in [NOTE 001](https://manrajchandpuri.com/lex/notes/is-a-403-a-reservation), where the RightSignal 1.1 audit dated 25 July 2026 recorded an automated Hypertext Transfer Protocol ("HTTP") 403 response while I could retrieve the instruction file manually, leaving two encounters that a later successful fetch should not collapse into one.

The Internet Archive’s Wayback Machine offers a useful model for looking backwards, and its [capture calendar already distinguishes successful responses, redirects and errors](https://help.archive.org/help/using-the-wayback-machine/), although saving a page does not capture an entire site or arrange continuing preservation. I would build the proposed history around the instruction locations and request conditions relevant to a particular observation, allowing a reader to inspect both what was collected and what the collection failed to reach.

## A specimen entry for the ANI observation

The table below illustrates how I would organise the retained research, rather than claiming that a new archive has already captured all the fields I propose. The [empirical report and underlying export](https://archive.org/download/empirical-report-permission-at-the-point-of-extraction/Empirical%20Report_Permission%20at%20the%20Point%20of%20Extraction.pdf#page=472) remain the sources for the automated observation and classification, while the [figures reproduced in NOTE 001](https://manrajchandpuri.com/lex/notes/is-a-403-a-reservation#ani-evidence-diagram) keep the manual retrieval as a separate item with its own provenance limits.

| Associated record | Material I would retain | Limit that must remain visible |
| --- | --- | --- |
| Automated encounter | The request for `aninews.in/robots.txt`, its recorded date and 403 result, with the export location. | Unrecorded client or network details cannot be reconstructed by assumption. |
| Manual encounter | The retained figure of the retrieved instruction file, associated with the author’s account. | It cannot be substituted for the automated response or represented as a simultaneous controlled comparison. |
| Instrument interpretation | The version 1.1 silence-based `ALLOW` outputs for the four audited uses. | The output identifies a resolver result without establishing permission. |
| Research interpretation | The report’s indeterminate classification, linked to the observations it assesses. | The classification does not establish what OpenAI received during earlier collection. |

The report’s distinction between eleven unreadable probes and five underlying requests also belongs in this record, because several checks that depend on the same failed response do not become independent observations simply because they occupy separate rows ([section 4, table 5](https://archive.org/download/empirical-report-permission-at-the-point-of-extraction/Empirical%20Report_Permission%20at%20the%20Point%20of%20Extraction.pdf#page=4)). I would connect each interpretation to its underlying request so that a reader can assess the actual evidential weight rather than count repeated dependencies as corroboration.

## What a future capture should preserve

For a newly designed capture, I would retain the requested address, method, relevant request headers and client configuration, start and completion times in Coordinated Universal Time ("UTC"), redirects, response status, response headers and returned bytes, while marking a timeout or missing body explicitly. The instruction file, linked policy and article-level response should remain separate associated records within a disclosed collection window, because successive requests are not a perfectly simultaneous account of a website and may encounter different configurations.

The [Web ARChive ("WARC") format](https://iipc.github.io/warc-specifications/specifications/warc-format/warc-1.1/) already accommodates requests, responses and associated metadata, so my proposal concerns the organisation and interpretation of evidence rather than the invention of another container. I would add an index that connects signal locations, capture conditions and instrument versions, allowing a later reader to distinguish a change in the website from a change in the rules used to interpret the same retained material.

If a corrected resolver reaches a different result, I would publish that as a new interpretation with its own date and explanation, leaving the original response and earlier result available for inspection. A later capture should likewise receive a separate reference, because replacing yesterday’s refusal with today’s success would remove the very difference the archive was meant to preserve.

## What preservation can and cannot prove

A digest and an independently issued timestamp would help detect subsequent alteration and support the prior existence of retained material, but the [time-stamping standard](https://www.rfc-editor.org/rfc/rfc3161.html#section-2) does not establish that a publisher served the bytes, authorised their contents or sent them to a particular recipient. I would distinguish the capture clock from any server-supplied time and independent timestamp, while recording their sources rather than treating several differently generated times as mutually confirming by default.

The same distinction matters to the litigation that prompted my interest, because [*ANI Media Pvt. Ltd. v. Open AI OpCo LLC*](https://indiankanoon.org/doc/93327052/) ('*ANI*') records the ability to block crawlers within its assessment of interim relief [[¶262]](https://indiankanoon.org/doc/93327052/#blockquote_228). A dated record may help investigate what a particular collector encountered, but two captures on either side of a disputed date do not prove that the same instruction remained continuously in force between them, and my later audit cannot settle the historical facts in *ANI*.

For an Indian proceeding, section 63 and the Schedule to the [Bharatiya Sakshya Adhiniyam, 2023](https://indiacode.gov.in/act/9e461245-2c13-4171-9d9d-267425c01e1d) ("BSA") require attention to the production and certification of electronic evidence, which makes provenance and the people able to explain the relevant systems important alongside the stored objects. I would preserve that information when collecting the record, without describing a hash, a screenshot or an Internet Archive deposit as a substitute for the applicable admissibility requirements.

A public version could contain documented redactions of credentials or personal information while controlled originals remain preserved, with the relationship between the two versions explained rather than concealed. The archive would then support the separate inquiries into authority in [ESSAY 001](https://manrajchandpuri.com/lex/essays/who-speaks-for-the-server) and interpretation in [ESSAY 003](https://manrajchandpuri.com/lex/essays/the-human-readable-problem), giving each an inspectable history without claiming to decide either question merely by storing it.
