October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
AI policy

Parsing TDMRep and AI.txt: Purpose-Based Scraping Controls

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Short answer: TDMRep and ai.txt are policy signals that tell compliant crawlers how a site wants its content used. TDMRep is a W3C Community Group protocol focused on text-and-data-mining reservations and licensing. The proposed ai.txt format covers a wider set of AI activities, including training, scraping, indexing, caching, retrieval, agent overrides, attribution, disclosure and audits. Neither file technically blocks a crawler. Use authentication, access controls or network blocking when prevention is required.

What TDMRep is—and what it is not

TDMRep (Text and Data Mining Reservation Protocol) lets a rightsholder declare reservations and licensing policies for lawfully accessible web content. It is a W3C Community Group specification, not a W3C Recommendation or completed web standard. The vocabulary defines reservation as 1 (rights reserved) or 0 (rights not reserved), policy as a URL to a rightsholder policy, and policy values such as mine, research and non-research. The vocabulary page identifies revision 1.2 dated 2024-02-23.

A TDM agent must look for the origin declaration before it starts scraping. The site-wide location is /.well-known/tdmrep.json. TDMRep can also be carried in HTTP response headers, HTML metadata, EPUB metadata and PDF XMP metadata.

Minimum JSON model

The origin file is an array of rule objects. location and tdm-reservation are mandatory; tdm-policy is optional. A minimal site-wide reservation is:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
[
  {
    "location": "/",
    "tdm-reservation": 1
  }
]

A policy URL can be added when you publish terms for permitted mining:

[
  {
    "location": "/",
    "tdm-reservation": 1,
    "tdm-policy": "https://example.com/tdm-policy"
  },
  {
    "location": "/public-research/",
    "tdm-reservation": 0,
    "tdm-policy": "https://example.com/research-licence"
  }
]

How to parse tdmrep.json

1. Fetch the origin file first

Request the exact well-known path on the content origin, not a copied file on a documentation host. Check the HTTP status and JSON content type, then parse the array. Treat malformed JSON or invalid rule objects as an implementation error rather than silently granting permission.

2. Match the requested URL

Each rule’s location identifies the path to which it applies. For a requested path, select the most specific matching location. If no location matches, the result is unset; an unmatched URL is not implicitly allowed or denied.

import json
from urllib.parse import urlparse
import requests

def load_tdmrep(origin):
    endpoint = origin.rstrip('/') + '/.well-known/tdmrep.json'
    response = requests.get(endpoint, timeout=20)
    response.raise_for_status()
    rules = response.json()
    if not isinstance(rules, list):
        raise ValueError('TDMRep must be a JSON array')
    for rule in rules:
        if not isinstance(rule, dict):
            raise ValueError('Each TDMRep rule must be an object')
        if 'location' not in rule or 'tdm-reservation' not in rule:
            raise ValueError('location and tdm-reservation are required')
        if rule['tdm-reservation'] not in (0, 1):
            raise ValueError('tdm-reservation must be 0 or 1')
    return rules

def decision(rules, target_url):
    path = urlparse(target_url).path or '/'
    matches = [r for r in rules if path.startswith(r['location'])]
    if not matches:
        return {'state': 'unset'}
    rule = max(matches, key=lambda r: len(r['location']))
    return {
        'state': 'reserved' if rule['tdm-reservation'] == 1 else 'not_reserved',
        'location': rule['location'],
        'policy': rule.get('tdm-policy')
    }

rules = load_tdmrep('https://example.com')
print(decision(rules, 'https://example.com/articles/42'))

This example uses longest-prefix matching, which implements the protocol’s “most specific match” rule for ordinary path locations. Production agents should also validate location normalization, redirects and the origin they trust, and should retain the raw response for auditability.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

TDMRep precedence across files, headers and metadata

TDMRep values can be declared in several places. Processing is ordered, and later declarations supersede earlier values:

  1. Read the origin /.well-known/tdmrep.json file.
  2. Apply TDM-related HTTP response headers.
  3. Apply HTML metadata in the fetched document.
  4. Apply EPUB or PDF metadata, including the PDF XMP properties tdm:reservation and optional tdm:policy.

A later declaration replaces a value that was already set. Absence does not clear the current state: if a header omits tdm-policy, that omission does not erase a policy obtained from the origin file. Implementers should therefore carry state forward property by property instead of replacing an entire object with every layer.

TDM policies use an ODRL-based JSON-LD profile. Depending on the policy document, they can describe mining permissions, research versus non-research conditions, contact obligations and financial compensation. The protocol transports the policy reference; interpreting the linked policy remains a separate step.

What ai.txt proposes

ai.txt is an IETF Internet-Draft, not an adopted Internet standard. Its syntax and semantics may change, so publish a draft or retrieval version in your operational documentation. In production, the draft specifies https://example.com/.well-known/ai.txt with Content-Type: text/plain; charset=utf-8.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The format is block-based and inspired by robots.txt. Each line is a key: value pair; # begins a comment; indented lines belong to the preceding block. A minimal example is:

Spec-Version: 0.1
Site-Name: Example Publishing
Site-URL: https://example.com
Training: deny

Site-wide controls

The draft defines Training, Scraping, Indexing and Caching. Their values are allow or deny. Training may also be conditional, which activates path rules:

Spec-Version: 0.1
Site-Name: Example Publishing
Site-URL: https://example.com
Training: conditional
Training-Allow: /licensed/**
Training-Deny: /private/**
Scraping: deny
Indexing: allow
Caching: allow

Training-Allow and Training-Deny accept glob patterns. When patterns overlap, the more specific pattern takes precedence. Other fields include Training-License (an SPDX identifier), Training-Fee (a licensing or pricing URL), agent blocks with per-agent overrides and advisory rate limits, plus Attribution, AI-Disclosure, Audit and Audit-Format.

A small parser for block syntax

from collections import defaultdict

def parse_ai_txt(text):
    blocks = []
    current = None
    for number, raw in enumerate(text.splitlines(), 1):
        if not raw.strip() or raw.lstrip().startswith('#'):
            continue
        indented = raw[:1].isspace()
        line = raw.strip()
        if ':' not in line:
            raise ValueError(f'Line {number}: expected key: value')
        key, value = [part.strip() for part in line.split(':', 1)]
        if indented:
            if current is None:
                raise ValueError(f'Line {number}: indented field has no block')
            current.setdefault('fields', []).append((key, value))
        else:
            current = {'type': key, 'value': value, 'fields': []}
            blocks.append(current)
    return blocks

with open('ai.txt', encoding='utf-8') as f:
    for block in parse_ai_txt(f.read()):
        print(block)

This parser preserves repeated and indented fields so an agent can apply the draft’s block semantics without losing information. It intentionally does not assume that an Internet-Draft field will remain unchanged in a future revision.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

TDMRep versus ai.txt

Axis TDMRep ai.txt
Primary purpose Text-and-data-mining reservations and licensing Broader AI-use policy, including training, retrieval indexing, caching and disclosure
Declaration surface Well-known JSON, HTTP headers, HTML, EPUB and PDF metadata Well-known plain-text file
Granularity Path/location and individual assets through metadata surfaces Site-wide fields, path globs and agent-specific blocks
Precedence Origin file, then headers, HTML, EPUB/PDF; later values supersede earlier ones Draft block and pattern rules; semantics remain subject to draft changes
Licensing expression ODRL-based policy references can express permissions, conditions, contacts and compensation License identifier and fee URL fields, plus attribution and disclosure concepts
Enforcement A declaration is a signal for agents that choose to comply; it does not authenticate or block requests
Status W3C Community Group specification, not a W3C Recommendation IETF Internet-Draft, not an adopted Internet standard

The scope distinction follows the fields and processing models defined by each document. Questions about inference, retrieval-augmented generation, search and discovery—including whether AI-boosted search is text and data mining—remain active standardization issues. W3C-versus-ISO work and interaction with IETF AIPREF are also still being discussed.

Can these files stop AI crawlers?

No. TDMRep, ai.txt and robots.txt communicate an intended policy; they do not impose an access-control decision at the network layer. The International Press Telecommunications Council describes robots.txt as a recommendation that does not guarantee compliance by AI providers in any jurisdiction. Its guidance points to HTTP-level blocking when prevention is required.

Use declarations and controls together

  • Publish a clear declaration that matches your legal and licensing position.
  • Use HTTP authentication, signed URLs, an application firewall, account-level authorization or network blocking for technical prevention.
  • Keep TDMRep, ai.txt, robots.txt and server controls consistent.
  • Log requests and monitor crawler user-agent changes; a user-agent string is not proof of identity.

IPTC recommends a site-wide TDMRep file with location: "/" and tdm-reservation: 1 when the objective is to reserve data-mining rights across a site. Its guidance also notes that detailed tdm-policy was not, to its knowledge, implemented by crawler bots, making the reservation value the practical current signal.

Deployment checklist for site owners

  1. Decide which activities you are addressing: text-and-data mining, model training, search indexing, retrieval, caching or all of them.
  2. Publish /.well-known/tdmrep.json with valid JSON, mandatory fields and the narrowest path exceptions you can maintain.
  3. If you use ai.txt, serve the draft file as UTF-8 plain text and record the draft version or retrieval date internally.
  4. Test path matching, overlapping globs, redirects and missing declarations with an automated checker.
  5. Verify that headers and embedded metadata do not accidentally override the origin policy.
  6. Apply technical controls to protected paths and test that unauthenticated requests receive the intended status code.
  7. Review policies when licensing terms, content classifications or crawler behavior changes.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common parser and deployment failures

The file returns 404

Check that the request uses the exact origin and path, including /.well-known/. A declaration on www does not automatically describe a different host or CDN origin.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Rules appear to conflict

For TDMRep, inspect the longest matching location and then apply the precedence layers. For ai.txt, compare overlapping glob patterns and choose the more specific match according to the draft.

A missing field seems to erase a policy

Do not treat omission as a reset in TDMRep. Carry the previous property forward unless a later declaration explicitly supplies a replacement.

Crawlers still fetch reserved pages

That is expected when a crawler ignores advisory policy. Enforce the requirement with authentication, authorization or network controls, then retain the declaration as a signal for compliant agents.

A PDF or HTML result disagrees with the site file

Check the defined precedence order. HTML and EPUB/PDF metadata are later layers than the origin file, so an embedded value can supersede it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Visual smoke tests without building a browser harness

If you maintain a human-facing policy or documentation page, a rendered screenshot can be a quick check that the page exposes the intended links and version labels. ScreenshotNeo is a website screenshot API and MCP server for developers; it is not a replacement for fetching and parsing the machine-readable files.

Or skip the browser setup

Use one request to capture a rendered page:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://freedom251.com -o shot.webp

See the ScreenshotNeo documentation for options. Before capture, it can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for the free ScreenshotNeo plan.

Frequently Asked Questions

Is TDMRep a W3C standard?

No. It is a W3C Community Group specification. Do not describe it as a W3C Recommendation.

Is ai.txt the same as robots.txt?

No. The proposed format is broader and includes AI-specific training, licensing, agent, attribution, disclosure and audit fields. It remains an IETF Internet-Draft.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should an unmatched TDMRep URL return?

Treat it as unset. TDMRep does not define an unmatched path as automatically allowed or reserved.

What if I need an actual block rather than a policy signal?

Use authentication, authorization, signed access or network-level controls. Keep the policy files aligned with those controls.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.