Pipeline & receipts

A record you can inspect for every fetch.

The pipeline is deliberately boring: six stages, each of which writes something into the receipt. Nothing reaches your dataset without a trace of how it got there.

  1. 01

    Resolve the target

    The URL is normalized, the host is resolved, and the exact request — agent, headers, method — is recorded before anything is sent.

    source_urlfetch_method
  2. 02

    Fetch the page

    We retrieve the page with a declared user agent and record the HTTP status, content type, and timing of the response.

    http_statuscontent_typefetched_at
  3. 03

    Normalize and hash

    Substantive content is extracted from navigation and boilerplate, then hashed so the exact bytes you received can be re-verified later.

    content_sha256content_check
  4. 04

    Read permission signals

    Robots.txt and machine-readable AI-use declarations are read for the declared agent and the requested path, and the outcome is stored verbatim.

    robots_txtai_use_signal
  5. 05

    Write the receipt

    Provenance, checks, signals, and the billable units are assembled into one inspectable record tied to the page.

    receipt_idbilled_pages
  6. 06

    Deliver content and evidence

    Content and receipt return together, so storing the data and storing the audit trail are the same action.

    unit_cost_usd
Receipt specification

Every field, and why it exists.

receipt_id
Stable identifier for this record, referenced by the job and by billing.
source_url
The exact URL requested after normalization, including query string.
fetched_at
UTC timestamp of the response, so freshness is never guessed.
fetch_method
How the page was retrieved: plain HTML or rendered browser.
http_status
The response code received, kept even when the fetch was unsuccessful.
content_sha256
Hash of the normalized content for later verification and change detection.
content_check
Whether the page held substantive content, or was thin, blocked, or an error.
robots_txt
The outcome of applying robots.txt to this agent and path.
ai_use_signal
What machine-readable AI-use declarations stated at fetch time.
billed_pages / unit_cost_usd
The units charged and the rate applied, in the same record.
SAMPLE RECEIPT / JSON
{
  "receipt_id": "rcpt_8f92c1_001",
  "source_url": "https://developer.mozilla.org/en-US/docs/Web",
  "fetched_at": "2026-09-27T10:42:18Z",
  "fetch_method": "http_html",
  "http_status": 200,
  "content_type": "text/html",
  "content_sha256": "b92c4f…7a05d3",
  "content_check": "substantive",
  "robots_txt": "allowed",
  "ai_use_signal": "review_required",
  "billed_pages": 1,
  "unit_cost_usd": 0.00012
}
Permission signals are evidence for your review. They are not a legal determination, and they do not grant rights over the content.
Answers

Pipeline questions

Still unsure? Write to dani@dataforgaio.com or read the full FAQ.

See a receipt in context.

Open the dashboard to inspect your own jobs, or read the API documentation to see the response shape.