APEX FLOWWEB SCRAPER

APEX FLOW WEB SCRAPER / REFERENCE LIBRARY

What is source provenance in scraped data?

Source provenance records where a value came from and how it was extracted. It makes a result reviewable without pretending that collection proves the source is correct.

By Apex Flow Labs · Reviewed 2026-10-08

Keep an extraction receipt

For each field, retain the page URL, selector, match count and extracted value. For each page, retain the collection timestamp and content hash. These details help a reviewer reproduce the rule and identify stale or ambiguous values.

Evidence is not verification

A selector matching a price proves only what that page displayed during collection. It does not prove that the product is in stock, the vendor is reliable or the price includes tax. Downstream decisions need their own checks.

What Apex stores today

The engine records field evidence, source URLs, page-level metadata and content hashes. Saved jobs can be inspected and compared. This is not an immutable full-page snapshot archive or a cryptographic timestamp service.

Useful review workflow

Choose ten records, open their sources, and check required fields. Log mismatches by type: selector error, stale source, ambiguous units, duplicate content or blocked page. Fix the cause instead of silently deleting failures.

← All reference articles · Current capabilities and limits →