APEX FLOWWEB SCRAPER

APEX FLOW WEB SCRAPER / REFERENCE LIBRARY

Validating scraped data with schemas

Validation checks whether extracted values meet a declared structure. It can catch missing fields and wrong types; it cannot decide whether a source is truthful.

By Apex Flow Labs · Reviewed 2026-10-08

Define the contract first

Describe the fields a consumer needs before crawling. A catalog record might require title and source_url, while price and availability are optional. Decide how missing values, arrays and duplicate matches should be represented.

Validate at the boundary

Run validation before exporting records to a consumer. Keep validation errors with the source evidence. A partial record may still be useful, but it must not be confused with a complete one.

Current engine restrictions

Apex supports required recipe fields and a restricted JSON Schema validation path. Remote schema references and regular-expression patterns are rejected in the current schema interface. These restrictions bound work on untrusted inputs; they are not support for every JSON Schema feature.

Example review

If a field is declared numeric but contains "Contact us", the right result is a validation issue. Replacing it with 0 would create a false price. Preserve the original value, explain the failure and adjust the consumer contract if the value is legitimate.

← All reference articles · Current capabilities and limits →