APEX FLOW WEB SCRAPER / REFERENCE LIBRARY
Validating scraped data with schemas
Validation checks whether extracted values meet a declared structure. It can catch missing fields and wrong types; it cannot decide whether a source is truthful.
Define the contract first
Describe the fields a consumer needs before crawling. A catalog record might require title and source_url, while price and availability are optional. Decide how missing values, arrays and duplicate matches should be represented.
Validate at the boundary
Run validation before exporting records to a consumer. Keep validation errors with the source evidence. A partial record may still be useful, but it must not be confused with a complete one.
Current engine restrictions
Apex supports required recipe fields and a restricted JSON Schema validation path. Remote schema references and regular-expression patterns are rejected in the current schema interface. These restrictions bound work on untrusted inputs; they are not support for every JSON Schema feature.
Example review
If a field is declared numeric but contains "Contact us", the right result is a validation issue. Replacing it with 0 would create a false price. Preserve the original value, explain the failure and adjust the consumer contract if the value is legitimate.
← All reference articles · Current capabilities and limits →