Own the execution
Run collection on your infrastructure. No mandatory scraping vendor account or external model call in the core engine.
01 AN APEX FLOW LABS PRODUCT
Open it.
Turn public websites into structured data, with the source attached and the results in your hands.
Open engine · Your infrastructure · No scraping subscription
02 REVEAL THE SIGNAL
Pages become fields. Fields become answers. Collect the useful details from public articles, catalogs, feeds, and documentation.
HTML · Markdown · JSON · RSS · Text PDFs
03 EACH FACET, A FUNCTION
Saved queues, live progress, cancellation, and recovery.
Reusable field recipes with required-value validation.
Source URLs, selectors, match counts, and content hashes.
Explore the interactive mineral core
04 YOUR WORK. YOUR EVIDENCE.
Compare saved runs. Search your collected data. Export JSON or CSV. Connect the same private engine to your assistant workflow.
title → h3 → 1 match
source → attached
missing fields → visible05 BUILD WITH WHAT YOU DISCOVER
Apex Flow Web Scraper. An owned engine for turning public web content into something useful.
02 / SEE THE SHAPE OF YOUR DATA
Try a sample extraction in your browser. This demonstration uses the example page below, without contacting a website.
Sample content · No account required
FIELD COLLECTION
24.00 USD
A weather-ready notebook for ideas worth keeping.
View specifications ↗Both fields are read from this page's built-in example. No remote request is made.
03 / THE APEX APPROACH
Our focus is practical control: a system you can understand, inspect and use across your workflow.
Run collection on your infrastructure. No mandatory scraping vendor account or external model call in the core engine.
Page counts, progress events, content hashes and explicit errors make each run inspectable.
Use saved jobs and recipes through the MCP interface once your assistant is connected. Search the data you've already collected.
Compare two saved runs to find changed fields and text. Keep the earlier evidence alongside the new result.
Download structured JSON or spreadsheet-ready CSV. Your collected results remain portable.
Site access rules, budgets and validation are part of the engine. Blocked access is reported plainly.
THE KNOWLEDGE LAYER
Practical reference guides, ten implemented strengths, twenty alternatives and transparent pricing.
05 / STRAIGHT ANSWERS
The self-hosted 0.1.0 engine is available on GitHub. Public hosted accounts and billing are not available. The documentation describes the current engine and its limits.
No hosted scraping service is required by the core engine. Your own computer, hosting, storage and network still have operating costs.
No scraper has universal access. This release supports public HTTP and HTTPS content and isolated page rendering. It reports access challenges and respects robots.txt. Authenticated account scraping, OCR and a global search index are not included.
In the private engine workspace, separate from this public website. Results are available through the workspace and authorized assistant connections.
We use repeatable fixtures for extraction, rendering, cancellation, recovery and resource limits. We do not claim a market-leading success rate before a representative comparative benchmark has been run.