Extract data from any URL
Send a URL or upload a file, get back clean markdown, metadata, JSON-LD schemas, and AI-powered extractions. Handles JavaScript-rendered pages, PDFs, and office documents out of the box.
# Response — extracted content
"markdown": "# Example Domain\nThis domain is for use in...",
"metadata": { "title": "Example Domain", "language": "en" },
"credits_used": 1
}
From URL to structured data
A single API call handles fetching, parsing, and extraction.
Fetch
The page is fetched via HTTP or rendered with a headless browser for JS-heavy sites.
Parse
HTML is cleaned, content extracted, and metadata/schemas parsed automatically.
Return
Clean markdown, metadata, schemas, and optional AI enrichments are returned instantly.
Multiple output formats
Markdown
Clean, readable markdown with preserved heading structure, links, and formatting.
HTML
Raw or cleaned HTML. Great for custom parsing pipelines or archival.
Metadata
Title, description, language, Open Graph tags, Twitter cards, and more.
JSON-LD / Schema
Structured data extracted from JSON-LD, microdata, and RDFa embedded in the page.
PDFs & documents
PDF, Word, Excel, PowerPoint, OpenDocument, EPUB and CSV to markdown, tables included. Optional OCR for scanned pages.
AI enrichment
LLM-generated summaries, and structured extraction with a custom prompt or schema.