80% of enterprise data exists in unstructured or semi-structured formats: HTML pages, PDF documents, images, email bodies, and free-text fields. Turning those formats into structured, queryable records is not a single technique — it is a selection problem. The right extraction approach depends on the source format, the volume, the variability of structure, and whether the content is static or dynamically rendered.
Structured data extraction is the process of producing machine-readable, schema-conformant records from sources that do not natively expose that structure. This guide covers the full technical spectrum: HTML DOM parsing with CSS selectors and XPath, JSON and XML processing, PDF extraction including OCR, handling JavaScript-rendered content, LLM-based extraction for unstructured text, and the production pipeline architecture that makes extraction reliable at scale. By the end, you will know which technique applies to which source type and how to build extraction pipelines that survive source format changes.
Structured Data Extraction: HTML DOM Parsing
DOM Tree and Parser Selection
An HTML document is represented as a tree (the Document Object Model). Each HTML element is a node; parent-child relationships reflect nesting. Parsing converts a raw HTML string into this tree structure, which can then be queried using CSS selectors or XPath expressions.
Parser options and tradeoffs:
| Parser | Language | Speed | Malformed HTML tolerance |
|---|---|---|---|
| lxml | Python | Fastest (C implementation) | Good |
| BeautifulSoup (with lxml) | Python | Fast | Excellent |
| html5lib | Python | Slow | Perfect (HTML5 spec) |
| Cheerio | JavaScript | Fast | Good |
| goquery | Go | Fast | Good |
For most production use cases, lxml directly (or BeautifulSoup backed by lxml) is the right choice. html5lib is appropriate when source HTML is severely malformed and must be parsed exactly as a browser would parse it.
Real-world HTML is frequently invalid — unclosed tags, improperly nested elements, missing attributes. Tolerant parsers handle these cases without crashing; strict parsers fail on input they encounter in every production scraping pipeline. Never choose a strict parser for web-sourced HTML.
CSS Selectors
CSS selectors are the primary extraction technique for most HTML parsing tasks. They are concise, readable, and match what web developers already know from front-end work.
Selector patterns for structured data extraction:
- Element selector —
div,span,table— matches by tag type - Class selector —
.product-price,.listing-title— matches by CSS class - ID selector —
#main-content— matches by unique identifier - Attribute selector —
[data-product-id],[href^="https://shop"]— matches by attribute presence or value - Descendant combinator —
div.product span.price— matches elements nested within other elements - Child combinator —
ul.items > li— matches direct children only (not deeper descendants)
Pseudo-selectors for positional targeting: :nth-child(n), :first-of-type, :last-child, :not(.excluded). These are essential for extracting values from tables (target every other row, skip the header row) and lists (skip the first promotional item).
XPath: Advanced Extraction Queries
XPath (XML Path Language) provides extraction capabilities that CSS selectors cannot match: upward traversal (select a parent or ancestor of a matched element), text content matching, and arithmetic predicates.
When XPath excels over CSS selectors:
- Upward traversal —
//span[text()='Price']/../following-sibling::span— select the sibling span that follows the span containing "Price". No CSS equivalent. - Text content matching —
//h2[contains(text(), 'Product Details')]— select headings that contain specific text. CSS:contains()is a non-standard extension not universally supported. - Conditional expressions —
//div[@class='product'][.//span[contains(text(), 'In Stock')]]— select product divs that contain an in-stock indicator. Complex predicates are natural in XPath.
XPath axes enable navigation in any direction through the DOM tree: child::, parent::, ancestor::, following-sibling::, preceding-sibling::, descendant::. These axes are essential for extraction from poorly structured HTML where the target value is not directly selectable but can be located relative to a reliably identifiable anchor element.
Practical guidance: Use CSS selectors by default. Switch to XPath when CSS selectors require you to locate a parent element, match on text content, or express conditional logic. Most production parsers use both.
JSON and XML Parsing
JSON: Schema Validation and JSONPath
REST APIs and modern web applications return JSON. Simple JSON structures can be accessed by key directly. Complex nested structures benefit from JSONPath (for Python, the jsonpath-ng library; for command-line, jq).
JSONPath notation: $.store.books[?(@.price < 20)] — extract all books from the store array where price is below 20. $..name — extract all name values at any depth.
For large JSON files (hundreds of MB to GB), streaming parsers avoid loading the entire document into memory. Python's ijson library parses JSON incrementally, yielding records as they are read rather than materializing the full document.
JSON Schema validation — use jsonschema (Python) or ajv (JavaScript) to validate extracted records against a defined schema before they enter the transformation pipeline. Catch missing required fields, wrong data types, and out-of-range values at extraction time, not in the reporting layer.
XML and Feed Parsing
XML requires namespace-aware parsing when documents use namespaces (common in SOAP responses, RSS feeds, and enterprise data exchange formats). XPath queries against namespaced XML must qualify element names with namespace prefixes; unqualified queries silently return zero results.
RSS and Atom feeds are the most common XML extraction use case in data pipelines. Python's feedparser library abstracts the differences between RSS 1.0, RSS 2.0, and Atom, providing a consistent interface regardless of feed version. Feed freshness monitoring (checking <lastBuildDate> or <updated> against the last-seen value) enables efficient incremental collection.
PDF Data Extraction
PDF is the format where structured data extraction most frequently requires format-specific tooling. Two distinct cases:
Text-Layer PDF Parsing
PDFs created from digital documents (not scanned) contain a text layer — the character sequences that map to visual positions on the page. Python libraries for text-layer extraction:
- pdfplumber — best for table extraction from structured PDFs. Uses spatial reasoning to reconstruct table structure from text positions.
- PyMuPDF (fitz) — fastest for general-purpose text extraction; good for financial reports, legal documents, and technical manuals.
- pdfminer.six — lower-level; provides access to character position data needed for custom layout reconstruction.
Table extraction from PDFs is inherently imprecise because PDF format does not encode table structure — tables are rendered as positioned text objects that happen to look like a grid. pdfplumber uses two approaches: lattice (detects tables by visible ruling lines) and stream (detects tables by whitespace patterns). Real-world PDFs often require both, and manual verification of table extraction results is necessary before relying on automated output.
OCR for Scanned Documents
Scanned PDFs (images of physical documents) contain no text layer. OCR (Optical Character Recognition) converts the visual pixel data into text. Key tools:
Tesseract — the dominant open-source OCR engine, supporting 100+ languages. The LSTM-based engine (default in Tesseract 4+) achieves 95%+ accuracy on clean, high-resolution scans of printed text. Accuracy drops significantly with low contrast, skew, or degraded originals.
Cloud OCR services — Google Cloud Vision, Amazon Textract, and Azure Form Recognizer offer higher accuracy than Tesseract, especially for handwritten text, complex layouts, and form field extraction. Amazon Textract includes specialized extractors for tables and forms that return structured output rather than raw text.
Pre-processing for OCR accuracy improvement:
- Deskew (correct rotation) — a 1–2 degree skew reduces accuracy by 10–15%
- Binarization (convert to black-and-white with adaptive thresholding)
- Denoising (remove scan artifacts using Gaussian or median filters)
- Resolution upsampling (minimum 300 DPI for Tesseract; 400+ DPI for challenging documents)
In document processing pipelines at Smart Maple, pre-processing has consistently improved OCR success rates on degraded historical documents by 25–35%, reducing the need for manual review.
JavaScript-Rendered Content
Modern web applications render content client-side after the initial HTML response. The server returns an application shell; JavaScript makes API calls and populates the DOM with data. Standard HTTP + HTML parsing extracts only the shell.
Approach 1: Headless Browser Rendering
Playwright (or Puppeteer) executes JavaScript in a real browser engine, waits for content to render, then extracts from the fully rendered DOM. The extraction logic (CSS selectors, XPath) is identical to static HTML parsing — the difference is the tool used to obtain the rendered HTML.
The cost is high: browser startup overhead, 50–200 MB RAM per session, and processing rates of 20–100 pages per minute versus 500–2,000 pages per minute with Scrapy. Reserve headless browsers for pages that genuinely require JavaScript execution.
Approach 2: API Discovery (Preferred)
When a JavaScript application renders content, it typically fetches that content from a JSON API. The browser developer tools Network panel reveals these API calls. Making requests directly to the discovered API endpoints returns clean, structured JSON — no HTML parsing required, faster, and more reliable than rendering.
This "API discovery" technique is the preferred approach for many production scenarios because: the JSON structure is more stable than HTML markup, the data is already structured (no parsing required), requests are faster (no browser overhead), and the pattern is more resilient to front-end redesigns that change HTML class names but preserve API contracts.
Infinite Scroll and Pagination
Pagination parameters (?page=2, ?offset=20, ?cursor=abc123) can usually be enumerated directly without browser simulation. Infinite scroll triggered by scroll position can be handled in Playwright by simulating window.scrollTo(0, document.body.scrollHeight) repeatedly. For both cases, API discovery is preferable — find the paginated API endpoint and iterate through pagination parameters directly.
LLM-Based Extraction for Unstructured Text
Large language models handle extraction tasks where rule-based approaches require prohibitive engineering effort: job description parsing, legal clause extraction, unstructured product specification parsing, and any domain where the target information appears in free text rather than consistent HTML structure.
When LLM extraction is appropriate:
- Source text is free-form prose (not structured HTML with consistent patterns)
- The same information appears in many different phrasings and formats
- Extraction failure cost is low enough to accept occasional LLM errors
- Volume is low enough that per-call API cost is acceptable
Prompt engineering for structured output:
Defining a JSON schema in the prompt and using function calling (OpenAI) or structured output mode (Anthropic) constrains LLM output to valid JSON. Provide 2–3 examples (few-shot) of the target extraction. Define fallback values for fields the LLM cannot find. The result is a reliable extraction pipeline for inputs that would require hundreds of regex rules to handle with rule-based logic.
Hybrid approach: Use CSS selectors and XPath for structured HTML sections (product title, price, SKU — consistent across all pages). Use LLM extraction for free-text sections (product description, shipping policy, compatibility notes — variable prose). The hybrid approach optimizes cost (LLM calls only for genuinely unstructured sections) and reliability (deterministic extraction for structured sections).
Production Pipeline Design
Validation at Each Layer
Structured data extraction pipelines should validate at every stage transition:
- After fetch — HTTP status code, content type, minimum response size (empty page = missing content)
- After parse — required field presence, field type checks, value range validation
- After transform — schema conformance, business rule validation (price > 0, date within plausible range)
Use Pydantic (Python) for schema validation — define a model class with field types and constraints, pass extracted data through it, and handle validation errors explicitly. Validation failures should route to a dead-letter store for manual review, not silently propagate downstream.
Handling Parser Breakage
Source site structure changes are inevitable. A CSS class rename, a layout restructuring, or a site migration breaks selector-based parsers without warning. Detection strategy: monitor field completeness rates continuously. A sudden drop in completeness for a specific field — from 98% to 30% — is a strong signal that the selector targeting that field is broken.
Mitigation strategy: maintain fallback selectors (primary CSS selector, fallback XPath expression) for critical fields. If the primary selector fails, try the fallback before marking the record as failed.
Monitoring Metrics
Track per-source extraction metrics: success rate (target >95%), records extracted vs. prior-day baseline (alert on >25% deviation), field completeness for critical fields (alert below 98%), and parse error rate (alert above 2%). Structured logging with source identifier, URL, extraction timestamp, and field counts enables automated metric aggregation.
Conclusion
Structured data extraction is not one technique — it is a library of techniques applied based on source format and complexity. CSS selectors and XPath for HTML, JSONPath for nested JSON, pdfplumber and OCR for documents, headless browsers for JavaScript-rendered content, and LLM extraction for free-text fields. Production pipelines combine these techniques, validate output at every stage, monitor for parser breakage, and treat extraction reliability as an operational concern, not a one-time implementation.
Related Articles
MLOps Guide: Taking Machine Learning Models to Production [2026]
87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow
Read MoreLLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]
General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not
Read MoreComputer Vision Applications: Object Detection, OCR, and Industrial AI [2026]
Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl
Read More
