Extracting main content from HTML: where our approach differs from Trafilatura and Readability
Where our extractor diverges from Trafilatura and Readability, and the evaluation harness that tells us whether a change is an improvement.
Every content extractor fails on some pages. That is the honest starting point. Readability and Trafilatura are the two implementations most people reach for, and both are good. Danubia uses a third approach, and this post is about what that approach actually is and how we check whether it is an improvement or just a different set of mistakes.
The shape of the thing
Our extractor runs three passes over a parsed DOM:
- Mark the regions that look like main content, using semantic tags and class, id and role hints.
- Prune the regions that look like navigation, footers, related posts, cookie banners and other chrome, using link density.
- Collect the survivors into a structured tree, then choose the best subtree with a ladder of size and ratio thresholds.
There is no per-paragraph scoring pass and no candidate ranking. That is the main structural difference from a Readability-style implementation, which builds a score for every candidate node and keeps the winner. We apply rules in sequence, then choose between five pruning configurations based on how much text survives.
The difference matters most when something goes wrong. A scoring extractor fails because the winning score landed on the wrong node, and the fix is to retune weights. A rule-based extractor fails because a rule matched something it should not have, and the fix is to make that rule narrower. The second kind of failure is easier to reason about, and it is why our discard rules carry comments naming the specific pages that caused them.
Two vocabularies for chrome
The single most useful thing we learned is that the word "sidebar" means different things in different parts of a page.
Outside the main content, a generous vocabulary is right. A div with sidebar in its class is almost
always chrome, and so is anything matching footer, related, newsletter, breadcrumb, cookie,
consent, paywall or a role of nav.
Inside a region already marked as main content, that same vocabulary is dangerous. Real articles
contain elements whose classes say widget, button, border, tags or banner, and those
elements are part of the article. Two examples from pages that broke us:
- A Hackaday article that rendered its post listing inside a div classed
widget widget-recent-postswith the idsidebar-mobile-1, positioned inside<main>. - A Monnaie de Paris page whose content sits inside elements classed
item-border,with-label-linkandbutton-container, again inside<main>.
So the discard rules come in two versions. The broad set applies outside main content. A much narrower set applies inside it, where only unambiguous chrome words are allowed to discard.
We also removed a rule rather than keep it. A bare /search/ pattern looked useful until it matched
Wikipedia's body class skin-vector-search-vue and started discarding article text. The comment in
the source records why it is no longer there.
Link density, with the actual numbers
Link density is the ratio of link text to total text in a subtree. It is the classic signal for "this block is a menu". The interesting part is the threshold logic, because a naive ratio deletes content cards and product teasers.
Our rule works like this:
- A block is eligible for removal by density only if it holds a minimum amount of text. That floor is 60 bytes for a paragraph with no same-tag sibling after it, 30 bytes for a paragraph that has one, 300 bytes for other elements without a same-tag sibling, and 100 bytes otherwise.
- Above the floor, the block is removed when at least 80 percent of its text is link text.
- Below the floor, a different rule applies. The block is removed only when it contains at least two links and at least 80 percent of those links are short, where short means the link's own text is under 50 bytes.
That last clause is what saves content cards. A single short link with a little text beside it is a teaser, and a menu is a list of them.
Two further details do real work. Any block containing a heading is exempt from link-density removal, because a heading is strong evidence that the block is content. And the same-tag sibling check relaxes the floor, on the theory that repeated siblings with the same tag are usually lists, feeds or schedules.
The fallback ladder
After marking and pruning, we still have to decide what to emit. We do that by collecting the tree five times with different pruning configurations, from the most aggressive to the most permissive, and accepting the first result that clears both an absolute byte floor and a relative share of the document's text.
| Pass | Main content only | Remove search discards | Remove link-density discards | Accepted when |
|---|---|---|---|---|
| 1 | yes | yes | yes | over 200 bytes and at least 15% of text |
| 2 | yes | yes | no | over 200 bytes and at least 15% of text |
| 3 | yes | no | no | over 200 bytes and at least 15% of text |
| 4 | no | yes | yes | over 150 bytes and at least 15% of text |
| 5 | no | yes | no | over 50 bytes and at least 10% of text |
If none of them clears the bar, we fall back to the whole document with search discards removed, which still strips the worst junk.
This is affordable because the marking pass is non-destructive. Marks live in a side table keyed by node, so we can re-collect with a different filter without re-parsing or re-marking. The expensive work happens once.
Small decisions that mattered more than expected
Several of the highest-impact changes were about units and unwrapping rather than clever heuristics.
Text length is counted in UTF-8 bytes. Every threshold in the extractor is a byte count. The heuristics were originally written against crawler code where string length means byte length, and the unit turns out to be right for a multilingual corpus anyway: a Japanese paragraph and an English paragraph with the same character count are very different amounts of text.
An anchor with no href is unwrapped. A link with no target, or with a fragment-only target, is not navigation. Leaving it in the tree inflates link density and can get a content block discarded. Unwrapping keeps the text and removes the false signal. Skip links are dropped entirely for the same reason.
A figure that wraps code is unwrapped. We delete figures, iframes, embeds, video and
canvas as non-content, and that rule broke documentation sites where <figure> is how
syntax-highlighted code gets wrapped. If a figure's subtree contains a <pre> or a <code>, we
unwrap it and keep the code.
Templates are traversed. Declarative shadow DOM puts real markup inside
<template>. If the cleaner skips template content, styles leak into the extraction and text can go
missing. We descend into it so it gets cleaned like anything else.
Trailing links and headings get trimmed, with a guard. A link or a heading at the very end of a document is usually a footer or a related-links block, so we trim trailing nodes that contain nothing else. The guard is that we skip the trim entirely when the trailing material is more than half the total text, because otherwise we would delete the content of exactly the pages that are made of links, such as homepages, forums and listings.
Tables and code blocks
These two deserve their own note, because they are where extraction output most often becomes useless.
Tables get a structure-preserving path: header cells become a header row and the rest become body rows. The escape hatch is that if any cell contains block content, such as a paragraph, a list or a nested table, we stop treating it as a table and linearize the cells instead. A page-layout table that was never data should not be forced into pipe-separated cells.
Code blocks get the opposite treatment. A syntax highlighter wraps every token in a span, so reading
only the direct text children of a <pre> returns fragments. We walk all descendants and join tokens
with spaces inside a line, starting a new line only at an actual line break. Without that, every
highlighted code block comes out as a single run-on line.
What comes out the other end
The extractor does not emit text directly. It builds a tree of typed nodes, with headings, paragraphs, lists, links, images, code blocks, quotes and tables, and every node carries flags recording whether it descended from marked main content and whether it was discarded by the search or link-density rules. Markdown, cleaned HTML and plain text are three separate serializers over that same tree.
Keeping the structure means the markdown can be correct in ways a text dump cannot. Link targets are resolved to absolute URLs while the original authored value is kept alongside. Tables synthesize a header row when the page did not have one. Lists indent their nested blocks. Layout tables linearize.
If you are building a retrieval pipeline, the companion post on turning HTML into clean markdown for RAG covers what to do with this output.
How we keep it honest
An extractor that only its author tests is a story rather than a measurement. We built a harness for exactly this reason.
The corpus is 799 pages, committed to the repository as a JSON list, spread across twelve labels (articles, documentation, news, homepages, forums, products, listings, glossaries, recipes and more) and twelve languages. Each entry records what a good extraction of that page should contain, which is what the scoring compares against.
Two scoring systems run over the same pages.
A deterministic metric suite. It normalizes the reference and the extractor output into token streams, then computes unigram and bigram precision, recall and F1, plus structure recall for headings, link targets and code blocks. It also raises mechanical flags that need no judgment call: output truncated, code fences left empty, headings missing, reference stale because the page changed. The headline number is the mean of the unigram and bigram F1 scores.
An LLM judge. It scores recall, precision, boilerplate and formatting fidelity on a 1 to 5 scale. Three things keep it from being a vibe check. The extractor labels are shuffled per page and the judge never sees which output came from which tool. Temperature is zero and the response is schema-validated. The judge runs on a different model family than the one that produces the reference extraction, so the two biases do not line up.
We also know the judge's limits. Its noise floor is around one point of mean score, so anything smaller than that is treated as noise and re-run rather than celebrated.
The build gate is separate and stricter. A set of golden-file tests, committed snapshots of extraction output for specific pages, runs on every change. A change that breaks them is reverted before it is ever measured. That is the part that stops the harness from being optional.
Where this goes wrong
We keep a list of known defects and it is not short. Current entries include:
- A page where the extractor lost the title and every heading, and emitted empty code fences for Haskell code blocks. It scored 2 out of 5 on recall and 1 out of 5 on fidelity.
- A crash in the collector on complex nested tables, which surfaces as an extractor error on those pages.
- Truncation on very long pages.
- Missing bylines and dates on news and documentation pages.
The corpus has blind spots too. Pages that block our fetcher, sit behind a paywall, or render only in JavaScript are excluded from the default corpus, so the benchmark understates how hard those pages are. It is also English-heavy: 687 of the 799 pages are English, so the multilingual results carry less weight than the label count suggests.
We plan to publish the corpus and the harness so you can reproduce the numbers and argue with them. A benchmark you cannot inspect is marketing.
Try to break it
The most useful thing you can send us is a page that other extractors get wrong. We will run it and show you the output, including on the pages where we lose.
Sign up for 500 free credits, no card required, and send a URL through the fetch API. If you would rather compare implementations on your own corpus, the extractor is a library with its own tests and you can drive it directly. The REST API reference lists every format and option it supports.