The API that turns web pages into data, a founder letter
This week I am launching the free public beta of Danubia, a content extraction API. Let me tell you more!
Hi đź‘‹
Danubia is an API that turns web pages into high-quality data. You give it a URL, it sends back the page's main content as clean markdown, on demand or in batch.
I used to be an engineer on the crawl side of a large search engine. Now I keep watching scraping projects reinvent the same hard, fragile parts of a crawler, and I keep thinking to myself:
"Hey, I think I know a better way"
So I started applying what I know about running crawlers in production to a brand new product. I'd be pleased if you gave Danubia a try. Here's why I think it's worth your time.
Who is this for?
Here's why you may be interested in something like this:
You feed LLMs with fresh web data. For RAG, agents, or training and evaluation sets, clean markdown is what you want to work with, not the raw HTML with cookie banners, nav menus and JavaScript stubs. A cleaner input means more effecient token usage and better answers.
You build products on top of scraped content. Chatbots grounded in live sources, internal
search, market research, newsletter digests. You don't want to think about robots.txt or rate
limits, and you don't want to build the crawling infrastructure, you want a dependable data source.
You need the web, but as data. This is the long-term vision for this product: clean text today, structured data extraction at scale next. The end goal is to automate the process of creating clean datasets from web data.
Here's a little playground so you can give it a try:
RequĂŞte
curl -X POST undefined/api/v1/demo \ -H "Content-Type: application/json" \ -d '{"url": "https://github.com/trending"}'Réponse en direct
Appuyez sur Lancer pour récupérer une vraie page.
If this doesn't sound like something you might need, that's alright! If it does, please bear with me for a while as I tell you where I'm headed.
The secret sauce is the content extraction
To get the main content of a page without the boilerplate, Danubia runs its own content extraction library. It was inspired by state-of-the-art solutions like Trafilatura, but rebuilt from scratch for speed and accuracy. I believe that's where most of the value of Danubia will come from, which is why you should consider giving it a try before building your own thing: best-in-class extraction.
I keep a benchmark that combines traditional metrics and LLM-as-judge, comparing Danubia's extraction against well-established alternatives (Trafilatura, Readability, Turndown, Firecrawl) across a wide set of pages. On most of these pages, my extraction performs at least as well, and often better, than the alternatives. And I'll keep working on it.
For AI pipelines the payoff is direct: cleaner markdown in, better answers and fewer wasted tokens out.
A good bot beats an aggressive one
Most scraping pipelines get rejected, and the first reflex is to get cleverer at sneaking past blocks. My experience building a real crawler taught me the opposite: a polite bot has a better chance of getting the data.
Being polite means more than respecting robots.txt. It means coordinating fetches so a single
website never gets hammered, and you never see a wall of "429 Too Many Requests" from your side. You
just get the data. That's genuinely hard when everyone runs their scraper from a laptop.
The solution, I think, is to build something similar to what general-purpose search engines run: a large-scale crawler that enforces per-domain concurrency limits and crawl delays at scale. Because the fetch layer is shared by everyone, we all benefit from one well-behaved bot instead of thousands of noisy ones. When two users of Danubia ask for the same page in a short window, we only fetch that page once.
Where we're at, and where this is going
Danubia is currently in a free public beta. Sign up and get 500 credits to play with, and if you want to keep building I'll top you up in exchange for honest feedback. When billing finally turns on, you'll be able to buy more credits as you need them.
There's a batch API too, if you'd rather scrape a large dataset in one go: hand it a big JSONL of
URLs and get a big JSONL of clean text back.
Beyond that, here's what I'm excited about:
A structured data API. Today you hand me a URL and get back markdown. The upgrade I'm building next: you hand me a JSON schema and a list of URLs, and I hand you a clean dataset, shaped exactly as you asked. Give it a schema, get a dataset.
import { danubia } from "@danubia/client";
const result = await danubia.extract({
schema: {
name: "string",
price: "number",
rating: "number",
topReviews: [{ text: "string", score: "number" }],
},
urls: ["https://example.com/product/1", "https://example.com/product/2"],
});
result.rows;
// [
// { name: "…", price: 29.9, rating: 4.6, topReviews: [{ text: "…", score: 5 }] },
// ]
Crawling, too. Today you bring the URLs. Soon you won't have to: give me a starting point, I'll crawl the site and hand you a complete map of it, plus its content.
In any case, I'm building in public. I'll publish the roadmap and regular progress reports here, and early adopters will shape the product: if you're building on Danubia, I want to talk to you. I need your honest take on what's clunky, what's missing, and what a structured-data API needs to look like. Feel free to reach out to me at [email protected], I read every email.
A small business built in the EU
Finally, Danubia is built in the EU by a solo French developer and hosted on European cloud providers.
If you care about supporting European alternatives for digital services, if you're a European builder willing to support local businesses, or if you simply care about small businesses having a chance to compete, that might be a good reason to give my product a go.
Sign up to get your first 500 credits for free. If you're hitting a limit, reach out to me and I'll top you up.
See ya. Guillaume, founder of Danubia