Public beta now - Read our announcement

Turn any web page into clean text.

Danubia is a scraping and content extraction API for developers. Send a URL and get clean markdown back, on demand or in batch.

One request. Clean text back.

Request

curl -X POST https://app.danubia.tech/api/v1/demo \
  -H "Content-Type: application/json" \
  -d '{"url": "https://github.com/trending"}'

Live response

Press Run to fetch a real page.

LLM-ready data

The page's main content as clean markdown, HTML or plain text, plus structured metadata, ready to feed your models, indexes and agents.

Polite by design

robots.txt respected and enforced, with strict per-domain concurrency limits. We crawl the way sites want to be crawled.

Built in the EU 馃嚜馃嚭

Servers in the European Union, GDPR-compliant by default. Your data never leaves the Union.

Use cases

What you can build with Danubia

If your product needs content from the web, Danubia fetches and cleans it for RAG, aggregation and datasets.

Feed your RAG pipeline

Send a URL and get clean markdown with ads, nav and boilerplate stripped, the exact input your embeddings and agents need.

Aggregate content from anywhere

Pull clean articles from dozens of sources into one feed, without writing a scraper per site.

Create datasets

Send a batch of URLs and get a full dataset of clean text back, without building or running your own crawling infrastructure.

How it works

From URL to clean text in three steps

You send URLs. We handle the fetching.

  1. 01

    Send a URL

    Call the API with a URL, or drop thousands into the batch queue.

  2. 02

    We crawl it

    A worker fetches the page, follows redirects, and extracts the main content.

  3. 03

    Get clean text

    You get markdown and structured metadata in one response.

FAQ

Questions, answered

The short version. Need more detail? Email us at [email protected].

Is Danubia free during the public beta?
Danubia is in public beta. Everyone who signs up gets free credits to try our best-in-class content extraction.
Do you respect robots.txt?
Yes. Every fetch is checked against robots.txt, including redirect targets, and crawl delays are honored per host. We never bypass blocks.
What do I get back?
A JSON response with the page's main content as markdown, plus structured metadata (title, description, language) and fetch metadata like the final URL and response time.
Do you render JavaScript?
Yes. When you call our API, you can choose to use a real browser. Browser rendering costs more credits.
How does batch pricing work?
Batch jobs use the same prepaid credits as on-the-fly requests, at half the cost: a fresh fetch costs 0.5 credits in batch mode, cache hits 0.5 credit, failures free. Your credits are shared across both.
Is my data used to train models?
Never. Content you scrape through Danubia is only stored to serve your requests and is never used for training or shared with third parties.
What are the rate limits?
Every account works from a shared pool of prepaid credits. Rate limits protect the service and scale with your account; a zero balance blocks new work until you top up.
Where is my data stored?
In the European Union. Danubia is built and hosted in the EU. Fetching, storage and backups all run on EU servers. We are GDPR-ready (a data processing agreement is available on request), and content you scrape is never used to train models or shared with third parties.
What is a content extraction API?
It is an API that returns the main content of a page instead of the whole document. You send a URL and get the article text back as clean markdown, with the navigation, ads, cookie banners and script stubs removed. That is the input a retrieval pipeline or a language model actually wants.
Is Danubia a crawling API?
Danubia fetches the URLs you give it, one at a time or as a batch of thousands, and it does not follow links on its own yet. Crawling from a starting point is on the roadmap. To turn a URL list into clean text today, the batch API is the path.

Start scraping today.

Get your API key, send your first URL, and get clean text back.

New accounts get 500 credits on us.

Built & hosted in the EU 路 GDPR-ready 路 No training on your data

漏 2026 Danubia. All rights reserved.

Built & hosted in the EU

Built with daisyUI