Docs
Supported sources
What anymd converts today (web pages, X, YouTube, GitHub, Reddit, Hacker News, PDFs, Office files, images), what each returns, limits, and what's planned.
anymd picks a converter per URL. Specialised adapters go first; everything else goes through the general web pipeline. The kind field in the frontmatter (and the X-Anymd-Kind header) tells you which one ran.
At a glance
| Source | kind |
What you get | Credits |
|---|---|---|---|
| Any web page | web |
Main content with clutter removed | 1 |
| X / Twitter | x |
Posts, long-form Articles, quotes, polls, media, engagement stats | 1 |
| YouTube | youtube |
Title, channel, thumbnail and a timestamped transcript | 3 |
| GitHub | github |
READMEs, issues, PRs, discussions | 1 |
reddit |
Threads with comments | 1 | |
| Hacker News | hackernews |
Story plus the top of the discussion, nested | 1 |
pdf |
Text with layout and tables kept | 3 | |
| DOCX, XLSX, XLS, ODS, ODT, CSV | document |
Documents as Markdown, sheets as tables | 3 |
| Images (JPEG, PNG, WebP, SVG) | image |
A vision-model description and transcription | 5 |
| Plain text, Markdown, JSON URLs | text |
Passed through (JSON in a code block) | 1 |
Cached results cost 0. Failed conversions cost 0. Full pricing: Billing & credits.
Web pages
The web pipeline fetches the page, parses it with linkedom, and runs a content-extraction engine that finds the main content, removes navigation, ads and boilerplate, standardises headings, code blocks and footnotes, and emits Markdown.
anymd adds a few things on top:
- Retry for empty shells. If a client-rendered page comes back with almost no text, anymd retries once with its bot user agent and keeps the richer result. It always identifies as
anymd. - GitHub gets a bot user agent first, because it serves server-rendered content to bots.
- Documents behind URLs. If a URL serves a PDF, Office file or image, it is routed to the document converter automatically.
- Your overrides.
selectorpicks the content block when automatic detection misses;images=0strips images;langsetsAccept-Language. See URL API.
Limit: 5 MB of HTML per page.
Site rules anymd follows
Before fetching a page itself, anymd checks the site's robots.txt for the anymd token (or *) and refuses disallowed pages with 403 robots_disallowed. If robots.txt cannot be read because of a server error, the request fails with 503 robots_unreachable. Each site also has a shared per-minute fetch budget (429 domain_rate_limited), and domains whose owners opted out are refused on every source with 403 site_opted_out. None of these use credits. Site owners: see Site Owners & Abuse.
X / Twitter
Status URLs on x.com and twitter.com (including mobile.) go through the FxTwitter API. You get the post text, long-form Articles, quoted posts, polls, photos and videos, plus likes, retweets, replies and views in the frontmatter.
YouTube
Watch, youtu.be, Shorts and Live URLs. Title, channel and thumbnail come from YouTube's oEmbed endpoint. The transcript comes from a third-party caption provider and is grouped into readable paragraphs with timestamps:
## Transcript
**0:00** · Welcome back. Today we're looking at…
**0:42** · The first thing to notice is…
Pass lang=vi (or any language code) to prefer a transcript language. When no transcript is available, the output says so and still includes the video's metadata. That conversion is still charged.
Hacker News
news.ycombinator.com/item?id=… uses the official Hacker News API: the story (points, author, comment count, link, text) and the top 20 comment threads, each with up to 3 replies, nested as blockquotes up to 3 levels deep.
Files and documents
PDFs, images, Office documents and spreadsheets are converted with Cloudflare Workers AI toMarkdown:
- PDF: text layout kept, tables preserved.
- DOCX, ODT: document text as Markdown.
- XLSX, XLS, ODS, CSV: sheets become Markdown tables.
- Images: a vision model describes and transcribes the image; the Markdown starts with the image itself.
Convert a file behind a URL with the URL API, or upload one with POST /api/v1/convert/file or anymd file <path>. Limit: 20 MB per file.
What won't convert
- Private networks,
localhost, internal hostnames and URLs with credentials. They're blocked on purpose. - Pages behind a login or paywall. anymd sees what an anonymous visitor sees.
- Content types other than HTML, text, JSON and the document types above return
415 unsupported_type.
If a public page converts badly, open an issue with the URL. That's how adapters get better.
Planned
These are on the roadmap and not available yet:
- Audio and podcasts
- Any video (speech-to-text, beyond YouTube captions)
- Threads
- TikTok
- Notion
- Google Docs
- EPUB
Until then, the general web pipeline may still extract something useful from some of these, with no guarantee of quality.
Updated 2026-09-26 · Edit on GitHub