Markdown converter for documents and web pages.
A Markdown converter for documents and public web pages: PDF, Word, PowerPoint, Excel, images and HTML become text you can inspect, reuse and pass to an AI workflow.
Anonymous trial · 10 MB · 3 conversions / hour · 10 / day
Convert to Markdown from
- DOCX
- PPTX
- XLSX
- PNG / JPG
- HTML
- URL
How it works
How the Markdown converter works
markitdown.ai is a markdown converter that turns PDFs, Office files, images, and web pages into clean, AI-ready Markdown for ChatGPT, Claude, RAG pipelines, and agents through an anonymous trial or a paid account with developer API access.

What happens
Upload a file or paste a URL. The conversion runs on our servers, not as a browser copy-paste hack, so the same file produces the same Markdown wherever it comes in. Each page is handled by what is on it, and headings, lists and tables come out as real Markdown structure.
What you get
- Text-layer pages are read directly; scans and images go through OCR automatically
- URL fetches run from our servers with private and internal addresses blocked
- Public converter: up to 10 MB and 25 pages, no sign-in
- Signed-in plans raise both limits, up to 200 MB and 200 pages on Pro
What happens
The viewer places the source document and the generated Markdown next to each other, so you can check headings, tables and reading order before you use the result. Switch between the rendered preview and the raw Markdown to see exactly what a model or a pipeline will receive.
What you get
- Source and Markdown render side by side in the viewer
- Switch the Markdown pane between rendered preview and raw source
- Copy the result for free, no account needed
- Download the .md file once you're signed in
What happens
The same settings that worked for one file carry over to a stack of files or a pipeline call, and every result lands in your Library. Batch tracks each file with its own status and retry, and the API hands the same Markdown straight to your code.
What you get
- Same parsing engine for the web app, Batch (ZIP or a list of URLs) and the API
- POST /v1/convert/{name} returns Markdown in one call within 60 seconds, or a conversion you poll or receive by webhook
- Library keeps each run for your plan's history window: 30, 60 or 90 days
Credits
Markdown accuracy, priced per page.
A text-layer invoice and a scanned balance sheet cost the same per page: the pipeline picks direct extraction or OCR for you. AI extras for figures and images are opt-in and metered separately.
- per page, text or scan
- 1 credit
- per page, text or scan
Text-layer PDF, Office files and HTML
OCR for scanned pages, photos and screenshots
Charts and figures described in PDFs · paid plans
Image content described, not just OCR · paid plans
Markdown API
A Markdown API for your RAG pipeline.
- The same parsing pipeline as the web app
- Batch thousands of files with webhooks on completion
- Scoped API keys and a full activity log
- Markdown in one call, or a conversion to poll when a file needs more than 60 seconds
curl -X POST https://api.markitdown.ai/v1/convert/word \ -H "x-api-key: $MARKITDOWN_API_KEY" \ -F file=@vendor-security-review-q3-2026.docx # If HTTP 202, GET the Location URL with the same API key.# Verified local parser output for vendor-security-review-q3-2026.docx:# ## Review Findings# # Risk ratings follow the standard High / Medium / Low scale defined in the vendor risk management policy.# # | Vendor | Risk Rating | Finding | Owner |# | --- | --- | --- | --- |# | CloudScan OCR | Medium | No documented data retention limit on uploaded scans | Security |# | PayBridge | Low | SOC 2 Type II renewed; no open items | Finance |# | LingoTrans API | High | Sub-processor list not disclosed on request | Legal |# | ArchiveNow Storage | Medium | Encryption at rest confirmed; key rotation overdue | Security |Pricing
Simple pricing for Markdown conversion.
The problem
Documents aren't AI-ready
Most documents were built for people and printers, not for language models. Before a file can be embedded, retrieved, or reasoned over, it has to become predictable text. That is harder than it looks, and it is exactly where naive extraction breaks down.

Reading order breaks
Multi-column PDFs, headers, and footnotes get interleaved, so the text an LLM reads no longer matches the document a human sees.
Tables and headings flatten
Copy-paste collapses tables into runs of numbers and drops the heading hierarchy that gives a document its structure.
Scans need OCR
Image-based pages and scanned contracts carry no text layer at all until they are run through optical character recognition.
Raw text is noisy
Ad-hoc extraction leaves page numbers, broken hyphenation, and stray characters that pollute prompts and retrieval results.
RAG pipelines and agents need consistent structure, not just extracted characters. Predictable Markdown is what makes chunking, embedding, and retrieval behave the same way across thousands of documents instead of failing quietly on the messy ones. The parser that handles the ugly files is what keeps the clean ones trustworthy.
Formats in
Convert any document to Markdown
Every input format has its own converter page with limits, format notes, and a sample output. They all share one parsing engine, so pick the page that matches your file:
- PDF to Markdown — reports, papers, manuals, and contracts. Native text PDFs parse directly with multi-column reading order rebuilt; scanned PDFs go through OCR page by page.
- Word to Markdown — .docx and legacy .doc files with headings, lists, tables, and comments preserved, so a policy or a spec reads as an outline instead of one block.
- PPT to Markdown — .pptx and .ppt decks slide by slide, with slide titles as headings and bullets and tables kept, which makes a 40-slide deck searchable.
- Excel to Markdown — .xlsx and .xls workbooks where each sheet becomes a Markdown table under its sheet name.
- Image to Markdown — PNG and JPG screenshots and scans through OCR. AI image understanding is a separate, signed-in option.
- HTML to Markdown — saved pages, exports, and fragments with scripts and styles stripped and the article structure, links, and tables kept.
- URL to Markdown — paste a public web address and get the page as Markdown, fetched server-side with a size cap and a timeout so a slow site cannot hang the request.
Formats out
From Markdown to any format
Markdown is the interchange format, so conversion also runs the other way. The Markdown editors below render as you type in the browser and never upload your text:
- Markdown to HTML — clean semantic HTML with GitHub-flavored tables, task lists, and code highlighting, ready to paste into a CMS or an email.
- Markdown to text — plain text with headings, lists, and tables flattened readably for chat windows, tickets, and forms.
- Markdown to PDF — a local print preview; use Print / Save as PDF in your browser and review the saved pages.
Structured JSON extraction and audio and video transcription are planned layers on top of Markdown. They are listed as roadmap, not as shipped features.
Use cases
Made for every team
The Markdown converter offers an anonymous trial for checking a first result. The reason teams keep using it is everything after the first file: history, batches, larger documents, and an API that folds conversion into the systems you already run. The same Markdown works whether the consumer is a person, a search index, or a model.
RAG and AI pipelines
Convert source documents into clean Markdown before chunking, embedding, or agent processing. Predictable structure means your splitter sees real headings and tables instead of a wall of extracted text, which keeps retrieval relevant and prompts free of layout noise.
Research and reports
Turn papers, whitepapers, and slide decks into editable, searchable Markdown. Researchers and analysts can quote, annotate, summarize, and reuse findings without retyping tables or losing the reading order of a dense PDF.
Operations documents
Process policies, contracts, invoices, and internal files without manual copy-paste cleanup. A saved Markdown history makes recurring document work easier to review, compare, and reuse across a team's day-to-day operations.
Developer automation
Use the API to convert files inside ingestion jobs, internal tools, and document workflows. Submit a request, receive Markdown, and feed it into a CMS, knowledge base, support bot, or agent pipeline with output you can depend on.
Knowledge bases and support
Bring product manuals, release notes, and internal wikis into one Markdown corpus that a support bot or an internal assistant can search. Because headings and tables survive, answers point at the right section instead of a page-sized blob.
Agents and tools
Give an agent a Markdown converter for supported files and public URLs, and check completion before it consumes the result. Read this site the same way: every public page answers Accept: text/markdown with a Markdown twin, and llms.txt lists them all.
Output quality
Structure, not just text
A useful Markdown converter preserves more than isolated characters from a PDF. The goal here is Markdown that keeps the structure your downstream tools depend on, produced by a server-side parsing pipeline instead of fragile browser copy-paste.
- Preserves headings, lists, tables, links, and reading order where the source has them.
- Handles PDF, Word, PowerPoint, Excel, images, HTML, and URLs through one pipeline.
- Processes scanned PDFs and images with OCR today; recognition quality depends on the source image.
- Runs long documents asynchronously, so large reports and decks finish instead of timing out.
Standard pages and OCR pages cost 1 credit per page; paid-plan accounts can use AI image understanding at 5 credits per image. Output follows CommonMark with GitHub-flavored tables, so it renders in any Markdown viewer and parses with any Markdown library. Every conversion should produce an artifact you can inspect, save, and reuse — that is the bar.

Three ways in
Browser, API, or agent
The Markdown converter in the web app covers one-off and recurring manual conversion. Batch covers stacks of files. The API covers everything that should happen automatically, including agents that call it as a tool. All three share one account, one credit balance, and one parsing engine. Most teams start in the browser and move to the API once conversion becomes part of a product or an internal workflow.
| Aspect | Browser | Batch | API |
|---|---|---|---|
| Best for | One file at a time | Many files in one run | Pipelines, tools, and agents |
| Sign-in | Not needed for the first file | Paid plans | Paid plans, API key |
| Input | Upload or URL | A list of uploads or file IDs | Upload, URL, inline text, or file ID |
| Output | Preview, copy, download .md | Library with per-file status | JSON with Markdown, per-page output, webhooks |
| Limits | 10 MB and 25 pages without sign-in | Plan concurrency | 20 to 60 requests per minute by plan |
| Wait | Seconds, in the page | Background | Sync up to 60 s, then a conversion you poll |
Web app
- Upload files or paste URLs and preview rendered Markdown
- Save conversion history and download .md results
- Batch convert repeated work with per-file status
- Track credits and usage per plan
Developer API
- POST /v1/convert/{name} returns Markdown in one call, or a conversion when a file needs more than 60 seconds
- Poll conversions or receive webhooks when they settle
- Send Accept: text/markdown to read any public page as Markdown
- Machine-readable spec at /v1/openapi.json
FAQ
Questions teams ask
What does a Markdown converter do?
A document to Markdown converter is a tool that reads a PDF, Word, PowerPoint, Excel, image, or web page and rewrites its content as Markdown: headings become # lines, tables become pipe tables, and lists stay lists. The point is predictable plain text that language models, RAG pipelines, and agents can chunk, embed, and cite without layout noise.
Is the free converter really free?
Yes. The public converter pages need no account and no card. Files up to 10 MB and 25 pages convert in the browser, and you can preview and copy the Markdown. Lite, Pro, and Max subscriptions start at $20 per month and add a saved library, batch, and the API.
How do scanned PDFs and images work?
Image-only pages carry no text layer, so the pipeline runs OCR automatically and returns Markdown the same way it does for digital text. Standard pages and OCR pages cost 1 credit per page. Recognition quality depends on the source image: a 300 dpi scan reads far better than a phone photo of a page.
Can I run this from code or from an agent?
Yes. POST /v1/convert/{name} accepts an upload, a URL, or inline text with an API key and returns Markdown in one call when the file finishes within 60 seconds, or a conversion you can poll or receive by webhook. Agents can request Accept: text/markdown on any public page, and the machine-readable spec lives at /v1/openapi.json.
What happens to my files after conversion?
Deleting a conversion removes it from your library at once and erases its stored files within an hour; a source file shared with another conversion is kept until that one is deleted too. Deleting your account erases all your stored source and result files. Paid plans keep history for 30, 60, or 90 days depending on the plan.