NewMarkdown to HTML, text and PDF tools

Markdown converter for documents and web pages.

A Markdown converter for documents and public web pages: PDF, Word, PowerPoint, Excel, images and HTML become text you can inspect, reuse and pass to an AI workflow.

Anonymous trial · 10 MB · 3 conversions / hour · 10 / day

Convert to Markdown from

  • PDF
  • DOCX
  • PPTX
  • XLSX
  • PNG / JPG
  • HTML
  • URL

How it works

How the Markdown converter works

markitdown.ai is a markdown converter that turns PDFs, Office files, images, and web pages into clean, AI-ready Markdown for ChatGPT, Claude, RAG pipelines, and agents through an anonymous trial or a paid account with developer API access.

Markdown Converter turning a DOCX document into a Markdown window

What happens

Upload a file or paste a URL. The conversion runs on our servers, not as a browser copy-paste hack, so the same file produces the same Markdown wherever it comes in. Each page is handled by what is on it, and headings, lists and tables come out as real Markdown structure.

What you get

  • Text-layer pages are read directly; scans and images go through OCR automatically
  • URL fetches run from our servers with private and internal addresses blocked
  • Public converter: up to 10 MB and 25 pages, no sign-in
  • Signed-in plans raise both limits, up to 200 MB and 200 pages on Pro

Credits

Markdown accuracy, priced per page.

A text-layer invoice and a scanned balance sheet cost the same per page: the pipeline picks direct extraction or OCR for you. AI extras for figures and images are opt-in and metered separately.

per page, text or scan
1 credit
per page, text or scan
Documents

Text-layer PDF, Office files and HTML

1 credit / page
Scans & images

OCR for scanned pages, photos and screenshots

1 credit / page
AI figure enhancement

Charts and figures described in PDFs · paid plans

3 credits / figure
AI image understanding

Image content described, not just OCR · paid plans

5 credits / image

Markdown API

A Markdown API for your RAG pipeline.

  • The same parsing pipeline as the web app
  • Batch thousands of files with webhooks on completion
  • Scoped API keys and a full activity log
  • Markdown in one call, or a conversion to poll when a file needs more than 60 seconds
Read the API reference
curl -X POST https://api.markitdown.ai/v1/convert/word \  -H "x-api-key: $MARKITDOWN_API_KEY" \  -F file=@vendor-security-review-q3-2026.docx # If HTTP 202, GET the Location URL with the same API key.# Verified local parser output for vendor-security-review-q3-2026.docx:# ## Review Findings# # Risk ratings follow the standard High / Medium / Low scale defined in the vendor risk management policy.# # | Vendor | Risk Rating | Finding | Owner |# | --- | --- | --- | --- |# | CloudScan OCR | Medium | No documented data retention limit on uploaded scans | Security |# | PayBridge | Low | SOC 2 Type II renewed; no open items | Finance |# | LingoTrans API | High | Sub-processor list not disclosed on request | Legal |# | ArchiveNow Storage | Medium | Encryption at rest confirmed; key rotation overdue | Security |

The problem

Documents aren't AI-ready

Most documents were built for people and printers, not for language models. Before a file can be embedded, retrieved, or reasoned over, it has to become predictable text. That is harder than it looks, and it is exactly where naive extraction breaks down.

Document to Markdown: copy-pasted text beside clean Markdown structure

Reading order breaks

Multi-column PDFs, headers, and footnotes get interleaved, so the text an LLM reads no longer matches the document a human sees.

Tables and headings flatten

Copy-paste collapses tables into runs of numbers and drops the heading hierarchy that gives a document its structure.

Scans need OCR

Image-based pages and scanned contracts carry no text layer at all until they are run through optical character recognition.

Raw text is noisy

Ad-hoc extraction leaves page numbers, broken hyphenation, and stray characters that pollute prompts and retrieval results.

RAG pipelines and agents need consistent structure, not just extracted characters. Predictable Markdown is what makes chunking, embedding, and retrieval behave the same way across thousands of documents instead of failing quietly on the messy ones. The parser that handles the ugly files is what keeps the clean ones trustworthy.

Formats in

Convert any document to Markdown

Every input format has its own converter page with limits, format notes, and a sample output. They all share one parsing engine, so pick the page that matches your file:

  • PDF to Markdown — reports, papers, manuals, and contracts. Native text PDFs parse directly with multi-column reading order rebuilt; scanned PDFs go through OCR page by page.
  • Word to Markdown — .docx and legacy .doc files with headings, lists, tables, and comments preserved, so a policy or a spec reads as an outline instead of one block.
  • PPT to Markdown — .pptx and .ppt decks slide by slide, with slide titles as headings and bullets and tables kept, which makes a 40-slide deck searchable.
  • Excel to Markdown — .xlsx and .xls workbooks where each sheet becomes a Markdown table under its sheet name.
  • Image to Markdown — PNG and JPG screenshots and scans through OCR. AI image understanding is a separate, signed-in option.
  • HTML to Markdown — saved pages, exports, and fragments with scripts and styles stripped and the article structure, links, and tables kept.
  • URL to Markdown — paste a public web address and get the page as Markdown, fetched server-side with a size cap and a timeout so a slow site cannot hang the request.

Formats out

From Markdown to any format

Markdown is the interchange format, so conversion also runs the other way. The Markdown editors below render as you type in the browser and never upload your text:

  • Markdown to HTML — clean semantic HTML with GitHub-flavored tables, task lists, and code highlighting, ready to paste into a CMS or an email.
  • Markdown to text — plain text with headings, lists, and tables flattened readably for chat windows, tickets, and forms.
  • Markdown to PDF — a local print preview; use Print / Save as PDF in your browser and review the saved pages.

Structured JSON extraction and audio and video transcription are planned layers on top of Markdown. They are listed as roadmap, not as shipped features.

Use cases

Made for every team

The Markdown converter offers an anonymous trial for checking a first result. The reason teams keep using it is everything after the first file: history, batches, larger documents, and an API that folds conversion into the systems you already run. The same Markdown works whether the consumer is a person, a search index, or a model.

RAG and AI pipelines

Convert source documents into clean Markdown before chunking, embedding, or agent processing. Predictable structure means your splitter sees real headings and tables instead of a wall of extracted text, which keeps retrieval relevant and prompts free of layout noise.

Research and reports

Turn papers, whitepapers, and slide decks into editable, searchable Markdown. Researchers and analysts can quote, annotate, summarize, and reuse findings without retyping tables or losing the reading order of a dense PDF.

Operations documents

Process policies, contracts, invoices, and internal files without manual copy-paste cleanup. A saved Markdown history makes recurring document work easier to review, compare, and reuse across a team's day-to-day operations.

Developer automation

Use the API to convert files inside ingestion jobs, internal tools, and document workflows. Submit a request, receive Markdown, and feed it into a CMS, knowledge base, support bot, or agent pipeline with output you can depend on.

Knowledge bases and support

Bring product manuals, release notes, and internal wikis into one Markdown corpus that a support bot or an internal assistant can search. Because headings and tables survive, answers point at the right section instead of a page-sized blob.

Agents and tools

Give an agent a Markdown converter for supported files and public URLs, and check completion before it consumes the result. Read this site the same way: every public page answers Accept: text/markdown with a Markdown twin, and llms.txt lists them all.

Output quality

Structure, not just text

A useful Markdown converter preserves more than isolated characters from a PDF. The goal here is Markdown that keeps the structure your downstream tools depend on, produced by a server-side parsing pipeline instead of fragile browser copy-paste.

  • Preserves headings, lists, tables, links, and reading order where the source has them.
  • Handles PDF, Word, PowerPoint, Excel, images, HTML, and URLs through one pipeline.
  • Processes scanned PDFs and images with OCR today; recognition quality depends on the source image.
  • Runs long documents asynchronously, so large reports and decks finish instead of timing out.

Standard pages and OCR pages cost 1 credit per page; paid-plan accounts can use AI image understanding at 5 credits per image. Output follows CommonMark with GitHub-flavored tables, so it renders in any Markdown viewer and parses with any Markdown library. Every conversion should produce an artifact you can inspect, save, and reuse — that is the bar.

Markdown File Converter with stacked Markdown file cards and quality check marks

Three ways in

Browser, API, or agent

The Markdown converter in the web app covers one-off and recurring manual conversion. Batch covers stacks of files. The API covers everything that should happen automatically, including agents that call it as a tool. All three share one account, one credit balance, and one parsing engine. Most teams start in the browser and move to the API once conversion becomes part of a product or an internal workflow.

AspectBrowserBatchAPI
Best forOne file at a timeMany files in one runPipelines, tools, and agents
Sign-inNot needed for the first filePaid plansPaid plans, API key
InputUpload or URLA list of uploads or file IDsUpload, URL, inline text, or file ID
OutputPreview, copy, download .mdLibrary with per-file statusJSON with Markdown, per-page output, webhooks
Limits10 MB and 25 pages without sign-inPlan concurrency20 to 60 requests per minute by plan
WaitSeconds, in the pageBackgroundSync up to 60 s, then a conversion you poll

Web app

  • Upload files or paste URLs and preview rendered Markdown
  • Save conversion history and download .md results
  • Batch convert repeated work with per-file status
  • Track credits and usage per plan

Developer API

  • POST /v1/convert/{name} returns Markdown in one call, or a conversion when a file needs more than 60 seconds
  • Poll conversions or receive webhooks when they settle
  • Send Accept: text/markdown to read any public page as Markdown
  • Machine-readable spec at /v1/openapi.json

FAQ

Questions teams ask

What does a Markdown converter do?

A document to Markdown converter is a tool that reads a PDF, Word, PowerPoint, Excel, image, or web page and rewrites its content as Markdown: headings become # lines, tables become pipe tables, and lists stay lists. The point is predictable plain text that language models, RAG pipelines, and agents can chunk, embed, and cite without layout noise.

Is the free converter really free?

Yes. The public converter pages need no account and no card. Files up to 10 MB and 25 pages convert in the browser, and you can preview and copy the Markdown. Lite, Pro, and Max subscriptions start at $20 per month and add a saved library, batch, and the API.

How do scanned PDFs and images work?

Image-only pages carry no text layer, so the pipeline runs OCR automatically and returns Markdown the same way it does for digital text. Standard pages and OCR pages cost 1 credit per page. Recognition quality depends on the source image: a 300 dpi scan reads far better than a phone photo of a page.

Can I run this from code or from an agent?

Yes. POST /v1/convert/{name} accepts an upload, a URL, or inline text with an API key and returns Markdown in one call when the file finishes within 60 seconds, or a conversion you can poll or receive by webhook. Agents can request Accept: text/markdown on any public page, and the machine-readable spec lives at /v1/openapi.json.

What happens to my files after conversion?

Deleting a conversion removes it from your library at once and erases its stored files within an hour; a source file shared with another conversion is kept until that one is deleted too. Deleting your account erases all your stored source and result files. Paid plans keep history for 30, 60, or 90 days depending on the plan.

Convert your first file to Markdown.

Free to try — no sign-up needed.