High-fidelity document parsing for AI, self-hosted

Structure, tables and layout preserved — not just text. No cloud account: docker run and go.

  • < 5min

    from zero to production document parsing
  • -40%*

    fewer LLM tokens with DocLang
How it works

Three steps to structured data

From a raw document to clean DocLang or JSON, entirely inside your own infrastructure.

  1. 01 Docker Run

    Pull the container and run it with your license file. Mount a volume for the results.

    $ docker run -d -p 8080:8080 \
      -e FINEPARSER_LICENSE_DATA="$(cat acme.fineparserlicense)" \
      -v fineparser-output:/app/output \
      abbyyteam/fineparser
  2. 02 Submit

    POST a document and the output type. You get a job ID back right away.

    curl -X POST localhost:8080/parse \
      -F file=@form.pdf \
      -F outputType=doclang
    # {"jobId":"6c4f1f0e-…","file":"form.pdf"}
  3. 03 Download

    Poll the job until it succeeds, then download DocLang or JSON.

    <doclang>
      <heading level="1">Form 1040 (2025)</heading>
      <table>
        <ched/>Line<ched/>Description<ched/>Amount<nl/>
        <fcel/>1a<fcel/>Wages (W-2 box 1)<fcel/>$68,400<nl/>
        <fcel/>11<fcel/>Adjusted gross income<fcel/>$71,250<nl/>
        <srow/>Total tax<lcel/><lcel/><fcel/>$9,842<nl/>
      </table>
    </doclang>
Developer experience

One small API, any language

POST a file to /parse, poll the job, download the result. Plain HTTP, so nothing to install and no SDK to keep up to date.

# 1. Submit. FineParser accepts the file and returns a job ID right away.
curl -s -X POST http://localhost:8080/parse \
  -F "file=@invoice.pdf" \
  -F "outputType=doclang"
# {"jobId":"6c4f1f0e-…","file":"invoice.pdf"}

# 2. Poll until status is "successful".
curl -s http://localhost:8080/jobs/6c4f1f0e-…
# {"status":"successful","file":"6c4f1f0e-…/invoice.pdf.doclang"}

# 3. Download the result.
curl -OJ http://localhost:8080/jobs/6c4f1f0e-…/invoice.pdf.doclang

Writing the integration with Claude Code or Codex? Point it at the ABBYY docs MCP server, or drop the repo's AGENTS.md runbook into your project, and it writes against the real API. See how →

Built for LLMs

DocLang speaks fewer tokens

DocLang's markup maps cleanly to LLM tokens. Extend your tokenizer with it, and the same document costs up to 40% less to feed into your model.*

1

Add the special tokens

Extend any Hugging Face tokenizer with DocLang's vocabulary in a few lines:

from transformers import AutoTokenizer
import doclang.tokenization as dt

tokenizer = AutoTokenizer.from_pretrained("openai/gpt-oss-20b")
tokenizer.add_special_tokens(
    {"additional_special_tokens": dt.get_special_tokens()}
)
2

Fewer tokens per document

Real reductions across two production-grade tokenizers:

-41%

granite-4.0

-45%

gpt-oss-20b

*Full methodology and benchmark details in the FAQ.

About

Built on three decades of OCR expertise

FineParser is made by ABBYY, the document recognition company founded in 1989, named a Leader in the 2026 Gartner® Magic Quadrant™ for Intelligent Document Processing Solutions. Learn more about ABBYY →

Founded in 1989
500+ employees worldwide
Pricing

Start free. Scale when you're ready.

All plans are self-service monthly subscriptions — except Enterprise. Pull the Docker image and parse your first document in under five minutes.

Processing stops when the monthly page limit is reached. No overages — upgrade or wait for the next billing period.

Enterprise

Custom pricing

Unlimited pages, orchestrated throughput, fully air-gapped deployment and engine tuning.

  • Everything in Business
  • Engine tuning: full access to recognition settings
  • Orchestrated throughput of 500+ pages / second
  • Unlimited pages
  • Air-gapped deployment, zero connectivity
  • Enterprise SLAs
  • ABBYY technical support

The FineReader Engine behind FineParser is trusted by Volkswagen and enterprises worldwide.

Talk to sales

See full pricing and feature comparison →

Support

Get help on GitHub

Every self-service tier is community-supported. Report a bug, ask a question or request a feature, all in the open on GitHub Issues.

Enterprise customers get contractual SLAs and ABBYY technical support. Compare support levels →

FAQ

Frequently asked questions

What is FineParser?

FineParser is a self-hosted, Docker-native document-parsing engine for AI applications, powered by ABBYY FineReader Engine.

Is it self-hosted?

Yes. FineParser runs entirely on your own infrastructure via docker run. Sign up to get your .fineparserlicense file and pass its contents to the container in the FINEPARSER_LICENSE_DATA environment variable. No key to activate and no license file to mount; just mount a volume for the results.

How much does it cost?

FineParser starts free — 1,000 pages/month for a year. Paid plans run $119–$2,000/month for 25,000 to 1,000,000 pages, and Enterprise offers custom, volume-based licensing with no monthly ceiling. See pricing for the full tier and feature comparison.

What formats does it export?

Three: DocLang, JSON and TXT. DocLang is the structure-preserving default built for LLMs; JSON carries the same structure with pixel-accurate bounding boxes.

Does DocLang reduce LLM token costs?

Yes. DocLang's markup is a controlled vocabulary built to map cleanly to LLM tokens, and extending your tokenizer with its special tokens cuts token consumption by roughly 40% on average.

*Methodology: benchmarked across 8,030 DocLang documents, comparing each tokenizer's base vocabulary against the same tokenizer extended with DocLang's special tokens via the open-source doclang package (doclang.tokenization.get_special_tokens()). Results: -41% mean / -40% median on granite-4.0, and -45% mean / -45% median on gpt-oss-20b. Actual savings depend on your documents and choice of base tokenizer. See DocLang speaks fewer tokens for the integration code.

What languages does it support?

200+ languages, including Latin, Cyrillic, CJK, and Arabic scripts.

Can it run fully offline / air-gapped?

FineParser needs outbound access to ABBYY's license server to validate your license and meter pages; without it, nothing parses. It also sends usage telemetry. Document content never leaves your infrastructure. The Enterprise tier supports fully air-gapped deployment with zero connectivity.

Start parsing documents in minutes.

$ docker pull abbyyteam/fineparser
Start for free