PDF to Markdown: A Failure-First MinerU Workflow That Actually Survives Real Dependencies
If you need to convert a PDF to Markdown locally and the document contains columns, formulas, tables, or scanned pages, use a layout-aware parser rather than treating the PDF as plain text. For a current local workflow, MinerU is the more reliable starting point. A reproducible setup looks like this:
uv venv ~/venv/mineru --python 3.13
uv pip install --python ~/venv/mineru/bin/python "mineru[pipeline]"
~/venv/mineru/bin/mineru-models-download \
-s huggingface \
-m pipeline
MINERU_MODEL_SOURCE=local \
~/venv/mineru/bin/mineru \
-p input.pdf \
-o output \
-b pipeline
For simple single-column PDFs, a plain extractor may still be faster. But once the document becomes two-column, formula-heavy, table-heavy, or scanned, the main problem is no longer "extract text." It is reconstructing reading order and document structure. That distinction matters because many PDF-to-Markdown failures are not parsing failures at all. They are dependency, model, backend, or version failures that happen before the first useful Markdown line is written.
Start by deciding whether you actually need MinerU
A basic text PDF can often be extracted with:
pdftotext -layout input.pdf -
This is extremely fast when the document is single-column and mostly text. The trouble starts with academic papers and technical PDFs. In one real test, a 14-page single-column application note was extracted in roughly 0.03 seconds with correct reading order. The same method was then used on a 13-page two-column paper. It finished just as quickly, but the output mixed text from the left and right columns, preserved line-break hyphenation such as:
com-puter
syn-chronization
and dropped formulas and images entirely. That is the point where "PDF to text" stops being equivalent to "PDF to Markdown."
For complex documents, move to a structural parser.
Use an isolated environment first
Do not install a PDF parsing stack into a Python environment that already contains unrelated ML packages. An isolated environment turns dependency failures into something you can reason about:
uv venv ~/venv/mineru --python 3.13
uv pip install \
--python ~/venv/mineru/bin/python \
"mineru[pipeline]"
One tested installation produced:
mineru 3.4.5
along with its pipeline dependencies. The important part is the extra: [pipeline]. Installing only the base package can leave the parser present while optional inference dependencies are missing. That can produce errors such as:
ModuleNotFoundError: No module named 'transformers'
or:
No module named 'torchvision'
The failure is easy to misdiagnose because the main package appears to be installed correctly.
Download only the models you intend to use
The model downloader can fetch pipeline models, VLM models, or both. If you only need pipeline parsing, specify that explicitly:
~/venv/mineru/bin/mineru-models-download \
-s huggingface \
-m pipeline
Do not rely on interactive defaults when disk usage matters. A tested version of the downloader used:
model_type = click.prompt(
"Please select the model type to download: ",
type=click.Choice(['pipeline', 'vlm', 'all']),
default='all'
)
Pressing Enter without choosing a model type therefore selected:
all
which also pulled the VLM weights. For a machine intended only for pipeline parsing, that is unnecessary storage and another source of confusion.
A misleading dependency error can hide the real missing module
One particularly nasty failure occurred while trying the hybrid HTTP backend. The command failed with:
mineru.cli.common.HybridDependencyError:
`hybrid-http-client` requires local pipeline dependencies
(`mineru[pipeline]`, including `torch`).
Install `mineru[pipeline]` or `mineru[core]`.
The problem was that mineru[pipeline] was already installed. The useful clue came from inspecting the import path directly. The underlying exception was:
ModuleNotFoundError: No module named 'six'
Installing the missing package fixed that stage:
uv pip install \
--python ~/venv/mineru/bin/python \
six
The broader lesson is important: when a wrapper catches both ImportError and ModuleNotFoundError and converts them into one generic dependency message, the visible exception may describe the wrapper, not the real missing package. When a dependency error contradicts your environment, import the failing module directly before reinstalling the entire stack.
"Peer closed connection" may actually be a server queue problem
After fixing the missing module, the next attempt successfully connected to the local OpenAI-compatible inference server. Then it stalled and ended with:
httpx.RemoteProtocolError:
peer closed connection without sending complete message body
(incomplete chunked read)
That looks like a transport error. It was not. The server log showed:
omlx.exceptions.SchedulerQueueFullError:
Scheduler waiting queue full: 32 >= 32
The client was capable of issuing a large burst of concurrent inference requests. The local server, meanwhile, had been configured with:
max_concurrent_requests = 8
and its waiting queue worked out to 32. Temporarily increasing the server concurrency to 32 expanded the queue and allowed the parsing flow to finish. But that did not make it a good long-term architecture. Changing a carefully tuned inference server just to accommodate one document parser created more operational cost than benefit. The final choice was to stop bridging through that server and use the parser's native inference path instead.
This is a useful PDF-to-Markdown debugging rule: If an HTTP backend dies after requests have already started, inspect the inference server logs before changing timeouts, TLS settings, or retry counts.
A model can load successfully and still be the wrong model
The next attempt avoided the HTTP bridge and pointed the local parser at an already-downloaded MLX model. It failed with:
Missing 391 parameters:
vision_tower.blocks.0.attn.proj.bias,
vision_tower.blocks.0.attn.proj.weight,
...
vision_tower.patch_embed.proj.weight.
The model was not simply "corrupted." The converted model contained the language-model side but did not contain the complete vision tower required for document parsing. The safer fix was to download the parser's own complete VLM weights:
~/venv/mineru/bin/mineru-models-download \
-s huggingface \
-m vlm
There was another trap. Because the configuration already contained a path under:
models-dir.vlm
the downloader treated an existing configured model directory as reusable and skipped the download. The path had to be cleared before the downloader would actually fetch the complete model. That kind of failure is easy to miss because:
- the path exists;
- the model directory is non-empty;
- some model files load;
- and the actual failure appears much later as hundreds of missing vision parameters.
An existing model directory is not proof that the model is complete for the selected backend.
A successful native parse
After downloading the complete model, the native hybrid path finally completed the PDF:
Layout Predict: 100%|██████████| 13/13
Predict: 100%|██████████| 277/277
OCR-det: 100%|██████████| 155/155
Processing pages: 100%|██████████| 13/13
Completed batch 1/1 | Processed 13/13 pages
The resulting Markdown preserved heading structure, converted inline formulas to $...$, converted display formulas to $$...$$, and extracted images with relative references. A comparison with pipeline mode also exposed why hybrid parsing can matter. Pipeline output contained OCR noise such as:
distributed evęnts ... their reàl time precedence
while the enhanced output restored:
distributed events ... their real time precedence
Another fragment changed from broken LaTeX:
$\mathrm { o f } \ ^ { \mathrm { \left. } }
\mathrm { f a l s e } ^ { \mathrm { \right. } }$
to ordinary text:
of "false"
For documents that already contain clean embedded text, pipeline mode may be enough. For damaged OCR, scans, formulas, or complex layouts, the extra model stage can materially improve the Markdown.
Old magic-pdf installations have a completely different class of traps
If you inherit an older magic-pdf environment, do not assume commands from one tutorial apply to your installed version. One real environment ran:
magic-pdf -p example.pdf -o output_dir -m auto
and got:
Error: No such option: -p
Checking help showed an older command structure:
magic-pdf, version 0.6.1
with commands such as pdf-command, json-command, local-json-command. The PDF call looked like:
magic-pdf pdf-command \
--pdf "pdf_path" \
--inside_model true
Then another environment, after switching to Python 3.10.16 and reinstalling, reported:
magic-pdf, version 1.1.0
and the CLI had changed again:
-p, --path
-o, --output-dir
-m, --method
So a command-line error may be a version mismatch, not a typo. Always run:
magic-pdf --version
magic-pdf --help
before copying a command from an older guide.
detectron2 can reveal a Python-version mismatch
The same older setup next failed because detectron2 was missing. The attempted installation returned:
No matching distribution found for detectron2
The workaround was not another pip flag. The environment was moved to Python 3.10.16, where the required build was available. After that, the parser still reported missing modules one by one:
No module named 'paddle'
No module named 'cv2'
No module named 'ultralytics'
No module named 'doclayout_yolo'
No module named 'timm'
No module named 'unimernet'
No module named 'paddleocr'
No module named 'rapid_table'
No module named 'struct_eqtable'
No module named 'rapidocr_onnxruntime'
This is a signal to stop treating each missing import as an independent problem. When a parser throws a sequence of optional-module errors, first verify that you installed the correct dependency extra for that exact version. Otherwise you can spend an hour reconstructing the package's dependency set manually.
The fitz package is not the same thing as PyMuPDF
Another installation used:
pip install "magic-pdf[full]==0.7.0b1"
and immediately failed during startup:
AttributeError: module 'fitz' has no attribute 'Document'
The environment contained a conflicting fitz installation. The recovery path was to remove it and install PyMuPDF instead. When the first reinstall still did not resolve the environment cleanly, PyMuPDF was removed and installed again at the latest available version. This distinction matters because Python imports PyMuPDF through:
import fitz
so an unrelated package named fitz can satisfy the import while exposing the wrong API. If you see:
module 'fitz' has no attribute 'Document'
check the installed package identity before changing your parser code.
RapidTable changed an argument name across versions
A separate table-recognition failure looked like this:
TypeError:
RapidTableInput.__init__() got an unexpected keyword argument 'model_path'
The relevant constructor call was effectively:
RapidTableInput(
model_type=table_sub_model_name,
model_path=slanet_plus_model_path
)
A later API used:
model_dir_or_path
instead. That means a parser package and its table-recognition dependency can both install successfully while still disagreeing at runtime about a constructor signature. A historical working environment recorded:
magic-pdf 1.3.12
rapid-table 1.0.5
rapidocr 3.2.0
Treat that exact trio as a forensic record, not a universal pin for a fresh installation. Dependency contracts evolve, and current package compatibility rules may differ from the environment that originally produced the successful run. The durable debugging technique is to inspect the installed constructor signature and compare it with the caller.
Illegal instruction (core dumped) is not a normal Python exception
During another old magic-pdf run, the process exited with:
Illegal instruction (core dumped)
That matters because the interpreter did not raise a normal Python exception. The process executed a CPU instruction the host could not use. With native numerical libraries, OCR runtimes, inference engines, and prebuilt wheels involved, chasing Python stack traces is pointless once the process dies this way. Check the CPU instruction-set support and the native wheels you installed. For older servers in particular, AVX/AVX2 support can be the dividing line between a package importing normally and the process exiting immediately.
A practical PDF-to-Markdown decision path
For simple text PDFs:
pdftotext -layout input.pdf -
Use it when you only need readable text and the document is structurally simple.
For complex digital PDFs:
MINERU_MODEL_SOURCE=local \
mineru -p input.pdf -o output -b pipeline
Use pipeline mode first. It is easier to deploy and avoids unnecessary VLM complexity.
For scans or documents where pipeline output contains OCR noise, formula corruption, or serious layout mistakes:
MINERU_MODEL_SOURCE=local \
mineru -p input.pdf -o output -b hybrid-engine
If it fails, debug in this order:
- Verify the exact parser version.
- Verify the exact Python environment.
- Run the installed CLI's own
--help. - Verify the correct dependency extra is installed.
- Check the real import exception behind wrapper errors.
- Verify model directories contain the models required by the selected backend.
- Inspect the inference server logs if using an HTTP backend.
- Only then start changing parser code.
The biggest time sink in PDF-to-Markdown work is usually not the PDF. It is assuming that an installed package, an existing model directory, a successful import, or a copied command proves the whole parsing stack is compatible. It does not.