Learn / Documents and data
How do you parse scanned PDFs and tables for RAG?
Short answer
Parsing turns files into clean, structured text before they are indexed. Scanned pages need OCR, which reads the pixels, and tables need layout analysis so that rows and columns survive instead of being flattened into a stream of numbers.
With troveGEN
troveGEN checks every upload before choosing how to read it, reads scans and drawn-text PDFs with OCR, keeps tables as tables, and reports a file it cannot read instead of indexing nothing.
Not every PDF is the same
- Text PDFs have a real text layer that can be extracted directly.
- Scanned PDFs are pictures of pages with no text at all; they need OCR.
- Drawn-text PDFs look like text on screen but store the letters as vector shapes, so extraction returns nothing and the page must be rendered and read like a scan.
- Mixed PDFs combine all of the above in one file.
What OCR does and where it struggles
Optical character recognition converts an image of text into characters. Accuracy falls with low resolution, skew, stamps over text, handwriting and unusual scripts. Language support matters: a model without the right script data will return noise. Page images should be rendered at a sufficient resolution before recognition.
Tables are the hard part
A table is only useful if the link between a value and its row and column labels survives. Flatten it to text and "1,200" no longer means anything. Good parsing detects the table, keeps its structure, and stores it in a form that can be searched row by row and, where numbers are involved, calculated over exactly.
Quality checks to insist on
- Fail loudly: a scan that yields no text should be reported, not indexed as an empty document.
- Detect garbled text layers and fall back to OCR.
- Keep headings, page numbers and reading order so passages can be cited.
- Record how each document was parsed so problems can be traced.
Key takeaways
- Identify what kind of PDF you have before choosing how to read it.
- Preserve table structure; flattened tables are unusable.
- A document that cannot be read should be reported, not silently indexed.
How troveGEN helps with OCR and document parsing
troveGEN checks each upload before choosing how to read it, uses fast local parsers for clean text and a heavier layout engine only when needed, renders pages that are drawn or scanned and reads them with OCR in the language you choose, and keeps tables as tables. If a file cannot be read, you are told why instead of getting an empty document.
What troveGEN provides
- A pre-scan that picks the lightest reliable way to read each file
- OCR in the language you set, including pages drawn as shapes
- Table structure kept for search and exact calculation
- Clear errors when a document is unreadable
- A step-by-step timeline for every document
Frequently asked questions
Is OCR always needed for PDFs?
No. If a PDF has a real text layer, direct extraction is faster and more accurate. OCR is for pages that are images or drawn text.
Can handwriting be read?
Standard OCR is built for printed text. Handwriting needs specialised models and is less reliable.
Why did my table come out as one line of numbers?
The parser did not detect the table structure. Layout-aware parsing keeps rows and columns together.
How does troveGEN help with OCR and document parsing?
troveGEN checks every upload before choosing how to read it, reads scans and drawn-text PDFs with OCR, keeps tables as tables, and reports a file it cannot read instead of indexing nothing. It provides: A pre-scan that picks the lightest reliable way to read each file; OCR in the language you set, including pages drawn as shapes; Table structure kept for search and exact calculation; Clear errors when a document is unreadable; A step-by-step timeline for every document.