Portimg Insights

The Ultimate Guide to Extract Text from PDF Using OCR

Date
Read time5 min read
The Ultimate Guide to Extract Text from PDF Using OCR

Learn how to extract text from PDF using OCR easily. Follow this step-by-step guide and try it for free at Portimg.com!

Scanned PDFs present a specific problem: they look like documents but behave like images. The text is visible on screen but not selectable — you can read it but you can't copy it, search within it, or edit it. This happens because the file contains a photograph of a page rather than the text itself. OCR converts that photograph back into real, editable characters.

Here's how to extract text from a scanned PDF using Portimg's OCR tool, and what determines how clean the output will be.

Why scanned PDFs are different

Not all PDFs are the same. A PDF exported from Word or created by a modern printer contains actual text data — you can select it, copy it, and search it immediately. A scanned PDF is different: it's typically a photograph of a printed page, embedded inside a PDF container. The file format says "PDF" but the content is an image.

This is common with older documents, signed contracts that were printed and scanned back in, forms filled out by hand, and anything that originated on paper. OCR is the only way to get editable text out of these files without retyping everything manually.

The extraction process

Go to portimg.com/image-to-text-ocr and upload your file. For scanned PDFs, the most reliable workflow is to convert the PDF pages to images first using Portimg's PDF to image tool, then upload the resulting images to the OCR tool. This gives you control over each page and produces cleaner results than feeding a multi-page PDF directly.

Once uploaded, the tool processes the image and returns extracted text in a results panel within seconds. You can copy it directly or download it as a text file. No account is needed, and files are deleted from the server immediately after processing.

What affects the output quality

The accuracy of the extracted text depends almost entirely on the quality of the scan. A crisp, well-lit scan of typed text on white paper will convert with very high accuracy — perhaps one or two corrections needed per page. A low-resolution photograph of a faded document taken at an angle in poor light will produce output that needs significant review. The OCR engine is doing its best with whatever pixels it receives.

Resolution is the baseline. 300 DPI is the minimum for reliable results on standard body text. Below that, the fine details of letterforms start to disappear at the pixel level, and visually similar characters — '1' and 'l', '0' and 'O', 'rn' and 'm' — become ambiguous. If you're scanning documents specifically to OCR them, set your scanner to at least 300 DPI; 600 DPI for anything with small print or degraded paper.

Even lighting eliminates the most common errors. A shadow crossing part of the page drops local contrast and increases substitution errors in that region. Flat, diffuse light — natural light from the side, or two light sources positioned symmetrically — keeps contrast consistent across the whole page. Avoid flash, which creates a bright centre and dark edges.

Straight alignment matters on multi-line documents. Text running at an angle is harder for OCR to segment into lines correctly. If a scanned page is visibly tilted, rotate it before uploading. Most phone camera apps and basic image editors have a straighten function.

Grayscale can help on colour scans. If your document was scanned in colour and has a non-white background — yellowed paper, coloured forms — converting to grayscale before OCR sometimes improves contrast and reduces background noise the model would otherwise have to ignore.

Multi-page documents

For documents with multiple pages, convert each page to a separate image, run them through OCR individually, and combine the text outputs in order. This takes slightly more steps than uploading a single file, but gives you the ability to check and correct each page before moving on — useful for long documents where a problem on one page shouldn't hold up the rest.

Once you have the extracted text, it behaves like any other text file: paste it into a Word document, a Google Doc, a spreadsheet, or a translation tool. You can search within it, reformat it, run it through spell check, and edit it as needed. For documents that also need to be merged or reorganised after extraction, Portimg's PDF merge tool handles combining multiple files into one.

When OCR is and isn't the right tool

If your PDF was created digitally — exported from Word, generated by software, saved from a website — the text is already in the file and you don't need OCR. Try selecting text in the PDF first. If it selects cleanly, copy it directly. OCR is specifically for cases where the PDF contains images of text rather than text itself.

For handwritten content, accuracy depends heavily on the clarity of the handwriting. Clear, consistently-formed printing comes through well. Informal cursive or hurried notes will need more correction. Test a sample page before committing to a large batch.

If you have a scanned PDF you've been unable to extract text from, running a page through the tool takes under a minute and will quickly show you what to expect from your specific document.

Image ToolsPDFTutorial
Found this helpful?

Try Portimg's free tools

Convert, compress, and edit images and PDFs — no sign-up needed.

Explore tools