A scanned document looks like a PDF but doesn't behave like one. You can open it, zoom in, and read it, but you can't search for a word, select a sentence, or copy a paragraph. The file contains a photograph of text rather than text itself — and until you run it through OCR, that's all it will ever be.
Converting a scanned image to a searchable PDF solves this. The OCR process analyses the image, recognises the text, and embeds it into the PDF so that the words become selectable and searchable while the original visual appearance of the document is preserved. The result looks identical to the scan but behaves like a real document.
When this matters most
The practical value varies depending on what you're doing with the document. For a single receipt you're filing away, searchability might not matter much. For a contract you'll need to reference repeatedly, being able to search for a clause by keyword rather than reading through every page saves real time. For an archive of hundreds of scanned pages — legal records, medical files, historical documents, old correspondence — searchable PDFs are the difference between a usable digital archive and a folder of images you can only browse manually.
Accessibility is another case worth noting. Screen readers used by visually impaired users work on text, not images. A scanned PDF that hasn't been processed with OCR is effectively invisible to a screen reader. Converting to searchable PDF makes the content accessible to assistive technology.
How to convert with Portimg
Go to portimg.com/image-to-text-ocr, upload your scanned image — JPG, PNG, or WebP — and the tool extracts the text via OCR. For a searchable PDF specifically, the workflow is to extract the text, then use the image to PDF tool to combine your original scan with the extracted text layer into a single document.
For documents that already exist as scanned PDFs, convert the pages to images first using the PDF to image tool, then run the resulting images through OCR. No account is required at any step, and files are deleted from the server immediately after processing.
Getting clean OCR output
The quality of the searchable text layer depends on the quality of the scan. A clear, high-contrast scan of typed text will produce accurate, reliable output. A dark, blurry, or skewed photograph will introduce errors that make the search layer less useful — you might search for "invoice" and miss results because the OCR read "lnvoice" with a lowercase L instead of an I.
A few things consistently improve results. Sharp focus matters more than anything else — a slightly lower resolution image that's in focus will outperform a high-resolution blurry one. Even, flat lighting eliminates the shadow gradients that reduce local contrast and cause substitution errors. If the scan is visibly tilted, straighten it before uploading; OCR segments text into horizontal lines, and a skewed page produces less accurate line breaks. Dark ink on white or near-white paper gives the model the clearest signal to work from.
For typed documents in good condition, you can expect very high accuracy — typically few enough errors that the searchable layer is fully reliable. For older documents with faded ink, unusual fonts, or damaged paper, accuracy will be lower and the output may need review before you rely on the search functionality.
Handwriting
OCR on handwritten content produces more variable results than on printed text, and a "searchable" PDF from a handwritten scan should be treated as an approximation rather than a reliable index. Clear, consistently-formed printing in good lighting often comes through well enough to be useful. Casual cursive or hurried notes will produce a text layer that may miss or misread a significant portion of words. For handwritten archives where searchability matters, it's worth testing a representative sample before processing the full collection.
Multi-page documents
For documents with multiple pages, process each page as a separate image and combine the results. Run each image through OCR, then use the PDF merge tool to assemble the pages into a single searchable document in the correct order. This approach also lets you review and correct each page's OCR output before finalising the document, which is worth doing for anything you'll rely on heavily.
If you have a stack of documents to process, the consistent workflow — scan at 300 DPI minimum, even lighting, straight alignment, then OCR — produces reliable results across different document types and saves correction time later.
