How to search scanned PDFs by what’s inside

A scanned PDF is a picture of a page, so Ctrl-F finds nothing. This guide covers making scans searchable by their contents: the tools, the tradeoffs, and how to ask real questions instead of guessing keywords.

If searching your scanned documents turns up nothing, it isn’t you. A scan is an image, and an image has no text to match. Fixing it takes three moves: add a text layer, search across the whole pile at once, and then the part most guides skip: asking real questions instead of hunting keywords.

The short version

A scanned PDF is an image of a page with no underlying text, so search has nothing to match. To make it searchable you run OCR (optical character recognition), which adds an invisible text layer behind the image. After OCR, Ctrl-F works, text is selectable, and with the right tool you can search across thousands of scans at once and even ask questions of them in plain English.

How to tell if you have a scanned PDF

Open the file and try to select a line of text. Three outcomes tell you what you’re dealing with: if the cursor grabs nothing and just draws a box over blank space, it’s a scan and needs OCR. If you can select text but it pastes as gibberish, it has bad OCR and should be re-run. If text selects cleanly and pastes correctly, it’s already a digital PDF and needs nothing.

Two more checks settle the ambiguous cases. Open the file's properties or document info: scanners and multifunction copiers usually stamp their own name into the producer field, while a PDF exported from Word or a browser names the software that made it. And look at the page under magnification; scanned text shows the soft, uneven edges of an image, where digital text stays crisp at any zoom.

Mixed files are common and easy to misjudge. A twelve-page agreement can be digital for eleven pages and scanned for the signature page, and a report can carry searchable body text with scanned exhibits attached at the back. Searching finds the digital pages and silently misses the rest, which is how people conclude a document does not mention something it plainly does.

What OCR actually does

OCR reads the image, recognizes the characters, and writes them into an invisible text layer positioned under the visible page. The document looks identical; the difference only shows up when you search, select, or copy. Good tools place that text accurately beneath each word, so highlighting and copy-paste line up with what you see on the page.

Accuracy depends on inputs more than on brand. Scans at 300 DPI in grayscale or black and white, straight on the glass with the page flat, produce clean results from almost any engine. Photographs taken at an angle, low-contrast faxes, tenth-generation copies, and pages with heavy background shading are where errors cluster, and no amount of software choice fully rescues a bad capture.

It is worth knowing what OCR does not do. It does not understand the document, so it cannot tell you that a clause is unusual or that a number contradicts another page. It converts pictures of words into words, and everything after that, searching, extraction, or answering questions, is a separate layer working on the text OCR produced.

Make one scan searchable, by tool

Adobe Acrobat Pro

Open the PDF, go to the Scan & OCR tools, and choose Recognize Text. Acrobat embeds a searchable text layer while preserving the page image, and it processes files locally, which matters for sensitive documents. Downsides: it needs a paid Acrobat Pro subscription (roughly $15–25/month) and can be slow on large files.

OCRmyPDF (free, open-source)

OCRmyPDF wraps the Tesseract engine and adds a text layer in a single command, offline. It’s the best option for batch work and automation: ocrmypdf input.pdf output.pdf, with flags like --deskew --clean to straighten and de-noise pages and -l eng+fra to set languages. It outputs searchable PDF/A and keeps the original image resolution. The catch is a command line and a few minutes of setup.

macOS Live Text

Recent macOS versions recognize text in images and PDFs with Live Text, so you can select and copy from a scan on the fly. Handy in a pinch, but it doesn’t write a text layer back into the file, so the PDF itself stays unsearchable. Treat it as a viewer feature, not a way to convert an archive.

Google Drive / Docs

Opening a scanned PDF with Google Docs runs OCR and drops the recognized text into a new Doc. That’s useful for grabbing text out of one file, but it doesn’t give you a searchable PDF: you lose the original layout, and it doesn’t scale to an archive.

Online OCR tools

Upload-and-download services turn a scan into a searchable PDF in about a minute with no install. Fine for a one-off, non-sensitive document, but think twice before uploading contracts, financials, or anything confidential to a free web tool, since you’re handing the file to a third party.

Whatever tool you use, keep the searchable version and discard the original only deliberately. Most tools write the text layer into the same PDF, so you end up with one file that looks identical and now searches, which is the outcome you want. Where a tool produces a separate text or Word file instead, you have created a second document that will drift from the first, and in a year nobody will know which one governs.

Batch is a different problem from single files. Every tool here works one document at a time, so a folder of two hundred scans becomes two hundred operations plus the clicks between them. That is the practical ceiling of the manual approach, and it is where people either give up or go looking for something that processes an archive without being asked twice.

Which tool should you use?

The honest answer depends on volume and sensitivity more than on features. For one document you already have a tool: Preview on a Mac and most scanner software both run OCR, and neither costs anything extra. For a few dozen documents, a paid desktop tool that batches is worth the money once. For an archive, per-document tools stop being the answer at all, whatever their per-file quality.

Sensitivity narrows it further. Free web converters are convenient and they are also an upload to a company you have no relationship with, which is a fine trade for a scanned recipe and a poor one for a lease, a tax return, or anything covered by a confidentiality obligation. When the document matters, keep the processing local or use a service whose terms you have actually read.

ToolCostBest forWatch out for
Acrobat ProPaid (~$15–25/mo)A few files, locally, with a licenseCost; slow on big files
OCRmyPDFFreeBatch and automation, offlineCommand line; setup
macOS Live TextFreeGrabbing text on screenDoesn’t make the PDF searchable
Google DocsFreePulling text from one fileNo searchable PDF; loses layout
Online OCRFree / freemiumA one-off, non-sensitive filePrivacy: you upload the file

The real problem: searching across hundreds of scans

OCR-ing one file is easy. The pain starts when the answer could be in any of hundreds of scanned invoices, contracts, or reports and you don’t know which. Per-file OCR doesn’t solve that; you need everything OCR’d and searchable in one place. Options range from OCR-ing a folder in bulk (OCRmyPDF handles batches) and letting your operating system index it, to a document tool that OCRs on import and searches the whole archive together. The goal is one search box over every document, not a folder-by-folder hunt.

Two properties matter more than raw speed once you cross a few hundred documents. Search has to cover the contents of every file, including the ones nobody remembered to OCR, since a single unprocessed scan is invisible and you will never learn it was missed. And results have to point at the page, because finding the right document inside a two-hundred-page file only moves the problem.

The structure question settles itself. Folder trees built for filing are rarely built for finding, and reorganizing an archive to make search work is effort spent at the wrong end of the problem. Leave the structure alone and put a search over it: the folders keep their meaning for the people who built them, and the search answers everyone else.

Beyond keywords: asking questions of your scans

Even perfect search only finds words. “What’s the total on this invoice?” isn’t a keyword, and neither is “which of these leases allows pets.” Tools that read the OCR’d text with AI let you ask in plain English and get the answer with the page it came from, so you can verify it. That’s the layer DocuStrata adds: it OCRs scans automatically on import and lets you ask across the entire archive, every answer citing its source. And unlike a free web converter, your documents are never used to train a model.

The difference shows up in the questions you can ask. Keyword search answers where a word appears; asking the archive answers what a document says, which is what someone actually wants when they type a question about a renewal date or a deductible. The useful version cites its source, because an answer you cannot trace has to be verified by hand, which returns you to reading.

This is also where scanned material stops being second-class. Once the text layer exists, a photographed receipt and a native invoice behave identically, and the archive stops being divided into the parts you can query and the parts you have to remember to open.

Getting OCR results you can trust

Accuracy is mostly set before OCR even runs, by the scan itself.

Spot-check instead of trusting or despairing. Open three or four processed documents, search for a word you know is present, and confirm it is found where you expect. If a batch shows systematic errors, the fix is almost always at capture: rescan the worst originals instead of reprocessing bad images repeatedly.

Keep the originals. A scanned page is a photograph of a document, the text layer is an interpretation of that photograph, and interpretations improve with time and better engines. Discarding the image to save space forecloses ever redoing the work, and storage is the cheapest part of the exercise.

Frequently asked questions

Why can’t I search my scanned PDF with Ctrl+F?

Because the PDF contains a picture of text, not text. Ctrl+F searches the text layer, and a scan doesn’t have one until you run OCR. Once OCR adds that layer, Ctrl+F behaves normally, and the page looks unchanged because the layer sits invisibly behind the image.

How do I make a scanned PDF searchable without Adobe Acrobat?

OCRmyPDF does it free from the command line and processes files locally. Google Drive can extract text through Google Docs, though it won’t write a text layer back into your PDF. An AI document tool like DocuStrata runs OCR automatically on import, so every scan you add becomes searchable without a separate step.

Does OCR work on handwriting?

Printed text OCRs well; handwriting is much harder. Modern AI models read clear, consistent handwriting with fair accuracy, but cursive, faint pencil, and messy field notes still miss often. Treat OCR’d handwriting as a finding aid and verify anything important against the page image.

Can OCR read documents in other languages?

Yes. Good OCR engines handle dozens of languages, and accuracy in major European languages is close to English. Set or confirm the document language where the tool allows it: recognition against the wrong language model produces confident nonsense, especially around accented characters.

Can you search a scanned PDF without OCR?

No. A scanned PDF is an image with no text layer, so search has nothing to match. You have to run OCR first to add a searchable text layer; after that, normal search works.

Does Google Drive make a scanned PDF searchable?

Not exactly. Opening a scan with Google Docs extracts the text into a new document, but it does not produce a searchable PDF and it loses the original formatting. It is fine for pulling text out of one file, not for keeping a searchable archive.

How do I search hundreds of scanned PDFs at once?

OCR them all, then keep them somewhere with one search over the whole set: a bulk OCR pass plus your operating system’s index, or a document tool that OCRs on import and searches the entire archive together. Per-file OCR alone will not scale.

Is OCR accurate enough to trust?

For clean, printed text scanned at 300 DPI, modern OCR is very accurate. Accuracy drops with low-resolution scans, handwriting, stamps, and multi-column layouts, so always verify critical figures against the original page.

Is it safe to use a free online OCR tool for sensitive documents?

Be careful. Free online tools require uploading your file to a third party. For contracts, financial records, or anything confidential, use a local tool like Acrobat or OCRmyPDF, or a service that clearly states it does not retain or train on your documents.

Make your scans answerable

DocuStrata OCRs scans on import and lets you ask across the whole archive, with the source behind every answer, and nothing is ever used to train a model.

See how it works Request admission