Understanding PDFs

Scanned PDFs vs text-based PDFs

Two PDFs can look identical on screen and behave completely differently. One contains real text; the other is a photograph of text. Knowing which you have explains most "why did it do that?" moments with PDF tools.

Three kinds of PDF

  • Text-based ("born-digital") — created by software: exported from a word processor, saved from a browser, generated by a billing system. The letters are stored as text in a font.
  • Scanned (image-based) — created by a scanner or phone camera. Each page is a single picture. The words are just pixels.
  • Scanned with OCR — a scan that has had text recognition applied. It's still a picture, but an invisible text layer sits behind it so you can search and select.

Many real documents mix these — a text report with a scanned signature page, for example.

How to tell which one you have

  1. Try to select a word. Drag across a line of text in your PDF reader. If individual words highlight, there's text. If you get a rectangle over the whole page (or nothing), it's an image.
  2. Search for a word you can see. Press Ctrl+F (Cmd+F on Mac). No results for a visible word means no text layer.
  3. Zoom in to 400%. Text-based pages stay razor-sharp at any zoom. Scanned pages go soft or pixelated.
  4. Check the size per page. Text pages are often a few to tens of kilobytes; scanned pages are commonly hundreds of kilobytes or more.

Why it matters for each tool

TaskText-based PDFScanned PDF
CompressUse Standard. Extra usually makes it bigger and removes selectable text.Extra often shrinks it a lot. Check small print afterwards.
Merge, split, extract, deleteWork normally; text stays selectable.Work normally; pages stay images.
RotateLossless.Lossless — ideal for fixing scanner orientation.
PDF to JPGWorks; text becomes part of the image.Works; the image quality is capped by the original scan.
Copy, search, screen readersYes.No — unless OCR has been applied.

Making a scanned PDF searchable

Turning pictures of words into real text requires OCR (optical character recognition). The tools on this site don't perform OCR. Many scanner apps, desktop PDF editors and office suites include it. After OCR, the page looks the same, but you can search, copy text, and screen readers can read it aloud — a significant accessibility improvement.

Note that OCR is recognition, not a copy: it can misread characters, especially in poor scans, handwriting or unusual fonts. Check important numbers.

Scanning tips for better image-based PDFs

  • 200–300 dpi is right for most documents; higher resolutions mainly add size.
  • Use greyscale or black-and-white for plain text documents; colour only where colour matters.
  • Clean the scanner glass — specks and lines repeat on every page.
  • Feed pages straight; tilted scans can't be fixed by 90° rotation.

Questions and answers

Not directly — the words are part of an image. You'd need OCR software to convert them into text first.

Image-based compression lowers the resolution and quality of the page images. Keep the original if fine detail matters, or use gentler options.