Why PDFs Get So Big
Text is almost never the problem. A three-hundred-page report of plain text compresses to a few hundred kilobytes, because the text in a PDF is stored as glyph descriptions that shrink extremely well. The weight comes from images: photographs, exported illustrations, and scanned pages. A single scan saved at 300 dpi can easily weigh two to five megabytes, so a hundred scanned pages turn a simple memo into a half-gigabyte file that no mail server will accept.
Two quieter causes add up over time. First, incremental saves: every time a PDF is edited, the old object versions can remain in the file alongside the new ones, so the document carries its own history. Second, embedded fonts — a file may ship an entire typeface family when only one weight is used, and re-exported documents often accumulate duplicate fonts from every source they were assembled from.
Downsample Images Strategically
Every image has a pixel size and a placement size, and the ratio between them decides the file weight. A photo placed four inches wide needs about 1,200 pixels across for crisp 300 dpi print output, so keeping a 4,000-pixel scan in that slot stores sixteen times the data no reader will ever see.
Match resolution to the destination. For files that will be read on a screen or attached to an email, 96 to 150 ppi is plenty; for print, hold at 300 ppi. Also match the format to the content: photographs compress well as JPEG, while line art, diagrams, and screenshots full of text should stay lossless or vector, where compression costs nothing in quality.
Strip What the Document Does Not Need
Font subsetting embeds only the characters a document actually uses instead of the whole typeface. Most exporters do this automatically, but documents assembled from many sources may still carry several complete fonts, some of them never rendered on a single page. Removing them is invisible to readers and measurable in kilobytes.
Metadata hides in the same way: scanner bookmarks, preview thumbnails, annotation history, and leftover form data all live inside the file. Re-saving a cleaned copy rebuilds the object tree from scratch, dropping everything no longer referenced — including the revisions that incremental saves left behind.
Five Checks That Save the Most
Compression is always a trade between size and fidelity, so work through the checks in order. The first four cost you nothing visually; the last one is the guard rail that keeps the file honest.
- Downsample images to the resolution the document needs, not the resolution the scanner or camera produced.
- Run a lossless pass first — Flate compression on existing streams shrinks files without touching a single pixel.
- Subset fonts and drop unused ones so each typeface is embedded once, and only as far as it is used.
- Re-save the document as a fresh copy to discard incremental revisions, orphaned objects, and stale thumbnails.
- Compare the result at 100% zoom before sending it: a smaller file is only a win if text, diagrams, and scans still read correctly.
Browse more articles about working with PDF files.
Back to the Blog