CleanPDFPages
CORE CAPABILITIES & ARCHITECTURE

Engineered for Accurate, Private PDF Cleanup

CleanPDFPages is purpose-built to eliminate blank and duplicate pages from PDF batches without sending a single byte of your data over the network.

Bulk PDF Analysis

Analyze up to 20 documents simultaneously

Upload and analyze entire document sets in a single pass. Process up to 20 PDF files (up to 500 MB combined batch size) with non-blocking, asynchronous rendering.

Blank Page Detection

Smart margin-aware luminance analysis

Detect blank pages in both digital exports and scanned documents. Features automated margin-masking that ignores scanner shadows, feed roller marks, and binding hole punches.

Duplicate Page Detection

Dual text signature & perceptual dHash classifier

Combines bigram Dice similarity for digital text with 128-bit 2D visual difference hashing and 32×32 Pearson correlation to detect identical copies and separate recurring form templates.

Visual Review Board

High-resolution side-by-side inspection

Examine any flagged page in full resolution before making changes. Side-by-side comparison displays original versus duplicate pages with keyboard navigation (arrow keys, Esc).

100% Local Processing

Client-side WebAssembly & zero cloud uploads

Every byte of your PDF documents stays inside your web browser. Page rendering, classification, and output assembly happen entirely on your CPU via client-side JavaScript.

Safe & Non-Destructive Cleanup

Original files remain untouched on your machine

CleanPDFPages never modifies your source files or automatically deletes pages without review. Export freshly sanitized documents with retained vector fonts, bookmarks, and color spaces.

Detection Engine

How CleanPDFPages Analyzes Your Documents

STAGE 1

Text Extraction

Extracts text stream tokens, calculates lexical density, and generates bigram signature vectors for instant digital duplicate comparison.

STAGE 2

Luminance & dHash

Renders page canvases off-screen to compute standard deviation of luminance (for blank pages) and 128-bit gradient difference hashes.

STAGE 3

Pearson Correlation

Cross-references low-Hamming-distance candidates with 32×32 grayscale intensity matrices to distinguish exact duplicates from recurring form templates.

Test CleanPDFPages on your documents

100% free, runs entirely in your browser, with no file size limits or subscriptions. Have questions? Read our FAQ guide.

Clean PDF Online Free