Extract Text from a PDF

Get the words out of a PDF as plain text, page by page, without copy and paste. Handy for quoting, searching or scripting.

Files are never uploaded

How to use it

Add your PDF

Drop the document onto the box above, or click to browse. The text layer of each page is read in turn.

Choose page markers

Keep the headers on when you plan to search or quote across a long document, and turn them off when you want a clean block of text.

Extract

Press Extract Text. The progress bar names the page being read, which is useful on long documents.

Check the text

Open the downloaded file and skim a couple of pages. Reading order and spacing are approximate, so tables and multi-column layouts need a quick look.

Save the text file

Download the plain UTF-8 text file and open it in any editor, spreadsheet or script.

Common issues

Expecting OCR

If the page is a scan, the extracted text is empty or nearly empty. That is not a fault in the tool; there is simply no text layer to read.

Tables coming out scrambled

Rows and columns flatten into a word stream. Treat the output as a starting point and clean it up before it goes into a spreadsheet.

Right-to-left and non-Latin scripts

Reading order for Arabic, Hebrew and some Asian scripts can come out reversed or broken, because a PDF stores glyphs in drawing order rather than reading order.

Very large documents taking a while

A several hundred page file produces a big text file and takes time to read page by page. The progress bar shows how far it has reached.

Frequently asked questions

Does this work on a scanned PDF?
No. A scanned page is an image with no text layer, so there is nothing to read. Those files need OCR, which this tool does not do.
Does it keep the original layout?
Only roughly. Text is returned in reading order with line breaks where the PDF marks them. Columns and tables come out as a stream of words rather than a grid.
What encoding is the file in?
UTF-8, which every modern editor and spreadsheet opens correctly.
Can I copy the text without downloading?
Open the downloaded text file and copy from there. This tool writes a single plain file so the result is reproducible.
Why are some characters wrong?
A PDF can embed unusual fonts and ligatures. The extractor does its best, but ligatures and symbol fonts sometimes arrive as the wrong character.
Do hyperlinks survive?
No. You get the visible link text, not the address behind it.
Can I extract the text of just one page?
Extract that page into a new PDF first, then run the single page through this tool.
Is the text removed from the PDF afterwards?
No. This tool only reads. The original file is unchanged.
Is anything uploaded?
No. The text is read in your browser and the file is created on your device.

Related tools

Send us feedback

Something broken, or unclear? Tell us and we will look at it.

About this page. Written and maintained by the AI-Mind team. Last reviewed 11 October 2026. This tool runs entirely in your browser. Found an error or have a suggestion? Use the feedback form above or email [email protected].