Extract Text from a PDF
Get the words out of a PDF as plain text, page by page, without copy and paste. Handy for quoting, searching or scripting.
Files are never uploadedHow to use it
Add your PDF
Drop the document onto the box above, or click to browse. The text layer of each page is read in turn.
Choose page markers
Keep the headers on when you plan to search or quote across a long document, and turn them off when you want a clean block of text.
Extract
Press Extract Text. The progress bar names the page being read, which is useful on long documents.
Check the text
Open the downloaded file and skim a couple of pages. Reading order and spacing are approximate, so tables and multi-column layouts need a quick look.
Save the text file
Download the plain UTF-8 text file and open it in any editor, spreadsheet or script.
Common issues
Expecting OCR
If the page is a scan, the extracted text is empty or nearly empty. That is not a fault in the tool; there is simply no text layer to read.
Tables coming out scrambled
Rows and columns flatten into a word stream. Treat the output as a starting point and clean it up before it goes into a spreadsheet.
Right-to-left and non-Latin scripts
Reading order for Arabic, Hebrew and some Asian scripts can come out reversed or broken, because a PDF stores glyphs in drawing order rather than reading order.
Very large documents taking a while
A several hundred page file produces a big text file and takes time to read page by page. The progress bar shows how far it has reached.
Frequently asked questions
Does this work on a scanned PDF?
Does it keep the original layout?
What encoding is the file in?
Can I copy the text without downloading?
Why are some characters wrong?
Do hyperlinks survive?
Can I extract the text of just one page?
Is the text removed from the PDF afterwards?
Is anything uploaded?
Related tools
Send us feedback
Something broken, or unclear? Tell us and we will look at it.