Extract text from a PDF
Pull the words out of a PDF so you can search, quote or edit them. Nothing is uploaded.
Free ยท No sign-up ยท Your file never leaves your computer
No file handy? โ generated here, not downloaded.
How to use it
- 1
Choose a PDF that has real text in it โ one you can select in a reader.
- 2
Extract. The text appears below, page by page.
- 3
Copy it, or download it as a .txt file.
When you would use this
Text in a PDF is often locked in a shape that is fine to read and useless to work with. You need a clause out of a contract, a specification into a spreadsheet, or a paragraph you have to quote back with the wording exactly right.
The one thing to know first: a scan has no text in it. A PDF produced by a scanner or a phone camera is a picture of a page. There are no words in the file, only pixels arranged to look like words, and no amount of extracting will find them. The test takes two seconds โ open it in any reader and try to select a sentence. If your cursor will not grab it, there is nothing there, and what you need is OCR rather than extraction.
Layout is rebuilt, not stored. A PDF does not contain paragraphs; it contains fragments of text with coordinates saying where each one sits on the page. Lines here are reconstructed by grouping fragments that share a baseline, which recovers ordinary prose reliably. Columns, tables and anything with an unusual reading order come out approximately โ the words are all present, the shape is a best effort.
Pages are separated with a marked heading rather than run together. That matters more than it sounds when you are quoting from a long document and need to say which page something came from.
One caution about what you do next. Text extracted from a signed or issued document is a copy, not the document โ the moment it leaves the PDF it loses whatever authority the original had. Quote from it, search it, work with it; do not send it on as though it were the thing itself.
Questions
- Nothing came out โ why?
- The PDF is almost certainly a scan: a picture of a page, with no text layer in it at all. Extracting text from an image needs OCR, which this tool does not do. If you cannot select the text in a PDF reader, there is none to extract.
- Why is the layout different?
- A PDF stores where each fragment of text sits on the page, not paragraphs. Lines are rebuilt from those positions, which recovers reading order well and columns and tables imperfectly.
- Are the page breaks marked?
- Yes โ each page is separated by a marked heading so you can tell where one ended, which matters when you are quoting from a long document.
- Can I extract from a password protected PDF?
- Remove the password in your PDF reader first. A locked file cannot be read without it, and the file never reaches us to ask.
- Is my document uploaded?
- No. The text is extracted inside your browser and nothing is sent anywhere.
Related tools
Retyping from a PDF because the words live in the wrong place
Ofivio keeps scopes, specifications and BOQ lines as records rather than as documents, so the text is already data and nothing has to be extracted from anything.
See Ofivio