TinyFileLab

Home›PDF›PDF to text

Nothing is uploaded — it all runs on your device

PDF to text converter

Pull the text out of a PDF as plain text: every page at once, with table columns kept apart and fake-bold duplicates removed. The file is read in your browser and never uploaded.

Drop a PDF here or tap to choose — text-based PDFs, any size

Choose a PDF and its text appears below. Nothing is uploaded.

Read in your browserStatements and contracts are opened by this page, not uploaded to one.
No gibberish passed off as textCharacters that cannot be decoded are counted and left out, not guessed.
Hundreds of pages in secondsA 210-page filing comes out in about a second and a half.

What comes out, and in what shape

A PDF does not store its text as sentences. It stores a drawing program: choose this font at this size, move to these coordinates, draw these glyph codes, move again. To get the words back, this page runs that program — tracking the text matrix, the font size, character and word spacing and every glyph's width — until it knows which character each code stands for and where on the page it lands. Then it puts the pieces back together into lines.

By default you get the lines as they sit on the page, with a blank line wherever the gap between two lines is clearly bigger than the page's normal line spacing. That keeps addresses, tables and headings readable. Tick Join wrapped lines into paragraphs when you want flowing prose instead: lines inside a paragraph are joined with a space, and a word that was split across two lines with a hyphen (infor- / mation) is glued back together, as long as the second half starts with a lower-case letter.

Tabs between table columns turns a wide horizontal gap into a tab character instead of a space. It is on by default because the most common reason to pull text out of a PDF is a statement or a report with columns in it, and tab-separated lines paste straight into a spreadsheet as separate cells. Switch it off if you only want prose.

Why a PDF can have no text at all

If a document was scanned, photographed, or “printed” through an image-only driver, each page is a single picture. There are no characters in the file to find, only pixels that happen to look like letters. This tool tells you exactly which pages are in that state rather than handing you a blank result without explanation. Getting text out of them needs optical character recognition, which is a different job and is not done here.

Scanners and phone apps often run OCR themselves and hide the recognised text behind the image, drawn in invisible ink. That hidden layer is real text as far as the file is concerned, and it is extracted like any other. Its quality is whatever the scanner's OCR managed, so a scan of a creased page can produce text with odd substitutions that were already in the file.

Gibberish, and how it is avoided

You may have copied text out of a PDF and pasted something like pUHYLRXV\HDU¶V instead of “previous year’s”. That is not corruption. The font in the file numbers its glyphs in its own order, and the file forgot to include the table (a ToUnicode map) that says which glyph is which letter. A naive extractor reads the glyph numbers as if they were character codes and every letter comes out shifted.

When that table is missing or empty, this page goes looking inside the embedded font file itself: its own character map, read backwards, and failing that its glyph names. For the Microsoft core fonts — Arial, Times New Roman, Courier New, Verdana, Tahoma and Georgia, which all share one glyph order — it can fall back on that order when iPhones and Macs strip everything else out, which they regularly do. On a set of iPhone-generated forms where a well-known PDF library returned shifted nonsense, this recovers the actual sentences.

When none of that works, the characters are counted and left out, and the report under the text tells you how many. A gap you know about is more useful than a paragraph of plausible-looking garbage that you only notice after you have pasted it into an email.

Reading order

Text comes out in the order the file draws it, which for almost every word processor, report generator and browser is the reading order: columns are drawn one after another, not interleaved line by line. That is the right default, because sorting everything top to bottom would splice the left and right columns of a two-column page into nonsense.

The cost is that a few generators draw a page in an odd order — the footer first, or a table cell by cell down each column — and the extracted text follows them. When that happens the text is all there, just in a sequence you need to rearrange. Rotated text, such as a sideways label on a chart, comes out as its own line.

Two things that look like text are not included. What someone typed into a fillable form's fields lives in the form layer, not in the page's drawing, and comments and sticky notes are annotations laid on top of the page. Both are left out, so a filled-in form gives you its printed labels without the answers.

Fake bold and repeated text

Some programs make text look bold by drawing it several times, each copy nudged a fraction of a point sideways. Bank statement systems do this constantly. A simple extractor faithfully reports every copy, so “Deposit $31,746.61” appears four times in a row. Here, a run of text that repeats the previous one at almost exactly the same position is recognised as the same text and kept once. The tolerance is deliberately smaller than the narrowest letter, so two genuine letter l’s in “William” are never mistaken for a copy of each other.

Letter-spaced headings get the reverse treatment. When a heading is drawn one glyph at a time with extra tracking, every letter has a small gap after it; the tool measures the usual gap on that line and only calls it a space when a gap is clearly bigger. Otherwise “YOUR PRIVATE LINK” would come out as a row of single letters.

How it was checked

Before this page went up, its output was compared against PyMuPDF, a widely used open-source PDF library, on 288 real documents from more than a hundred different producers: Word, Excel and PowerPoint exports, Chrome and Safari print-to-PDF, LaTeX, bank statement systems, iText, ReportLab, scanners with OCR, and more. For the median document the characters matched exactly. Where they did not, the difference was usually on the reference side — the repeated fake-bold copies described above, or the letter-shifted text from forms whose fonts had lost their mapping.

Common questions

Why is my extracted text empty?

Because the PDF has no text layer: the pages are pictures of text, typically from a scanner or a phone camera. The report under the output says which pages are affected. Turning a picture into text needs OCR, which this tool does not do. If the scanner ran OCR itself, the hidden text layer it added is extracted normally.

Can I get the text of just a few pages?

Yes. Type a range in the Pages box, such as 3, 1-5 or 2,7,10-. The output updates as you type. Page numbers are positions in the file counting from 1, not the numbers printed on the pages.

Does it keep tables?

It keeps each row on its own line, and with Tabs between table columns switched on, wide gaps between columns become tab characters. Pasted into Excel, Numbers or Google Sheets, each tab starts a new cell, which is usually enough to rebuild a statement or a price list. It does not detect cell borders or merged cells.

Why do some characters come out missing?

Those glyphs use a font that carries no information about which letters they are, not even inside the embedded font file. Rather than print them as shifted nonsense, the tool leaves them out and tells you how many there were. Opening the same PDF in another reader and copying usually produces gibberish for the same characters, which is a sign the information really is not in the file.

Can it read a password-protected PDF?

Yes. A PDF that opens without asking for a password (one that only restricts printing or copying) is read straight away. If the file asks for a password to open, a password box appears; the password is checked inside this page and is not sent anywhere. There is no guessing or cracking of passwords.

Does it work for Chinese, Japanese, Arabic or Russian?

Chinese, Japanese, Korean, Cyrillic and Greek text extract correctly when the PDF maps its glyphs to Unicode, which modern files nearly always do, and older CJK encodings such as Shift-JIS and GBK are decoded with the browser's own converters. Right-to-left scripts come out with the characters correct but in the order they were drawn, which can be reversed compared with how you read them.

Is my PDF uploaded?

No. The file is read by this page in your browser and never leaves your device; there is no server to receive it. Load the page, disconnect from the internet, and extraction still works. That matters for the documents people most often need text from: statements, payslips, contracts and medical letters.

How is this different from copying and pasting from a PDF reader?

Copying from a reader gives you one page's selection at a time, often with a line break after every visual line, duplicated text from fake bold, and garbled letters where a font lacks its Unicode map. This reads every page in one go, removes the duplicates, recovers characters from the embedded fonts where it can, and can reflow lines into paragraphs.