We check whether you need it
Most PDFs already have text. If yours does, we say so and do nothing, it is already searchable, exactly, with no guessing involved.
A scan has no text in it, nothing to search, nothing to copy. This reads what is on the page and puts it back into the file, so the document works everywhere, not just here.
A scan has no text in it, nothing to search, nothing to copy. This reads what is on the page and puts it back into the file, so the document works everywhere, not just here.
Most PDFs already have text. If yours does, we say so and do nothing, it is already searchable, exactly, with no guessing involved.
One page at a time, so a long document becomes searchable from the first page rather than only at the end. Roughly three seconds a page.
Behind the scan and invisible, at the position of each word. The page looks exactly as it did; the text is simply there now.
Everything you need is ready in the browser.
Most PDFs already have text. If yours does, we say so and do nothing, it is already searchable, exactly, with no guessing involved.
One page at a time, so a long document becomes searchable from the first page rather than only at the end. Roughly three seconds a page.
Behind the scan and invisible, at the position of each word. The page looks exactly as it did; the text is simply there now.
A scanned PDF is a picture of a page. To a computer it is no different from a photograph of a wall, there is nothing to search, nothing to select, nothing to copy. OCR reads the shapes and works out what they say.
What we do with the answer matters more than most tools admit. The recognized text is written back into the PDF as an invisible layer sitting exactly where the ink is. Your file still looks identical. But Preview, Acrobat, Chrome, Spotlight and Windows Search can all now find text inside it.
That is the difference between a document you own and a feature you rent. A tool that keeps the text in its own database gives you search for as long as you keep paying. A searchable PDF is yours, and it still works years later on a machine that has never heard of us.
On a clean scan of printed text, very. On a phone photograph of a creased contract in bad light, less so, and no OCR is different in that respect, including the expensive ones.
So we show you the confidence for every page rather than presenting the result as fact. Under about 70 per cent usually means a misreading rather than a difficult word, and those pages are worth a glance.
Words we are genuinely unsure of are left out of the file rather than guessed at. Invisible text cannot be proofread, nobody will ever notice it is wrong , so a bad guess becomes a silent false result in every search from then on. A gap you can find is better than a mistake you cannot.
No. The recognition runs on your own machine, in your browser. There is no OCR service involved and no per-page cost, which is also why we can offer it without metering it.
If the document is in your workspace it is already stored with us, so this is not a claim that it never leaves your computer. It is a narrower and more honest one: reading it does not involve a third party.
That is the part a downloaded file cannot do. Alongside the text in the PDF, we keep the position of every word, so you can search a whole project and land on the right page of the right document, with the words highlighted.
A filing cabinet of scanned contracts is the case this exists for. One searchable PDF is useful; four hundred of them, searchable together, is a different thing entirely.
English at the moment. Each additional language is a separate model of a few megabytes, downloaded the first time it is used, so we would rather add the ones people ask for than ship a dozen nobody wants.
If you need another, tell us which and it is a small change.
It does not read handwriting. Tesseract is trained on printed type, and a handwritten note comes back as noise or as nothing. If you need handwriting, you need a different kind of model and we do not have one.
It does not correct itself. A word read wrongly stays wrong, and because the text is invisible you will not see the mistake, you will simply not find the document when you search for it. Check the confidence figures on anything that matters.
It does not preserve layout. Columns, tables and forms come back as text in reading order, which is usually right and sometimes scrambles a two-column page. The positions are correct even when the order is not, so search still lands in the right place.
It only reads English at the moment.
Open the editor to work on your file in the browser, with no account or watermark.
No. The page image is untouched, we compared the before and after pixel by pixel, and not one differs. The text is added behind it and never drawn.
No. It is drawn at zero opacity rather than in white, which is the mistake that produces a document looking fine on screen and covered in duplicated text on paper.
Around three seconds a page on a normal laptop. A fifty-page scan is a couple of minutes, and each page becomes searchable as it finishes rather than at the end, so you can start reading page one while page forty is still going.
Yes, on Pro. That is the part a downloaded file cannot do: the word positions are kept alongside the file, so a search across a project lands on the right page of the right document with the words highlighted.
We tell you it does not need this, and do nothing. Its existing text is exact; anything OCR produced would be a guess at something already known.
Everything you need to take a document from raw file to reviewed, signed and filed.