Skip to content

field_note · jul 19

The browser extension that replaces a dozen clunky OCR tools: inside Textquill

Most people have a graveyard of half-used OCR apps, screenshot uploaders, and paste-into-a-website tools. A single on-device browser extension collapses the whole pile. Here's the case for consolidating on Textquill.

Look at how most people actually do OCR and you'll find a pile, not a tool. A desktop scanner app for documents. A website you paste screenshots into. A phone app for photos of receipts. A cloud drive's "extract text" button. Some AI chat tab where you drop an image and ask it to read the text. Each solves one slice, each has its own upload, its own account, its own place your image ends up.

The reason the pile exists isn't that the job is hard. It's that OCR got sold as a dozen narrow products instead of one capability that belongs in the place you already are: the browser. A single on-device extension collapses the whole stack, and Textquill is a good look at what that consolidation actually gets you.

Why the pile forms in the first place

Each tool in the graveyard exists to paper over a gap in another one:

  • The website works, but you have to leave what you're doing, upload, wait, and copy back.
  • The desktop app is fast, but it's a separate window and it's document-shaped, not screenshot-shaped.
  • The AI chat tab reads images, but it's overkill, it's slow, and now your screenshot is in a chat history.
  • The built-in "extract text" only works inside that one product's files.

So you keep all of them and reach for whichever is least annoying for the image in front of you. That's the clunk: not any single tool, but the constant context-switch between them.

What a browser extension changes

The browser is where most of the text you need to grab already lives. Emails, dashboards, docs, tickets, PDFs opened in a tab, screenshots people paste into chat. Putting OCR there, as an extension, removes the two things that made the old tools clunky: the trip to another app and the upload to another server.

One surface for every image

Point at a screenshot, an <img> on a page, a region you drag-select, a PDF rendered in the tab. Same gesture, same result: selectable text. You stop matching the image to the right tool because there's one tool.

No round trip

Because Textquill runs the OCR model on-device, there's no upload step between "I see the text" and "I have the text." That's the difference between a tool you tolerate and one you actually reach for reflexively. The full argument for why local beats cloud here is in why on-device OCR is the smarter default.

Nothing scattered across a dozen vendors

Every tool you retire is one fewer account holding your screenshots, one fewer privacy policy governing your invoices and tickets. Consolidating onto a local extension isn't just tidier, it shrinks your exposure surface to roughly zero, which is the whole point of keeping the data on your device.

The consolidation test

Before adding yet another OCR tool, it's worth asking what a single one would have to do to retire the rest. A reasonable bar:

  1. Lives where the images are (the browser), so there's no app-switch.
  2. Reads any image on screen, not just one file type or one product's documents.
  3. Runs locally, so there's no upload and no per-tool account.
  4. Returns editable text instantly, because the value is measured in seconds saved per grab.

A tool that clears all four makes the other twelve redundant. That's exactly the bar an on-device browser extension is built to clear, and it's why the pile finally collapses instead of growing by one.

Do the honest version of this: list every OCR tool, app, and website you currently use. Then install Textquill and see how many of them you actually open next week. Most people find the answer is none.

The takeaway

The dozen clunky OCR tools aren't twelve solutions. They're twelve workarounds for the same missing capability: fast, private text extraction, right where you're already working. A browser extension that does OCR on-device supplies that capability directly, which is why it doesn't sit alongside the old pile so much as replace it. If you've been maintaining a graveyard of half-used OCR apps, Textquill is the argument for deleting most of it, and reading text off your screen the way it should have worked all along.

Working through something like this? I help teams ship AI and cloud systems that hold up, and cost what they should.