

We've tested Speechdub, a document listening app that turns PDFs, web pages and pasted text into audio you can follow along with, rewrite, translate and question.
Welcome to this Speechdub review β¨
Everyone has the pile. The whitepaper you saved in March, the forty-page essay someone swore was worth it, the school PDF your kid is supposed to get through. None of it is hard to read. It's just that reading it means sitting down and giving it your eyes, and your eyes are already booked.
Text-to-speech has promised to solve this for about fifteen years and mostly delivered a robot reading a table of contents. Speechdub is a newer attempt, and its pitch is narrower and more interesting than "we read your PDFs": it wants to be a workspace where a document becomes something you listen to, follow along with, rewrite for the ear, translate, and interrogate, all in one place.
I signed up and spent an afternoon in it, importing two documents: a 72-minute essay by URL and a two-page school handout as a PDF. Here's what actually happens.

The app opens on a refreshingly empty screen. A headline reading "Your documents, finally speaking", one large box saying "Drop a file, paste text, or a URL...", an Add file button, and nothing else. No project setup, no workspace naming, no import wizard. The left rail has three items (New, Search, Library) and a credit meter pinned to the bottom.
I started with a URL, partly because URL and text imports don't consume credits and partly because it's the harder test. Speechdub showed "Fetching page...", and about ten seconds later I was looking at Paul Graham's How to Do Great Work rendered as a clean, paragraph-by-paragraph reading view with a player docked at the bottom. It reported 71 minutes and 40 seconds of audio, split into 761 individually addressable sentences.
Two things about that import. The app auto-titled the document "Universal work mastery recipe" rather than keeping the page's own title, which is a slightly odd choice: it's a fair summary, but finding the thing again means remembering Speechdub's name for it rather than the author's. And it carried a few markdown footnote artifacts into the body (\[[1](#f1n)\] sitting at the end of a sentence), which the reading view displays and would presumably try to pronounce. The built-in editor exists partly to clean exactly that up.

PDF is the input the marketing leads with, so I made a real one: a two-page Year 9 biology revision handout with a title, a subtitle, seven numbered sections, and a chemical equation. Dropped in, it came back noticeably cleaner than the web import.
Headings survived as actual headings, with the title, subtitle and section headers in a visible hierarchy rather than flattened into body text. The equation was the surprise. My source file had it as flat ASCII, 6CO2 + 6H2O -> C6H12O6 + 6O2, and the reading view rendered it as 6 COβ + 6 HβO β CβHββOβ + 6 Oβ, with real subscripts and a real arrow. Nothing about listening to a document requires that, which is precisely why it's a nice signal about how much care went into the import path. It came out at 50 sentences and an estimated 3:05 of audio, with no stray artifacts of the kind the URL import left behind.
One rough edge showed up here. Because the handout has headings, the Outline panel finally has something to show, and it lists the title and subtitle correctly, then renders my seven numbered sections as bare 1. 2. 3. down to 7. with the titles missing. It appears to treat the leading number as the entire heading. On a textbook chapter or a contract, where numbered headings are the norm, that turns the outline into a useless column of digits.
Fair warning that I tested a small, well-formed, text-based PDF. A scanned page or a two-column academic paper is a different problem, and I can't tell you how it copes with either.
The player is a single strip: previous sentence, play, next sentence, a speed control, a scrubber, the current voice, a microphone, and a chat button. Speed runs from 0.5Γ to 2.5Γ on a slider, wider than most readers bother with.
Every sentence in the document is a real button. Clicking any line seeks the player straight to it, and the sentence being read carries a faint highlight. That sounds minor until you're forty minutes into something and want to re-hear one paragraph. Most listening apps make you scrub blindly.
The voice picker is where Speechdub's central idea lives. Open it and you get a language selector (mine read "English DETECTED", worked out from the import), ten English voices with illustrated avatars (Nathan, Leo, Marcus, Alex, Ethan, Sophie, Emma, Clara, Luna, Nora), and a toggle between Cloud, tagged 1 credit/min, and Eco, tagged with a little leaf.

Eco is the interesting one. It's local, in-browser synthesis: same voices, no credits, unlimited. Speechdub claims up to 20Γ less power than cloud playback, which is their figure and not one I could test.
What the marketing doesn't tell you is the entry cost. The first time I hit play in Eco mode, the app began downloading voice models, with a progress line reading vector_estimator.onnx Β· 1% Β· 2.3 / 256.5 MB. Once installed, the site was holding 380 MB in browser storage. So: 256 MB to download, 380 MB parked on your disk afterwards.
That's not a scandal, it's the honest price of running a neural voice on your own machine, and it's a one-off. But it happens before you hear a single word, and the UI presents it as an ordinary loading spinner rather than a download you might want to start deliberately. There's a Download models button in the voice picker for doing it on purpose, which is the right idea; it just isn't where a first-time user looks.
This is the feature that separates Speechdub from a text-to-speech button, and the one I'd lead with if I were the founder.
A document can have multiple versions, and creating one opens a panel with four tabs. Each states its intent in a sentence, then shows a capability checklist with the traits it gives you ticked and the ones it doesn't struck through. It's a genuinely clear piece of UI, because it tells you what you're trading away:

Pick a tab, hit Continue, and you land on a second step with an Output Language dropdown and a live credit estimate. Keeping my English essay in English was quoted as Free. Switching the output to French quoted 3 credits, which for a 68,000-character document is close to nothing.
Two notes on how versions behave. Generating one is a background job rather than something you wait on, so it appears in the switcher straight away and fills in when ready, and you can keep listening to the source meanwhile. A whole-document rewrite of that size takes a while, so it's a feature you kick off and come back to.
The thing I'd change is the naming. A new version arrives labelled "Original" exactly like the source, with no language or mode tag to tell them apart. Once you have a French copy and a Lecture rewrite of the same document, the switcher is three entries with the same name. Stamping the mode and output language on the label would fix it.
The chat panel opens on the right, scoped to the document you're in. I asked it something that required actually reading the essay rather than skimming a summary: the four steps the author lays out for a project, plus what he says about curiosity.
It came back with the four steps correctly named and explained (choose a field, learn enough to get to the frontier, notice gaps, explore promising gaps), then a separate section on curiosity that picked up the essay's "engine and rudder" framing and its point that excessive curiosity is a signal about what to pursue. No hallucinated structure, no generic self-help filler.
The nice touch is that the answer is attributed to your chosen voice (mine came from Sophie) and carries a read-aloud button, so an answer can be spoken back to you without leaving the listening flow. There's also a microphone in the player for asking by voice, which is the version of this that makes sense with headphones on and hands full.
Exporting is one click from the document toolbar, and the modal is admirably blunt. It showed the version name, ~72 min, the voice (Sophie, tagged Narrator), the speed, and then: Estimated credits 72 credits.

That single number tells you more about Speechdub's economics than the pricing page does. Cloud playback and export both run at roughly a credit a minute. Premium's 1,200 monthly credits therefore buy about 20 hours of cloud listening or exported audio a month, and the free tier's 10 credits buy ten minutes.
I can confirm how fast that goes. One French translation, one chat question, a few minutes of cloud playback and one small PDF import took me from 10 credits to 0 in a single afternoon. The free plan is not a trial of cloud playback in any meaningful sense. It's a trial of Eco playback, which is unlimited and free, with a handful of credits so you can see what the paid features do. Once you understand that, the pricing stops looking confusing and starts looking deliberate. There's also a detail in the footnotes that softens the export cost: cloud-cached sentences export at 0 credits, so audio you've already listened to isn't billed twice.
A quick roundup of what else is in there:
Cmd+Shift+O for a new document, Cmd+K for search.If you've searched for this kind of tool you already have Speechify or ElevenLabs Reader in mind, and your browser probably reads pages aloud for free. The built-in readers in Edge and Safari are genuinely fine for a single article, but they read a page; they don't give you a library, a place in a document you can come back to, versions, or a chat scoped to what you're listening to.
Against the dedicated apps, Speechdub's distinguishing bet is Eco. Speechify and ElevenLabs are cloud-first, so listening is metered against a plan, and heavy listening means a bigger plan. Speechdub's local mode moves the cost of routine listening to a one-off download on your own machine and charges only for the things that genuinely need a server: cloud voices, rewrites, translation, chat and export. If you listen a lot and don't need much else, that's a meaningfully different shape.
Prices follow your location and the figures don't change with it: the same page showed me β¬6 and showed a US visitor $6.
| Plan | Price | Documents | Credits / month |
|---|---|---|---|
| Free | 0 forever | 10 | 10 |
| Premium (monthly) | 10 / month | Unlimited | 1,200 |
| Premium (annual) | 6 / month (72 / year, 40% off) | Unlimited | 1,200 |
| Credit Pack | 7 once, never expires | n/a | 500 added |
Two things matter more than the table. Eco playback is unlimited and free on every plan including Free, which is the load-bearing promise of the whole model. And 1,200 credits is roughly 20 hours of cloud audio a month, which is the number to hold against your own habits. Credits are spent on cloud playback (~1 per minute at 1Γ), PDF import and audio export (0.5 per started MB or 30 seconds), and document chat, AI versions and enhance (tokens plus 1 credit).

Worth flagging: at the time of writing, the pricing page renders its plan limits as literal template placeholders. Free reads "{10} documents" and "{10} credits / month", Premium reads "{1,200} credits / month", the Credit Pack reads "{500} credits added to your balance". The comparison table further down the same page shows the correct numbers, so nothing is wrong with the plans themselves. It's the headline cards, on the exact screen where someone decides whether to pay, that are visibly broken.
The last pricing note is the one I liked most. Speechdub offers reduced pricing for accessibility needs, applied for from inside the app, described as trust-based with no verification required. For a product whose most obvious users are people with dyslexia, ADHD or low vision, choosing not to make them prove it is a real decision and a good one.

Speechdub fits well if you're:
It's the wrong tool if you want a general-purpose voice generator for video or podcast production. There's no multi-speaker timeline, no SSML control, no project structure. Speechdub is built around a document you are trying to get through, and that focus is why the rest of it is as simple as it is. It's also desktop-first for the free path, since Eco mode needs a modern desktop browser, so a phone-only user is on cloud playback and therefore on credits from the start.
There's a version of this product that's just a play button on a PDF, and Speechdub is clearly trying not to be that. Making sentences individually addressable, letting a document be rewritten four ways for four listening jobs, scoping chat to the document, typesetting a chemical equation nobody will ever look at, pricing local playback at zero forever: these are the decisions of someone who has thought about what listening to a document is actually like.
The rough edges are all in the packaging rather than the thinking. Unrendered placeholders on the pricing page, versions that all answer to the same name, an outline that eats numbered headings, a quarter-gigabyte download served as an anonymous spinner: none of that is architecture, it's a week of polish that hasn't happened yet.
What I liked:
Things to keep in mind:
{10}, {1,200} and {500} as unsubstituted placeholdersIf you have a reading pile you keep not getting to, Speechdub is worth an evening and costs nothing to find out.



