Aug 12, 2026-Reviews
Speechdub Review - Your Reading Pile, Finally Out Loud

Speechdub Review - Your Reading Pile, Finally Out Loud

We've tested Speechdub, a document listening app that turns PDFs, web pages and pasted text into audio you can follow along with, rewrite, translate and question.

Welcome to this Speechdub review ✨

Everyone has the pile. The whitepaper you saved in March, the forty-page essay someone swore was worth it, the school PDF your kid is supposed to get through. None of it is hard to read. It's just that reading it means sitting down and giving it your eyes, and your eyes are already booked.

Text-to-speech has promised to solve this for about fifteen years and mostly delivered a robot reading a table of contents. Speechdub is a newer attempt, and its pitch is narrower and more interesting than "we read your PDFs": it wants to be a workspace where a document becomes something you listen to, follow along with, rewrite for the ear, translate, and interrogate, all in one place.

I signed up and spent an afternoon in it, importing two documents: a 72-minute essay by URL and a two-page school handout as a PDF. Here's what actually happens.

The Speechdub home page: 'Turn any document into audio', with the four-step strip and a mock of the app's single paste box.
The Speechdub home page: 'Turn any document into audio', with the four-step strip and a mock of the app's single paste box.

Getting in: one box, three kinds of input

The app opens on a refreshingly empty screen. A headline reading "Your documents, finally speaking", one large box saying "Drop a file, paste text, or a URL...", an Add file button, and nothing else. No project setup, no workspace naming, no import wizard. The left rail has three items (New, Search, Library) and a credit meter pinned to the bottom.

I started with a URL, partly because URL and text imports don't consume credits and partly because it's the harder test. Speechdub showed "Fetching page...", and about ten seconds later I was looking at Paul Graham's How to Do Great Work rendered as a clean, paragraph-by-paragraph reading view with a player docked at the bottom. It reported 71 minutes and 40 seconds of audio, split into 761 individually addressable sentences.

Two things about that import. The app auto-titled the document "Universal work mastery recipe" rather than keeping the page's own title, which is a slightly odd choice: it's a fair summary, but finding the thing again means remembering Speechdub's name for it rather than the author's. And it carried a few markdown footnote artifacts into the body (\[[1](#f1n)\] sitting at the end of a sentence), which the reading view displays and would presumably try to pronounce. The built-in editor exists partly to clean exactly that up.

The reading view mid-playback: the current sentence carries a faint highlight, and the player shows position in the 71-minute essay.
The reading view mid-playback: the current sentence carries a faint highlight, and the player shows position in the 71-minute essay.

The PDF path, which is the better one

PDF is the input the marketing leads with, so I made a real one: a two-page Year 9 biology revision handout with a title, a subtitle, seven numbered sections, and a chemical equation. Dropped in, it came back noticeably cleaner than the web import.

Headings survived as actual headings, with the title, subtitle and section headers in a visible hierarchy rather than flattened into body text. The equation was the surprise. My source file had it as flat ASCII, 6CO2 + 6H2O -> C6H12O6 + 6O2, and the reading view rendered it as 6 COβ‚‚ + 6 Hβ‚‚O β†’ C₆H₁₂O₆ + 6 Oβ‚‚, with real subscripts and a real arrow. Nothing about listening to a document requires that, which is precisely why it's a nice signal about how much care went into the import path. It came out at 50 sentences and an estimated 3:05 of audio, with no stray artifacts of the kind the URL import left behind.

One rough edge showed up here. Because the handout has headings, the Outline panel finally has something to show, and it lists the title and subtitle correctly, then renders my seven numbered sections as bare 1. 2. 3. down to 7. with the titles missing. It appears to treat the leading number as the entire heading. On a textbook chapter or a contract, where numbered headings are the norm, that turns the outline into a useless column of digits.

Fair warning that I tested a small, well-formed, text-based PDF. A scanned page or a two-column academic paper is a different problem, and I can't tell you how it copes with either.

The player, and the thing nobody tells you about Eco mode

The player is a single strip: previous sentence, play, next sentence, a speed control, a scrubber, the current voice, a microphone, and a chat button. Speed runs from 0.5Γ— to 2.5Γ— on a slider, wider than most readers bother with.

Every sentence in the document is a real button. Clicking any line seeks the player straight to it, and the sentence being read carries a faint highlight. That sounds minor until you're forty minutes into something and want to re-hear one paragraph. Most listening apps make you scrub blindly.

The voice picker is where Speechdub's central idea lives. Open it and you get a language selector (mine read "English DETECTED", worked out from the import), ten English voices with illustrated avatars (Nathan, Leo, Marcus, Alex, Ethan, Sophie, Emma, Clara, Luna, Nora), and a toggle between Cloud, tagged 1 credit/min, and Eco, tagged with a little leaf.

The voice picker: ten English voices, the detected language at the top, and the Cloud toggle showing its 1 credit/min rate next to free Eco.
The voice picker: ten English voices, the detected language at the top, and the Cloud toggle showing its 1 credit/min rate next to free Eco.

Eco is the interesting one. It's local, in-browser synthesis: same voices, no credits, unlimited. Speechdub claims up to 20Γ— less power than cloud playback, which is their figure and not one I could test.

What the marketing doesn't tell you is the entry cost. The first time I hit play in Eco mode, the app began downloading voice models, with a progress line reading vector_estimator.onnx Β· 1% Β· 2.3 / 256.5 MB. Once installed, the site was holding 380 MB in browser storage. So: 256 MB to download, 380 MB parked on your disk afterwards.

That's not a scandal, it's the honest price of running a neural voice on your own machine, and it's a one-off. But it happens before you hear a single word, and the UI presents it as an ordinary loading spinner rather than a download you might want to start deliberately. There's a Download models button in the voice picker for doing it on purpose, which is the right idea; it just isn't where a first-time user looks.

AI versions: four ways to rewrite a document for the ear

This is the feature that separates Speechdub from a text-to-speech button, and the one I'd lead with if I were the founder.

A document can have multiple versions, and creating one opens a panel with four tabs. Each states its intent in a sentence, then shows a capability checklist with the traits it gives you ticked and the ones it doesn't struck through. It's a genuinely clear piece of UI, because it tells you what you're trading away:

  • Original: "Translate or restore the original version." Keeps original phrasing and chapter breaks, available in 31 languages, explicitly does not do natural-sounding phrasing or memorisation.
  • Natural: "Moderate rephrasing for better flow and natural speech patterns." Adds natural phrasing, removes redundancy, cleans up document artifacts.
  • Lecture: "Structured for dense information and focused listening." Paced for information density, cleans up artifacts, drops the natural phrasing.
  • Conversational: "Dialogue-style rewrite for podcast-like listening." Natural phrasing and redundancy removal, but not paced for density.
The Natural version mode, with the four traits it adds ticked and the two it deliberately skips struck through.
The Natural version mode, with the four traits it adds ticked and the two it deliberately skips struck through.

Pick a tab, hit Continue, and you land on a second step with an Output Language dropdown and a live credit estimate. Keeping my English essay in English was quoted as Free. Switching the output to French quoted 3 credits, which for a 68,000-character document is close to nothing.

Two notes on how versions behave. Generating one is a background job rather than something you wait on, so it appears in the switcher straight away and fills in when ready, and you can keep listening to the source meanwhile. A whole-document rewrite of that size takes a while, so it's a feature you kick off and come back to.

The thing I'd change is the naming. A new version arrives labelled "Original" exactly like the source, with no language or mode tag to tell them apart. Once you have a French copy and a Lecture rewrite of the same document, the switcher is three entries with the same name. Stamping the mode and output language on the label would fix it.

Document chat, tested on a 72-minute essay

The chat panel opens on the right, scoped to the document you're in. I asked it something that required actually reading the essay rather than skimming a summary: the four steps the author lays out for a project, plus what he says about curiosity.

It came back with the four steps correctly named and explained (choose a field, learn enough to get to the frontier, notice gaps, explore promising gaps), then a separate section on curiosity that picked up the essay's "engine and rudder" framing and its point that excessive curiosity is a signal about what to pursue. No hallucinated structure, no generic self-help filler.

The nice touch is that the answer is attributed to your chosen voice (mine came from Sophie) and carries a read-aloud button, so an answer can be spoken back to you without leaving the listening flow. There's also a microphone in the player for asking by voice, which is the version of this that makes sense with headphones on and hands full.

Audio export, and the credit maths that follows

Exporting is one click from the document toolbar, and the modal is admirably blunt. It showed the version name, ~72 min, the voice (Sophie, tagged Narrator), the speed, and then: Estimated credits 72 credits.

The export modal quoting 72 credits to render the 72-minute essay as an MP3, before you commit to it.
The export modal quoting 72 credits to render the 72-minute essay as an MP3, before you commit to it.

That single number tells you more about Speechdub's economics than the pricing page does. Cloud playback and export both run at roughly a credit a minute. Premium's 1,200 monthly credits therefore buy about 20 hours of cloud listening or exported audio a month, and the free tier's 10 credits buy ten minutes.

I can confirm how fast that goes. One French translation, one chat question, a few minutes of cloud playback and one small PDF import took me from 10 credits to 0 in a single afternoon. The free plan is not a trial of cloud playback in any meaningful sense. It's a trial of Eco playback, which is unlimited and free, with a handful of credits so you can see what the paid features do. Once you understand that, the pricing stops looking confusing and starts looking deliberate. There's also a detail in the footnotes that softens the export cost: cloud-cached sentences export at 0 credits, so audio you've already listened to isn't billed twice.

The rest of the product

A quick roundup of what else is in there:

  • Import formats: PDF, web pages, plain text files, markdown, and pasted text. Paste and URL imports are free; PDF import costs 0.5 credits per started megabyte.
  • 31 languages for both import and playback, detected automatically.
  • Voice characters: Narrator, Questioner and Expert presets, which change the register rather than just the timbre.
  • Skip sentences: mute repeated instructions, citations or code blocks without editing the source. It's stored per version as a list of sentence indices, which is a tidy way to do it.
  • Built-in editor: fix imports and format headings inline, before or during listening.
  • Resume playback: the app stores your last-played sentence, so the library remembers your place in every document and the home screen has a "Continue listening" tab.
  • Library: folders, search, starring, an archive toggle, and grid or list views. Each card shows how many versions a document has.
  • Data hosted in the European Union, stated plainly in the footer.
  • Keyboard shortcuts: Cmd+Shift+O for a new document, Cmd+K for search.

How it sits against the obvious alternatives

If you've searched for this kind of tool you already have Speechify or ElevenLabs Reader in mind, and your browser probably reads pages aloud for free. The built-in readers in Edge and Safari are genuinely fine for a single article, but they read a page; they don't give you a library, a place in a document you can come back to, versions, or a chat scoped to what you're listening to.

Against the dedicated apps, Speechdub's distinguishing bet is Eco. Speechify and ElevenLabs are cloud-first, so listening is metered against a plan, and heavy listening means a bigger plan. Speechdub's local mode moves the cost of routine listening to a one-off download on your own machine and charges only for the things that genuinely need a server: cloud voices, rewrites, translation, chat and export. If you listen a lot and don't need much else, that's a meaningfully different shape.

Pricing and plans

Prices follow your location and the figures don't change with it: the same page showed me €6 and showed a US visitor $6.

PlanPriceDocumentsCredits / month
Free0 forever1010
Premium (monthly)10 / monthUnlimited1,200
Premium (annual)6 / month (72 / year, 40% off)Unlimited1,200
Credit Pack7 once, never expiresn/a500 added

Two things matter more than the table. Eco playback is unlimited and free on every plan including Free, which is the load-bearing promise of the whole model. And 1,200 credits is roughly 20 hours of cloud audio a month, which is the number to hold against your own habits. Credits are spent on cloud playback (~1 per minute at 1Γ—), PDF import and audio export (0.5 per started MB or 30 seconds), and document chat, AI versions and enhance (tokens plus 1 credit).

The pricing page, where the plan limits are still rendering as unsubstituted template placeholders.
The pricing page, where the plan limits are still rendering as unsubstituted template placeholders.

Worth flagging: at the time of writing, the pricing page renders its plan limits as literal template placeholders. Free reads "{10} documents" and "{10} credits / month", Premium reads "{1,200} credits / month", the Credit Pack reads "{500} credits added to your balance". The comparison table further down the same page shows the correct numbers, so nothing is wrong with the plans themselves. It's the headline cards, on the exact screen where someone decides whether to pay, that are visibly broken.

The last pricing note is the one I liked most. Speechdub offers reduced pricing for accessibility needs, applied for from inside the app, described as trust-based with no verification required. For a product whose most obvious users are people with dyslexia, ADHD or low vision, choosing not to make them prove it is a real decision and a good one.

The accessibility page, which leads straight to an 'Apply for accessibility pricing' button with no credit card required.
The accessibility page, which leads straight to an 'Apply for accessibility pricing' button with no credit card required.

Who should use it

Speechdub fits well if you're:

  • A parent turning homework PDFs and class notes into something a child can listen to
  • A student who wants a dense chapter re-paced as a lecture, or reviewed at 1.5Γ— before an exam
  • Someone with dyslexia, ADHD or eye strain who reads better with sentence highlighting and their ears
  • A language learner who wants the same text in two languages with native voices
  • Anyone with a backlog of long-form reading they'd get through on a commute

It's the wrong tool if you want a general-purpose voice generator for video or podcast production. There's no multi-speaker timeline, no SSML control, no project structure. Speechdub is built around a document you are trying to get through, and that focus is why the rest of it is as simple as it is. It's also desktop-first for the free path, since Eco mode needs a modern desktop browser, so a phone-only user is on cloud playback and therefore on credits from the start.

Conclusion

There's a version of this product that's just a play button on a PDF, and Speechdub is clearly trying not to be that. Making sentences individually addressable, letting a document be rewritten four ways for four listening jobs, scoping chat to the document, typesetting a chemical equation nobody will ever look at, pricing local playback at zero forever: these are the decisions of someone who has thought about what listening to a document is actually like.

The rough edges are all in the packaging rather than the thinking. Unrendered placeholders on the pricing page, versions that all answer to the same name, an outline that eats numbered headings, a quarter-gigabyte download served as an anonymous spinner: none of that is architecture, it's a week of polish that hasn't happened yet.

What I liked:

  • Eco mode: unlimited, free, local playback on every plan, including Free
  • Every sentence is clickable, so seeking to a specific line is one click
  • PDF import preserved heading hierarchy and typeset a plain-text equation with real subscripts
  • The AI version picker shows what each mode doesn't do, with traits struck through
  • Document chat answered a comprehension question about a 72-minute essay accurately
  • The export modal quotes the exact credit cost before you commit
  • Trust-based accessibility pricing with no verification required
  • Data hosted in the EU, stated plainly

Things to keep in mind:

  • Eco downloads 256 MB of voice models on first play and keeps 380 MB on disk, behind a plain spinner
  • The pricing page still shows {10}, {1,200} and {500} as unsubstituted placeholders
  • The outline renders numbered headings as bare digits, losing the titles
  • Versions all carry the same name, so a translation and a Lecture rewrite are hard to tell apart
  • Rewriting or translating a long document is a background job measured in minutes
  • Free tier's 10 credits is about ten minutes of cloud audio, so it's an Eco trial, not a cloud one

If you have a reading pile you keep not getting to, Speechdub is worth an evening and costs nothing to find out.


Related Articles

Instagram Transcript Generator Review - Every Word of a Reel, No Signup
Reviews

Instagram Transcript Generator Review - Every Word of a Reel, No Signup

We've tested the Instagram Transcript Generator, a free tool that turns any public Reel into a full text transcript without an account.
Postiv AI Review - A Content System, Not a Post Generator
Reviews

Postiv AI Review - A Content System, Not a Post Generator

We've tested Postiv AI, a LinkedIn content system that turns your pillars into a weekly plan, drafts and carousels, and feeds performance back into next week's briefs.
Jamdesk Review - Documentation Your AI Agent Can Actually Read
Reviews

Jamdesk Review - Documentation Your AI Agent Can Actually Read

We've tested Jamdesk, a documentation platform that ships MDX docs, an API playground, and a built-in MCP server so your AI agent can write and maintain the docs with you.
Fomr Review - A Free, Unlimited Form Builder for Small Teams
Reviews

Fomr Review - A Free, Unlimited Form Builder for Small Teams

We've tested Fomr, a free unlimited form builder with a visual drag-and-drop editor, an AI form generator, and workflows that push responses to Google Sheets.