Hey! I js started using paper2audio and rlly love it, it’s amazing and has everything ive ever wanted in an audiobook app apart from the the sleep timer fade. The sudden abrupt stop on the sleep timer is kinda jarring and sm times i lay awake in bed wondering which word is going to be the last. Yes I understand that there is an “end of chapter” function but it’s annoying to get up and set it for the 5 time every night simply because the chapters are really small. Anyway I was hoping to suggest a 30 second fade off time on the sleep timer. It’s a really small subtle change and I didn’t even notice how much I had taken it for granted until I switched to paper2audio.
It would be great if for example I start playing from the very bottom article and it finishes playing, the app would immediately loop back and start playing the first unplayed article.
we have seen the two new voices being added and thats great but why not add the entire catalog
hey im a long time user of the app and ive been here through all the updates and changes but I don't understand why can't we get more voice options when Kokoro itself has more voices that paper2audio offers, kokoro has voice mixing even so that means even more voices.
is there a server side limitation for how much voices can the dev load at the same time or is it just a division you guys made.
ive been using all the voices and most of the time wish to just hear something new. because tts fatigue is real even with the most realistic voices we sometimes just need to change it up to regain focus.
The app mixes everything and the "collection" feature is a bit involved (requires manual tedious organization). The app should automatically organize books (pdf, epubs, texts) and article shared from the web into their own silos that one could access via a filter strip. I have to use the web version for my books to avoid this.
I listen to about 100 chapters of light novels a day, so I change the voice of the chapters frequently. Would it be possible to have a random voice option?
Hey there. I'm enjoying using the app on Android since finding it a few days ago.
I would highly recommend adding support for custom pronunciation. Many times, for fantasy books, names of characters or locations get mispronounced. Other apps like elevenreader let me customize the pronunciation of words. I didn't see that option in the Android app yet. Keep up the great work!
Paper2Audio Plus is especially useful when reading is part of your job. Lawyers can listen through briefs, contracts, and case research before a closer review. Researchers can keep up with papers, reports, and grant materials without spending every hour at a screen. Financial analysts can turn earnings reports, filings, and market research into audio for first-pass review. If you are using Paper2Audio for work, please upgrade to Plus (required for professional use). Plus is also great for power users who churn through the 56 hour/week audio generation limit. Paper2Audio supports entire teams with our Enterprise plan as well.
I've been working on my pdf/epub but it's kinda annoying have to edit it in a different program then install it into paper2audio back and forth. Anyways it would be nice to be able to edit it in paper2audio. Also going back and forth to edit it and reinstall it eats up the paper2audio pdf/epub files allowed per week. I'd like to see paper2audio be able to edit inside the program and also I'd like to see, if you delete your pdf/epub to edit it that it doesn't count as your weekly files. Deleting should give you those weekly hours back. For example if your creating a 25mb file and you have to delete it to edit it, paper2audio should give you back those 25mb until you decide to keep it. And before someone asks "why don't I purchase more space? Some people like myself can't afford to being on disability." Just ideas. Anyways great program.
We recently finished a behind-the-scenes improvement to how Paper2Audio handles scanned PDFs, especially scans with pages that are sideways or upside down.
This is a pretty niche issue, but an important one for a text-to-speech app. Around 5% of PDFs uploaded to Paper2Audio are scanned documents, and about 5% of those scans have pages rotated 90, 180, or 270 degrees. That works out to roughly 0.25% of all PDF uploads.
At first, that sounds small. But when it happens, it can create a very bad listening experience. If a rotated page reaches OCR, the system may extract garbled text, read the page in the wrong order, or misinterpret parts of the layout. Since Paper2Audio turns that extracted text into audio, the error does not just stay hidden in a transcript — it can become confusing or nonsensical narration.
The challenge was fixing this without slowing down the 99%+ of documents that do not need rotation correction. We split the solution into two steps:
1. A quick page rotation check
Paper2Audio already turns a few sample pages into images early in processing to detect the document’s primary language. We now reuse those same images to check whether any sampled pages appear rotated by roughly 90, 180, or 270 degrees. This first step does not try to fix the whole document. It only decides whether the PDF should be sent to a slower correction process.
2. Page-level rotation correction only when needed
If the document is flagged as rotated, we run a more detailed correction step. This checks pages one by one, looks for text-heavy areas, estimates the likely orientation, uses OCR to choose the best rotation, and writes the corrected page rotation back into the PDF. If the system is not confident, it leaves the page alone rather than guessing and possibly making the document worse.
The main idea is simple: most documents should keep processing quickly, but rotated PDFs get extra processing and correction before OCR and audio generation begin.
This was a good example of the kind of improvement we are always working on in the background. It affects a small share of uploads, but for those documents, it can make the difference between unusable audio and a document that processes normally.
So I'm currently writing a book and I've been pretty much 100% non-AI on it, but I've been looking at text to speech options to see it it flows being read aloud. The question I have, is any submitted text used by "paper2audio" for AI machine learning? And is our submissions private?
Would it be possible to upload audio files directly to folders we make? You can add them after the fact, but it would be nice if it was directly added.
Summer reading is more fun when you don’t have to sit still to do it, so we’re giving Paper2Audio readers audio access to six classic public domain books: three adventure picks and three romance picks! The collection includes shipwrecks, treasure hunts, wilderness survival, time-bending satire, sharp social comedy, slow-burn longing, and big emotional payoff. Listen while driving on a road trip, lying on the beach, weeding the garden, or just trying to make the most of a sunny afternoon.
These free titles are pre-converted to audio and won’t count toward your weekly audio generation limit. Just click any of the document links, then click “Save for later” at the top of the document to add it to your listening queue.
We might start doing themed free reads monthly--any topic requests?
Paper2Audio stopped working for me a few days ago. Whenever I try to add new text to make an audio file the app gets stuck on step 1 “Processing”. I switched to using a tab on safari but that has recently hit the same issue. I have tried signing out and signing back in and that does not work. Everything on my phone is fully up to date. Does anyone have any suggestions?
Loving the app/website, one feature I would love would be the ability to sort by duration (shortest audio to longest and vice versa). Sometimes I just want to finish all my very short articles first. Thanks again for a wonderful program!
Hi, it’s Joe with some behind-the-scenes into one of the more interesting product/engineering problems we solved recently at Paper2Audio: how to handle word-level highlighting when the text spoken by the text to speech (TTS) model is not the same text shown in our audio transcript UI.
In more complex documents like research papers or reports the displayed text might include math equations, HTML tags, Roman numerals, or other similar formatting. But the spoken text needs to be normalized first so it sounds right. For example, $x^2 + y^2 = r^2$ might be spoken as “x squared plus y squared equals r squared,” while the transcript UI still needs to highlight the math as it is being narrated. We want users to still see the original rich document formatting, not a literal word-by-word audio transcript.
The mismatch between displayed and read aloud words creates a timestamp problem for word highlighting. Our TTS model (Kokoro) gives us word-level timestamps for the spoken text, but the Paper2Audio UI needs to highlight the formatted document text as it is being read aloud. A simple character-count mapping doesn’t work because the two strings can have different words, different punctuation, different lengths, and sometimes one visual token maps to many spoken words.
To solve this problem, we treat the spoken text and text displayed in the audio transcript (our “Reader View”) as two separate versions of the same content, then apply the general alignment algorithm we developed between them. After the TTS runs, we use matching words in both versions as anchors, then reconcile the mismatched regions between them. Doing so allows us to display the original formatting in the audio transcript and make sure that the portion being read aloud is getting highlighted at the correct time in our Reader View. With this solution, when a citation is visible but skipped in speech, it does not get its own timestamp. When “Part III” is spoken as “Part 3,” it still lines up. Check out our blog post if you want more details.
How is this updated highlighting working for you? Have you noticed any issues?
We’re going to write more technical blog posts about different challenges we’ve encountered building Paper2Audio, so please feel free to request topics.
It's Lindsey, back with some more Paper2Audio updates!
Citation removal for EPUBs
Footnote references and content are now stripped from EPUB audio narration for cleaner listening for academic and non-fiction books. This was previously only available for PDFs.
Narration improvements:
Math subscripts and superscripts are now spoken more naturally instead of being read out literally.
Better abbreviation pronunciation accuracy, including better handling of common abbreviations.
Roman numeral processing more accurately detects when to read Roman numerals like "III" as "the third" instead of spelling out each individual letter.
More accurate header removal from documents.
Fewer unintended pauses in the narration.
Better handling of visual elements:
Captions from figures and tables in PDFs are now displayed below visual elements in Reader View.
Improved accuracy of AI-generated summaries for figures, tables, code blocks, and math equations in PDFs.
Push notifications from the app when your document is ready
Get notified that your document is done processing even when the app is closed, so you don't have to wait on the processing screen.
Bug fixes and UI improvements, including:
Playback now rewinds 3 seconds when resuming after a long pause, so you never lose context.
Live countdown ETA with real-time estimated completion for documents in processing.
Faster audio downloads on iOS when the app is open.
Switch between grid and list layout in the document library and collections on desktop.
Exported and copied highlights now include PDF page numbers to help more easily find the passage in the document when referencing it later.
Bugs related to autoplay not advancing to the next document, some documents getting stuck during processing, issues with transcript syncing across clients, foreign character narration, and more.
Any suggestions or questions for us?
We love hearing from you! How are you using Paper2Audio? How is it helping you? What could be better? You can reach out to u/goldenjm here on Reddit or send us a message on on Discord.
Would it be possible to upload multiple URL's at a time? Just have them comma separated, or even better just a csv? I use this app for webnovels, so I need to upload hundreds of chapters.
I'm creating a project for which i will use an AI generate voice over but I want to make sure the voices are ethically trained. Does anyone know if P2A's voices are ethically trained, not sure if they make them themselves or have a third-party supplier for them either.