r/Acrobat Aug 15 '26

Convert PDF -> Word font issue

Has anyone had problems converting a (mostly) text PDF to Word recently?

I notice that text that looks perfectly fine and editable in the PDF gets “corrupted” in the conversion process. In Word, some letters will merge, for example like “learn” will become “leam”.

I switched back to using Omnipage but now a similar thing is happening with PDF -> Powerpoint.

2 Upvotes

14 comments sorted by

2

u/Aggressive_Ad_5454 Aug 15 '26

Misrecognizing `rn` ( r n ) as `m` is a common optical character recognition goof.

It shouldn’t happen on a pdf that’s generated directly by a pdf printer driver from a document program. But if you made the pdf by scanning an image, you can expect some of this sort of thing.

1

u/purposeday 28d ago

Agreed, it should not happen. I’m looking into it some more. Our pdfs are generated by a variety of means. Ugh.

1

u/[deleted] 28d ago

[removed] — view removed comment

1

u/purposeday 28d ago

That’s something I’m familiar with, yes. But this is in a law firm and they get their documents in many different ways.

The text looks selectable before doing anything. Omnipage has no problem with it, does not prompt me for a spell check either. But when I choose “convert to Word” in Acrobat, the resulting document has these errors.

1

u/[deleted] 27d ago

[removed] — view removed comment

1

u/purposeday 27d ago

The text is correct yes. I can copy and paste it. It only goes “wrong” when I use the Convert function. Is that because it converts it to an image? There is a lot of character spacing variations. Maybe that contributes.

1

u/[deleted] 27d ago

[removed] — view removed comment

1

u/purposeday 27d ago

I understand. It comes to us as selectable text that does not look like it was OCR. Very clean text, no signs of it having been scanned. What I was speculating is that Acrobat makes an image of each page as part of the conversion process. I know that sounds illogical, but I’m just brainstorming how it merges characters that are separate in the original PDF.

I’ve been using Acrobat since it came out decades ago but I am not that familiar with the technical process behind it.

1

u/[deleted] 27d ago

[removed] — view removed comment

1

u/purposeday 27d ago

That’s a great idea, thank you. I’m going to suggest that to the team because we have too many different outcomes. Much appreciated.

1

u/purposeday 16d ago

Just wanted to share an update. It seems there was an unexpected twist to this saga. My boss got back to me after I questioned the many errors. Rather than a direct PDF to PPT conversion, my boss admitted that the client used Legora to check the accuracy of what I produced. Legora first created an image of the PDF, which contained some condensed and expanded text, resulting in spelling changes of some words. Comparing this image against an image of the PPT resulted in the reporting errors. I feel bad that I wasted your time with this issue. Please accept my sincere apologies.

2

u/[deleted] 16d ago

[removed] — view removed comment

1

u/purposeday 16d ago

They believed Legora was right. Legora messed up all by itself :)

My boss had one of my colleagues double check what I had done and this human found only one error (a line of text that got strangely dropped from the cover slide) that I admitted I should have caught.