Why extracting data from PDFs is still a nightmare for data experts

misk@sopuli.xyz · 2 months ago

Why extracting data from PDFs is still a nightmare for data experts

FaceDeer@fedia.io · 2 months ago

This is silly.

Whether it’s “silly” or not is irrelevant, the problem described in the article is real. I have seen innumerable PDFs over the years that were atrocious when it came to the use of those accessibility features, the format’s design factors in to how people use it and people use it terribly. If plain old OCR were enough then this wouldn’t be such a problem.