Apple has their own variations on many of their frameworks it seems. For example, PDFKit in Preview isn’t 100% the same as we have.
Indeed.
I’ve also experienced significant degrading of pdfs. But, I’ve found it depends upon which pdf app used for ocr. On an iMac 2015, I used Adobe Acrobat Pro. When I changed to a Mac Studio and started using the latest version of Acrobat Pro, after a while I found the interface bloated and cumbersome - and made worse by its Ai popping up uninvited and unwanted.
For some things I’d used PDF Expert but found it too simplistic. However, earlier this year I found PDF Expert much improved for my needs so have ditched Acrobat Pro.
I’ve not experienced degrading since - but to an extent it does I think depend upon the quality of the original. I have approximately 100,000 pdfs indexed to DT4.
Another possibility is PDF Studio which I also use for some requirements.
I’m using the following settings for the OCR:
I’ve tested DT OCR on many PDFs, which vary a lot in scan quality. My results: it works perfectly. After performing an OCR, the documents are correctly labeled “PDF+text”, I can search in DT by the words, and I see no (visible to me) degradation in PDF scan quality. As for the document size — it’s something I’m really not paying much attention, but one of my databases has hundreds of PDFs and it’s 38GB in size.
p.s. I also saw this “searchable PDFs without running an OCR” in the DT4 release notes, but for myself I’ve decided to always run an OCR on a PDF.
This thread and my question at the top were not about whether OCR works, but rather how indexing works with the Apple Vision Framework in DT4, or rather what is not working at the moment.
