Re-Indexing Imported PDF Documents

(NB. Not a question about Importing vs Indexing)

I work with book scans (from 1800’s) that often do not have the text recognized. In days past, I would run DT’s built-in ABBYY and convert them (Kind: PDF Document) to a format that is has the words broken out in the Info panel (PDF+Text).

In more recent editions of DT – and concurrent with the ability to transcode audio speech to text – I have started to notice that there is an Indexing phase for imported PDFs. It does not use the OCR helper, it uses another one called DevonTHINK Helper. I am presuming that this an alternative workflow where DT is leveraging built-in capabilities. In my experiment, I imported ten “PDF Document” and after some time, they became “PDF+Text”

I am in a position now where I would like to go through my sizable databases and bring many “PDF Documents” up to scratch. It seems that ABBYY has some licensing issues on doing huge batches of PDFs, and so I am wondering if I can trigger the OS-based indexing on documents already inside DT.

As a last alternative, I could export my files and then re-import them. This could be advantageous as other file types that I’ve accumulated could get a more modern indexing.

I would appreciate any thoughts or suggestions, or even a pointer to the menu item that I’m missing to accomplish this task. Thank you!

You are probably using Apple’s Vision framework. That does recognize text, but the text is not added to DT’s index, afaik. There has been some discussion on this in the forum, so perhaps a search here will provide more in-depth insights.

Thank you for the feedback. I will go through other posts to see if I can glean some information about this.

Addressing a point you raisedI spoke a little too quickly about my “experiment” – after the indexing they all had words, though only about half of them changed their Kind to PDF+Text. But all of them had their Concordance populated and were searchable.

(I am assuming search will consult the Concordance and does not do something like OCR-on-the-fly)

The Vision indexing is not a replacement for OCR just as Apple’s Live Text is not (despite what people claim).

  • Can you find these documents in a database search? Most likely, yes.
  • Does it make the found files searchable? Yes, but the search occurrences may not show as highlighted text.
  • Are these files searchable after sending them to
    someone? No as there is not an added text layer.
1 Like

This is very clarifying. I agree, the ABBYY Finereader is a different workflow with different value. I have two scenarios where 1) I want to keep the PDF original yet search it and 2) make a sharable PDF that has high quality OCR.