(NB. Not a question about Importing vs Indexing)
I work with book scans (from 1800’s) that often do not have the text recognized. In days past, I would run DT’s built-in ABBYY and convert them (Kind: PDF Document) to a format that is has the words broken out in the Info panel (PDF+Text).
In more recent editions of DT – and concurrent with the ability to transcode audio speech to text – I have started to notice that there is an Indexing phase for imported PDFs. It does not use the OCR helper, it uses another one called DevonTHINK Helper. I am presuming that this an alternative workflow where DT is leveraging built-in capabilities. In my experiment, I imported ten “PDF Document” and after some time, they became “PDF+Text”
I am in a position now where I would like to go through my sizable databases and bring many “PDF Documents” up to scratch. It seems that ABBYY has some licensing issues on doing huge batches of PDFs, and so I am wondering if I can trigger the OS-based indexing on documents already inside DT.
As a last alternative, I could export my files and then re-import them. This could be advantageous as other file types that I’ve accumulated could get a more modern indexing.
I would appreciate any thoughts or suggestions, or even a pointer to the menu item that I’m missing to accomplish this task. Thank you!