In DEVONthink > Preferences > OCR uncheck the preference “Enter metadata after text recognition” and DEVONthink will stop pestering you.
You might want to hold off merging your documents because for some uses (e.g. See Also & Classify) the discrete OCRd pages might have an advantage.
A script that looks in your catalog and determines how to merge a given document is going to be complex and likely to fail from time to time, IMO. Personally, I wouldn’t trust it. The worse thing is that the catalog is wrong, or things the TIFFs are out of order or named incorrrectly, and the script would start munging together pieces that don’t belong together, and you’d need to browse every document to figure that out. ![]()
I’d suggest you do not import all the TIFFs into DEVONthink, index them instead, and run OCR against the indexed TIFFs. The resulting PDF+Text will be imported into the database – which is what you want – but not the TIFFs. You won’t need the TIFFs in your database – right? Make sure the setting in Preferences to trash the original is OFF, so your valuable collection isn’t at risk. I assume you have at least two backups of the TIFFs - one of them is off site?
Before working this process with 24,000 documents – try it on a dozen or so. DEVONthink’s ABBYY OCR software can render pretty lousy resolution on graphics - regardless of the resolutions set in OCR preferences, in my experience. Over here, I prefer Acrobat X Pro for OCR. YMMV.
See also this recent dialog: indexing v. Add to database question