I’m using DEVONthink To Go version 4 on my iPhone to scan documents. The OCR is set to make scans searchable and to move the original to the trash. I also have the quality at 100%.
The results I am getting are very, very poor. Here is an excerpt from a typical scan:
Why would the OCR process render the scan unintelligible? Is there a setting to improve this?
The non-OCR’d image is good enough, but notably seems to be default to a colour document scan. I found where to adjust this to a black and white scan (a button on the scanning page you quickly adjust before the scan takes place).
My next processing of such documents is in DEVONthink on the Mac. I guess an alternative would be to scan the document with OCR off on DTTG, sync to the Mac, and then somehow have the Mac OCR the document. Perhaps I could create a smart rule looking at the inbox to OCR any document that is imported in the inbox that’s a PDF and not a PDF + Text? I am concerned that “on import” may not be the right setting because I’m not sure if a sync from DEVONthink To Go to DEVONthink on the Mac is actually an import. Is it?
Which version of DTTG are you running? Whilst I haven’t been able to reproduce the problem scanning several receipts it is likely to be caused by the OCR’s image preprocessing. What are your OCR settings?
That’s a pretty hard scan to get recognised as OCR often relies on much higher contrast between text and background.
The “bad quality” scan is simply what the OCR is recognising as dark enough to be text – it needs more help.
You might need to adjust it to make the text as close to black and the background as close to white as possible. Perhaps just increasing contrast will do the trick.
Magenta is the most fugitive pigment in color printing. That’s why when you see old posters in the windows of abandoned storefronts, they’re mostly yellow, blue (technically cyan), and black. Sorry, my printing history is showing
My key concern with the OCR process on DTTG4 on the iPhone is that the image is converted to something that is unintelligible to the human eye.
As to DEVONthink’s smart rule functions such as date and amount recognition, I find this reads well from a photo quality document exactly like the one I posted at the top (the human readable one) (that’s been OCRd in DEVONthink on the Mac). It’s just how DTTG is converting the image, making it illegible to the human eye which is my concern.
I note that other apps don’t convert a human readable image scanned in a grey-scale or photo setting to an unintelligible to human image on being OCRd. I appreciate that DEVONthink is not a specialised scanning app and may not have the capability of those apps dedicated to such a function. I don’t expect that to be the case. I was merely wondering whether a change in setting somewhere might resolve the issue. If not, I can do the OCR in DEVONthink on the Mac, which doesn’t seem to have this issue.