Poor scanning results DTTG4

I’m using DEVONthink To Go version 4 on my iPhone to scan documents. The OCR is set to make scans searchable and to move the original to the trash. I also have the quality at 100%.

The results I am getting are very, very poor. Here is an excerpt from a typical scan:

If I scan the same document with Make Scan Searchable off, the result is better as follows:

Why would the OCR process render the scan unintelligible? Is there a setting to improve this?

The non-OCR’d image is good enough, but notably seems to be default to a colour document scan. I found where to adjust this to a black and white scan (a button on the scanning page you quickly adjust before the scan takes place).

My next processing of such documents is in DEVONthink on the Mac. I guess an alternative would be to scan the document with OCR off on DTTG, sync to the Mac, and then somehow have the Mac OCR the document. Perhaps I could create a smart rule looking at the inbox to OCR any document that is imported in the inbox that’s a PDF and not a PDF + Text? I am concerned that “on import” may not be the right setting because I’m not sure if a sync from DEVONthink To Go to DEVONthink on the Mac is actually an import. Is it?

Which version of DTTG are you running? Whilst I haven’t been able to reproduce the problem scanning several receipts it is likely to be caused by the OCR’s image preprocessing. What are your OCR settings?

That’s a pretty hard scan to get recognised as OCR often relies on much higher contrast between text and background.

The “bad quality” scan is simply what the OCR is recognising as dark enough to be text – it needs more help.

You might need to adjust it to make the text as close to black and the background as close to white as possible. Perhaps just increasing contrast will do the trick.

Sean

1 Like

Very true and the inks used (called fugitive pigments) and non-archival paper user in receipts make scanning them as soon as possible a must.

How this isn’t, at the very least, a song title, let alone a book or film title is beyond me.

3 Likes

You could try to scan to greyscale, not B/W.

Magenta is the most fugitive pigment in color printing. That’s why when you see old posters in the windows of abandoned storefronts, they’re mostly yellow, blue (technically cyan), and black. Sorry, my printing history is showing :wink:

3 Likes

Book spines, too.

I think most of that receipt is below 50% grey so would need brightness and contrast settings adjusted, too (less brightness, more contrast).

Sean

1 Like

4.2 (Chandra) (18220)

Yes, it is not an easy scan.

My key concern with the OCR process on DTTG4 on the iPhone is that the image is converted to something that is unintelligible to the human eye.

As to DEVONthink’s smart rule functions such as date and amount recognition, I find this reads well from a photo quality document exactly like the one I posted at the top (the human readable one) (that’s been OCRd in DEVONthink on the Mac). It’s just how DTTG is converting the image, making it illegible to the human eye which is my concern.

I note that other apps don’t convert a human readable image scanned in a grey-scale or photo setting to an unintelligible to human image on being OCRd. I appreciate that DEVONthink is not a specialised scanning app and may not have the capability of those apps dedicated to such a function. I don’t expect that to be the case. I was merely wondering whether a change in setting somewhere might resolve the issue. If not, I can do the OCR in DEVONthink on the Mac, which doesn’t seem to have this issue.