Lately the autodetection of language in OCR has caused a bunch of problems. A lot of receipts I have been scanning get detected as Indonesian or Vietnamese. But it doesn’t seem possible to force the language back to English, so chat suggestions and similar screw up all the metadata. I enabled the Language field in Preferences → Data, and that makes it appear in the Data panel in the sidebar, but nothing I change there has any effect in the Language column of the actual database.
In the screenshot you can see I set the document to English but the Language column remains Vietnamese, and all chat suggestions are permanently locked to Vietnamese. Nothing I can think of can convince the system that the receipt is in English.
DevonTHINK 4.3.2
Tested with Gemma-4-e4b and Gemma-E26-A4B-QAT through LM Studio
Recent versions of DEVONthink use the latest features of macOS to detect the language of documents. Unfortunately this seems to be less reliable, the next release will therefore switch back to the formerly used solution.
In Settings > AI > Summarization it’s possible to choose the desired language for summaries at least. Other cases would require custom prompts (see e.g. the Rename to Chat suggestion batch processing config)
Yes, it seems strange that a database field that has that much of an effect on so many operations is beholden to an automated process with a notable error rate.
To say nothing of my discovery of a SECOND language field in my database that seems to do nothing.