Annotation Files and OCR

I’m a historian who has thousands of newspaper clippings from the early 20th century. Finding and using DT has been critically important to me in setting up a system so that I can find information in those clippings. I’ve recently finished processing 4,368 clippings and then started cleaning up the inevitable errors I had made. I discovered something that I wanted to share. Perhaps others already know this but…

Annotation File

This is what you will see if the text layer of the clipping contains the search word. The flag by the incoming and outgoing link icons in the annotation file indicates that the search term is contained in the annotation file. The lack any such indication for the image file simply means that the image file contains the search term. If you get something like this:

Annotation File - 1

it means that DT could not find the search term in the associated image file. That could be for any number of reasons. If you want to double-check, convert the image file to plain text to see if you can find the search term.

I had not paid attention to the flag icon in previous searches so I was a bit puzzled by what it meant until I started digging. DT apparently adds the flag icon to the annotation file by default, which is great. I just had not noticed that behavior before.

Anyway … if you have a lot of old documents that OCR doesn’t do a very good job with, this will be an interesting exercise. An additional tip: when looking for the existence of an OCR text layer, relying on word count doesn’t always work because some of my clippings only have the copyright information from newspapers.com and no further text. So the word count might be 30 or 50 but that is no guarantee that the search term is included in that number.

Another note that may be of interest: I have created the annotation files by using Gemini and then cleaning them up to correct misspellings of names that either Gemini created or were in the original document. Gemini, like any AI engine, is not to be trusted. It hallucinates and that is often very annoying. Even a prompt that tells Gemini to provide a verbatim transcript does not work. Do not trust AI - it is a useful tool but it is not intelligent in any way, shape or form.

Which “search word”? You didn’t mention anything about search before.

That would be news to me. A flag, as far as I know, indicates that the record has been marked with a flag. By the user or perhaps a smart rule. Again – what search term are you referring to?

Which “image file”? All I can see are “PDF+Text” and “RTF” files. None of them are “image files”. And how does a flag at one time indicate the presence of a “search term” and at the next time its absence indicates the presence of that term? In my mind, that would just mean that the flag does not indicate any such thing at all. But maybe I’m just confused.

I have only four annotations files, and not a single one has this flag. Which does not seem to speak for your evaluation that DT adds it “by default”.

What will be an interesting exercise? Adding OCR? Adding annotations?

chrillek:

Thank you for all of the questions. A search term is just that - a word or phrase entered into the search box at the top of the window. I enter the search term because I want to find all instances of it that appear in my database.

Bluefrog or cgrunenberg can correct me, because I may well be wrong but my take on the flag icon, which I do not recall doing anything to cause its appearance, is that DT, by default, flags an annotation file that has the search word in it. I’m not sure why DT doesn’t flag the image file, which is the original clipping from newspapers.com. I may be using the wrong terminology by naming that file as an “image file” but nothing else came to mind when I created the post. Perhaps there is a better way to describe it. I’m fine with DT not flagging the image file - that would defeat what I surmise is the purpose of the action. But, again, I may be all wet.

“PDF+text” is the clipping imported (not indexed) into DT. It has been OCRed, thus the name “PDF+text” instead of “PDF document.”

I hope that others comment on my supposition that DT flags the appearance of the search term in an annotation file by default. Perhaps I am wrong.

You say you only have four annotation files and none are marked with the flag. The flag does not appear until you execute a search on a term. Find a word or phrase that is common to all four files and see what you get. If I’m right, you’ll see a flag by the annotation files that contain the search term.

As to the exercise I suggested: many, if not most, people think OCR is the cat’s meow - it will solve all of their searching problems. It does not. In a perfect world, with a typewritten document in black text on a white background, I think OCR does, indeed, live up to its reputation. But it struggles with any kind of degraded document. Apple Vision does a much better job but the problem with Vision is that there is no text layer involved and an exported file will not be searchable. ABBYY, the OCR engine included in DT, does a good job in most instances. It has not been particularly helpful for my purposes, though. Gemini, with all of its issues, has been a game-changer for me.

As far as I know, a flag is permanent, that is, it belongs to the record, not to the fact that is or isn’t part of the current search result. You can easily check that yourself by clearing your search and taking a look at the flagged record in the inspector. If the flag is gone, I stand corrected.

Oh well…


I searched for the lovely word “Gebührenbescheid” in a database containing an annotation file which in turn contains that word. And, lo and behold, the record is found, but there’s no flag. At all.

I beg to differ. Apple Vision is a nice thing, but it’s lacking compared with any OCR product I’m aware of. Starting with its limitations language-wise, continuing with its problems to line up words correctly that are on the same row.

You can, if you feel so inclined, add the Vision text as a finder comment to the file. In which case it is at least searchable with Spotlight in the AppleVerse, and that even works for real image files like JPG, PNG and so on.

chrillek:

If I clear the search, the flag disappears - it is not to be found anywhere in DT. Perhaps I clicked on something to make this happen but if I did so, I’m happy with the result because it shows me just how faulty OCR can be. Creating Gemini transcriptions, which became annotation files, was the answer for me. It was a tedious and time-consuming task, but it works for me, regardless of the flag issue. I just like the fact that the appearance or non-appearance of a pair of two files of the same content, one with a flag (annotation) and one without (the imported file) validates all the work that I did. Again, maybe I’m wrong about this “feature” - maybe I’m on drugs - but this is what I see. Maybe bluefrog or cgrunenberg will have an explanation…

chrillek:

I continue to play with this new found feature of DT which for some reason you can’t duplicate. Here is my latest: I entered the search term “Howard Robinson” and found a file that had the flag by the annotation file but there was no corresponding imported file. I looked at the text layer for the imported file and found out why: the search term in the text layer was H o w a r d Robinson, which didn’t match what I had entered, thus no file was returned. Then, I entered “Bazemore,” another name in the same file and got the pair of files - one flagged and one not flagged. I’m not sure why DT doesn’t show a flag by the imported file but I suspect that is because if both files were flagged, the user wouldn’t know why. With the annotation file flagged and the imported file not flagged, DT is telling the user that the search term does not exist in the text layer of the imported file. Conversely, if the term is found in the imported file but not in the annotation file, it indicates that DT can’t find the term in the annotation file. This has happened to me on occasion and when I check the annotation file, the term is included. Why DT doesn’t pick the term up in the annotation file on occasion (it happens rarely), I don’t know. But I’m happy with what DT tells me - it is very useful to me.

No one else has commented on this feature but I think it is pretty doggone cool! It absolutely highlights the limitations of OCR for the user interested in finding information in their files.

I looked into this topic this morning out of interest as I’d never noticed this behaviour and it didn’t make sense, and I’ve confirmed that search does not flag items. As stated in the manual, flags are user defined and the DT interface doesn’t interfere with this (which makes logical sense - flagging files is usually done for a reason, and if we can’t trust flags to appear where they’re meant to the mechanic would be redundant). I don’t use flags at all in my DT setup, and they do not appear in searches.

Furthermore, there’s a flaw in your understanding of the default search which I think warrants me pointing out in case it affects your research: DEVONthink does not undertake fuzzy word searches by default (it’s not looking for similar-sounding words when you enter a search term - it’s looking for an exact match to what you typed). If you search “Howard” and a file comes up in the search results, that’s because the term appears in the document. There isn’t ambiguity here, you asked DT to look for a term and it found it and shows it to you. HOWEVER, you can enable fuzzy word searches by clicking on the magnifying glass in the search bar, along with searching for related words, and it’s possible you have one or both enabled, in which case DT may be surfacing a file it thinks is related to your search term but where the term itself doesn’t appear. The heat map to the left of each search results tells you how likely DT thinks it is that the file matches what you’re looking for.

If these files are flagged and you don’t know why, it’s either because you’ve been accidentally clicking the flag icon on files (possible via the main window, the inspector or the toolbar), or because you’ve have a rule or script running that’s applying a flag when criteria are met.

3 Likes

MsLogica:

Thank you for your kind and thoughtful reply. You are correct. When I clicked on the Annotation database at the left side of the main window and then selected “Mark,” I saw that the option to flag had been checked. How I did that, I do not know. It was not deliberate.I unchecked “Flag” and all of the flags disappeared.

I will say this, though. Flagging all of the annotation files did help me to understand how poorly OCR works on old newspaper clippings. I’m not criticizing DT, ABBYY or OCR in general - I’m just saying that for my purposes, OCR does not work very well. Gemini has been a game-changer for me. Gemini is alleged to be the best of the LLMs for transcriptions and it does work pretty well but it has to be scrutinized also. It has its faults and quirks and is probably 95% accurate.

I rarely use fuzzy word searches and I am aware that is not the DT default. I clean up the annotation files as I go, correcting both Gemini errors and typographical/reporter errors transcribed by Gemini from the original import. Gemini, like all LLMs, hallucinates. Sometimes it is a “Wow!!” error and other times just a minor mistaking of a ‘c’ for an ‘o’ or an ‘e.’

When I find a file that DT reports does not have the search term in it, I sometimes look at the text layer and almost always find out why DT can’t find the term. Either OCR spells it differently than what is in the imported clipping, has spaces between letters or is simply not there is invariably the case.

Thanks again for putting me back on the path to discovering more of the power of DT. It is an essential tool for my work.