Best Practices for [Finder vs DevonThink4]

I use DEVONthink as my primary Personal Knowledge Manager. URLs, articles, PDFs, docs, etc)

I have a robust tag hierarchy, and use both tags and smart groups to organize items

IN ADDITION, I also have a 2TB hard drive and Finder contains a robust folder hierarchy of files.

I have scans, PDFs, photos, videos, bank statements, utility bills, contracts, legal documents, etc etc. Some in DT. Some in Finder.

The duality is frustrating me, because I have to constantly consider “hey will this thing live in my curated folder hierarchy in Finder? or should it live in DT?”

I have had bad luck with Indexing in general. Have to constantly re-index to force updates to changes. Every time I Empty Trash in DT I have to wonder if I might have deleted an Indexed Finder file, in which case I should not “delete all”. It’s a bummer.

There are tradeoffs, 2 simple examples:

+ Things in Finder get backed gracefully with TimeMachine and Backblaze. DT’s backup is for the monolithic dbase.

+ very large files seem to bloat DT’s dbase, so more natural to leave in Finder, and index if needed?

--

I am curious for other folks’ systems. How you power users delineate:

1. what you import to and have live in DT

2. what you have live in Finder

3. what you leave in Finder, but also Index into DT

--

I am using DT4, with all features, and my Macs are all on the latest OS.

2 Likes

For me, generally everything is imported and lives in DT
I only index from Finder if there is a reason
example: a google drive folder shared with others

Exceptions
. music files; I use the native music app
. photos; duplicate DT and the native photo app; cross-linked
. geneology data; duplicate DT and dedicated app MacFamilyTree

>DT’s backup is for the monolithic dbase
The database is not “monolithic”
It’s a Mac OS Finder package, with our files stored in a folder

9 Likes

I index anything I am having directly manipulated/edited/whatever by AI. I have a feeling that via MCP you could edit internal files and not run into too many issues, but that is not a best practice and I don’t want to risk something catastrophic happening. But I do not have any index issues and haven’t really over the years. Little problems here and there but nothing major, and not lately.

1 Like

Nearly everything goes in DT for me, with a couple of exceptions that are clearly defined and don’t leaving me wondering where something should go (images are in Finder, files for a couple of specific apps are in Finder). There’s no reason for these to be in DT - I don’t do any management of my images that would warrant DT tools, and the files for specific apps are unique to those apps.

So in your use-case, scan, PDFs, bank statements, contracts, legal docs, etc. - it’s all in DT. I have a database for life admin which has all that stuff (I use groups to organise it all, but that’s not necessary and is down to personal preference). I’d leave the images and videos in Finder, but there are folks on here who store this in databases too and it depends on what you do with them.

For me, having everything in DT has the added bonus that doing a manual backup of all my files is very easy, since I just create an archive my databases every few weeks, which produces a zip file I store on an external drive. It’s music quicker than creating a manual backup of Finder.

9 Likes

What is very important to me lives in DT. In many databases, segmented by interest (e.g., Photography, Cooking).

Legacy (stuff I want to keep but is not that important, or I don’t need ready access, etc) lives in the Finder ( 8 TB). What does that include? Decades of work stuff, archived apps, assorted stuff I’ve read, saved (mostly pdf)

A special category is files that are very large (by that I mean 50 - 100 MB or more) - that I do use - e.g., manuals, or large scanned reference books. These live in the Finder, but have pointers in DT. By pointer I mean a link to these, in the right database, placed in an md or rtf file in DT.

I back up to several external drives, Backblaze, and even (very specific, e.g., images) M-Discs. And Time Machine is on as well.

5 Likes

For me the biggest hesitation about DT by far is, as you noted, the lack of granular backup and recovery.

In Finder it’s trivial to search for and recover files or folders via Time Machine, Carbon Copy Cloner, etc. In DT it’s very challenging unless you know exactly when the problem occurred so you can recover the DT database of as of that date.

For this reason I still keep essential documents, e.g. archived tax returns, in Finder. I hope DT gets granular backup and recovery at some point, not sure how difficult that’d be to implement given DT’s architecture.

3 Likes

thanks! yes this “just create a copy every few weeks” is not optimal for me. My databae is many GBs, and creating a snapshot daily or weekly just consumes too much disk space for the backups themselves. And Finder’s TimeMachine (and I use Backblaze for offsite) run continuously and incrementally by file, so it’s crazy simple to back to a version of a file on a particular date or time.

thanks. That’s interesting to index docs you’re accessing via AI. I’ve been experimenting with ollama within DEVONthink, though I haven’t yet found a compelling internal use.

can you run incremental backups on the Package?!? If so that is awesome – it didn’t used to be the case. :crossed_fingers: this is now possible and this makes my decision much easier

1 Like

Sure about that? DEVONthink databases are actually individual files inside of a macOS “package” presented as a “file” in Finder, but it’s thousands of folders and files. I’m pretty sure, and others can confirm (?) TimeMachine only copies files that have changed since the last backup run. The entire database is not backed up on each run, AFAIK. And with TimeMachine you can restore databases to the date/time in the TimeMachine backup.

Perhaps you are thinking of the DEVONthink database archives which are indeed zip-ed versions of the entire database. Me, I take these every so often when I think of doing it and indeed they are backed up also with TimeMachine (and not excluded).

Overall, I use a 3-2-1 backup regime for my machines.

3 Likes

thx, clarifying - confirmed the package components do backup via TimeMachine.

in order to Restore, however, I need the whole DB to navigate to find the file b/c normal backups of the package do not maintain the Group nor the tag hierarchy, so browsing to find a specific file to recover is hard for a high volume of files

or is there a way to find and apply the metadata in TimeMachine restore? if that’s the case I’m able to just put everything in DT!

I only remember doing a DEVONthink database recovery from TimeMachine once. If I were to do it today, I would restore the entire database from the date/time wanted into a separate and new database and after restored, do a Verify and Repair to check. I would not dig into the DEVONthink package and selectively restore.

Perhaps re-read the section in the “DEVONthink Manual”, “A word about backups” which gives sage advice carefully written by @bluefrog.

5 Likes

I’m another of those for whom everything lives in DEVONthink apart from photos and music – with the exception of a few images that I want to keep together with related material. The only things I index are my Obsidian vault and my Bookends attachments folder.

At the moment I have a single, large datase for everything. I’ve experimented with multiple databases and I can see the advantages, but at the moment I’m happier with lobbing everything into one database. That way I know where things are. My backup strategy is to use Time Machine, Backblaze and Arq as well. The latter two are not that expensive and having both feels just a little bit more secure.

5 Likes

Yes, I have automated hourly incremental backups via Arq Premium
Two backup copies; external HDD and the cloud
Also TimeMachine, to an external HDD

navigate to find the file

Search on filename works

4 Likes

That’s why DT creates a zipped file. But in any case, that’s why I do it on an external drive. It’s not my only backup. But if I’m anticipating a scenario where I will want to retrieve one specific file and am not going to load the whole database from Time Machine, I’d just use these manual backups instead. Time Machine for me is for a catastrophic failure, not me losing one file.

1 Like

(This is hypothetical for me, I’ve never used my backups. But I assume I’ll appreciate the choice :joy:)

1 Like

After lots of experimentation, mistakes, and suboptimal setups, I have settled on the following breakdown.

In DT:

  • Working database. Contains all current financial records, research data, temporal and structural data. 177GB stored on internal drive. NB: This should not exceed 250GB to ensure responsiveness.
  • Music database. As a musician, while I use other apps for playback, DT lets me replicate extensively to create practice lists and multiple versions of songs and backing tracks. This DB is 2.7TB stored on a USB SSD, where responsiveness is not as important.

In Calibre:

  • All ebooks (EPUB, PDF) without attached .zip files, about 500GB, stored on internal SSD.

In Finder:

  • Software archive and media archive, both about 3TB, both stored on networked SMB drive. I index these using NeoFinder.
  • Ebooks with attached .zip files or audio books, indexed with NeoFinder.

Thanks all for the engagement on this thread. Feels like there isn’t a right answer here.

If I want the efficiency of file-level, incremental backups and restore, I need to keep in Finder.

and if I’m okay taking large snapshots of my entire database, it might as well be in DevonThink4.

This leaves me wondering if there is any value at all in TImeMachine’s incremental backups of the content of the database package. If not, for example, should I exclude the .dbase files from TimeMachine? Curious if folks have thoughts on this.

I’m wondering what problem you are fixing?

How often do you expect to restore individual items, be they in the macOS file system and displayed by Finder (e.g. not “in” Finder) or in DEVONthink and displayed by DEVONthink? As backups by TimeMachine are in background and automatic, what is your concern? What problem are you fixing?

I’m not sure what you mean by “taking large snapshots” unless you are talking about creating Archives of the DEVOThink databases.

as I said above, I backup entire macOS file system using TimeMachine. Actually, use 3-2-1 Backup Regime with multiple local and remote copies (Arq, NAS, etc.). I don’t often restore as luckily failures don’t happen that often.

thanks again. IYKYK. You share above that you’ve never had to find and restore a time snapshotted file. When you have to, you have to. TimeMachine and BackBlaze both make this trivially easy and granular.