Yes, Agree. Not sure what point you are making, though with regard to using files imported into DEVONthink, or indexed, or just left in the macOS and interact with them with Finder. Please clarify and maybe we can give better advice.
Imo, that’s a misconception. DT databases are just folders. So, any backup program worth it’s name can do a file-level, incremental backup and restore with them.
we talked about this above. the problem is restoring from within the package does not display the metadata to allow you to browse by group hierarchy or tag hierarchy. If you know the name of the file, you can search for it. but with hundreds of thousands of files it’s not uncommon that you might have e.g. multiple files matching a name search.
I have had to do file-specific restores from history with timemachine backups, and it’s very easy. it wouldn’t be for a file buried inside my big DT database. that’s just a tradeoff in weighing Finder vs DT for different types of docs
It’s still possible to restore the complete database to a separate folder, open it in DT and copy the relevant file(s) to the correct place.
I’m not saying that it’s a nice solution. But saying that you need to keep files in Finder for file-level, incremental backups and restore is simply not correct.
One way to ease the burden of identifying the correct file is to use a naming convention that include for example the date.
It is also possible that different files outside of DT are named identically, stored in different locations, and the user does not remember which is which… One such example would be intentional copies of different states of the same document (doc-1.pdf, doc-2.pdf, doc-3.pdf…)
thx. I get what you’re suggesting. we’re not speaking the same backup/restore language.
it would be not too dissimilar to say I can take a full backup of my entire Mac hard drive, keep it around somewhere, and if I ever need a file, just restore the whole hard drive from a date around then and then go find the file. IMO that’s not a backup/restore solution, nor is keeping an entire DEVONthink database.
Incremental backup solutions exist to make this automatic and easy.
anyway, for me, this is helpful and a helpful delineator to what lives in Finder vs. what lives in DT. This has been a very helpful discussion appreciate everyone sharing their workflows with us
The good news you can do what you want, but I’m not convinced you as yet understanding how the backups work and how, IMHO, un-important it is when considering putting files in or out of DEVONthink and using it vs. Finder to view those files. But up to you, but it doesn’t appear you are landing on what you seek–“best practice”.
FYI, I just did a test to restore my biggest DEVONthink database (30.58 GB, 23 Group, 46,489 files, 1,638,849 words, …). Not small but perhaps not big in DEVONthink world.
40 seconds total restore from a TimeMachine backup on a Synology NAS on my local network. I’m guessing that if I used a USB disk, it would be even quicker. 40 seconds for infrequent restores is not a burden, in my view. At this point I can open the new backup database and move any individual files or groups that I need to individually restore. 40 seconds.
thx. i understand what you’re saying.
how many copies of your 30 GB database do you keep? going back how far? at what interval?
The Time Machine default is hourly, going back as far as there is room on the disk. Since these are incremental backups, a decent sized disk will give you several years of history.
FWIW, the primary advantage of incremental backups is that they make the backup step more efficient. For almost all situations, backup operations will be very frequent, while restore operations will be very rare. Moreover, you backup because you “want” to, while you restore because you “need” to. Making the backup operation easier makes it more likely to actually happen.
My TimeMachine backup, from which I did this 40 second restore of my biggest DEVONthink database, is from a full system backup of a month ago – everything is backed up and includes DEVONthink databases, those occasional archive zips, and everything else on the MacBook’s disk (few “excludes”). Restating–the backups are not just of DEVONthink, but full system.
Backups to the Synology NAS machine are taken automatically every day and as of now go back three years (when the machine was new). Still more space on the NAS so as time moves forward to have longer retention. Restores of course can be made from these TimeMachine backups.
I occasionally test restoring from these backups, so this 40 second test was not wasted time for me.
TimeMachine is not my only backup. I use the 3-2-1 methodology, which you can read about on the “interweb”. I take special backups of important files using Chronosync, which includes the latest DEVONthink archive zip, and those are kept–so far–forever.
I have some of these special backups stored on the NAS going back more than 10 years. Probably way over the top, but so far have had no pressing need to delete them.
If not already done, read about what DEVONthink Support says about backups in the “DEVONthink Manual”.
How backups are restored should have really nothing to do with how you decide what to import, or index, or not into DEVONthink.
I’ve read the manual. And the thread here. And I believe we are speaking 2 different backup/restore languages. My hope with this thread, to tease out the DT philosophy/thinking re: backup/restore and I feel we’ve done that.
Let’s consider a basic scenario:
1. your DEVONthink DB is 300 GB and contains 500,000 docs and folders
2. you need to restore a 1MB file to the state it was in on Feb 2, 2024.
I’m hearing those of you on this thread prescribe 2 paths:
(Path 1 - Full DEVONthink Database archive/zip)
a. Find a version of your 300GB backup from around that time,
b. let’s say you have one from Jan 2024.
c. Restore the full 300GB.
d. Search within that restored database amongst the 500,000 items for the version of the file you need,
e. Export it.
f. hope the date range for that particular .zip was close enough
This is what I referred to as a monolithic backup/restore. In this case, how much backup space is required to allow you to do that? If you’re restoring a 300GB database to then browse and find a 1MB file, that’s a lot of copies of a 300 GB database you have to keep(!), which requires a lot of space, and that restore operation is cumbersome.
(Path 2 - TimeMachine file incrementals within the DEVONthink package)
a. Open TimeMachine,
b. look inside the 300GB package, and
c. try to find the 1MB file you are seeking to restore, and then
d. scroll back in time to Feb 2024 to find it and restore it.
In this case, the searching-for-the-specific-file is cumbersome, because you do not have the group hierarchy or a tag hierarchy to assist.
--
The compare for me is a simple Finder folder/tag hierarchy, which is fully available and browsable in TimeMachine restores.
I’m not saying every file requires this level of specificity in restore. I’m saying (1) is space-inefficient and cumbersome, and (2) is more cumbersome to retrieve than (3).
All of this is just input to tradeoffs in what a user chooses to keep in DT (related to backup/restore) vs what a user chooses to keep in Finder.
--
In a perfect world it would be great if there was an option (4) = a DEVONthink-native implementation of TimeMachine (or Backblaze, or …) that preserved the DT metadata and performed continuous incremental backups at the file level. This would be space-efficient, performance-efficient, and elegantly usable on both backup and restore.
I gather from this thread that this community does not feel that’s warranted. Fair enough.
Actually, the Time Machine solution would be:
- Scroll back in time to approximately the date you want.
- Restore the full DT database as of that date.
- Search within that database using the full array of DT metadata to assist.
Now we’re talking. That does solve the metadata-browse-to-restore.
Though requires space to restore a 300 GB database in search of a 1 MB file ![]()
Q: do we know if DT support this approach? Specifically, can we rely on hourly incremental backups of DT’s TimeMachine packages to perform file-level incremental backups of the package contents?
I can’t see this called out in the DT4 manual, though it does still encourage full-zip archives. I recall in prior DT versions, TimeMachine was specifically called out as not reliable for this.
If this is the case, then this works. And maybe warrants a paragraph in the manual ![]()
thx for the brainstorming
Note this is beyond the generally recommended limits.
If you routinely deal with 300 GB databases, then you probably want that much free workspace anyway. Or you might consider breaking the database into smaller chunks for performance reasons.
I don’t speak for DT, but the current (DT 4) manual includes no warnings about Time Machine. I would guess that – like any backup solution – it is less reliable if a file is actively being edited while the backup runs. But part of the point of hourly backups is that if the backup at 1 PM misses a file, the one at 2 PM will get it.
I’m still baffled as to what the actual use-case is here. I.e. What is the scenario in which you specifically want to search for the edits of a file you made 2 years ago, and how frequently do you imagine doing this? If you’re doing it that frequently, why aren’t you just saving copies of the files with different dates like people have done for versioning for decades?
Also are you aware as of last year DT has version history? So all the prep you’re doing now will be redundant for edits made after this time (at which point you’d just use version history).
I’m baffled as well.
The meta topic here is: DT database vs. MacOS filesystem.
This thread started with my request for folks to share use cases describing how they decide what items to keep in DT, vs Finder, vs in Finder but indexed in DT. We explored a few of those.
It then evolved into the “just put everything into DT” camp, which as you’d expect in this forum reliably has the most vocal advocates.
Which prompted me to explore what are the tradeoffs of that path?
Pros: better meta-data for sure. Tagging, Grouping, etc
Cons: I attempted to explore this, but the DT-everything advocates don’t see any cons. If they did they wouldn’t be DT-everything advocates.
To me, this includes at least a handful of cons. It’s effectively choosing a file system (DT’s package) within a file system (MacOS and Finder). This prompts me to ask, if I went all-in to DT, what would I be giving up?
What MacOS file system tools do I rely on? Quite a few actually. Terminal. Scripting. Disk Utility. Encryption. TimeMachine. I chose to explore TimeMachine here b/c it’s in my mind the simplest to discuss. I poked at this with specifics (you need to rollback 1 file in a large database) to get folks to consider granular backups, a problem that for a long time has been elegantly solved in MacOS.
I’m a little baffled that the DT-everything answer here came back several times as “just restore the entire database”
That’s the exact equivalent to saying “just restore the entire hard drive”
…so I’m missing something. No one would advocate that restoring an entire hard drive to retrieve a file is a viable backup/restore solution. But in this thread we argue restore-full-database seems totally fine. I’m not sure what I’m missing on that front.
I am aware “Version History” has been added in DT4. I haven’t researched this at all, and on the backup/restore front specifically it obvi warrants a look. Again this would be opting into the file-system-within-the-filesystem, and learning yet another versioning solution on top of MacOS/Finder/TM. Is it worth it?
It’s not yet clear to me (and it’s not anyone’s job here to make it clear to me) why we choose to replicate file system functionality. Using backup/restore as one illustrative example, someone chose to implement version history within DT rather than to leverage MacOS/TimeMachine. Presumably for a benefit.
I appreciate all the discussion and shared thinking in here. I feel this thread has largely exhausted the idea behind my original post - and I thank you.
You keep saying that but not at all true. We are explaining how to restore individual, many, or entire database when you find it needed.
Who’s “we” here? I don’t think anyone is suggesting that you replicate the entire Mac OS filesystem inside DT. The DevonTech folks certainly don’t view their product as a Finder replacement. And even if you did want to store “everything” in DT, there’s no reason to store it all in a single enormous database.
okay. fwiw, this thread was in no way a troll. i am a (happy) multi-year DT user, trying to understand how to expand my usage further. clearly i do not.
No worries and no one thinks you’re trolling
Just people weighing in with reminders or suggestions that may not be known to you (or others).