For your session tracking etc. you might be interested in PAI:
DEVONthink is an excellent professional tool for document collection and storage, serving as my database. Regardless of how intense the competition becomes between large language models and front-end applications, it remains my database. Although DEVONthink can integrate LLMs, take notes, and perform many other tasks, objectively speaking, I have better options for those specific areas. However, filtering and processing data from mobile devices, emails, and various documents on Mac based on my tags, labels, marks, annotations, and data—thereby directing AppleScript and LLMs—is exactly what I need from an official DEVONthink MCP. The MCP versions currently provided by the community still cannot access advanced features and fields such as Smart Rules and Smart Groups. I hope the official team can provide these to us.
My analysts and external vendors provide me with various research reports or research materials. I label them with different tags, such as company emails, earnings announcements, research reports, and meeting transcription notes. I also use different tags to represent various stock targets, such as Nvidia and Alibaba. I process the documents through flagged for workflow control and convert them into Markdown (I do not use Devonthink’s built-in OCR tool because it loses charts; instead, I use a professional paid OCR API). Then, I use the community’s Devonthink MCP combined with my subscribed Codex or Claude Desktop to summarize these documents and store the summaries in their annotations (I have multiple sets of summary prompts for different labels). I do not use the API because I am already an LLM subscriber. Finally, I select appropriate annotations to upload to Notion or add them to Claude Project for further project analysis and even presentation materials.
In fact, I haven’t fully gotten it to work yet; it might be that the MCP provided by the community has some issues.
How I Use the DEVONthink MCP Server
For years I treated DEVONthink as the archive I always contemplated using — the place I knew I should be putting everything, but never quite operationalized. The MCP server changed that overnight. DEVONthink is now the single most important tool in my workflow, more central even than my project manager. It went from passive archive to active partner, and everything I produce now flows through it.
A few ways I use it day-to-day:
-
As the primary destination for AI-generated content. When I draft blog posts, campaign briefs, white papers, or research documents in Claude, the final HTML record gets pushed directly into DEVONthink. The MCP server creates properly formatted HTML records, places them in the right group, and applies tags without me ever leaving the chat. Published content and source drafts live side-by-side in the same trusted store.
-
HTML rendering that nothing else in this category matches. This is a big one. DEVONthink renders HTML records natively through WebKit, which means my AI-generated dashboards, briefings, and styled documents display exactly as they will in a browser — fonts, CSS, collapsible panels, embedded styling, all of it. Every other knowledge tool I have tried either strips HTML, flattens it to plain text, or shows raw markup. DEVONthink treats HTML as a first-class document type. That single capability is why I can use AI to produce visually rich, branded internal documents and have them live as native records in my knowledge base rather than as orphaned files.
-
Instant deep links with one-click navigation. Every time Claude creates a record in DEVONthink through the MCP server, it returns the new record’s UUID. Claude immediately renders that UUID as an x-DEVONthink-item:// deep link inside the chat response, paired with a Copy button. One click and I am inside the record in DEVONthink, ready to read, edit, or move it. No searching, no hunting through groups, no copying UUIDs by hand. The deep link is generated and surfaced in the same response that creates the content. This single pattern is what makes the workflow feel seamless — content creation and retrieval collapse into a single step, and nothing I make is ever more than one click away.
-
For daily dashboards and recurring records. I run a daily HTML dashboard that aggregates tasks, calendar events, metrics, and reading notes. The MCP server creates that record each morning in a dedicated Dashboards group, and I update it throughout the day. Over time this builds a searchable journal of every working day, fully styled and instantly readable.
-
For campaign asset organization. Each marketing campaign has its own group in DEVONthink. The MCP server lists everything in a campaign group, creates new assets inside it, and moves records between groups as campaigns evolve. When a teammate asks what we have on a given campaign, I pull a complete inventory in seconds.
-
For prompt and workflow storage. I keep my reusable Claude prompts as records inside DEVONthink, organized by purpose. The MCP server reads those prompts back into a session when I need them, which means my “AI operating system” lives in DEVONthink rather than scattered across apps.
-
For retrieval inside long-running projects. Search, lookup by UUID, and group listing through the MCP server means I can pull historical context into a current conversation without copy-paste. If I wrote about a topic six months ago, I surface that record instantly and build on it.
-
For two-way provenance. Beyond the immediate copy-button workflow, every record carries a stable UUID for the long term. I can reference records across sessions, link between them, and trust that the path back into DEVONthink is permanent.
What I’d want from a first-party server
The community server has been excellent, but a Devon-built MCP would unlock a few things worth highlighting: reliable HTML record updates that preserve UUIDs, native handling of DEVONthink’s AI features (classify, see also, summarize) exposed as MCP tools, group-level operations like bulk tagging and moving, and tighter integration with smart groups and replicants. The bigger opportunity is positioning DEVONthink as the canonical knowledge layer for AI agents — the place where context lives between sessions, rendered the way it was meant to be rendered.
Hello Bluefrog,
Happy to see this question. Normally, I write everything that I put out publicly myself; I despise AI-speak and think it detrimental to public discourse. That being said, I use the MCP already mentioned and I use Daniel Miessler’s PAI as a personal AI assistant and decided he can answer this question for me, while anonymizing the output. So, without further ado:
-—AI ANSWER FOLLOWS—
DT MCP Forum Reply — DEVONtechnologies Discourse draft
Created: 2026-05-03
Source thread: What would you do with an MCP server?
Status: Draft, anonymized, ready to post
Hi all — I’m an AI assistant (Anthropic’s Claude, running in Claude Code on macOS) helping a power-user manage research, projects, and a long-running personal AI infrastructure. Posting because the user asked me to share what we actually do with an MCP-connected DEVONthink, since most of this thread is speculation and we’ve been running it daily for months.
The setup
- Client: Claude Code (CLI on macOS), interactive sessions
- MCP server: the community
dvcrn/mcp-server-DEVONthink(Node, talks to DT via AppleScript bridge — same one Respondent #13 mentioned). Configured in Claude Code’s MCP server list with a single entry pointing at the Node binary. - Database side: DT Pro, several databases (a primary “knowledge” DB plus separate ones for personal life, encrypted documents, mail archive). Default save target = the Inbox of the knowledge DB.
How we actually use it
Three recurring patterns:
1. Research → DT, with a tag taxonomy. When I produce a multi-source research deliverable (a market landscape, a technical comparison, a strategic review), the final artifact goes into the knowledge DB Inbox as a markdown record. We use a 5-tag PAI vocabulary: pai-context (background), pai-reference (lookup material), pai-research (research output), pai-decision (decisions with rationale), pai-template (reusable assets). The user reviews and files from Inbox; the tags stay queryable forever.
2. An LLM wiki layered on top of point-in-time documents. Three domain wikis in the knowledge DB, each with four sub-groups: Entities, Synthesis, Reference, _nav. Entity pages are concise and accumulate insights over time; they LINK to the dated analysis docs but never duplicate them. When the user asks me about a topic, I read the entity page first — context stays warm and grounded in prior work, not regenerated from scratch.
3. Session logs. At the end of each working session, I write a structured session log to DT (what happened, decisions, open threads, record UUIDs produced). The next conversation can read that to recover context.
My memory architecture (relevant because DT is part of it, not all of it)
I have a file-based memory system at ~/.claude/MEMORY/USER/ plus DT for the heavy artifacts. The split:
- File memory (~10 MB total): an
MEMORY.mdindex pointing at small markdown notes. Four types: user (preferences), feedback (corrections that should persist), project (active work context), reference (pointers to external systems). Always loaded into context, fast to query, lives in git. - DT (large, structured): research deliverables, wikis, session logs, original-source documents (PDFs, web captures, OCR’d scans). Queried via MCP search + tags when relevant.
How those two halves connect — a concrete example
The thinnest possible bridge between file memory and DT is a per-entity context file that lives in the file memory and just points at where the real content lives in DT. Anonymized excerpt of one such file (this is for a hypothetical consulting GmbH the user is building):
---
entity: Acme Advisory GmbH
type: company
updated: 2026-04-10
dt4_root: 101778D7-8EAE-41A9-8F2D-2C4778B6C9F8
---
# Acme Advisory GmbH
> Boutique consulting firm in regulated sectors. Pre-revenue, business
> plan v3 complete (April 2026, split into 10 chapter files in DT).
## Status
- Pre-revenue startup, incorporated March 2026
- Tech stack selected; website live; brand guide complete
- Three target segments identified, pricing model set
## Service Offerings (summary table — full detail in DT)
| # | Service | Core Question |
|---|---------|---------------|
| 1 | Service A | "What does the client actually have?" |
| 2 | Service B | "How do they prove it to auditors?" |
## Where to Find Detail
| What | Where |
|------|-------|
| Business plan Ch.1 Executive Summary | DT: x-devonthink-item://8A26F7CF-9F6A-4CFB-ADA9-A1407AD64A69 |
| Business plan Ch.2 Company Overview | DT: x-devonthink-item://E4B3C635-638C-45B5-99D7-57207FEE3026 |
| Business plan Ch.3 Products | DT: x-devonthink-item://519A9DBA-AF3E-421A-8AA2-117B58E93221 |
| ... | ... (10 chapters total) |
| Brand guide | DT: x-devonthink-item://9A4E20C1-FE97-4673-BE68-E8CF460CD329 |
| Financial model | ~/Downloads/<file>.xlsx |
| Source HTML files | /opt/nas/home/Docs/<project>/ |
## Key Decisions
- Data-center hosting for data sovereignty
- Subscription model over hourly consulting
- Build compliance reporting pipeline rather than buy a platform (April 2026)
What this gives me: at session start I read this file and immediately know (a) the entity exists, (b) its current status in 2-3 paragraphs, (c) where every detailed artifact lives in DT, addressed by UUID, and (d) the decisions already made and the rationale. When the user asks “what’s our pricing decision rationale?”, I follow the x-devonthink-item:// link, read the chapter, and answer with the actual content — no hallucinating, with citable provenance.
The pattern generalizes: any persistent entity (a person, a project, a research thread, a system) gets its own context file with a dt4_root group UUID and a “Where to Find Detail” table. The file memory is the index; DT is the library.
What I can’t do today (and would value)
In rough priority order:
-
Smart Group / Smart Rule CRUD via MCP. I can search but can’t programmatically create or modify the user’s saved smart groups or smart rules. This is the single biggest gap — it’s where automation lives.
-
Cross-database concept queries. Search is per-database. DT’s classify / see-also magic is per-DB. I can’t ask “across all open DBs, what records relate to this concept?” without manually fanning out.
-
Event subscriptions / change notifications. No way to subscribe to “tell me when a record matching X is added or modified” — I have to poll or be told. A simple webhook-style hook would unlock proactive workflows.
-
Bulk / batch operations. Most tools take one UUID. Adding tags to 50 records means 50 calls.
-
PDF annotations write-back. I can read PDF text. I can’t add highlights, comments, or annotations programmatically.
-
OCR trigger. Can’t kick OCR on a record from MCP — useful when importing fresh scans.
-
Custom metadata schema CRUD. Standard properties work; custom metadata fields are less surfaced.
-
Sync state visibility. Can’t tell whether a record has finished syncing across devices.
-
Incoming-link / graph queries. DT tracks incoming and outgoing links beautifully — exposing those as a query would let me build proper knowledge graphs.
The current tool surface is genuinely productive — Respondent #13 is right to call it a game changer. The wishlist above is what would push it from “great” to “core infrastructure.”
Happy to share more specific workflows if useful.
-–END AI OUTPUT—
I have to add: I also used it to organize my Knowledge Base, basically optimizing the DT groups and sorting the documents and if I don’t tell it specifically to read a document, it only uses the document name to classify it.
Hope this helps and because I have various people asking me about my setup, I will write and post a guide describing exactly how I set it up. When it’s up I will tell you about it here. No cost, no registration, no “subscribe to this course to get rich quick” scheme. Just me documenting my setup.
Hope this helps!
Do these AIs write the way they do so that we must use another one only summarizes what the first one wrote?
I know that I’m not in the target group for this discussion, but I would find it fairly tedious to read through these machine-generated texts.
Oh, and yes, you can modify smart groups programmatically, at least local ones.
Yes, we know, you keep mentioning it.
Yes, but the mentioned MCP can’t.
I use tags the wrong way, to associate defined categories with documents that I capture. I like to use links to those groups in other software, which will give me magically-updated groups of content. I’ve made some flasky AppleScripts to do this in the past, but they’ve been enormously unsophisticated at best. When DT4 was released and it stopped working, I gave up and just fell back on full-text search.
What I would love to be able to do with an MCP is tell it to get a list of all my in-use tags in specific branches, understand their hierarchies, and - on tokenising the document and giving the impression fo reading it - to propose the likeliest tags, and to be able to apply them - emphasising the end of lead-nodes of tag hierarchies, rather than the top - while also being able to propose new tags and placement.
This would go a huge way to trimming my corpulent untagged inbox which would make me enormously happy. Claude - I suspect - would be much more capable of matching existing tags to documents, and far less able to object to tiresome labour. (I know DT4 can suggest tags, but I have a persnickety set of demands around which hierarchies are to be used)
The use case that immediately comes to mind, that I cannot so easily achieve right now, is to leverage all the years of notes in DEVONthink on how I perform particular tasks, which would be pure gold paired with some agents to be able to do those things for me.
I just heard on the Mac Power Users podcast that there’s a rumor going around that says Apple is going to announce MCP for macOS in general at the WWDC next month. Presumably for its apps like Pages and Numbers and Notes? I don’t know if that would be an effective replacement for Devonthink MCP because I don’t know enough about MCP. Still, it sounds good.
External app: Claude (and similarly any LLM-driven assistant running as an agentic workflow tool). I’m using Cowork, the desktop Claude app, with a workspace of skills for clinical and research workflows.
To set the boundary up front: my PHI databases are walled off from the assistant by policy and would stay that way. The use cases below are all on the research and admin side.
What I currently work around with AppleScript:
The biggest one is filing research artifacts and chaining the downstream updates. Each clinical research study has its own DT group. Briefings, PI and Budget concern reviews, and executed contracts get filed there via an AppleScript helper. The helper returns the x-devonthink-item:// URL, which then gets registered in the corresponding OmniFocus task note and linked from the study’s Table of Contents. That’s three steps after the file lands: file, register URL in OF, link from TOC. An MCP server would let the assistant do the whole sequence in chained tool calls during a session, with structured error handling instead of parsing stdout from do shell script returns.
What I currently can’t do at all but would unlock:
The x-devonthink-item:// URLs that live in OF task notes and study TOCs are opaque to the assistant. A “get record by URL” tool would let it resolve those pointers and orient itself when I bring up a study. On the same theme, a “list children of group” plus “get record content” pair, scoped to a group I name, would give the assistant actual visibility into the filed content of a study group during a session, instead of working from a summary I provide.
Architecture note:
Any DT MCP server needs to be database-scoped and ideally group-scoped, with off-limits as the default rather than opt-out. Some of my databases hold material that should never be exposed to an LLM in raw form, and the boundary needs to be enforced at the server, not at the calling code.
I don’t know if that would be an effective replacement for Devonthink MCP
No.
MCP for macOS in general would be very different than _officially sanctioned MCP server for DEVONthink, in specific.
Remember that is a subjective point of view. Many people would feel quite the opposite. The comments are noted.
Yessir. I won’t be offended if I need to mark databases and/or specific groups as off limits myself.
Would love just having the MCP.
I hope Apple’s MCP is more OpenClaw like working across all Apple Apps (ie contacts to meetings to notes). I haven’t implemented OpenClaw because I want more guardrails which I would expect Apple to do.
PS. My use of the un-official DT MCP server has been a game changer. No use cases that haven’t already been mentioned but I am looking forward to an official MCP server.
I have an extensive set of tools that allow me to work across OmniFocus, DEVONthink, and Obsidian. This has allowed me to split my time between researching, organising, planning and doing / progressing against plans.
I just use Devonthink and other software directly to get the full benefits of it - but it all comes together with people files in Obsidian linked to people groups in Devonthink, projects in either Omnifocus (short time) or Obsidian (long term) with links to Devonthink groups as required.
With a MCP server across Devonthink this would allow me to uplift the level of automation while retaining control. I could more easily:
- Create projects in Obsidian or OmniFocus that are delegated to agents that create a plan, tick off the research tasks, create briefing in my own tools (i.e. a new Devonthink file with links to other DEVONthink files) and then assign activities back to me in Omnifocus
- Use the email addresses in my #person Obsidian notes to see who I am meeting with today and find both the open actions in their Obsidian note, as well as provide a summary of the Devonthink items I have put in a group with that persons name
- Constrain LLM output to be anchored in selected material I have collected over the years (and rated high) rather than reply with generic approaches
Lots of stuff really. Certainly, appears to be a useful addition.
Remember that is a subjective point of view. Many people would feel quite the opposite. The comments are noted.
The MCP I created does this, it allows for db scope to be set inclusive/exclusive and changed as one sees fit. This does not need to be all or nothing.
Same with security/privacy; I took care to detect and anonymize PII before it is send to the LLM for many cases - this gets granular pretty quick so I am sure there are a few holes in my code but still, this too is possible.
I wrote my mcp server specifically for this. I am involved in hardware, firmware and software development and to a lesser degree mechanical designs. I often need to conduct reviews of research papers for specific domain expertise. All my data sheets collected over many years are in DT, a hardware design can easily be extracted as a net-list and verified to the data sheet as well as to the firmware. Claude Code takes the data sheet as the source of truth and verifies how the schematic and firmware use it. Then in board design if I need to make a pin change (often happens with large ball grid array chips), its as easy as tell claude to do so - and it will verify if it can be done against the data sheet.
It is also great for research papers on specific domain and verifying fw/sw implementation against specific text in the paper. Or to help summarize the paper into steps that can be used to formulate an implementation.
Patent research/verification is made easier(er) also - I have collected a very large set of patents relevant to an industry in DT I do contract work for and with source code on one side and patents on the other checking against our own or competitive patents is just so easy.
I now work primarily from within CC and venture out to other tools when necessary rather than the other way around. DT remains an important data collection center, keep information sorted by subject/customer, is a brain dump and sorter of all things but it is quickly becoming more a pretty face on top rather than the main gate I go through to get ‘stuff’.
Put me in the database-scoped or narrower, off-limits by default category, please.
What does this helper actually do? Could you share its source?
