Liberate and own your data. Run powerful AI tools on it, on your terms.
datalib mirrors your personal data — chats, email, messages, contacts, documents, photos, health and fitness — out of the services that hold it and into one place you own: a folder on your own computer, in open formats, with history. Once it is there you can search it, join it across sources, and point whatever tools you like at it, without asking anyone's permission.
Think of it as the data warehouse and pipelines a big company runs over its own data, in laptop-sized packaging. Hat tips to Perkeep and Dogsheep, which got there first.
datalib is an early Imbue project. It is read-only today: it brings data in and never writes anything back to a source.
Three ways in, from least to most hands-on:
- The desktop app (macOS, Apple Silicon). Download the
.dmgfrom the latest release. It asks which folder to keep your data in, then opens the app; you add your first source from the Data sources card. - The command-line tools (macOS or Linux). One
curl | shinstalls them. The first-time user guide walks through install, credentials, the config file, and your first sync. - Hand it to your agent. Point an AI coding agent at this
repository and ask it to set you up.
docs/agent_user.mdis written for an agent running datalib on your behalf;AGENTS.mdis for one working on the code.
Cautious? The Docker image keeps the binaries and your credentials inside a container that sees only the folders you mount, and it comes with a demo library already loaded, so you can look before you hand it anything of yours.
datalib is Plain Old Software. Running a sync invokes no cloud AI model and no agent, and nothing leaves your machine: the only network traffic is datalib reading from the services you configured, plus one-time downloads the first time they are needed: the search models, and (for the command-line install) the Node runtime that search and latchkey run on.
What it produces, though, is a very valuable pile of private data in one place — and most of it was written by other people. Three things follow:
- Credentials. Web sources authenticate through latchkey, which keeps live session cookies and API tokens on your machine. Any process that can run commands as you can use them to act as you on those services, with no further prompt. Only do this on a machine you trust.
- Agents. Before you let an agent loose on the mirror, read about the lethal trifecta. Treat everything in your mirror as both private and untrusted content. And remember that an agentic harness sends what it reads to a model provider: ask yourself whether the people who wrote you those messages would be fine with that.
- Terms of service. The Claude.ai, ChatGPT and Garmin sources talk to the same undocumented APIs those services' own apps use, signed in as you. It is your data, but check the terms of the services you use, and know that those APIs can change without notice.
Every source links to its section of Getting your data: what it mirrors, and how to get at yours — a login kept by latchkey, an export, or a backup pulled off a phone.
A source's type says what is being mirrored (claude, whatsapp,
…). Its ingest step says how, with one table named for the method:
[steps.params.api] reads the product's own API, [steps.params.export]
an unpacked export, [steps.params.backup] a phone backup,
[steps.params.fswalk] a folder on disk, and so on (jmap, mbox,
caldav, …). So a claude source pulled
from the API and one read from an export share a type and differ only
in that table. Every shape, fully commented, is in
all_sources.toml.
Something not here? Any program that speaks the
step protocol is a source.
Two layers.
The lower layer is a small, unopinionated pipeline runner. A data
store is a file in a folder. A data processor is a program, in whatever
language you like. datalib-dag arranges those programs into a graph
(a DAG), runs it, and re-runs only the steps whose inputs changed. Any
executable that speaks a small NDJSON protocol can be a step — see
docs/dev/step_protocol.md.
The upper layer is the batteries. Each source above comes with ready-made steps:
ingestbrings the raw data in.render_markdownturns the raw records into readable markdown (most sources have one).keyword_indexadds that markdown to a keyword search index, built with qmd.embedadds semantic search, which matches on meaning rather than words. It is slow, so you can turn it off per source.
Across all the sources, grid_index builds one SQL table of every
message and document (grid_rows), and a single search box reads it
together with the qmd index. A local web UI, also shipped as a desktop
app (Tauri), manages your sources and syncs, searches, and browses the
results.
The stores are doltlite: SQLite's engine over a versioned, content-addressed file format, so a database file is also a git-shaped history of itself. Every sync is a commit. That is what makes the pipeline incremental — each stage asks "what changed since the commit I last read?" rather than rescanning — and it is what keeps the record of what a source changed or deleted between syncs.
Deltas are things too. If you can render a collection of things, consider rendering the difference between two versions of that collection — what was added, what was removed, what changed and how. Code has had this for fifty years; almost nothing else does, and most of the questions people bring to their own data are questions about change. Two mechanisms carry it here:
- Every store keeps its history. A raw store's commits are the
syncs; a render store's commits are the renders.
datalib-doltlitereads either at any commit or diffs any two (dolt_log,dolt_diff), and a source's commit history in the app shows what each commit did to each table. - A comparison is a source of its own. "Compare two versions…" on a
source makes a diff group: the source's own renderer run at both
commits and subtracted, written as an ordinary source. Its documents
carry the changes marked — added and removed sections on green and
red, edited words inside a modified one — and its grid rows say
added,removedormodifiedand which columns moved, so everything that works on a source works on the difference.
Near term, we want to be able to ingest and understand many data sources:
- Big tent — popular and unpopular sources alike, discovering each one's schema rather than forcing it into ours.
- Local-first — file-based storage; views and processing work offline.
- Incremental — cheap to keep up to date.
- Stable identity — a message keeps its id through content edits, so links to it survive.
- Versioned — notice when the upstream loses or edits data. The history is in the stores, and a diff group renders what changed between two syncs as a source of its own.
- Legible — render raw data from many schemas into markdown.
- Findable — search by metadata, keywords, and vectors.
- Read-only (for now) — ingest-only views of every source.
Longer term: all your data in one place instead of one app per data type; your own apps that join data across sources; as much of your data as possible in your own hands, in formats you can use; and publishing it back into other apps.
A mirror you can't leave is just another silo, so the exits are plain:
-
Markdown —
<source id>/render_markdown/is ordinary.mdfiles, one per conversation or document. Nothing to export. -
SQL —
datalib-doltliteships with the tools and is asqlite3shell that understands the versioned format. One pipe writes a plain SQLite file for any tool that wants one:datalib-doltlite -readonly unified_index/grid_index/db.doltlite_db .dump | sqlite3 grid.sqliteOr browse a store in a GUI: a build of DB Browser for SQLite patched to open doltlite files is at https://github.com/thadd3us/sqlitebrowser/releases (macOS).
Details, and what a snapshot does and doesn't carry, in
docs/dev/doltlite.md.
- First-time user guide — install the CLI and mirror your own data.
- Getting your data — per-source credentials and exports.
- Running in Docker — the sandboxed way in: the demo, your own exports, credentials in the container.
- Agent user guide — for AI agents operating datalib on a user's behalf: config, sync, querying, custom steps.
- First-time dev guide — build and hack on datalib from source.
- Contributor runbook — for humans and AI agents working on datalib: the doc map, repo layout, testing rules, and conventions.
MIT. A release carries the notices of the third-party
software it bundles under licenses/ (in the .app,
Contents/Resources/licenses/).