Skip to content

Latest commit

 

History

2,649 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

datalib — Project Data Liberation ✊

Liberate and own your data. Run powerful AI tools on it, on your terms.

datalib mirrors your personal data — chats, email, messages, contacts, documents, photos, health and fitness — out of the services that hold it and into one place you own: a folder on your own computer, in open formats, with history. Once it is there you can search it, join it across sources, and point whatever tools you like at it, without asking anyone's permission.

Think of it as the data warehouse and pipelines a big company runs over its own data, in laptop-sized packaging. Hat tips to Perkeep and Dogsheep, which got there first.

datalib is an early Imbue project. It is read-only today: it brings data in and never writes anything back to a source.

Get started

Three ways in, from least to most hands-on:

  1. The desktop app (macOS, Apple Silicon). Download the .dmg from the latest release. It asks which folder to keep your data in, then opens the app; you add your first source from the Data sources card.
  2. The command-line tools (macOS or Linux). One curl | sh installs them. The first-time user guide walks through install, credentials, the config file, and your first sync.
  3. Hand it to your agent. Point an AI coding agent at this repository and ask it to set you up. docs/agent_user.md is written for an agent running datalib on your behalf; AGENTS.md is for one working on the code.

Cautious? The Docker image keeps the binaries and your credentials inside a container that sees only the folders you mount, and it comes with a demo library already loaded, so you can look before you hand it anything of yours.

Read this before you point an agent at it

datalib is Plain Old Software. Running a sync invokes no cloud AI model and no agent, and nothing leaves your machine: the only network traffic is datalib reading from the services you configured, plus one-time downloads the first time they are needed: the search models, and (for the command-line install) the Node runtime that search and latchkey run on.

What it produces, though, is a very valuable pile of private data in one place — and most of it was written by other people. Three things follow:

  • Credentials. Web sources authenticate through latchkey, which keeps live session cookies and API tokens on your machine. Any process that can run commands as you can use them to act as you on those services, with no further prompt. Only do this on a machine you trust.
  • Agents. Before you let an agent loose on the mirror, read about the lethal trifecta. Treat everything in your mirror as both private and untrusted content. And remember that an agentic harness sends what it reads to a model provider: ask yourself whether the people who wrote you those messages would be fine with that.
  • Terms of service. The Claude.ai, ChatGPT and Garmin sources talk to the same undocumented APIs those services' own apps use, signed in as you. It is your data, but check the terms of the services you use, and know that those APIs can change without notice.

Supported data sources

Every source links to its section of Getting your data: what it mirrors, and how to get at yours — a login kept by latchkey, an export, or a backup pulled off a phone.


AirVisual

Apple Messages

Apple Photos

Beeper

CalDAV

ChatGPT

Claude

Claude Code

Codex

Contacts

Email

Facebook

Fastmail

Fastmail Calendar

Fastmail Contacts

Garmin

GitHub

GitLab

Gmail

Google Calendar

Google Chat

Google Takeout

Google Voice

Lightroom

LinkedIn

Local files

Media

Notion

PDFs

Perseus

Signal

Slack

SMS Backup & Restore

WhatsApp

YoLink

A source's type says what is being mirrored (claude, whatsapp, …). Its ingest step says how, with one table named for the method: [steps.params.api] reads the product's own API, [steps.params.export] an unpacked export, [steps.params.backup] a phone backup, [steps.params.fswalk] a folder on disk, and so on (jmap, mbox, caldav, …). So a claude source pulled from the API and one read from an export share a type and differ only in that table. Every shape, fully commented, is in all_sources.toml. Something not here? Any program that speaks the step protocol is a source.

How it works

Two layers.

The lower layer is a small, unopinionated pipeline runner. A data store is a file in a folder. A data processor is a program, in whatever language you like. datalib-dag arranges those programs into a graph (a DAG), runs it, and re-runs only the steps whose inputs changed. Any executable that speaks a small NDJSON protocol can be a step — see docs/dev/step_protocol.md.

The upper layer is the batteries. Each source above comes with ready-made steps:

  • ingest brings the raw data in.
  • render_markdown turns the raw records into readable markdown (most sources have one).
  • keyword_index adds that markdown to a keyword search index, built with qmd.
  • embed adds semantic search, which matches on meaning rather than words. It is slow, so you can turn it off per source.

Across all the sources, grid_index builds one SQL table of every message and document (grid_rows), and a single search box reads it together with the qmd index. A local web UI, also shipped as a desktop app (Tauri), manages your sources and syncs, searches, and browses the results.

The stores are doltlite: SQLite's engine over a versioned, content-addressed file format, so a database file is also a git-shaped history of itself. Every sync is a commit. That is what makes the pipeline incremental — each stage asks "what changed since the commit I last read?" rather than rescanning — and it is what keeps the record of what a source changed or deleted between syncs.

Deltas are things too. If you can render a collection of things, consider rendering the difference between two versions of that collection — what was added, what was removed, what changed and how. Code has had this for fifty years; almost nothing else does, and most of the questions people bring to their own data are questions about change. Two mechanisms carry it here:

  • Every store keeps its history. A raw store's commits are the syncs; a render store's commits are the renders. datalib-doltlite reads either at any commit or diffs any two (dolt_log, dolt_diff), and a source's commit history in the app shows what each commit did to each table.
  • A comparison is a source of its own. "Compare two versions…" on a source makes a diff group: the source's own renderer run at both commits and subtracted, written as an ordinary source. Its documents carry the changes marked — added and removed sections on green and red, edited words inside a modified one — and its grid rows say added, removed or modified and which columns moved, so everything that works on a source works on the difference.

What we are aiming for

Near term, we want to be able to ingest and understand many data sources:

  • Big tent — popular and unpopular sources alike, discovering each one's schema rather than forcing it into ours.
  • Local-first — file-based storage; views and processing work offline.
  • Incremental — cheap to keep up to date.
  • Stable identity — a message keeps its id through content edits, so links to it survive.
  • Versioned — notice when the upstream loses or edits data. The history is in the stores, and a diff group renders what changed between two syncs as a source of its own.
  • Legible — render raw data from many schemas into markdown.
  • Findable — search by metadata, keywords, and vectors.
  • Read-only (for now) — ingest-only views of every source.

Longer term: all your data in one place instead of one app per data type; your own apps that join data across sources; as much of your data as possible in your own hands, in formats you can use; and publishing it back into other apps.

Getting your data out again

A mirror you can't leave is just another silo, so the exits are plain:

  • Markdown — <source id>/render_markdown/ is ordinary .md files, one per conversation or document. Nothing to export.

  • SQL — datalib-doltlite ships with the tools and is a sqlite3 shell that understands the versioned format. One pipe writes a plain SQLite file for any tool that wants one:

    datalib-doltlite -readonly unified_index/grid_index/db.doltlite_db .dump | sqlite3 grid.sqlite

    Or browse a store in a GUI: a build of DB Browser for SQLite patched to open doltlite files is at https://github.com/thadd3us/sqlitebrowser/releases (macOS).

Details, and what a snapshot does and doesn't carry, in docs/dev/doltlite.md.

Documentation

License

MIT. A release carries the notices of the third-party software it bundles under licenses/ (in the .app, Contents/Resources/licenses/).

About

Liberate and own your data ✊ Run powerful AI tools on it, on your terms!

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages