Skip to content

How it works

The vault runs as a database, modelled on Databricks and Unity Catalog. A row is any note whose table property names a table, wherever it lives, so the database can be the whole vault or one corner of it. Every write is checked against its table, journaled and can be undone.

  • Metastore — <root>/database.yaml (default root Database) holds config plus catalogs → schemas → tables (columns with comment, nullable, generated, identity, default incl. {now}/{today}; constraints as CHECK filters; properties for partition, file name, archive, excluded folders and history; row_template) and views. Tables are named catalog.schema.table; shorter names resolve through config.default_catalog.
  • Rows — any note whose table property names a table. New rows go to <root>/<catalog>/<schema>/<table>/, named by the table’s filename pattern.
  • Row id — every metastore table starts with the system column id (identity: uuid): a UUIDv7 assigned when the row is created and never changed (renames, moves and alter_table keep it). It cannot be set, dropped or altered (only its label, description, aliases). A row created by hand without an id gets one; a duplicated note gets a new one (recompute_columns).
  • Transaction log (time travel; off by default: configure_database with history: true for all tables, alter_table with history: "on" | "off" | "inherit" per table) — .datanotes/delta_log/<catalog.schema.table>/NNNNNNNNNNNNNNNNNNNN.json, one commit per change (commitInfo, metaData, add with the path, hash and full row content, remove with the path and hash). A checkpoint (full table) is written once the commits since the last one outweigh it (2× the table size, at least 256 KB, or 1000 versions), so the log grows with what is written. Hand edits are recorded as EXTERNAL_EDIT. Requests with the X-Datanotes-Actor header (the client library’s actor option) record the actor: the journal entry gets actor and the commit gets actor and userName (“harness / model”), so the history and the migrations list (by) show who changed what.
  • Graph — notes and their links as a labelled graph: edges from link columns carry the column name (people, event), other resolved links are link, and rows carry their table. graph_neighbors, graph_traverse (breadth-first, by direction, edge labels, result and pass-through tables, depth) and graph_path (shortest path).
  • Performance — row files are cached by stat (mtime + size) with their hash and parsed frontmatter, so operations only read what changed. SIZES=20000 npm run bench (one table, rows ~600 B, in-memory file index) measured at 20 000 rows: insert_row ~0.12–0.16 s, a filtered select_rows ~50 ms (the first one after start ~1 s), check_table ~0.5 s.
  • Operations (the HTTP API, and the same calls through the client library): table and column DDL, constraints, views, clone_table, insert / update / merge / delete / move rows, select_rows (joins, group by, views — including UNION ALL views over several tables, rows tagged _source —, system.information_schema.*, time travel by version or timestamp), history, change data feed, restore, vacuum, migrations list and undo. Every write is journaled and undoable.
  • Private store — journal, transaction log and logs live in a dot-folder at the vault root (.datanotes/, setting storeFolder), which Obsidian does not index (no file explorer, search, graph, Dataview or Bases entries). journal/<YYYY-MM>.jsonl: one line per operation (undo data; undoing appends an undone line); texts ≥ 2 KB (e.g. database.yaml before/after) are stored once in blobs/<hash>.txt. When the engine starts, entries older than journalRetentionDays (90) are dropped, journal notes of the old format (journalFolder/*.md) are imported and removed, and log files older than logRetentionDays (30) are deleted. logs/<YYYY-MM-DD>.<source>.jsonl: the engine’s log records (fileLogEnabled). The store stays with the vault folder the engine runs on: Obsidian Sync skips dot-folders and git sync leaves .datanotes/ to .gitignore, so the journal and the history belong to that machine.

What the operations do beyond reading and writing

Section titled “What the operations do beyond reading and writing”
  • Batches: batch runs several writes that succeed or fail together: in order, each seeing the previous ones; if one is refused, the ones already applied are undone. One journal entry, so undo_migration reverts the whole batch. A dry run checks each operation against the current data.
  • Polymorphic queries: select_rows with table: "base:record" returns the rows of every table that extends the base record (directly or through other bases), on the base’s columns, each tagged _source. A new table that extends the base is included without changing any view.
  • Inverse links: a link column with inverse keeps the other side in step — { name: friends, type: list, items: link, ref: self, inverse: friends } (symmetric), children ↔ parents, partner ↔ partner (single value: setting it replaces the other row’s previous partner), or across tables (project.members ↔ person.projects). Declared on one side is enough. Writes add or remove the reciprocal links in the same operation (one undo); links edited by hand are mirrored a couple of seconds later; check_table reports links that are not reciprocated and align_relations adds the missing ones.
  • Schema inference: infer_table with a folder (or paths) reads the notes’ properties and returns a create_table proposal — types, enums with their values, links with the table they point to (ref), required and unique columns, a file-name rule — plus the adopt_rows call that brings the notes in. With a table, it lists the keys the table’s rows use without declaring them. Read-only.
  • Attribution: the X-Datanotes-Actor header (harness, model, session, tool) records on whose behalf a write is made, in the journal and the history.