vibe coded with ❤️
Rune · Data sharing

Urd's Well

Backup as Data Contract

Treat a tool's on-disk output as a documented, versioned contract with a reference reader, so other tools can use it as an offline data source without linking the producer's code.

Problem

Several tools need the same data from Jamf Pro: a backup tool, a simulator, a console. If each one pulls it live, each needs its own API client with privileges, its own rate limiting and its own error handling. History is lost, too: a live API shows only the present. Sharing code between the tools couples their release cycles and dependencies.

Context & forces

  • One tool already reads everything, regularly and read-only (the backup).
  • Consumers differ: sandboxed or not, Swift or Python, different dependencies.
  • The data is large; storing every snapshot in full would be expensive.
  • Backups land in shared or synced folders (OneDrive, SharePoint, network shares), where files arrive late and half-written files are visible.
  • Consumers must never mistake an incomplete backup for a complete one.

Solution

The producer writes a documented folder format and treats it as a public contract: a short normative reference, a longer implementation guide and a small reference reader in another language. Consumers read the files directly.

 instance folder/
   instance.json            which server
   store/                   content-addressed objects, shared by all backups
   backups/<UTC time>/
     manifest.json          written last: its presence means "complete"
     sections/<id>.json     which objects belong to this backup

Full snapshots, deduplicated

Every backup is a full snapshot: its index lists every object that existed at that time, so reading one backup never requires an older one. Deduplication is physical: objects are named by their content hash and stored once in a store shared by all backups. A backup where little changed costs a few index files plus the changed objects.

Completeness by construction

A backup is complete when its folder has a manifest and no .partial suffix. The producer writes objects first, then indexes, the manifest last, then renames the folder. In synced folders it stages the whole run locally and moves it in one step, packing new objects into a single pack file so the sync client uploads a few files instead of thousands.

For developers

  • Version field in the manifest; readers refuse newer versions they don't understand, and ignore unknown keys (new optional keys come without a version bump).
  • Deterministic JSON (sorted keys, UTC dates), sortable backup IDs (UTC timestamps).
  • Section status is explicit (complete, notPermitted, disabled, failed). Consumers report "no data", not "zero objects".
  • A lock file with a heartbeat keeps two Macs from writing the same instance; stale locks are taken over.
  • Optional encryption: objects and indexes are sealed with a data key wrapped for the organisation key (see Fáfnir); the summary stays readable so lists and retention work.
  • Consumers adapt the format to their own model: both Janus and Jarl answer their API requests from the backup, so one parsing code path serves live and offline sources.

Consequences

  • One read-only API client serves many tools, and history comes for free.
  • Consumers stay independent: no shared code, no dependency on the producer's build.
  • The format becomes hard to change; a format rebuild means a coordinated migration (Janus deliberately waited for Gimle's packed format instead of supporting version 1).
  • Gaps in what the producer backs up become gaps in every consumer.
  • The documentation and reference reader are part of the product and need maintenance.

Known uses

  • Gimle produces the backups and documents them in BACKUP-FORMAT.md and IMPLEMENTING-GIMLE-BACKUPS.md, with a Python reference reader tested against plain, packed and encrypted backups.
  • Janus links an organisation to a live server, a Gimle backup folder, or both; it imports the newest backup at start and lets the user switch to older ones like Time Machine.
  • Jarl lets each organisation show a chosen Gimle backup instead of the live server, read-only. Unlike Janus it links the producer's library and answers its normal API requests from the backup, so the whole console works unchanged on a snapshot. Both approaches meet the same contract: Janus stays independent, Jarl reuses more.
  • Lynceus scans Gimle backups for security problems through the producer's reader. A live scan first writes a snapshot in the same format with the producer's engine, then scans that, so live and offline scans share one code path and one history.

Name

Urd's Well, beneath Yggdrasil, where the Norns keep what has been. The backup is the record of what was, and other tools draw from it.