On October 6, underlay.org moved to version 2 of the Underlay protocol, a redesign of how collections are stored that substantially improves what Underlay can do. Collections can now grow far larger, to billions of records by design, while the cost of updating one stays proportional to what changed. Large collections are faster to browse and can be streamed at any size, and any copy of a collection, wherever it is hosted, can be verified independently against the original. For a nonprofit maintaining public data infrastructure, the new design also makes that infrastructure considerably less expensive to run.
Existing collections came across intact. Each was converted before the switch with its version numbers, publication dates and files preserved, and record, schema and file hashes carry over unchanged. Version hashes now take a new form with the prefix ulv2:.
Why a new design
The first version of Underlay kept file bytes in object storage and record contents in Postgres, each record stored once under its hash. What tied the two together was a membership table: each version recorded one row for every record it contained. The arrangement was easy to reason about, but its cost grew with the size of the collection rather than the size of the change. A version of a 5.5-million-record collection added roughly 2 GB of indexed membership rows, whether one record or a million had changed since the version before it. Computing a version hash required a full scan, and requests deep into large collections could time out.
In June we described a partial remedy in Content-Addressed Records, in which a version becomes a manifest of record hashes rather than a set of rows. Protocol v2 adopts that approach and extends it, so that the manifest itself is stored as a tree, and moves records and trees alike into object storage.
Records as a tree in object storage
In v2, the records of each type are sorted by id and divided into leaves that average about a thousand records each. Leaves are grouped into interior nodes, and interior nodes into higher levels, until a single root remains. Every node is stored in object storage under the SHA-256 hash of its contents. A version is a short document listing the root of each type’s tree, and the hash of that document is the version hash. The database is reserved for the small amount of state that does change, such as accounts, collection settings and the pointer to each collection’s latest version.
Leaf boundaries are determined by the record ids alone: a leaf closes after any record whose id hashes to a value ending in ten zero bits, and the interior levels apply the same test with a stricter threshold. Structures of this kind, known as prolly trees, were introduced by Noms, an open-source decentralized database, and are used today in the version-controlled database Dolt. AT Protocol, which holds the data of every Bluesky account, relies on a close relative, the Merkle Search Tree, for the same reason: its shape depends only on its contents, not on the history that produced them. They give Underlay three useful properties. The same set of records always yields the same tree, regardless of the order in which the records were written, so two collections that publish identical public data arrive at the same version hash independently. A change to one record rewrites only its leaf and the path from that leaf to the root, and everything else in the new version is a reference to nodes that already exist. And because boundaries do not depend on position, any range of ids can be rebuilt on its own, which allows a large commit to be divided among parallel workers.
A tree of 5.5 million records is four levels deep, and a tree of a billion would be five. We also evaluated a size-aware boundary rule, of the kind Dolt adopted in 2025, which produces more uniform leaves; because its boundaries depend on where the preceding leaf began, it would have ruled out parallel commits, and we kept the fixed rule.
Cost of common operations
| Operation | Cost in v2 |
|---|---|
| Changing one record, in a collection of any size | One leaf, its path to the root and a new root: a few small writes and milliseconds of computation |
| Editing a README or other metadata | A new root that reuses every existing tree |
| Forking a collection | A new root and one database row, with no records copied |
| Comparing two versions | Proportional to the number of differences, since subtrees with equal hashes are skipped unread |
| Opening an arbitrary page of records | The same as opening the first page, since each node records how many entries lie beneath it |
The protocol now fixes the limits on individual records (8 MB per record, 1 KB per id, 64 levels of nesting) and places none on the number of records in a collection. A push transmits only changed records, in resumable sessions. When a commit involves more than two million changes, the server divides it into units of about a million records, processes them as separate jobs and assembles the result, which a property test confirms is byte-identical to a serial commit. Files upload directly to object storage and never pass through Underlay’s servers.
Where Underlay runs
Underlay v2 was designed for edge worker platforms, which execute code in data centers close to the people requesting it. Its core is a protocol library that reaches storage only through a small interface, so it works with any S3-compatible object store and any SQLite-compatible database, and the same code can also run on an ordinary Node server. Because nearly all of a collection is immutable and addressed by hash, reads can be served and cached wherever a deployment runs. Where Underlay is hosted is therefore an operational choice rather than an architectural one, which leaves room to distribute it as widely as its use requires.
Edge environments also impose tight limits, typically on the order of 128 MB of memory per request and no local disk, and designing within them required every read path to stream. A version can be read as newline-delimited JSON of any length, resumable from any record. Its compressed form is especially cheap to serve: each leaf’s records are stored as a separate gzip member, and since concatenated gzip members form a valid gzip stream, the server returns stored objects in sequence without recompressing them. Because each stored object already holds about a thousand compressed records, a full export comes down to roughly one read per thousand records, with no work on the server for any individual record.
Verifiable copies and mirrors
For replication, the protocol defines packs, tar archives holding exactly the objects a receiver lacks relative to a version it already has. The receiver checks each object against its hash, and rebuilds the affected trees to confirm the root, before writing anything. Each collection also keeps an append-only log of its versions, in which every entry is signed and carries the hash of the entry before it, an approach adapted from Certificate Transparency. Together these mean that a copy of a collection obtained from any server, including a mirror run by a third party, can be verified to the same standard as one obtained from the origin. Records a publisher marks private are held apart from public ones and appear in a public version only as a salted commitment, so public versions remain verifiable without exposing them.
Organizations can also attach their own S3-compatible storage to a collection as a mirror. Each mirror is a complete Underlay repository in a documented layout, readable from the bucket alone. For libraries, archives and other institutions concerned with long-term preservation, their data then does not depend on the continued operation of underlay.org.
Prior work
Underlay’s use of hashing is more modest than that of many systems, since each collection has a single authoritative server and there is no consensus mechanism or peer-to-peer discovery. That allowed us to draw on the experience of other projects without their surrounding machinery. Underlay fixes one set of tree parameters per protocol version, after IPFS found that letting tools choose their own chunking gave identical files different identifiers. It hashes records over RFC 8785 canonical JSON and refuses inputs on which implementations in different languages tend to disagree. Its nodes carry their type, level and entry counts, avoiding the ambiguity behind Bitcoin’s CVE-2012-2459, and its leaves are large, in light of the read amplification Ethereum clients met when storing many small trie nodes by hash.
The specification and next steps
Protocol v2 is documented in the protocol specification. It is intended to be precise enough that an independent implementation could reproduce the reference implementation’s hashes exactly. Our next priorities are a load test with many concurrent large pushes, a command-line client that clones a collection, commits locally and pushes only its changes, and support for importing a collection from a mirror.
Since we revived the project in April, Underlay has rested on the view that public knowledge should be held in a structured, versioned form that anyone can access and verify. Protocol v2 gives that view an infrastructure we expect to support it for a long time, at a cost a nonprofit can sustain. If you work with public datasets, institutional collections or research archives and would like to discuss bringing them to Underlay, we would be glad to hear from you at partnerships@knowledgefutures.org.
