Skip to main content

Decision

Every blob is identified by its BLAKE3 hash. Clients and nodes verify received bytes against the known hash on every transfer. The hash→backend mapping is internal to each origin-backed node and never shared — no participant in the network can learn or bypass the node’s backing storage.

Why BLAKE3

  • Fast. Several GB/s on modern hardware — faster than SHA-256, often limited by memory bandwidth rather than CPU.
  • Tree-structured. The BLAKE3 hash tree enables streaming verification: a client can verify chunks as they arrive rather than waiting for the full blob.

Verified streaming

Delivery uses bao — an interleaved verified-stream encoding of the BLAKE3 tree. The paid-delivery payload is always bao bytes: content interleaved with the hash proofs needed to verify each 16 KiB chunk group against the root the client expects. The bao decoder verifies every chunk group as it arrives. A malicious node cannot insert a single corrupt byte and still earn — the client rejects the stream before paying for that chunk group. Because every chunk group carries its own proof, a resumed or ranged fetch verifies independently — the decoder aligns the requested offset down to its 16 KiB chunk-group boundary, so there’s no need to re-read from the start. Payment meters all transmitted bytes, including the proof bytes: the verification overhead is roughly 0.4% of blob size, not free metadata. Where an origin publishes the blob’s outboard (its BLAKE3 tree) as a sidecar, an origin-cold node streams the origin’s bytes straight through to the paying client while teeing them into its own store, verifying each group against the root as it goes. Only when the origin publishes no outboard must the node import the blob and build one first, at a one-time import latency. Node-to-node cache-miss pulls stay bao-encoded and pipelined end to end.

Publishers and namespaces

A hash answers “what is this blob?”. It does not answer “who is responsible for serving it?” — that is a namespace, a publisher-owned identifier for a content set and the unit of origin addressing. A delivery request pairs the hash with the namespace its content is published under, and the node routes on it: The namespace is a routing hint, not a trust anchor: bytes are verified against the hash independently, so a wrong or hostile namespace can only cause a failed fetch, never corrupt or mis-attributed delivery. Cache-only serving stays permissionless — any staked operator may re-serve bytes it holds, for any hash, under any namespace. Only the origin role is namespace-gated (takedown).

Hash-to-object-key mapping

Origin backends speak their own namespace: S3 objects are addressed by key, NFS by path, local disk by filename. Blobs are addressed by BLAKE3 hash. The mapping between them is a private catalog held by the operator and queried on cache miss to locate the bytes. It is not on-chain and not part of the protocol — the protocol only requires that a node deliver the correct bytes for a given hash.

Why the origin is hidden

Leaking the origin URL would bypass the pay-per-byte economic model — anyone could download directly from S3 and skip the CDN entirely. A stream response carries no route hint of any kind, so the backend topology of origin-backed nodes is fully hidden from the network, and the origin backend is not addressable from outside the node process. This also protects origin operators from direct egress cost attacks: an adversary cannot generate S3 traffic by pointing clients at the bucket URL.

What clients verify vs. trust