What Is Content-Addressed Storage?

Content-addressed storage names a file by its bytes. Instead of “server.com/files/photo.jpg,” you get a SHA-256 hash encoded as a multihash. Change one byte and the identifier changes completely. That gives cryptographic proof the data wasn’t modified. No third party is required.
Location-addressed storage gives you none of that natively.
How Is Content-Addressed Storage Different From Regular Storage?
In regular storage, you address data by location: a URL, a bucket key, a filesystem path. You trust the server to return the right file, and if someone changes it on the server you have no automatic way to know.
The identifier, called a CID (Content Identifier), is a hash of the data itself. You don’t trust the server; you verify the hash. A matching CID means the file is exactly what was requested; a mismatch means it has been altered or corrupted.
With regular web storage you might have ETags that track whether something changed, but proving a whole website hasn’t been changed is hard. HTTP has no immutable DAG or Merkle foundation.
| Feature | Regular Storage (Location-Addressed) | Content-Addressed Storage |
|---|---|---|
| Address type | Path/URL (chosen by user) | CID (derived from content hash) |
| Modification detection | ETags, timestamps | Automatic (hash changes = new CID) |
| Trust model | Trust the server | Verify the hash |
| Snapshot capability | No native support | Root CID = verifiable snapshot |
| P2P distribution | Server-dependent | Any peer with the blocks can serve |
What Is a CID and How Does It Work?
A CID (Content Identifier) is IPFS’s native identifier. It usually encodes a SHA-256 content hash in a multihash format that records which hash function was used, how long the digest is, and the digest itself. It’s the address you use to retrieve content, but unlike a URL, the address itself is the verification.
The data behind a CID is organized as a DAG (Directed Acyclic Graph): blocks of data connected by their relationships. In IPFS, the underlying data spec is IPLD, which defaults to roughly 256 KB blocks, each holding references to its child blocks. Files and folders both resolve to child blocks, and each block CID hashes its contents plus those child links. Change any block and its parent hashes change, all the way to the root CID.
How Does Content Addressing Enable Verifiable Website Snapshots?
A root CID is a cryptographically verifiable snapshot of an entire website. The DAG structure means one root identifier represents the full site state; recording root hashes gives you a change history.
Traditional web storage doesn’t expose this. You’d have to build your own tracking layer on top of HTTP, and even then you can’t prove the server didn’t serve different content at different times. Immutability has to start at the storage layer.
How Does Pinner Implement Content-Addressed Storage?
Pinner uses Sia for persistence and an IPFS/IPLD layer for content addressing and distribution.
Sia handles persistence. “Pinned” means the network keeps data available persistently rather than just caching it temporarily. It stores IPLD blocks as raw blobs through its object interface, so Pinner can persist any data Sia accepts. Sia erasure-codes data into 30 shards (10 data, 20 parity); any 10 recover the file. At the low level it uses Merkle trees to track sectors.
The IPFS/IPLD layer generates CIDs and builds the DAG. Pinner’s portal and gateway use the boxo framework. Metadata lives in a SQL database; raw blocks live as Sia objects. When boxo requests a block, the blockstore fetches it from Sia on demand, with an in-memory cache for hot content.
| Layer | Responsibility | Technology |
|---|---|---|
| Content addressing | CID generation, DAG structure, block linking | IPFS / IPLD |
| Distribution | Serving blocks, caching, P2P retrieval | boxo framework, edge gateway |
| Storage | Persistence, erasure coding | Sia (30 shards, Reed-Solomon) |
| Database | Pin records, upload metadata | Internal DB |
What Happens When a Storage Node Goes Offline?
When a host goes offline, the renter software detects the missing shards and autonomously uploads new ones to healthy hosts. That means up to 20 hosts can fail and the data is still recoverable.
Data is encrypted client-side with ChaCha20 before it leaves your machine, then split into shards. No single host can read what they’re storing. They only hold encrypted fragments and can’t piece them together.
If Pinner’s infrastructure goes offline, data can’t be distributed unless it’s cached. It isn’t gone, though. The Sia network still holds it, and when connectivity is restored the blocks become available again through boxo.
The CID always resolves to the same content. Reachability depends on the storage layer, not the address.
FAQ
Is content-addressed storage the same as IPFS?
No. IPFS is one implementation. Content addressing is the underlying idea.
Can I access content-addressed data with just the CID?
Yes, as long as the blocks are pinned and available on the network. The CID is both the address and the verification. You can retrieve the data from any peer that has those blocks, not just the original server.
What happens if the data changes?
Any modification produces a new root CID while the old one still resolves to the original content, so nothing is overwritten.
Does Sia use CIDs?
No. Sia uses Merkle trees to track sectors, but it doesn’t use IPFS CIDs natively. Pinner stores IPLD blocks as raw blobs in Sia and handles the CID layer above it.
Is content-addressed storage slower than regular storage?
Retrieval speed depends. The hash computation adds overhead, but content-addressed data can be fetched from any peer that has it, not just one origin server. Pinner’s cache layer on top of Sia keeps common content fast.