Data Availability Sampling (DAS) is a blockchain technique that allows network participants to verify with high probability that block data is available without downloading the entire block or dataset. Instead of requiring every node to receive all published data, participants request small, randomly selected samples. If enough samples can be retrieved successfully, nodes gain strong statistical confidence that the complete dataset was made available.
DAS addresses a fundamental blockchain scaling problem. Increasing the amount of data that a network can process usually increases bandwidth and storage requirements for nodes. If every node must download every byte, larger blocks can eventually make independent verification impractical for ordinary participants.
Sampling changes this relationship. A blockchain can distribute larger quantities of data across its peer-to-peer network while individual nodes handle only a fraction of the total. The objective is not to prove that one particular node possesses the entire block. It is to make systematic data withholding detectable and ensure that the complete dataset can be reconstructed by the network when required.
The technique has become particularly important for modular blockchains and rollup-centric scaling. Networks such as Ethereum need to provide increasingly large amounts of data availability capacity to Layer 2 systems without forcing every Ethereum node to download all Layer 2-related data.
The Problem DAS Is Designed to Solve
Blockchains depend on more than consensus over block headers or state commitments. Participants also need the underlying data required to verify and reconstruct the state represented by those commitments.
Imagine that a block producer creates a block, calculates a valid-looking commitment, but distributes only part of the underlying data. A node that sees the commitment cannot necessarily determine from the commitment alone whether every transaction or data segment is actually obtainable.
The straightforward solution is for every full node to download the complete block. This provides a strong availability check because each verifying node directly receives the information it needs.
The problem appears when block size increases. A network processing ten times more data can require approximately ten times as much data transmission under a simple full-download model. Continually increasing these requirements can raise the cost of operating nodes and reduce the number of independent participants capable of verifying the network.
DAS asks whether every participant really needs to download everything.
If data is encoded with redundancy and many nodes independently request random pieces, a producer attempting to hide a meaningful portion of the dataset faces a high probability that one or more nodes will request missing pieces. Repeated independent sampling can therefore transform a small amount of downloaded data into a strong probabilistic availability guarantee.
This makes DAS fundamentally different from simply asking a trusted server whether a block is available. The security comes from cryptographic commitments, redundant encoding, random sampling, and many independent participants.
How Sampling Detects Withheld Data
The central insight behind DAS is statistical. A node does not need to inspect every piece of a dataset if unavailable portions cannot be predicted or hidden from all random checks.
Suppose, purely as a simplified example, that half of an encoded dataset is unavailable. A node requesting one uniformly random sample would have a 50% chance of encountering unavailable data. If it independently requested ten samples, the probability of selecting only available pieces would be:
0.5¹⁰ = 0.0009765625
That is less than 0.1%.
After 20 independent samples, the probability would fall to approximately 0.000095%. Real DAS protocols have more complicated assumptions, encoding structures, networking behaviour, and sampling rules, but the example illustrates why random sampling can provide strong confidence without complete downloading.
The process generally relies on several components:
- The original block data is divided into smaller pieces.
- Erasure coding expands the dataset with redundant information.
- A cryptographic commitment binds the producer to the encoded data.
- Nodes independently request randomly selected samples.
- Samples are checked against the relevant commitment.
- Successful retrieval of enough independent samples provides statistical confidence that sufficient data is available.
- The redundant encoding allows the complete original dataset to be reconstructed when enough encoded pieces are obtained.
The number of samples required depends on the protocol’s assumptions and desired security level. DAS therefore provides probabilistic rather than absolute verification by every individual node.
Probabilistic does not mean arbitrary. Security parameters can be selected so that successful withholding becomes extremely unlikely to escape detection across a sufficiently decentralised population of samplers.
Why Erasure Coding Matters
Random sampling alone is not enough. A malicious block producer could potentially withhold only a small but critical portion of the original data, making it difficult for samplers to detect the missing information.
Erasure coding addresses this problem by transforming the original dataset into a larger encoded dataset containing redundancy. The original information can be reconstructed from a sufficient subset of the encoded pieces.
A simplified analogy is a file converted into many fragments in such a way that the complete file can still be recovered without possessing every fragment. The important difference from ordinary replication is that erasure coding mathematically generates redundant pieces rather than simply making identical copies.
This creates an important threshold property. If enough encoded data is available, the original dataset can be reconstructed. To make reconstruction impossible, an attacker must withhold a substantial portion of the encoded information.
That larger missing portion is much easier for random samplers to detect.
Cryptographic commitments are then used to prevent the producer from responding to different nodes with inconsistent pieces. A sample needs to be verifiably connected to the data commitment associated with the block.
The combination of erasure coding and commitments therefore converts availability into something that can be tested efficiently across a distributed network.
DAS Compared With Full Data Download
DAS changes the amount of work an individual verifier needs to perform, but it also changes the security model from direct observation to statistical confidence.
| Characteristic | Full Data Download | Data Availability Sampling |
| Data downloaded by each verifier | Entire dataset | Small subset of the dataset |
| Availability confidence | Direct for downloaded data | Probabilistic |
| Bandwidth per node as data grows | Grows with total data | Can grow much more slowly |
| Need for redundant encoding | Not inherently required | Typically important |
| Random sampling | Not required | Core mechanism |
| Scalability | Limited by per-node resources | Designed for larger data volumes |
| Network participation | More demanding at high throughput | Can reduce individual bandwidth burden |
| Reconstruction | Node already has complete data | Network reconstructs from sufficient encoded pieces |
The full-download model remains conceptually simpler. Every full node sees everything and can independently evaluate availability.
DAS deliberately removes that requirement. Its value appears when the network wants to increase total data capacity substantially while keeping individual verification requirements manageable.
This does not mean that nobody stores or reconstructs complete data. Some network participants may handle much larger portions of the dataset. DAS is primarily about removing the requirement that every verifying participant must do so.
DAS in Modular Blockchain Design
Data Availability Sampling became especially relevant as blockchain architecture moved away from the assumption that one chain must perform execution, consensus, settlement, and data availability in the same way for every transaction.
Rollups illustrate this change. A rollup can execute thousands of transactions outside its settlement layer, but information required to reconstruct its state still needs to be made available somewhere.
As rollup throughput increases, the amount of data produced by these systems can become much larger than the amount of computation the settlement blockchain performs directly.
Simply increasing Layer 1 block sizes would push the resulting bandwidth requirements onto every full node. DAS provides another scaling path. The network can increase total data capacity while dividing data-handling responsibilities among participants.
This makes data availability an independently scalable blockchain resource.
Celestia was among the blockchain projects designed around this modular concept, using Data Availability Sampling as an important part of its architecture. Ethereum has pursued related ideas as part of its rollup-centric roadmap, where scaling data availability is necessary to support substantially higher Layer 2 throughput.
The implementations are not identical. DAS is a general technique, while each network determines its own encoding, sampling, networking, and security architecture.
Ethereum, PeerDAS and Blob Data
Ethereum’s use of sampling needs to be understood in the context of its transition towards specialised data availability infrastructure.
EIP-4844 introduced blob-carrying transactions with the Dencun upgrade on 13 March 2024. Blobs gave rollups a dedicated mechanism for publishing data without using ordinary calldata for the same purpose. This reduced the direct connection between Layer 2 data demand and Ethereum execution gas.
However, increasing blob capacity creates another problem. If every Ethereum node must download every blob as the number of blobs grows, node bandwidth eventually becomes a scaling bottleneck.
PeerDAS was designed to change that model. Instead of every node downloading all blob data, the network distributes responsibility for data across peers and uses sampling to verify availability.
PeerDAS was deployed as part of Ethereum’s Fusaka upgrade in December 2025. This represented an important step beyond the original EIP-4844 architecture. Blob capacity could increasingly be scaled without requiring every node to handle the entire volume of blob data.
In simplified terms, the evolution can be viewed as a sequence from calldata-based rollup data, to dedicated blobs, to distributed sampling of blob data. Each stage changes how Ethereum handles the growing data requirements created by Layer 2 systems.
This is also why DAS is closely associated with Ethereum’s longer-term Danksharding direction. The objective is not merely to create larger data objects. It is to distribute the work of verifying their availability so total data capacity can grow faster than individual node requirements.
Security Assumptions Behind DAS
DAS provides strong scalability benefits, but its guarantees depend on several assumptions working together.
Sampling must be sufficiently random. If a malicious producer can predict exactly which pieces will be requested, it could potentially reveal those pieces while withholding others.
There must also be enough independent sampling activity. One node requesting a handful of samples provides weaker collective protection than thousands of independently sampling participants.
Networking is another important factor. Data can theoretically exist while still being difficult for particular nodes to retrieve because of network partitions, censorship, or connectivity problems. DAS therefore depends not only on cryptography but also on effective peer-to-peer data distribution.
Important conditions for a robust DAS system include:
- unpredictable or sufficiently independent sample selection;
- enough sampling participants to make withholding difficult;
- correct erasure coding of the original dataset;
- cryptographic commitments that prevent inconsistent responses;
- sufficient peer-to-peer bandwidth and connectivity;
- mechanisms for reconstructing data from available encoded pieces;
- security parameters that make undetected withholding statistically negligible.
Implementation complexity is consequently much higher than the basic “download a random piece” description suggests. Encoding must be efficient, proofs or commitments must be practical to verify, and networking protocols need to distribute samples without creating new bottlenecks.
The behaviour of light clients is also important. DAS can allow resource-constrained participants to contribute meaningfully to availability verification, but only if they can obtain samples independently rather than relying on a small number of centralised gateways.
What DAS Changes for Blockchain Scaling
The most important contribution of Data Availability Sampling is that it weakens the traditional relationship between blockchain data capacity and the bandwidth required from every individual verifier.
Without sampling, increasing block data generally means asking each full node to process a proportionally larger stream of information. That approach eventually collides with practical hardware and network limits.
With DAS, total network capacity can increase by distributing data responsibility across a larger population of nodes. Individual participants verify small pieces, while the network collectively provides confidence that enough information exists to reconstruct the whole dataset.
This does not make data availability free. Larger datasets still need to be transmitted, encoded, stored for the required period, and served across the network. DAS changes how that workload is distributed rather than eliminating it.
Its significance is therefore architectural. Blockchain scaling no longer has to depend entirely on making every node powerful enough to process larger and larger blocks.
For rollup-centric and modular systems, this creates a path towards much higher data throughput while preserving relatively accessible node requirements. Execution can scale on Layer 2, data can be distributed across a DA network, and participants can verify availability without personally downloading every transaction generated by the broader ecosystem.
Data Availability Sampling is ultimately a method for turning many small independent checks into a network-wide security guarantee. That property makes it one of the key technologies behind modern attempts to scale blockchain data without sacrificing decentralised verification.