Data Availability (DA) is the property that the data required to verify a blockchain state transition is accessible to network participants. A block can contain a valid commitment or proof, but users still need access to the underlying transaction data in many blockchain designs to independently reconstruct state, verify balances, detect invalid behaviour, or recover the system if an operator stops functioning.
The problem is particularly important for Layer 2 rollups and modular blockchains. These systems can execute transactions outside a Layer 1 blockchain while relying on another network to publish or distribute the data needed for verification. Separating execution from data availability can increase throughput, but it also creates a fundamental question: how can users know that the information behind a state commitment is actually available?
Data availability should not be confused with permanent data storage. A DA system is primarily concerned with making the data needed for verification available during the period in which network participants need it. Long-term archival storage is a related but separate function.
Ethereum has made data availability a central part of its rollup-centric scaling strategy. Rollups use Ethereum not only for settlement but also as a place where transaction-related data can be made available. The introduction of blob-carrying transactions through EIP-4844 in March 2024 significantly changed how Ethereum provides this function to Layer 2 systems.
Why Blockchains Need Available Data
A blockchain does not become verifiable simply because a block producer publishes a block header.
Consider a block containing hundreds or thousands of transactions. The producer can calculate the resulting state and publish a compact cryptographic commitment to that state. Other participants may be able to confirm that the commitment has the correct format, but they cannot necessarily reconstruct what happened without the transaction data behind it.
This creates the data availability problem. A malicious producer could theoretically publish a commitment while withholding some of the information needed to independently evaluate the underlying state transition.
The consequences depend on the blockchain architecture. In some systems, missing data can prevent participants from constructing fraud proofs. In others, it can prevent users from reconstructing balances or continuing operation after a sequencer failure.
A DA mechanism therefore needs to give participants confidence that the relevant block data has actually been published and can be retrieved.
The issue becomes more important as blockchains scale. Requiring every node to download and store increasingly large quantities of data provides strong availability guarantees, but it also increases hardware and bandwidth requirements. If running a node becomes too expensive, decentralisation can suffer.
Modern DA systems attempt to increase the amount of data a blockchain can process without requiring every participant to download every byte.
Data Availability in Rollups
Rollups separate transaction execution from the base blockchain. Transactions are generally processed by Layer 2 infrastructure, while compressed transaction data and commitments are published through a settlement or data availability layer.
This architecture can reduce the amount of computation Ethereum performs directly. However, rollups still need a reliable source of data because users must be able to reconstruct the Layer 2 state independently of the rollup operator.
The requirements differ somewhat between optimistic and zero-knowledge rollups.
Optimistic rollups assume submitted state transitions are valid unless challenged. Data availability is therefore essential for independent parties that need to inspect transactions and identify invalid state transitions during the applicable challenge mechanism.
ZK-rollups use validity proofs to demonstrate that state transitions satisfy the rollup’s rules. This reduces the need to re-execute every transaction simply to determine whether a state transition is valid. Nevertheless, data availability remains important if users need to reconstruct the state, prove ownership, or recover from operator failure.
A simplified rollup data flow can be described as follows:
- Users submit transactions to Layer 2 infrastructure.
- The rollup executes those transactions and updates its state.
- Transaction-related data is compressed or otherwise prepared for publication.
- The rollup publishes the required data through its chosen DA mechanism.
- A state commitment and, depending on the design, a validity proof or other verification information is submitted.
- Independent participants use available data to reconstruct or verify the Layer 2 state.
This is why a rollup’s execution capacity is closely connected to DA capacity. Faster execution alone does not solve scaling if the system has nowhere sufficiently secure and inexpensive to publish the resulting data.
On-Chain DA, DA Layers and Committees
Different blockchain systems make different trade-offs between cost, throughput and security when providing data availability.
| DA Approach | Where Data Is Made Available | Main Advantage | Main Trade-Off |
| Layer 1 data availability | Base blockchain network | Strong security tied to the L1 | Limited capacity and potentially higher cost |
| Dedicated DA layer | Specialised blockchain or network | Higher DA throughput and specialisation | Additional network and security assumptions |
| Data Availability Committee | Selected group of participants | Low cost and high efficiency | Trust depends on a limited committee |
| Validium-style external DA | Outside the settlement chain | Lower L1 data costs | Weaker availability guarantees than fully on-chain rollups |
| Sampling-based DA | Distributed across network participants | Can scale without every node downloading everything | Requires additional protocol and networking mechanisms |
Publishing data directly through the same Layer 1 used for settlement provides strong guarantees because the rollup can rely on the base network’s consensus and availability mechanisms. This approach has historically been central to Ethereum rollups.
Dedicated DA networks take another approach. Projects such as Celestia were designed specifically around data availability rather than general-purpose smart contract execution. Other ecosystems provide their own specialised DA infrastructure.
Some Layer 2 systems use Data Availability Committees, or DACs. A limited group of parties stores and serves the required information. This can significantly reduce costs, but the security model changes because users depend on the committee to keep the data available.
The distinction is important when comparing rollups with validiums. Both can use validity proofs, but a rollup generally publishes the required transaction data to its underlying blockchain, whereas a validium can keep that data elsewhere. Validity of execution and availability of data are separate properties.
Data Availability Sampling
One of the major techniques for scaling DA is Data Availability Sampling, commonly abbreviated as DAS.
The basic objective is to avoid requiring every participant to download an entire large data object. Instead, nodes request small, randomly selected portions of encoded data. If enough independent samples are successfully retrieved, participants can obtain strong statistical confidence that the complete dataset has been made available.
This approach relies on techniques such as erasure coding. Data is transformed with additional redundancy so that the original information can be reconstructed even when only a sufficient subset of the encoded pieces is available.
Sampling changes the relationship between blockchain capacity and node requirements. In a traditional design, increasing block size directly increases the amount of information every full node needs to receive and process. With an effective sampling architecture, the network can distribute verification responsibilities more broadly.
Important elements of sampling-based DA include:
- erasure coding that adds redundancy and enables reconstruction from subsets of data;
- random sampling so a producer cannot easily predict which portions participants will request;
- cryptographic commitments linking samples to the original dataset;
- sufficient independent participants to make withholding difficult;
- peer-to-peer networking capable of distributing data efficiently;
- reconstruction mechanisms for recovering complete datasets when required.
The security is probabilistic rather than based on every participant personally downloading everything. As the number of independent samples increases, the probability that unavailable data remains undetected can become extremely small.
This makes DAS particularly attractive for modular blockchain architectures where data throughput needs to increase much faster than the resource requirements of individual nodes.
Ethereum’s Changing DA Architecture
Ethereum originally provided rollups with data availability largely through transaction calldata. Rollups could publish compressed Layer 2 transaction information inside Ethereum transactions, where it became part of the data processed and stored by the network.
Calldata was not designed specifically for rollup data, however. Rollups needed information to be available for verification without requiring that every byte remain part of Ethereum’s permanent execution-layer state model.
EIP-4844, also known as Proto-Danksharding, addressed this mismatch by introducing blob-carrying transactions. The upgrade was activated as part of Dencun on 13 March 2024.
Blobs created a dedicated data path for rollups with separate fee accounting from ordinary execution gas. Blob contents are not directly accessible to EVM contracts in the same way as calldata. Instead, Ethereum uses cryptographic commitments to connect the data with the transaction while providing the availability properties required by Layer 2 systems.
This represented an important architectural shift. Ethereum began treating rollup data availability as a specialised resource rather than forcing rollup data to compete directly with ordinary EVM execution for the same type of block space.
Ethereum’s subsequent DA development has focused on distributing the burden of handling increasing quantities of blob data. PeerDAS, deployed with the Fusaka upgrade in December 2025, introduced a sampling-based approach in which nodes can participate in verifying data availability without every node downloading all blob data.
The direction is central to Ethereum’s longer-term Danksharding roadmap: increase the amount of data available to rollups while avoiding a proportional increase in bandwidth requirements for every node.
Data Availability Is Not Data Storage
The terms availability and storage are sometimes used interchangeably in discussions of blockchain data, but they solve different problems.
Data availability asks whether the information necessary to verify or reconstruct a blockchain state can be obtained when required. Storage asks where that information remains over longer periods and who is responsible for preserving it.
Ethereum’s blob architecture demonstrates the distinction. Blob data is intended to be available through the protocol for a limited period rather than permanently stored by Ethereum consensus nodes. This is sufficient for the primary verification requirements of rollups because old data does not need to remain inside Ethereum’s active DA mechanism indefinitely.
Historical information can still be preserved by rollup operators, archival services, explorers, indexers, or other storage systems. The fact that consensus nodes do not retain a particular data object forever does not imply that the information must disappear from the broader ecosystem.
This separation helps blockchain systems avoid turning every validator into a permanent archive of all historical Layer 2 activity. As transaction volume grows, requiring permanent replication of every piece of data across every node would create continuously increasing storage requirements.
The Data Availability Trilemma
DA design involves a practical tension between throughput, decentralisation and security.
A network can increase throughput by allowing very large blocks, but if every node must download those blocks in full, bandwidth requirements increase. Eventually, only operators with expensive infrastructure may be able to participate effectively.
A system can reduce costs by storing data with a small external committee, but this creates stronger trust assumptions. If enough committee members fail or collude, users may lose access to information even though the corresponding state commitments remain visible on-chain.
Sampling-based designs attempt to improve this trade-off by spreading data across a larger network and allowing individual nodes to verify availability using only portions of the total dataset.
The appropriate model depends on the application. A high-value financial rollup may prioritise Ethereum-backed DA despite higher costs. A gaming or social application processing very large transaction volumes may accept different security assumptions in exchange for substantially cheaper data publication.
For this reason, transaction throughput alone is not enough to compare Layer 2 systems. Two networks advertising similar execution performance can have very different security properties if one publishes data through Ethereum while the other depends on a smaller external committee.
Why DA Has Become Core Blockchain Infrastructure
As blockchain architecture becomes modular, data availability is increasingly treated as an independent infrastructure layer alongside execution, consensus and settlement.
A monolithic blockchain performs most of these functions within one protocol. Modular systems can separate them. A rollup might execute transactions in one environment, settle commitments on Ethereum, and use Ethereum or a specialised network for data availability.
This separation allows each component to be optimised for a narrower purpose. Execution systems can concentrate on processing transactions, while DA infrastructure can focus on distributing large quantities of verifiable data.
The security implications are substantial. If transaction data disappears, having a state root or proof on another blockchain may not be enough for users to reconstruct the system. Data availability therefore determines whether decentralised verification remains practical when execution moves away from the base layer.
This is why DA has evolved from a relatively specialised blockchain research problem into one of the central components of Layer 2 and modular blockchain design. Rollup scaling ultimately requires more than faster execution. It requires enough verifiable data capacity to support that execution without making network participation prohibitively expensive.
The development of blobs, specialised DA layers, erasure coding and Data Availability Sampling reflects the same objective: make substantially more blockchain data available while preserving the ability of ordinary network participants to verify that the data exists.