Data as Infrastructure: Can Science Function Without It Today?

August 12, 2026

Just over a decade ago, the main challenge for many scientific disciplines was obtaining enough data to answer a research question. Today, the exact opposite is true. Science generates such vast amounts of information that the challenge is no longer collecting it, but storing, sharing, processing, and transforming it into useful knowledge.

This transformation has led to a quiet yet profound paradigm shift: data is no longer simply an output of research but has become an essential infrastructure underpinning modern science. Just as a laboratory needs electricity, instrumentation, or Internet access to operate, research today depends on ecosystems capable of managing volumes of information that would have been unimaginable only a few years ago.

The scale of this shift is clear. The Large Hadron Collider (LHC) at CERN generates around one trillion (10¹²) particle collisions per second. Although only a small fraction of these events is retained for analysis, the experiments produce tens of petabytes of scientific data every year, which must be processed through a distributed computing infrastructure spanning more than 170 computing centers across over 40 countries. Known as the Worldwide LHC Computing Grid, this network is one of the largest scientific computing systems in the world.

This example illustrates a reality shared by virtually every scientific discipline. Astronomy continuously observes the universe through ground- and space-based telescopes; genomics sequences millions of DNA fragments every day; Earth observation satellites constantly collect information on oceans, forests, and the atmosphere; and laboratories generate ever-growing volumes of images, simulations, and experimental results. In every case, data is what makes the research itself possible.

When Data Stops Being a Resource and Becomes Infrastructure

When we think of scientific infrastructure, particle accelerators, telescopes, or supercomputers usually come to mind. However, an increasing number of institutions now consider data itself to be an infrastructure of equal strategic importance.

But what does it really mean to think of data as infrastructure?

It is not simply a matter of having large storage capacities. Data infrastructure encompasses all the elements required to make information usable by the scientific community: acquisition systems, secure storage, high-speed communication networks, computing platforms, common standards, analytical tools, and mechanisms that enable results to be shared efficiently and reproducibly.

In other words, data only becomes valuable when it can be found, understood, combined, and reused.

This is the principle behind the FAIR principles, promoted by the international scientific community, which establish that data should be Findable, Accessible, Interoperable, and Reusable. Far from being merely a best practice, these principles have become a reference framework for universities, research centers, and European programs seeking to maximize the impact of publicly funded research.

This shift in approach also responds to the need for greater scientific efficiency. When data is properly documented and reusable, other research groups can validate results, develop new lines of research, or combine information from entirely different disciplines. Research no longer advances in isolation but builds upon shared knowledge.

It is no coincidence that the European Union is promoting initiatives such as the European Open Science Cloud (EOSC), a digital infrastructure designed to enable researchers across different countries to securely access, share, and reuse scientific data in a standardized way.

Science That Moves Petabytes

The volume of information generated by modern science is difficult to imagine. A few decades ago, a researcher could store years of work on a handful of hard drives. Today, some scientific projects generate more data in a single day than many institutions once produced over an entire decade.

One of the most representative examples is the Square Kilometre Array Observatory (SKAO), the largest radio telescope ever built. Once fully operational, it will produce several hundred petabytes of scientific data every year, making it one of the largest generators of information on the planet.

Earth observation provides another particularly relevant example. The European Copernicus program, coordinated by the European Commission and the European Space Agency (ESA), provides continuous, open data on the state of our planet through its constellation of Sentinel satellites. This information is essential for monitoring wildfires, droughts, air quality, sea ice evolution, agriculture, and responses to natural disasters.

What is particularly significant is that none of these projects would be possible through better telescopes or satellites alone. Their success also depends on infrastructures capable of transporting, storing, processing, and distributing massive amounts of information almost in real time.

In this context, the ability to manage data has become a scientific advantage as important as having better observational instruments. Those capable of extracting knowledge faster and more accurately are also better positioned to address the major scientific and technological challenges of our time.

Artificial Intelligence Needs Data, but Above All, Quality Data

The rise of artificial intelligence has placed data at the heart of scientific and technological debate. However, there is a widespread misconception that the success of an AI model depends solely on the sophistication of its algorithm. In reality, its performance is largely determined by the quality, diversity, and reliability of the data on which it has been trained.

In scientific research, this is particularly critical. A model may analyze millions of medical images, physical simulations, or experimental records within minutes, but its results will only be as robust as the information it receives. Incomplete, poorly labeled, or biased data can lead to incorrect conclusions, even when the most advanced tools are used.

This is why the concept of data stewardship is becoming increasingly important. It refers to the responsible management of data throughout its entire life cycle. From acquisition and validation to storage, documentation, and reuse, every stage directly influences the quality of research and the ability to reproduce its results.

This need is also driving the development of new methodologies for integrating data from very different sources. A single project may combine satellite imagery, IoT sensors, physical models, computational simulations, and experimental laboratory data. In many cases, the ability to connect all this information coherently is what enables new discoveries.

The Next Scientific Infrastructure Will Not Be Only Physical

For decades, investment in science was primarily associated with the construction of major infrastructures: particle accelerators, astronomical observatories, laboratories, and supercomputing centers. All of these remain essential, but it is becoming increasingly clear that their true potential depends on another, less visible but equally critical infrastructure: the systems that manage and process the data they generate.

This evolution explains the growth of High Performance Computing (HPC) centers, cloud computing platforms for research, and international scientific data-sharing networks. European supercomputers such as MareNostrum 5, in Barcelona, make it possible to run simulations and analyses that only a few years ago would have required months of computation. Today, these infrastructures are strategic assets for disciplines as diverse as new materials design, weather forecasting, biomedical research, and space exploration.

But the horizon extends even further. Quantum computing promises to address problems whose complexity exceeds the capabilities of classical systems, particularly in areas such as optimization, molecular simulation, and drug development. Although this technology is still under development, its potential will likewise depend on our ability to integrate, process, and manage enormous volumes of scientific information.

In other words, the infrastructure of the future will consist not only of buildings or instruments, but also of digital ecosystems capable of connecting data, algorithms, and knowledge on a global scale.

An Invisible Infrastructure for Tackling the Challenges of the Future

The work carried out by centers such as ARQUIMEA Research Center, within the framework of its QCIRCLE project, reflects precisely this transformation. Research in areas such as artificial intelligence, photonics, quantum computing, biotechnology, and space systems generates and uses large volumes of information that must be integrated, processed, and analyzed to transform it into useful knowledge.

In this context, simply having access to data is no longer enough. The real challenge lies in developing the technological capabilities required to manage it securely, interoperably, and efficiently, fostering collaboration across disciplines and accelerating the transfer of knowledge into new scientific and technological applications.

Just a few decades ago, having access to large amounts of information was a competitive advantage. Today, it has become an essential requirement for doing science. Yet the real value lies in building the infrastructures that make it possible to transform data into knowledge, share it efficiently, and reuse it to answer increasingly complex questions.

In the coming years, some of the greatest scientific advances will likely depend on our ability to connect millions of data points from different disciplines and transform that information into new ideas, new technologies, and new ways of understanding the world.

Because just as electricity enabled the Industrial Revolution and the Internet transformed global communication, data has become the silent infrastructure on which 21st-century science is built.