Speaker
Description
Environmental and climate research increasingly depends on large, heterogeneous, and distributed datasets. Managing these data requires not only sufficient storage and computing capacity, but also reliable services for publication, discovery, access, provenance, and long-term reuse. The DETECT project addresses these requirements through a distributed infrastructure operated across participating institutions. This poster focuses specifically on the DETECT data and computing infrastructure hosted at the Jülich Supercomputing Centre (JSC).
The JSC infrastructure combines institutional high-performance storage, cloud resources, Kubernetes-based service deployment, the Earth System Grid Federation (ESGF), and THREDDS Data Server. Together, these components support the publication, discovery, and standardized access of large environmental datasets and their integration into scientific processing workflows. The infrastructure complements the DETECT GeoNetwork catalogue operated on servers at the University of Bonn, which provides an additional project-level mechanism for metadata management and dataset discovery.
Operating the JSC infrastructure has revealed several practical research data management challenges. One major issue is maintaining stable and transparent access when datasets are migrated between storage systems or when services are redeployed. Changes in storage paths, mounted filesystems, service endpoints, certificates, metadata records, or publication states can affect data availability even when the scientific data remain unchanged. Additional challenges include coordinating metadata across distributed services, republishing datasets after infrastructure changes, validating large data collections, preserving provenance, and ensuring that services remain maintainable beyond individual project phases.
To address these challenges, the JSC infrastructure uses containerized applications, Kubernetes-based orchestration, standardized ESGF publication workflows, THREDDS-based data access, automated validation, and structured integration with HPC and cloud resources. These approaches improve portability and operational reproducibility, but several open questions remain. How can metadata remain synchronized between independently operated catalogues and data services? How can storage migration be made transparent to users and downstream workflows? Which provenance information should be captured automatically? How should responsibilities be divided among data producers, metadata catalogues, HPC centres, and long-term repositories?
The poster will present the architecture and current status of the DETECT infrastructure at JSC, lessons from its deployment and maintenance, and open questions concerning distributed environmental research data management. Although the case study originates in climate and atmospheric science, the underlying challenges are relevant to many disciplines operating large datasets across institutional, cloud, and HPC infrastructures.