Speakers
Description
Modern biomedical research increasingly relies on large and heterogeneous datasets generated across different experimental platforms, research groups, and core facilities. Within our Research Training Group, these range from sequencing and spatial transcriptomics to proteomics, imaging, electrophysiology, and data derived from mouse models and human tissue. Managing such diverse and potentially large-scale datasets throughout their entire life cycle requires more than sufficient storage capacity: data need to remain structured, traceable, accessible, and reusable while accommodating different data types, computational requirements, and regulatory constraints.
Here, we present a research data management concept specifically designed for the interdisciplinary Research Training Group (RTG) Development and Epileptogenesis of Dysplasias in the Interplay of Distinct CNS Cell Types. The concept integrates research data management from the planning stage and follows the FAIR principles of making data Findable, Accessible, Interoperable, and Reusable. Key elements include a common data management plan, identification and harmonization of data sources, appropriate consent structures for human data, and integration of existing institutional infrastructures.
The proposed architecture separates but connects three major components: data storage, metadata management, and computational analysis. Data generated across individual RTG projects, Open Toolboxes (OTBs), and core facilities are curated with support from data stewardship and integrated into centralized institutional infrastructures. Large research files are maintained in dedicated storage systems, while associated metadata are managed separately to facilitate data discovery, interpretation, and traceability. Dedicated analysis platforms and high-performance computing resources provide scalable computational access according to data type and analytical requirements.
Finally, the architecture connects active research data management with long-term preservation and publication through domain-specific, general-purpose, and institutional repositories. By considering the complete data life cycle before large-scale data generation begins, this concept aims to establish a scalable and sustainable framework for managing heterogeneous biomedical datasets within the RTG while facilitating collaboration, reproducibility, and future data reuse.