Speaker
Description
Large comparative studies increasingly depend on the availability of consistent, high-quality public datasets. We encountered this challenge while investigating the evolution of plant genes. To maximize the number of species in our analysis, we surveyed the publicly available genome sequences of flowering plants in the largest database at the NCBI. While genome sequences of many species are available, information about the position of genes in these sequences is largely missing. Therefore, we decided to identify genes in >3000 plant genome sequences ourselves and to make them accessible to the community. What initially appeared to be a biological analysis consequently became a large-scale data-processing problem. The genome sequences originate from many different sources and vary considerably in size and quality. Processing each dataset requires substantial computational resources and we also need to keep track of thousands of inputs, outputs, intermediate files, failures, and quality-control measures. This project illustrates challenges that arise when transforming a large, heterogeneous public dataset into a consistent and reusable research resource. Beyond computational throughput, questions of reproducibility, spreading workload, provenance, metadata, quality control, storage, and failure handling become central at this scale. We present this project as a case study in large-scale scientific data processing and invite discussion on how such workflows can be designed to remain reproducible, maintainable, and reusable as datasets, software, and computational infrastructure continue to evolve.