Speakers
Description
Machine Learning (ML) in agriculture is often approached from a model-centric perspective, where the primary focus lies on developing and optimizing model architectures and learning algorithms, rather than on the data itself. This is reflected, for example, in the common practice of training and evaluating ML models on static benchmark datasets with the goal of achieving high predictive performance and computational efficiency. However, these benchmark datasets are far from being representative of the real-world conditions in agricultural contexts, leading to poor generalisability of the trained models.
Agricultural data are particularly challenging as they are heterogeneous, spatially and temporally variable, costly to acquire, and often only partially representative of the conditions encountered during deployment. Data-centric machine learning shifts the focus from optimising models on static benchmark datasets towards systematically improving the data used throughout the machine learning lifecycle.
This spans data creation, such as compiling task-specific datasets from diverse sources or scaling in-situ measurements to mapped areas; curation strategies such as detecting and handling noisy or influential samples; training methods that explicitly leverage quality information; and evaluation approaches that go beyond aggregated accuracy metrics, for example by adding uncertainty estimates.
This poster addresses questions that extend beyond agriculture. Therefore, we aim to use the poster to discuss these questions across disciplines and identify common concepts, methods, and infrastructure needs for moving from simply collecting more data to creating and maintaining data that enable reliable, reproducible scientific conclusions.