6 October 2026
Wissenschaftszentrum Bonn
Europe/Berlin timezone

DQ4Agri: From Data Quality Assessment to Reusable Metadata

Not scheduled
20m
Wissenschaftszentrum Bonn

Wissenschaftszentrum Bonn

Ahrstraße 45, 53175 Bonn

Speaker

Sundar Sripada Venugopalaswamy Sriraman (Geoinformation, IGG, University of Bonn)

Description

Data science in the agricultural domain faces unique data quality challenges. Agricultural datasets are often collected from sensors in the field, processed, published, and made accessible through Research Data Infrastructures (RDIs). The FAIRagro consortium, with more than 30 partners, is building a FAIR (Findable, Accessible, Interoperable, Reusable) research data management system for the agrosystems research community. The FAIR principles provide a blueprint to improve the management and accessibility of agricultural data. Specifically, FAIR principle R1 on the reusability of data states that "(meta)data are richly described with a plurality of accurate and relevant attributes" (citation).

For sensor-based time-series data, data quality directly impacts subsequent analysis and research outcomes. Researchers must therefore invest substantial time and effort in ensuring the quality of collected data. While vocabularies such as the W3C Data Quality Vocabulary (DQV) provide a framework for representing data quality, many quality characteristics relevant to agricultural sensor data still lack widely accepted domain-specific descriptors.

To help address this gap, DQ4Agri provides an algorithmic toolbox for assessing data quality early in the agricultural data lifecycle. It enables researchers to detect and react to potential data quality issues during the data collection process. In particular, DQ4Agri enables researchers to quickly monitor the quality of sensor data in the field and allows the resulting quality information to be exported in machine-readable XML and JSON-LD formats, supporting its subsequent incorporation into dataset metadata.

This poster highlights several lightweight algorithms built into DQ4Agri that provide a quick overview of a dataset's data quality. These include profiling algorithms that calculate global and sliding-window statistics, as well as outlier detection using z-scores and a custom local outlier factor (LOF)-based method. Furthermore, we present ongoing work on a Validator module that extends DQ4Agri with user-defined data quality rules. The Validator allows researchers to specify requirements for a dataset and check whether the collected data satisfy those requirements, complementing the existing profiling and outlier-detection functionality.

In summary, DQ4Agri enables researchers to extract data quality information during data collection and makes this information available for later reuse. Specifically, this information can subsequently be preserved alongside a dataset through its incorporation into the corresponding dataset metadata.

Author

Sundar Sripada Venugopalaswamy Sriraman (Geoinformation, IGG, University of Bonn)

Co-author

Jan-Henrik Haunert (PhenoRob, University of Bonn)

Presentation materials

There are no materials yet.