Speaker
Description
Data-sharing of clinical patient-level data has the potential to accelerate research but is strongly restricted by data protection regulations such as the GDPR. Synthetic data generation (SDG) is a promising alternative to direct data-sharing, but release of synthetic data requires evaluation of fidelity, privacy, and downstream utility. These assessments are non-trivial and often require technical expertise of domain experts, which limits the accessibility of SDG in clinical research.
To address this, we developed Syndat, a user-friendly dashboard for automated assessment of synthetic patient-level data. Syndat computes normalised scores for both synthetic data fidelity and privacy through an automated pipeline that requires no prior expert knowledge. It also provides visualizations such as outlier plots, feature distributions, and pairwise correlations, as well as post-processing functions to improve synthetic data quality. Syndat is fully open source and available both as a dashboard and as a Python package.
Using Syndat, we evaluated several privacy-enforcing SDG and anonymization methods on three real-world clinical datasets. Differential Privacy notably reduced feature correlations, while non-DP synthetic data retained high fidelity in two of three studies, with no strong evidence of privacy breaches. K-anonymization, in contrast, introduced notable privacy risks. High-fidelity synthetic data showed utility comparable to real data in two downstream tasks.
Our results show that high fidelity data measured by Syndat metrics can successfully be applied to solve real world tasks. Syndat can support synthetic data release in strict clinical settings, by serving as a tool for both experts and non-experts, providing easily interpretable measures for synthetic data quality.
Ongoing future work on Syndat aimed to address current limitations focuses on extending existing metrics and adding support for utility evaluation and potential on-site SDG.