TY - JOUR
T1 - Learning inherent genetic patterns and trait associations with deep generative models for discrete genotype simulation
AU - Xie, Sihan
AU - Tribout, Thierry
AU - Boichard, Didier
AU - Hanczar, Blaise
AU - Chiquet, Julien
AU - Barrey, Eric
N1 - Publisher Copyright:
© The Author(s) 2026. Published by Oxford University Press on behalf of GigaScience. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
PY - 2026/1/1
Y1 - 2026/1/1
N2 - Background: Deep generative models open new avenues for simulating realistic genomic data while preserving privacy and addressing data accessibility constraints. While previous studies have primarily focused on generating gene expression or haplotype data, this study explores generating genotype data in both unconditioned and phenotype-conditioned settings, which is inherently more challenging due to the discrete nature of genotype data. Results: We developed and evaluated commonly used generative models, including Variational Autoencoders, Diffusion Models, and Generative Adversarial Networks, and proposed adaptation tailored to discrete genotype data. We conducted extensive experiments on large-scale datasets, including all chromosomes from cow and multiple chromosomes from human. Model performance was assessed using a well-established set of metrics drawn from both deep learning and quantitative genetics literature. Our results show that these models can effectively capture genetic patterns and preserve genotype–phenotype association. Conclusions: As deep generative models are able to reproduce key characteristics of genotype data, they can serve as direct tools for genotype–phenotype simulation, while also enabling privacy-preserving data sharing. Our findings provide a comprehensive evaluation of these models and offer practical guidance for future research in genotype-phenotype simulation.
AB - Background: Deep generative models open new avenues for simulating realistic genomic data while preserving privacy and addressing data accessibility constraints. While previous studies have primarily focused on generating gene expression or haplotype data, this study explores generating genotype data in both unconditioned and phenotype-conditioned settings, which is inherently more challenging due to the discrete nature of genotype data. Results: We developed and evaluated commonly used generative models, including Variational Autoencoders, Diffusion Models, and Generative Adversarial Networks, and proposed adaptation tailored to discrete genotype data. We conducted extensive experiments on large-scale datasets, including all chromosomes from cow and multiple chromosomes from human. Model performance was assessed using a well-established set of metrics drawn from both deep learning and quantitative genetics literature. Our results show that these models can effectively capture genetic patterns and preserve genotype–phenotype association. Conclusions: As deep generative models are able to reproduce key characteristics of genotype data, they can serve as direct tools for genotype–phenotype simulation, while also enabling privacy-preserving data sharing. Our findings provide a comprehensive evaluation of these models and offer practical guidance for future research in genotype-phenotype simulation.
KW - SNP
KW - deep generative models
KW - genomics
KW - genotype-phenotype simulation
KW - quantitative genetics
UR - https://www.scopus.com/pages/publications/105038552321
U2 - 10.1093/gigascience/giag044
DO - 10.1093/gigascience/giag044
M3 - Article
C2 - 41980277
AN - SCOPUS:105038552321
SN - 2047-217X
VL - 15
JO - GigaScience
JF - GigaScience
M1 - giag044
ER -