DRIVE v3: Command Line Application for Identity-by-Descent Haplotype Clustering in Large Biobank Scale Data.
JT, B., HH, C., GF, E., AC, S., RJ, B., CD, H., QS, W., DC, S., & JE, B. (2026). DRIVE v3: Command Line Application for Identity-by-Descent Haplotype Clustering in Large Biobank Scale Data.. Genetic epidemiology. https://doi.org/10.1002/gepi.70048
JT B, HH C, GF E, AC S, RJ B, CD H, et al. DRIVE v3: Command Line Application for Identity-by-Descent Haplotype Clustering in Large Biobank Scale Data.. Genetic epidemiology. 2026; doi: 10.1002/gepi.70048
JT B, HH C, GF E, et al. DRIVE v3: Command Line Application for Identity-by-Descent Haplotype Clustering in Large Biobank Scale Data.[J]. Genetic epidemiology. 2026. DOI: 10.1002/gepi.70048.
@article{jt2026,
author = {Baker JT and Chen HH and Evans GF and Scartozzi AC and Bohlender RJ and Huff CD and Wells QS and Samuels DC and Below JE},
title = {DRIVE v3: Command Line Application for Identity-by-Descent Haplotype Clustering in Large Biobank Scale Data.},
journal = {Genetic epidemiology},
year = {2026},
doi = {10.1002/gepi.70048},
note = {PMID: 42363641},
}
TY - JOUR AU - Baker JT AU - Chen HH AU - Evans GF AU - Scartozzi AC AU - Bohlender RJ AU - Huff CD AU - Wells QS AU - Samuels DC AU - Below JE TI - DRIVE v3: Command Line Application for Identity-by-Descent Haplotype Clustering in Large Biobank Scale Data. T2 - Genetic epidemiology PY - 2026 DO - 10.1002/gepi.70048 AN - PMID:42363641 ER -
There is a need for genetic analytical methods that integrate multi-individual identity-by-descent (IBD) tools with phenotypic enrichment testing to discover novel shared haplotypes contributing to disease traits. Existing tools are designed to identify IBD sharing and leave interpretation and phenotype association tests to further analyses. Here we present Distant Relatedness for Identification and Variant Evaluation (DRIVE) v3, a python command-line interface tool that identifies networks of participants who share an identical haplotype at a given genomic location. Given phenotypic data, DRIVE additionally estimates significant enrichment of dichotomous traits within networks. DRIVE is designed for efficient use across large-scale genetic data resources, featuring a versatile application programming interface and a backend structure designed for flexible integration into existing analytical pipelines. In this work, we describe the implementation of DRIVE v3 and illustrate two applications of the tool to an autosomal dominant condition and to an autosomal recessive condition, cardiomyopathy and cystic fibrosis, respectively. These applications highlight the substantial performance improvements between v1 and v3 and demonstrate practically how the newer features of DRIVE such as the enrichment test can be used in the interpretation of the identified networks.