Innovation Spotlight: The Phenotype-Genotype Reference Map

AI&DHI Innovation Spotlight
 

 

Dr. Bastarache (left) and Dr. Zawistowski (right)

This edition of the AI&DHI Innovation Spotlight features an award-winning collaboration between the University of Michigan and Vanderbilt University.

Teams led by Dr. Lisa Bastarache, Research Associate Professor of Biomedical Informatics at Vanderbilt, and Dr. Matthew Zawistowski, Clinical Associate Professor of Biostatistics at the U-M School of Public Health, have landed the NIH Replication Prize for their work developing the Phenotype-Genotype Reference Map (PGRM)—an open-science tool for evaluating data quality and integrity across large-scale electronic health record (EHR)-linked biobanks.


Biobanks allow researchers to study connections between genetics, health and disease across large populations. Some biobanks, such as the Michigan Genomics Initiative (MGI) at U-M, link participants’ genetic information with data from their electronic health records (EHRs).

While this combination can support a wide range of studies by providing investigators with a rich and continually evolving picture of health, it is also susceptible to selection bias and data errors that can potentially lead to false discoveries. For example, a biobank that recruited participants through an academic medical center may disproportionately include people with chronic or severe illnesses and/or who have better access to specialty care, while potentially underrepresenting healthier individuals and people who face barriers to care.

This selection process can create misleading associations. A genetic variant may appear to be associated with a disease partly because those with the disease interact with the health care system more often and, as a result, have had more opportunities to be tested, diagnosed and enrolled.

“Without careful validation and adjustment for health care utilization and enrollment patterns, researchers may mistake patterns in documentation and biobank participation for true biological relationships,” said Dr. Zawistowski.

So, how can researchers who are conducting biobank-based studies feel confident that their results are true biological discoveries rather than technical artifacts?

With this question in mind, the U-M and Vanderbilt teams developed the PGRM to provide a benchmark for evaluating the quality and integrity of EHR-linked biobanks by testing whether a dataset can reproduce genetic associations that have already been well established through prior research.

Graphical representation of a circular genome sequence map.

(Image credit: Adobe Stock) Each association represents a relationship between a genetic variant and a health-related characteristic that researchers should reasonably expect to observe again in a sufficiently large biobank with high-quality data and sound analytical procedures.

The tool contains nearly 6,000 “true positive” associations spanning 149 diseases selected from a catalog maintained by the National Human Genome Research Institute and the European Bioinformatics Institute. Each association represents a relationship between a genetic variant and a health-related characteristic that researchers should reasonably expect to observe again in a sufficiently large biobank with high-quality data and sound analytical procedures.

“If a biobank analysis reproduces many of these known associations, then researchers can have greater confidence that the data and methods are working as expected,” said Zawistowski. “If it does not, then the results can point to potential problems that should be investigated before researchers draw conclusions.”

Interdisciplinary teams across U-M and Vanderbilt built and tested PGRM across several biobanks, including MGI, Vanderbilt’s BioVU, BioBank Japan and the UK Biobank. The project received support from AI&DHI, with the team using existing results generated by AI&DHI researcher Anita Pandit, MS, through her routine work.

”Having access to these results really sped up the project,” said Zawistowski. “It also highlighted the importance of the work being done by AI&DHI staff and the valuable data they are making available to the U-M research community.”

 

Accomplishments of the PGRM

The NIH Replication Prize that the researchers were awarded was launched to recognize and reward progress in making key areas of biomedical research more replicable. It also aims to encourage a culture in which replication is treated as a standard and essential part of the scientific process.

The team was recognized under the competition’s “Track 2: Replication Exemplars” category. According to the NIH, this track honors “pioneering researchers who have creatively and successfully integrated replication into their standard research practice,” with strategies that have demonstrated success in improving research rigor and strengthening trust in scientific outcomes.

“Having access to these results really sped up the project. It also highlighted the importance of the work being done by AI&DHI staff and the valuable data they are making available to the U-M research community.”
— Matthew Zawistowski, PhD

The Phenotype-Genotype Reference Map has since been used as the primary method for validating data quality and integrity for the NIH’s landmark All of Us dataset. The All of Us program is collecting health data from one million or more people in the United States, with an emphasis on building a participant group that reflects the country’s diversity. The dataset is designed to support research across many diseases and areas of health.

“Matt and the Michigan team played an essential role in this project,” Dr. Bastarache said. “MGI is a premier biobank and an important benchmark dataset for the PGRM. It was a pleasure to collaborate with scientists who care about rigor and reproducibility.”

 Those interested in learning more about PGRM are encouraged to view the team’s paper published in The American Journal of Human Genetics. PGRM itself is also freely available to the research community as an R software package with its code and reference datasets accessible through GitHub.

 


Referenced Study

Bastarache, L., Delozier, S., Pandit, A., He, J., Lewis, A., Annis, A. C., LeFaive, J., Denny, J. C., Carroll, R. J., Altman, R. B., Hughey, J. J., Zawistowski, M., & Peterson, J. F. (2023). The phenotype-genotype reference map: Improving biobank data science through replication. The American Journal of Human Genetics, 110(9), 1522–1533. https://doi.org/10.1016/j.ajhg.2023.07.012

 

About AI & Digital Health Innovation

AI & Digital Health Innovation (formerly Precision Health at U-M) is dedicated to empowering researchers at the University Michigan to change the future of digital healthcare. They work with multi-disciplinary teams of health providers, basic scientists, engineers, and administrators to tackle the most difficult research problems and help rapidly bring ideas to the bedside. For more information visit aidhi.umich.edu.

Next
Next

AI&DHI researchers land $2.9 million grant to advance the science of AI at the bedside