preprintbioRxiv (Cold Spring Harbor Laboratory)Aug 10, 2020GREEN OA

The MRC IEU OpenGWAS data infrastructure

University of Bristol · National Institute for Health and Care Research · +2 more institutions

Indexed incrossref

Abstract

Abstract Data generated by genome-wide association studies (GWAS) are growing fast with the linkage of biobank samples to health records, and expanding capture of high-dimensional molecular phenotypes. However the utility of these efforts can only be fully realised if their complete results are collected from their heterogeneous sources and formats, harmonised and made programmatically accessible. Here we present the OpenGWAS database, an open source, open access, scalable and high-performance cloud-based data infrastructure that imports and publishes complete GWAS summary datasets and metadata for the scientific community. Our import pipeline harmonises these datasets against dbSNP and the human genome…

Citation impact

954
total citations
FWCI
Percentile
References
36
Citations per year

Authors

14

Topics & keywords

Keywords
  • Python (programming language)
  • Computer science
  • Genome-wide association study
  • Biobank
  • Data mapping
  • Metadata
  • Data mining
  • Database
UN Sustainable Development Goals
  • Industry, innovation and infrastructure
No related works found for this paper.

Funding