preprintbioRxiv (Cold Spring Harbor Laboratory)Jun 20, 2019GREEN OA

Evaluating Protein Transfer Learning with TAPE

Berkeley College · University of California, Berkeley · +2 more institutions

Indexed incrossref

Abstract

Abstract Protein modeling is an increasingly popular area of machine learning research. Semi-supervised learning has emerged as an important paradigm in protein modeling due to the high cost of acquiring supervised protein labels, but the current literature is fragmented when it comes to datasets and standardized evaluation techniques. To facilitate progress in this field, we introduce the Tasks Assessing Protein Embeddings (TAPE), a set of five biologically relevant semi-supervised learning tasks spread across different domains of protein biology. We curate tasks into specific training, validation, and test splits to ensure that each task tests biologically relevant generalization that transfers to real-life…

Citation impact

683
total citations
FWCI
Percentile
References
63
Citations per year

Authors

8

Topics & keywords

Keywords
  • Computer science
  • Artificial intelligence
  • Machine learning
  • Generalization
  • Transfer of learning
  • Set (abstract data type)
  • Task (project management)
  • Representation (politics)
UN Sustainable Development Goals
  • Industry, innovation and infrastructure
No related works found for this paper.

Funding