Logo image
Twelve quick steps for genome assembly and annotation in the classroom
Journal article   Open access   Peer reviewed

Twelve quick steps for genome assembly and annotation in the classroom

Hyungtaek Jung, Tomer Ventura, J. Sook Chung, Woo-Jin Kim, Bo-Hye Nam, Hee Jeong Kong, Young-Ok Kim, Min-Seung Jeon and Seong-il Eyun
PLoS Computational Biology, Vol.16(11), pp.1-25
2020
PMCID: PMC7660529
PMID: 33180771
pdf
Twelve quick steps for genome assembly and annotation in the classroom1.17 MBDownloadView
Published Version Open Access CC BY V4.0
url
https://doi.org/10.1371/journal.pcbi.1008325View
Published Version

Abstract

genome annotation computational pipelines genome sequencing genome analysis plant genomics DNA sequencing
Eukaryotic genome sequencing and de novo assembly, once the exclusive domain of well-funded international consortia, have become increasingly affordable, thus fitting the budgets of individual research groups. Third-generation long-read DNA sequencing technologies are increasingly used, providing extensive genomic toolkits that were once reserved for a few select model organisms. Generating high-quality genome assemblies and annotations for many aquatic species still presents significant challenges due to their large genome sizes, complexity, and high chromosome numbers. Indeed, selecting the most appropriate sequencing and software platforms and annotation pipelines for a new genome project can be daunting because tools often only work in limited contexts. In genomics, generating a high-quality genome assembly/annotation has become an indispensable tool for better understanding the biology of any species. Herein, we state 12 steps to help researchers get started in genome projects by presenting guidelines that are broadly applicable (to any species), sustainable over time, and cover all aspects of genome assembly and annotation projects from start to finish. We review some commonly used approaches, including practical methods to extract high-quality DNA and choices for the best sequencing platforms and library preparations. In addition, we discuss the range of potential bioinformatics pipelines, including structural and functional annotations (e.g., transposable elements and repetitive sequences). This paper also includes information on how to build a wide community for a genome project, the importance of data management, and how to make the data and results Findable, Accessible, Interoperable, and Reusable (FAIR) by submitting them to a public repository and sharing them with the research community.

Details

Metrics

170 File views/ downloads
37 Record Views

InCites Highlights

These are selected metrics from InCites Benchmarking & Analytics tool, related to this output

Collaboration types
Domestic collaboration
International collaboration
Web Of Science research areas
Biochemical Research Methods
Mathematical & Computational Biology

UN Sustainable Development Goals (SDGs)

This output has contributed to the advancement of the following goals:

#3 Good Health and Well-Being

Source: SDGs from InCites

Logo image