SciELO - Scientific Electronic Library Online

 
vol.33 issue2Clustering Residential Electricity Consumption Data to Create Archetypes that Capture Household Behaviour in South Africa author indexsubject indexarticles search
Home Pagealphabetic serial listing  

Services on Demand

Journal

Article

Indicators

    Related links

    • On index processCited by Google
    • On index processSimilars in Google

    Share


    South African Computer Journal

    On-line version ISSN 2313-7835Print version ISSN 1015-7999

    Abstract

    MPANGASE, Phelelani T. et al. nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data. SACJ [online]. 2021, vol.33, n.2, pp.1-16. ISSN 2313-7835.  https://doi.org/10.18489/sacj.v33i2.830.

    The rate of raw sequence production through Next-Generation Sequencing (NGS) has been growing exponentially due to improved technology and reduced costs. This has enabled researchers to answer many biological questions through "multi-omics" data analyses. Even though such data promises new insights into how biological systems function and understanding disease mechanisms, computational analyses performed on such large datasets comes with its challenges and potential pitfalls. The aim of this study was to develop a robust portable and reproducible bioinformatic pipeline for the automation of RNA sequencing (RNA-seq) data analyses. Using Nextflow as a workflow management system and Singularity for application containerisation, the nf-rnaSeqCount pipeline was developed for mapping raw RNA-seq reads to a reference genome and quantifying abundance of identified genomic features for differential gene expression analyses. The pipeline provides a quick and efficient way to obtain a matrix of read counts that can be used with tools such as DESeq2 and edgeR for differential expression analysis. Robust and flexible bioinformatic and computational pipelines for RNA-seq data analysis, from QC to sequence alignment and comparative analyses, will reduce analysis time, and increase accuracy and reproducibility of findings to promote transcriptome research. Categories: · Applied computing ~ Bioinformatics

    Keywords : bioinformatics; pipelines; workflows; nextflow; RNA-seq; singularity; container; reproducible.

            · text in English     · English ( pdf )