Speaker
Description
Gene expression analyses are necessary to understand gene function and unravel the genetic basis of biological processes. Omics technologies like transcriptomics enable such expression analyses via techniques like RNA-Seq. Transcriptomic technologies have been subjected to a number of transition stages and have witnessed rapid advancements. Falling prices and enhanced throughput of short read sequencing technologies are powering more transcriptomics studies. This in turn has resulted in public repositories like the sequence read archive (SRA), teeming with RNA-seq datasets that are valuable for various biological analyses like co-expression, differential expression analyses, and gene regulatory analyses to name a few. However, processing of these raw RNA-seq datasets requires large compute power, and is energy intensive. Therefore, availability of an updated expression collection that encompasses the new datasets will enable biological analyses at a fine-grained level and depth. Here, we introduce XpBrew, a Python workflow used to process RNA-Seq data and generate PanXpresso, a massive collection of gene expression datasets encompassing the different domains of life including plants, animals, fungi, archaea, and bacteria. Such large-scale expression data can facilitate a large number of studies pertaining to gene expression distribution analysis and gene function elucidation. Hosted in bonndata with a simple downloadable interface, PanXPresso provides a readily available transcriptomic resource for biologists to undertake a variety of analyses to address the myriads of biological questions. XpBrew is available for use in: https://github.com/PuckerLab/XpBrew and PanXPresso can be accessed at: https://doi.org/10.60507/FK2/OBIGQH
Key words: PanXpresso, Bonndata, RNA-seq, transcriptomics, big data