login

A MULTIVARIATE EXTENSION OF THE GENE SET ENRICHMENT ANALYSIS

Journal of Bioinformatics and Computational BiologyPublished 1 October 2007
Lev B. Klebanov, Galina Glazko, Peter Salzman, Andrei Yakovlev, Yuanhui Xiao
Citations38
SJR quartileQ4
SJR score0.23
SNIP0.27

TL;DR

Simulation studies and analysis of biological data confirm the conjecture that the N-statistic is a much better choice for multivariate significance testing within the framework of the GSEA.

Abstract

A test-statistic typically employed in the gene set enrichment analysis (GSEA) prevents this method from being genuinely multivariate. In particular, this statistic is insensitive to changes in the correlation structure of the gene sets of interest. The present paper considers the utility of an alternative test-statistic in designing the confirmatory component of the GSEA. This statistic is based on a pertinent distance between joint distributions of expression levels of genes included in the set of interest. The null distribution of the proposed test-statistic, known as the multivariate N-statistic, is obtained by permuting group labels. Our simulation studies and analysis of biological data confirm the conjecture that the N-statistic is a much better choice for multivariate significance testing within the framework of the GSEA. We also discuss some other aspects of the GSEA paradigm and suggest new avenues for future research.

Keywords

Biochemistry, Genetics and Molecular Biology