One of the core aims of differential expression analysis is to understand which pathways and areas of biology any differentially expressed genes represent. One method to determine this is Over Representation Analysis (ORA). In ORA the significantly different genes are compared to a database (e.g. GO terms, KEGG) of pre-defined lists of genes (gene-sets) - each of which represents a specific pathway or area of biology (e.g. Cell Cycle, TNF signalling). A hyper-geometric test is then used to determine whether each gene-set is enriched or not for the significantly differential genes. The most enriched gene-sets have the lowest p-values.
<br />
<br />
As individual genes can have several different functions, and several different gene-sets in a database might describe a similar biology (e.g. cell-cycle, G1 phase, replication) there is often considerable redundancy in gene-set databases and thus enrichment results. I.e. often there are many enriched gene-sets that contain very similar genes, and so are themselves very similar. This can often mean that the ten most enriched gene-sets are highly similar, and only represent a portion of the enriched biology. It is therefore highly useful to view all gene-set enrichment results as a network, where each node is a gene-set and each edge links gene-sets with highly similar genes. This often gives rise to clusters of gene-sets with similar function. When viewing these plots it is more important to consider the general biological theme of each cluster rather than the exact details of the gene-sets or genes within them. Networks are given for the enriched gene-sets for: all significantly differential genes, significantly upregulated genes and significantly downregulated genes.
