Efficient Single-Sample Gene Set Enrichment Scoring Methods for Large-Scale Datasets

Published on July 23, 2025
Written by Antonino Zito
⏱ 4 min read

The prominent problem of scalable single sample gene set enrichment scoring

Single-gene studies have uncovered individual dysregulated factors and undoubtedly advanced biomedical research.

However, complex diseases are driven by multi-factorial perturbations in gene sets which underlie the complex pathophysiological profiles. Emerging studies suggest that complex diseases’ therapeutic approaches would be more effective by targeting multiple members of a dysregulated gene set in the aim to restore the function of the entire pathway.

Thus, identifying biologically meaningful gene sets in a biological sample (an individual or a single cell, is critical).

Unfortunately, current computational methods demand excessive runtime and memory resources, becoming impractical for large datasets. Overcoming this limitation is crucial to support basic and clinical research in academia and the pharmaceutical industry.

PLAID: BigOmics Analytics’ direct solution to the problem

We recognize how serious the problem is. At BigOmics Analytics, we developed PLAID, a computational method that delivers a tailored solution. PLAID stands for Pathway Level Average Intensity Detection.

PLAID scores each gene set using the average intensity of its constituent genes within each sample. By relying on sparse matrix operations, PLAID is an ultrafast and memory optimized single sample gene set enrichment scoring algorithm.

It surpasses the performance of current methods in single-cell and bulk transcriptomics, and proteomics data.

How PLAID’s gene set enrichment results compare with existing methods

Best practices suggest cross-checking results from distinct methods. A strength of PLAID is its concordance in gene sets scoring with other, widely used enrichment scoring methods.

For example, in single-cell RNA-sequencing data, PLAID’s enrichment score values are well correlated with singscore, UCell and AUCell.

On the other end, the concordance with GSVA and ssGSEA can be improved by centering gene sets across samples, which results in a generalized higher concordance with other methods. These data assure the validity of PLAID and reinforce its reliability for cross-methods comparison and validation.

PLAID as a multi-methods platform for gene sets scoring

Biomedical research requires robust results. Robustness is ensured by reproducibility of research results with independent methods. Yet, for gene sets scoring, the bottleneck is the computational inefficiency of existing methods in large data.

PLAID resolves this issue. PLAID is equipped with the most widely used single-sample gene set enrichment scoring algorithms, namely singscore, scSE, ssGSEA, GSVA, UCell and AUCell, which use PLAID’s back-end of efficient sparse matrix calculation, thereby gaining significant computational efficiency.

How modern biomedical research may greatly benefit from PLAID

Modern biomedical research increasingly relies on large-scale data. Large-scale data enable the power of identifying subtle but biologically meaningful molecular profiles uniquely associated with a single patient and a specific condition.

These profiles often represent a group of genes, i.e., a gene set, exerting coordinated molecular activity within a sample. Ultimately, these represent individual-specific signatures usable in personalized medicine approaches.

By delivering matched accuracy and superior computational efficiency compared to existing single-sample enrichment scoring methods, PLAID provides a great benefit to analyses of large-scale data in modern biomedical research.

By also offering functionalities for cross-methods replication, PLAID may serve as a standalone single-sample gene set enrichment scoring platform in biomedical research studies.

Is PLAID available and usable?

Here is the good news! PLAID is currently undergoing peer-review. It will be soon available in  Omics Playground. Eager to read it and see all technical details?

Read more in our BioRxiv preprint “PLAID: ultrafast single-sample gene set enrichment scoring“!

Scientist with lens exploring the drug connectivity analysis tab in Omics Playground

Unlock the full potential of your RNA-seq and proteomics data!

About the Author

Antonino Zito

Antonino is a senior bioinformatics engineer at BigOmics with a strong background in bioinformatics and biostatistics. With a PhD in genetics and bioinformatics and an MSc in biotechnology, he has made significant contributions to computational analysis in numerous projects during his previous research at Harvard Medical School and King’s College London.