Recently, the research group of Feng Shengzhong at the Guangdong Institute of Intelligent Science and Technology, in collaboration with Macao Polytechnic University and Jiangsu Cancer Hospital, has made new progress in the analysis of single-cell transcriptomic data.
The research team proposed a regularized adaptive graph‑based method for rare‑cell identification, termed RAG. This method estimates cell‑specific neighborhood radii and models local‑scale affinities, enabling the graph structure to adaptively accommodate variations in sampling density across the feature space. Evaluations on multiple single‑cell RNA‑sequencing datasets demonstrate that RAG substantially improves the accuracy and robustness of rare‑cell identification overall. The related work has been published as “RAG: a regularized adaptive graph‑based method for rare‑cell identification from single‑cell expression data” in Briefings in Bioinformatics, a JCR Q1 journal.
Background
Rare cells generally refer to cell subpopulations that are present in extremely low abundance within a sample and are difficult to clearly distinguish from adjacent major cell populations. Nevertheless, they may participate in critical biological processes such as immune responses, tumour drug resistance, and cell differentiation. Owing to their scarcity, limited sampling coverage, and uneven local sampling density in the feature space, the identification of such cells remains a technical challenge.
Graph algorithms, which characterize nodes and their interconnections, are widely employed in bioinformatics, social networks, traffic networks, recommendation systems, and knowledge graphs. In single‑cell analysis, graph algorithms provide neighbourhood structural information for representation learning and community detection.
However, existing methods commonly adopt fixed‑size *k*‑nearest‑neighbour graphs, i.e., retaining the same number of neighbours for each cell. In a feature space with uneven sampling densities, this fixed strategy tends to cause instability in adjacency relationships and population boundaries, particularly leading to the misincorporation of rare cells into adjacent dominant populations. Therefore, constructing a graph structure that can adaptively reflect local density differences is a key direction for improving rare‑cell identification performance.

Figure 1. Comparison between fixed‑size neighbourhood and RAG regularized adaptive neighbourhood.
Core Innovations
To address the above challenges, the research team proposed the RAG method and systematically integrated it into both the representation learning and clustering stages of single‑cell data analysis.
Instead of using a fixed number of neighbours, the method first takes the union of Euclidean‑distance and cosine‑distance neighbours as candidate neighbours, thereby balancing proximity in expression magnitude and similarity in expression pattern direction. Subsequently, in a locally normalized mixed‑difference space, it estimates a cell‑specific neighbourhood radius for each cell and retains valid adjacencies accordingly. Finally, mixed affinities of the retained edges are computed based on local scales, rendering the affinity values comparable across regions with different sampling densities.
Through this adaptively regularized graph structure, RAG markedly reduces weak bridging and cross‑population connections, mitigating the risk of rare cells being incorrectly embedded into dominant populations.

Figure 2. Overall workflow of the RAG method.
Experimental Validation
The research team systematically compared RAG with six representative methods on ten real single‑cell RNA‑sequencing datasets. The results indicate that, relative to the best‑performing baseline methods, RAG achieves average relative improvements of 42%, 26%, and 35% in precision, F1‑score, and rare‑type coverage, respectively. In stratified subsampling experiments on two large‑scale datasets, the runtime of RAG scales approximately linearly with sample size and is more efficient than the baseline method with the best accuracy.

Figure 3. Performance distribution of various rare‑cell identification methods on real datasets.
Biological Application Validation
In the single‑cell analysis of the colorectal cancer metastasis sample CRC G1, RAG successfully identified all annotated rare cell populations in the dataset and further distinguished a NK‑related candidate subpopulation from T cells that expressed canonical NK marker genes. In the mouse airway epithelium data, the rare cell clusters resolved by RAG likewise recovered known rare populations and additionally identified two proliferative substates.

Figure 4. Rare cell populations resolved by RAG based on the colorectal cancer metastasis sample CRC G1.

Figure 5. Rare cell populations resolved by RAG based on mouse airway epithelium.
Significance and Future Perspectives
Against the backdrop of the rapidly advancing paradigm of AI‑driven scientific research (AI4S), AI methods have demonstrated substantial potential in life sciences, chemistry, materials science, and other fields. The Guangdong Institute of Intelligent Science and Technology provides an exceptional platform for such interdisciplinary research.
Seizing this opportunity, the research group has deeply integrated biology‑driven problem formulation with AI algorithmic innovation. Targeting the key bottleneck of rare‑cell identification in single‑cell transcriptomic data, they proposed the regularized adaptive graph method RAG and validated its effectiveness and robustness across multiple real‑world datasets.
This study not only furnishes a more adaptive graph‑learning tool for single‑cell data analysis but also lays a methodological foundation for investigating biological mechanisms such as immune responses, tumour drug resistance, and aberrant differentiation. It further highlights the unique value of AI4S in empowering precision medicine and fundamental research.
Wang Xingsu, a joint doctoral student between Macao Polytechnic University and the Guangdong Institute of Intelligent Science and Technology, is the first author of the paper. Professors Feng Shengzhong and Associate Professor Huang Dian from the Guangdong Institute of Intelligent Science and Technology are the co‑corresponding authors. This work was supported by the High‑level Talent Innovation Team Project of the Guangdong‑Macao In‑depth Cooperation Zone in Hengqin, the National Natural Science Foundation of China, and other funding sources.
Article link:
https://academic.oup.com/bib/article/27/4/bbag379/8736956
COPYRIGHT © 2021
Copyright Guangdong Institute of Intelligence Science and Technology
粤ICP备2021109615号 KCCN
Official Account