Briefings in Bioinformatics GDIIST in Collaboration with Macao Polytechnic Univ
2026.08.06

Recently, the research group of Feng Shengzhong at the Guangdong Institute of Intelligent Science and Technology, in collaboration with Macao Polytechnic University and Jiangsu Cancer Hospital, has made new progress in the analysis of single-cell transcriptomic data.

 

The research team proposed a regularized adaptive graphbased method for rarecell identification, termed RAG. This method estimates cellspecific neighborhood radii and models localscale affinities, enabling the graph structure to adaptively accommodate variations in sampling density across the feature space. Evaluations on multiple singlecell RNAsequencing datasets demonstrate that RAG substantially improves the accuracy and robustness of rarecell identification overall. The related work has been published as RAG: a regularized adaptive graphbased method for rarecell identification from singlecell expression datain Briefings in Bioinformatics, a JCR Q1 journal.

 

Background

 

Rare cells generally refer to cell subpopulations that are present in extremely low abundance within a sample and are difficult to clearly distinguish from adjacent major cell populations. Nevertheless, they may participate in critical biological processes such as immune responses, tumour drug resistance, and cell differentiation. Owing to their scarcity, limited sampling coverage, and uneven local sampling density in the feature space, the identification of such cells remains a technical challenge.

 

Graph algorithms, which characterize nodes and their interconnections, are widely employed in bioinformatics, social networks, traffic networks, recommendation systems, and knowledge graphs. In singlecell analysis, graph algorithms provide neighbourhood structural information for representation learning and community detection.

 

However, existing methods commonly adopt fixedsize *k*nearestneighbour graphs, i.e., retaining the same number of neighbours for each cell. In a feature space with uneven sampling densities, this fixed strategy tends to cause instability in adjacency relationships and population boundaries, particularly leading to the misincorporation of rare cells into adjacent dominant populations. Therefore, constructing a graph structure that can adaptively reflect local density differences is a key direction for improving rarecell identification performance.

 图片1(1).png

Figure 1. Comparison between fixedsize neighbourhood and RAG regularized adaptive neighbourhood.

 

Core Innovations

 

To address the above challenges, the research team proposed the RAG method and systematically integrated it into both the representation learning and clustering stages of singlecell data analysis.

 

Instead of using a fixed number of neighbours, the method first takes the union of Euclideandistance and cosinedistance neighbours as candidate neighbours, thereby balancing proximity in expression magnitude and similarity in expression pattern direction. Subsequently, in a locally normalized mixeddifference space, it estimates a cellspecific neighbourhood radius for each cell and retains valid adjacencies accordingly. Finally, mixed affinities of the retained edges are computed based on local scales, rendering the affinity values comparable across regions with different sampling densities.

 

Through this adaptively regularized graph structure, RAG markedly reduces weak bridging and crosspopulation connections, mitigating the risk of rare cells being incorrectly embedded into dominant populations.

 图片2(1).png

Figure 2. Overall workflow of the RAG method.

 

Experimental Validation

 

The research team systematically compared RAG with six representative methods on ten real singlecell RNAsequencing datasets. The results indicate that, relative to the bestperforming baseline methods, RAG achieves average relative improvements of 42%, 26%, and 35% in precision, F1score, and raretype coverage, respectively. In stratified subsampling experiments on two largescale datasets, the runtime of RAG scales approximately linearly with sample size and is more efficient than the baseline method with the best accuracy.

 图片3.png

Figure 3. Performance distribution of various rarecell identification methods on real datasets.

 

Biological Application Validation

 

In the singlecell analysis of the colorectal cancer metastasis sample CRC G1, RAG successfully identified all annotated rare cell populations in the dataset and further distinguished a NKrelated candidate subpopulation from T cells that expressed canonical NK marker genes. In the mouse airway epithelium data, the rare cell clusters resolved by RAG likewise recovered known rare populations and additionally identified two proliferative substates.

 图片4.png

Figure 4. Rare cell populations resolved by RAG based on the colorectal cancer metastasis sample CRC G1.

 图片5(1).png

Figure 5. Rare cell populations resolved by RAG based on mouse airway epithelium.

 

Significance and Future Perspectives

 

Against the backdrop of the rapidly advancing paradigm of AIdriven scientific research (AI4S), AI methods have demonstrated substantial potential in life sciences, chemistry, materials science, and other fields. The Guangdong Institute of Intelligent Science and Technology provides an exceptional platform for such interdisciplinary research.

 

Seizing this opportunity, the research group has deeply integrated biologydriven problem formulation with AI algorithmic innovation. Targeting the key bottleneck of rarecell identification in singlecell transcriptomic data, they proposed the regularized adaptive graph method RAG and validated its effectiveness and robustness across multiple realworld datasets.

 

This study not only furnishes a more adaptive graphlearning tool for singlecell data analysis but also lays a methodological foundation for investigating biological mechanisms such as immune responses, tumour drug resistance, and aberrant differentiation. It further highlights the unique value of AI4S in empowering precision medicine and fundamental research.

 

Wang Xingsu, a joint doctoral student between Macao Polytechnic University and the Guangdong Institute of Intelligent Science and Technology, is the first author of the paper. Professors Feng Shengzhong and Associate Professor Huang Dian from the Guangdong Institute of Intelligent Science and Technology are the cocorresponding authors. This work was supported by the Highlevel Talent Innovation Team Project of the GuangdongMacao Indepth Cooperation Zone in Hengqin, the National Natural Science Foundation of China, and other funding sources.

 

Article link:

https://academic.oup.com/bib/article/27/4/bbag379/8736956


COPYRIGHT © 2021

Copyright Guangdong Institute of Intelligence Science and Technology 

粤ICP备2021109615号 KCCN

公众号

Official Account