Graph Learning · Research survey

Emerging Topics

As graph learning matures, a range of emerging topics are expanding its scope and deepening its integration with cutting-edge developments in AI. Moving beyond traditional tasks like node classification and link prediction, researchers are now exploring more versatile, scalable, and intelligent approaches that push the boundaries of what graph-based models can achieve [3, 63]. These emerging directions not only address new technical challenges but also reflect the broader AI agenda of building more generalizable, distributed, and knowledge-driven systems.

This section discusses several key frontiers in graph learning. We begin with graph foundation models, which aim to develop large-scale, pretrained graph encoders capable of generalizing across diverse tasks and domains. We then explore graph reinforcement learning, where graph structures are used to model complex environments and interactions for sequential decision-making. We also discuss federated graph learning, which enables collaborative model training across decentralized graph data sources while preserving privacy. Additionally, we cover advances in learning on knowledge graphs (KGs) and knowledge-infused graph learning, which leverage symbolic information to enhance reasoning and generalization. Finally, we examine the emerging field of quantum graph learning, which explores quantum computing techniques for more efficient and expressive graph representations. These topics represent the next wave of innovation, positioning graph learning as a central pillar in the development of intelligent, explainable, and scalable AI systems.

Graph Foundation Models

Graph Foundation Models (GFMs) are large-scale pre-trained models specifically designed for processing graph-structured data. These models are generally pre-trained on vast amounts of data across various domains [395, 396]. Besides, GFMs can adapt to varying demands with minimal fine-tuning or even in few-shot/zero-shot scenarios. There exist many open issues to be addressed, and we discuss them from three distinct perspectives: pre-training, architecture design, and adaptation strategies.

Pre-training Strategies

Pre-training involves training a model on large-scale data in advance. The goal is to create a model with transferable knowledge that can be easily fine-tuned for specific downstream tasks, improving performance even with limited target task data. Current GFMs mainly adopt supervised pre-training and self-supervised pre-training.

While pre-training has shown promise, several challenges remain that hinder its broader applicability. Below, we highlight key areas that require further exploration and innovation.

Architecture Design

GFMs have evolved through a variety of architectural paradigms, each designed to capture different aspects of graph data. From classic message-passing frameworks to more recent innovations like Graph Transformers and LLMs, these architectures have significantly advanced. The existing architectures for GFMs are as follows.

However, several challenges in architectural design continue to limit their effectiveness. Below, we outline key areas that require further innovation.

Tuning, Adaptation and Compression

GFMs for specific downstream applications requires a combination of adaptation strategies. This means that pre-trained models can be effectively fine-tuned for targeted tasks and efficiently deployed in various environments. The existing transfer strategies for GFMs are as follows:

Despite the advances in tuning, adaptation, and compression strategies for GFMs, several open issues persist that limit their effectiveness in applications.

GFMs are emerging as a unified framework for graph-based learning, combining large-scale pretraining, GTs, and flexible fine-tuning to support diverse applications from social networks to scientific discovery. As the field matures, GFMs may become core components of graph learning and become a unified tool across graph tasks.

Graph Reinforcement Learning

Deep Reinforcement Learning (DRL) is a subfield of machine learning that integrates deep learning with reinforcement learning [406]. DRL training agent to make decisions by learning optimal policy through interactions with the environment, using deep neural networks to handle high-dimensional state spaces and approximate complex value functions [407]. The sequential decision-making process can be modeled using a Markov Decision Process (MDP) or an extended MDP, such as a partially observable Markov decision process (POMDP) or Semi-MDP. Specifically, an MDP can be defined as \(<S,A,T,R>\), where \(S\) is state space, \(A\) is the action space, \(T\) is state transition probability, and \(R\) is the reward value obtained after actor \(a\) is performed. This approach enables the agent to maximize cumulative rewards over time in sequential decision-making tasks.

Graph Reinforcement Learning (GRL) is a novel method that combines DRL with graph representation, enabling the solution of decision-making problems based on graph data [408]. In DRL problems, graph-structured data provides rich relational information, especially in multi-agent scenarios, which is unavailable in non-graph-structured states. Additionally, incorporating DRL into graph learning introduces a new training approach that differs from existing graph learning tasks, which often require substantial expert knowledge. DRL allows for learning without prior knowledge. We will introduce GRL from two aspects: graph-enhanced reinforcement learning and reinforcement learning enhanced graph learning.

Graph Enhanced DRL

In multi-agent DRL, a group of agents cooperate or compete to achieve a common goal, where the key is to understand the mutual interplay between agents. This architecture is widely used in various tasks, such as traffic signal control [409], vehicle control systems [410], and network routing [411]. The communication between agents provides information about the states and environments of other agents. Various attention-based methods have been proposed to address communication issues in multi-agent systems. In G2ANet [412], the authors used a complete graph to model the relationships between agents in large-scale multi-agent scenarios and designed a two-stage attention model as the communication model. Similarly, in Graphocomm [413], the authors applied a relational graph module to model explicit and implicit relationships through explicit and implicit relation layers. Specifically, explicit relationships are built from prior knowledge, while implicit relationships are learned through interactions among agents. In SRI-AC [414], the authors deployed a Variational Autoencoder (VAE) model to predict interactions and learn a state representation from observational data.

In multi-task DRL, states and actions can vary significantly across tasks. However, this challenge can be effectively addressed by leveraging the generalization capabilities of GNNs, which facilitate efficient multi-source policy transfer learning in settings with state-action mismatches. In TURRET [415], GNNs are used to learn the intrinsic properties of agents, creating a unified state embedding space for all tasks. This approach enables TURRET to achieve more efficient transfer and stronger generalization across tasks, making it easily compatible with existing DRL algorithms.

DRL Enhanced Graph Learning

Most graph representation methods either use global pooling techniques to generate a single global representation vector, which overlooks the semantics of substructures, or rely on expert knowledge to extract local substructures [416]. This often results in poor interpretability and limited generalization ability [417]. By utilizing reinforcement learning, it is possible to capture the most significant subgraphs without the need for expert knowledge [418], offering a more flexible and adaptive approach to graph representation. By defining the update value \(k\) within the MDP framework and using Q-learning to learn the optimal policy, where \(k\) serves as the pooling ratio in top-k sampling. Similarly, the identification of an optimal subset can be formulated as a combinatorial optimization problem, where the goal is to select the best combination of elements from a larger set. Deep Q-learning can be applied to solve this problem by learning the value of each possible action and selecting the combination that maximizes the overall reward [415].

Optimizing the GNN framework and hyperparameters can significantly enhance learning capabilities. However, modifying these components requires extensive domain knowledge. Utilizing Neural Architecture Search (NAS) allows for the selection of suitable GNN models and their corresponding hyperparameters tailored to specific tasks. Due to the highly diverse nature of graph classification applications, designing data-specific pooling methods based on human expertise presents a significant challenge [419]. To address this, PAS [420] introduces a unified pooling framework that simplifies the design of the search space. Two variants, PAS-G and PAS-NE, are also proposed, each offering distinct approaches to improve the effectiveness of pooling in different scenarios.

Federated Graph Learning

Although graph learning techniques have significantly progressed across various domains in recent years, most existing graph neural networks still rely on centralized storage of large-scale graph data for training. However, growing concerns over data security and user privacy make this requirement impractical in real-world scenarios [421]. In many practical applications, graph data is often distributed across multiple data holders, and due to privacy concerns and regulatory constraints, no party can directly access the data of others. This data distribution across different devices and owners, coupled with privacy concerns, makes it increasingly challenging to learn a global model while preserving data privacy [422].

For instance, in the financial sector, a third-party company may need to train a graph learning model for multiple financial institutions to detect potential financial crimes and fraudulent activities. Each financial institution holds its local customer dataset, including demographic information and transaction records between customers. Based on this data, each institution forms a customer graph, where the edges represent transaction relationships. Due to strict privacy policies and industry competition, these local customer datasets cannot be directly shared with third-party companies or with other institutions. At the same time, there may be interactions between different financial institutions, representing structural information across institutions. The main challenge for the third-party company lies in training a graph learning model for financial crime detection based on both the local customer graphs and the structural information between institutions, without directly accessing the local data of each institution [423].

To address this challenge, federated learning (FL) has emerged as a promising distributed learning framework that tackles the issues of data isolation and privacy preservation [424]. FL allows participants to collaboratively train a global model without sharing their local data, and it has been widely applied to Euclidean data tasks, such as image classification, by aggregating model updates from multiple clients [425, 426]. However, the standard FL framework struggles to handle the complex relationships inherent in graph-structured data, limiting its direct application to graph learning scenarios [427].

In response, Federated Graph Learning (FGL) has emerged as a promising solution [428, 429, 430]. FGL combines the strengths of federated learning and graph neural networks, enabling multiple data holders to collaboratively train a graph learning model without sharing their local graph data. This method is particularly suitable to distributed graph data scenarios, such as analyzing customer transaction data across multiple financial institutions. Each institution holds a local customer graph, and the interactions between institutions form cross-institutional structural information. By integrating local graphs with inter-institutional structural information while ensuring data privacy, FGL can construct a global graph model, significantly enhancing the detection of financial crimes and other related tasks.

FGL can be further categorized into three types based on how graph structure information is distributed among participants: graph federated learning, subgraph federated learning, and graph structured federated learning. By incorporating structural information from graph data, FGL not only enhances distributed graph learning but also boosts performance in cross-institutional and cross-domain collaboration, while maintaining privacy protection [423, 421].

Graph Federated Learning

Graph federated learning is a natural extension of standard FL, where each participant locally holds graph structured data, and the global model is trained to perform graph level tasks. A typical application of graph federated learning is in AI4Science, such as drug discovery and molecular property prediction. In these cases, GNNs are used to study the graph structure of molecules, where molecules are represented as graphs, with atoms as nodes and chemical bonds as edges [429]. Each bio-pharmaceutical company has its private dataset, containing molecular nodes and their corresponding properties. Historically, commercial competition and data privacy concerns have hindered collaboration between these companies. However, with the graph federated learning framework, it is now possible to jointly train molecular structure models, enabling clients to collaborate on building robust models while preserving data privacy.

Formally, each client \(k\) owns its private local graph data \(\mathcal{D}_k=\left\lbrace \mathcal{G}_1, \mathcal{G}_2, \cdots \right\rbrace\), where each \(\mathcal{G}_i=(\mathcal{V}, \mathcal{E})\) is a graph comprising a set of nodes \(\mathcal{V}\) and a set of edges \(\mathcal{E}\). Each client collaborates with others to train a robust global graph model \(\mathcal{F}\) based on its local dataset \(\mathcal{D}_k\), while ensuring that \(\mathcal{D}_k\) remains local and private. A significant challenge in this setting is the heterogeneity of the data across clients, which may differ significantly in terms of node features and graph structures [431]. This data heterogeneity can lead to severe model divergence during the federated process, thereby degrading the performance of the global model. To address this, various methods have been proposed to either train a single global model with better generalization or develop personalized models that can adapt to the unique characteristics of each client’s data [428, 432].

Subgraph Federated Learning

In subgraph federated learning, each participant’s local data is treated as a subgraph of a larger, more comprehensive global graph [433]. Participants focus on using the nodes and edges within their subgraphs as training samples. Without directly sharing data, the participants collaboratively train a global graph model. Subgraph federated learning has broad application prospects in fields like social network analysis, recommendation systems, and financial risk control, where graph data plays a central role. It offers a solution to handle large-scale, distributed graph data while safeguarding privacy and improving model performance.

However, subgraph federated learning faces several challenges [434]. First, each client only holds a subgraph of the original global graph, and due to privacy concerns and the isolated storage of data, each node can only access the information of neighboring nodes within its subgraph, lacking access to information from other clients. This missing cross-client information leads to biased node embeddings on each client, thereby reducing the performance of the graph learning model [435]. A key challenge is to reconstruct the missing cross-client information to accurately compute the node embeddings. Another challenge is that the same node may belong to multiple clients. In such cases, during collaborative training, the embeddings of overlapping nodes from different clients may originate from different embedding spaces. The core idea to address this is to learn global node embeddings based on the local embeddings of each client, and an overlapping instance alignment technique has been introduced to ensure the consistency of node embeddings across clients [436].

Graph Structured Federated Learning

In real-world scenarios, clients are often interconnected. For instance, in traffic flow prediction tasks, sensor devices distributed across different geographical locations may have connections, and these links often contain rich information, such as spatial dependencies [422].

In graph structured federated learning, each client is treated as a node, and together, all the clients form a graph \(\mathcal{G} = (\mathcal{V}, \mathcal{E})\), where \(\mathcal{V}\) represents the set of clients and \(\mathcal{E}\) represents the edges between them. This approach takes into account the topological structure between clients and uses GNNs to aggregate models based on these topological relationships [423]. It is important to note that the data on the client side is not necessarily graph-structured.

Typically, clients upload their local model parameters or local embeddings to the server, just as in standard FL. The server then aggregates the uploaded models using graph-based algorithms, taking into account the client graph \(\mathcal{G}\), and finally distributes the updated model parameters back to the clients [437, 438, 439].

Learning on Knowledge Graphs

KGs represent structural relations between entities, storing human knowledge facts in intuitive graph structures and being towards human-level intelligence [440, 441]. Learning on KGs refers to combining the structured data represented in knowledge graphs with graph learning techniques [442, 443, 7]. Although KGs have made significant success in various real-world fields, how to use these semantically structured data to assist graph learning still faces challenges. Firstly, KGs are often incomplete with missing links between entities, resulting in a degradation of the quality of the knowledge graph and impacting model performance. Thus, it poses challenges in the knowledge extraction process to construct high-quality KGs. Secondly, KGs grow exponentially due to the quick updating of information in the real world. Thus, it poses challenges in proposing efficient representation methods to handle the scalability of KGs.

Knowledge Extraction

Knowledge extraction (or knowledge acquisition) tasks can be divided into three categories: knowledge graph completion (KGC), relation extraction, and entity discovery [441]. KGC aims to predict missing links in existing KGs, and embedding-based ranking, relation path reasoning, rule-based reasoning, and meta relational learning are main categories for KGC. Compared to KGC, relation extraction and entity discovery are for discovering new knowledge (e.g., relations and entities) from the text. Relation extraction models utilize attention mechanisms, GCNs, adversarial training, reinforcement learning, deep residual learning, and transfer learning. KGC, relation extraction, and entity discovery are three separate tasks for knowledge extraction. In order to simplify development and be convenient for knowledge extraction, proposing a unified framework is one of the challenges in this field. For example, [444] proposed a joint learning framework with mutual attention for data fusion between knowledge graphs and text, which simultaneously solves KGC and relation extraction from text. Additionally, knowledge extraction also face other challenges in KGC, relation extraction, and entity discovery as follows.

Knowledge Representation

Knowledge representation learning (KGL) is a critical research issue of KGs, which paves the way for many downstream applications [459, 460]. Existing KGL focuses on modeling the semantic interaction of facts and utilizing external information. GNNs are introduced to learn the connectivity structure in KGs under an encoder-decoder framework. R-GCN proposes a relation-specific transformation to model the directed nature of knowledge graphs, and GCN acts as a graph encoder [461]. SACN introduces weighted GCN to define the strength of two adjacent nodes with the same relation type, capturing the structural information within KGs [462].

Although embedding representations can capture relationships between knowledge elements, it is challenging to infer explicit logical relationships directly from these vector representations. Designing knowledge representation models that can provide interpretable reasoning while maintaining efficient computational performance remains an important research direction. Additionally, a unified understanding of knowledge representation is less explored. The field of knowledge representation still faces numerous challenges that hinder its broader application. Below are some key areas that require further exploration:

Knowledge-infused Graph Learning

Graph learning has achieved significant success in knowledge graph tasks, such as knowledge extraction and knowledge representation [448, 449]. In the domain of graph learning, models are operated under the assumption that all data is of high quality, and graph learning techniques are extended to enhance the performance of downstream tasks. However, challenges related to low-quality data and poor generalization when applying these methods are encountered, leading to subpar model performance on instances not seen during training [463]. To address these limitations, the concept of knowledge-infused graph learning has been introduced. This approach integrates external knowledge into various components of the graph machine learning pipeline, aiming for more accurate results. Knowledge can be categorized into formal scientific knowledge (e.g., established laws or theories that govern the behaviour of target variables) and informal experimental knowledge (e.g., well-known facts derived from long-term observations or human reasoning).

Compared to traditional graph machine learning techniques, incorporating human knowledge provides several advantages: (i) it allows for the successful integration of vast amounts of information, (ii) it enhances reliability in results, and (iii) it facilitates meaningful inferences, which can assist individuals lacking domain expertise. Knowledge-infused graph learning methods can be seen as employing knowledge in two ways: as prompts or as augmentation. “Knowledge as prompts” refers to using external knowledge to guide model learning and decision-making processes. Conversely, “knowledge as augmentation” involves analyzing a graph \(\mathcal{G=(V,E)}\) in conjunction with a relevant knowledge dataset \(\mathcal{D}\). The primary objectives of this approach are twofold: to learn a function that generates vector representations for graph elements, \(f:(\mathcal{G,D}) \rightarrow \textbf{Z}\), capturing both the structure and semantics of the graph alongside the knowledge stored in the knowledge database, and to retrieve direct and precise answers (or explanations) from the dataset \(\mathcal{D}\) to address user queries.

Despite the considerable progress of knowledge-infused graph learning in downstream applications such as large language models pre-training and drug discovery, challenges remain, including knowledge database composition and effective knowledge integration. These challenges will be discussed in the following sections.

Knowledge as Prompt

The prompt-based fine-tuning method originates from natural language processing and has been widely used to help pre-trained language models adapt to various downstream tasks [464]. Knowledge, as the initial guidance or prompt in the graph learning process, aims to assist the model in understanding the structure and relationships within the data, thereby generating more effective learning outcomes. There are two main types of prompt methods:

Although knowledge as a prompt has achieved some success in the field of graph learning, several challenges still remain:

Knowledge as Augmentation

Typical graph analysis applications include node role identification, personalised recommendation, social healthcare, academic network analysis, graph classification and epidemic trend study [468, 469]. Regarding external knowledge as data augmentation helps incorporate with graph learning for more accurate performance. For example, knowledge-augmented graph learning methods have achieved more precise drug discovery with limited training data [463]. To illustrate it well, we will discuss challenges from these four directions: incorporating knowledge in preprocessing, pretraining, training, and interpretability.

Quantum Graph Learning

Quantum mechanics is increasingly being explored to enhance graph learning techniques, giving rise to the emerging field of Quantum Graph Learning (QGL) [478]. Traditional graph learning methods, which focus on understanding and predicting relationships in graph-structured data, face challenges like scalability and capturing intricate node interactions. By integrating principles of quantum mechanics (e.g., superposition and entanglement), QGL aims to revolutionize this process. Quantum algorithms have the potential to process complex graph structures more efficiently, enabling faster computations and deeper insights into the connectivity patterns. As QGL is still in its infancy, researchers are actively exploring quantum models and algorithms that can be adapted to current graph learning frameworks.

Message-passing GNNs face significant limitations when it comes to efficiently capturing long-range dependencies between nodes. As the distance between nodes increases, the efficiency of information transfer diminishes, leading to issues such as oversmoothing, where node representations become indistinguishable. When representing data in quantum states, the quantum coherence or entanglement transfers long-distance information with a faster time. Quantum graph learning provides a potential solution by leveraging quantum phenomena like quantum coherence and entanglement [479, 480]. Quantum coherence allows quantum systems to exist in multiple states simultaneously. In this framework, entanglement plays a pivotal role by facilitating the transfer of information between distant nodes through non-local effects, allowing long-range dependencies to be captured with much greater precision and efficiency.

Non-Euclidean graph-structured data have rigorous requirements in hardware space and speed. Thus, classical hardware limitations induce bottlenecks for graph learning. Qubits carry exponentially more information than bits, and quantum algorithms solve graphs with lower time complexity [481, 482, 483]. Additionally, quantum search algorithms, such as Grover’s algorithm [484], have lower time complexity compared to their classical counterparts, making it possible to retrieve information from graph structures much faster. Thus, quantum hardware has exponentially accelerated access to data in hardware, and thus eliminates the fundamental hardware bottleneck for graph learning.

QGL harnesses these quantum advantages to intensify traditional graph learning processes. Classical methods (GNNs) are time-consuming to deal non-Euclidean graph-structured data, and thus cannot utilize hardware resources sufficiently. The use of superposition allows for the simultaneous exploration of multiple graph paths, while entanglement facilitates capturing complex, multi-node relationships with a level of intricacy that classical methods cannot easily achieve. As a result, QGL has the potential to drastically improve the efficiency and accuracy of tasks. Researchers are exploring quantum-enhanced graph neural networks (QGNNs) [485, 486] and quantum-inspired models [487, 488, 489] to adapt these concepts to existing machine learning frameworks.

Graph learning has a property of black-box. Existing models rarely provide a human-understandable explanation for tasks, and less explainable output results of models are discussed. Quantum mechanics explains the existence of our physical world. With the help of quantum theory, the explainable or interpretable problem will be resolved in graph learning. Graph learning criticized for its "black-box" nature, where models provide little insight into their decision-making processes. This lack of transparency can hinder trust and understanding, especially in critical applications like healthcare and finance. Quantum mechanics, known for its precise explanation of physical phenomena, offers a promising solution to this issue. By leveraging principles such as superposition and entanglement, QGL models can create features and relationships that are more comprehensible to humans. Quantum algorithms can highlight specific quantum states or transitions that correspond to patterns in the data offering clearer explanations for predictions.