Emerging Topics
As graph learning matures, a range of emerging topics are expanding its scope and deepening its integration with cutting-edge developments in AI. Moving beyond traditional tasks like node classification and link prediction, researchers are now exploring more versatile, scalable, and intelligent approaches that push the boundaries of what graph-based models can achieve [3, 63]. These emerging directions not only address new technical challenges but also reflect the broader AI agenda of building more generalizable, distributed, and knowledge-driven systems.
This section discusses several key frontiers in graph learning. We begin with graph foundation models, which aim to develop large-scale, pretrained graph encoders capable of generalizing across diverse tasks and domains. We then explore graph reinforcement learning, where graph structures are used to model complex environments and interactions for sequential decision-making. We also discuss federated graph learning, which enables collaborative model training across decentralized graph data sources while preserving privacy. Additionally, we cover advances in learning on knowledge graphs (KGs) and knowledge-infused graph learning, which leverage symbolic information to enhance reasoning and generalization. Finally, we examine the emerging field of quantum graph learning, which explores quantum computing techniques for more efficient and expressive graph representations. These topics represent the next wave of innovation, positioning graph learning as a central pillar in the development of intelligent, explainable, and scalable AI systems.
Graph Foundation Models
Graph Foundation Models (GFMs) are large-scale pre-trained models specifically designed for processing graph-structured data. These models are generally pre-trained on vast amounts of data across various domains [395, 396]. Besides, GFMs can adapt to varying demands with minimal fine-tuning or even in few-shot/zero-shot scenarios. There exist many open issues to be addressed, and we discuss them from three distinct perspectives: pre-training, architecture design, and adaptation strategies.
Pre-training Strategies
Pre-training involves training a model on large-scale data in advance. The goal is to create a model with transferable knowledge that can be easily fine-tuned for specific downstream tasks, improving performance even with limited target task data. Current GFMs mainly adopt supervised pre-training and self-supervised pre-training.
Supervised pre-training: In supervised pre-training, models are trained on labeled datasets, often incorporating multi-task learning to optimize performance across various tasks. This approach allows the model to learn representations that are specific to particular tasks, enhancing its effectiveness in downstream applications.
Self-supervised pre-training: Self-supervised learning has become a dominant strategy in graph foundation models, as it allows for training without labeled data. This approach includes techniques like contrastive learning, where the model learns by distinguishing between positive and negative samples, and generative models, where the model learns to generate or reconstruct graph structures.
While pre-training has shown promise, several challenges remain that hinder its broader applicability. Below, we highlight key areas that require further exploration and innovation.
Task Design. Current pre-training tasks tend to be tailored to specific domains, which limits the ability of models to perform effectively on diverse types of graph data. The development of more universal pre-training tasks that enable models to learn transferable features applicable across various graph domains remains a critical area of research.
Data Efficiency. Pre-training often demands large, high-quality datasets, which are not always accessible. Enhancing the data efficiency of pre-training models is crucial, particularly in data-scarce environments. This challenge involves optimizing the use of available data, including exploring large-scale datasets, leveraging cross-modal data, and implementing robust data augmentation techniques. Additionally, assessing and ensuring the quality of pre-training data is essential to maximize model performance.
Graph Complexity. The complex nature of graph structures poses troubles for pre-training. Designing effective pre-training methods for various types of graphs, such as dynamic, homogeneous, heterogeneous graphs, or even hypergraphs, is an area that requires further exploration.
Architecture Design
GFMs have evolved through a variety of architectural paradigms, each designed to capture different aspects of graph data. From classic message-passing frameworks to more recent innovations like Graph Transformers and LLMs, these architectures have significantly advanced. The existing architectures for GFMs are as follows.
Message-Passing Models: Message-passing models are the most common designs in graph models. These architectures capture local graph structures by iteratively passing and aggregating information between nodes. Classic models in this category include GCN [163], GAT [35], and GraphSAGE [397]. Despite the success, message-passing struggles with expressivity, often failing to capture long-range dependencies due to over-smoothing and limited receptive fields.
Graph Transformer: Inspired by the success of transformers in natural language processing, Graph Transformers (GTs) [398, 399, 400] have been extended to graph-structured data. These models apply global attention mechanisms to capture long-range dependencies in graphs, making them well-suited for handling complex relations. Their token-based graph representations are well-suited for pretraining tasks, although scalability and structural priors remain open challenges.
Generative Models: Generative graph models, such as VAEs [401] and GAN-based methods [402], learn latent graph distributions to support tasks like generation, completion, and anomaly detection. Besides, recent diffusion and autoregressive approaches [285] enable high-quality graph synthesis, capturing structural and semantic priors that enhance the versatility of Graph Foundation Models through generative pretraining.
Large Language Models: Recently, LLMs have been adapted for graph data by encoding graphs into tokens [403, 404]. This approach leverages the power of LLMs to process graphs, opening up new avenues for graph learning. While promising, this line of work raises open questions on how to design good tokenization schemes, inject structural priors, and mitigate hallucination. As GFMs evolve, the convergence of LLM-based approaches and GTs may play a central role in more areas.
However, several challenges in architectural design continue to limit their effectiveness. Below, we outline key areas that require further innovation.
Dynamic Graph Modeling. Real-world graphs are often dynamic, with nodes and edges evolving over time. Designing GFMs capable of handling dynamic graphs, especially in scenarios where the graph structure continuously updates, remains an area requiring further exploration.
Scalable Architecture Design. Maintaining GFMs’ effectiveness and efficiency becomes increasingly challenging as graph data scales to millions of nodes and edges. Future research can focus on developing scalable architectures that perform well on large-scale graphs without compromising on computational efficiency.
Balancing Global and Local Information. Current graph learning architectures often focus on either local or global structural information, rather than integrating both. Developing methods that effectively combine global and local information in GFMs is a promising direction for performance gain.
Handling Complex Graph Structures: Most existing graph models are designed for simpler graph structures, such as homogeneous graphs. However, many real-world graphs are heterogeneous or hypergraphs, which present additional challenges. Developing GFMs to process these complex structures remains challenging.
Tuning, Adaptation and Compression
GFMs for specific downstream applications requires a combination of adaptation strategies. This means that pre-trained models can be effectively fine-tuned for targeted tasks and efficiently deployed in various environments. The existing transfer strategies for GFMs are as follows:
Fine-Tuning and Adaptation: Fine-tuning is a common approach for applying pre-trained GFM to specific downstream tasks. By further training the model parameters on a small amount of task-specific data, fine-tuning enhances the model’s performance on the target task. This can be achieved through full-parameter fine-tuning or by freezing most of the model’s layers and only fine-tuning a few selected layers.
Model Compression: In resource-constrained environments, model compression techniques are crucial for deploying GFMs. Compression methods include pruning, quantization, and knowledge distillation. These techniques aim to reduce the model’s parameters or complexity, thereby lowering computational and storage requirements while retaining as much performance as possible.
Despite the advances in tuning, adaptation, and compression strategies for GFMs, several open issues persist that limit their effectiveness in applications.
Generalizable Fine-Tuning Methods. While fine-tuning is effective for specific tasks, developing methods that generalize across a wide range of downstream applications for GFMs remains challenging. The key issue is finding the right balance between retaining pretrained knowledge and adapting quickly to the unique requirements of new tasks.
Automated Fine-Tuning. Future research should explore automated tools like AutoML [405] to streamline the GFM fine-tuning process. Automation has the potential to reduce human intervention, enhance efficiency, and produce more optimal tuning outcomes, making fine-tuning more accessible and effective.
Maintaining Performance After Compression. Model compression techniques, such as pruning and quantization, often lead to performance degradation, particularly in complex graph tasks. A critical area of research is developing compression strategies on GFMs that minimize this loss of performance while ensuring that compressed models remain effective in practical applications.
Few- and Zero-Shot Scenatios. In many real-world scenarios, labeled data is limited. Effectively fine-tuning pre-trained models in few-shot settings is a significant challenge. Exploring how pre-trained models can be leveraged for zero-shot or few-shot learning, particularly in data-scarce environments, will be an important focus for future GFM tuning strategies.
GFMs are emerging as a unified framework for graph-based learning, combining large-scale pretraining, GTs, and flexible fine-tuning to support diverse applications from social networks to scientific discovery. As the field matures, GFMs may become core components of graph learning and become a unified tool across graph tasks.
Graph Reinforcement Learning
Deep Reinforcement Learning (DRL) is a subfield of machine learning that integrates deep learning with reinforcement learning [406]. DRL training agent to make decisions by learning optimal policy through interactions with the environment, using deep neural networks to handle high-dimensional state spaces and approximate complex value functions [407]. The sequential decision-making process can be modeled using a Markov Decision Process (MDP) or an extended MDP, such as a partially observable Markov decision process (POMDP) or Semi-MDP. Specifically, an MDP can be defined as \(<S,A,T,R>\), where \(S\) is state space, \(A\) is the action space, \(T\) is state transition probability, and \(R\) is the reward value obtained after actor \(a\) is performed. This approach enables the agent to maximize cumulative rewards over time in sequential decision-making tasks.
Graph Reinforcement Learning (GRL) is a novel method that combines DRL with graph representation, enabling the solution of decision-making problems based on graph data [408]. In DRL problems, graph-structured data provides rich relational information, especially in multi-agent scenarios, which is unavailable in non-graph-structured states. Additionally, incorporating DRL into graph learning introduces a new training approach that differs from existing graph learning tasks, which often require substantial expert knowledge. DRL allows for learning without prior knowledge. We will introduce GRL from two aspects: graph-enhanced reinforcement learning and reinforcement learning enhanced graph learning.
Graph Enhanced DRL
In multi-agent DRL, a group of agents cooperate or compete to achieve a common goal, where the key is to understand the mutual interplay between agents. This architecture is widely used in various tasks, such as traffic signal control [409], vehicle control systems [410], and network routing [411]. The communication between agents provides information about the states and environments of other agents. Various attention-based methods have been proposed to address communication issues in multi-agent systems. In G2ANet [412], the authors used a complete graph to model the relationships between agents in large-scale multi-agent scenarios and designed a two-stage attention model as the communication model. Similarly, in Graphocomm [413], the authors applied a relational graph module to model explicit and implicit relationships through explicit and implicit relation layers. Specifically, explicit relationships are built from prior knowledge, while implicit relationships are learned through interactions among agents. In SRI-AC [414], the authors deployed a Variational Autoencoder (VAE) model to predict interactions and learn a state representation from observational data.
In multi-task DRL, states and actions can vary significantly across tasks. However, this challenge can be effectively addressed by leveraging the generalization capabilities of GNNs, which facilitate efficient multi-source policy transfer learning in settings with state-action mismatches. In TURRET [415], GNNs are used to learn the intrinsic properties of agents, creating a unified state embedding space for all tasks. This approach enables TURRET to achieve more efficient transfer and stronger generalization across tasks, making it easily compatible with existing DRL algorithms.
DRL Enhanced Graph Learning
Most graph representation methods either use global pooling techniques to generate a single global representation vector, which overlooks the semantics of substructures, or rely on expert knowledge to extract local substructures [416]. This often results in poor interpretability and limited generalization ability [417]. By utilizing reinforcement learning, it is possible to capture the most significant subgraphs without the need for expert knowledge [418], offering a more flexible and adaptive approach to graph representation. By defining the update value \(k\) within the MDP framework and using Q-learning to learn the optimal policy, where \(k\) serves as the pooling ratio in top-k sampling. Similarly, the identification of an optimal subset can be formulated as a combinatorial optimization problem, where the goal is to select the best combination of elements from a larger set. Deep Q-learning can be applied to solve this problem by learning the value of each possible action and selecting the combination that maximizes the overall reward [415].
Optimizing the GNN framework and hyperparameters can significantly enhance learning capabilities. However, modifying these components requires extensive domain knowledge. Utilizing Neural Architecture Search (NAS) allows for the selection of suitable GNN models and their corresponding hyperparameters tailored to specific tasks. Due to the highly diverse nature of graph classification applications, designing data-specific pooling methods based on human expertise presents a significant challenge [419]. To address this, PAS [420] introduces a unified pooling framework that simplifies the design of the search space. Two variants, PAS-G and PAS-NE, are also proposed, each offering distinct approaches to improve the effectiveness of pooling in different scenarios.
Federated Graph Learning
Although graph learning techniques have significantly progressed across various domains in recent years, most existing graph neural networks still rely on centralized storage of large-scale graph data for training. However, growing concerns over data security and user privacy make this requirement impractical in real-world scenarios [421]. In many practical applications, graph data is often distributed across multiple data holders, and due to privacy concerns and regulatory constraints, no party can directly access the data of others. This data distribution across different devices and owners, coupled with privacy concerns, makes it increasingly challenging to learn a global model while preserving data privacy [422].
For instance, in the financial sector, a third-party company may need to train a graph learning model for multiple financial institutions to detect potential financial crimes and fraudulent activities. Each financial institution holds its local customer dataset, including demographic information and transaction records between customers. Based on this data, each institution forms a customer graph, where the edges represent transaction relationships. Due to strict privacy policies and industry competition, these local customer datasets cannot be directly shared with third-party companies or with other institutions. At the same time, there may be interactions between different financial institutions, representing structural information across institutions. The main challenge for the third-party company lies in training a graph learning model for financial crime detection based on both the local customer graphs and the structural information between institutions, without directly accessing the local data of each institution [423].
To address this challenge, federated learning (FL) has emerged as a promising distributed learning framework that tackles the issues of data isolation and privacy preservation [424]. FL allows participants to collaboratively train a global model without sharing their local data, and it has been widely applied to Euclidean data tasks, such as image classification, by aggregating model updates from multiple clients [425, 426]. However, the standard FL framework struggles to handle the complex relationships inherent in graph-structured data, limiting its direct application to graph learning scenarios [427].
In response, Federated Graph Learning (FGL) has emerged as a promising solution [428, 429, 430]. FGL combines the strengths of federated learning and graph neural networks, enabling multiple data holders to collaboratively train a graph learning model without sharing their local graph data. This method is particularly suitable to distributed graph data scenarios, such as analyzing customer transaction data across multiple financial institutions. Each institution holds a local customer graph, and the interactions between institutions form cross-institutional structural information. By integrating local graphs with inter-institutional structural information while ensuring data privacy, FGL can construct a global graph model, significantly enhancing the detection of financial crimes and other related tasks.
FGL can be further categorized into three types based on how graph structure information is distributed among participants: graph federated learning, subgraph federated learning, and graph structured federated learning. By incorporating structural information from graph data, FGL not only enhances distributed graph learning but also boosts performance in cross-institutional and cross-domain collaboration, while maintaining privacy protection [423, 421].
Graph Federated Learning
Graph federated learning is a natural extension of standard FL, where each participant locally holds graph structured data, and the global model is trained to perform graph level tasks. A typical application of graph federated learning is in AI4Science, such as drug discovery and molecular property prediction. In these cases, GNNs are used to study the graph structure of molecules, where molecules are represented as graphs, with atoms as nodes and chemical bonds as edges [429]. Each bio-pharmaceutical company has its private dataset, containing molecular nodes and their corresponding properties. Historically, commercial competition and data privacy concerns have hindered collaboration between these companies. However, with the graph federated learning framework, it is now possible to jointly train molecular structure models, enabling clients to collaborate on building robust models while preserving data privacy.
Formally, each client \(k\) owns its private local graph data \(\mathcal{D}_k=\left\lbrace \mathcal{G}_1, \mathcal{G}_2, \cdots \right\rbrace\), where each \(\mathcal{G}_i=(\mathcal{V}, \mathcal{E})\) is a graph comprising a set of nodes \(\mathcal{V}\) and a set of edges \(\mathcal{E}\). Each client collaborates with others to train a robust global graph model \(\mathcal{F}\) based on its local dataset \(\mathcal{D}_k\), while ensuring that \(\mathcal{D}_k\) remains local and private. A significant challenge in this setting is the heterogeneity of the data across clients, which may differ significantly in terms of node features and graph structures [431]. This data heterogeneity can lead to severe model divergence during the federated process, thereby degrading the performance of the global model. To address this, various methods have been proposed to either train a single global model with better generalization or develop personalized models that can adapt to the unique characteristics of each client’s data [428, 432].
Subgraph Federated Learning
In subgraph federated learning, each participant’s local data is treated as a subgraph of a larger, more comprehensive global graph [433]. Participants focus on using the nodes and edges within their subgraphs as training samples. Without directly sharing data, the participants collaboratively train a global graph model. Subgraph federated learning has broad application prospects in fields like social network analysis, recommendation systems, and financial risk control, where graph data plays a central role. It offers a solution to handle large-scale, distributed graph data while safeguarding privacy and improving model performance.
However, subgraph federated learning faces several challenges [434]. First, each client only holds a subgraph of the original global graph, and due to privacy concerns and the isolated storage of data, each node can only access the information of neighboring nodes within its subgraph, lacking access to information from other clients. This missing cross-client information leads to biased node embeddings on each client, thereby reducing the performance of the graph learning model [435]. A key challenge is to reconstruct the missing cross-client information to accurately compute the node embeddings. Another challenge is that the same node may belong to multiple clients. In such cases, during collaborative training, the embeddings of overlapping nodes from different clients may originate from different embedding spaces. The core idea to address this is to learn global node embeddings based on the local embeddings of each client, and an overlapping instance alignment technique has been introduced to ensure the consistency of node embeddings across clients [436].
Graph Structured Federated Learning
In real-world scenarios, clients are often interconnected. For instance, in traffic flow prediction tasks, sensor devices distributed across different geographical locations may have connections, and these links often contain rich information, such as spatial dependencies [422].
In graph structured federated learning, each client is treated as a node, and together, all the clients form a graph \(\mathcal{G} = (\mathcal{V}, \mathcal{E})\), where \(\mathcal{V}\) represents the set of clients and \(\mathcal{E}\) represents the edges between them. This approach takes into account the topological structure between clients and uses GNNs to aggregate models based on these topological relationships [423]. It is important to note that the data on the client side is not necessarily graph-structured.
Typically, clients upload their local model parameters or local embeddings to the server, just as in standard FL. The server then aggregates the uploaded models using graph-based algorithms, taking into account the client graph \(\mathcal{G}\), and finally distributes the updated model parameters back to the clients [437, 438, 439].
Learning on Knowledge Graphs
KGs represent structural relations between entities, storing human knowledge facts in intuitive graph structures and being towards human-level intelligence [440, 441]. Learning on KGs refers to combining the structured data represented in knowledge graphs with graph learning techniques [442, 443, 7]. Although KGs have made significant success in various real-world fields, how to use these semantically structured data to assist graph learning still faces challenges. Firstly, KGs are often incomplete with missing links between entities, resulting in a degradation of the quality of the knowledge graph and impacting model performance. Thus, it poses challenges in the knowledge extraction process to construct high-quality KGs. Secondly, KGs grow exponentially due to the quick updating of information in the real world. Thus, it poses challenges in proposing efficient representation methods to handle the scalability of KGs.
Knowledge Extraction
Knowledge extraction (or knowledge acquisition) tasks can be divided into three categories: knowledge graph completion (KGC), relation extraction, and entity discovery [441]. KGC aims to predict missing links in existing KGs, and embedding-based ranking, relation path reasoning, rule-based reasoning, and meta relational learning are main categories for KGC. Compared to KGC, relation extraction and entity discovery are for discovering new knowledge (e.g., relations and entities) from the text. Relation extraction models utilize attention mechanisms, GCNs, adversarial training, reinforcement learning, deep residual learning, and transfer learning. KGC, relation extraction, and entity discovery are three separate tasks for knowledge extraction. In order to simplify development and be convenient for knowledge extraction, proposing a unified framework is one of the challenges in this field. For example, [444] proposed a joint learning framework with mutual attention for data fusion between knowledge graphs and text, which simultaneously solves KGC and relation extraction from text. Additionally, knowledge extraction also face other challenges in KGC, relation extraction, and entity discovery as follows.
Challenges of KGC: As knowledge graphs are by nature incomplete, discovering new triples in knowledge graphs is the prime mission. The first challenge is that most existing KGC methods focus on one-hop reasoning but fail to capture multi-step relationships for complex reasoning. For example, embedding-based ranking methods usually use one-hop neighbour information to predict missing entities or relations [445, 446, 447]. However, these methods cannot consider the influence of multi-hop relationships and model complex relation paths in knowledge graphs. To solve this problem, relation path reasoning methods and rule-based methods attempt to leverage path information over the graph structure [448, 449]. Another two challenges of KGC are that: 1) the long-tail phenomena exist in the relations, 2) and unseen triples usually emerge because of knowledge being dynamic in the real-world scenario. Some efforts have been made to target these two challenges, and meta-relational learning is a solution to them. The principle behind meta-relational learning is based on meta-learning and local graph structures. It encodes one-hop neighbors to capture the structural information with R-GCN and then takes the structural entity embedding for multi-step matching guided by meta-learning [450].
Challenges of relation extraction: Relation extraction usually aims to extract unknown relational facts from plain text and add them into knowledge graphs. Traditional methods highly depend on feature engineering with prior knowledge to explore the inner correlation between features [451]. However, the main challenge of relation extraction is how to learn richer representations as possible. Thus, using deep neural networks like GCN can leverage relational knowledge in graphs to effectively extract relation [452, 453, 454]. For example, [455] applied GCN for relation embedding in knowledge graphs for sentence-based relation extraction.
Challenges of entity discovery: Entity recognition is the first step of entity discovery, which tags entities in text. Its technique has evolved from hand-craft to applying neural architectures such as LSTM-CNN methods, attention-based methods, K-BERT and so on [456]. Entity typing is the step after entity recognition. However, arranging a proper entity typing is not a simple task. Because entity typing not only includes coarse types, but also contains fine-grained types. And the latter is typically regarded as multi-class and multi-label classification. Thus, label noise is inevitably introduced to the entity typing process. The challenge of entity discovery is reducing label noise in entity typing. Embedding-based approaches have been proposed to tackle this problem. JOIE and ConnectE are embedding-based methods, which explore local typing and global triple knowledge structure information to enhance joint embedding learning [457, 458].
Knowledge Representation
Knowledge representation learning (KGL) is a critical research issue of KGs, which paves the way for many downstream applications [459, 460]. Existing KGL focuses on modeling the semantic interaction of facts and utilizing external information. GNNs are introduced to learn the connectivity structure in KGs under an encoder-decoder framework. R-GCN proposes a relation-specific transformation to model the directed nature of knowledge graphs, and GCN acts as a graph encoder [461]. SACN introduces weighted GCN to define the strength of two adjacent nodes with the same relation type, capturing the structural information within KGs [462].
Although embedding representations can capture relationships between knowledge elements, it is challenging to infer explicit logical relationships directly from these vector representations. Designing knowledge representation models that can provide interpretable reasoning while maintaining efficient computational performance remains an important research direction. Additionally, a unified understanding of knowledge representation is less explored. The field of knowledge representation still faces numerous challenges that hinder its broader application. Below are some key areas that require further exploration:
Dynamic Knowledge Representation: Knowledge in the real world is constantly changing, and the temporal changes in knowledge, along with the emergence of new knowledge, require knowledge representations to have the capability to handle dynamic changes. Designing dynamic knowledge graphs or knowledge representation models that can update over time and perform real-time reasoning is a challenging problem that needs to be addressed.
Cross-Domain Knowledge Representation: Many knowledge representation methods are often constructed for specific domains, making it difficult to represent and reason across domains. The concepts and relationships in different domains may vary significantly, and how to unify and integrate cross-domain knowledge remains a challenge.
Knowledge-infused Graph Learning
Graph learning has achieved significant success in knowledge graph tasks, such as knowledge extraction and knowledge representation [448, 449]. In the domain of graph learning, models are operated under the assumption that all data is of high quality, and graph learning techniques are extended to enhance the performance of downstream tasks. However, challenges related to low-quality data and poor generalization when applying these methods are encountered, leading to subpar model performance on instances not seen during training [463]. To address these limitations, the concept of knowledge-infused graph learning has been introduced. This approach integrates external knowledge into various components of the graph machine learning pipeline, aiming for more accurate results. Knowledge can be categorized into formal scientific knowledge (e.g., established laws or theories that govern the behaviour of target variables) and informal experimental knowledge (e.g., well-known facts derived from long-term observations or human reasoning).
Compared to traditional graph machine learning techniques, incorporating human knowledge provides several advantages: (i) it allows for the successful integration of vast amounts of information, (ii) it enhances reliability in results, and (iii) it facilitates meaningful inferences, which can assist individuals lacking domain expertise. Knowledge-infused graph learning methods can be seen as employing knowledge in two ways: as prompts or as augmentation. “Knowledge as prompts” refers to using external knowledge to guide model learning and decision-making processes. Conversely, “knowledge as augmentation” involves analyzing a graph \(\mathcal{G=(V,E)}\) in conjunction with a relevant knowledge dataset \(\mathcal{D}\). The primary objectives of this approach are twofold: to learn a function that generates vector representations for graph elements, \(f:(\mathcal{G,D}) \rightarrow \textbf{Z}\), capturing both the structure and semantics of the graph alongside the knowledge stored in the knowledge database, and to retrieve direct and precise answers (or explanations) from the dataset \(\mathcal{D}\) to address user queries.
Despite the considerable progress of knowledge-infused graph learning in downstream applications such as large language models pre-training and drug discovery, challenges remain, including knowledge database composition and effective knowledge integration. These challenges will be discussed in the following sections.
Knowledge as Prompt
The prompt-based fine-tuning method originates from natural language processing and has been widely used to help pre-trained language models adapt to various downstream tasks [464]. Knowledge, as the initial guidance or prompt in the graph learning process, aims to assist the model in understanding the structure and relationships within the data, thereby generating more effective learning outcomes. There are two main types of prompt methods:
Discrete Graph Prompt: In the discrete graph prompt method, prompts are explicitly marked or modified within the graph structure through specific nodes, edges, or subgraphs. These prompts are typically predefined or interpretable, directly embedded into the graph’s structure, and the prompt information is discrete.
Continuous Graph Prompt: In the continuous graph prompt method, prompts are embedded into the graph’s features or edge weights in the form of vectors or continuous values. The prompt information does not directly change the graph structure but subtly influences the model’s behavior by adjusting the graph representation or the parameters of the graph neural network.
Although knowledge as a prompt has achieved some success in the field of graph learning, several challenges still remain:
Limited Fine-Tuning Methods: In the domain of graph neural networks, existing research on graph prompts is still limited, and there is a lack of a general approach to cater to different downstream tasks. Some approaches have leveraged prompt-based fine-tuning of pre-trained models via edge prediction. [465, 466].
Generalization: Some current research methods design specific prompt functions for pre-training tasks, but these approaches are often tailored to particular pre-trained GNN models and may not perform well on others. Future research should explore a more general prompt-based fine-tuning method such as GPF [467].
Knowledge as Augmentation
Typical graph analysis applications include node role identification, personalised recommendation, social healthcare, academic network analysis, graph classification and epidemic trend study [468, 469]. Regarding external knowledge as data augmentation helps incorporate with graph learning for more accurate performance. For example, knowledge-augmented graph learning methods have achieved more precise drug discovery with limited training data [463]. To illustrate it well, we will discuss challenges from these four directions: incorporating knowledge in preprocessing, pretraining, training, and interpretability.
Incorporating Knowledge in Preprocessing: Taking drug discovery as an example, GNN models employ a limited number of message-passing steps \(L\), that is less than the diameter of the biomedical graph \(diam(G)\). It means that entities that are more than \(L\) hops apart will not receive messages about each other, which hampers the prediction of biomedical properties that are heavily dependent on global features. Thus, how to design new features from external domain knowledge for input data \(\mathcal{X}\) is a challenge. [470] developed a molecular feature-generation model to map molecular and fingerprint features. In addition, some studies propose capturing additional domain knowledge from external biomedical knowledge datasets to enrich the semantics of entities’ embeddings for precise drug-target interaction prediction [471, 472].
Incorporating Knowledge in Pretraining: One of the main challenges in the use of graph learning techniques is the construction of an appropriate objective function for model training, especially in the presence of limited supervision labels in biomedical datasets. One notable example of this approach is the PEMP mode, which identifies a group of useful knowledge to enhance the capability of GNN models in capturing relevant information through pretraining [473]. Additionally, knowledge from large-scale knowledge databases can be exploited as well.
Incorporating Knowledge in Training The sparsity of training data and complexity of process modeling are significant limitations. To address these challenges, leveraging external knowledge to guide the message-passing process is a promising solution. This approach involves leveraging domain knowledge to adjust the internal processes of GNN models [474]. However, the optimal method for adaptively utilising auxiliary knowledge by considering the uncertainty remains an open research question.
Incorporating Knowledge in Interpretability: The ability to understand and explain how a model works is a crucial step towards Trustworthy Artificial Intelligence (TAI). This highlights the need to improve interpretability and trustworthiness from the user’s perspective. By utilising machine-readable domain knowledge represented in KGs, relevant knowledge can be highlighted for downstream applications. [475] extracted local subgraphs in an external KG to extract biomedical knowledge based on a self-attention mechanism. Nevertheless, there remains a vast scope for further research in the area of advanced interpretability, with the aim of providing more holistic and adaptable explanations, such as advanced reasoning and question-answering capabilities [476, 477].
Quantum Graph Learning
Quantum mechanics is increasingly being explored to enhance graph learning techniques, giving rise to the emerging field of Quantum Graph Learning (QGL) [478]. Traditional graph learning methods, which focus on understanding and predicting relationships in graph-structured data, face challenges like scalability and capturing intricate node interactions. By integrating principles of quantum mechanics (e.g., superposition and entanglement), QGL aims to revolutionize this process. Quantum algorithms have the potential to process complex graph structures more efficiently, enabling faster computations and deeper insights into the connectivity patterns. As QGL is still in its infancy, researchers are actively exploring quantum models and algorithms that can be adapted to current graph learning frameworks.
Message-passing GNNs face significant limitations when it comes to efficiently capturing long-range dependencies between nodes. As the distance between nodes increases, the efficiency of information transfer diminishes, leading to issues such as oversmoothing, where node representations become indistinguishable. When representing data in quantum states, the quantum coherence or entanglement transfers long-distance information with a faster time. Quantum graph learning provides a potential solution by leveraging quantum phenomena like quantum coherence and entanglement [479, 480]. Quantum coherence allows quantum systems to exist in multiple states simultaneously. In this framework, entanglement plays a pivotal role by facilitating the transfer of information between distant nodes through non-local effects, allowing long-range dependencies to be captured with much greater precision and efficiency.
Non-Euclidean graph-structured data have rigorous requirements in hardware space and speed. Thus, classical hardware limitations induce bottlenecks for graph learning. Qubits carry exponentially more information than bits, and quantum algorithms solve graphs with lower time complexity [481, 482, 483]. Additionally, quantum search algorithms, such as Grover’s algorithm [484], have lower time complexity compared to their classical counterparts, making it possible to retrieve information from graph structures much faster. Thus, quantum hardware has exponentially accelerated access to data in hardware, and thus eliminates the fundamental hardware bottleneck for graph learning.
QGL harnesses these quantum advantages to intensify traditional graph learning processes. Classical methods (GNNs) are time-consuming to deal non-Euclidean graph-structured data, and thus cannot utilize hardware resources sufficiently. The use of superposition allows for the simultaneous exploration of multiple graph paths, while entanglement facilitates capturing complex, multi-node relationships with a level of intricacy that classical methods cannot easily achieve. As a result, QGL has the potential to drastically improve the efficiency and accuracy of tasks. Researchers are exploring quantum-enhanced graph neural networks (QGNNs) [485, 486] and quantum-inspired models [487, 488, 489] to adapt these concepts to existing machine learning frameworks.
Graph learning has a property of black-box. Existing models rarely provide a human-understandable explanation for tasks, and less explainable output results of models are discussed. Quantum mechanics explains the existence of our physical world. With the help of quantum theory, the explainable or interpretable problem will be resolved in graph learning. Graph learning criticized for its "black-box" nature, where models provide little insight into their decision-making processes. This lack of transparency can hinder trust and understanding, especially in critical applications like healthcare and finance. Quantum mechanics, known for its precise explanation of physical phenomena, offers a promising solution to this issue. By leveraging principles such as superposition and entanglement, QGL models can create features and relationships that are more comprehensible to humans. Quantum algorithms can highlight specific quantum states or transitions that correspond to patterns in the data offering clearer explanations for predictions.