Network Based Big Data Analytics

Academic Year 2026/2027 - Teacher: ALFREDO PULVIRENTI

Expected Learning Outcomes

This course introduces methods and tools for the large-scale analysis of (primarily biomedical) data using network models, with a particular focus on interactomes, co-expression networks, drug-target networks, knowledge graphs, and patient similarity networks. It integrates methodological and computational aspects of network science, graph algorithms, machine learning, Graph Neural Networks, and large-scale graph analysis with applications in Network Medicine, including the identification of disease modules, the study of comorbidities, drug repurposing, and patient stratification. Theoretical lectures are supplemented by Python labs using real biomedical data.

Knowledge and understanding:
Students will understand the structural properties of real-world networks and the appropriate null models; community detection, diffusion, and link prediction algorithms; the principles of graph representation methods (embedding, GNNs, knowledge graph embedding); the architecture of systems for computing on large-scale graphs; and the postulates and limitations of Network Medicine.

Ability to apply knowledge and understanding:
The student is able to build a reproducible pipeline that, starting from heterogeneous biomedical data, produces an integrated network, evaluates its properties, identifies modules associated with a phenotype, prioritizes candidate genes or drugs, and quantifies the uncertainty of the result.

Independent Judgment:
The student can recognize biases in biomedical resources (study bias in the interactome, incompleteness, circular annotation), identify the most common forms of leakage in the evaluation of graph models, and distinguish a topological association from causal or clinically useful evidence.

Communication skills:
The student can present the results of network analysis to a mixed audience of computer scientists and clinicians, using appropriate visualizations and providing an honest description of the limitations.

Learning skills:
The student can critically evaluate primary literature in network biology and graph machine learning and reproduce a published result using the available data and code.

Course Structure

Frontal lectures. 

Should teaching be carried out in mixed mode or remotely, it may be necessary to introduce changes with respect to previous statements, in line with the programme planned and outlined in the syllabus.
 

Learning assessment may also be carried out on line, should the conditions require it.

 

Required Prerequisites

Python programming.

Algorithms and data structures: complexity, graph traversal, minimum paths, priority queues.

Probability and statistics: random variables, hypothesis testing, correction for multiple testing.

Linear algebra: matrices, eigenvalues/eigenvectors, decompositions.

Useful but not required (reviews provided in class)

Fundamentals of machine learning (regularization, validation, metrics for imbalanced data).

Basic concepts of molecular biology and genomics: gene, transcript, protein, pathway, genetic variant.

Detailed Course Content

Fundamentals.

  • From tabular data to relational data.
  • Characterization of biomedical big data: volume, heterogeneity, sparsity, non-random incompleteness, temporal dynamics.
  • Overview of the case studies covered throughout the course: the human interactome, the diseasome, the drug–target–disease graph, and the patient similarity network.
  • Taxonomy of problems: prediction at the node, edge, subgraph, and graph levels.
  • Ethics and governance of health data as a project constraint, not an afterthought: GDPR and special categories of data, pseudonymization and re-identification from graph structures, FAIR principles.

Network science: structure and models

  • Representations and data structures: adjacency lists, formats for sparse graphs; directed, weighted, bipartite, multilayer, multiplex, and temporal graphs.
  • Local and global measures: degree distribution, clustering coefficient, assortativity, distances, and diameter.
  • Centrality: degree, betweenness, closeness, eigenvector, PageRank, k-core; biological interpretation (essentiality, bottlenecks, driver genes) and limitations.
  • Scale-free and small-world networks: Erdős–Rényi, Watts–Strogatz, Kleinberg, Barabási–Albert models, configuration model; correct estimation of the power exponent and the debate on the true ubiquity of scale-free networks.
  • Null models: degree-preserving randomization, switching algorithm, constrained graph sampling.
  • Robustness, percolation, targeted vs. random attacks; an introduction to the structural controllability of networks.
  • Motifs and graphlets; approximate triangle count.
Communities, modules, and diffusion processes

  • Modularity and its limitations; the Louvain and Leiden algorithms.

  • Spectral clustering, normalized Laplacian, conductance; the Stochastic Block Model and its degree-corrected version.
  • Overlapping and hierarchical communities (clique percolation, Infomap, dendrograms); consensus clustering and stability.
  • Validation: agreement indices, functional enrichment (GO, Reactome, KEGG) with FDR control.
  • Diffusion and propagation: random walk, random walk with restart, customized PageRank, diffusion kernel, heat kernel; algebraic formulation and iterative vs. direct solution; power iteration.
  • Information propagation as a prioritization tool: from a set of seed genes to the ranking of the entire network.

Scalability and Learning on Graphs

  • Graph data models: property graph vs. RDF; Cypher and SPARQL;
  • graph databases (Neo4j).
  • Programming models: Spark GraphX/GraphFrames, Pregel API, distributed PageRank and connected components.
  • Graph partitioning (edge-cut vs. vertex-cut), imbalance caused by heavy tails in the degree distribution.
  • Approximate and streaming algorithms: HyperLogLog sketching for cardinality estimation, Count-Min, MinHash/LSH, graph sampling (node/edge/forest fire, random walk sampling), and introduced distortions.
  • Approximate distances; approximate pattern counting.
  • Machine learning on graphs
  • Node embeddings: proximity matrix factorization, DeepWalk and node2vec, LINE; metapath2vec for heterogeneous graphs; properties and limitations (transductivity, instability, degree bias).
  • Graph Neural Networks: message-passing framework; GCN, GraphSAGE, GAT, GIN; oversmoothing, depth, normalization; mini-batching and neighborhood sampling for large graphs.
  • Heterogeneous and relational graphs: R-GCN, HGT; knowledge graph embedding: TransE, DistMult, ComplEx, RotatE; scoring, negative sampling.
  • Tasks: node classification, link prediction, graph classification, graph regression.
  • Evaluation: transductive vs. inductive splits, leakage through shared edges, temporal and entity-based splits (by drug, by disease), selection of negative examples, appropriate metrics for extreme class imbalance (AUPRC, recall@k).
  • Interpretability: GNNExplainer, PGExplainer.
  • An introduction to graph foundation models and pre-training on biomedical graphs.

Network Medicine: Interactome and disease modules

  • A concise overview of molecular biology for computer scientists: from the genome to the phenotype, pathways, and types of interactions.
  • The human interactome: sources and their nature (systematic Y2H and HuRI, co-complex/AP-MS, curated literature, pathway databases, computational predictions); study biases, incompleteness, and how results change depending on the network used.
  • The postulates of Network Medicine: the disease module hypothesis, the local pathway hypothesis, and the hypothesis of disease as a perturbed subnetwork.
  • Localization of disease genes in the interactome; significance of observed connectivity relative to degree-preserving null models.
  • Module identification: seed-and-extend methods (DIAMOnD), Steiner trees and minimum spanning trees, propagation-based approaches, diffusion-based methods.
  • Proximity and separation metrics between gene sets: network separation, nearest-neighbor/center proximity; interpretation of disease–disease relationships and comorbidities observed in clinical data.
  • The diseasome and disease–gene networks; integration with GWAS, rare variants, and tissue-specific expression.

Network Medicine: Network Pharmacology and Drug Repurposing

  • Drug–target networks, polypharmacology, and moving beyond the “one drug–one target” paradigm.
  • Network-proximity-based repurposing: quantifying the distance between the drug’s target module and the disease module; retrospective validation protocols and necessary controls.
  • Approaches based on transcriptional signatures and reversion signatures (CMap/LINCS); integration with network evidence.
  • Drug combinations: theory of complementary exposure, synergy, and antagonism from a topological perspective.
  • Prediction of adverse effects and drug interactions as link prediction on multi-relational graphs (Decagon approach); SIDER/OFFSIDES resources.
  • Current platforms and models: NeDRex, Hetionet/Rephetio, TxGNN, and foundational models for drug repurposing; 2025–2026 literature on knowledge graphs and GNNs for drug repurposing and on foundational models of network medicine.
  • From computational hypothesis to evidence: what is needed for a candidate to reach experimental validation or a clinical trial; how to critically evaluate a repurposing claim.

Network Medicine: Multi-omics, Single-Cell Analysis, and Clinical Networks

  • Co-expression networks and WGCNA: soft thresholding, modules, eigengene, correlation with traits.
  • Inference of gene regulatory networks: methods based on mutual information and trees (GENIE3/GRNBoost2), SCENIC, causal approaches; what benchmarks (BEELINE and subsequent ones) reveal about the actual reliability of these methods; developments for 2025–2026 based on foundational single-cell models and graph contrastive learning.
  • Single-cell: cell-specific networks, cell–cell communication (ligand–receptor), networks on spatial data.
  • Multi-omic integration as a problem in multilayer networks: Similarity Network Fusion, latent factor approaches, alignment between layers.
  • Patient similarity networks and stratification: similarity construction, clustering, association with clinical outcomes; validation on an independent cohort.
  • Graphs derived from real-world clinical data: EHRs as patient–diagnosis–drug–procedure graphs (MIMIC-IV), phenotypic knowledge graphs (HPO, UMLS), population-level comorbidity networks.
  • LLMs and knowledge graphs: GraphRAG and agents on biomedical knowledge graphs for clinical question answering and drug safety verification; strengths and risks.
  • Overview of microbiome networks and host–pathogen interaction networks.

Reproducibility, Privacy, and Critical Evaluation

  • Reproducibility of a graph-based pipeline: data and resource versioning (STRING/DrugBank releases affect results), seeds, environments, and experiment tracking.
  • Privacy: re-identification based on graph structures, graph anonymization and its weaknesses, federated and privacy-preserving graph learning in a clinical context.
  • Bias and equity: who is represented in cohorts and interaction databases, and how bias propagates in a ranking.
  • Regulation: GDPR and the European Health Data Space; an overview of the AI Act for clinical decision support systems.
  • Critical Reading Session: collective analysis of two articles, one methodological and one applied.

Textbook Information

  • A.-L. Barabási, Network Science, Cambridge University Press, 2016.  disponibile gratuitamente online.
  • J. Loscalzo, A.-L. Barabási, E. K. Silverman (eds.), Network Medicine: Complex Systems in Human Disease and Therapeutics, Harvard University Press, 2017
  • W. L. Hamilton, Graph Representation Learning, Morgan & Claypool, 2020 disponibile gratuitamente online. 
  • J. Leskovec, A. Rajaraman, J. D. Ullman, Mining of Massive Datasets, 3ª ed., Cambridge University Press  disponibile gratuitamente online. Capitoli su sketching, LSH e link analysis.
  • M. Newman, Networks, 2ª ed., Oxford University Press, 2018. Testo di Approfondimento.

Course Planning

 SubjectsText References
1A.-L. Barabási, Network Science, Cambridge University Press, 2016 
2A.-L. Barabási, Network Science, Cambridge University Press, 2016
3J. Leskovec, A. Rajaraman, J. D. Ullman, Mining of Massive Datasets, 3ª ed., Cambridge University Press
4·W. L. Hamilton, Graph Representation Learning, Morgan & Claypool, 2020
5J. Loscalzo, A.-L. Barabási, E. K. Silverman (eds.), Network Medicine: Complex Systems in Human Disease and Therapeutics, Harvard University Press, 2017
6J. Loscalzo, A.-L. Barabási, E. K. Silverman (eds.), Network Medicine: Complex Systems in Human Disease and Therapeutics, Harvard University Press, 2017
7J. Loscalzo, A.-L. Barabási, E. K. Silverman (eds.), Network Medicine: Complex Systems in Human Disease and Therapeutics, Harvard University Press, 2017