Research Focus

Laboratory of Integrative Cancer Genomics

The long-term goal of our lab is to understand somatic evolution in tumour and normal tissues. Specifically, we are interested in comprehensively charting genomes and bridging genotype to molecular phenotypes at the single-molecule, single-cell, and spatial levels to boost biological interpretation. To this end, we develop novel approaches and apply them to clinical samples to deliver novel insights with a translational impact.

Technology development

Single-cell technologies are revolutionizing our understanding of biology. However, few of these comprehensively co-capture genomic variation. To address this gap, we recently developed Single-cell Profiling of LONG-read Genome, Epigenome, and Transcriptome (SPLONGGET; Figure 1; doi: 10.1101/2025.09.08.674950), a novel single-cell multiomics approach integrating high-throughput droplet-based barcoding using 10X Genomics with fragment-length independent nanopore sequencing. SPLONGGET produces full-length transcriptomes alongside high-quality open chromatin and whole-genome data. Libraries can also be used for targeted post-hoc genotyping. With our collaborators, we are applying SPLONGGET to profile the evolution and heterogeneity of a range of malignancies, including paediatric B-cell Acute Lymphoblastic Leukaemia, rare T-cell lymphomas, sarcoma, and thyroid tumours, as well as normal tissues.

A diagram of a test tube

AI-generated content may be incorrect.

Figure 1: Schematic of the SPLONGGET workflow. See also doi: 10.1101/2025.09.08.674950 for more info

In addition, we are integrating single-molecule multiomics with deep learning to deliver functional insights into (complex somatic) variation. At the core of this lie technologies such as Fiber-Seq, which simultaneously captures long-range DNA sequence, CpG methylation, and chromatin accessibility at single-molecule and single-base resolution, genome-wide (Stergachis et al., Science, 2020; Figure 2). To unlock the full potential of these single-molecule datasets, we are developing deep learning frameworks, including sequence-to-function models and autoencoders. These models effectively internalize regulatory logic and accurately predict molecular phenotypes for any given (variant) sequence. Critically, we integrate explainable AI (xAI) techniques to ensure biological interpretability. Together, these approaches directly link any genotype to molecular phenotypes and enable systematic decoding of how (complex) variation reshapes cellular programs through altered chromatin architecture, transcriptional and translational control. With this, we aim to deliver a new era of AI-based personalized cancer genomics.

A close-up of a graph

AI-generated content may be incorrect.

Figure 2: Deep learning for single-molecule multiomics. (A) Model prediction (dotted line) of direct RNA-seq coverage (full line) at a test locus. (B) Allele-specific model predictions (dotted, REF green; ALT black) and observed chromatin accessibility (full line) at a test heterozygous chromatin accessibility quantitative trait locus. (C) Raw Fiber-Seq data for the locus in (B), individual Fiber-Seq reads (grey horizontal bars; grouped by haplotype, right inset) confirm allele-specific chromatin accessibility. REF reads show open chromatin (6mA labelling; green), while ALT reads reflect closed chromatin. Periodic 140bp footprints correspond to nucleosomes.

Forward Translation to Improve Sarcoma/Lymphoma Diagnosis and Care

By applying these multiomic bulk long-read approaches to sarcoma and rare lymphomas, we are specifically seeking to address unmet clinical needs. Current diagnostics for these complex malignancies relies on costly and slow cascade testing involving assays which provide only partial insights. Our comprehensive approach consolidates these diverse readouts into a streamlined, cost-effective workflow that promises to boost research and transform clinical practice. Critically, our deep learning models and classifiers further integrate these multimodal datasets, enabling automated detection of complex molecular patterns and facilitating real-time clinical interpretation. 

Somatic evolution in normal tissues and its impact.

Building on our expertise in somatic evolution and single-cell omics, we are exploring how somatic mutations can shape tissue biology and human health. In a collaborative effort, we are uncovering the molecular mechanisms driving haematoinflammatory disorders. By applying (deep) error-corrected sequencing and single-cell whole-genome analysis, we are revealing somatic mosaicism as a previously unrecognized driver of late-onset autoinflammatory diseases. Our data reveal that variants in specific hematopoietic lineages can trigger systemic inflammation through driver-passenger mechanisms—a paradigm shift that positions somatic evolution as central to understanding immune dysregulation. This work will also aid diagnosis and treatment of enigmatic inflammatory conditions affecting thousands of patients worldwide.

In parallel, we are revealing how prenatal chemotherapy exposure creates lasting mutagenic imprints in children's hematopoietic systems. We know that platinum-based and ABVD regimens induce up to 4-fold increases in mutational burden in newborn hematopoietic stem cells, with characteristic signatures. Now, we will specifically look at their persistence through childhood. We use error-corrected sequencing and single-cell multiomics to track chemotherapy-induced mutations and clones across timepoints to understand long-term health implications. This research establishes a framework for understanding how early-life mutagenic exposures may shape lifelong disease risk, informing preventive medicine approaches.