TOP
Search the Dagstuhl Website
Looking for information on the websites of the individual seminars? - Then please:
Not found what you are looking for? - Some of our services have separate websites, each with its own search option. Please check the following list:
Schloss Dagstuhl - LZI - Logo
Schloss Dagstuhl Services
Seminars
Within this website:
External resources:
  • DOOR (for registering your stay at Dagstuhl)
  • DOSA (for proposing future Dagstuhl Seminars or Dagstuhl Perspectives Workshops)
Publishing
Within this website:
External resources:
dblp
Within this website:
External resources:
  • the dblp Computer Science Bibliography


Dagstuhl Seminar 25351

Computational Proteomics

( Aug 24 – Aug 29, 2025 )

(Click in the middle of the image to enlarge)

Permalink
Please use the following short url to reference this page: https://www.dagstuhl.de/25351

Organizers

Contact

Shared Documents


Press Room

Summary

In 2025 the Dagstuhl Seminar “Computational Proteomics” (25351), part of a series of Dagstuhl Seminars with the same name, brought together researchers in proteomics, glycomics, computational biology, translational biomarker research, mass spectrometry, statistics, and machine learning for a week of intense discussions and collaboration. Building on the Dagstuhl Seminar “Computational Proteomics” (23301) in 2023, we extended the agenda in four directions that reflect where our field is now pushing hardest:

Translational Proteomics

The translational proteomics group defined translational proteomics as a continuum from discovery to clinical implementation, spanning basic model systems (cell lines, mouse), human biospecimens, clinical decision support, and ultimately population health. The group emphasized that translation is not just “applying proteomics in the clinic,” but instead structuring the entire value chain: standardized sample handling, acquisition, annotation, processing, interpretation, and delivery of actionable outputs (e.g. patient stratification, tumor board support). Major barriers identified include a lack of interoperable and wellannotated datasets, underpowered cohorts (especially in rare diseases), weak incentives for repetitive but clinically necessary assays, and difficulty converting molecular readouts into clinical recommendations. The group proposed ENIGMA, a staged, global-scale effort to generate and harmonize >100,000 proteomics datasets, starting in controlled mouse models and extending to human samples, as an AI-ready foundation for translation.

Machine Learning (in Proteomics and Glycomics)

The machine learning (ML) group concluded that the current culture of “incremental performance improvements” is unsustainable and often scientifically marginal. Instead, the group argued for community standards around software quality, reproducibility, interpretability, and dataset (ML/AI) readiness. Discussions focused on updating and extending existing recommendations (e.g. DOME, FAIR4RS) to address maintainability, testability and bias, and on defining what actually constitutes a publishable ML contribution in proteomics or glycomics. The group also highlighted the need for well-annotated, uncertainty-aware training and benchmarking datasets, including glycopeptide data, and began drafting two manuscripts – one on software quality and reporting expectations for ML in proteomics, and one on explainable and interpretable AI in MS-based proteomics.

Glycomics and Glycoproteomics

The “glyco” group focused on two tightly linked goals: improving confidence and comparability in glycan/glycopeptide identification and quantification, and lowering the barrier of entry for new researchers. First, the group outlined a plan for harmonizing glycan search spaces and reporting. A key recommendation is that outputs should carry standardized GlyTouCan identifiers and clearly encode the level of structural specificity (composition-only, topology, full linkage) so that results from different software tools can be compared on a common specificity level. The group also emphasized the need for explicit false-discovery rate (FDR) frameworks for glycan assignments, including topology- and isomer-sensitive scoring. Second, the working group substantially advanced two manuscripts: a best-practices/tutorial document for new glycosylation researchers (terminology, pitfalls, reporting standards), and a focused manuscript on how glycan structure affects glycopeptide signal intensity and the downstream challenges for quantification and biological interpretation. Writing responsibilities, timelines, and revision plans were agreed, and a first integrated draft was produced on-site.

Cross-cutting Topics: Federated Learning, Data Sharing and Credit, Multi-Omics Integration, and Quantitative Glyco-Proteomics

In the second half of the seminar week, people rotated between groups to discuss a number of cross-cutting topics and common challenges: Federated learning and controlled-access clinical data: While federated learning is still rarely used in proteomics due to widespread centralized data deposition, this is expected to change as clinical data are increasingly held locally for regulatory and privacy reasons. The group concluded that now is the time to define incentives, governance, and credit mechanisms for data generators so that high-value but unpublished datasets can still drive model development without leaving institutional boundaries. Multiomics integration: Participants discussed how to integrate proteomics with transcriptomics, phosphoproteomics, glycoproteomics, metabolomics/lipidomics, immunopeptidomics, and mass spectrometry imaging (and other spatially resolved data). The consensus was that most current “integration” is actually late-stage comparison of separate analyses. True multi-omics fusion is blocked by non-synchronous sampling, heterogeneous sample preparation, lack of common quality control (standards), differing biological timescales, and underdeveloped statistical control across layers. Guidance is needed on realistic experimental design and on levels of integration (early, mid-level latent, late/pathway). Quantitative glyco-proteomics statistics: The group outlined a plan to benchmark methods for glycoform quantification, normalization, missing value handling, and site occupancy estimation, leveraging tools such as MSstats/MSstats-PTM and experimentally perturbed datasets.

The 2025 Dagstuhl Seminar crystallized a shift in the field: from chasing marginal analytical improvements toward building scalable, interpretable, clinically relevant, and socially sustainable proteomics. The seminar concluded with concrete action items: manuscripts on ML standards and explainability; a best-practices/tutorial manuscript and a quantitative glycoform analysis manuscript from the glyco group; an ENIGMA proposal for large-scale, AI-ready translational proteomics data; a position statement on data sharing, incentive structures, and federated learning; and guidance on realistic multi-omics integration and QC. Together, these efforts define a forward-looking agenda for computational and translational proteomics in the coming years.

The discussions on trustworthiness and quality control in proteomics also fed directly into the planning of a Lorentz Center Workshop on “Trustworthiness in Proteomics”, successfully co-organized by Mathias Wilhelm and Magnus Palmblad in Leiden, in February 2026.

Copyright Rebekah Gundry, Magnus Palmblad, and Mathias Wilhelm

Motivation

In proteomics, novel algorithms are drivers of innovation, enhancing extraction of actionable insights from collected data and thereby advancing the field. We have identified three highly interesting computational challenges (and thus opportunities) that have come to the foreground in recent years: First is the application of machine learning in proteomics, which is already contributing to advancements in data analysis and interpretation in areas such as protein identification, functional annotations, and personalized medicine. Second is the rapidly growing field of glycoproteomics, which links to therapeutics as many therapeutic proteins are glycosylated. The focus here will lie on the characterization of the glycan moieties attached to proteins, the analysis of which poses specific computational challenges. Third is a cross-cutting topic at the interface between proteomics and other omics in the context of translational research. This includes both mass spectrometry and new approaches to generate, interrogate, and integrate data related to proteins, their transcripts and related metabolites. These topics are at the forefront of arising technologies and require novel solutions to emerging, highly complex problems.

As different experts need to be brought together to tackle these three topics, this Dagstuhl Seminar is thoroughly interdisciplinary: computer scientists, bioinformaticians, and statisticians, who develop algorithms and software for data interpretation; experimental life scientists that rely on proteomics as a key means to elucidate biology; and analytical chemists and engineers that develop new instruments and approaches to deliver ever more comprehensive and accurate data. Throughout, industry plays crucial roles as instrument and software vendors, and as advanced users driving applications, including in the development of new diagnostics and pharmaceuticals. Industry participation is therefore explicitly included in this seminar.

Attendance will therefore be built from a diverse group of participants, including computational and experimental scientists, and academic and industrial researchers across the three relevant domains of this seminar. The goal is to uncover as-yet unexplored synergies from interactions within and across these various backgrounds; a process which has already proven highly effective and inspiring in previous Dagstuhl Seminars on Computational Proteomics. A strong focus throughout the week thus will be on the free exchange of ideas between participants of different backgrounds to maximally benefit from obvious as well as less obvious synergies and to provide maximal opportunity for cross-fertilisation of ideas. To accommodate this, the seminar structure will be quite flexible, allowing for spontaneous working groups to emerge alongside pre-planned ones, and providing opportunity for any interested participant to start a discussion on any related topic of interest.

This interdisciplinary seminar is therefore poised to enable novel, breakthrough developments in computational proteomics around the three topics. The breadth and interrelatedness of the topics will provide a unique opportunity to bring together experimental, computational, translational, and clinical researchers from academic, government, and industry to come together to develop and inform future solutions in these topics.

Copyright Rebekah Gundry, Magnus Palmblad, and Mathias Wilhelm

Participants

Please log in to DOOR to see more details.

  • Charlotte Adams (University of Antwerp, BE)
  • Kiyoko Aoki-Kinoshita (Soka University - Tokyo, JP) [dblp]
  • Gad Armony (Bruker Nederland - Leiderdorp, NL)
  • Wout Bittremieux (University of Antwerp, BE) [dblp]
  • Isabell Bludau (Unviversitätsklinikum Heidelberg, DE) [dblp]
  • Robert Chalkley (University of California - San Francisco, US) [dblp]
  • Tine Claeys (Ghent University, BE)
  • Stephanie Cologna (University of Illinois - Chicago, US)
  • Eric Deutsch (Institute for Systems Biology - Seattle, US) [dblp]
  • Patrick Emery (Matrix Science Ltd. - London, GB) [dblp]
  • Melanie Föll (Universitätsklinikum Freiburg, DE) [dblp]
  • Wassim Gabriel (TU München - Freising, DE)
  • Paula González Menéndez (University of Adelaide, AU)
  • Rebekah Gundry (University of Nebraska - Omaha, US) [dblp]
  • Devon Kohler (Northeastern University - Boston, US)
  • Lev Levitskiy (University of Southern Denmark - Odense, DK)
  • Klaus Lindpaintner (Bruker - Concord, US)
  • Frédérique Lisacek (Swiss Institute of Bioinformatics - Geneva, CH) [dblp]
  • Zhiwei Liu (Westlake University - Hangzhou, CN)
  • Sriram Neelamegham (University at Buffalo - SUNY, US) [dblp]
  • Magnus Palmblad (Leiden University Medical Center, NL) [dblp]
  • Daniel Polasky (University of Michigan - Ann Arbor, US)
  • Rene Ranzinger (University of Georgia, US)
  • Tobias Schmidt (MSAID - Garching, DE) [dblp]
  • Nicola Ternette (University of Dundee, GB) [dblp]
  • Sergey Vakhrushev (University of Copenhagen, DK)
  • Hans Wessels (Radboud University Nijmegen, NL)
  • Mathias Wilhelm (TU München - Freising, DE) [dblp]
  • Dirk Winkelhardt (Ruhr-Universität-Bochum, DE)
  • Bernd Wollscheid (ETH Zürich, CH) [dblp]

Related Seminars
  • Dagstuhl Seminar 05471: Computational Proteomics (2005-11-20 - 2005-11-25) (Details)
  • Dagstuhl Seminar 08101: Computational Proteomics (2008-03-02 - 2008-03-07) (Details)
  • Dagstuhl Seminar 13491: Computational Mass Spectrometry (2013-12-01 - 2013-12-06) (Details)
  • Dagstuhl Seminar 15351: Computational Mass Spectrometry (2015-08-23 - 2015-08-28) (Details)
  • Dagstuhl Seminar 17421: Computational Proteomics (2017-10-15 - 2017-10-20) (Details)
  • Dagstuhl Seminar 19351: Computational Proteomics (2019-08-25 - 2019-08-30) (Details)
  • Dagstuhl Seminar 21271: Computational Proteomics (2021-07-04 - 2021-07-09) (Details)
  • Dagstuhl Seminar 23301: Computational Proteomics (2023-07-23 - 2023-07-28) (Details)

Classification
  • Emerging Technologies
  • Machine Learning
  • Other Computer Science

Keywords
  • proteomics
  • bioinformatics
  • machine learning
  • mass spectrometry
  • glycomics