📣 Help Shape the Future of UKRI's Gateway to Research (GtR)

We're improving UKRI's Gateway to Research and are seeking your input! If you would be interested in being interviewed about the improvements we're making and to have your say about how we can make GtR more user-friendly, impactful, and effective for the Research and Innovation community, please email gateway@ukri.org.

ProtFunAI: AI based methods for functional annotation of proteins in crop genomes

Lead Research Organisation: UNIVERSITY COLLEGE LONDON
Department Name: Structural Molecular Biology

Abstract

Our project will 'build on existing links and deepen existing relationships' between the two groups pioneering the development of AI/Deep-Learning models for proteins (Rost Group, TUM) and the application of these to protein domain families (Orengo Group, UCL). It will leverage world leading expertise in protein Language Models (pLMs) in order to accelerate the scientific discovery of protein functions in the genomes of key agricultural crops important for food security. However, our approaches will be generic and rolled out to all UniProt proteins through existing collaborations.

Synergies between both groups have evolved over several collaborations. Since 2019, ground-breaking results tuned the pLMs developed in the Rost Group (e.g. the ProtTrans series, incl. ProtT5, ProtTucker) with protein family and functional family data (CATH superfamilies and FunFams) generated and maintained by the Orengo Group.

The partnership proposed, here, would allow researchers in the Rost and Orengo Groups to intensify exchanges through visiting each others labs and interacting more comprehensively to design more effective protocols that enhance (1) protein homologue detection (2) protein function prediction and (3) protein functional site prediction.

The Orengo and Rost Groups began collaborating in 2000 when working together on protein family analysis for target identification in the NIH-funded USA Structural Genomics initiative (PSI), which ended in 2015 [21-23]. Subsequently funding from the German BMBF (Federal German Research Ministry) and DFG (German Research Foundation) supported visits of PhD and Masters students from both groups and resulted in the development of new approaches for protein function prediction [14,15]. This application seeks funds to continue these collaborations to leverage the latest advances in AI/Deep Learning. The Rost Group recently enhanced their pLMs significantly (ProstT5 [18]) and the funding would allow us to apply ProstT5 to exploit the hugely expanded CATH classification, which is currently integrating hundreds of millions of predicted protein structures from the AlphaFold portal (AFDB).

The application is very timely as it will address key BBSRC strategic priorities around data intensive biology and AI and the important challenge of food security. We will apply improved function prediction methods to significantly increase the functional annotations of plant genomes. This will bring 'new knowledge about key biological principles and mechanisms using AI-based approaches' and bring 'AI in sustainable agriculture and food' and enable 'smart agriculture' by identifying genes implicated in biological systems associated with growth and stress resistance e.g. drought and antimicrobial resistance. Most genes (typically >90%) from plants valuable as crops (e.g. wheat, maize, rice, sorghum) are experimentally uncharacterized or very poorly annotated. Our methods will be state-of-the-art to accurately guide experimental validation.

We will disseminate the annotations using our established web-based CATH resource accessed by over 27,000 users/month. Since CATH data is also disseminated by PDB, UniProt and InterPro the predictions will be accessible to >900,000s of users/month. We will also work closely with collaborators in the UK researching plant genomes to get feedback and solicit experimental validation where possible.

The project will significantly enhance the AI/ML skills of UK based researchers in the Orengo Group, whose prior training was largely in biology. On the flip side, the more AI-focused members from the Rost group will deepen their understanding of individual proteins, organisms, and evolution. German scholars will also dive deeper into the workings of UK-based resources.

Publications

10 25 50
 
Description ContraSTED:
Modern protein structure prediction has created hundreds of millions of protein structures, but many cannot yet be linked to known evolutionary families. ContrasTED learns a specialized representation of protein domains so that evolutionarily related proteins appear close together in a learned space. This makes it possible to quickly classify known families and discover entirely new ones among millions of structures.

ProFam learns from groups of related proteins instead of individual sequences. By seeing how proteins vary across evolution within a family, the model can better understand which mutations are allowed and which are important for function. This helps it both predict how mutations affect proteins and generate new protein sequences that still resemble naturally occurring proteins. The data ia available via below:
https://github.com/alex-hh/profam/ ; https://zenodo.org/records/17713590 ; https://huggingface.co/judewells/ProFam-1
Exploitation Route ? Guided Sequence Design: Generate novel, functional, and stable protein sequences, accelerating the design of new proteins like enzymes or antibodies.
? Exploring Family Diversity: Identify mutation-tolerant regions and discover novel functionalities within protein families. Also identifying novel superfamilies in the AFDB (UniProt based), ESMAtlas (MGnify based), BFVD (viral sequence database) datasets.
? Improving Fitness Prediction: Use ProFam's likelihood scores to guide experimental efforts in directed evolution or library design.
? De novo Protein Design: Design entirely new protein families with desired properties based on structural and functional constraints.

ContraSTED outcomes are useful to the community in terms of following:
• Large-scale protein annotation: Assign newly predicted protein structures to evolutionary superfamilies quickly and accurately.
• Discovery of new protein families: Detect clusters of proteins that likely represent previously unknown evolutionary groups.
• Improving structural classification systems: Identify cases where existing superfamily boundaries may need revision.
Sectors Education

Environment

Healthcare

URL https://github.com/alex-hh/profam/
 
Title ContrasTED 
Description ContrasTED is a supervised contrastive learning framework designed to classify protein domains into evolutionary superfamilies at large scale. The method projects structure-aware protein language model embeddings into a latent space optimized for superfamily discrimination. By learning representations where homologous domains cluster together, ContrasTED enables efficient nearest-neighbor classification of protein structures. The approach is designed to address the growing challenge posed by hundreds of millions of predicted protein structures that lack evolutionary annotations. 
Type Of Material Improvements to research infrastructure 
Year Produced 2025 
Provided To Others? No  
Impact • Achieves state-of-the-art accuracy in CATH superfamily classification using nearest-neighbor lookup in the learned embedding space. • Outperforms sequence-based, structure-alignment, and general embedding baselines for remote homology detection. • Scales to tens of millions of structures, enabling classification across the rapidly expanding structural proteome. • Applied to 20.8 million unannotated domains from the TED (The Encyclopedia of Domains), identifying 20,661 candidate novel superfamilies spanning 9.8 million domains. • Suggests potential revisions to existing classifications by identifying superfamily pairs with overlapping embedding neighborhoods. 
 
Title ProFam model 
Description ProFam-1 is a 251M-parameter autoregressive protein family language model (pfLM) designed for modelling and generating proteins in an evolutionary context. The model is trained with next-token prediction on millions of protein families represented as concatenated, unaligned sets of homologous sequences. By conditioning on entire protein families rather than single sequences, ProFam-1 explicitly captures evolutionary relationships such as residue conservation and covariance. The model supports tasks including zero-shot fitness prediction and homology-guided sequence generation, producing diverse sequences while preserving structural and evolutionary constraints. 
Type Of Material Computer model/algorithm 
Year Produced 2025 
Provided To Others? Yes  
Impact ProFam integrates multiple family definitions, enabling multimodal control in generation and employs efficient training. Demonstrates competitive performance on the ProteinGym benchmark for zero-shot fitness prediction (Spearman 0.47 for substitutions and 0.53 for indels). Releases a large curated training dataset (ProFam Atlas) and full training/inference pipelines as open source, enabling reproducibility and further research in protein family language models. Generates diverse sequences with predicted structural similarity while maintaining evolutionary conservation patterns. 
URL https://github.com/alex-hh/profam
 
Description ProtFunAI - Collababoration with Burkhard Rost Team 
Organisation Technical University of Munich
Country Germany 
Sector Academic/University 
PI Contribution Development of deep learning algorithms for protein function prediction, protein classification and analysis
Collaborator Contribution Training in deep learning protocols and protein language models. Contributions to project design. Novel protein language models to generate protein embeddings for protein function prediction and other protein based prediction tasks.
Impact Project has just started so no outputs yet
Start Year 2024