Preprint / Version 1

Machine Learning methods for Small Wet-Lab Data Challenges in Enzyme Engineering

##article.authors##

  • Yumeng Zhang Wuya College of Innovation, Shenyang Pharmaceutical University

DOI:

https://doi.org/10.31224/8043

Keywords:

enzyme engineering, Data scarcity, Protein language models, Transfer learning, Zero shot prediction, Active learning

Abstract

Enzyme engineering is fundamentally constrained by the scarcity of high-quality sequencefunction data, and the vastness of protein sequence space is severely mismatched with the limited through put of experimental characterization. Traditional machine learning models, which rely heavily on massive labeled datasets, further exacerbate this dilemma. This review surveys machine learning methods specifically designed to address the small-data challenge in enzyme engineering, providing an in-depth analysis of each approach and a detailed discussion of recent progress. Furthermore, we consolidate existing achievements in this field and evaluate the effectiveness of these methods in tack ling small-data problems from the perspective of specific enzymatic properties, while also summariz ing the remaining challenges. This review serves as a practical methodology reference for enzyme engineers working under limited experimental data, and offers a clear and comprehensive theoret ical framework for computer scientists to develop more powerful solutions. Looking forward, it is foreseeable that machine learning will propel enzyme engineering toward achieving superior results with fewer experimental requirements.

Downloads

Download data is not yet available.

Downloads

Posted

2026-08-24