Machine Learning methods for Small Wet-Lab Data Challenges in Enzyme Engineering
DOI:
https://doi.org/10.31224/8043Keywords:
enzyme engineering, Data scarcity, Protein language models, Transfer learning, Zero shot prediction, Active learningAbstract
Enzyme engineering is fundamentally constrained by the scarcity of high-quality sequencefunction data, and the vastness of protein sequence space is severely mismatched with the limited through put of experimental characterization. Traditional machine learning models, which rely heavily on massive labeled datasets, further exacerbate this dilemma. This review surveys machine learning methods specifically designed to address the small-data challenge in enzyme engineering, providing an in-depth analysis of each approach and a detailed discussion of recent progress. Furthermore, we consolidate existing achievements in this field and evaluate the effectiveness of these methods in tack ling small-data problems from the perspective of specific enzymatic properties, while also summariz ing the remaining challenges. This review serves as a practical methodology reference for enzyme engineers working under limited experimental data, and offers a clear and comprehensive theoret ical framework for computer scientists to develop more powerful solutions. Looking forward, it is foreseeable that machine learning will propel enzyme engineering toward achieving superior results with fewer experimental requirements.
Downloads
Downloads
Posted
License
Copyright (c) 2026 Yumeng Zhang

This work is licensed under a Creative Commons Attribution 4.0 International License.