Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
OpenMed built an end-to-end protein AI pipeline covering structure prediction, sequence design, and codon optimization. They trained and compared multiple transformer architectures for codon-level language modeling, finding CodonRoBERTa-large-v2 as the best with perplexity 4.10 and Spearman CAI correlation 0.40. They scaled to 25 species, training 4 production models in 55 GPU-hours, and built a species-conditioned system that no other open-source project offers.
From the source
CodonRoBERTa-large-v2 emerged as the clear winner with a perplexity of 4.10 and a Spearman CAI correlation of 0.40, significantly outperforming ModernBERT. We then scaled to 25 species, trained 4 production models in 55 GPU-hours, and built a species-conditioned system that no other open-source project offers.
huggingface.co