← Back to all development cases
Development case · blind model judgment pending independent human review; not a formal Benchmark conclusion.
AutocompleteEvidence version preferred18 / 162 · 905ac4ed89fc6ffc

Generative AI for Controllable Protein Sequence Design: A Survey

Quantitative biology · 2402.10516v1

FLOWING EVIDENCE BENCHMARK

How do we tell whether Evidence helps?

At the same writing position, with the same model and task, how does supplying retrieved paper passages change the first output? We compare matched versions and retain ties, unusable outputs and incomplete reviews.

SAME DRAFT · EVIDENCE ON OR OFF

01 · FIXED WRITING POSITION

MANUSCRIPT

Same writing position ▌

Matched autocomplete at the same draft position

02 · TWO MATCHED INPUTS

SHARED BY BOTH

Manuscript context, model, task and prompt

A · With Evidence

Retrieved paper passages supplied

B · Without Evidence

No retrieved passages supplied

LLM

Same model and version

A → first output

B → first output

Blind judge agent

First outputs are anonymized as X and Y

Continuation review: accuracy, fit to the writing task and usability
Returns: X preferred / tie / Y preferred / both unusable

Order check: X / Y → Y / X

HOW DOES BLIND JUDGING WORK?

① Anonymize both outputs
The judge sees the same draft and both first outputs without knowing which received Evidence.

② Compare and swap order
The judge applies task-specific criteria in X/Y and then Y/X order.

③ Review disagreements
A third pass resolves disagreements. Incomplete reviews remain in the denominator.

How are autocomplete positions stratified?

Before seeing generation outcomes, we check whether a retrieved passage contains a specific proposition that directly supports the next writing move. Those positions appear in the left opportunity group; the rest are ordinary positions on the right. We select a balanced sample from admitted papers in each field. The 50/50 split is experimental, not a measure of how often either type occurs in writing.

AUTOCOMPLETE · 110

Biology, statistics and astrophysics

55 positions on each side; statistics uses the ten-paper rerun.

AUTOCOMPLETE · 52

Psychology and climate science

26 positions on each side; psychology includes nine papers and climate science four.

How are the table percentages calculated?

Across five fields, 51 of 81 left-column positions preferred the Evidence version. Ties, pairs where both versions were unusable and incomplete reviews remain in the denominator.

51Evidence version preferred
÷
81All positions in this group
=
63%Evidence preference in this group

Source contribution is a separate review: 16 of 22 Evidence wins entered into source review directly used retrieved papers; another 29 wins await review.

These are development-stage model judgments pending independent human review. They are not formal Benchmark conclusions and do not, on their own, establish causality.

Manuscript writing position

Text excerpt · not a PDF page

Research manuscript · excerpt

Generative AI for Controllable Protein Sequence Design: A Survey

1 Introduction

…gorithms, have been propelling the protein design field towards an unprecedented revolution. In this survey, we systematically review recent advances in generative AI for controllable protein sequence design. To set the stage, we first outline the foundational tasks in protein sequence design in terms of the constraints involved and present key generative models and optimization algorithms. We then offer in-depth reviews of each design task and discuss the pertinent applications. Finally, we identify the unresolved challenges and highlight research opportunities that merit deeper exploration.

…ns, designing novel amino acid sequences that encode proteins with desired functions has been a central challenge in bioengineering. Nevertheless, due to the vast exploration space of possible proteins, conventional experimental methods, such as directed evolution Arnold 1998 , are prohibitively time-intensive and money-consuming. To overcome these challenges, recent years have witnessed the trend of leveraging generative AI technologies to explore the biochemical space intelligently. Fueled by advances in generative AI, the field of protein design is experiencing an unprecedented revolution.

↓ Next paragraph was generated here

The highlighted text is near the writing position. This is a layout of the evaluation input, not a PDF screenshot.

A matched comparison at the same writing positionJudgment: Evidence version preferred

Both arms use the same model, task and manuscript context. Only the retrieved passages differ. The first recorded output from each arm is shown below.

Evidence

Same model and position · retrieved passages supplied · first output

Preferred by blind review

Generative models have thus emerged as promising candidates for sequence-data-driven protein design, enabling both the generation of novel sequences and the extraction of structural and functional information from rapidly growing sequence databases. Unlike conventional experimental methods, these models learn the underlying distribution of protein sequences directly from data, offering a computational route to explore the vast biochemical space at substantially lower cost.

Passages supplied to the Evidence version

Efficient generative modeling of protein sequences using simple autoregressive models

Generative models emerge as promising candidates for novel sequence-data driven approaches to protein design, and for the extraction of structural and functional information about proteins deeply hidden in rapidly growing sequence databases. Here we propose simple autoregressive models as highly acc…

Read full passage excerpt

Generative models emerge as promising candidates for novel sequence-data driven approaches to protein design, and for the extraction of structural and functional information about proteins deeply hidden in rapidly growing sequence databases. Here we propose simple autoregressive models as highly accurate but computationally extremely efficient generative sequence models. We show that they perform similarly to existing approaches based on Boltzmann machines or deep generative models, but at a substantially lower computational cost. Furthermore, the simple structure of our models has distinctive mathematical advantages, which translate into an improved applicability in sequence generation and evaluation. Using these models, we can easily estimate both the model probability of a given sequence, and the size of the functional sequence space related to a specific protein family.

No Evidence

Same model and position · no retrieved passages · first output

Not preferred in this comparison

This survey is organized as follows. Section 2 introduces the foundational tasks in protein sequence design and the key generative models and optimization algorithms that underpin them. Sections 3 through 5 then review generative approaches for each design task in turn, before Section 6 discusses applications and Section 7 outlines unresolved challenges and future opportunities.

Blind model judgment

The blinded model review preferred the Evidence version. Its original rationale appears below.

Read the original model judgment

Output B provides a more coherent continuation of the specific paragraph context. The draft shows a paragraph ending with 'Fueled by advances in generative AI, the field of protein design is experiencing an unprecedented revolution' and the next paragraph beginning with 'designing novel amino acid sequences...' with a gap in between. Output B naturally bridges this transition by elaborating on how generative models enable protein design and extract information from databases, directly supported by the source snippet which states 'Generative models emerge as promising candidates for novel sequence-data driven approaches to protein design, and for the extraction of structural and functional information about proteins deeply hidden in rapidly growing sequence databases.' Output B also correctly notes the 'substantially lower computational cost,' supported by the source. Output A, while factually accurate about survey structure, reads more like a table of contents or meta-commentary on the paper's organization rather than a flowing continuation of the scientific narrative. The draft already has an introductory paragraph stating the survey's organization ('In this survey, we systematically review...'), making Output A somewhat redundant. Output B advances the substantive argument about generative AI for protein design, matching the draft's trajectory from the 'unprecedented revolution' claim toward the specific challenge of sequence design.

Source review confirmed that this continuation directly used specific, verifiable information from the retrieved paper.

The passages are input material; their presence does not mean an output used them or that they support every claim. Draft and source passages are excerpted; no simulated scores or PDF appear here.