← Back to all development cases
Development case · blind model judgment pending independent human review; not a formal Benchmark conclusion.
AutocompleteBoth unusable19 / 162 · b5f36b40ee124f25

Generative AI for Controllable Protein Sequence Design: A Survey

Quantitative biology · 2402.10516v1

FLOWING EVIDENCE BENCHMARK

How do we tell whether Evidence helps?

At the same writing position, with the same model and task, how does supplying retrieved paper passages change the first output? We compare matched versions and retain ties, unusable outputs and incomplete reviews.

SAME DRAFT · EVIDENCE ON OR OFF

01 · FIXED WRITING POSITION

MANUSCRIPT

Same writing position ▌

Matched autocomplete at the same draft position

02 · TWO MATCHED INPUTS

SHARED BY BOTH

Manuscript context, model, task and prompt

A · With Evidence

Retrieved paper passages supplied

B · Without Evidence

No retrieved passages supplied

LLM

Same model and version

A → first output

B → first output

Blind judge agent

First outputs are anonymized as X and Y

Continuation review: accuracy, fit to the writing task and usability
Returns: X preferred / tie / Y preferred / both unusable

Order check: X / Y → Y / X

HOW DOES BLIND JUDGING WORK?

① Anonymize both outputs
The judge sees the same draft and both first outputs without knowing which received Evidence.

② Compare and swap order
The judge applies task-specific criteria in X/Y and then Y/X order.

③ Review disagreements
A third pass resolves disagreements. Incomplete reviews remain in the denominator.

How are autocomplete positions stratified?

Before seeing generation outcomes, we check whether a retrieved passage contains a specific proposition that directly supports the next writing move. Those positions appear in the left opportunity group; the rest are ordinary positions on the right. We select a balanced sample from admitted papers in each field. The 50/50 split is experimental, not a measure of how often either type occurs in writing.

AUTOCOMPLETE · 110

Biology, statistics and astrophysics

55 positions on each side; statistics uses the ten-paper rerun.

AUTOCOMPLETE · 52

Psychology and climate science

26 positions on each side; psychology includes nine papers and climate science four.

How are the table percentages calculated?

Across five fields, 51 of 81 left-column positions preferred the Evidence version. Ties, pairs where both versions were unusable and incomplete reviews remain in the denominator.

51Evidence version preferred
÷
81All positions in this group
=
63%Evidence preference in this group

Source contribution is a separate review: 16 of 22 Evidence wins entered into source review directly used retrieved papers; another 29 wins await review.

These are development-stage model judgments pending independent human review. They are not formal Benchmark conclusions and do not, on their own, establish causality.

Manuscript writing position

Text excerpt · not a PDF page

Research manuscript · excerpt

Generative AI for Controllable Protein Sequence Design: A Survey

1 Introduction

…ns, designing novel amino acid sequences that encode proteins with desired functions has been a central challenge in bioengineering. Nevertheless, due to the vast exploration space of possible proteins, conventional experimental methods, such as directed evolution Arnold 1998 , are prohibitively time-intensive and money-consuming. To overcome these challenges, recent years have witnessed the trend of leveraging generative AI technologies to explore the biochemical space intelligently. Fueled by advances in generative AI, the field of protein design is experiencing an unprecedented revolution.

…t of the protein design space, lay the groundwork for sequence-based protein design. Deep generative models offer a promising tool for capturing the distribution underlying natural protein sequences to generate novel and diverse proteins. Moreover, making protein design controllable is crucial for the transition from theory to practice. Ideally, a practical method ought to generate proteins that adhere to task-specific structural and functional constraints. To this end, researchers have utilized conditional generative models and optimization algorithms to realize such tailored protein design.

↓ Next paragraph was generated here

The highlighted text is near the writing position. This is a layout of the evaluation input, not a PDF screenshot.

A matched comparison at the same writing positionJudgment: Both unusable

Both arms use the same model, task and manuscript context. Only the retrieved passages differ. The first recorded output from each arm is shown below.

Evidence

Same model and position · retrieved passages supplied · first output

Judged unusable

No continuation text was produced.

The model requested more evidence and did not continue.

Passages supplied to the Evidence version

Benchmarking deep generative models for diverse antibody sequence design

Computational protein design, i.e. inferring novel and diverse protein sequences consistent with a given structure, remains a major unsolved challenge. Recently, deep generative models that learn from sequences alone or from sequences and structures jointly have shown impressive performance on this…

Read full passage excerpt

Computational protein design, i.e. inferring novel and diverse protein sequences consistent with a given structure, remains a major unsolved challenge. Recently, deep generative models that learn from sequences alone or from sequences and structures jointly have shown impressive performance on this task. However, those models appear limited in terms of modeling structural constraints, capturing enough sequence diversity, or both. Here we consider three recently proposed deep generative frameworks for protein design: (AR) the sequence-based autoregressive generative model, (GVP) the precise structure-based graph neural network, and Fold2Seq that leverages a fuzzy and scale-free representation of a three-dimensional fold, while enforcing structure-to-sequence (and vice versa) consistency.

No Evidence

Same model and position · no retrieved passages · first output

Judged unusable

Despite this progress, the survey landscape remains fragmented: existing reviews tend to organize generative protein design either by model architecture or by application domain, and rarely by the control mechanisms that determine whether a generated sequence satisfies a specified constraint. This gap motivates the taxonomy adopted in this survey, which categorizes controllable protein sequence design methods according to the conditioning signals they accept and the design objectives they target.

Blind model judgment

The blind model judge marked both outputs unusable; this is counted separately from a tie.

Read the original model judgment

Output A is deliberately empty with action 'needs_evidence', which is unusable for an autocomplete task that clearly needs content. Output B provides a coherent academic paragraph that continues the manuscript's flow, transitioning from controllable protein design to a survey/taxonomy framing. However, Output B contains unsupported claims: it states that 'existing reviews tend to organize generative protein design either by model architecture or by application domain, and rarely by the control mechanisms' and introduces 'this survey' with a specific taxonomy — neither the draft nor the single supplied source mentions any survey, review organization patterns, or a taxonomy of control mechanisms. The source only benchmarks three specific deep generative models for antibody sequence design and discusses their limitations in modeling structural constraints and sequence diversity. The claim about how reviews are organized and the existence of 'this survey' is an invented framing with no source support. Despite this unsupported claim, Output B is the only output that provides usable content; Output A fails the basic requirement of providing completion text.

The passages are input material; their presence does not mean an output used them or that they support every claim. Draft and source passages are excerpted; no simulated scores or PDF appear here.