← Back to all development cases
Development case · blind model judgment pending independent human review; not a formal Benchmark conclusion.
AutocompleteEvidence version preferred46 / 162 · 614c35575aa1f1db

Large Language Models in Bioinformatics: A Survey

Quantitative biology · 2503.04490v3

FLOWING EVIDENCE BENCHMARK

How do we tell whether Evidence helps?

At the same writing position, with the same model and task, how does supplying retrieved paper passages change the first output? We compare matched versions and retain ties, unusable outputs and incomplete reviews.

SAME DRAFT · EVIDENCE ON OR OFF

01 · FIXED WRITING POSITION

MANUSCRIPT

Same writing position ▌

Matched autocomplete at the same draft position

02 · TWO MATCHED INPUTS

SHARED BY BOTH

Manuscript context, model, task and prompt

A · With Evidence

Retrieved paper passages supplied

B · Without Evidence

No retrieved passages supplied

LLM

Same model and version

A → first output

B → first output

Blind judge agent

First outputs are anonymized as X and Y

Continuation review: accuracy, fit to the writing task and usability
Returns: X preferred / tie / Y preferred / both unusable

Order check: X / Y → Y / X

HOW DOES BLIND JUDGING WORK?

① Anonymize both outputs
The judge sees the same draft and both first outputs without knowing which received Evidence.

② Compare and swap order
The judge applies task-specific criteria in X/Y and then Y/X order.

③ Review disagreements
A third pass resolves disagreements. Incomplete reviews remain in the denominator.

How are autocomplete positions stratified?

Before seeing generation outcomes, we check whether a retrieved passage contains a specific proposition that directly supports the next writing move. Those positions appear in the left opportunity group; the rest are ordinary positions on the right. We select a balanced sample from admitted papers in each field. The 50/50 split is experimental, not a measure of how often either type occurs in writing.

AUTOCOMPLETE · 110

Biology, statistics and astrophysics

55 positions on each side; statistics uses the ten-paper rerun.

AUTOCOMPLETE · 52

Psychology and climate science

26 positions on each side; psychology includes nine papers and climate science four.

How are the table percentages calculated?

Across five fields there are 81 source-opportunity positions. The original blind review preferred Evidence in 51; a task check moved one empty Evidence continuation to both unusable, leaving 50 in public counts. In the ordinary group, another pair of empty outputs moved from no winner to both unusable. Original verdicts remain visible on case pages.

50Evidence version preferred
÷
81All positions in this group
=
62%Evidence preference in this group

Source contribution is a separate review: 16 of 21 Evidence wins entered into source review directly used retrieved papers; another 29 wins await review.

These are development-stage model judgments pending independent human review. They are not formal Benchmark conclusions and do not, on their own, establish causality.

Manuscript writing position

Text excerpt · not a PDF page

Research manuscript · excerpt

Large Language Models in Bioinformatics: A Survey

1 Introduction

…l data. This survey provides a systematic review of recent advancements, focusing on genomic sequence modeling, RNA structure prediction, protein function inference, and single-cell transcriptomics. Meanwhile, we also discuss several key challenges, including data scarcity, computational complexity, and cross-omics integration, and explore future directions such as multimodal learning, hybrid AI models, and clinical applications. By offering a comprehensive perspective, this paper underscores the transformative potential of LLMs in driving innovations in bioinformatics and precision medicine.

…ve demonstrated remarkable progress in the domain of natural language processing (NLP), with applications that span a wide array of tasks Min et al. 2023 ; Jiang et al. 2023 ; Raiaan et al. 2024 ; Jiang et al. 2025a . However, the nature of biological data and the associated tasks differ significantly from text data, presenting unique challenges. The accurate and precise handling of biomedical data to effectively form features and embeddings suitable for LLMs is an ongoing challenge that necessitates innovative solutions Chen et al. 2022 ; Wang et al. 2023a ; ESM Team 2024 ; Liu et al. 2025 .

↓ Next paragraph was generated here

The highlighted text is near the writing position. This is a layout of the evaluation input, not a PDF screenshot.

A matched comparison at the same writing positionJudgment: Evidence version preferred

Both arms use the same model, task and manuscript context. Only the retrieved passages differ. The first recorded output from each arm is shown below.

Evidence

Same model and position · retrieved passages supplied · first output

Preferred by blind review

Recent work has begun to assess how far LLMs can be pushed toward the data science tasks that arise in clinical research, using collections of manually crafted and cross-verified coding tasks drawn from published studies as reference solutions.

Passages supplied to the Evidence version

Can Large Language Models Replace Data Scientists in Clinical Research? ↗

Our objective was to evaluate the practical utility of LLMs in handling complex clinical research data and performing the associated data science tasks. To this end, we identified 39 clinical studies published in medical journals that were linked to patient-level datasets (Fig. 1 a). We started by e…

Read full passage excerpt

Our objective was to evaluate the practical utility of LLMs in handling complex clinical research data and performing the associated data science tasks. To this end, we identified 39 clinical studies published in medical journals that were linked to patient-level datasets (Fig. 1 a). We started by extracting and summarizing the analyses performed in these studies, such as patient characteristic exploration and Kaplan-Meier curves. We then developed the code necessary to reproduce these analyses and the reported results in these studies. These coding tasks, along with their reference solutions, were all manually crafted and cross-verified to ensure accuracy. The result was a collection of 293 diverse, high-quality data science tasks, covering primary tools used in Python and R, e.g., lifelines for survival analysis in Python and Bioconductor for biomedical data analysis in R.

No Evidence

Same model and position · no retrieved passages · first output

Not preferred in this comparison

The product produced no insertable continuation text.

The model explicitly returned a no-completion marker.

Judgment and task check

The blinded model review preferred the Evidence version. Its original rationale appears below.

Read the original model judgment

In this rationale, Output A = No Evidence; Output B = Evidence.

Output A is empty and provides no content for the autocomplete task, which clearly needs a paragraph continuation. Output B provides a coherent academic paragraph that transitions from the discussion of biological data challenges to a relevant, source-grounded observation about clinical research data science tasks. The B output draws directly from the supplied source about evaluating LLMs in clinical research with manually crafted coding tasks from published studies, matching entity (LLMs, clinical research, data science tasks), relation (evaluation/assessment), and scope. The B output does not invent this paper's own methods or results; it cites external work in a general way appropriate for a survey paper's introduction/related work section. No citation markers are generated that would violate insertion contract.

Source review confirmed that this continuation directly used specific, verifiable information from the retrieved paper.

The passages are input material; their presence does not mean an output used them or that they support every claim. Draft and source passages are excerpted; no simulated scores or PDF appear here.