← Back to all development cases
Development case · blind model judgment pending independent human review; not a formal Benchmark conclusion.
AutocompleteBoth unusable33 / 162 · a09bfd3621e258b6

A Review of BioTree Construction in the Context of Information Fusion: Priors, Methods, Applications and Trends

Quantitative biology · 2410.04815v2

FLOWING EVIDENCE BENCHMARK

How do we tell whether Evidence helps?

At the same writing position, with the same model and task, how does supplying retrieved paper passages change the first output? We compare matched versions and retain ties, unusable outputs and incomplete reviews.

SAME DRAFT · EVIDENCE ON OR OFF

01 · FIXED WRITING POSITION

MANUSCRIPT

Same writing position ▌

Matched autocomplete at the same draft position

02 · TWO MATCHED INPUTS

SHARED BY BOTH

Manuscript context, model, task and prompt

A · With Evidence

Retrieved paper passages supplied

B · Without Evidence

No retrieved passages supplied

LLM

Same model and version

A → first output

B → first output

Blind judge agent

First outputs are anonymized as X and Y

Continuation review: accuracy, fit to the writing task and usability
Returns: X preferred / tie / Y preferred / both unusable

Order check: X / Y → Y / X

HOW DOES BLIND JUDGING WORK?

① Anonymize both outputs
The judge sees the same draft and both first outputs without knowing which received Evidence.

② Compare and swap order
The judge applies task-specific criteria in X/Y and then Y/X order.

③ Review disagreements
A third pass resolves disagreements. Incomplete reviews remain in the denominator.

How are autocomplete positions stratified?

Before seeing generation outcomes, we check whether a retrieved passage contains a specific proposition that directly supports the next writing move. Those positions appear in the left opportunity group; the rest are ordinary positions on the right. We select a balanced sample from admitted papers in each field. The 50/50 split is experimental, not a measure of how often either type occurs in writing.

AUTOCOMPLETE · 110

Biology, statistics and astrophysics

55 positions on each side; statistics uses the ten-paper rerun.

AUTOCOMPLETE · 52

Psychology and climate science

26 positions on each side; psychology includes nine papers and climate science four.

How are the table percentages calculated?

Across five fields, 51 of 81 left-column positions preferred the Evidence version. Ties, pairs where both versions were unusable and incomplete reviews remain in the denominator.

51Evidence version preferred
÷
81All positions in this group
=
63%Evidence preference in this group

Source contribution is a separate review: 16 of 22 Evidence wins entered into source review directly used retrieved papers; another 29 wins await review.

These are development-stage model judgments pending independent human review. They are not formal Benchmark conclusions and do not, on their own, establish causality.

Manuscript writing position

Text excerpt · not a PDF page

Research manuscript · excerpt

A Review of BioTree Construction in the Context of Information Fusion: Priors, Methods, Applications and Trends

1 Backgrounds

…pproaches to multimodal integration and prior knowledge incorporation. Finally, the review discusses diverse applications of BioTrees in various biological disciplines, from phylogenetics to developmental biology, and outlines future trends in leveraging DL to advance BioTree research. By addressing the challenges of data complexity and prior knowledge integration, this review aims to inspire interdisciplinary innovation at the intersection of biology and DL. Keywords : Biological Tree Analysis, Deep Learning Information Fusion, Cell Differentiation Analysis, Biological Evolutionary Analysis,

↓ Next paragraph was generated here

The highlighted text is near the writing position. This is a layout of the evaluation input, not a PDF screenshot.

A matched comparison at the same writing positionJudgment: Both unusable

Both arms use the same model, task and manuscript context. Only the retrieved passages differ. The first recorded output from each arm is shown below.

Evidence

Same model and position · retrieved passages supplied · first output

Judged unusable

No continuation text was produced.

Model action:complete

Passages supplied to the Evidence version

Prior Knowledge Enhances Radiology Report Generation

Though achieving good results, few works considered incorporating prior knowledge, which can provide supplementary information for accurate reporting. For example, medical observations presented in a chest X-ray image are usually not isolated from each other, where underlying mutual influences may e…

Read full passage excerpt

Though achieving good results, few works considered incorporating prior knowledge, which can provide supplementary information for accurate reporting. For example, medical observations presented in a chest X-ray image are usually not isolated from each other, where underlying mutual influences may exist. Compared to experienced radiologists aware of such relationships, deep learning methods tend to suffer from the lack of knowledge if not explicitly taught, which limits the generation accuracy.

No Evidence

Same model and position · no retrieved passages · first output

Judged unusable

No continuation text was produced.

Model action:complete

Blind model judgment

The blind model judge marked both outputs unusable; this is counted separately from a tie.

Read the original model judgment

Both outputs A and B return empty text with action 'complete'. The task requires continuing the manuscript with one coherent academic paragraph. The draft ends with keywords, suggesting the manuscript is incomplete or the blank paragraph needs content before the keywords section. The provided source discusses prior knowledge enhancing radiology report generation, which has thematic overlap with 'prior knowledge incorporation' mentioned in the draft, but neither output attempts to generate any paragraph content. An empty response when the task clearly needs content is unusable per the rules.

The passages are input material; their presence does not mean an output used them or that they support every claim. Draft and source passages are excerpted; no simulated scores or PDF appear here.