← Back to all development cases
Development case · blind model judgment pending independent human review; not a formal Benchmark conclusion.
AutocompleteBoth unusable5 / 162 · 0dcb80855f23c0b6

Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade

Quantitative biology · 2202.03534v2

FLOWING EVIDENCE BENCHMARK

How do we tell whether Evidence helps?

At the same writing position, with the same model and task, how does supplying retrieved paper passages change the first output? We compare matched versions and retain ties, unusable outputs and incomplete reviews.

SAME DRAFT · EVIDENCE ON OR OFF

01 · FIXED WRITING POSITION

MANUSCRIPT

Same writing position ▌

Matched autocomplete at the same draft position

02 · TWO MATCHED INPUTS

SHARED BY BOTH

Manuscript context, model, task and prompt

A · With Evidence

Retrieved paper passages supplied

B · Without Evidence

No retrieved passages supplied

LLM

Same model and version

A → first output

B → first output

Blind judge agent

First outputs are anonymized as X and Y

Continuation review: accuracy, fit to the writing task and usability
Returns: X preferred / tie / Y preferred / both unusable

Order check: X / Y → Y / X

HOW DOES BLIND JUDGING WORK?

① Anonymize both outputs
The judge sees the same draft and both first outputs without knowing which received Evidence.

② Compare and swap order
The judge applies task-specific criteria in X/Y and then Y/X order.

③ Review disagreements
A third pass resolves disagreements. Incomplete reviews remain in the denominator.

How are autocomplete positions stratified?

Before seeing generation outcomes, we check whether a retrieved passage contains a specific proposition that directly supports the next writing move. Those positions appear in the left opportunity group; the rest are ordinary positions on the right. We select a balanced sample from admitted papers in each field. The 50/50 split is experimental, not a measure of how often either type occurs in writing.

AUTOCOMPLETE · 110

Biology, statistics and astrophysics

55 positions on each side; statistics uses the ten-paper rerun.

AUTOCOMPLETE · 52

Psychology and climate science

26 positions on each side; psychology includes nine papers and climate science four.

How are the table percentages calculated?

Across five fields, 51 of 81 left-column positions preferred the Evidence version. Ties, pairs where both versions were unusable and incomplete reviews remain in the denominator.

51Evidence version preferred
÷
81All positions in this group
=
63%Evidence preference in this group

Source contribution is a separate review: 16 of 22 Evidence wins entered into source review directly used retrieved papers; another 29 wins await review.

These are development-stage model judgments pending independent human review. They are not formal Benchmark conclusions and do not, on their own, establish causality.

Manuscript writing position

Text excerpt · not a PDF page

Research manuscript · excerpt

Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade

1 Introduction

…oil Dupont2016 ; Mahe2017 , and air ( Clare2022 ) , which is known as environmental DNA ( Deiner2017 ; Ruppert2019 , eDNA, ) , or from the human body Huttenhower2012 ; Methe2012 ; Matsen2015 ; Wang2015 and other sources Hanson2016 ; Gohli2019 ; Lorimer2019 ; ElRakaiby2019 . Crucially, this enables the ecological survey of a community of organisms in their immediate environment (i. e., in situ ), and allows to directly study the genetic composition of species communities (from viruses to megafauna); a field known as metagenomics ( Thomas2012 ; Oulas2015 ; Escobar-Zepeda2015 ; Lindgreen2016 ) .

…neration Sequencing ( Logares2012 ; Mardis2013 , NGS,) , as well as later generations ( Pareek2011 ; Niedringhaus2011 ; Mignardi2014 ; Heather2016 ; Mardis2017 ) . For a sample of biological material, these technologies typically produce thousands to millions or even billions of short genetic sequences (also called “reads”) with a length of some hundred base pairs length each. Over the past decades, decreasing costs and increasing throughput of sequencing technologies have caused an exponential growth in sequencing data ( Muir2016 ) , which has now passed the peta-scale barrier ( Katz2022 ) .

↓ Next paragraph was generated here

The highlighted text is near the writing position. This is a layout of the evaluation input, not a PDF screenshot.

A matched comparison at the same writing positionJudgment: Both unusable

Both arms use the same model, task and manuscript context. Only the retrieved passages differ. The first recorded output from each arm is shown below.

Evidence

Same model and position · retrieved passages supplied · first output

Judged unusable

No continuation text was produced.

The model requested more evidence and did not continue.

Passages supplied to the Evidence version

Clustering pipeline for determining consensus sequences in targeted next-generation sequencing

However, with the large number of reads obtained from NGS technologies, computing pairwise alignment scores for all pairs of reads is computationally prohibitive.

No Evidence

Same model and position · no retrieved passages · first output

Judged unusable

The scale and complexity of this data have driven the development of computational methods that can assign taxonomic or phylogenetic identity to individual reads without the need for full assembly or alignment. Phylogenetic placement, in particular, has emerged as a scalable alternative to de novo assembly and reference-based classification, enabling the identification of organisms directly from short metagenomic reads.

Blind model judgment

The blind model judge marked both outputs unusable; this is counted separately from a tie.

Read the original model judgment

Order-swapped judges disagreed. Pass 1: Output A provides a coherent academic paragraph that flows logically from the preceding text about sequencing data growth, discussing computational methods for handling large-scale metagenomic data. However, Output A contains unsupported claims: it introduces 'phylogenetic placement' as 'a scalable alternative to de novo assembly and reference-based classification' and claims it enables 'identification of organisms directly from short metagenomic reads' — none of which are supported by the supplied source, which only discusses computational challenges of pairwise alignment in NGS clustering pipelines. The source does not mention phylogenetic placement, taxonomic assignment without assembly, or identification of organisms from short reads. Output B returns empty with action 'needs_evidence', which is unusable for an autocomplete task that clearly needs content, but avoids unsupported claims. Between an empty unusable output and a content-filled output with material unsupported claims, Output A is more useful for insertion but must be marked unsupportedClaim=true due to the ungrounded technical claims about phylogenetic placement. However, per rules, any output with unsupportedClaim=true must have usable=false, which creates a conflict. Re-evaluating: Output B is deliberately empty and marked 'needs_evidence' — this is not a valid deliberate empty evidence-card since this is autocomplete, not evidence-card generation. The task clearly needs content. Output B is unusable. Output A, despite its unsupported claims, provides content that could be edited. But per strict rules, unsupportedClaim=true requires usable=false. Given internal consistency requirements, I must mark A unusable due to unsupported claims, making both unusable. Pass 2: Output A returns an empty text with action 'needs_evidence', which fails the task of continuing the manuscript with one coherent academic paragraph. The task clearly needs content, making this output unusable. Output B provides a relevant, coherent paragraph that logically follows from the draft: after discussing exponential growth of sequencing data passing peta-scale, it transitions to computational challenges and solutions (phylogenetic placement as alternative to assembly/alignment). While the sources provided do not directly support the specific claims about phylogenetic placement, the draft itself establishes the context of metagenomics and NGS data growth, and the paragraph represents a natural scholarly continuation without introducing unsupported factual claims about specific studies. The paragraph is topically appropriate and could be inserted with minimal editing. No generated citation markers are present.

The passages are input material; their presence does not mean an output used them or that they support every claim. Draft and source passages are excerpted; no simulated scores or PDF appear here.