BENCHMARK

Can Evidence make autocomplete more accurate and better aligned with the draft?

The Benchmark compares first autocomplete outputs with and without Evidence at the same writing position. Evidence means paper passages retrieved and screened by Flowing from the local library. Only autocomplete results are currently public; polishing and verification remain internal development data.

FLOWING EVIDENCE BENCHMARK

How do we tell whether Evidence helps?

At the same writing position, with the same model and task, how does supplying retrieved paper passages change the first output? We compare matched versions and retain ties, unusable outputs and incomplete reviews.

SAME DRAFT · EVIDENCE ON OR OFF

01 · FIXED WRITING POSITION

MANUSCRIPT

Same writing position ▌

Matched autocomplete at the same draft position

02 · TWO MATCHED INPUTS

SHARED BY BOTH

Manuscript context, model, task and prompt

A · With Evidence

Retrieved paper passages supplied

B · Without Evidence

No retrieved passages supplied

LLM

Same model and version

A → first output

B → first output

Blind judge agent

First outputs are anonymized as X and Y

Continuation review: accuracy, fit to the writing task and usability
Returns: X preferred / tie / Y preferred / both unusable

Order check: X / Y → Y / X

HOW DOES BLIND JUDGING WORK?

① Anonymize both outputs
The judge sees the same draft and both first outputs without knowing which received Evidence.

② Compare and swap order
The judge applies task-specific criteria in X/Y and then Y/X order.

③ Review disagreements
A third pass resolves disagreements. Incomplete reviews remain in the denominator.

How are autocomplete positions stratified?

Before seeing generation outcomes, we check whether a retrieved passage contains a specific proposition that directly supports the next writing move. Those positions appear in the left opportunity group; the rest are ordinary positions on the right. We select a balanced sample from admitted papers in each field. The 50/50 split is experimental, not a measure of how often either type occurs in writing.

AUTOCOMPLETE · 110

Biology, statistics and astrophysics

55 positions on each side; statistics uses the ten-paper rerun.

AUTOCOMPLETE · 52

Psychology and climate science

26 positions on each side; psychology includes nine papers and climate science four.

How are the table percentages calculated?

Across five fields, 51 of 81 left-column positions preferred the Evidence version. Ties, pairs where both versions were unusable and incomplete reviews remain in the denominator.

51Evidence version preferred
÷
81All positions in this group
=
63%Evidence preference in this group

Source contribution is a separate review: 16 of 22 Evidence wins entered into source review directly used retrieved papers; another 29 wins await review.

These are development-stage model judgments pending independent human review. They are not formal Benchmark conclusions and do not, on their own, establish causality.
Autocomplete

Autocomplete development evaluation

Overview of all 5 fields in the table

Evidence version preferred

63%51/81

Where a source could directly support the next continuation

Wins entered into source review: verified direct use of the retrieved paper

73% · 16/22

16/22 · 73%

Of all 51 Evidence wins, 29 still await source review. The percentage above covers only wins entered into review.

No-Evidence version preferred

23%19/81

Same position and model, without retrieved passages

Evidence preferred 51Tie 3No-Evidence preferred 19Unusable or incomplete 8

This overview matches the five-field table below, with 81 positions on the left. Statistics uses the ten-paper rerun. Source review covers 35 positions; 46 are pending.

See results by field and methodology ↓

Detailed autocomplete results

The table splits positions by pre-generation source opportunity and shows completed autocomplete development evaluations across fields.

Quantitative biology

10/10 papers · 40 positions

Evidence can directly support the continuation · n = 20

Evidence version preferred

65% · 13/20

Verified direct use of the retrieved paper

45% · 9/20

Tie

5% · 1/20

No-Evidence version preferred

10% · 2/20

No directly supporting Evidence found · n = 20

Evidence version preferred

35% · 7/20

Tie

20% · 4/20

No-Evidence version preferred

25% · 5/20

Statistics

10/10 papers · 40 positions

Evidence can directly support the continuation · n = 20

Evidence version preferred

60% · 12/20

Verified direct use of the retrieved paper

Source review pending for this run

Tie

0% · 0/20

No-Evidence version preferred

35% · 7/20

No directly supporting Evidence found · n = 20

Evidence version preferred

45% · 9/20

Tie

0% · 0/20

No-Evidence version preferred

20% · 4/20

Astrophysics

10/10 papers · 30 positions

Evidence can directly support the continuation · n = 15

Evidence version preferred

60% · 9/15

Verified direct use of the retrieved paper

47% · 7/15

Tie

7% · 1/15

No-Evidence version preferred

27% · 4/15

No directly supporting Evidence found · n = 15

Evidence version preferred

33% · 5/15

Tie

7% · 1/15

No-Evidence version preferred

53% · 8/15

Psychology and cognitive neuroscience

9/10 papers · 36 positions

Evidence can directly support the continuation · n = 18

Evidence version preferred

67% · 12/18

Verified direct use of the retrieved paper

Source review pending for this run

Tie

0% · 0/18

No-Evidence version preferred

28% · 5/18

No directly supporting Evidence found · n = 18

Evidence version preferred

22% · 4/18

Tie

6% · 1/18

No-Evidence version preferred

44% · 8/18

Climate and environmental science

4/10 papers · 16 positions

Evidence can directly support the continuation · n = 8

Evidence version preferred

63% · 5/8

Verified direct use of the retrieved paper

Source review pending for this run

Tie

13% · 1/8

No-Evidence version preferred

13% · 1/8

No directly supporting Evidence found · n = 8

Evidence version preferred

50% · 4/8

Tie

0% · 0/8

No-Evidence version preferred

38% · 3/8

Notes for reading the table

  • “Evidence preferred” compares overall quality. The purple count checks whether those winning continuations directly used the retrieved paper. The overview’s source-use percentage covers only wins entered into source review; pending cases are not treated as failures.
  • Positions were labelled for direct source support before generation, then sampled equally from both groups for comparison. The split does not represent how often either kind occurs in everyday writing.
  • Positions where both continuations were unusable or judging was incomplete are not listed separately, but remain in the denominator. The three displayed percentages therefore may not add up to 100%.

REAL DEVELOPMENT CASES

Explore all 162 autocomplete comparisons

This index shows all 162 autocomplete pairs across five fields, matching the denominators above. Every position includes the draft, retrieved passages, both first outputs and model review record. Ties, both-unusable pairs and incomplete judgments remain visible.

These are development records pending independent human audit. Drafts and source passages are excerpted; output text, individual edits and verification judgments come from the original records. Failed judgments have no winner.

Result
Field

Showing 162 / 162 pairs · page 1 / 14

AutocompleteNo-Evidence version preferred2c19f4cf93463f45

Bridging from single to collective cell migration: A review of models and links to experiments

Quantitative biology · 2011.10873v1

… d cell shape and function, in both normal and diseased cells. Recent efforts aim to link single cell behavior to collective behavior of many cells and emergent dynamics of tissues.

Open comparison →
AutocompleteNo-Evidence version preferrede6bfbd7a9561cfcb

Bridging from single to collective cell migration: A review of models and links to experiments

Quantitative biology · 2011.10873v1

… , have fostered many generations of computational models. By comparison, the level of basic scientific computational research on multiscale cell biology modeling is still emerging.

Open comparison →
AutocompleteEvidence version preferredce22995352c0cdc1

Bridging from single to collective cell migration: A review of models and links to experiments

Quantitative biology · 2011.10873v1

… PDEs and patterns inside a single cell, it can then be simplified, depicted by a direction vector [ 25 , 15 ] in place of a full internal gradient of Rho or Rac for multiple cells.

Open comparison →
AutocompleteEvidence version preferredcebee9387c3b261c

Bridging from single to collective cell migration: A review of models and links to experiments

Quantitative biology · 2011.10873v1

… incorporate this information with the intracellular networks to form a signaling network on a multicellular level to study how a group of cells processes signals collaboratively.”

Open comparison →
AutocompleteBoth unusable0dcb80855f23c0b6

Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade

Quantitative biology · 2202.03534v2

… and increasing throughput of sequencing technologies have caused an exponential growth in sequencing data ( Muir2016 ) , which has now passed the peta-scale barrier ( Katz2022 ) .

Open comparison →
AutocompleteTie7dce670e0a924e05

Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade

Quantitative biology · 2202.03534v2

… or space to examine the composition of a species community as a whole, for instance as a means of sample ordination and visualization, and association with environmental variables.

Open comparison →
AutocompleteEvidence version preferreddecdba86c8631092

Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade

Quantitative biology · 2202.03534v2

… mental metadata variables. Despite the growing popularity of phylogenetic placement, there are several methodological and usage aspects that will benefit from further developments.

Open comparison →
AutocompleteTiebc30ba197adadd03

Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade

Quantitative biology · 2202.03534v2

… e case where the placed data actually contains yet undescribed species. A classification based on the LWR and pendant length of the placement locations might offer a solution here.

Open comparison →
AutocompleteEvidence version preferred199cc69dbcc484e8

Predictive Coding: a Theoretical and Experimental Review

Quantitative biology · 2107.12979v4

… ictive coding and the widely-used backpropagation of error algorithm, as well as surveying the close relationships between predictive coding and modern machine learning techniques.

Open comparison →
AutocompleteTie7192878e6a884576

Predictive Coding: a Theoretical and Experimental Review

Quantitative biology · 2107.12979v4

… g motor commands into action. Predictive coding, then, would be a model of these inbuilt controllers at the periphery rather than the core action selection mechanisms in the brain.

Open comparison →
AutocompleteEvidence version preferred27f2ba81533c87f2

Predictive Coding: a Theoretical and Experimental Review

Quantitative biology · 2107.12979v4

… lemented within the predictive coding paradigm but instead likely relies on a complex set of machinery specialised for performing reinforcement learning ( Sutton & Barto (2018) ) .

Open comparison →
AutocompleteEvidence version preferred6f54e7abd741955a

Predictive Coding: a Theoretical and Experimental Review

Quantitative biology · 2107.12979v4

… expressive generative models to handle these kinds of temporal dependencies in continuously varying inputs is an open challenge in both neuroscience as well as in machine learning.

Open comparison →
1 / 14

← Back to home