← Back to all development cases
Development case · blind model judgment pending independent human review; not a formal Benchmark conclusion.
AutocompleteEvidence version preferred87 / 162 · 47b26df9e7aa733b

White Paper on Light Sterile Neutrino Searches and Related Phenomenology

Astrophysics · 2203.07323v3

FLOWING EVIDENCE BENCHMARK

How do we tell whether Evidence helps?

At the same writing position, with the same model and task, how does supplying retrieved paper passages change the first output? We compare matched versions and retain ties, unusable outputs and incomplete reviews.

SAME DRAFT · EVIDENCE ON OR OFF

01 · FIXED WRITING POSITION

MANUSCRIPT

Same writing position ▌

Matched autocomplete at the same draft position

02 · TWO MATCHED INPUTS

SHARED BY BOTH

Manuscript context, model, task and prompt

A · With Evidence

Retrieved paper passages supplied

B · Without Evidence

No retrieved passages supplied

LLM

Same model and version

A → first output

B → first output

Blind judge agent

First outputs are anonymized as X and Y

Continuation review: accuracy, fit to the writing task and usability
Returns: X preferred / tie / Y preferred / both unusable

Order check: X / Y → Y / X

HOW DOES BLIND JUDGING WORK?

① Anonymize both outputs
The judge sees the same draft and both first outputs without knowing which received Evidence.

② Compare and swap order
The judge applies task-specific criteria in X/Y and then Y/X order.

③ Review disagreements
A third pass resolves disagreements. Incomplete reviews remain in the denominator.

How are autocomplete positions stratified?

Before seeing generation outcomes, we check whether a retrieved passage contains a specific proposition that directly supports the next writing move. Those positions appear in the left opportunity group; the rest are ordinary positions on the right. We select a balanced sample from admitted papers in each field. The 50/50 split is experimental, not a measure of how often either type occurs in writing.

AUTOCOMPLETE · 110

Biology, statistics and astrophysics

55 positions on each side; statistics uses the ten-paper rerun.

AUTOCOMPLETE · 52

Psychology and climate science

26 positions on each side; psychology includes nine papers and climate science four.

How are the table percentages calculated?

Across five fields there are 81 source-opportunity positions. The original blind review preferred Evidence in 51; a task check moved one empty Evidence continuation to both unusable, leaving 50 in public counts. In the ordinary group, another pair of empty outputs moved from no winner to both unusable. Original verdicts remain visible on case pages.

50Evidence version preferred
÷
81All positions in this group
=
62%Evidence preference in this group

Source contribution is a separate review: 16 of 21 Evidence wins entered into source review directly used retrieved papers; another 29 wins await review.

These are development-stage model judgments pending independent human review. They are not formal Benchmark conclusions and do not, on their own, establish causality.

Manuscript writing position

Text excerpt · not a PDF page

Research manuscript · excerpt

White Paper on Light Sterile Neutrino Searches and Related Phenomenology

3 Interpretations of the Anomalies > 3.4 Summary of Interpretations

…ns the indication in favor of a short-baseline disappearance of electron neutrinos, possibly due to neutrino oscillations. Figure 40 shows the 90% contours in the | U e ​ 4 | 2 |U_{e4}|^{2} – Δ ​ m 41 2 \Delta{m}^{2}_{41} plane obtained from the analysis of the measured and expected Ge 71 {}^{71}\rm{Ge} event rates, considering the neutrino survival probability Eq. ( 7 ), where one can see that the squared-mass difference Δ ​ m 41 2 \Delta{m}^{2}_{41} allowed values are ∼ 1 \sim 1 eV 2 or larger, for the different approaches to compute the ν e − Ga 71 \nu_{e}\,-\,{}^{71}\rm{Ga} cross section.

…+1 neutrino oscillation hypothesis with the Reactor Antineutrino Anomaly and with the inclusion of data from tritium experiments and from experiments measuring solar neutrinos [Giunti:2022btk]. Remarkably, it was found that the Gallium Anomaly is in strong tension with bounds obtained from the other data sets. In addition, when all data are combined, the corresponding parameter goodness of fit is below 0.042%, implying a tension of ∼ 5 ​ σ \sim 5\sigma , leading to the conclusion that it should be necessary to seek for alternative solutions to the short-baseline oscillations for this anomaly.

↓ Next paragraph was generated here

The highlighted text is near the writing position. This is a layout of the evaluation input, not a PDF screenshot.

A matched comparison at the same writing positionJudgment: Evidence version preferred

Both arms use the same model, task and manuscript context. Only the retrieved passages differ. The first recorded output from each arm is shown below.

Evidence

Same model and position · retrieved passages supplied · first output

Preferred by blind review

The 3+2 neutrino mixing framework reduces the tension in the interpretation of the data, though at the cost of introducing a second sterile neutrino, and the improvement in the parameter goodness of fit is largely a statistical effect of the increased number of parameters.

Passages supplied to the Evidence version

3+1 and 3+2 Sterile Neutrino Fits ↗

In spite of the presence of a tension in the interpretation of the data, 3+1 neutrino mixing is attractive for its simplicity and for the natural correspondence of one new entity (a sterile neutrino) with a new effect (short-baseline oscillations). The allowed regions in the oscillation parameter sp…

Read full passage excerpt

In spite of the presence of a tension in the interpretation of the data, 3+1 neutrino mixing is attractive for its simplicity and for the natural correspondence of one new entity (a sterile neutrino) with a new effect (short-baseline oscillations). The allowed regions in the oscillation parameter space can be tested in near-future experiments. In the framework of 3+2 neutrino mixing there is less tension in the interpretation of the data, at the price of introducing a second sterile neutrino. Moreover, the improvement of the parameter goodness of fit is mainly a statistical effect due to an increase of the number of parameters.

No Evidence

Same model and position · no retrieved passages · first output

Not preferred in this comparison

The product produced no insertable continuation text.

The model generated prose, but the product admission rule blocked it as needing evidence. That raw prose has not been independently reviewed here.

Judgment and task check

The blinded model review preferred the Evidence version. Its original rationale appears below.

Read the original model judgment

In this rationale, Output A = Evidence; Output B = No Evidence.

Output A provides a coherent, source-grounded continuation that directly addresses the tension problem raised in the draft's conclusion about seeking 'alternative solutions to the short-baseline oscillations for this anomaly.' The content is fully supported by the source snippet, which explicitly states that 3+2 mixing 'reduces tension in the interpretation of the data' with the caveat that this comes 'at the price of introducing a second sterile neutrino' and that 'the improvement of the parameter goodness of fit is mainly a statistical effect due to an increase of the number of parameters.' Output A captures all these elements with appropriate academic phrasing. Output B returns an empty string with action 'needs_evidence,' which is not a valid completion for a task that clearly needs content; the source directly supports the relevant content, so an empty response is not justified as 'no supplied source directly supports a card'—this is an autocomplete task, not an evidence card task.

Source review confirmed that this continuation directly used specific, verifiable information from the retrieved paper.

The passages are input material; their presence does not mean an output used them or that they support every claim. Draft and source passages are excerpted; no simulated scores or PDF appear here.