← Back to all development cases
Development case · blind model judgment pending independent human review; not a formal Benchmark conclusion.
AutocompleteTie10 / 162 · 7192878e6a884576

Predictive Coding: a Theoretical and Experimental Review

Quantitative biology · 2107.12979v4

FLOWING EVIDENCE BENCHMARK

How do we tell whether Evidence helps?

At the same writing position, with the same model and task, how does supplying retrieved paper passages change the first output? We compare matched versions and retain ties, unusable outputs and incomplete reviews.

SAME DRAFT · EVIDENCE ON OR OFF

01 · FIXED WRITING POSITION

MANUSCRIPT

Same writing position ▌

Matched autocomplete at the same draft position

02 · TWO MATCHED INPUTS

SHARED BY BOTH

Manuscript context, model, task and prompt

A · With Evidence

Retrieved paper passages supplied

B · Without Evidence

No retrieved passages supplied

LLM

Same model and version

A → first output

B → first output

Blind judge agent

First outputs are anonymized as X and Y

Continuation review: accuracy, fit to the writing task and usability
Returns: X preferred / tie / Y preferred / both unusable

Order check: X / Y → Y / X

HOW DOES BLIND JUDGING WORK?

① Anonymize both outputs
The judge sees the same draft and both first outputs without knowing which received Evidence.

② Compare and swap order
The judge applies task-specific criteria in X/Y and then Y/X order.

③ Review disagreements
A third pass resolves disagreements. Incomplete reviews remain in the denominator.

How are autocomplete positions stratified?

Before seeing generation outcomes, we check whether a retrieved passage contains a specific proposition that directly supports the next writing move. Those positions appear in the left opportunity group; the rest are ordinary positions on the right. We select a balanced sample from admitted papers in each field. The 50/50 split is experimental, not a measure of how often either type occurs in writing.

AUTOCOMPLETE · 110

Biology, statistics and astrophysics

55 positions on each side; statistics uses the ten-paper rerun.

AUTOCOMPLETE · 52

Psychology and climate science

26 positions on each side; psychology includes nine papers and climate science four.

How are the table percentages calculated?

Across five fields, 51 of 81 left-column positions preferred the Evidence version. Ties, pairs where both versions were unusable and incomplete reviews remain in the denominator.

51Evidence version preferred
÷
81All positions in this group
=
63%Evidence preference in this group

Source contribution is a separate review: 16 of 22 Evidence wins entered into source review directly used retrieved papers; another 29 wins await review.

These are development-stage model judgments pending independent human review. They are not formal Benchmark conclusions and do not, on their own, establish causality.

Manuscript writing position

Text excerpt · not a PDF page

Research manuscript · excerpt

Predictive Coding: a Theoretical and Experimental Review

5 Discussion and Future Directions

…e ( A.K. Seth et al. (2012) ; A.K. Seth & Critchley (2013) ; A.K. Seth (2013) ) . The key advantage of predictive coding is that it offers a well-tested and mathematically principled framework in which to model these phenomena. These advantages may allow the modelling of considerably more complex interoceptive feedback loops than can be straightforwardly interpreted in terms of feedback control ( Barrett & Simmons (2015) ; Pezzulo et al. (2018) ; Tschantz et al. (2021) ) , and especially allows for the relatively straightforward construction of nested and hierarchical feedback control models.

…tive of a similar scheme widely used in robotics, whereby complex planning or reinforcement learning algorithms determine the desired forces to send to the actuators, however the actuators themselves are equipped with onboard PID controllers, to ensure that the expected force is actually produced by the motors, and to correct for minor deviations ( Johnson & Moradi (2005) ) occuring at the lower levels of translating motor commands into action. Predictive coding, then, would be a model of these inbuilt controllers at the periphery rather than the core action selection mechanisms in the brain.

↓ Next paragraph was generated here

The highlighted text is near the writing position. This is a layout of the evaluation input, not a PDF screenshot.

A matched comparison at the same writing positionJudgment: Tie

Both arms use the same model, task and manuscript context. Only the retrieved passages differ. The first recorded output from each arm is shown below.

Evidence

Same model and position · retrieved passages supplied · first output

No continuation text was produced.

The model requested more evidence and did not continue.

Passages supplied to the Evidence version

Predictive Processing in Cognitive Robotics: a Review

However, as the authors point out, a limitation of the architecture is that it only uses the bottom-up flow of information, without integrating top-down predictions to modulate the activation of lower levels. Tani, (2019) provides a thorough review of related neurorobotics experiments, many of which…

Read full passage excerpt

However, as the authors point out, a limitation of the architecture is that it only uses the bottom-up flow of information, without integrating top-down predictions to modulate the activation of lower levels. Tani, (2019) provides a thorough review of related neurorobotics experiments, many of which carried out in the authors’ laboratory. A very interesting implementation is described in Hwang et al., (2018) , which the authors refer to as a predictive coding model. The adopted network is a multi-layer hierarchical architecture encoding visual and proprioceptive information. Although the work is far from the formulations laid in the free-energy principle ( Friston,, 2009 ) , the VMDNN (Predictive Visuo-Motor Deep Dynamic Neural Network) performs very similar operations.

No Evidence

Same model and position · no retrieved passages · first output

No continuation text was produced.

The model requested more evidence and did not continue.

Blind model judgment

The blind model judge did not prefer either output.

Read the original model judgment

Both outputs are identical and deliberately empty with 'needs_evidence' action. The task requires continuing a paragraph about predictive coding as a model of peripheral controllers (like PID controllers in robotics) versus core action selection mechanisms. The single supplied source discusses predictive processing in cognitive robotics but focuses on visuo-motor deep dynamic neural networks, bottom-up vs top-down information flow, and neurorobotics experiments. It does not directly support claims about: (1) predictive coding as a model of peripheral/inbuilt controllers specifically, (2) the distinction between peripheral controllers and core action selection mechanisms in the brain, or (3) the PID controller analogy for interoceptive feedback. The source mentions hierarchical architectures and predictive coding in robotics, but the specific analogy being developed in the draft (predictive coding as peripheral controller vs. central action selection) is not directly supported by this source. A deliberate empty response with 'needs_evidence' is appropriate when no supplied source directly supports the content needed. Both outputs correctly identify this gap.

The passages are input material; their presence does not mean an output used them or that they support every claim. Draft and source passages are excerpted; no simulated scores or PDF appear here.