← Back to all development cases
Development case · blind model judgment pending independent human review; not a formal Benchmark conclusion.
AutocompleteBoth unusable31 / 162 · 364ed65b4abf2883

From Words to Molecules: A Survey of Large Language Models in Chemistry

Quantitative biology · 2402.01439v1

FLOWING EVIDENCE BENCHMARK

How do we tell whether Evidence helps?

At the same writing position, with the same model and task, how does supplying retrieved paper passages change the first output? We compare matched versions and retain ties, unusable outputs and incomplete reviews.

SAME DRAFT · EVIDENCE ON OR OFF

01 · FIXED WRITING POSITION

MANUSCRIPT

Same writing position ▌

Matched autocomplete at the same draft position

02 · TWO MATCHED INPUTS

SHARED BY BOTH

Manuscript context, model, task and prompt

A · With Evidence

Retrieved paper passages supplied

B · Without Evidence

No retrieved passages supplied

LLM

Same model and version

A → first output

B → first output

Blind judge agent

First outputs are anonymized as X and Y

Continuation review: accuracy, fit to the writing task and usability
Returns: X preferred / tie / Y preferred / both unusable

Order check: X / Y → Y / X

HOW DOES BLIND JUDGING WORK?

① Anonymize both outputs
The judge sees the same draft and both first outputs without knowing which received Evidence.

② Compare and swap order
The judge applies task-specific criteria in X/Y and then Y/X order.

③ Review disagreements
A third pass resolves disagreements. Incomplete reviews remain in the denominator.

How are autocomplete positions stratified?

Before seeing generation outcomes, we check whether a retrieved passage contains a specific proposition that directly supports the next writing move. Those positions appear in the left opportunity group; the rest are ordinary positions on the right. We select a balanced sample from admitted papers in each field. The 50/50 split is experimental, not a measure of how often either type occurs in writing.

AUTOCOMPLETE · 110

Biology, statistics and astrophysics

55 positions on each side; statistics uses the ten-paper rerun.

AUTOCOMPLETE · 52

Psychology and climate science

26 positions on each side; psychology includes nine papers and climate science four.

How are the table percentages calculated?

Across five fields, 51 of 81 left-column positions preferred the Evidence version. Ties, pairs where both versions were unusable and incomplete reviews remain in the denominator.

51Evidence version preferred
÷
81All positions in this group
=
63%Evidence preference in this group

Source contribution is a separate review: 16 of 22 Evidence wins entered into source review directly used retrieved papers; another 29 wins await review.

These are development-stage model judgments pending independent human review. They are not formal Benchmark conclusions and do not, on their own, establish causality.

Manuscript writing position

Text excerpt · not a PDF page

Research manuscript · excerpt

From Words to Molecules: A Survey of Large Language Models in Chemistry

1 Introduction

Humans understand and describe their environment using natural language, which reflects the complexity of human thought. The emergence of Large Language Models (LLMs) marks a significant advancement in artificial intelligence, showcasing remarkable abilities in various domains. These models excel at understanding and generating complex text, making them crucial for tasks that demand deep textual analysis and creation.

…similar to how syntax operates, while molecules are formed within specific physical constraints, echoing the principles of grammar. This parallel suggests the potential for encoding chemical information into LLMs in a manner comparable to natural language. Despite the conceptual parallels, the languages of chemistry and human communication differ substantially in their semantics. Consequently, incorporating chemical knowledge into LLMs presents a complex challenge, with numerous approaches being explored to leverage LLMs in the field of chemistry, making it a subject of considerable interest.

↓ Next paragraph was generated here

The highlighted text is near the writing position. This is a layout of the evaluation input, not a PDF screenshot.

A matched comparison at the same writing positionJudgment: Both unusable

Both arms use the same model, task and manuscript context. Only the retrieved passages differ. The first recorded output from each arm is shown below.

Evidence

Same model and position · retrieved passages supplied · first output

Judged unusable

This survey reviews the emerging landscape of LLM applications in chemistry, organizing existing approaches by the chemical tasks they target and the modalities they employ. We first examine how molecular representations are adapted for language models, then survey applications spanning property prediction, retrosynthesis, and molecular generation, before discussing the open challenges that remain.

Passages supplied to the Evidence version

BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

Safety : While large-scale language models serve as a novel technique, their generated outputs are determined by a certain probability distribution, resulting in unforeseen dangers of generating bias, discrimination, or harmful content. Though we have endeavored to reduce the potential risk of BioMe…

Read full passage excerpt

Safety : While large-scale language models serve as a novel technique, their generated outputs are determined by a certain probability distribution, resulting in unforeseen dangers of generating bias, discrimination, or harmful content. Though we have endeavored to reduce the potential risk of BioMedGPT by fine-tuning on meticulously curated English biomedical corpus, it is hard to fully eliminate this problem. It is essential to ensure the responsible and ethical use of BioMedGPT. While BioMedGPT is endowed with expertise in biomedicine and chemistry, we emphasize that it should NOT be employed for research scenarios that endanger human life, and any further real-world applications should undergo cautious and professional supervision and comprehensive experiments.

No Evidence

Same model and position · no retrieved passages · first output

Judged unusable

No continuation text was produced.

Model action:complete

Blind model judgment

The blind model judge marked both outputs unusable; this is counted separately from a tie.

Read the original model judgment

Output A provides a coherent academic paragraph that continues the manuscript's flow, transitioning from the conceptual parallel between chemistry and language to a survey structure. However, it contains unsupported claims: it describes a specific survey structure ('organizing existing approaches by the chemical tasks they target and the modalities they employ,' 'molecular representations,' 'property prediction, retrosynthesis, and molecular generation') that is not grounded in any supplied source. The draft discusses general conceptual parallels and challenges but does not establish that this paper is a survey with these specific sections or that these particular chemical tasks are covered. The source provided (BioMedGPT) discusses safety concerns and biomedical/chemistry expertise but does not support any of the specific survey content in Output A. Output B is empty when the task clearly needs content (a coherent academic paragraph), making it unusable. Despite A's unsupported claims, B is completely empty and thus fails the basic task requirement. However, per instructions, any output with unsupportedClaim=true must have usable=false. Re-evaluating: Output A has unsupportedClaim=true, so usable must be false. Output B is empty when content is needed, so usable=false. This leads to both_unusable.

The passages are input material; their presence does not mean an output used them or that they support every claim. Draft and source passages are excerpted; no simulated scores or PDF appear here.