Back to blog

The Bottleneck Has Moved: What Matters When AI Can Already Write

By Jim Lau9 min read

Fluent academic prose is becoming cheap. The harder problem is deciding what should be written in the first place.

Not long ago, one of the hardest things for an AI writing tool to do was simply produce a paragraph that sounded as if it belonged in a paper.

That problem has mostly changed.

Give a strong model a paragraph from a manuscript today and it can usually continue in competent academic English. The terminology will often be right. The transition will be smooth. The sentence may even sound exactly like something the author could have written.

And yet it can still be the wrong sentence.

We ran into this repeatedly while looking at continuations from a chemistry manuscript on metal-halide perovskites. At one point, the manuscript was discussing Mn²⁺ incorporation and its effect on photoluminescence.

The argument had already narrowed to a fairly specific question: how Mn²⁺ incorporation might affect halide-vacancy-related defect states, and whether those local changes could help explain the optical behaviour observed in the system.

ChatGPT continued in a broader direction, discussing cavity engineering, defect passivation and compositional optimisation.

There was nothing obviously bad about the continuation. All of those are legitimate ideas in the broader perovskite literature. Read on its own, the paragraph sounded technically competent and would not have looked particularly strange in a general Discussion section.

The problem was that it answered a different question.

The manuscript was no longer asking, broadly, what approaches can improve optical performance in perovskites? It was already trying to understand what Mn²⁺ was doing in this particular system.

Moving back toward general optimisation strategies widened the discussion precisely when the paper needed to stay narrow.

This is the kind of failure I find more interesting than a grammatical mistake. The model has not misunderstood the vocabulary of the field. It has misunderstood why this paragraph exists here, and where the argument is trying to go next.

A sentence can be fluent, technically plausible and individually defensible, yet still pull the manuscript in the wrong direction.

When plausible language becomes abundant

Academic writing is often described as if the difficult part were turning ideas into sentences. Sometimes it is. But once a researcher is deep inside a paper, the harder decisions usually happen before the sentence appears.

A paragraph about an experimental result can move in several reasonable directions. The author might explain a mechanism, compare the result with earlier work, introduce a limitation, shift from efficiency to stability, or decide that the paragraph is already finished.

A language model can write all of those continuations.

Its ability to generate them does not tell us which one belongs there.

Generative AI has dramatically increased the supply of plausible language. A researcher no longer has to struggle to produce three possible continuations; a model can produce ten before they have finished thinking through the first one.

More choice is useful only up to a point. After that, the work becomes selection.

The researcher still has to judge whether the proposed direction fits the role of the paragraph, whether the evidence is strong enough, whether a mechanism is being overstated, and whether an apparently relevant paper really supports what the new sentence claims.

The irony is that better prose can make this harder.

If an AI produces an awkward sentence, we inspect it. If it produces a polished sentence using the right vocabulary and a familiar scientific explanation, our attention can shift very quickly from “Is this justified?” to “Does this read well?”

The Mn²⁺ example illustrates the problem. A statement about defect passivation can be scientifically sensible while still failing in several ways: the experiment may have been performed under different conditions, the source may support the general mechanism but not the strength of the claim, or the sentence may simply take the manuscript somewhere the author does not want it to go.

These are not really problems of language.

A thousand PDFs do not give you a thousand papers’ worth of working memory

There is another reason generation is becoming the wrong place to focus all our attention.

Researchers already possess far more potentially useful information than they can actively remember.

A literature library grows in a strange way. At first, almost every paper is familiar. After a few years there may be hundreds or thousands of PDFs. Many were read carefully at some point. Some contain exactly the experiment, caveat or comparison needed for the paragraph being written today.

But the researcher no longer remembers enough to retrieve them cleanly.

The memory is often incomplete: there was a paper that showed something similar under continuous excitation. Or: I remember a figure where the dopant seemed to occupy a particular site. Or simply: I’ve seen this mechanism somewhere before.

At this stage the normal search box becomes surprisingly demanding. It asks the researcher to convert a partial memory into the right terminology before the system can help.

That is frequently not how memory works.

Recognition often arrives before language.

You can fail to remember a paper’s title, author and keywords, then recognise the relevant paragraph almost instantly when you see it again.

This is why a small passage from a paper can sometimes be more useful than a perfectly ranked list of ten titles.

A snippet is not only a search result. It can be a cue.

Sometimes the reaction is straightforward: yes, that’s the paper I meant. But the more interesting case is when the resurfaced passage changes the direction of thought.

Perhaps you were looking for evidence about efficiency and a passage reminds you that the same study reported a stability effect you had forgotten. Perhaps a mechanism described in another material system suggests a comparison you had not planned to make. Perhaps the paper turns out to be less supportive than you remembered, and you decide not to write the sentence at all.

None of these outcomes looks particularly impressive in a conventional AI demo. No paragraph may have been generated.

But something useful has happened: knowledge that had fallen out of active memory has returned to the work.

Large literature collections are usually treated as storage and retrieval problems. We organise papers into folders, tag them, index full text and improve search.

Those things matter. But once the collection becomes large enough, another problem appears.

The question is no longer only whether the information exists somewhere in the library.

It is whether the right fragment can come back into the researcher’s attention at the moment when it becomes useful.

Search begins with a need that has already been articulated.

Recall does not always have to.

It can begin with the manuscript.

Context is useful only when it becomes selective

A common response to AI limitations is to add more context.

Give the model the full paper. Give it twenty papers. Increase the context window. Add notes, previous conversations and eventually the whole literature library.

There are obvious benefits to this, but a research library exposes the limit of the idea.

A thousand PDFs already contain an enormous amount of context. Almost all of it is irrelevant to the paragraph currently being written.

The interesting problem is not how much information the system can hold. It is how well it can decide what deserves to become active.

Imagine a researcher writing three sentences about a particular experimental observation. Somewhere in the library are two passages from papers read eighteen months ago that bear directly on the mechanism, another passage that challenges the obvious interpretation, and hundreds of papers sharing the same broad keywords while contributing nothing useful to this particular moment.

Putting all of them into a larger context window does not solve the problem.

Finding the few fragments worth looking at might.

This is where the distinction between generation and decision support starts to matter.

A useful system could surface those fragments while the researcher is still deciding what the paragraph means. They do not have to dictate the next sentence. One of them may support it. Another may weaken it. Another may persuade the researcher to abandon the sentence altogether.

Academic tools have traditionally treated the manuscript and the literature library as two different places. You write in one and search the other when necessary.

AI gives us an opportunity to connect them more closely, but simply connecting them is not enough.

What matters is what crosses that boundary, and when.

A paper title is sometimes enough. Often it is not.

A paragraph, a figure caption, or a few lines around a result may contain the thing that actually changes what the researcher thinks next.

Once generation is cheap, the workflow itself becomes the interesting problem

This is the part I find increasingly important.

If AI can already produce the sentence, perhaps the next question is not how to make it generate even more.

It is how to reorganise what happens before generation.

Today, a fairly normal research-writing workflow still looks something like this: write → remember something vaguely → leave the manuscript → search the library → try several queries → open PDFs → locate the passage → return to the manuscript → ask the AI → continue writing.

Each individual step is manageable.

The friction comes from having to reconstruct the connection between the manuscript and the literature every time.

A different workflow becomes possible if the manuscript itself is treated as a signal.

The paragraph already tells us quite a lot: what the researcher is discussing now, what has just been established, which concepts are active, and often what kind of evidence might matter next.

That means the researcher should not always have to stop writing and formulate a search query before the literature can become useful.

A small number of relevant passages could resurface from the researcher’s own library while the argument is still taking shape.

Not fifty recommended papers.

Not the entire library compressed into a context window.

Just a few fragments with a real chance of mattering here.

One might recover a paper the researcher was already trying to remember.

Another might reveal that the remembered evidence was weaker than expected.

A third might introduce a comparison the researcher had not been looking for at all.

At that point the workflow begins to look less like prompt → generation, and more like manuscript → relevant fragments → recognition → evidence → judgment → writing.

The AI still has an important role at the end of that chain. It can help articulate what the researcher has decided, propose a continuation, or improve the expression.

But generation no longer has to be the first interesting thing the system does.

Sometimes the most useful intervention happens before a sentence exists.

What becomes valuable then?

I would not measure the next generation of academic AI mainly by how much text it can produce.

I would pay more attention to whether it can bring forward material that matters to the local argument, rather than merely the broad topic of the manuscript.

Whether a suggestion can be traced back to the original evidence without ten minutes of detective work.

Whether literature that has fallen out of active memory can become useful again.

And whether the system leaves the final judgment where it belongs: with the researcher deciding whether two experiments are comparable, whether a mechanism is justified, whether the evidence is strong enough, and whether the manuscript should move in that direction at all.

These capabilities do not always result in more text.

Sometimes the result is a deleted sentence.

Sometimes it is ten minutes spent rereading an old paper.

Sometimes a resurfaced passage changes the framing of an entire paragraph.

Measured in words per minute, none of that looks especially productive.

But research has never worked particularly well under that metric.

AI is already becoming very good at producing the sentence.

What interests me more now is whether it can make everything immediately around that sentence easier: remembering what we have read, seeing the evidence that matters, and deciding what is actually worth writing next.

Substack

Exploring the next generation of writing tools

Follow our Substack for essays on evidence-grounded AI writing, local research libraries, and where academic tools go next. Subscribe to join the conversation.

Join on Substack

Related

Also published on Substack