Research Writing in the Age of AI: Beware the Bubble of Beautiful but Empty Prose
AI has become very good at writing sentences that look right.
Grammar is rarely the problem anymore. The terminology often sounds appropriate. Give a model a paragraph from a manuscript and ask it to continue, and it usually knows when to introduce a mechanism, when to qualify a result, and when to connect one observation to another.
The harder problem is that some of these sentences are simply too convincing.
They read well enough that we stop asking a more basic question: how much of this sentence is actually supported by the evidence?
Suppose you are writing about Mn²⁺ incorporation and defect passivation in a perovskite system.
The experimental results are already on the page. What you need next is an explanation for the improvement in photoluminescence.
An AI model might naturally produce something like: “Mn²⁺ incorporation improves photoluminescence by suppressing defect-mediated non-radiative recombination.”
It is clean. The terminology is plausible. The causal chain is easy to follow. It would not look out of place in a Discussion section.
Then you go back to the paper that is supposed to support it.
What the paper actually shows may be narrower: Mn²⁺ incorporation is accompanied by a reduction in vacancy-related defect states, while photoluminescence also improves. The authors may suggest that these changes are related, but the experiment does not necessarily establish the full causal chain implied by the generated sentence.
The evidence may support: the improvement may be associated with reduced defect states.
It may not support: the improvement is caused by defect suppression.
Only a few words have changed.
In research writing, those few words can carry most of the scientific responsibility.
What AI is best at can also hide what is still unresolved
Before AI-assisted writing became routine, many problems announced themselves through bad prose.
If a sentence was difficult to write, you stopped.
If the mechanism was unclear, the paragraph often sounded unclear too.
If two results did not fit together, the awkwardness sometimes forced you to look at the data or literature again.
Now a model can remove that awkwardness in seconds.
That is useful. But it can also remove some of the signals that used to tell us that the thinking was unfinished.
A vague interpretation can quickly become a polished paragraph.
The transitions work. The terminology is correct. The logic feels complete. There may even be a citation attached.
Yet linguistic completeness is not the same as scientific completeness.
Sometimes we have simply made an unresolved problem sound resolved.
This is a different kind of empty prose from the old cliché of academic writing full of decorative language and little substance.
AI-generated emptiness can be much harder to spot.
It may contain field-specific terminology, a coherent argument, and a real reference. What it lacks is not the appearance of substance, but enough evidence and judgment underneath the language.
A real citation does not mean the sentence is supported
More academic AI tools now attach citations to generated claims, and that is clearly better than unsupported generation.
But a citation is a very coarse unit.
A paper can be relevant to a sentence without supporting that sentence as written.
The part that actually matters may be one experiment in the Results section, a figure, a sentence in the Discussion, or even the conditions described in a caption.
Within the same paper, several things may coexist: what the experiment directly observed, what the authors inferred from those observations, what they proposed as one possible explanation, and what remained unresolved.
When all of that is compressed into a smooth sentence, those boundaries are easy to lose.
The citation can be real. The paper can be relevant. The topic can match perfectly. And the sentence can still say more than the source allows.
So “does this sentence have a citation?” is only the beginning.
A more useful test is: if the source passage were placed in front of me right now, would I still be comfortable writing the sentence this strongly?
Some sentences should not exist yet
The default AI-writing workflow usually begins with generation.
Generate the sentence. Read it. If it seems plausible, verify it or add a citation afterwards.
For many tasks, that order works well.
Grammar correction, shortening, tone adjustment, title variants, or turning an already settled idea into a cleaner abstract are mostly problems of expression.
The researcher already knows, roughly, what they want to say. AI helps say it better.
Scientific claims are different.
Here the unresolved question may be whether the statement should be made at all.
Is this correlation or causation? Does this experiment establish the mechanism? Are the conditions in the two papers actually comparable? Should the verb be “demonstrates,” or only “suggests”? Does the conclusion apply to this system, or only to the one in the source paper?
When the complete sentence appears first, there is a subtle change in the direction of reasoning.
Instead of asking “What does the evidence allow me to conclude?”, we begin asking “I already have a sentence. Which paper can support it?”
That reversal is easy to miss, especially when the sentence is very good.
Put the evidence in front first, and the sentence often changes
Return to the Mn²⁺ example.
Imagine that before writing the explanation, you first see two short passages from papers already in your library.
One reports Mn²⁺ incorporation together with a reduction in vacancy-related defect states.
Another is much more cautious, saying that this change may contribute to reduced non-radiative loss.
Now the sentence is likely to come out differently: “The improved photoluminescence may be associated with the reduction of vacancy-related defect states following Mn²⁺ incorporation.”
It is less decisive. Perhaps less elegant. But it is closer to what the evidence actually permits.
Sometimes reading the source leads to a stronger outcome: the mechanism should not be written at all.
The two observations may occur together without the experiment being able to show that one caused the other.
In that case, the best result is not a more carefully hedged AI sentence.
It is deleting the sentence.
Research writing contains many moments like this.
Sometimes the best continuation is no continuation.
Literature can stop a sentence, not just support one
There is another possibility.
You begin by looking for evidence explaining why photoluminescence improved.
While reading the source, you notice that the authors also discuss stability under continuous excitation.
Now the question in your own paragraph starts to shift.
Perhaps the more interesting issue is not only why emission improved, but whether the same intervention affected both efficiency and stability.
The literature has not helped you finish the sentence you intended to write.
It has interrupted it.
That matters because literature retrieval is often designed as if its purpose were to justify claims that already exist.
But research does not always move in that order.
A figure can suggest a comparison you had not considered. An old mechanism can make your current explanation look too narrow. A forgotten experimental condition can weaken what looked like a clean conclusion.
Sometimes the most useful piece of evidence is the one that makes you stop.
The pause after seeing the evidence is worth keeping
There is plenty of friction in research writing that should disappear.
Searching folders. Opening the same PDFs repeatedly. Trying five versions of a keyword. Finding a paper you know you have read. Scrolling for the page where a result was reported.
AI can reduce much of that, and it should.
But once the relevant evidence is actually on screen, there is another kind of friction that I would rather keep.
The few seconds spent asking: Did this paper really demonstrate that? Are the experimental conditions comparable? Is the author reporting an observation or interpreting one? Am I saying more than the evidence says?
Those pauses are not inefficiency.
A surprising amount of scientific rigor lives inside them.
The useful role for AI is to reduce the mechanical work required to reach the evidence, not to eliminate the judgment that begins once the evidence is there.
Some academic tools are beginning to move in this direction
A few newer tools are already experimenting with ways to bring sources closer to the writing process.
Jenni allows researchers to bring PDFs and reference libraries into AI-assisted writing and citation workflows, so generated text can be grounded in sources rather than produced in isolation.
Flowing takes a somewhat different approach. It connects the current manuscript with a researcher’s own literature library, resurfacing relevant source passages as the researcher writes and allowing those passages to inform later polishing or continuation.
The implementations differ, but the underlying direction is interesting.
The next useful step for academic AI may not be to produce an even more fluent sentence.
It may be to make the relevant evidence easier to see before the sentence becomes settled.
Then the researcher can decide whether the claim should be strong, cautious, reframed, or removed entirely.
The harder part is still the claim
AI can help with the wording afterwards.
The harder part is still deciding how much the evidence actually allows us to say.
Substack
Exploring the next generation of writing tools
Follow our Substack for essays on evidence-grounded AI writing, local research libraries, and where academic tools go next. Subscribe to join the conversation.
Join on Substack