Research Writing: How Our Memory Gets Buried Under Thousands of PDFs
A research library is only useful if the right part of it can come back into view when you need it.
After a few years of research, something strange starts to happen to a personal literature library.
The number of papers keeps growing. The number you can readily recall does not.
When you first enter a field, the library may contain only a few dozen papers. You roughly remember what each one is about, where a particular figure came from, which experiment used which method. Finding something is rarely difficult because much of the literature is still sitting somewhere in working memory.
A few years later, there may be hundreds or thousands of PDFs. Some were read carefully. Some were saved for a single figure. Some were annotated or discussed in a group meeting. Others stayed in the library because, at the time, you had the vague feeling that this might be useful someday.
The papers are still there.
What fades is your connection to them.
When writing, the problem is often not “I’ve never seen anything about this.”
It is more like “I’m sure I’ve seen something like this somewhere.”
Or: “There was a paper that did something similar.”
Sometimes all that remains is: “I remember a figure that might be relevant to this mechanism.”
You do not remember the title, the author, the year, or even a reliable keyword.
You remember a shape.
Search works best once memory has become language
Suppose you are writing about defect passivation in a perovskite system.
You vaguely remember a paper in which doping reduced defect states and affected non-radiative recombination.
You try: dopant defect passivation perovskite. Too broad.
Then: Mn vacancy passivation perovskite. Still not right.
You open one paper. Wrong material system.
Another discusses Mn²⁺, but mainly in the context of luminescence tuning rather than the mechanism you remember.
You try “defect suppression.” Then “halide vacancy.” Perhaps an author name you think you remember.
This kind of back-and-forth is familiar to anyone with a large literature library. What consumes attention is not always the quality of the search engine. Often it is the work of translating a half-formed memory into a query the search system can understand.
Human memory does not always arrive as language.
Recognition comes before recall can be put into words.
You may spend ten minutes unable to remember the title of a paper, then see three or four lines from the middle of it and know immediately: yes. That’s the one.
A short passage can sometimes do what a list of titles cannot. It brings back not only the paper, but the reason you cared about it: the figure before it, the experimental condition, perhaps even the judgment you made when you first read it.
What gets lost is not only the paper
A literature library also accumulates things that never became proper notes.
A method you wanted to revisit. A figure that suggested a control experiment. A paragraph in the Discussion that had little to do with your project at the time, but seemed worth remembering. A result that bothered you, although you never followed the thought.
Months later, the paper itself may have disappeared from active memory while something faint remains: I think I’ve seen a similar explanation before.
This is why a large library can feel both rich and strangely inaccessible.
The files are intact. The thoughts attached to them are not.
It also explains why someone can have a thousand papers in a library and still return, while writing, to the same familiar few dozen. The rest are not necessarily less useful. They are simply harder to bring back at the moment they matter.
Sometimes you don’t know that you should be searching
There is a harder case than forgetting the right keyword.
Sometimes you are not looking for anything at all.
You are writing.
Perhaps the Results section is moving into Discussion. You are explaining an improvement in photoluminescence and naturally following the argument toward defect passivation.
Then you come across an old passage mentioning that the same intervention also affected stability under continuous excitation.
The paragraph may now need to go somewhere else.
What began as “Why did the emission improve?” may turn into “Did this intervention affect both efficiency and stability? Is that relationship more interesting than the original framing?”
The paper has not answered the question you were asking.
It has changed the question.
This happens often enough in real research that I am wary of treating literature retrieval only as a way to support claims that have already been formed. A forgotten condition can weaken a conclusion. A figure can suggest a comparison. A mechanism from another system can make the current explanation look too narrow.
Sometimes the literature is useful precisely because it interrupts the direction you were already taking.
A different way for the library to enter the writing process
This is also why I am not convinced that the answer to a large literature library is simply more recommendations.
Researchers rarely suffer from a shortage of papers.
Sometimes what matters more is three forgotten lines from a paper they already read six months ago.
Imagine keeping the manuscript in front of you while a small number of passages from your own library surface alongside it. Not fifty recommended papers, and not the entire library compressed into a context window. Just a few snippets with a plausible connection to the paragraph you are currently writing.
One gets ignored.
Another confirms that the experimental conditions really are comparable.
A third looks familiar enough to make you stop and open the original PDF.
The manuscript already provides a surprisingly rich signal: what you are discussing, what has just been established, whether the paragraph is explaining a mechanism or making a comparison, and what kind of evidence may matter next. In some situations, that is more informative than a keyword you invent after leaving the writing flow.
The usual sequence is something like: remember something vaguely → stop writing → invent keywords → search → open PDFs → find the passage → return to the manuscript.
I would rather have the option of: keep writing → a few relevant fragments resurface → recognise / remember → judge → inspect the source if needed → decide what to write next.
Search still matters. If you know the author, title, method or material system, a search box is exactly what you want.
This other mode is useful when the memory is incomplete — or when you have not yet realised that an old paper may matter at all.
There are already tools moving toward parts of this workflow. Jenni lets researchers bring PDFs and sources from tools such as Zotero or Mendeley into the writing process for citation, questions and AI-assisted drafting. Flowing takes a more passage-oriented approach, resurfacing snippets from a researcher’s own library according to the manuscript context while they write.
What interests me is less which implementation wins than the shift in the workflow itself.
The library should not wait for the right keyword
A personal literature library should not have to wait until I remember the right keyword before it becomes useful.
Sometimes the most valuable thing it can do is return a fragment I had forgotten — at exactly the point where it changes what I was about to write next.
Substack
Exploring the next generation of writing tools
Follow our Substack for essays on evidence-grounded AI writing, local research libraries, and where academic tools go next. Subscribe to join the conversation.
Join on Substack