Alex Lugovskoy

The current discussion of generative artificial intelligence in academia is largely focused on the wrong question. We ask whether researchers should be allowed to use AI in writing papers, how its use should be disclosed, whether AI-generated text can be detected, and where the boundary between legitimate assistance and misconduct should be drawn.
These are understandable questions. They may also become irrelevant remarkably quickly.
Not using AI in academic work is rapidly becoming comparable to writing with a quill pen when a computer with voice input is available. Literature searches, language editing, summarization, elementary statistical analysis, generation and debugging of code, comparison of alternative interpretations, and even adversarial examination of one’s own arguments can all be performed faster and, in many cases, better with AI assistance. Insisting that researchers refrain from using these tools would preserve neither scientific integrity nor intellectual achievement. It would mostly preserve inefficiency.
The more serious problem is almost the opposite: generative AI destroys the evidential value of academic text as evidence of academic work.
The paper was never merely a paper
Traditionally, a scientific or scholarly paper performed several functions simultaneously. It communicated a result; documented that research had been performed; demonstrated the competence of its authors; provided material for peer evaluation; and served as a unit of academic reputation.
These functions were never perfectly aligned. Ghostwriting, honorary authorship, plagiarism, selective reporting and outright fabrication existed long before generative AI. Nevertheless, producing twenty pages of coherent scholarly prose normally required substantial effort. A competent literature review was therefore at least indirect evidence that somebody had searched, read, organized and understood a significant body of literature.
That inference is no longer justified. A polished twenty-page review can now be generated without the nominal author having performed anything remotely resembling the intellectual process that the resulting text appears to document. The same increasingly applies to introductions, theoretical discussions, interpretations of results and responses to reviewers.
The important distinction is therefore not between human-written and AI-written text. It is between a scholarly claim and the intellectual and empirical provenance of that claim. A literature review has not become useless. What has approached zero is the evidential value of the mere fact that somebody has written one.
Experimental science has an escape route
For an experimental scientist, the solution is conceptually straightforward. Consider a materials-science paper reporting that changing a deposition parameter modifies coating morphology and consequently changes its corrosion behaviour. Behind this statement there should exist a chain such as research question → experimental design → specimen → instrument → raw data → processing → physical interpretation → claim.
AI may legitimately assist at almost every stage of this chain. It may search literature, suggest experimental designs, write analysis code, perform statistical tests, identify alternative explanations, improve figures, challenge interpretations and produce excellent English.
None of this necessarily compromises the science. What matters is that the chain remains inspectable. If a paper claims that a microstructural change has occurred, it should ultimately be possible to trace that claim to measurements, raw instrument files, acquisition conditions, specimen identity and experimental records. The paper then becomes an interface to the research rather than the research itself.
This distinction also suggests a different role for peer review. A competent reviewer is not particularly valuable because he or she can detect an incorrect comma, locate three missing references or recalculate a routine statistical test. Machines can increasingly perform such tasks. The reviewer is valuable when able to say: No. These data do not support this conclusion. That is a much more interesting use of expert time.
There is, however, an uncomfortable complication. AI is available not only to authors and reviewers but also to fabricators. It dramatically reduces the cost of producing internally consistent fraudulent research: plausible datasets, supplementary tables, spectra, statistical analyses, methodological descriptions and responses to reviewers can all be generated and cross-checked.
Thus, experimental science is protected not because experimental papers are intrinsically trustworthy, but because their claims can, at least in principle, be connected to physical reality. The defensible object is no longer the manuscript. It is the provenance chain behind the manuscript.
What, then, is a philologist supposed to do?
The problem becomes more interesting in disciplines in which the principal research product is itself an intellectual interpretation.
Imagine a scholar writing:
In the late poetry of Mandelstam, spatial imagery undergoes a transition from topographical specificity toward…
A sufficiently capable language model can examine a corpus, consult a substantial body of secondary literature, propose several competing interpretations, identify supporting passages and counterexamples, relate the argument to previous scholarship, and produce a polished article defending one of those interpretations.
What, exactly, remains the scholar’s contribution? “Writing the paper” is no longer an adequate answer. Humanities scholarship may therefore require its own equivalent of experimental provenance. A possible chain might be source/corpus → selection procedure → observations → classification → argument → interpretation → claim.
Instead of merely asserting that a phenomenon occurs in Mandelstam, a scholar might increasingly be expected to show where it occurs, how the relevant corpus was defined, which cases were included or excluded, according to what criteria they were classified, which observations contradict the proposed interpretation, and why one interpretation is preferred over plausible alternatives.
AI may perform much of the mechanical work here as well. It may find the passages, suggest categories, search for counterexamples and formulate competing hypotheses. Again, that is not necessarily a problem. The scholarly contribution moves elsewhere: toward asking the question, defining meaningful distinctions, evaluating evidence, choosing between competing interpretations, and accepting responsibility for the resulting claim.
This works reasonably well for corpus-based linguistics, history, philology and many forms of textual scholarship. It becomes considerably more difficult in purely hermeneutic work, philosophical argument or scholarship whose principal product is precisely a novel and persuasive interpretation.
If an AI can generate an interpretation that is as original, informed and persuasive as that produced by the scholar, it becomes difficult to explain why the production of the prose itself should retain its traditional academic value. There may be no comfortable answer to this problem.
An arms race without an obvious winner
We are entering an unusual academic arms race: author + AI → editor + AI → reviewer + AI → fabricator + AI. Every participant becomes more productive, yet the signal-to-noise ratio may deteriorate rather than improve. An editor may be able to screen five hundred manuscripts with AI assistance. Authors may simultaneously become capable of submitting five thousand. Reviewers may detect inconsistencies more efficiently, while fabricators use the same technology to eliminate precisely the inconsistencies by which fraudulent papers were previously detected.
The problem is therefore not simply that AI makes misconduct easier. It makes production of the outward appearance of scholarship extraordinarily cheap.
Academic institutions cannot solve this by prohibiting AI. Detection of AI-generated prose is unlikely to provide a durable solution either. Nor does disclosure of AI assistance address the fundamental epistemic problem. A meticulously disclosed AI-assisted paper may represent excellent scholarship, while an entirely human-written paper may be incompetent or fraudulent.
The relevant question is not: Who wrote these sentences? It is rather: Why should we believe this claim?
From authorship to responsibility
Perhaps the most useful consequence of the present disruption is that it forces academia to distinguish activities that have historically been bundled together. Writing is not research. Searching literature is not necessarily research. Performing an analysis is not necessarily research. Even producing a plausible interpretation is not necessarily sufficient to constitute research. A possible working definition for the AI era is instead:
Academic research is the production of claims for which the researcher can provide an inspectable chain of empirical or intellectual justification and for which the researcher accepts intellectual responsibility.
For an experimental scientist, that chain may run from specimen to measurement to raw data to model to claim. For a philologist, it may run from source to observation to analytical criterion to interpretation to claim.
AI can participate extensively in either chain. The crucial requirement is not that every intellectual operation was performed unaided by a human being. It is that the researcher understands the chain, can defend its transitions, can expose its evidential basis to scrutiny and can identify what evidence or argument would cause the claim to be abandoned.
This last condition may prove particularly useful. If a scholar cannot describe any observation, text, experiment or argument that could make the proposed interpretation untenable, generative AI may not have created the underlying problem. It may merely have made it impossible to ignore.
Author’s note on AI-assisted composition
This commentary was developed by Alex Lugovskoy in creative collaboration with ChatGPT (OpenAI). The initial thesis, disciplinary examples, argumentative direction and substantive judgments originated in a discussion initiated by the author. ChatGPT was used as an interlocutor in developing the argument and subsequently assisted in structuring, drafting and refining the English text. The author reviewed the resulting text and assumes responsibility for the claims and arguments presented here.
This disclosure is intentionally more explicit than a conventional statement of language or editorial assistance. The manner in which this commentary was produced is itself an example of the distinction argued for above: the relevant question is not who physically generated each sentence, but who can defend the argument and accepts responsibility for it.