B. Narrative structure: what storytelling transfers, and what it does not

Narrative is not embellishment

Advice to “tell a story” is among the most commonly given and most commonly misapplied suggestions in scientific writing, and the misapplication follows directly from the loose sense of the word. In the loose sense, a story is something added: colour, drama, a human angle, a sense of journey. Nothing in that sense transfers to an article, and attempts to import it produce the failures catalogued at the end of this section. One property does transfer, and it is structural rather than decorative: tension and resolution — a question the reader wants answered, held open long enough to be felt as open, and then closed. That property is worth isolating precisely because it is not about liveliness. A completely plain article can have it, and a vividly written one can lack it.

The four moves of II.D are already an arc

The narrative frame does not introduce a new structure, and it would be a misreading of this section to conclude that an article needs one. Section II.D’s four moves — establish the territory, narrow to the problem and open the gap, state the approach, preview the contribution — are already the shape under discussion: a situation, a complication in it, a response to the complication, and an outcome. What the narrative frame adds is a criterion for judging whether those moves are working, which II.D’s structural description could not supply. A structurally complete Introduction can execute all four moves and still leave a reader unmoved, and the diagnosis in such cases is nearly always that move two named a gap without making the reader feel it as one. Structure guarantees the slot; tension is what fills it.

The Introduction's four moves read as a rising tension arc: an expectation installed, held open, and resolved beyond it.

Tension requires an expectation the reader already holds

An expectation cannot be violated unless it has first been installed, and this is the single most useful consequence of the narrative frame. A finding is surprising only relative to a belief, and if the article does not state the belief, the finding arrives as an isolated fact. The HANS study is instructive here for a reason distinct from its modality discipline (III.B): before reporting that strong models fail on its diagnostic set, it makes explicit the belief the field held — that high accuracy on standard natural language inference benchmarks reflects genuine entailment reasoning. The reader is given the expectation, and the result then lands against it. Had the article opened directly with its diagnostic construction, the same numbers would have carried far less force, not because they would have been less valid but because the reader would have had nothing for them to contradict.

The installation must be honest, which distinguishes this from manufactured suspense. The belief stated must be one the field actually holds, stated in the form the field actually holds it, and an article that installs a straw expectation in order to overturn it is doing to prior work what III.B described as overclaiming pointed outward.

One question, held open

Tension requires not only an expectation but a single through-line. An article that opens three questions and answers them in sequence has three short arcs and no long one, and reads as a report of activity rather than as an argument. This is the same requirement II.A imposed from a different direction when it asked for the discourse to be drafted as a single section-free paragraph: a discourse that cannot be compressed into one paragraph without becoming a list is usually a discourse with no through-line. Where an article genuinely contains several independent findings, the honest response is to find the question they jointly answer, or to accept that the article is a report and forgo the arc rather than impose one — an imposed arc is detectable, and it costs more credibility than a plainly organized report ever would.

The article’s narrative is the reader’s path to conviction, not the researcher’s path to the result

This is the rule that separates the useful sense of narrative from the harmful one, and nearly every misapplication of storytelling advice violates it. The research had a chronology: an initial idea, a pilot that failed, a pivot, an unexpected observation, a final experiment. The article has an order too, but it is not that one. The article is ordered by what the reader needs in order to be convinced, which is a dependency ordering, not a temporal one. A reader does not need to know that the first architecture was abandoned; they need to know why the reported architecture is the right one to evaluate. Process narration — “we first attempted X, which did not work, so we then tried Y” — feels like storytelling and is in fact the substitution of the researcher’s chronology for the reader’s dependency chain, which is why it so reliably makes an article harder to follow rather than easier.

The exception proves the rule. Where the chronology is the dependency chain, chronological order is correct, and I.B recorded exactly this in the resource/dataset paper’s Data Collection section, whose procedure is described in roughly the order the resource was built. There the reader’s understanding of step three genuinely depends on step two, and the temporal order and the logical order coincide.

Resolution must resolve the question that was opened

An arc that opens one question and closes another has not resolved anything, and this failure is already familiar from a different vocabulary: I.C identified a visible mismatch between Introduction and Conclusion as one of the more reliable signals that an article’s central argument shifted during writing without the earlier component being revised. Under the narrative frame, the same mismatch is a broken arc. Section II.G’s first Conclusion move — restate the central claim, now as an established result — is the operation that closes the loop, and it is checkable in the most literal way available anywhere in the course: the question stated in move two of the Introduction and the claim restated in move one of the Conclusion should be recognizably the same object.

Where narrative is a defect

Narrative does not belong everywhere in an article, and this is not a matter of taste but of the component’s function under I.C. The Method’s function is reconstruction: it must let a competent reader understand and in principle reproduce what was done. Reconstruction is served by predictability, and a Method that withholds, builds, or surprises is failing at its job while succeeding at something the writer mistook for its job. A reader working through a Method should be able to anticipate what comes next; where they cannot, the ordering is wrong.

The Abstract is the sharper case, and the rule there is the exact inverse of narrative practice elsewhere. The Abstract’s function is self-selection (I.C, II.C): it must let a reader decide relevance without reading further, and for the large majority of readers who see nothing else, it is the article. An Abstract that withholds the finding to preserve interest has failed at its only function, and has failed it most severely for the readers who most needed it to work. Whatever “no spoilers” means elsewhere, in a scientific abstract the finding must be spoiled, stated plainly and early, and the tension the article carries lives in the Introduction, where a reader has already committed.

Matching to the structural variant

The empirical/analysis paper is the natively narrative variant, since a research question held open and then answered is its structure rather than an overlay on it, and III.E already noted that its prose “stays closer to a narrative connecting each sub-finding back to the motivating question.” The system/method paper has a weaker native arc — problem, solution, measurement — and the characteristic failure specific to it is the novelty parade, in which the enumerated contributions list I.B described as a table of contents becomes the article’s only through-line, and the reader is offered a sequence of new things rather than a question. The resource/dataset paper has the weakest arc of the three, because its contribution is an object rather than an answer; its tension, where it has any, must be built from the need the resource meets, and it is the one variant with a section where chronological order is correct.

A worked example

Take the invented scenario that has recurred since III.B: a model fine-tuned on a syntactically diverse corpus scores four points higher on one benchmark, in one language, at one scale, than the same model fine-tuned on a standard corpus.

Before: “Data selection for fine-tuning has received increasing attention. In this work, we investigate the effect of syntactic diversity in fine-tuning corpora. We construct a diverse corpus, fine-tune on it, and compare against a standard corpus, observing an improvement of 4.1 points.” After: “Fine-tuning corpora are ordinarily selected for domain match, on the assumption that resemblance to the target task is what makes fine-tuning data useful. Syntactic diversity cuts against that assumption: a corpus can be syntactically varied and domain-mismatched at once, and it is not obvious which property matters more. Fine-tuning on a syntactically diverse corpus improves benchmark accuracy by 4.1 points over a domain-matched standard corpus.”

The rewrite adds no evidence and one claim. What it buys is an installed expectation — that domain match is the operative property — against which the same four points now register as a correction rather than as a measurement. It also converts a description of activity (“we investigate,” “we construct,” “we compare”) into a statement of the question at issue, which is the difference between reporting what was done and stating what is at stake in the doing. The Before version is not badly written; it is arcless, and no sentence-level revision would fix that.

Common failure modes

Process narration, discussed above, is the most common and is usually defended as transparency, though what it transmits is the researcher’s chronology rather than the reader’s dependency chain. The mystery abstract is rarer but more damaging per instance, since it defeats the one component whose function every reader depends on. False drama is the sentence-level residue of the same impulse: “Surprisingly,” attached to a result nobody expected to go the other way, “Interestingly,” attached to a result whose interest the sentence does not demonstrate — both are intensifiers in the sense III.F catalogued, asserting a reaction the reader is meant to have instead of supplying grounds for it. Beyond these: an arc imposed on findings that are genuinely independent, which reads as forced precisely because the reader can feel the connective tissue bearing no load. And, most consequentially, the narrated discovery that conceals a post-hoc question — “we were surprised to find,” where the finding in fact came first and the question was constructed afterwards to frame it. That is a narrative device with an ethical dimension rather than only a rhetorical one, and it is taken up under ethical writing in Part V.

A practical check

State the article in one sentence of the form: X was believed, or needed; this article shows, or provides, Y; therefore Z. All three slots must be fillable from the article as drafted. An empty first slot means there is no installed expectation and the finding will arrive unattached. An empty third slot means there is no resolution, only a result — which is IV.A’s consequence test arriving from the other direction, and the two checks failing together is the usual case rather than the exception.


This site uses Just the Docs, a documentation theme for Jekyll.