A. Framing the contribution: why the reader should care
Stakes are relational, not intrinsic
Section II.D described the Introduction’s first move as establishing the territory: stating why the problem area matters, broadly enough that a reader outside the immediate subfield can see the stakes. It described the slot without describing what fills it well, and the reason is that the question “why does this matter” has no answer in the abstract. It has an answer only relative to a reader. Three notions are habitually confused under the single word importance, and separating them is most of the work of this section. Importance is the claim that the problem area is large or consequential in general. Relevance is the claim that the problem bears on what a particular reader is doing. Consequence is the claim that something changes if the finding holds.
Importance is cheap, because it is nearly always true and nearly always available: any problem in natural language processing can be described as important without the description being false. This is precisely what makes it useless as a framing device — a claim that could be made with equal justice about every article in the proceedings distinguishes none of them. Relevance is not cheap, but it is also not fully under the writer’s control: it is a fact about the reader’s current work, and the most the article can do is make relevance easy to detect. Consequence is the one the article can genuinely supply, and it is the one most often left out.
The consequence test
The operative rule is a single one, and the rest of this section applies it: a stake is real when something a reader does, believes, or builds would change if the finding holds. Its usefulness lies in being checkable by negation. Suppose the article’s central finding were reversed — the method did not improve on the baseline, the diagnostic revealed nothing, the resource was never built. What would a reader do differently? If the honest answer is nothing, the framing has not located a stake, and no amount of emphatic prose in the opening paragraph will install one.
The test has a corollary that is worth stating separately, because it catches a distinct failure: the beneficiary must be nameable. “This is an important problem for the community” names no one, and a claim about the community is not a claim about a reader. “A practitioner deploying an entailment model in a setting where the inputs do not resemble the training distribution currently has no way to tell whether the reported accuracy will transfer” names someone, states what they cannot currently do, and thereby specifies exactly what the article would have to deliver to matter. The second sentence is longer than the first and does far more work, which is the usual relation between a real stake and an asserted one.
The cost of the generic opener
“Natural language processing has seen rapid progress in recent years.” The sentence is true, it is grammatical, it satisfies every rule Part III established, and it is the single most common opening sentence in the field. Its defect is not inaccuracy but that it is free to skip: it carries no information a reader of the article did not already have, and it occupies the highest-attention position the article will ever have — the first line a reader who has decided to continue past the Abstract will encounter. An article that spends that position on a proposition its readers already believe has spent it on nothing.
The repair is not to make the opening more dramatic. It is to move the specific problem forward. Section II.D warned that move one should be “already pointed toward the specific problem the article addresses rather than the field in general,” and the generic opener is exactly the failure that warning anticipates. A useful discipline is that the opening should be a sentence a reader could disagree with, or at least one they might not have thought about — a statement about what the field currently assumes, what it currently cannot measure, or what it currently treats as settled. Disagreement is a form of attention; agreement with a truism is not.
Framing is a choice among true framings
A single set of results almost always admits several accurate descriptions, and choosing among them is a rhetorical act performed on unchanged evidence. A method that scores three points higher than the strongest baseline while training in half the time can be framed as an accuracy contribution, as an efficiency contribution, or as evidence that a particular architectural assumption was unnecessary. All three descriptions are true of the same tables. They are not interchangeable, because each selects a different reader population, and with that population comes a different set of baselines the reader will expect to see, a different venue at which the article is best placed, and a different set of experiments whose absence the reader will register as a gap.
This is why framing belongs to the preparation stage described in II.A rather than to final polishing. The framing determines which experiments the article needs, and discovering that after the experiments are complete is expensive. It is also where the boundary with spin lies, and the boundary is checkable: a framing is legitimate when the article’s evidence answers the question the framing raises. An efficiency framing obliges the article to report training cost under a stated protocol, not merely to mention speed in the Introduction; an architectural-assumption framing obliges an ablation that isolates the assumption. A framing whose question the article does not answer is not a framing but a promise, and IV.E takes up what happens to unkept promises.
Matching to the structural variant
The three variants from I.B differ in where a legitimate stake is most easily found. The system/method paper most often frames on capability or cost — something that could not be done, or could not be done affordably, and now can — which is why I.B recorded its tone as confident and comparative; the risk specific to this variant is that the stake migrates from the problem to the method’s novelty, a failure treated below. The empirical/analysis paper has the highest natural stakes of the three and the least need to manufacture them, because its contribution is a revision to something the reader believes, and a revised belief is a consequence by definition. The resource/dataset paper is the hardest case, for the reason III.B already identified in the SQuAD abstract: utility is exactly what a resource paper cannot measure directly. Its framing must therefore rest on the research the resource unblocks — the question that could not previously be asked for want of data — rather than on the resource’s own properties, since size and coverage are attributes, not consequences.
A worked example
The HANS study (I.B) can be framed in at least three ways that its evidence would support, and comparing them shows how much a framing decides. Framed as a new diagnostic set, the contribution is an artifact: a controlled evaluation suite isolating three syntactic heuristics. This framing addresses readers looking for evaluation resources, and obliges the article to document construction and coverage. Framed as a critique of existing benchmarks, the contribution is a defect report about MNLI and its relatives; this addresses benchmark builders, and obliges the article to show the defect is systematic rather than incidental. Framed as a revision of what benchmark accuracy licenses — which is the framing the paper actually adopts — the contribution is a change in what a reader is entitled to conclude from a number they see in every results table in the subfield. That framing addresses nearly everyone in the subfield, and it obliges the article to show that strong models fail on the diagnostic, which is precisely what its central result does.
The third framing is the most ambitious and also the best supported, which is not a coincidence: the framing was chosen to match evidence the article could actually produce. The Transformer paper (I.B) shows the same discipline in a smaller way by carrying two stakes at once. Its abstract claims both that the model is “superior in quality” and that it requires “significantly less time to train,” and the second claim is not decoration — it names a distinct consequence for a distinct reader, the one whose constraint is compute rather than accuracy, and the article’s training-cost figures are what make that second framing legitimate under the rule above.
Common failure modes
The generic opener, discussed above, is the most frequent and the least noticed, because it feels like a proper beginning. Beyond it: importance asserted without a beneficiary, where the Introduction claims consequence for “the community” or “downstream applications” without naming a single one, and where the vagueness is usually a symptom that the consequence has not been worked out rather than that it has been compressed. A related failure locates the stake in the method rather than the problem — “existing approaches have not applied contrastive objectives to this setting” states a gap in the literature, not a reason the gap matters, and a gap that nobody had a reason to close is not a stake but an observation. And, most consequentially, there is the framing the evidence does not answer, where the Introduction promises an efficiency contribution and the article reports no timings, or promises a claim about generality and evaluates one language; this is overclaiming relocated from the sentence to the article’s architecture, and it is the subject of IV.E.
A practical check
Delete the first two sentences of the Introduction and read what remains. If nothing was lost — if the third sentence is a perfectly good opening — the deleted sentences were establishing territory that needed no establishing, and the check has recovered two sentences of the article’s most valuable space. Then perform the second half: name, in one sentence and without using the words community, field, or applications, the reader who is worse off if this article does not exist, and state what they cannot currently do. If that sentence is hard to write, the difficulty is not one of expression, and it will not be solved in the Introduction.