E. Balancing rigor and excitement without overselling

The tension this course set up in its own rationale

The opening of Introduction.md observed that the quality of research and the quality of its written account are logically independent, and that a well-constructed article can lend undue persuasiveness to weak results. Sections IV.A through IV.D have spent four sections teaching devices that make an article more persuasive: framings that install stakes, arcs that hold a question open, headings and questions that direct attention, figures that deliver a claim in one look. Every one of these works on a reader whose result is weak exactly as well as on a reader whose result is strong. This section states the constraint that keeps them honest, and it is not an appendix to Part IV but the condition under which the rest of Part IV is teachable at all.

Every engagement device makes a promise

The constraint can be stated as a single rule, and it is III.B’s modality rule lifted from the sentence to the device: each engagement device makes a promise to the reader, and overselling is a promise the article does not keep. The promises are specific and can be enumerated. A framing promises that the stake it names is one the article’s evidence bears on. A narrative promises that the question it opens is the question the article closes. A question promises an answer. A heading promises that what it names is what the section contains. A figure promises that the pattern it shows holds as widely as it appears to — across seeds, across scales, and between the plotted points.

Each engagement device makes a specific promise, kept by a specific part of the article; a promise with no keeper is overselling.

The rule’s value is that it makes overselling checkable rather than a matter of tone. A writer asking whether an Introduction is too enthusiastic has no way to answer; a writer asking which table keeps the promise the Introduction’s third sentence makes has a question with a determinate answer. And it locates the failure correctly: an emphatic article whose promises are all kept is not overselling, and a flatly written one that promises generality it never tests is, however modest its adjectives.

Who catches an unkept promise

The costs of overselling are distributed unevenly across readers, which is why the practice persists. A reader skimming the abstract and the first figure will not detect a broken promise, because detecting one requires reading the section that was supposed to keep it. Reviewers do exactly that, and so do the readers who know the area best — the ones who will decide whether the article is cited, built on, or quietly regarded as unreliable. Section III.B made the same observation about overstatements concerning prior work: the readers most likely to notice are the ones who know that work well. Overselling therefore buys attention from the readers whose attention matters least and spends credibility with the readers whose judgment matters most, and it does so on a delay long enough that the trade is easy to miss.

Excitement is a property of the finding, not of the prose

The most useful reframing available here is that a writer cannot add interest to a finding. What a writer can do is remove the obstacles between the reader and whatever interest the finding actually has — an unframed stake, an uninstalled expectation, a heading that hides the content, a caption that labels rather than claims. Every device in IV.A through IV.D is an obstacle-removal device, and none is an interest-generation device, though several can be misused as though they were.

The practical consequence concerns modest results, where the temptation is strongest. A precisely framed modest finding is more engaging than an inflated one, and this is not a consolation but an observation about how readers work: a reader can locate a precise claim in the space of things they already know, decide whether it bears on their work, and remember it. An inflated claim resists all three operations, because a reader who cannot tell how big a claim actually is cannot tell whether it applies to them. Inflation makes a finding harder to use, and a finding that is hard to use is not exciting.

The hard cases

Three kinds of contribution are genuinely difficult to frame without either underselling or overselling, and each has a characteristic honest framing. A negative result has as its stake the belief it corrects, which means the framing work is the expectation-installation IV.B described, done carefully enough that the reader recognizes the belief as one they held — a negative result framed as an absence of finding is unpublishable, and framed as a correction to a widely held expectation is often the more valuable article. An incremental improvement has as its stake the condition under which the increment holds: not that the method is better, but that it is better under a constraint the reader may be operating under, at a cost the reader may be able to pay. And a resource paper faces the problem III.B identified in the SQuAD abstract, where the central value claim — that the resource is useful — is precisely what cannot be measured at publication time; the honest form is SQuAD’s own, marking the utility claim as an inference rather than a result, and framing on the question the resource makes askable.

Limitations is what makes an ambitious framing survivable

There is an apparent conflict between an engaging framing and a well-derived Limitations section, and it dissolves once the two are seen as parts of the same claim. Section I.C described Limitations as bounding the claim by scope, and II.H insisted the conditions be derived from the protocol rather than collected as caveats. A framing states how far a claim reaches; Limitations states where it stops. Without the second, the first is unbounded, and an unbounded claim is one a reader cannot evaluate and a reviewer must treat sceptically. A framing may be as strong as its scoping permits — which means that a well-written Limitations section does not constrain an ambitious framing but licenses it, and that the writer who wants to claim more should be writing a more precise Limitations section rather than a vaguer one.

Matching to the structural variant

Each of I.B’s variants is exposed at a different point, and the exposure follows from where its central claim lives. The system/method paper is most exposed through its comparative claim, as III.F noted when discussing intensifiers: the promise that a margin is meaningful is the easiest of all promises to make in an Introduction and the hardest to keep in a table, particularly where variance across seeds is of the same order as the reported gain. The empirical/analysis paper is most exposed through capability language, III.B’s third form of overclaiming, and the exposure is structural rather than accidental — the most engaging framing of a behavioral result is nearly always the capability framing, so the framing that best satisfies IV.A is often the one that most violates III.B, and the writer feels the pull of both at once. The resource/dataset paper is most exposed through the utility claim, since its stake and its unmeasurable quantity are the same thing.

A worked example

Consider one finding written at three levels of framing — the invented syntactic-diversity result, again. Flat: “We fine-tune on a syntactically diverse corpus and observe a 4.1-point improvement on the benchmark.” Well-framed: “Fine-tuning data is ordinarily selected for domain match; we find that syntactic diversity yields a 4.1-point gain over a domain-matched corpus of equal size, suggesting that variety of structure contributes to fine-tuning value independently of domain resemblance.” Oversold: “We show that syntactic diversity, not domain match, is what makes fine-tuning data effective, and that models trained on diverse data acquire more robust representations of grammatical structure.”

The promises are what separate them. The flat version promises only a number and keeps it, which is why it is not dishonest but is also not engaging: it gives the reader nothing to do with the number. The well-framed version promises a comparison against a domain-matched corpus of equal size, and a contribution independent of domain resemblance — both promises the article can keep, the first by its experimental design and the second at the modality suggesting rather than as a demonstration. The oversold version makes two promises the article cannot keep: that diversity rather than domain match is what makes fine-tuning effective, which requires ruling out domain match as a factor rather than comparing against one instance of it, and that the models acquire more robust representations, which is III.B’s capability language attached to a behavioral measurement. Nothing in the evidence changed across the three.

Tenney et al. is the instructive real borderline. Its title asserts that BERT “rediscovers the classical NLP pipeline,” which is a strong and memorable framing, and the abstract’s own language is notably more careful than the title’s: it claims that the model “represents the steps of the traditional NLP pipeline in an interpretable and localizable way,” that the regions “appear in the expected sequence,” and that the model “can and often does adjust this pipeline dynamically.” The title makes a promise the abstract immediately scopes, and the article’s per-layer evidence is what keeps it. This is the pattern the rule above predicts: an ambitious framing is survivable exactly to the extent that the scoping around it is done well, and a title carrying more force than its abstract is not a failure so long as the abstract does the bounding within a few lines of it.

Common failure modes

The unkept framing promise, discussed throughout, is the general case, and the specific ones are worth naming. The Abstract that claims generality the Experiments never test is the most common, and it is the more damaging for appearing in the component every reader reads. The contribution list that enumerates four contributions where the article supports two, with the remaining two being descriptions of work performed rather than claims established. The figure whose visual assertion exceeds its evidence, treated in IV.D. The Discussion that treats a plausible interpretation as a finding, which III.B’s four strengths already covers but which engagement pressure specifically encourages, since a strength-3 statement written at strength 1 is invariably the more exciting sentence. And, distinct from all of these, the article whose framing and Limitations section contradict one another — an ambitious claim in the Introduction, a Limitations section that quietly withdraws it — which is not a compromise between the two but a failure of both, since the reader who reads both learns that the article knew.

A practical check

List every promise the article makes in its Title, Abstract, first figure, and Introduction — every stake asserted, every question opened, every claim of generality, every comparison implied. Against each, name the specific table, figure, or section that keeps it. Promises with no keeper are the article’s overselling, enumerated exhaustively and without any judgment of tone being required. The check is deliberately mechanical, and it is placed last in Part IV because it audits everything the four preceding sections teach: each device those sections describe appears in this list as a promise, and the article is engaging and honest at once precisely when the list has no unmatched entries.


This site uses Just the Docs, a documentation theme for Jekyll.