C. Active vs. passive voice, and when each serves the argument

Neither rule is right

Two blanket rules circulate, and both are wrong in the same way. “Never use the passive” is style-guide advice imported from general expository prose, where the agent is usually the point; “always use the passive, science is impersonal” is a convention inherited from an older tradition that took agentlessness as a marker of objectivity. Each prescribes a form without reference to function, which is exactly the move Introduction.md set the course against. The functional test is available and simple: the subject position of a sentence is its most prominent slot, and the right voice is the one that puts there the element the reader needs to be tracking. Voice is a device for controlling what a sentence is about, not a marker of rigor.

Topic position and the attribution of agency

Two considerations follow from that test. The first is topical continuity: if consecutive sentences are about a dataset, keeping the dataset in subject position — which usually means the passive — lets a reader follow one thread, whereas alternating between “we filtered the data” and “the data then passed through” forces the reader to re-anchor at every sentence. This is a cohesion effect, and III.E returns to it. The second is agency attribution: the passive can omit the agent entirely, and whether that omission is appropriate depends on whether the identity of the agent is part of what the reader needs. When the agent is obvious or irrelevant — the authors did it, as they did everything else in the Method — omitting it costs nothing. When the agent is the claim, omitting it destroys the sentence’s point.

“We” is a separate question

The choice to write “we” is often treated as the same decision as the choice of voice, and it is not. What “we” does in “we show,” “we hypothesize,” “we release,” and “we conjecture” is mark the epistemic stance of III.B: it identifies a claim as this article’s, made at a stated strength, and distinguishes it from what prior work established or from what is generally accepted in the field. An article that systematically avoids the first person loses that marking, and its claims become harder to separate from the background it reports. “It is hypothesized that” does not attribute the hypothesis to anyone, which is a real loss of information in a sentence whose whole purpose is to say who is predicting what.

Which components lean which way, and why

The tendencies follow from I.C’s functions rather than from convention. The Method and a resource paper’s Data Collection lean passive, because their function is reconstructive and reconstruction does not depend on who acted: “the corpus was deduplicated at the document level” tells a reader everything they need. Results usefully splits: reporting sentences, answerable to the data, sit comfortably in the passive or in an inanimate-subject active (“accuracy drops by nine points on the subsequence subset”), while interpretive sentences, answerable to the argument, take the first person (“we attribute this to…”). Keeping the two voices distinct makes visible on the page the report/interpret separation that II.H insisted on. The Introduction and Conclusion lean active, since their moves (II.D, II.G) are claims and promises that belong to someone. And Limitations and the Ethics statement lean active most strongly of all, because they are statements of authorial responsibility: “it should be noted that certain biases may be present in the data” declines to say who noted it, who is responsible for the data, or whether the biases were actually found — an agentless construction doing the work of disclaiming the very responsibility the component exists to take.

Voice runs from the Method's agentless passive to Limitations' and the Ethics statement's strongly active, authorial voice.

Matching to the structural variant

The pattern is visible across I.B’s three variants without much elaboration. The resource/dataset paper carries the highest density of agentless passives in its annotation-procedure prose, where the process rather than its operator is the object of description, and then switches sharply to the first person for release and recommendation statements. The empirical/analysis paper is the most first-person of the three, since “we observe,” “we find,” and “we attribute” are precisely the epistemic markers its RQ-driven structure needs. The system/method paper sits between them, passive in its architecture description and active in its contribution and comparison sentences.

A worked example

The same content, written both ways, in two different components. In the Method — passive: “Sentences longer than 128 tokens were truncated.” Active: “We truncated sentences longer than 128 tokens.” The passive is at least as good here and arguably better, because the sentence is about the preprocessing pipeline the reader is reconstructing, and the fact that the authors did it is already given. In Limitations — passive: “It should be acknowledged that the findings may not extend to typologically distinct languages.” Active: “We tested only English, and we do not know whether the finding holds for languages with freer word order.” The active version names the agent, states the condition concretely rather than as a possibility floating free of anyone, and is checkable against the protocol — which is exactly what II.H asked Limitations to be derived from.

The published abstracts show the same split. The HANS abstract opens agentlessly and generically — “A machine learning system can score well on a given test set by relying on heuristics” — because that opening sentence is about a phenomenon, not about the authors; it then turns first-person for every claim the article itself makes: “We hypothesize,” “we introduce,” “We find,” “We conclude.” The Transformer abstract does the same, describing the prior landscape without agents (“The dominant sequence transduction models are based on complex recurrent or convolutional neural networks”) and switching to “We propose” at the sentence where the contribution starts. In both cases the voice change marks the boundary between what is being reported about the world and what is being claimed by this article — the same boundary III.B asked writers to keep legible.

Common failure modes

The passive can be used to hide who made a decision, which is a substantive rather than stylistic problem when it appears in Limitations, in the Ethics statement, or anywhere a choice is being reported that a reader might reasonably question. The agentless passive can also make a protocol unreconstructible: “the data was filtered” omits both the agent and, usually, the criterion, and II.F’s practical check — could a competent reader rebuild this without asking a question — fails on exactly such sentences. In the other direction, unbroken first-person can flatten a section into a sequence of sentences all beginning “we,” which loses the topical continuity described above and makes the genuinely agentive claims no more prominent than the procedural ones. And nominalization is frequently mistaken for the passive and criticized in its place: “the performance of an evaluation of the model was undertaken” is bad because its verb has been buried in a noun, not because it is passive, and rewriting it as an active with the same buried verb fixes nothing.

Toward III.D

Sections III.A through III.C have treated three properties of the sentence — its register, the strength it asserts, and where it places its agent — as instruments of the functions established in Parts I and II. What all three have taken for granted is that the terms the sentences are built from are stable: that a word naming a construct in the Introduction names the same construct in the Results. Section III.D takes up that assumption directly, as a question of precision and consistency in terminology.


This site uses Just the Docs, a documentation theme for Jekyll.