B. Collaborative writing and the seams it leaves
A co-authored draft is diagnosable
Most NLP articles have several authors, and a reader can usually tell. This is not in itself a defect — a reader is not owed the illusion of a single hand — but co-authorship produces three specific failures that a single-author draft does not, and all three are detectable in the text. Treating them as text problems rather than as interpersonal ones is what brings this section inside the course’s framework, and it also makes them fixable, since a seam in a document can be repaired by whoever notices it.
The three seams
The first is register discontinuity. Section III.A established that register varies legitimately by component — an Abstract and a Model Architecture section are properly written differently — but a shift within a component, or a shift between two adjacent sections at the same level of technicality, is a seam rather than a variation. The reader’s inference is not that two people wrote it, but that nobody read it through.
The second is terminological drift, and it is the most consequential. Section III.D established that a term is a pointer and that the article owns the referent; four co-authors independently naming one construct produce exactly the four-name failure III.D described, with the added property that each author’s usage is internally consistent, so no one rereading their own contribution will find it. Drift is invisible from inside a section and obvious across sections, which makes it the seam most reliant on a deliberate pass.
The third is the unreconciled claim. Section I.C identified a mismatch between Introduction and Conclusion as a reliable signal that the argument shifted during writing; the collaborative version is more specific. When one author writes the Introduction and another the Conclusion, the promise and its closure are drafted from two different mental models of the contribution, and nothing in the process forces them to meet. II.G’s correspondence check exists precisely for this, and it is the check most often skipped, because each author has checked their own half.
Division by section is the worst available division of labour
The obvious way to divide the writing is by section, and it is the arrangement that produces all three seams most reliably. The reason is structural rather than social: I.C’s components are interdependent in specific, known ways. The Introduction makes a promise the Results must keep. Experiments fixes a protocol the Method must have described. Related Work’s differentiation sentences determine which baselines the Results must contain. Dividing by section places the boundary between writers exactly on top of these dependencies, so every dependency now has to survive a handoff, and the handoffs are invisible in the finished text.
Two divisions work better. Dividing by claim keeps each dependency inside one writer’s material, at the cost of every writer touching several sections. Alternatively, one writer owns the article’s spine — the discourse, the Introduction, and the Conclusion — while others supply the technical content of Method, Experiments, and analysis, which keeps the promise and its closure in one head. Neither is free, but both put the seams somewhere other than on the load-bearing joins.
The discourse paragraph is the coordination artifact
Section II.A asked for the discourse to be drafted as a single section-free paragraph — problem, gap, approach, finding, implication — before any section is written. For a single author that paragraph is a planning device. For a group it is the cheapest coordination artifact available, and its value rises with the number of authors. It is short enough that everyone will actually read it, concrete enough that disagreement about the contribution surfaces immediately rather than in the third draft, and it fixes the single through-line IV.B required before anyone has written prose that would have to be discarded to change it.
Disagreements that surface at the paragraph stage cost an afternoon. The same disagreements surfacing after drafting appear as a Conclusion that does not match the Introduction, and by then the disagreement is expensive to distinguish from a writing problem.
One writer must own the seams
The final pass over a co-authored draft has to be made by one person, and its purpose is not style. Uniformity of voice is not worth much and is not what a reader notices. What one reader must do is check terminology across sections, check that the Introduction’s promise and the Conclusion’s closure are the same object, and check that the claims made in the Abstract are the claims the Results support — the three seams above, in the order in which they are most easily missed. This is a distinct task from each author revising their own material, and it cannot be performed by the group, because the failures it hunts are only visible from a position that reads all the sections consecutively.
Credit, order, and contribution statements
The social apparatus around authorship is largely conventional and varies by group, so this course describes rather than prescribes. In NLP, first authorship conventionally marks the person who did the most of the work and typically wrote most of the article, and last authorship frequently, though not universally, marks the senior supervising author; intermediate positions are ordered by contribution in some groups and negotiated in others. Contribution statements, which several venues now invite, make these arrangements explicit and are worth writing for the reason I.C gave about Acknowledgments: they carry no argumentative weight, which is exactly why the article loses nothing by being precise in them. The one substantive point is temporal. Authorship order is far easier to agree at the discourse-paragraph stage than after a submission deadline has begun to approach, and the same is true of who owns the final pass.
Matching to the structural variant
The three variants from I.B carry different collaborative loads. The resource/dataset paper is the most multi-authored of the three and the most seam-prone, since annotation, guideline development, quality analysis, and baseline experiments are commonly done by different people, and I.B’s list of its apparatus — statistics, agreement, pipeline, examples, baselines — is very nearly a list of separable work packages. The system/method paper divides more cleanly, because its Method is subdivided by component (I.B) and the components are genuinely modular, but its Introduction-to-Results dependency is the tightest of the three and suffers most when the two ends are written separately. The empirical/analysis paper typically has the fewest authors and the strongest through-line, and is correspondingly the least exposed — but for the same reason, terminological drift damages it most, since III.E noted that its argument lives in the linkage between findings rather than in any single number.
A worked example
A constructed instance of the second seam. In a four-author draft, the Introduction motivates the work in terms of syntactic diversity; the Method section describes the construction of a corpus with high structural variety; the Results report gains attributable to linguistic richness; and the Limitations section notes that diversity was measured on a single corpus. Each section is internally consistent, and each author would defend their own term. The reader is left to decide whether four properties are in play or one, and cannot resolve it from the text — which is III.D’s construct-validity problem, produced here by process rather than by carelessness.
The repair is the terminology inventory III.D described, applied across authors rather than across a draft: one term per construct, one definition, one location where it is introduced, and a search for every synonym. The inventory is also the artifact that prevents recurrence, because the disagreement it exposes — whether structural variety and syntactic diversity are the same construct — is a substantive question about the contribution that the group had not previously noticed it had never answered.
Common failure modes
Division by section, discussed above, is the root cause of most of the rest. Beyond it: the merge that nobody reads through, in which each author reviews their own contribution and no one reads the assembled article consecutively; the late-arriving co-author, whose contribution is inserted without the surrounding claims being adjusted, and which is detectable as a paragraph that could be deleted without loss; the Introduction written first and never revised, which II.G’s check catches and which is the collaborative default because the Introduction is the natural first task and the least natural to revisit; and the terminology negotiated in comments rather than in the text, where a resolution reached in discussion is never propagated to the three sections that still use the old term.
A practical check
Build the terminology inventory from III.D as a table with one row per construct and one column per section, and fill it in by searching rather than by memory. Any row with more than one distinct term is drift, and any row with a term the group cannot agree on is a substantive disagreement about the contribution that has been hiding inside a vocabulary problem. Then run II.G’s correspondence check between the Introduction’s contribution list and the Conclusion’s first move, with the specific instruction that whoever performs it must not be an author of either.