G. Writing assistance and disclosure

What the policy distinguishes

The outline this course follows predates the current disclosure regime, and the topic belongs here because it is now part of every submission rather than because it is novel. The ACL policy on writing assistance draws its distinctions by what the tool contributes, not by which tool is used, and the gradations are worth stating accurately. Assistance with language alone — paraphrasing and polishing text the authors wrote — is treated as comparable to a spell checker or a thesaurus and requires no disclosure, as does short-form input assistance such as predictive text. Using a model to help find relevant literature is permitted, but the authors must read, discuss, and cite the work themselves. Generated text describing widely known concepts must be flagged, verified for accuracy, and properly attributed. Where a model contributes new research ideas that the authors then develop, the policy suggests acknowledging its use and checking the ideas are not already in the literature. And where a model contributes new ideas and new text together, the policy states that this “seems like the definition of a co-author, which the models cannot be,” and discourages it.

Disclosure obligation rises with how much of the idea and the text a tool contributed, from language polish to ideas and text together.

Checklist item E1

Disclosure is made through the Responsible NLP checklist’s Section E, alongside the artifact, computational, and human-participant sections V.D described. Three properties of that mechanism are worth stating plainly because they are commonly misunderstood. The answers are visible to reviewers. They are published with accepted papers. And disclosure is explicitly not automatic grounds for rejection — the mechanism is a transparency instrument rather than a prohibition, and a writer who conceals a disclosable use is generally trading a small and permitted disclosure for a much larger integrity problem.

The boundary is accountability, not tooling

The policy’s gradations look arbitrary until read functionally, and then they follow from a single principle. Authorship is answerability for the claims an article makes: an author is someone who can be asked why a sentence is true and is obliged to have an answer. Language polishing does not touch answerability, since the claims and the evidence for them are unchanged and the authors remain exactly as able to defend them. Idea generation touches it partially, which is why acknowledgment is suggested. Ideas and text together produce a contribution whose provenance is a party that cannot be asked anything and cannot be held to anything — which is precisely the policy’s own reasoning that this describes a co-author the model cannot be.

This is the same criterion I.C used to organize the entire taxonomy of components. Every component’s function was stated in terms of what it does for a reader trying to evaluate a claim, and evaluation requires someone answerable at the other end. The disclosure regime is that requirement, applied to how the text came to exist.

Verification is the obligation the tool does not discharge

The characteristic failure in practice is not that text was generated but that a claim went unchecked, and it lands precisely on V.E’s rule that a citation is a checkable claim. Generated related-work prose reliably produces two specific errors: references that do not exist, and references that exist but do not say what the surrounding sentence claims. The second is far more dangerous, because it survives every automated check — the paper is real, the authors are real, the venue is real — and fails only against V.E’s practical check of opening the source and locating the supporting sentence.

Section III.F’s unattributed attribution has an exact counterpart here. “Recent work has shown” borrows the authority of a literature without identifying it; a fabricated or misdescribed citation identifies a literature that does not support what is claimed, which is the same failure with the appearance of rigour added. Both are detected the same way, and V.E’s outward audit is the pass that catches them.

What no tool supplies

Two operations in this course resist assistance for a structural reason rather than a technical one. Section II.A’s discourse paragraph requires knowing what the evidence will actually support, which is knowledge held by the people who ran the experiments and is not recoverable from a draft. Section IV.A’s framing decision requires choosing among true descriptions of a result, and IV.A established that the choice determines which experiments the article needs — a decision that is only sound when made by someone who knows what the results can bear. Both are judgment about evidence rather than production of text, and both are the points at which an article stops being a document and becomes a claim someone is making.

Matching to the structural variant

Little varies here, since the policy is variant-independent, but the exposure differs. The system/method paper is most exposed on generated descriptions of widely known components, which is exactly the low-novelty text the policy asks be flagged and verified, and where an inaccurate standard description is easy to miss because it reads as familiar. The empirical/analysis paper is most exposed on generated related-work prose, since it typically surveys a larger and more interpretive literature and its claims about that literature are load-bearing for its gap. The resource/dataset paper is least exposed in its prose and most exposed in its documentation, where V.D’s artifact obligations require statements about licenses, provenance, and consent that are matters of fact rather than of writing.

A worked example

A constructed case. A draft’s Related Work contains: “Several studies have found that data diversity improves generalization more reliably than data volume (Smith et al., 2021; Chen and Park, 2022).” The sentence is fluent, appropriately hedged, and correctly formatted. Three things must be true for it to be usable: both articles must exist, both must be about data diversity rather than a neighbouring topic, and both must support a comparison against data volume rather than merely study diversity. The third condition is the one that fails most often and the one no formatting check detects. If only the first two hold, the correct repair is not to delete the sentence but to state what the cited work does support and to mark the comparison as this article’s own claim — which is V.E’s third framing, arrived at from a different direction.

Common failure modes

The undisclosed disclosable use is the one with the largest downside relative to its benefit, given that disclosure is not itself grounds for rejection. Beyond it: the citation that exists but does not support the sentence, described above; the generated standard description that is subtly wrong about a widely known method, which reviewers who know the method catch immediately; generated Limitations sections, which II.H’s requirement that limitations be derived from the protocol rules out by construction, since the protocol is not in the text; and the assumption that fluency indicates correctness, which is the general form of all of these and which the course’s own diagnostic addresses directly — a well-formed sentence can fail every function I.C identified while remaining well-formed.

A practical check

For every claim about prior work in the article, open the cited article and locate the sentence that supports it. This is V.E’s check, repeated here without modification and deliberately, because it is the single check that catches this section’s characteristic failure, and because a check that must be run for two independent reasons is one worth running.


This site uses Just the Docs, a documentation theme for Jekyll.