A. Writing rebuttals and responding to reviewers

A rebuttal has two readers, and they need different things

Every component treated in Parts I through IV addresses a reader who chose to read it. A response to reviewers addresses two readers who did not, and who differ in what they need. The reviewer needs their specific concern addressed, in terms recognizable as a response to what they wrote. The area chair or action editor needs something else entirely: a basis for deciding, across several review threads at once, which concerns survived the response and which did not. The second reader is the one who decides, and the second reader will very often not re-read the article.

This produces the governing constraint of the form. A response must be readable by someone who has the reviews and the response in front of them and not the article, which rules out any answer whose force depends on the reader turning back to a section to check it. Quoting the relevant sentence, naming the table by number, and stating the change that will be made are not courtesies; they are what makes a response usable by the reader who will act on it. A response that is perfectly persuasive to a reviewer who re-reads the paper, and opaque to a chair who does not, has been addressed to the wrong reader.

A review is a set of claims, not a verdict

The most useful preliminary move is to stop reading a review as a judgment and start reading it as a document making claims of varying strength — which is III.B’s ladder, applied to someone else’s text. Four kinds recur, and each takes a different response.

  1. A factual error about the article. The reviewer states that something is absent or was not done, and it is present. This is the easiest to answer and the most important to answer without heat: the location is given, and nothing more is required.
  2. A real weakness. The reviewer has identified something the article genuinely does not establish. The response is to concede it precisely, state its scope, and say where the article will acknowledge it — usually in Limitations.
  3. A misunderstanding. The reviewer has read the article and drawn a conclusion the authors did not intend. This is treated below, because it is the case most often mishandled.
  4. A difference of judgment. The reviewer thinks a different experiment would have been more informative, or that the contribution is smaller than claimed. This cannot be resolved by information, only argued, and it is where a response’s limited credit is best spent.

Sorting the points before drafting anything is worth the time it takes, because the four categories have different costs and different chances of success, and a response that treats them all the same distributes its space badly.

Four kinds of review point -- factual error, real weakness, misunderstanding, difference of judgment -- each taking a different response.

A reviewer’s misunderstanding is usually a writing defect

This is the section’s central claim, and it is the point at which reviewing becomes the most direct feedback on Parts I through IV that a writer will ever receive. When a competent reviewer, reading in good faith, concludes something the article did not intend to say, the most probable explanation is not that they read carelessly but that the article permitted the reading. A reviewer who thinks the contribution is a claim about models in general has usually been given a sentence that says so, and III.B’s overclaiming section describes exactly how such a sentence gets written. A reviewer who thinks a baseline is missing has usually not been told, in II.E’s differentiation sentence, why the omitted comparison is not the relevant one.

The consequence for the response is concrete. The strongest answer to a misunderstanding names the sentence that produced it and states what it will be changed to, because that answer is simultaneously a correction, a demonstration that the concern was understood, and a commitment the camera-ready can be checked against. “The reviewer has misunderstood our claim” is weaker than the same content delivered as “our phrasing invited this reading; the sentence will be changed to read as follows” — not because the second is more polite, but because it repairs the defect rather than assigning it to the reader.

The constraints are functional

The rules governing responses look arbitrary until they are read against the chair’s reading budget, at which point they follow. At ACL Rolling Review, responses are text-only and no external links are permitted; new experimental results may be included when they answer a reviewer’s question directly, but only as “minor add-ons” rather than “unsolicited new results or fresh results that would indicate substantial additional work”; and area chairs read at most two author responses per review thread. Each of these protects the same scarce resource. The link prohibition keeps the decision material self-contained for a reader working through many threads. The limit on new results prevents a response from becoming a different article that was never reviewed. The two-response cap makes continued argument literally unreadable past a point, which is why ARR’s own guidance warns against “repeatedly arguing your point of view to a strongly opinionated reviewer.”

Section IV.C described a reading path and argued that an article’s navigational surfaces carry more traffic than its prose. A response is that argument in its most compressed form: it is read once, under load, by someone comparing it against several others, and every device IV.C recommends — informative structure, front-loaded claims, extractable first sentences — applies with more force here than it does in an article.

Concession is an instrument, not a surrender

A response that contests every point contests none, and the reason is the one III.B gave about hedges: a marker is informative only if its absence also occurs. If a reader cannot find a single point the authors accepted, they have no way to calibrate the points the authors rejected, and the rational response is to discount all of them equally. Conceding a genuine weakness precisely — naming its scope, and stating where the article will record it — is what makes the contested points readable as considered rather than reflexive.

Precision matters as much in concession as in claim. “We agree this is a limitation and will add a discussion” concedes nothing checkable. “The result is established at one model scale; we will state in Limitations that the trend at larger scales is untested” concedes a specific thing, bounds it, and produces a sentence that will exist in the camera-ready. The first is a gesture; the second is II.H’s protocol-derived scoping arriving through a different door.

“Not novel” is II.E arriving late

The single most common substantive criticism in NLP reviewing is that the contribution is insufficiently distinct from existing work, and the structure of that criticism is worth recognizing. If the response must explain the delta between this article and a piece of prior work, then Related Work did not contain the differentiation sentence II.E required for each cluster of prior work — because if it had, the reviewer would have encountered the explanation in the article. The response can supply the explanation, and should. But the repair belongs in the article, and a response that wins the point without changing the text has fixed the reviewer and left the defect in place for every subsequent reader.

The same diagnosis applies to a reviewer who asks for a baseline the authors considered and rejected. The reasoning existed; it simply was not written down, and its absence is a gap in the article rather than an oversight by the reader.

Ordering the response

Reviews arrive numbered, and responses are conventionally organized to match. That ordering serves the reviewers and not the chair, who is trying to determine which concerns are decisive. The better ordering is by decision weight: the concern most likely to determine the outcome first, addressed at the greatest length, with the minor and factual points collected briefly afterwards. Where the same concern appears in two reviews, it should be answered once, prominently, and cross-referenced rather than answered twice at half strength.

Section IV.C named the strongest legitimate use of a question in an article as one that states an objection a sceptical reader is already forming, immediately before the article addresses it. That device is cheaper than a rebuttal and works on the same objection, in the same words, before a reviewer has committed the objection to writing. An anticipated objection costs two sentences in a Discussion; the same objection raised in review costs a review cycle.

Matching to the structural variant

The three variants from I.B attract characteristically different criticisms, and knowing which is coming is most of the anticipation IV.C recommended. The system/method paper draws missing-baseline and insufficient-ablation demands, because its claim is comparative and every comparison invites a further one; the productive pre-emption is I.B’s ablation table, made to isolate the contribution of the specific design choices claimed as novel. The empirical/analysis paper draws generalization challenges — one dataset, one language, one scale — which are III.B’s bounded-evidence problem arriving as a review comment, and which are best answered before the fact in Limitations rather than after it in a response. The resource/dataset paper draws annotation quality, licensing, and utility challenges, the last being the one III.B identified as structurally unanswerable, since a resource paper cannot measure downstream usefulness at publication time and should not pretend to.

A worked example

Take the invented syntactic-diversity article used since III.B and suppose four review points. First: “The paper does not report variance across random seeds.” If the article reports three seeds in a table, this is a factual error, and the answer is one sentence naming the table — with a note that the seed information will move into the main results table, since a reviewer missing it is evidence it was placed where readers do not look, which is IV.D’s argument about where evidence is actually read. Second: “Results are shown for one language only.” This is a real weakness and the concession should be exact: the claim holds for the language tested, Limitations will state that cross-lingual transfer is untested, and the Abstract’s phrasing will be scoped to match.

Third: “The authors claim diverse data produces more robust representations, but they measure only benchmark accuracy.” This is a misunderstanding produced by the article, and III.B named the mechanism — capability language attached to behavioral evidence. The response should concede the phrasing, quote the offending sentence, and give its replacement. Fourth: “A stronger contribution would have been to test whether the effect holds under domain shift.” This is a difference of judgment; it cannot be answered with information, and the honest response states what the present design establishes, acknowledges the proposed experiment as a genuine extension, and declines to reframe the article around it. Four points, four different kinds of answer, and only one of them an argument.

Common failure modes

The aggrieved register is the most damaging and the easiest to fall into, since a response is typically drafted shortly after the reviews are read; III.A’s advice about register applies unchanged, and the practical form of it is that a response should be drafted and then left for a day. Beyond it: equal length per point, which spends the chair’s attention uniformly on concerns of wildly unequal weight; promising revisions the page limit cannot hold, which converts a response into a commitment that the camera-ready will visibly break; new results exceeding the “minor add-ons” the process permits, which invites the objection that the article under review is not the article being defended; and continued argument with a reviewer whose position is fixed, which the process explicitly warns against and which consumes the space the decisive point needed.

A practical check

For every point in every review, write one line naming the sentence, table, or section that will change, or stating that nothing will change and why. Points with no entry have not been answered, only discussed. The list produced by this exercise is also the camera-ready’s task list, which is why V.F treats the response and the final revision as one operation rather than two.


This site uses Just the Docs, a documentation theme for Jekyll.