Blog I-9 Home You automated the review and outsourced the reading
Post

Article Software Engineering

You automated the review and outsourced the reading

A stack of repeated sheets emerges from a printer beside a short review with distinct marks and a magnifying glass
A stack of repeated sheets emerges from a printer beside a short review with distinct marks and a magnifying glass

Artificial intelligence is writing for people who do not want to write and summarizing for people who do not want to read.

Somewhere in the middle, someone still needs to understand.

I find that irony particularly good when it appears in a code review: examining a software change, looking for problems, and discussing what needs to change before incorporating it.

You save a few minutes generating the assessment. Your colleague receives a small encyclopedia and has to work out what is a real problem, a repetition, a hypothesis, or merely a paragraph that stands up straight.

Congratulations. The writing has been automated. The reading has a new owner.

TL;DR

AI can investigate code, test hypotheses, and make a review clearer. Publishing its output without reading, checking, and removing duplicates can transfer the work to the rest of the team. The criterion for a good review is a verifiable contribution: what problem exists, where, under what condition, with what consequence, and what the comment adds to what has already been raised. A long explanation may be necessary. Volume alone does not prove analysis. Whoever publishes the review needs to take responsibility for what they chose to keep.

The comment was already there. More similar ones arrived.

In one review, I carried out a focused analysis and used AI to make the wording more concise. The judgment remained mine: what to investigate, what was supported, and what deserved to enter the discussion.

Then a colleague published an extensive review generated by another AI. As I read it, it repeated the same already-raised points three or four times.

This is my recollection and interpretation of the episode, not an independent audit. The company, people, and details of the change stay outside this article. The habit interests me more than the colleague’s identity.

My concern was the contribution: more text did not mean, in that case, more problems discovered or more decisions clarified.

The review needed to review what had already been reviewed too.

One person sees the savings. The team gets the bill.

Producing text with AI can be fast. Receiving that text does not make its evaluation instantaneous.

Someone needs to open the change, understand the context, locate the relevant section, verify whether the risk exists, and compare the comment with what the discussion already contains.

If the same problem appears in four formulations, the recipient needs to recognize that they are not four problems. If a suggestion depends on an unconfirmed premise, someone needs to discover it. If a comment arrives after a correction, someone needs to notice that its subject has already disappeared.

Generation was cheap for the sender. Triage was left to the recipient.

That cost can become interrupted attention, duplicate responses, repeated investigation, and rework. I have no percentage to assign to it. I have a criterion: the sender’s savings should not be evaluated while ignoring the work of the person who has to check the output.

“My review was ready faster” describes one stage. We still need to know whether the team reached a better change faster.

Repeating the finding is not confirming the problem

A finding is the specific problem a review claims to have found.

If two analyses reach the same suspicion and one brings a test that reproduces it, a previously overlooked condition, or a different consequence, there is a new contribution.

If the second merely changes the vocabulary and repeats the conclusion, there is more text.

A hypothesis can also be useful, provided it stays labeled as a hypothesis and explains what remains to be verified. What does not help is giving a suspicion the tone of a proven defect because that makes the comment more imposing.

I want a review to help us decide. That requires reading the code, the change’s criteria, the existing comments, and the corrections that have already happened.

Generating an assessment of an isolated section and dropping it into a living discussion without reconciling that context is an efficient way to arrive late with great conviction.

“Then ask AI to summarize it”

Of course I can.

The first AI writes too much. The second summarizes. A third compares the summaries. Meanwhile, someone tries to discover whether the first was right.

The scene is ironic. The capability can be useful.

A good summary helps people navigate a necessary assessment. But reducing its length does not verify its contents. It may even erase the very condition that separated a real risk from a generic suspicion.

I do not want to replace an enormous text with a short, wrong comment either.

The work remains selecting, checking, and judging. Responsibility for what gets published stays with the person who brings the assessment into the discussion.

A long review can be the right review

Some changes are complex. Sometimes a long explanation is necessary to show a sequence of failures, a compatibility constraint, or a risk spanning several components.

My criterion is not counting lines or imposing a minimalist comment on every problem.

It is asking what each passage does for the reader and the decision.

Google’s code review guide on comments recommends explaining the reasoning and making clear why a suggestion matters. It is a communication reference, not a measurement of the episode I described.

A good justification may need space. Four versions of the same demand still need editing.

Being concise means delivering enough information for the next decision, with as little unnecessary burden as possible.

Useful AI comes before the publish button

I want to use AI to investigate better and write more precisely.

I can ask it to look for cases my analysis missed, challenge a hypothesis, prepare a test, compare the change against its requirements, or find duplicates in my own comments. I can ask for a clearer explanation once the reasoning is supported.

But the result needs to come back for a concrete check:

  • Which section and condition support this finding?
  • Was the effect demonstrated, inferred, or does it still need a test?
  • Has this already been raised or corrected? What does this analysis add?
  • Does the proposal fit the scope and expected behavior of the change?
  • Have I read the comment I am about to publish, and can I defend it?

If several observations describe the same problem, I consolidate them and preserve the new evidence. If there is no additional contribution, I discard the duplicate. If context is missing, I say which context, rather than turning uncertainty into a technical accusation.

This can also become part of the review’s workflow. Generation does not need to be the last stage before sending.

Someone needs to understand before signing

The provocation about writing less and reading less describes a shortcut I want to question. It is not a statistic about everyone or a diagnosis of the colleague’s intent.

There are excellent uses of AI for reading, investigating, and reviewing more. I want those uses.

What I reject is treating the ease of producing an assessment as permission to leave its understanding to the recipient.

Before the next review, check what has already been said, remove what only repeats, label what remains a hypothesis, and take responsibility for what is left.

You can automate an enormous part of the work.

Just do not call the review complete when the package still needs to be reviewed by the person who received it.

Reference and limits of use

Google’s guide to review comments supports the guidance to explain the reasoning and relevance of suggestions. It does not evaluate AI tools, measure productivity, or prove what happened in the episode. The account belongs to the author; the transfer of attention costs and the proposed criteria are this article’s practical interpretation.

The cover is a synthetic editorial illustration, with no text and independent of language. It does not represent the workplace or the people in the account.

This post is licensed under CC BY 4.0 by the author.

Open conversation

Continue the conversation

Disagree, spot a gap, or have an experience that adds to the subject? You can comment without creating an account or use one of the available sign-in methods. New comments may be moderated; when published, they are public. Do not publish personal data, credentials, or sensitive information.

When comments load or are submitted, technical data may be processed as described in our privacy policy. Privacy.