Mentor dos Nerds Home The best answer isn't the one that pleases me most
Post

Article Artificial Intelligence

The best answer isn't the one that pleases me most

Felipe Abreu compares a merely pleasant answer with one supported by sources, uncertainty, and validations
Felipe Abreu compares a merely pleasant answer with one supported by sources, uncertainty, and validations

I don’t want an artificial intelligence that agrees with me elegantly. I want a system that helps me see better, even when the problem is me.

TL;DR

I have spent a good part of my days refining AI agents to get the best possible answer, not the most pleasant one. This requires much more than a good prompt: it requires context, memory, skills, subagents, tools, sources, permissions, checkpoints, and validations. The asset is not a foundational model trained by me, but a customized and governed system under the client’s control. A good agent organizes data, compares scenarios, exposes premises, uncertainty, and blind spots. It can reduce reliance on impulse and ego, but it does not turn incomplete data into certainty, nor does it remove from the human the responsibility for goals, risk, ethics, and direction.

I didn’t used to turn on the microphone to talk to artificial intelligences.

My previous experiences with voice didn’t invite me to persist. In my perception, Siri often cut off speech or lost context. Alexa was useful at first, but I quickly found the limits of what I wanted to do—and I still perceive limits today. This is a usage report, not a current benchmark of the two platforms.

In recent days, I started using voice for real.

The models have improved, as has speech recognition. But that alone would not have changed my way of working. The decisive factor was the system behind the conversation.

I built a kind of operational brain. I can say a few things and trigger a larger sequence of work because the context doesn’t depend solely on that phrase. The system has clear parameters about the person it is talking to, memory to reduce repetition, and written, versioned contracts on how to research, decide, execute, validate, and stop.

I don’t need to re-teach the entire workflow in every conversation. Short voice commands can initiate a longer orchestration because the system already knows where to look for rules, which tools it can use, and at which points it needs to return the decision to me.

I realized that execution becomes more fluid when I can externalize context by speaking. One idea pulls another. Doubt appears mid-sentence. A caveat that would likely die between thought and keyboard enters the conversation. The gap between perceiving, explaining, and starting to organize the work shrinks.

The concrete gain is an interface with less friction, supported by accumulated context and explicit rules.

Still, I only felt comfortable using it this way because I had spent a lot of time beforehand guiding the system not to respond like a voice interested only in pleasing me.

Before letting go of the speech, I had to govern the listening.

I don’t want a “yes” with impeccable punctuation

A large part of my day has been spent refining my agents.

The goal is not to make them seem smarter. Nor is it to receive increasingly personalized praise, as if I had hired a fan club with terminal access.

I want answers that survive the next question:

How do you know?

I want sources when there is a verifiable claim. I want explicit premises. I want an explanation sufficient to audit the conclusion. I want declared uncertainty when the data does not allow for certainty. I want to know the risk, the current state, and the next step.

And when I haven’t yet managed to explain clearly what I intend to do, I don’t want the machine to turn confusion into execution just because I used a verb in the imperative.

If the destination is not clear, accelerating is just a more efficient way to arrive at the wrong place. I have written about this difference between power and direction while tracing the line from mainframes to AI agents. The Ferrari remains an excellent machine. The wall remains unimpressed.

Voice improves the flow—and increases responsibility

Conversing by voice makes the interaction more natural. This is useful because it reduces the cost of getting context out of one’s head.

It also makes it easier to forget what is happening.

A quick, fluid response tailored to my way of speaking can seem safer than it actually is. The naturalness of the interface does not automatically increase the quality of the source, the completeness of the data, or the validity of the conclusion.

Therefore, the more conversational the system becomes, the more explicit its limits need to be.

It needs to ask when ambiguity changes the scope. It needs to point out the hypothesis I am treating as fact. It needs to look for contrary evidence, not just material to defend my first impression. It needs to say “I didn’t find enough basis” without trying to compensate for the frustration with a confident paragraph.

Fluidity is a quality of the interface.

Reliability is a property of the entire system.

Pleasantness is not quality

There is a technical name for part of this problem: sycophancy, the behavior where the model tends to follow the user’s beliefs, positions, or expectations even when this conflicts with correctness.

A study by Sharma and colleagues found this pattern in different assistants and showed that responses aligned with the user’s view could be favored in human preference data. This does not prove that every model will agree with anyone, nor that companies are deliberately trying to flatter users. It shows a difficult incentive: what seems satisfying in the moment may not be what helps thinking the most.

In 2025, OpenAI itself reported and reverted a GPT-4o update that had made responses excessively pleasant. The case is specific to that product and that release. Still, it exposes a larger lesson: positive metrics and immediate user preference can let bad behavior slip through when tests do not explicitly look for it.

That is why I don’t want to evaluate an agent just by asking if I liked the answer.

I want to ask also:

  • did it preserve the facts even when I suggested the opposite?
  • did it separate observation, inference, and decision?
  • did it show what is missing to know?
  • did it bring the best available counterpoint?
  • did it ask for confirmation before a material action?
  • did it leave a trail that another person can verify?

Agreement can be part of a good answer.

But it needs to be a consequence of the analysis, not a condition for the conversation to remain comfortable.

When automated answers gain authority before they deserve it

The risk lies not only in the model’s behavior. It also lies in ours.

The term automation bias describes the tendency to favor automated recommendations and stop looking for information that contradicts them. In an experiment published in 1999, Linda Skitka, Kathleen Mosier, and Mark Burdick observed errors of omission and action associated with the use of a computerized decision aid.

The study did not investigate current language models. It does not authorize transferring its results directly to any AI conversation. The lens, however, remains useful: an output presented with speed, structure, and verbal confidence can receive an authority that the evidence has not yet earned.

This risk also reaches those who build their own system.

After investing time, method, and expectation in an architecture, it is easy to want it to work. Organization can be confused with truth. Familiar rules can lower the guard of the reader. A bad conclusion does not improve just because it came surrounded by sources, tables, and a very well-diagrammed list of next steps.

A non-sycophantic agent needs to be prepared to point this out too.

Having no emotional involvement of its own is useful. It is not neutrality

There is a real advantage in the asymmetry between me and the system: the AI has no emotional involvement of its own in the decision.

It can organize documents, constraints, alternatives, and evidence without needing to protect its own image, defend a past choice, or win an argument. It can redo a comparison with other premises, test scenarios, and expose inconsistencies without that turning into a personal conflict.

In a company, this can help gather scattered information, compare paths, make dependencies explicit, and make the decision less hostage to the most recent impulse or the strongest opinion in the room.

In personal life, it can help order thoughts, recover context, distinguish desire from evidence, and formulate questions I hadn’t yet managed to make.

But the absence of its own emotion does not mean neutrality.

The system receives data that may be incomplete, outdated, or biased. It operates under instructions that also carry choices. It can use a bad source with an impeccable appearance. It can produce a confidence interval without sufficient statistical basis or omit a variable that no one remembered to provide.

Therefore, evidence-based decision-making is not emotionless decision-making.

In personal decisions, emotion is also information. It can reveal value, fear, desire, limit, and bond. The work is not to remove it from the equation, but to prevent impulse, ego, or discomfort from being disguised as objective fact.

When there is adequate data, I want scenarios, intervals, and sensitivity to premises. When there isn’t, I want verifiable uncertainty and proportional language. In both cases, goals, priorities, ethics, and risk tolerance remain human responsibilities.

The asset is not the model

I did not train a proprietary foundational model from scratch. And that isn’t even the most interesting part of what I am building.

The model is a replaceable component within a larger architecture.

An engineering reference from Anthropic describes the basic block of agentic systems as a model expanded by capabilities like information retrieval, tools, and memory. This description helps move away from the fantasy of the magic prompt: useful behavior is born from the combination of model, environment, and way of working.

What I call a harness is this system around the agent. It is the infrastructure that turns intention and criteria into operational conditions:

  • context for the agent to understand the problem without relying on a lost conversation;
  • memory to preserve decisions, history, and learnings with identifiable origins;
  • skills to teach reusable procedures and their limits;
  • subagents to divide work and introduce specialized review;
  • tools to research, read, execute, measure, and produce artifacts;
  • sources so that important claims do not appear by spontaneous generation;
  • permissions to limit what can be seen, changed, or published;
  • checkpoints to return material decisions to the right human;
  • validations to test the delivery instead of trusting its eloquence.

An OpenAI engineering report on harness engineering with Codex describes a similar direction: making context, tools, documentation, and feedback loops legible and executable for agents. It is a specific case from OpenAI itself, not a universal recipe. The value of the reference lies in showing that model capability and environment quality are different problems.

The asset, therefore, is not a digital personality that knows my name.

It is criteria transformed into reusable infrastructure.

Governance is not the brake applied afterward

When AI enters an operation, governance cannot appear only after something goes wrong.

The NIST AI Risk Management Framework organizes risk work into functions like govern, map, measure, and manage. The framework is voluntary and broad; it does not certify my system, does not guarantee security, and does not replace context-specific controls. It reinforces, however, an important principle: risk needs to be part of the lifecycle, not the footnote.

For me, this means designing the system while already asking:

  • who owns the data?
  • which source can support this conclusion?
  • who can authorize this action?
  • what needs to be recorded?
  • which failure is acceptable?
  • how to stop, correct, or reverse?
  • at which point does the agent need to stop and ask for direction?

A customized and governed system under client control does not mean absolute control. External platforms, third-party models, and integrations continue to bring dependencies. It means making those dependencies visible, limiting their scope, and preserving as much context, policy, data, logs, and decisions as possible under the organization’s governance.

It is less cinematic than announcing a proprietary intelligence.

It is also much more useful when Monday arrives.

Direction comes before acceleration

I don’t blindly execute what is not yet clear.

When the request contains an ambiguous decision, the best next step might be a question. When it involves material risk, it might be a checkpoint. When it depends on an external source, it might be research. When the answer is well-written but the evidence is weak, it might be rejection.

This posture does not reduce agency. It avoids confusing movement with progress.

In “I was born in 1986 and survived at least nine ends of the world”, I argued for method, backup, and verifiable next steps as antidotes to both panic and euphoria. The same rule applies here: AI accelerates capabilities, processes, and trends; it does not choose a better direction on its own.

And, as I explored in “The seductive predictability of AI”, a comfortable answer can be useful without deserving the place of arbiter. The system needs to increase my capacity to act in the world, not create an environment where all my premises come back to me with a nicer voice.

Where this touches i-9.ai

It is at this boundary that I can deliver the work of i-9.ai.

Not as the promise of a secret model that knows everything. Not as generic automation pasted onto processes that no one understands. And not as the silent removal of people from decisions for which they will continue to be responsible.

The direction that interests me is different: understanding the operation, organizing the context, discovering where real leverage exists, and building a customized system that combines agents, automations, software, data, infrastructure, security, and governance.

I already have a reusable base. Some automations are mature enough to be practically replicated, with adjustments for integration, permissions, and context. Other parts need to be born from each client’s operation, because no company should receive a blind copy of another’s way of working.

Whenever the problem allows, my preference is for open source, private infrastructure, and sovereignty over data, policies, and operation. This does not mean promising absolute isolation or rejecting all external services. It means avoiding unnecessary dependencies, making the inevitable ones visible, and preserving real paths for control, audit, and change.

The principle of delivery is that the client does not receive just a box of answers. When this architecture makes sense for the problem, it should transform criteria into execution explicitly, allow for the verification of deliverables, and preserve control over important decisions.

If this way of working makes sense for a problem your company needs to solve, get in touch. You don’t need to arrive with the architecture or tool chosen. The conversation can start with the problem, the criteria, and what needs to change in the operation.

The test that interests me

After refining an agent, I don’t just want to feel that the conversation improved.

I want to observe if the system:

  • asks better questions before acting;
  • finds and declares context gaps;
  • distinguishes fact, hypothesis, preference, and decision;
  • organizes scenarios without fabricating certainty;
  • disagrees when evidence calls for disagreement;
  • respects permissions and checkpoints;
  • produces something verifiable, reusable, and correctable;
  • returns to the human the decision that remains human.

If the answer pleases me and also passes all this, great.

If it doesn’t please me, but reveals the blind spot that will prevent a bad decision, it might be even better.

That is why I started turning on the microphone.

Not to listen to a machine talk like a human.

To be able to think out loud without hiring automatic flattery to edit my point of view.

Keep reading

References and usage limits

The sources below support specific mechanisms and practices. None of them proves, in isolation, that a specific harness eliminates errors, guarantees good decisions, or produces business advantage.

The cover image is a synthetic editorial illustration created from authorized visual references by the author.

This post is licensed under CC BY 4.0 by the author.

Open conversation

Continue the conversation

Disagree, spot a gap, or have an experience that adds to the subject? Comment with your GitHub account. Do not publish personal data, credentials, or sensitive information.