Mentor dos Nerds Home A Hypothesis Does Not Become Fact Because AI Repeated It
Post

Article Artificial Intelligence

A Hypothesis Does Not Become Fact Because AI Repeated It

An incomplete amber structure passes through a confirmation checkpoint, becomes a blue decision, and feeds a violet synthesis connected to applications and review
An incomplete amber structure passes through a confirmation checkpoint, becomes a blue decision, and feeds a violet synthesis connected to applications and review

I have no problem allowing an AI agent to work with hypotheses.

Requiring certainty before every step would sound responsible. It would also be an excellent way to stop doing anything at all.

Almost every real execution begins with something incomplete: an intention that still needs scope, data that has not arrived, a preference that appeared in three conversations but was never confirmed, a likely cause that still needs testing.

My problem begins when the hypothesis loses its label.

One agent finds a plausible explanation. It acts on it. Another agent receives the result. A third summarizes the context. By the time I notice, “it seems” has become “we know” without anyone ever making that decision.

It is a telephone game with persistence: each repetition seems to give the sentence more authority, when it has only increased its distance from the original evidence. The verifiable trail exposes that distance before the hypothesis becomes a rule.

It was not a spectacular, cinematic hallucination. It was worse: a reasonable inference received a fact badge and started governing the system in silence.

That is why, inside my harness, an applied hypothesis must leave a receipt.

TL;DR

Agents need to act under uncertainty, but they cannot promote a hypothesis to fact merely because it was repeated, summarized, or used successfully once. When a hypothesis influences an action, I record its origin, supporting and contrary evidence, confidence and limits, scope, affected consumers, provisional action, and review condition. My wiki separates three minimum objects: hypotheses hold possibilities that remain unconfirmed; decisions record the adopted verdict and its reach; syntheses preserve reasoning, sources, and disagreements. Plain-text Markdown files and Git version history make the trail readable and revisable, but they do not create truth by magic. Direction and the final verdict remain mine.

Uncertainty is not the defect. Its disappearance is

A good agent does not need to halt everything at the first gap.

If the impact is small, reversible, and local, it can declare an assumption and proceed. If two interpretations are plausible, it can choose one to produce a working version. If that choice changes price, security, publication, personal data, or another high-risk decision, it must return the question before acting.

The mistake is not forming hypotheses.

The mistake is hiding that execution depended on one.

There is an enormous difference between these two statements:

The client prefers to receive a weekly summary by email.

and:

I am assuming, with low confidence, that the weekly summary should be sent by email because that channel appeared in the last two conversations. I found no explicit confirmation. I will use this premise only to draft the message; I will not configure delivery until the owner confirms it.

The second statement is longer because it carries what the first erased: origin, uncertainty, scope, and action limit.

It also makes correction possible.

If the client says, “No, use the team messaging channel,” I know what must change. Without that trail, I have to search for every place where false certainty has already spread.

Every applied hypothesis needs a receipt

In my system, a useful hypothesis is not merely a sentence containing the word “maybe.” It must answer operational questions.

Field What must remain visible
Origin Which gap, conversation, file, or observation generated the hypothesis
Hypothesis Exactly what I am assuming, without language that pretends confirmation
Supporting evidence Signals that make the interpretation plausible
Contrary evidence Contradictions, exceptions, and missing data
Confidence and limits How fragile the interpretation is and what it does not allow us to conclude
Scope Where provisional use is allowed and where it is forbidden
Affected consumers Which texts, decisions, agents, automations, or files depend on it
Action taken What the agent did under that premise
Review if Which confirmation, denial, deadline, or evidence requires reopening the matter

I do not require a fabricated percentage just to make the record look scientific.

“Medium confidence,” accompanied by two pieces of evidence and one described contradiction, is more honest than an “82%” that came from nowhere.

The point of the record is not to decorate the hypothesis. It is to make action possible without erasing uncertainty—and to unwind propagation if the premise falls.

When I confirm or deny it, the system must know what to relearn

An isolated hypothesis is easy to correct.

The problem is the hypothesis that has already fed a post, a prompt—the instructions given to AI—a recommendation, a routing rule—the criterion that chooses which agent or flow receives the case—and three specialized agents.

That is why I record consumers: the places where a conclusion was used or may change behavior.

When I confirm a hypothesis, it does not merely receive a green stamp. The confirmed content must move to the proper owner: a decision, contract, specification, editorial rule, or another current source.

When I deny it, the work also does not end by writing “refuted.” The agent must propose updates to the dependents:

  • which text became inaccurate;
  • which decision used a premise that failed;
  • which agent must stop applying the rule;
  • which synthesis must preserve the disagreement;
  • which completed action needs review or reversal.

That preserves provenance—the path from origin to interpretation, application, and change—without pretending we never make mistakes.

I prefer a system that can show why it changed its mind over one that rewrites the past whenever it receives a correction.

My wiki is not magical memory

I use a wiki because I need to consolidate knowledge that deserves to be found and reviewed by people and agents.

That does not mean saving every conversation or dumping everything that happened into a folder.

Memory and wiki serve different roles in my system.

Memory helps recover continuity: what happened, where work stopped, which signals may matter now.

The wiki organizes knowledge that has already been worked through: explicit hypotheses, current decisions, syntheses, patterns, limits, and relationships to their sources and applications.

Neither sits above the file that governs the subject today. If a current specification contradicts an old synthesis, the specification wins. If I correct an interpretation about myself, the correction does not need permission from a summary produced by an agent.

The same distinction appeared when AI forgot precisely what I had already decided. Recovering context matters. Knowing which file has authority over each decision is something else.

A wiki reduces rediscovery.

It does not turn well-organized text into truth.

Three objects are enough to begin

A first wiki for agents does not need to be born with a taxonomy—an organized classification scheme—worthy of the Library of Alexandria.

I would begin with three objects.

1. Hypotheses

They hold possibilities that remain unconfirmed.

A hypothesis records the gap, its origin, evidence for and against it, allowed provisional use, dependents, promotion condition, and discard condition.

It may help an agent ask a better question or advance a reversible draft. It cannot silently remove a doubt or become a source of truth through repetition.

2. Decisions

They hold directions actually adopted by the responsible owner.

A decision records the verdict, relevant reasons, scope, exceptions, whom or what it affects, and when it must be revisited.

A decision is not the same as a recommendation. Three agents agreeing with one another still produces a recommendation. Direction changes when the decision owner—the person responsible for the verdict—adopts, corrects, or rejects the proposal.

3. Syntheses

They hold reasoning that does not fit inside a short verdict.

A good synthesis preserves sources, convergences, disagreements, open hypotheses, recommendation, and residual risk—the risk that still exists after the adopted controls or decisions. It does not need to reproduce the entire conversation. It needs to preserve what explains why the direction appeared reasonable and what remained unanswered.

When a synthesis leads to a decision, each should link back to the other.

The decision shows what governs now.

The synthesis preserves the path—including disagreements that the most convenient summary would try to erase.

Knowledge that never reaches application becomes a museum

A page can be immaculate and still change nothing.

If the system does not know who consumes a piece of information, the wiki becomes a beautiful place where decisions go to rest after they die.

That is why I connect knowledge to application.

A hypothesis about my style may affect the writing, translation, and voice-review agents. A privacy decision may affect tools, technical event records (logs), publication, and retention. A synthesis about a client may support a proposal, but it should not leak into another context.

The same principle appears in my CHARACTER.md. A blind spot only becomes valuable when it creates an observable counterweight in the system.

In the wiki, the equivalent question is:

If this information changes tomorrow, who needs to know?

If the answer is “nobody,” perhaps it is still only a note.

Not every observation deserves a page

Governance can also become a hiding place.

It is entirely possible to spend more time creating fields, indexes, and relationships than making the decision that would justify any of them.

I do not want to turn every provisional sentence into an administrative process.

An observation is worth promoting to the wiki when at least one of these conditions appears:

  • it will be reused across more than one session or by more than one agent;
  • it changes behavior, permission, priority, architecture, or communication;
  • it creates relevant risk if wrong;
  • it has already been rediscovered or debated more than once;
  • it has several consumers that will need coordinated updates;
  • losing its origin would make correction expensive or ambiguous.

A local, cheap, disposable question can remain in the current work.

The record should be proportional to the cost of forgetting, distorting, or repeating the mistake.

A starting instruction block for creating this wiki

The block below is a prompt—a set of instructions for guiding the task—that is intentionally small and tool-agnostic. It can be used with Codex, Claude, or another agent capable of working with files and Git.

It does not copy my harness. It is a first structure for you to test and adapt to your context.

AGENTS.md is the entry file: it presents purpose, authority, rules, validations, and navigation before the agent opens the rest. Reusable file structures—templates—define the minimum shape that each hypothesis, decision, or synthesis must fill in.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
Create an initial, revisable wiki using Markdown files versioned in Git.
It must be readable by people and agents without relying on magical memory.

Begin with:
- AGENTS.md as the entry point: explain purpose, authority, navigation,
  privacy rules, and validations;
- hypotheses/ for possibilities that remain unconfirmed;
- decisions/ for directions adopted by the responsible owner;
- syntheses/ for sources, reasoning, convergences, disagreements, and limits.

Define minimal reusable file structures (templates).

For each hypothesis, record: status, origin of the gap, explicit formulation,
supporting and contrary evidence, confidence and limits, provisional-use scope,
consumers or impacts, action taken, review condition, promote if, and discard if.

For each decision, record: verdict owner, adopted decision, relevant reasons and
evidence, scope, exceptions, affected consumers, related hypotheses and
syntheses, and review condition.

For each synthesis, record: question, sources, convergence points, preserved
disagreements, open hypotheses, related decisions, limits, and the next possible
application.

Mandatory rules:
1. repetition does not promote a hypothesis to fact;
2. confirmation or denial must generate a proposal to update affected consumers;
3. current sources and explicit human decisions outrank syntheses;
4. do not record secrets, credentials, personal data, or unnecessary sensitive
   content;
5. do not execute publication, deletion, or external action solely because the
   wiki changed;
6. keep one canonical source per decision and use Git history to review changes
   instead of creating "final-v2" files.

Before creating files, show the proposed minimum tree, the templates, and what
will deliberately remain outside this first version. This is an incipient wiki,
not a copy of another harness or a promise of universal architecture.

As an entry point, that file does not need to carry the whole wiki. It needs to teach the system how to find the smallest sufficient source.

Markdown and Git help because they make change visible

I prefer this knowledge to live in plain-text files using Markdown with history in Git.

Not because they are the only possible technologies.

But because I can read the material without a proprietary tool, compare changes, review a proposal, return to an earlier version, and discover when a hypothesis gained or lost reach.

Version control records changes over time and allows us to compare or recover states. That supports the technical trail.

It does not support the truth of the content.

Git can prove that a sentence changed. It cannot prove that the new sentence is correct.

For that, I still need sources, tests, an accountable owner, and review.

Direction remains mine

I built agents to seek contrary evidence, recall earlier decisions, and point out when an overly agreeable answer may be hiding a problem. That connects directly to the principle that the best answer is not the one I like most.

Even then, the agent does not become an arbiter of truth merely because it gained a wiki.

It can record that a hypothesis exists.

It can show where that hypothesis was applied.

It can prepare consumer updates if I confirm or deny it.

It can prevent “I do not know” from disappearing when a long conversation is summarized to fit the available space—a process known as context compaction.

The verdict still depends on whoever is responsible for the decision and its consequences.

This is the kind of system I build at i-9.ai: not an infinite memory that promises to know everything, but an operation able to show what it knows, what it assumes, who decided, and what must be reviewed when reality answers differently.

If your company is already using AI but decisions, assumptions, and corrections keep getting lost among chats, documents, and people, get in touch. We can begin with the decision your team repeats most and the mistake that would become expensive if a silent hypothesis started governing it.

A hypothesis does not need to be forbidden.

It needs to remain recognizable as a hypothesis until someone with sufficient authority and evidence decides what it becomes.

Continue reading

Go deeper

References and limits of use

  • AGENTS.md describes an open, readable format for guiding agents through context and instructions. It does not guarantee that every tool interprets precedence, commands, or nested files the same way; validate the behavior of the agent you choose.
  • OpenAI’s prompt engineering guide supports only the explanation that prompts guide model behavior and can be refined. It is provider documentation, not a universal standard for routing or quality.
  • The OpenTelemetry glossary, NIST entries for taxonomy and residual risk, and the W3C PROV family support narrow definitions. The NIST glossary is security-oriented; this article uses only the general definitions of a classification scheme and remaining risk, simplifies the concepts for editorial governance, and does not claim full compliance with those standards.
  • Pro Git, “About Version Control” supports the explanation that version control records changes and can recover earlier states. It does not validate the content of a decision or replace backups, access policy, or human review.
  • OpenAI’s compaction documentation supports the example of reducing conversation context while preserving useful state. It describes a specific OpenAI API implementation; it does not establish that every tool compacts, summarizes, or discards context in the same way.

The separation among hypotheses, decisions, and syntheses, the suggested fields, and the consumer-update flow are an authorial synthesis of how I operate agents. They are not a universal standard, an accuracy guarantee, or a complete copy of the private architecture I use. The version proposed in the prompt is deliberately initial and should be reduced, expanded, or discarded according to real context and risk.

The cover image is a synthetic, language-independent editorial illustration of an incomplete hypothesis passing through confirmation, decision, synthesis, application, and review.

This post is licensed under CC BY 4.0 by the author.

Open conversation

Continue the conversation

Disagree, spot a gap, or have an experience that adds to the subject? Comment with your GitHub account. Do not publish personal data, credentials, or sensitive information.