There is a kind of advice about artificial intelligence that seems useful until the third time you have to follow it.
“Use this prompt—the request, written or spoken, that guides what the AI should do.”
“First, tell the AI to do some research.”
“Then ask it to take on the role of an expert.”
“Do not forget to require sources, define the format, ask for a review, check the risks, preserve the context, and request the answer as a table.”
Great. Now people just have to remember how to run a miniature public procurement process every time they want to solve a problem.
I am not against prompts. A request is still the most direct way to tell a machine what I want. What bothers me is the expectation that every user should have to memorize the operational choreography of AI: when to research, which tool to use, how to hand off context, which data not to expose, which validation to run, and how far the AI agent—a system that receives context, uses tools, and executes steps to reach a goal—may go without asking for a decision.
If the work is recurring, those criteria should not continue to live only in the memory of the person making the request. They should live in the system.
Having a folder full of skills—instructions and reusable capabilities for performing a class of task— without knowing when to activate them is like collecting superpowers and leaving them offstage precisely when the problem appears. I do not know about you, but when I watched Heroes and saw Sylar accumulate abilities, what I felt above all was indignation—with some nervous frustration thrown in. The situation would tighten, I would remember an ability he had already acquired and that seemed perfect for it, and I would feel an almost physical urge to warn the character through the screen: “Come on, man! You already have the right tool; use it.” To make my indignation better, his arsenal even included supermemory. That was exactly what frustrated me as a viewer: capability listed in the inventory is not the same as capability available in practice.
A skill that is installed but never recognized by the agent in the right context remains a forgotten superpower. The user should not have to remember the skill’s exact name or prescribe every step to bring it onstage. The AI Harness—the set of context, rules, and criteria that decides when and how to use those capabilities—has to connect the natural request to the execution contract.
That is also where the resources and scripts used by skills, trigger rules, contracts, contextual data, workflows, validations, and limits come in. These are not sophisticated terms for selling a bigger prompt. They are parts of the operation.
The practical goal is simple: I make a natural request; the system recognizes the kind of work involved, loads the right skill, and follows the process we have already agreed on. I still provide the direction. I just do not need to relearn how to navigate the same roundabout every morning.
This does not happen because the AI acquired intuition. It depends on someone designing the trigger signals, testing real cases, correcting poor choices, and keeping rules and sources current.
The user should not have to navigate an invisible menu of hacks—/research, /handoff, /review, /now-be-an-expert—for the operation to work. They describe the result, provide the context only they know, and state the constraints. The Harness selects the applicable capability and procedure; if two routes are plausible, authorization is missing, or the decision has real consequences, it asks instead of choosing in secret.
A good prompt helps. A good system stops depending on my memory
An improvised prompt can produce an excellent response. It can also produce an excellent response on Monday, a mediocre one on Tuesday, and a minor work of corporate fiction on Wednesday.
The problem is not just the quality of the text I typed into the conversation box. It is everything I left out because I was in a hurry, tired, or simply thought it was obvious.
Compare this:
Research this topic and give me a comprehensive report.
With a request to an AI Harness that already knows the research contract:
I want to understand whether this solution makes sense in our situation.
In the second case, the sentence is shorter, but the work behind it may be greater. The research skill knows it needs to identify the decision at stake, prioritize primary sources, separate evidence from inference, record limitations, verify that the information is current, compare alternatives, and avoid treating a commercial page as independent proof.
I did not have to recite that checklist because it had already been discussed, written down, tested, and versioned.
GitHub’s documentation describes precisely this function of Agent Skills: instructions, scripts, and resources that are loaded when they are relevant to specialized work. The guidance itself distinguishes general rules, which belong in project instructions, from detailed procedures that should enter the context only when they are needed.
That matters because stuffing every rule into every request does not solve the problem either. It merely trades amnesia for congestion.
The right skill needs to appear without becoming a magic word
A useful skill is not an amulet named perfect-answer.md.
At a minimum, it needs to declare:
- what problem it knows how to solve;
- what signals should trigger it;
- what minimum information it requires;
- what sources and tools it may use;
- when to research before making a claim and when to hand the work to another specialist;
- which steps cannot be skipped;
- which output format it delivers;
- how quality will be verified;
- which authorization it received and which actions remain outside its scope;
- when it should stop and ask for human judgment.
The description is part of the behavior. If a skill says only that it “helps with research,” almost anything can trigger it, and almost nothing defines what it should do. If it says it should be used when a current claim, a vendor choice, or an architectural decision depends on external evidence, the agent has a more useful boundary.
GitHub documents that an agent can select a skill based on the task description. This does not prove that selection will be perfect in every tool or context. It proves something more modest and more important for this article: the idea of loading specialized instructions at the relevant moment, instead of forcing users to paste them into every conversation, already has a concrete implementation.
If Sylar represents the frustration of having a capability and failing to activate it, Batman represents the positive use. He is not Batman merely because he carries a utility belt; it is because he recognizes the right tool, knows how to use it, and puts it into action at the right moment. That is what a skill properly integrated into the Harness needs to do for the operation—without the cape, because budgets also have limits.
The natural request is still necessary. The entire ritual is not.
Deep research should not depend on me remembering to ask for deep research
If I ask which color works best for a button, we probably do not need to launch a scientific expedition.
If I ask whether an architecture should store personal data in another country, answering solely from what the model remembers would be irresponsible.
An ordinary user should not have to know the name of the research mode, decide how many searches to run, or tell the agent to “reason step by step.” The AI Harness can recognize the signals: a current topic, an expensive decision, legal risk, a specific reference, uncertain information, or a claim that will be published.
In that case, it triggers the appropriate process, says that it is going to research the question, and returns not only a conclusion but also the verifiable path that led to it.
This does not eliminate one essential question for the user:
What are you trying to decide with this research?
Without an objective, even impeccable research can solve the wrong problem with very attractive references.
A context handoff is not copying the entire conversation and praying
Another recurring prompt is: “Continue from where we left off.”
Where, exactly?
A long task accumulates decisions, files, hypotheses, test results, pending work, permissions, and things that seemed true three hours ago. Copying the entire conversation transfers volume, not necessarily useful context. It is the telephone game with an 80-page attachment.
A context-handoff skill may require a minimum package:
- the current objective;
- the actual state of the work;
- decisions already made;
- hypotheses that remain open;
- the files and systems involved;
- validations already run;
- risks, blockers, and the next step;
- what the new task is—and is not—authorized to do.
The user can say, “Take this to another task.” The AI Harness should know that “this” does not mean dumping everything. It means preserving the state required to continue without inventing continuity.
“Review” needs to mean more than rereading with the same confidence
When I ask for a review, I do not want the same answer to swap out three adjectives and come back wearing glasses.
A review skill needs to know both the object and the criteria. Code may require tests, diff analysis, compatibility checks, and risk assessment. An article may require authorial coherence, sources, links, strong claims, translation, mobile readability, and metadata. A decision may require someone to look for fragile assumptions and overlooked alternatives.
The work can follow different patterns. In its article on building effective agents, Anthropic describes patterns such as prompt chaining, routing, parallel execution, and an evaluator-optimizer loop in which one step produces an output and another evaluates it against defined criteria. The company also recommends starting with the simplest solution that works and adding complexity only when it produces measurable results.
That supports the technical pattern, not a promise of automatic quality. A bad evaluator can approve a bad answer with impressive solemnity. The criteria still need to be written, tested, and reviewed by people who understand the work.
Sensitive data needs to change the route before it enters the prompt
“Summarize this spreadsheet” may be an innocent request.
The spreadsheet may contain salaries, diagnoses, identification documents, contracts, trade secrets, or information about someone who never authorized it to be sent to an external provider.
It is not reasonable to expect every employee to remember to recite the company’s data policy in every conversation. It is equally unreasonable to conclude that a line in the prompt can make an unsafe workflow safe.
The AI Harness should classify the context before acting and apply rules such as:
- minimize the data being used;
- prefer local processing or an approved environment when required;
- remove identifiers that are not necessary;
- block providers or tools that fall outside policy;
- request authorization when the purpose is not covered;
- record what was transmitted without copying sensitive content into the log.
The generative AI risk profile published by the United States National Institute of Standards and Technology (NIST) treats risk management as a combination of governance, mapping, measurement, and management throughout the lifecycle. It does not provide a universal configuration for your AI Harness, nor does it certify that a policy written in Markdown will be followed. I use it here to support the need for context, responsibilities, evaluation, and monitoring—not to outsource to NIST a decision that remains the organization’s responsibility.
Publishing is not the same as finishing the writing
One of the most important differences between a chatbot and an operation is that generating an artifact does not automatically grant the right to put it out into the world.
If I say, “Prepare a post,” an editorial skill may research, write, review links, generate an image, validate the site, and open a version for review. Even then, publication may require my explicit approval.
If I say, “Publish the next approved article,” the AI Harness can check:
- whether human approval has been recorded;
- whether the technical review is complete;
- whether the date reflects the actual publication date;
- whether translations, image, and metadata are complete;
- whether the article depends on other content;
- whether the production environment is healthy;
- whether there is a path to reversal.
Notice the difference: the skill does not force me to remember every condition that must be satisfied before proceeding—every gate—but it also does not turn my sentence into unlimited authorization.
The same applies to a campaign, an infrastructure change, a payment, or an email. The final action may be simple; the conditions for reaching it are not.
What remains the responsibility of the person asking
I want to eliminate useless choreography, not responsibility.
The user still needs to provide four things that no instruction package should invent:
1. Objective
What needs to change in the world after the work is done?
“Analyze the data” is an activity. “I want to decide whether to keep this product” is direction.
2. Context only they know
A skill may know where to look for the files. It does not know that the vendor promised a term outside the contract, that the team lost someone, or that the priority changed in yesterday’s meeting—unless that context has been recorded and is accessible.
3. Real constraints
Deadline, budget, privacy, infrastructure, acceptable risk, affected people, and what cannot be changed. If a constraint changes the path, it needs to be in the request or in a trustworthy source of context.
4. Decision
The agent can organize evidence, present scenarios, and point out inconsistencies. A decision that commits money, people, reputation, or direction still requires an identifiable person who is responsible for it.
I have already written that data does not decide for us. Neither do skills.
What remains the organization’s responsibility
A company does not “install” skills and call it done.
It needs to observe how they are actually used:
- is the skill being triggered by the right requests?
- does it fail to appear when it should?
- does it ask for too much context or too little?
- are its sources still valid?
- do its scripts still work?
- do its validations catch real defects?
- do its limits reflect the current policy?
- do the cost and time still make sense?
- can people challenge and correct the result?
Context ages. Processes change. Tools are replaced. A skill that solved January’s problem can automate September’s retired rule with admirable efficiency.
In its article on Harness engineering, OpenAI reports that concentrating everything in a gigantic AGENTS.md led to too many instructions, the loss of useful context, stale guidance, and difficulty verifying the result. The direction it describes is to use that file as a map to smaller, structured sources of truth.
That is not a universal law, nor does it prove that the same architecture suits every company. It is a primary implementation account that matches the problem: operational criteria need to be discoverable, specific, and verifiable. A 40-page prompt pasted into every conversation is not maturity. It is a filing cabinet emptied onto the table.
This is not blind autonomy
When a system recognizes a task and applies the appropriate skill, it can seem as though the AI “already knows what to do.” The phrase is convenient, but it deserves scrutiny.
It does not know the way a person knows. It receives descriptions, rules, tools, context, and signals that increase the chance of choosing a useful path. That selection can fail. Execution can fail. The environment may have changed. Two rules may conflict.
That is why a reliable AI Harness also needs to include:
- states of uncertainty;
- stop criteria;
- permissions proportional to risk;
- evidence of what was executed;
- repeatable evaluations;
- human review at points of impact;
- reversal when an action changes something real.
The OpenAI documentation on Codex describes AGENTS.md as a way to record project organization, test commands, and expected practices. It also emphasizes configured environments, reliable tests, and verifiable evidence. This is not “write two rules and trust them forever.” It is precisely the opposite: turn expectations into contracts and results into something that can be checked.
How I would start without building NASA just to send an email
Choose a recurring task that currently depends on someone remembering many details. Do not start with the most critical task in the company.
It could be preparing research, reviewing a proposal, organizing the context from a meeting, or validating an article before review.
Then:
- record three good examples and three bad ones;
- write down when the skill should and should not be triggered;
- define the minimum inputs and what needs to be asked;
- turn the real process into short steps;
- mark the gates that require human decisions;
- automate only the checks that can be verified;
- run the skill against the examples;
- review the errors and update the contract;
- track whether it improves consistency, time, or rework;
- retire the skill if it only adds ceremony.
An initial prompt could be:
Help me turn this recurring task into a skill for my agent. First, identify the objective, trigger signals, minimum context, tools, steps, output format, validations, limits, and decisions that remain human. Use three real cases that went well and three that went badly. Do not write the final version before pointing out ambiguities and risks. Then propose a small, reversible, and measurable test.
That still does not produce a reliable skill. It produces a first artifact for discussion. Reliability begins when the text encounters real cases and survives them.
The best interface may go back to being a normal sentence
I invest a great deal of time in improving my agents so that, when it is time to work, I do not have to spend the same amount of time teaching them everything again.
That is the apparently contradictory part: behind a simple request, there may be a complex system. Not because I want to hide complexity behind an attractive demo, but because I have decided where that complexity should live.
Research has a research contract. Publishing has a publication gate. Sensitive data changes the route. Review looks for defects against explicit criteria. Context handoff preserves state. The agent receives tools that suit the task and stops when it reaches a decision that belongs to me.
The market will keep selling the definitive prompt of the week. Some of them will be useful. I will probably borrow ideas from several.
I just do not want to depend on remembering that the right incantation was Wingardium Leviosa—without delivering Ron’s version and having Hermione correct the emphasis—before every task.
I want to say what I need and provide the context only I have. The Harness recognizes the capability, applies the procedure, and respects the conditions we have already agreed on.
The prompt is still the conversation.
The AI Harness is what keeps that conversation from having to reinvent the entire company every time I open my mouth.
If this is the kind of operation you want to build—less ritual, more context, and verifiable results—learn more about I-9.ai.
Continue reading
- The Model Is Only the Engine. The AI Harness Is the Whole Car
- Does Your Agent Know When to Stop?
- If Nobody Knows What Happens Next, You Do Not Have a Workflow
- My Harness Learns, but It Does Not Get to Rewrite the Past
References and usage limits
- GitHub Docs — About agent skills: supports the definition of a skill as a loadable package of instructions, scripts, and resources for specialized work. It does not prove that activation will be correct for every agent or that a skill is good merely because it exists.
- GitHub Docs — Adding agent skills for GitHub Copilot: supports the role of the description in triggering a skill and the distinction between general instructions and specialized skills. It refers to GitHub Copilot’s documented behavior.
- Anthropic — Building Effective AI Agents: supports the patterns of routing, chaining, evaluation, and orchestration, as well as the recommendation to start simple. It is an account published by a vendor, not an independent comparison of business returns.
- Anthropic — Effective context engineering for AI agents: supports the distinction between optimizing an isolated prompt and managing the full set of instructions, tools, memory, history, and data available to an agent. It is technical guidance from the vendor itself.
- NIST AI 600-1 — Generative Artificial Intelligence Profile: supports the governance and risk-management approach throughout the lifecycle. It does not provide a ready-made architecture or certify the controls described in this article.
- OpenAI — Harness engineering: supports the account of the limitations of a monolithic instruction file and the use of structured documentation as a source of truth. It is an OpenAI implementation case, not a universal rule.
- OpenAI — Introducing Codex: supports the use of
AGENTS.md, configured environments, tests, and verifiable evidence in Codex. It does not prove autonomy or reliability for every task outside the context described.

Open conversation
Continue the conversation
Disagree, spot a gap, or have an experience that adds to the subject? Comment with your GitHub account. Do not publish personal data, credentials, or sensitive information.