You find an excellent skill. You install it.
You find another that seems to do the same thing, but has a better review step. You install that one too.
Then a third appears, with clearer examples and a description that seems written for your problem. Now you have more options and one more question.
Which one should I use?
A skill is a package of reusable instructions and resources that guides a class of tasks. An AI agent is the system that receives context, chooses steps and uses tools to carry them out. Having several skills available should expand an agent’s capabilities. Without clear criteria, it can also expand the confusion.
That management problem led me to build I-9 Skills, a public project from I-9.ai for discovering, creating, evaluating, organizing, maintaining and distributing skills. I use the tool in my local environment and opened the project to anyone who wants to experiment, evaluate their own skills and contribute improvements.
I also built the I-9 Skills website, available in Portuguese, English and Spanish, to introduce the library. The code and documentation are in the public repository.
My aim is to draw on the best contributions in the skills we encounter — the crème de la crème — and turn them into better packages for a defined job. Choose what works, understand why it works, combine it when appropriate and be able to improve it later.
Installing a skill is easy. Knowing whether it is still the right one takes work.
Three skills for the same job. No clear decision
Imagine you want to review an application. One skill prioritizes architecture. Another looks for security problems. A third is also called “review,” but was written to check code style.
The names look equivalent. The deliverables are different.
If the agent chooses the label closest to your request, you may get an impeccable answer to a question you never asked. Loading everything can mix competing criteria and consume context — the material available to guide the model — without clarifying the work.
I-9 Skills treats discovery and routing as separate responsibilities. Discovery finds and qualifies candidates. Routing chooses the procedure suited to the goal and constraints of that execution.
The skill-routing entry point can recommend one skill, a short sequence with distinct deliverables, a limited list of alternatives that need clarification, or none when no skill is necessary. The catalog helps find candidates; choosing requires reading SKILL.md, the file containing the package’s instructions and boundaries.
This returns to a question I explored about available skills that never reach the work: people should not have to memorize an invisible menu for the right capability to show up.
The skill that does everything, including what you never asked for
Another scenario: you request an analysis, but the skill also contains research, implementation, testing, installation and publication. Everything sits in the same package, with a description broad enough to fit almost any conversation.
It has become a small department. It just forgot to define the jobs.
That can make selection ambiguous: the whole procedure may not apply, even though part of it is useful. It also makes it harder to see where analysis ends and an action requiring authorization begins.
In I-9 Skills, one responsibility per skill and one main reviewable deliverable guide the design. skill-design defines the job; skill-authoring produces the package; skill-evaluator evaluates behavior. Auditing and refactoring — reorganizing while preserving what must keep working — help identify overlaps and propose divisions.
The idea is to let you choose, test or fix one piece without rewriting the whole collection. That is also why the library uses meta-skills: skills whose job is to care for other skills and their collection.
In another agent, the same skill became something else
Consider a skill installed in a project, another copy in the user’s environment and an adaptation in another agent. All carry the same name. One received a fix. Another kept an old rule. The third gained a step specific to that application.
When behavior diverges, where do you start looking?
“But I already fixed that” becomes a rather unhelpful sentence when nobody knows which copy did the work.
The project’s approach combines catalog, provenance and history. The catalog makes the collection inspectable; records distinguish sources and revisions, even when names coincide. Installation and migration have their own responsibilities, so moving a package does not mean losing its resources or origins.
The skills are written in English and follow the Agent Skills specification. Each package carries its Apache-2.0 license, references, resources and applicable scripts — small helper programs. Those resources are located relative to the skill’s own directory; results go to the workspace chosen by the person requesting the work.
This self-contained structure lets you take the package to another environment without depending on a file left on its creator’s machine. Required tools and configuration still need to be declared and prepared.
The skill was born with limited context. And stayed trapped there
A skill may have been excellent for the problem that inspired it. Then exceptions appear, documentation changes, you find a better approach or the agent fails in a situation nobody had anticipated.
If the correction stays in the conversation, the next agent may repeat the error. If each attempt changes the skill without comparison, it becomes hard to tell whether you improved the procedure or merely accommodated the last case.
This is where I want a process for evolution.
skill-evidence-collection organizes evidence; skill-evolution produces a candidate based on it; skill-optimization works on measurable improvements in bounded iterations. Evaluation compares behavior against cases and criteria defined before the verdict. Adopting the new candidate remains a separate decision.
A snapshot, an identifiable copy of a state, helps preserve the previous version. skills-snapshot handles creation, verification and restoration; skills-maintenance-scheduling prepares recurring maintenance proposals. You can organize reviews and keep a way back without turning a calendar into authorization to change everything.
That is the same concern behind learning without rewriting the past: a failure can justify an investigation. The investigation needs to support the change.
Several skills have good ideas. Do I have to choose just one?
Not always.
Imagine one skill with useful questions for gathering requirements, another with good test cases and a third that clearly explains when to stop. They make different contributions to the same goal.
That is the kind of reuse I want to enable. skills-discovery qualifies sources; skill-domain-research researches the subject through reliable references and the knowledge of whoever owns the process. When two or more skills contribute distinct elements, skills-synthesis organizes what to combine, what to discard and why.
Synthesis is not pasting three texts together and hoping length turns into quality. It means preserving useful procedures, license compatibility, references and boundaries, resolving contradictions before designing the new skill. A single contributing source can proceed to design without an artificial synthesis; subject research is still required.
For me, “the best of the best” expresses that ambition to select and test contributions. It is not a certification that we have found the best skill in the entire market.
Installed, forgotten or triggered at the wrong time?
A skill can be available and never chosen. Another can show up where it should not. A third can be read frequently without producing a useful deliverable.
These problems call for observation. That also led me to work on hooks, integrations triggered by application events, and telemetry, bounded records of what was observed.
Session hooks can provide a compact map of available skills. Read observation is optional and depends on the host, configuration and chosen storage. The documented Claude adapter distinguishes attempts from successful native reads; Codex hooks provide context without promising native read counts.
To investigate consumption and recurrence, we need to separate the questions:
- Was it found or read? That signals availability or consultation.
- Was its procedure used? That is an activation assertion requiring its own record.
- Did the work finish, and was the result good? Recorded completion and evaluated quality are different evidence.
Lifecycle records allow selection, activation and outcomes to be recorded with source identity and coverage boundaries. Those signals are reported by whoever records them. A read count does not demonstrate understanding; an absent record does not demonstrate an absence of use.
This helps formulate better questions about effectiveness. It does not turn frequency into quality or make a skill evolve on its own.
A complete package, with an identity for each piece
Management spans discovery, research, design, synthesis, authoring, evaluation, security, cataloging, evolution, installation and publication. Each responsibility has its place.
The skills follow guidelines and structural templates for declaring inputs, outputs, boundaries, resources and checks. Validators check format and additional collection criteria. Behavioral evaluation looks for errors that a well-formed file does not reveal.
There is also skill-icon-design, dedicated to creating icons. The collection’s packages include distinct icons and Codex interface metadata, allowing visual presentation in clients that support the feature. When you create your own skills, the workflow can produce that identity too.
Yes, the nice little icon is included. It helps you recognize the piece. You still need to find out whether it does the job properly.
The core is agent-agnostic: following the procedures does not require a specific vendor. Integrations add capabilities according to the host, the application running the agent. There are manifests for Codex, Claude Code and Copilot; discovery, interface, hooks and permissions need to be checked in the chosen environment. Package portability does not promise identical behavior across clients.
Experiment with a problem you already have
You can use the project as portable skills or as a plugin, which adds the collection and its integrations to the application. There is also a CLI, a terminal command interface, for the catalog, validation and management tools, plus an MCP server, using a protocol that connects AI applications to tools and resources.
To install the collection in a project, the README uses the external Skills CLI:
1
npx skills add i-9-ai/skills --yes
Here, --yes selects the available skills and skips the selection steps. The --global option changes the destination to the user scope. Check the scope and the agent chosen by the installer.
Alternatively, for Codex, with the README prerequisites prepared:
1
2
3
codex plugin marketplace add i-9-ai/skills --ref main
codex plugin add i9-skills@i9-skills
codex plugin list
The plugin installation guide covers enabling, updating and removing the plugin. Restart the session or client as directed.
Our CLI is a different package: @i-9.ai/skills. Node.js provides the environment for running these programs, and npm manages their packages. With the published prerequisites prepared, npx lets you run the CLI without a global installation; the first invocation may download the package and its dependencies:
1
2
3
npx --yes @i-9.ai/skills --help
npx --yes @i-9.ai/skills catalog overview
npx --yes @i-9.ai/skills catalog read --skill skill-authoring
These commands query the collection bundled with the CLI. They do not install skills in the agent. To discover packages in the current project without including global ones:
1
npx --yes @i-9.ai/skills context available-skills --project . --no-global
With the skills available in your agent, start with a bounded request:
Use
skills-auditto analyze my collection. Identify skills with overlapping responsibilities, divergent versions, missing resources and descriptions that make it difficult to choose the right procedure. Show the evidence and propose improvement priorities. Do not change files, install or publish anything at this stage.
Then choose one audit finding. If different sources contain useful contributions, request a synthesis plan. If the problem is an overly broad responsibility, request a split. If a skill is failing, bring concrete cases to its evaluation and evolution.
The project is experimental. I want you to be able to try a small target, preserve the previous state and evaluate the result in your own environment.
If that fits your collection, explore the I-9 Skills website, try the library and share what you found: the problem, the host you used, what you expected and what happened. A minimal example without credentials, personal data or client material helps turn feedback into a verifiable improvement.
A good skill deserves to be found, chosen in the right context and improved when reality reveals its limits. I built this library to care for that path.
Further reading
- The model is just the engine: places skills in the system surrounding the model.
- Agent Skills specification: explains the portable structure of instructions and resources.
References and limits of use
- I-9 Skills website: the library’s public introduction in this article’s language.
- Repository and README: source for commands, responsibilities, interface resources and distribution paths. Check prerequisites and integrations before installing.
- Wiki, especially validation and source research: documents the method and checks. Structural compliance does not certify production quality.
- Hooks and lifecycle evidence: define observation, configuration and declared records. They do not prove activation through reading or effectiveness through frequency.
- Native pilots: bounded integration checks, without universal certification of hosts or model behavior.
The situations in this article illustrate the problems that motivated the project. My local use is a usage account; it does not replace independent evaluation or demonstrate a quantitative productivity gain.

Open conversation
Continue the conversation
Disagree, spot a gap, or have an experience that adds to the subject? You can comment without creating an account or use one of the available sign-in methods. New comments may be moderated; when published, they are public. Do not publish personal data, credentials, or sensitive information.
When comments load or are submitted, technical data may be processed as described in our privacy policy. Privacy.