An AI agent skill is a reusable workflow packaged as a directory: a required SKILL.md file explains what to do, and optional files can provide references, templates, assets, or Python scripts. Start with clear instructions and add code only when it makes a step more reliable or repeatable. Then test both whether the skill is selected for the right requests and whether it produces the results you specified.
What an agent skill contains
OpenAI’s Skills guide describes a skill as a directory of files that includes a SKILL.md. The file combines identifying information with instructions for carrying out a workflow. A skill can also include supporting material, such as reference documents, scripts, assets, or test fixtures.
As an Amazon Associate I earn from qualifying purchases.
A minimal instruction-only bundle might look like this:
summarize-reports/
└── SKILL.md
If the workflow needs a Python helper and sample input, the directory could instead look like this:
#1 Best Overall
summarize-reports/
├── SKILL.md
├── run.py
├── requirements.txt
├── references/
│ └── output-format.md
└── assets/
└── example.csv
These are illustrative layouts, not mandatory filenames or a universal project scaffold. Include only the files the workflow actually uses.
Write the skill’s name and description first
The front matter in SKILL.md identifies the skill. Choose a concise, distinctive name, then describe both its task and the situations in which it should be used. That description helps the agent decide whether the skill fits a request. OpenAI’s article on systematically testing agent skills highlights the name and description as important invocation signals.
Rank #2
---
name: summarize-reports
description: Summarize a supplied CSV report into key trends and anomalies when the user asks for a concise analysis of tabular report data.
---
This example names a specific task and gives a recognizable trigger. A description that is too broad can invite the skill for unrelated requests; one that is too narrow may miss requests it is meant to handle. Avoid bundling unrelated workflows under one description.
Free tools Windows power users keep installed
One-click scans. No signup required.
Write instructions with observable completion checks
Keep the main workflow in SKILL.md. State what input the agent should expect, what steps to follow, what output to produce, and how to tell whether the task is complete. When a reference document or template is useful but would make the main instructions unwieldy, put it in a supporting file and point to it from the skill.
- Define the input. Specify the relevant file, data, or information the user must provide, and how to handle missing or malformed input.
- Describe the steps. Give the workflow in a practical order. Separate decisions the agent must make from deterministic work a script can perform.
- Specify the output. Name the format, required fields, and any constraints, such as whether to include units or cite rows used in a calculation.
- Set completion checks. Describe verifiable conditions—for example, that the output contains all required fields and that totals reconcile with the input.
Make the instructions concrete enough that you can evaluate them. “Analyze the file carefully” is difficult to test; “report the three largest categories, their totals, and the calculation basis” provides observable expectations.
Decide whether Python belongs in the skill
Python is optional. Use it when a workflow benefits from repeatable computation, file conversion, validation, or another deterministic operation. If the task is mainly interpretation or writing and a script would add setup without making the result more dependable, keep the skill instruction-only.
| Approach | Use it when | What to include |
|---|---|---|
| Instruction-only | The workflow is adequately expressed as steps and does not need a deterministic helper. | SKILL.md, plus only useful reference material or templates. |
| Script-backed | A repeatable computation or transformation should be performed consistently by code. | SKILL.md, the script, any dependency declaration it needs, and relevant fixtures or assets. |
When adding Python, make the invocation and working directory explicit in the instructions. Identify required dependencies and explain what the script reads and writes. Keep example data and fixtures separate from real user data. The OpenAI cookbook’s API skills example shows a bundle containing SKILL.md, run.py, requirements.txt, and a CSV asset; its packages and commands are specific to that example, not prerequisites for other skills.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Test invocation and results separately
A skill can produce good output when explicitly invoked and still have a weak description that causes it to be selected at the wrong time—or not selected when needed. Build a small evaluation set with intended triggers, non-triggers, and observable output requirements. OpenAI’s skill evaluation guidance describes evaluating skills systematically rather than relying only on one successful example.
Best Value
| Test case | Example request | What to check |
|---|---|---|
| Positive trigger | “Summarize the trends in this CSV report.” | The skill is selected, follows its analysis steps, and returns the specified fields. |
| Negative trigger | “Explain what a CSV file is.” | The skill is not selected if the request does not ask for analysis of supplied report data. |
| Output behavior | Provide a valid report that includes an obvious category total. | The result includes required fields and its reported total matches the input. |
| Failure behavior | Provide a file with a missing required column. | The skill follows its instructions for invalid input rather than silently inventing a value. |
For each case, record the request, whether the skill was invoked, and whether the output passed each check. A few manual spot checks are useful for finding obvious defects; a defined set makes comparisons more repeatable when you revise the description or instructions. Passing these cases is evidence about those cases in that setup, not a guarantee of behavior across every model or environment.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose the setup for the environment you will use
Packaging and discovery depend on where the skill will run. OpenAI’s API documentation distinguishes local execution from hosted, container-based use; in the Agents API, skill directories are discovered through configured capability directories. The cookbook’s script-backed API example likewise describes its own API-oriented setup. Do not assume that a local folder is automatically available to a hosted agent, or copy setup steps from one surface into another without checking the relevant instructions.
Before testing, confirm that the target environment can discover the directory and any supporting files, and that its execution context can run the Python command your skill specifies. Run local checks first, including script and fixture checks where applicable. If evaluation involves API requests, make those requests deliberately; the cookbook example explicitly opts in to API use after local checks.
The official guidance and examples are starting points, not proof that a particular skill works. Verify invocation, output requirements, and failure handling in the environment where you intend to use it, and report only results you actually observed.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




