This guide is for founders, CMOs and business leaders ready to put a first AI agent to work. It is the second of five parts.
How do you build an AI agent for your business?
Start by writing a one-page job description for the agent: the one job it owns, the tools it may use, which decisions it makes alone and which it hands to a person, and who checks its work. Then assemble five parts: a model, instructions, tools (usually connected through MCP), the knowledge it works from, and a loop that runs it. Use an agent built into software you already have, or configure a general agent with your instructions and documents, before you build one in code. Put guardrails in from day one, test it on about twenty real past cases, and run it beside a person for a month before giving it more freedom.
AI Summary
Building an agent is mostly writing, not coding. Start with a one-page job description: the job, the tools, the judgment it may use and who checks its work. Then assemble five parts: a model, instructions, tools, knowledge and a loop. Use or configure an existing agent before you build one, and build only when the job is central to your business. Put guardrails in from day one: least access, approval for anything that leaves the building, a never-do list, a log and an owner. Test it on twenty real cases before it touches live work, then run it beside a person for a month.
Background
In Part 1, we defined an agent as software that is given a goal and works out its own steps, and argued that the best way to think about one is as a new hire with a narrow job. This part is about turning that idea into a working agent. Part 3 covers where it should run, and Parts 4 and 5 cover what you can buy and the platforms for running your own.
Step 1: Write the Job Description
Every agent that works starts as a page of plain English. Here is a real example of the shape:
Job description: Inbound lead researcher
The job
Within 10 minutes of a new inbound lead, research the person and company
and add a five-line summary to the lead in the CRM.
Tools
Read: the CRM lead record, the company's website, public news, LinkedIn.
Write: one note on the lead record. Nothing else.
Judgment
On its own: what to research and what to include in the summary.
Hand to a person: anything that looks like an existing customer,
a competitor, or a company over 5,000 employees.
Never: email the lead, change the lead's owner or status, or guess at
facts it could not find. If unsure, write "unknown".
The boss
The SDR team lead reviews 10 summaries every Friday.
Good looks like: the rep would not have to open a browser before the call.
Measured by: time to first call, and how often reps edit the summary.
Notice what it does. It names one job. It lists what the agent may read and, separately, the one thing it may write. It spells out what the agent decides alone, what it hands to a person and what it must never do. It names a boss and a measure of good.
Writing this takes an hour. It is the most valuable hour in the project, because every later decision comes from it: which platform, which tools, which data, which tests. It also travels. If you change vendors next year, the job description moves with you.
Step 2: Assemble the Five Parts
The model
The model is the AI that reasons and writes. The leading families, Anthropic's Claude, OpenAI's GPT and Google's Gemini, are all capable of business agent work. Choose on three things: how well it follows long instructions, how reliably it uses tools, and cost at your volume. Larger models are better at judgment and long tasks; smaller ones are cheaper and faster for simple, high-volume steps. Many teams use a large model to plan and a small one for repetitive sub-tasks. Whatever you choose, keep your instructions and knowledge separate from the model so you can switch later.
The instructions
The instructions, often called the system prompt, are the job description rewritten for the agent. Good instructions:
- State the goal and who the work is for.
- Describe what good output looks like, with two or three real examples.
- List the rules: what to do when unsure, what never to do, when to stop and ask.
- Explain the why behind each rule, so the agent can handle cases you did not foresee.
Write them as you would brief a smart new hire on their first day. Vague instructions are the most common reason an agent disappoints.
The tools
Tools are what turn a chatbot into an agent: search, read a CRM record, update a spreadsheet, send a draft, book a meeting. Most agent platforms now connect to business software through MCP, the Model Context Protocol, an open standard introduced by Anthropic and now supported across the major AI platforms. An MCP connector lets an agent use a system, such as HubSpot, Salesforce, Google Drive or Slack, without custom code for each one.
Give the agent the fewest tools the job needs, and prefer read access over write access. Every tool is something it can get wrong.
The knowledge
Knowledge is what the agent works from: your website, product documents, case studies, pricing rules, playbooks, past examples. Most bad answers come from thin or outdated source material, not from the model. Writing down what your best people know is often the real work of building an agent, and it pays off for your people too.
Keep knowledge current and in one place. A website written clearly enough for an AI to read, with an llms.txt file and plain-text versions of key pages, doubles as knowledge for your own agents. MachineReady checks how well yours reads.
The loop and memory
The loop is how the agent works: take a step, look at the result, decide the next step. Platforms handle this for you; what you decide is how long it may run, how much it may spend and when it must stop and ask. Memory is what it carries between runs, such as past decisions, preferences and notes. Give it memory only for what helps the job, and remember that whatever it remembers, it may repeat.
Step 3: Use, Configure or Build
You rarely need to start from code. Climb only as far as the job needs:
Use when the job lives inside one product you already pay for. Salesforce, HubSpot, Microsoft 365, Google Workspace and most help desks now ship agents. Part 4 surveys what exists.
Configure when the job crosses a few tools but follows your playbook. A general agent, such as a Claude project or a custom GPT, given your instructions, your documents and a few connectors, is often enough. No-code agent builders sit here too.
Build when the job is core to how you win, faces customers, or needs your own systems and data. Building today means a model API or an agent platform that runs the loop for you. Part 5 covers the options, from Claude Managed Agents and the Claude Agent SDK to the platforms from OpenAI, Google, AWS and Microsoft.
Whichever rung you choose, keep the job description, instructions and knowledge as your own assets in your own files. They are what make the agent yours.
Step 4: Put Guardrails In From Day One
Guardrails are not a later phase. Build them in from the first version:
- Least access. Only the data and tools the job needs. An agent that writes reports does not need to send email.
- Approval steps. A person approves anything that spends money, cannot be undone or speaks to a customer, until the record says otherwise.
- A never-do list. Short, specific and tested: never quote a price, never promise a date, never email outside the company.
- Behavior when unsure. Tell it to say “unknown” or ask, never to guess.
- A log. A record of what it did and why, so you can review its work and find the cause of mistakes.
- An owner. One named person who reads its work and improves its instructions every week.
- Protection from hidden instructions. Agents that read email, web pages or uploaded files can meet text written to hijack them, known as prompt injection. Do not give an agent that reads untrusted content the power to send data out or take sensitive actions without approval.
Step 5: Test Before You Trust
Before an agent touches live work, collect twenty real cases from the last month, including a few awkward ones. Run the agent on all of them and compare its output with what your best person did. Note every miss and its cause: missing knowledge, unclear instructions, a missing tool or a bad call. Fix the cause, then run all twenty again.
Keep those cases. Every time you change the instructions, the model or the tools, run them again. This set of test cases, often called an eval, is how you know a change made the agent better rather than just different.
Step 6: Your First 30 Days
One agent, one job, one month. The aim is not a platform; it is one agent people actually use.
- Week 1: pick the job and write the job description. Choose a task that happens often, follows a pattern and has an obvious “good” answer.
- Week 2: gather what it needs. Collect the knowledge and the two or three tools. Fix the gaps you find; they were already hurting your people.
- Week 3: run it beside a person. The agent drafts, a person approves. Keep every case where the person changed the draft.
- Week 4: measure and decide. Compare time and quality against the old way. Then widen what it may do alone, move it closer to customers, or stop.
How We Built Ours
Our own agents followed these steps. The A2A agent that answers other companies' AI agents on buildmarketing.ai started as a page listing what it may discuss, the facts it may share and what it must hand to a person. Each of its skills is switched on deliberately in our admin, so adding a capability is a business decision, not a code change. Its knowledge is the same set of articles and pages our site chat uses, so improving one improves both. You can watch it working.
The agent behind MachineReady proposes new checks and recommendations from what it learns, and none of them appear in a public report until I approve them. That is level 1 autonomy by design, on the part of the job where a mistake would reach customers.
Test yourself
5 quick questions. Pick an answer to see if you are right.
1. What should you write before choosing any agent platform?
Show answer
The answer is A one-page job description for the agent. The job description decides everything else: the tools, the data, the tests and which platform fits. It also moves with you if you change vendors.
2. What is the most common cause of an agent giving bad answers?
Show answer
The answer is Thin or outdated knowledge to work from. Most bad answers come from missing or stale source material. Writing down what your best people know is often the real work.
3. What does MCP (the Model Context Protocol) do?
Show answer
The answer is Lets agents connect to business tools in a standard way. MCP is an open standard that lets an agent use systems such as a CRM, Google Drive or Slack without custom code for each one.
4. When should you build your own agent rather than use or configure one?
Show answer
The answer is When the job is core to how you win, faces customers or needs your own systems. Most first agents should be used or configured. Build when the job is central to your business or needs your own data and systems.
5. Before an agent touches live work, what should you do?
Show answer
The answer is Test it on about twenty real past cases and fix the causes of misses. A set of real test cases shows where it fails and why, and lets you check that every later change makes it better, not just different.
Now put it to work on your own website. Run MachineReady to see how AI agents read it, or send our A2A agent to see if it can talk to theirs. Both are free.
Conclusion
How do you build an AI agent? Write the job description first. Assemble a model, instructions, tools, knowledge and a loop. Use or configure before you build. Put guardrails in from day one, test on real cases before you trust it, and give it an owner who improves it every week.
Next in the series: Where Should Your AI Agents Run?.
Your agent is only as good as its knowledge. Run MachineReady on your website to see whether your content is ready for agents to work from, or get in touch for a second opinion on your agent's job description.
← Back to all articles