# Turning an Agent into an AI Employee

> What changes when an agent stops answering tasks and starts owning work.

Muhammad Ahmad · October 1, 2026 · https://devraftel.com/blog/turning-an-agent-into-an-ai-employee

The more we worked on AI agents, the more we realized we were asking the wrong question. We kept asking how to make an agent more capable. The harder question was: what would have to be true before we could actually give that agent a job?

That distinction sounds mostly semantic until you try to build the system. An agent can reason, call tools, and complete a task. But a business does not hire a collection of successful tool calls. It gives a person a role, responsibilities, access, context, and a standard of work, then expects that person to operate within clear boundaries.

We did not begin with a clean definition. We arrived at it by repeatedly hitting problems that the word *agent* did not describe well enough.

> An agent performs work. An employee is assigned work and is responsible for the result.

This is written for anyone trying to make that move with their own agent. It walks the properties an employee needs, roughly in the order you would add them.

## 1. The Agent Is Not the Employee

LangChain defines an AI agent as a system that uses an LLM to decide the control flow of an application. That is a useful engineering definition: the model decides what happens next, uses tools, observes results, and continues until the work is done. Anthropic uses a similar operational framing around agents dynamically directing their own process and tool use. [1][2]

But neither definition tells us who the agent is inside the organization, what it is responsible for, what it is allowed to change, or what happens after the current task ends. Those are not small details. They are exactly the things that make an employee an employee.

So we started separating the two. The agent is the execution mechanism. The employee is the organizational abstraction around that mechanism.

## 2. What Makes Something an AI Employee?

Our working definition became simple: **an AI employee is an AI system given a defined business role, ongoing responsibilities, bounded authority, and access to the context and systems required to perform that work.**

An employee is not a different kind of system. It is an agent plus what the organization has to supply:

```definition
AI Agent    = reasoning + tools + execution

AI Employee = Agent
            + Role
            + Responsibility
            + Identity
            + Context
            + System of Record          (how this company operates)
            + Skills
            + Authority
            + Tasks
            + Supervision
            + Verification
            + Auditability
            + Lifecycle
```

An employee contains an agent. The agent solves execution; everything wrapped around it is organizational.

All of it still has to be built, and most of it is real engineering: identity, task state, audit trails and lifecycle are not paperwork. The point is that none of it arrives from a better model.

Not every employee needs every property in exactly the same form. What matters is that a real employee is more than intelligence plus tools. It is a persistent unit of work inside an organization.

As we worked through the pieces, three distinctions became especially important because they are where otherwise capable systems tend to collapse different concepts together.

### Capability is not skill

This is our taxonomy rather than an industry standard, but it has been useful: a capability is an outcome the employee delivers, such as reconciling an account or closing the month. A skill is the procedure it uses to deliver it. Keeping them separate means you can version a procedure without changing what the employee is responsible for. Merge them and you get one prompt where every capability starts depending on everything else.

### Context is not memory

Context is what has to exist before work can start: the books, the bank feeds, the documents, each authoritative for its own slice. Memory is what the employee remembers from before. Treat them as the same thing and the employee starts guessing at facts it should have been given. Keep guessing and you have built a parallel ledger: convincing, internally consistent, and slowly diverging from what the company actually believes.

### A tool is not authority

Having access to a system is not permission to change it.

These distinctions also give us a useful dependency order for the rest of the system:

**system of record → identity → authority**

The system of record answers whose rules decide. Identity answers who is acting. Authority answers what they may do.

> The simplest test is also the most useful: could you assign this agent a job and reasonably hold it accountable for the work?

## 3. The Shift: Task Execution to Work Ownership

This was the biggest conceptual change for us.

> A task has a beginning and an end. A role does not.

An agent can be told, "reconcile these transactions." An employee can be responsible for keeping a company's books reconciled, notice what is missing, prioritize the next piece of work, handle routine exceptions, and escalate the cases it should not decide on its own.

That does not mean an AI employee must run without humans. It means responsibility is persistent while autonomy is bounded. The employee owns a scope of work; the human still owns the organization.

OpenAI describes agents changing the unit of knowledge work from short interactions to delegated, long-horizon tasks. Anthropic's Cowork is built around the same idea: handing Claude real, multi-step work rather than asking it questions. [3][6]

The infrastructure required to make agents run is increasingly becoming commoditized, so the interesting problem moves toward giving agents organizational responsibility.

## 4. The System of Record

**Memory is what the employee remembers. A system of record is what the organization has decided about how its work is done**, and it wins whenever the two disagree.

The employee asks the operational systems what the numbers are. It asks the system of record what the rule is: the standard for what is supposed to happen, not the data about what happened.

The system of record is the property we underestimated for longest.

A qualified accountant joining a firm already knows accounting. What they learn in their first month is how *this firm* does it. What the capitalization threshold is, the amount above which a purchase counts as an asset rather than an expense. Which clients override the firm default. What needs a partner's sign-off, and above what amount. How small a difference is too small to matter. What order the month-end close runs in.

None of that is in the model, and none of it is in the accounting system. It lives in policy documents, in partner judgment, and quite often only in people's heads.

An employee that does not have it has two options, and both are wrong. It can guess, which produces confident answers nobody can trace to a rule. Or it can ask every time, which means it is not an employee. It is a very fast intern.

So the firm's way of working has to become something the employee can actually use. In practice that means writing it down and giving it the properties it needs to be usable:

- **Published, not draft.** Only approved guidance reaches the employee.
- **Versioned.** So a decision can be traced to the rule that was in force when it was made.
- **Scoped.** Firm-wide by default, with client-specific policy overriding it on the same topic.
- **Reachable at the moment of judgment**, not loaded into a prompt and hoped for.
- **Citable.** The employee says which policy it applied, by name and version, in its answer and in the record.

This is where Anthropic's Agent Skills get interesting. Skills package procedural knowledge and organizational context into something the agent loads only when it is relevant, which is exactly the shape a policy has. A firm's capitalization rule is a short document that matters on the few occasions the question comes up. [8] You do not need to build that mechanism yourself anymore. The interesting part is what you put inside it.

The firm's way of working also includes its controls, and those bind the employee the same way they bind a person: approval thresholds, closed periods where a finished month's books are locked, who may approve what, and segregation of duties, the rule that whoever prepares an entry cannot be the one who approves it. An agent that decides and then acts is maker and checker in the same process, which quietly breaks a control that predates computing. An employee that cannot work inside those controls is not employable at that firm, regardless of how good its work is.

> The useful test: when the employee makes a judgment call, can you find out which rule it followed, and was that rule yours?

## 5. Governance

The system of record is what the employee answers to. **Governance is the other direction: what the organization decides about the employee itself.** Who it is, what it owns, what it may do, who supervises it, and how it can be changed or revoked. Three of those get sections of their own next.

## 6. Identity: Who Is Acting

Most agent products today act as the user. That is a reasonable default for an assistant and a problem for an employee.

For an assistant it is right. That is what user-owned credentials are for: the ticket your agent files carries your handle, not a bot's. [11]

For an employee it is not. Once two parties can act in the same systems, you need to know which one did.

If the employee acts as you, every line in the log is your line. You cannot separate your work from its work, revoke its access without revoking your own, or investigate an incident, because the record says a person did it.

BNY takes the other route: its digital employees have their own logins and human managers. [9]

The chain you want to be able to reconstruct is:

**human → employee → task → tool → action**

Who asked. Which employee acted. Under which task. Through which tool. What changed. Without that chain, attribution, revocation, and investigation are all guesswork.

## 7. Capability Is Not Authority

One of the easiest mistakes is to confuse tool access with permission to act.

An employee may know how to issue a payment and still not be authorized to issue one. The same should be true for an AI employee. The system needs a scope of authority that is separate from its technical capabilities.

The practical form is a small number of tiers.

- **Free.** Reads, reports, drafts, internal notes.
- **Gated.** Anything that changes a record in a business system, every time.
- **Prohibited.** Things the employee must never do even when asked, because they leave the organization. Filing a return, moving money, sending something to a client.

The third tier is the one people skip, and it matters most. If everything is *ask permission*, then *approve* becomes a universal override.

> Keeping prohibited separate from gated is what stops approval from becoming the master key.

BNY's example is useful because it draws the line in a real system. One of its digital employees identifies and patches code vulnerabilities, then submits the change for its human manager's approval. It has the capability all the way through. It does not have the authority to finish. [10]

Anthropic frames the same problem as containment: as agents gain capability and access, the worst case grows even as the likelihood of failure falls, so the job becomes bounding the damage rather than preventing every failure. [5] Which is why authority matters more as the runtime gets better, not less.

## 8. Supervision and Escalation

An AI employee should have a supervisor, even when most of its work is autonomous.

That does not mean a human checking every action. It means a defined supervision model: someone responsible for this employee's work, a stated set of things that escalate, and thresholds that decide which is which.

BNY's structure makes this concrete. Rachel Lewis, its head of payment operations, supervises nine digital employees alongside thousands of human ones. [10]

Verification is how supervision scales past what a person can watch. A probabilistic worker cannot be managed by assuming every successful-looking output is correct, so you need evaluation, observability, guardrails, recovery, and escalation. LangChain calls out observability, evals, memory, permissions, and safe execution as requirements that grow more important as autonomy increases. [7]

Two rules from our own evals are worth passing on.

- **Run a case several times before trusting a prompt change**, because a single sample says very little about a stochastic model.
- **Be honest about which rules are deterministic.** Some can be enforced in code. Others genuinely need judgment and can only be evaluated. Dressing the second kind up as the first, with keyword matching for instance, produces a control that looks real and is not, which is worse than an honest gap, because people stop looking.

> The goal is not zero autonomy. It is bounded autonomy with a named supervisor and measurable performance.

## 9. From Chat to Delegation

Chat is useful, but chat is not the definition of work.

A real employee can receive a request, inspect the relevant systems, perform several actions, stop when a decision requires approval, continue later, and report what happened. That is a task lifecycle, not a conversation.

For us that led to a task-first model. **The task is organization-owned, with an objective, a status, and a history.** It can be linked to a conversation but it does not belong to one, and it survives when that conversation ends. The conversation is context around the work; the task is the unit of execution and accountability.

Two rules earn their place here.

- **Creating a task grants no permission.** A task is the work record, not an authority source.
- **Task status is separate from output status.** A journal entry can be ready for review while the task around it is still in progress.

Conflate those two and a draft eventually gets reported as if it were already in the books.

Anthropic's Cowork shows the same shape from the product side: scheduled and recurring work, workflows across connected tools, and sensitive actions that wait for approval before running. [6][8]

## 10. The Architecture Around the Agent

With the properties established, the architecture falls out of them.

It splits in two: a management or control plane, and an employee runtime.

- **The control plane manages** organizations, humans, employees, roles, skills, permissions, integrations, tasks, policies, evaluations, and lifecycle.
- **The runtime handles** the model, context, tools, state, and execution.

Do not bake the entire meaning of an employee into one container or one agent process.

> The agent is the employee's execution engine, not the employee record itself.

## 11. What We Built

Our own build was the test of the definition. We started where everyone starts, with an agent that could reason, use tools, and finish work. The question was what had to exist around it before we could hand it a job.

We used a Staff Accountant on QuickBooks to find out. Not because an accountant is what an AI employee is, but because it forces every abstraction in this piece to become real: a role someone can write down, books that are authoritative whether you like it or not, permissions that mean something, a firm with its own rules, and a standard of correctness that is not a matter of opinion.

Underneath, it is:

- A portable definition file declaring the employee's identity, purpose, capabilities, skills, required context, authority, supervision, and quality criteria.
- A shared work protocol every employee follows.
- Skills pinned to reviewed versions.
- An approval gate in front of writes.
- An append-only record of proposals and decisions against the task.

## 12. A Practical Blueprint

If we were starting again, we would not begin by building an AI employee platform. We would start with a working agent and add the organizational layers in this order. Each one exists because skipping it produces a specific failure.

| Step | What you add | What breaks without it |
|---|---|---|
| **1. Role and responsibilities** | What it owns, and what is out of scope | It does whatever it is asked. Nobody can say whether a task was in bounds. |
| **2. Skills** | Procedures, separate from the agent and from each other | One prompt holding everything. You cannot version or review a procedure. |
| **3. Context** | The business systems that hold the data | A parallel ledger that slowly diverges from what the company believes. |
| **4. Identity** | The employee as a distinct actor | Every action looks like the user's. No attribution, no revocation, no investigation. |
| **5. Authority** | Free, gated, prohibited. Enforced outside the prompt | Capability becomes permission. An instruction is not a control. |
| **6. System of record** | The firm's policies, thresholds and controls, published and citable | It guesses or asks every time. Neither is an employee. |
| **7. Tasks and state** | Work as an object that outlives the conversation | Work dies with the thread. Duplicate effort, no handoff. |
| **8. Supervision and escalation** | A named supervisor, stated thresholds, an audit trail | Nobody is responsible. Failures surface from the customer. |
| **9. Verification** | Evals, observability, recovery | You trust it, and find out later. |
| **10. Lifecycle** | Provision, change, suspend, retire | An employee you cannot revoke or replace. Every model upgrade is a rewrite. |

> More tools improve capability. They do not create responsibility, authority, or accountability.

## 13. The Test: Can You Actually Hire It?

The term *AI employee* is only useful if it tells you something about how the system behaves. Here is a test. It is deliberately not a yes/no checklist, because almost everything answers yes to those. Each question asks you to show something.

1. **Does it own a role?** Show the written definition, not the system prompt. The document that says what it is responsible for and what it must not do.
2. **Does it have its own identity?** Show one action in the log and say who performed it. If the answer is a person, it has no identity.
3. **Is it grounded in the business data?** Take a number it reported and trace it to the business system.
4. **Does it have bounded authority?** Show the rule that fires when it crosses a boundary, and say where that rule lives. If the answer is *in the prompt*, it is a convention, not a control.
5. **Does it follow your organization's rules?** Take a judgment it made and find the policy it applied, by name and version.
6. **Does it persist?** Show a piece of work that survived a closed conversation.
7. **Who supervises it?** Name the person, and say what escalates to them.
8. **Can you reconstruct what it did?** Show the chain for one action: who asked, which employee, which task, which tool, what changed.
9. **Can you retire it?** Show revocation: what happens to its access, its tasks, and its history when it is switched off.

The more of these you cannot show, the closer you are to an agent or an AI feature. That is not a criticism. Most systems should be agents. Use the word that matches what you built.

> Persona is not employment. Responsibility is.

## 14. From Employees to Workforce

Once you have one employee that works, the next problem arrives quickly: what happens at ten, then a hundred?

You need to provision them, place them in an organization, manage permissions across them, coordinate their work, measure performance, update their skills, investigate failures, and retire them. That is workforce management, and it is where the control plane stops being an implementation detail and becomes the product.

BNY is already there. Its 2025 annual report describes 134 digital employees operating autonomously alongside human colleagues, each with a login and a manager. [4][9]

> The properties do not change at scale. They just stop being optional.

## The Part We Would Keep

If there is one thing we would keep from all the iterations, it is this: do not start with the word *employee* and work backward from a story. Start with responsibility and work forward.

What job are we assigning? What decisions does the worker need to make? What information does it need? Which systems can it touch? What is it allowed to change? Whose rules does it follow? How do we know the work is correct? What happens when it gets stuck? Who can inspect what it did?

Once those questions have good answers, the word *employee* starts to make sense, and most of the work turns out to be writing down things your organization already knew but had never written down before.

## Sources

1. LangChain, [*What is an AI agent?*](https://www.langchain.com/blog/what-is-an-agent) (Jess Ou, July 2026).
2. Anthropic, [*Building effective agents*](https://www.anthropic.com/engineering/building-effective-agents) (December 2024).
3. OpenAI, [*How agents are transforming work*](https://openai.com/index/how-agents-are-transforming-work/) (June 2026).
4. BNY, [*Annual Report 2025*](https://www.bny.com/corporate/global/en/investor-relations/annual-report-2025.html).
5. Anthropic, [*How we contain Claude across products*](https://www.anthropic.com/engineering/how-we-contain-claude) (2026).
6. Anthropic, [*The Future of AI at Work: Introducing Cowork*](https://www.anthropic.com/webinars/future-of-ai-at-work-introducing-cowork) (2026).
7. LangChain, [*Deep Agents overview*](https://docs.langchain.com/oss/python/deepagents/overview).
8. Anthropic, [*Cowork Workshop: Foundations*](https://www.anthropic.com/webinars/cowork-workshop-foundations) (June 2026).
9. Forrester, [*BNY Built Its Digital Workforce Backward — And It's Working*](https://www.forrester.com/blogs/bny-built-its-digital-workforce-backward-and-its-working/).
10. CNBC, [*Digital employees, AI bootcamps: America's oldest bank is spending billions on tech*](https://www.cnbc.com/2026/02/09/digital-employees-ai-bootcamps-americas-oldest-bank-spends-billions-on-tech.html) (February 2026).
11. LangChain, [*Connections: managed credentials and per-caller identity for Managed Deep Agents*](https://www.langchain.com/blog/connections-managed-credentials-and-per-caller-identity-for-managed-deep-agents) (September 2026).

### Further reading

- OpenAI, [*Introducing the Agents API*](https://openai.com/index/introducing-the-agents-api/) (September 2026).
- OpenAI, [*Introducing OpenAI Frontier*](https://openai.com/index/introducing-openai-frontier/) (February 2026).
- Anthropic, [*Scaling Managed Agents: Decoupling the Brain from the Hands*](https://www.anthropic.com/engineering/managed-agents) (2026).
- OpenAI, [*ChatGPT Enterprise and Edu release notes*](https://help.openai.com/en/articles/10128477-chatgpt-enterprise-and-edu-release-notes) (September 2026).
- LangChain, [*LangChain Skills*](https://www.langchain.com/blog/langchain-skills) (March 2026).
- Anthropic, [*Claude Sonnet 5.5*](https://www.anthropic.com/claude-sonnet-5-5) (September 2026).
