AI Agent Development: Agents That Take Real Actions, Safely

When an agent is warranted and when a fixed workflow is better, and how we design tools, approvals, sandboxes, scenario evals and audit logs for agents that act.

An agent loop circling a model through tool stations and passing a latched gate where risky actions await approval.

You want software that does more than answer questions. It should look things up, decide what to do next and take actions in your systems: update a record, issue a refund, reschedule an appointment, open a pull request. AI agent development is the work of giving a model that ability while keeping each action bounded, reviewable and, where it matters, reversible.

A good agent is uneventful in production. It works with a small set of well-designed tools, asks a person before anything consequential, stops and escalates when it is stuck instead of improvising, stays within a cost budget and leaves a record detailed enough to reconstruct exactly what it did and why.

Agent or workflow: deciding what you need

The word agent covers two very different designs. In a workflow, your code decides the sequence of steps and calls a model at specific points, for example to classify a request or draft a reply. In an agent, the model decides the next step in a loop, choosing among tools until it judges the task complete.

Aspect Fixed workflow Agent
Who decides the next step Your code The model, within limits you set
Predictability High: the same input follows the same path Lower: paths vary between runs
Cost and latency Bounded and easy to estimate Variable, so budgets must be enforced
Testing Conventional tests for each step Scenario suites, run many times
Best for Processes you can draw as a flowchart Tasks whose path depends on what is found along the way

If you can draw the flowchart, build the workflow. It is cheaper, faster and far easier to test, and most business processes fit. Agents earn their place when the route cannot be enumerated in advance: investigating a support case across several systems, reconciling records that disagree in unpredictable ways, working in a codebase.

Often the best design is a workflow with one agentic step inside a tightly bounded box. And if a wrong action would be expensive and nobody could review it in time, the task should not be automated yet.

Designing tools the model can use well

Tools are the agent’s interface to your business, and most agent quality problems turn out to be tool design problems. We design tools the way we would design a public API for a capable new colleague who has never seen your systems:

  • A few focused tools rather than one that does everything, each with a clear name and a description of when to use it.
  • Typed parameters with enumerations and validation, so an invalid call fails fast with a message the model can act on.
  • Compact, structured results with pagination, instead of raw database dumps that flood the context window.
  • Read tools kept separate from write tools, so their permissions can differ.
  • Idempotent writes with idempotency keys, so a retry never issues the same refund twice.
  • Preview modes that return what would change without changing it.

Validation happens on the server regardless of what the model sends. We expose tools through native function calling or open protocols such as the Model Context Protocol, and we do not hand an agent raw SQL or a shell outside a sandbox.

Permissions, human approval and sandboxing

An agent should act with the permissions of the person it serves, through scoped, short-lived credentials, never through an all-powerful service account. On top of that, we classify every action by impact and reversibility:

Reads pass freely, writes are logged and risky actions wait at a locked gate, beside an isolated sandbox.
  • Reads run automatically.
  • Low-impact, reversible writes, such as adding a note or a tag, run automatically and are logged.
  • Consequential or irreversible actions, such as payments, messages to customers and deletions, wait for approval. The approver sees exactly what will happen, with the real parameters, not a summary written by the model.

Limits such as maximum amounts or allowed recipients are enforced in code, where the model cannot talk its way past them. Agents that run code or browse the web do it in isolated containers or virtual machines with no production credentials, restricted network access, resource limits and disposable file systems.

Prompt injection matters most for agents, because they can act on what they read: a web page, email or document can contain text written to redirect the agent. Tool results are data, not instructions, and content from untrusted sources must never be able to trigger a privileged action without a person in the loop.

State, memory and failure handling

Run state belongs in your database, not only in the model’s context: the task, completed steps, tool results and pending approvals. Checkpoints make a run resumable after a crash, or after an approval that took a day. Long-term memory, meaning facts an agent keeps between sessions, is stored deliberately, scoped to a user or tenant, visible to them and deletable, with an expiry where appropriate.

An agent run with checkpoints filed to a state store, a loop cut short, a step budget dial and a hand-off at the end.

Failure paths are designed from the start:

  • A maximum number of steps and a wall-clock timeout per run, plus a timeout on every tool call.
  • Retries with backoff for transient errors, made safe by idempotent tools.
  • Loop detection for when the agent keeps calling the same tool with the same arguments.
  • An explicit way to give up: stop, summarize what was done and what is blocking, and hand the case to a person.

An agent that reports partial progress honestly is worth more than one that guesses its way to an answer.

Evaluating agents with scenario suites

Tools get conventional unit tests. The agent itself is evaluated with scenario suites: realistic tasks run against a simulated environment, such as a test CRM with seeded customers and orders, where actions have no real consequences. Each scenario is graded on the end state (was the right refund issued to the right account?), on the path (did it ask before paying, and did it stay away from tools it should not touch?) and on efficiency in steps, time and tokens.

Agents do not behave identically on every run, so each scenario runs several times, and we track pass rates across versions of prompts, tools and models. The suite includes adversarial cases: missing data, failing tools, contradictory records and documents carrying injected instructions. Every production failure becomes a new scenario.

Before an agent acts on its own, it runs in shadow mode, proposing actions that people carry out, until its proposals consistently match their decisions.

Cost control and audit logs

The orchestrator, not the model, enforces budgets for tokens, tool calls and time per run. We use smaller models for routine steps and larger ones for planning, trim and summarize context as runs grow, and cache tool results that cannot change within a run. The figure we report is cost per completed task, because a cheap run that fails is not cheap.

Every step is written to an append-only audit log: the input, the model and prompt version, each tool call with its arguments and result, each approval with the approver’s identity and time, and the outcome. Correlation IDs connect agent actions to your own system’s audit trail. In FinTech, MedTech and LegalTech products, this record is often what makes an agent acceptable to a compliance team at all, and it turns debugging into replaying a run instead of guessing.

How an AI agent development engagement works

We start with one job, not a platform. Our business analysts map how people do that job today, list the systems and actions involved, classify each action by risk and agree with you on an approval policy and a definition of success.

Engineers build and test the tools first, then a workflow version, and add agentic steps only where the workflow cannot cope, while QA builds the scenario suite alongside. Rollout moves from shadow mode to approval mode to autonomy for low-risk actions, as evaluation results justify each step. The model integration underneath follows the practices in LLM integration, and the agent usually sits inside a larger product, as our overview of AI app development describes. You receive the code, tools, scenario suite, dashboards and runbooks, with regular demos and written reports throughout.

Frequently asked questions

What is the difference between an agent and a chatbot?

A chatbot answers. An agent acts: it calls tools that read and change data in your systems, which is why it needs permissions, approvals, budgets and audit logs that a chatbot does not.

Can an agent work with our internal systems?

Yes, through APIs wrapped as tools. Where a system has no API, building a small one is ordinary backend development, and it is usually more reliable than an agent operating a user interface, which is slower and breaks when screens change.

Will the agent act without approval?

Only for actions you have classified as safe to automate. The approval policy is yours, it is enforced in code, and it can be tightened or relaxed per action as confidence grows.

Which agent framework do you use?

We choose per project and keep orchestration thin, often plain code around the provider’s tool-calling interface. Frameworks change quickly, while your tools, permissions, logs and scenario suite are the lasting assets, so we keep those independent of any framework.

How do we know the agent is ready?

When it passes its scenario suite consistently, agrees with human decisions in shadow mode, stays within its budget per task and fails by escalating rather than by acting wrongly.

If there is a job in your business whose path no flowchart can capture, it may be a good candidate for an agent, and if a flowchart can capture it, we will say so. Tell us about the task you want to automate, and we will propose a design with the right amount of control.