Skip to content

AI Transformation

What Is an Agentic Workflow? Harness First

By Robert Antolin · · 10 min read

An agentic workflow is a process where an AI model decides what happens next. Instead of following a checklist someone wrote in advance, the model reads the situation, picks a tool, runs it, looks at the result, and repeats until the job is done. Think of the difference between answering one pop quiz question and completing a full science fair project. A normal AI prompt is the quiz question: one answer, in one breath, and if part of it is wrong the answer ships anyway. An agentic workflow is the project: plan, act, check, fix, repeat.

The model is the smallest part of an agentic workflow. Almost everything that decides whether it works is a decision about what surrounds the model: what it can see, which tools it can use, what its standing instructions say, and what happens when it gets something wrong. Engineers call that surrounding machinery the harness. How much harness do you need on day one? Three pieces. Add a fourth only when a logged failure shows you need it.

What is an agentic workflow?

An agentic workflow is one where the model chooses the next step and the tool for it, then loops until a stopping condition is met, rather than filling one slot in a fixed sequence a person wrote. Automation with a model call inside it is not agentic: the model there is one part in a path someone else decided.

The loop, in one line:

Goal -> Plan the steps -> Act (use a tool) -> Check the result -> Done? If not, fix and repeat

How that works in practice:

  • Planning. Instead of guessing the whole answer at once, the model breaks the goal into smaller jobs.
  • Tool use. It does not rely on memory alone. If it needs current information it searches; if it needs exact math it runs a calculation; if it needs a record it queries the system that holds it.
  • Checking. After each action, the workflow inspects the result. If a test fails or two numbers disagree, it tries a different approach instead of carrying the error forward.
  • Repeating. It loops through those steps until the result meets the goal, then stops.

Three tiers, so you can place your own case:

  1. Fixed script. Every step is written in code and runs in the same order every time. The same input always takes the same path. Engineers call this deterministic, which just means predictable. A scheduled report or a Zapier flow lives here.
  2. AI-assisted step. The path is fixed and a model fills one slot in it: sort this support ticket, pull the total from this invoice, summarize this contract.
  3. Agentic loop. The model chooses which tool to call and when to stop, so the path can differ every run. You gain the ability to handle cases nobody planned for. You give up knowing in advance what will happen.

A quick comparison, using an invoice check as an illustration:

  • Non-agentic. You ask an AI whether an invoice matches its purchase order. It reads the invoice once and answers yes or no from that single look. If the order number is smudged, it guesses.
  • Agentic. The workflow reads the invoice, finds that the order number matches nothing, searches the order system by supplier and date, finds the likely order, compares the two line by line, flags the lines that differ, and hands the flagged pairs to a person to confirm. The person, not the model, makes the call on whether to pay.

Most processes belong in tier one or tier two. Which tier a process belongs in is the subject of the sibling piece on when not to use an agentic workflow, which carries the five-question test operators run before they build anything. Run that test first. Everything below assumes it came back yes.

What has to exist around the model for it to work?

Two published component lists answer this, and they disagree on size. Both come from parties with something to sell, which is worth holding in mind while you read them.

Three components, from Barry Zhang of Anthropic, in a talk on how they build effective agents:

  • The environment: the systems and data the model can see and act on.
  • The tools: the specific actions it is allowed to take, such as searching, reading a record, or running code.
  • The system prompt: the standing instructions the model reads before every run.

The talk notes that Anthropic's own three internal agent products share almost the exact same code. Anthropic sells the tokens a harness consumes. Tokens are the units of text a model reads and writes, and they are how AI usage is billed.

Six components, from the MongoDB authors writing on the agent harness:

  • State and persistence: remembering where a run got to, even after a crash.
  • Security and governance: rules about what the workflow may touch and change.
  • Orchestration and tools: the code that runs the steps and the actions the model may take.
  • Memory: what carries over from one run to the next.
  • Observability: logs a person can read after the fact to see what happened.
  • Evals: test cases that score whether the output was right.

MongoDB sells several of those platform pieces.

The same authors offer a rule of thumb, and they present it as a rule of thumb rather than a measurement: about one line of glue code (the ordinary code that connects the pieces) per model token for a weekend chatbot, and about fifty for a governed production platform. Use it to see the shape of the gap between a demo and a production system. Do not use it to estimate your own build.

How much harness does an agent need on day one?

Three. The two lists answer different questions, which is why they do not match. Three is the first-build answer, the minimum that makes a loop run at all. Six is the at-scale answer, what a governed platform serving many workflows ends up carrying.

Start with three, log every run in full, and pull in a fourth component when a logged failure names it. Memory earns its place when runs demonstrably need to remember. A governance gate, a rule in code that stops the workflow from writing where it should not, earns its place when something writes where it should not have. Cost tracking earns its place when a bill surprises you. Logs are cheap on day one; the platform can wait.

This is also Anthropic's published guidance: find the simplest solution possible and add complexity only when it is needed, and most successful implementations use simple, composable patterns (small pieces that plug together) rather than pre-built agent frameworks.

Which workflows deserve an agent at all, versus a script, is the first output of an AI readiness assessment, and it is cheaper to answer there than in a build.

Does the harness change results, or only cost?

Both, and the evidence points in more than one direction. Here it is in one place. A benchmark, where the word appears below, is a standard set of test tasks used to compare AI systems.

  • Vercel cut about 80% of its agent's tools, roughly 15 down to two (running commands and running database queries), on the same model. Task success went from 4 of 5 to 5 of 5, and tokens per task fell 37%, about 102k to 61k. That is five tasks and one team's before-and-after, not a benchmark. Treat it as a result worth testing against your own tool list rather than a number to expect.
  • LangChain's coding agent went from 52.8% to 66.5% on Terminal Bench 2.0, a public coding benchmark, with harness changes only and the same model (gpt-5.2-codex) throughout. A benchmark score, not a production outcome.
  • Princeton's CORE-Bench: the same model scored 42% under one harness and 78% under another. Read it for direction: the harness moves the score. Do not read either number as a target.
  • Counter-evidence, and it belongs in the same breath. Scale AI's SWE-Atlas found harness choice within the margin of error for some model families, meaning models from some makers barely cared which harness they ran in. METR found that established coding harnesses do not consistently beat a basic one.
  • NOOA, from NVIDIA labs, reports that the harness advantage is widest when the model's reasoning is turned off (reasoning is the setting that lets a model think longer before it answers) and narrows as reasoning effort rises. Their comparison used different budgets and the authors call the result indicative, so take the direction and not a size.

The rule that survives all of it: the harness matters most where the model is weakest at the task. Where the model is already strong, elaborate scaffolding may buy nothing.

Which patterns keep agentic workflows reliable?

Five, in the order they tend to matter:

  1. Fixed code for the routine work, the model only for judgment calls. Ordinary code fetches the data, moves it between systems, saves the results and retries failures, the same way every run. The model is called only where something has to be read and understood, like sorting an invoice that matches nothing.
  2. A human gate that verifies rather than approves. A human gate is a point where a person checks before anything goes out. A human in the loop shown the exact changes, a confidence score (how sure the model says it is), and the source record is doing verification. One shown a summary and an approve button is rubber-stamping.
  3. Progress saved outside the conversation. Write the run's state to a database or to files, not into the chat context, which is the running text the model can see and which gets cut off when it grows too long. Anything needed on run two has to survive run one.
  4. A hard limit on what the model may change. Read broadly, write narrowly. The set of records the workflow can modify should be short enough to write down, and enforced in code rather than in the prompt, because a prompt is an instruction and code is a wall.
  5. Knowledge that builds up across runs. What the workflow learned last week should be available this week as saved, structured data, not worked out again from scratch on every run.

Is more scaffolding a sign of maturity?

No, and it can be the opposite. Building the six-component platform before a single workflow runs in production is a way to look mature while shipping nothing, and it is one of the reasons why AI pilots fail at mid-market companies: the harness question gets answered in the abstract, ahead of any evidence about which failures actually occur.

ServiceNow's Enterprise AI Maturity Index, a vendor index, has the average score falling from 44 to 35 out of 100 year over year, with under 1% of organizations above 50. Read it as direction rather than as a measurement of any one company: more companies are trying this, and more of them are finding it harder than they expected.

The component worth spending on early is evals, the test cases that score the output. Anthropic's guidance is to start with 50 to 100 cases per user flow (one job the agent does end to end), grade the final state rather than the path the agent took, and use simulated-user evals, where a second model plays the customer, to find new cases rather than to score the agent. Grading the end state, was the result right, is what lets you change the harness and know whether the change helped.

Does a multi-agent setup change the answer?

A multi-agent setup splits one job across several agents, each with its own loop. It is the same argument at a larger scale. Anthropic reports that its multi-agent research system beat a single agent by 90.2% on its internal eval while using about 15 times the tokens of a chat turn, and that a single agent already uses about 4 times the tokens of a chat turn. Those figures are Anthropic's on Anthropic's own system, relayed through secondary write-ups. Treat the 90.2% as an upper bound under favorable conditions. The 15x is the part that survives contact with a budget, and it is the number to plan against.

The rule holds at every scale: add the machinery when a logged failure asks for it, and not before. That sequencing is most of what AI transformation work with mid-market operators consists of in practice.

If you have already decided a workflow should be agentic and the open question is how much machinery to build around it, a 45-minute working session on your actual process list is the fastest way to get to a build order and a first set of evals. Get in touch to book one.

Next step · AI working session

Put this to work in your operation

Bring the process you want to automate and the data it runs on, and we come prepared. First conversation, no fee.

Book an AI working session

Sources

More insights