The model is only part of the agent

Developers can use Claude Code, Codex, or Cursor while focusing most of their attention on the LLM behind the experience. Yet all three are closed-source agentic harnesses built around LLM models, and the surrounding software is responsible for much of what makes their behavior agentic. Its role can be expressed compactly as “Agent = Model + Harness.” The harness may remain largely invisible while determining how model output turns into action.

The equation separates two jobs that can otherwise look like one. A standalone model accepts inputs such as text, images, audio, or video and produces text in response. An agent adds a surrounding system that carries those responses into continuing, controlled activity. Engineering a useful agent therefore requires developers to understand both the model and the system connecting its output to state, tools, execution, and further model calls.

A harness turns model responses into ongoing work

That surrounding system is the agentic harness, also called an agent harness or agentic AI harness. It connects an LLM with external tools, data sources, memory, execution environments, and feedback loops. Those connections determine what can happen after a response: the system can retain information, invoke another resource, execute work, and return results for another model interaction. Through this cycle, an agent can keep working toward a task after its first generated response.

Memory shows how the harness extends a model’s basic behavior. An agent may need information to persist between interactions, and memory files and MCPs are among the mechanisms available for that purpose. MCP, or Model Context Protocol, provides a way for agent systems to connect models with external resources and tools. Persistence lets later work use information held outside a single model response, giving an agent continuing state across its work.

That continuing state becomes useful when the agent can also act. For code execution, the harness can provide a sandbox, an isolated environment where code runs. For work requiring several rounds, loops can return execution results to the system and trigger further steps. The loop lets an agent keep working until it solves the given task, with each result becoming input to the next stage.

Those loops may also need information missing from the current interaction. Web search gives the harness one way to retrieve external data and make it available for later work. Together, memory, sandboxes, loops, and web search separate model generation from agent operation: the model produces responses, while the harness creates the conditions for persistent work that can retrieve information and use tools.

Once those conditions exist, the engineering problem expands with them. An LLM that can retain state, retrieve outside information, execute code, and operate through several rounds needs rules governing which resources it can reach and when it can use them. The harness is where developers can implement those operational decisions. Its internal architecture shows the range of decisions involved.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

Inside the harness: the architecture that coordinates an agent

The first architectural concern is where people can see and control agent activity. An interface can expose what the agent is doing while letting a person approve or interrupt actions. Human control becomes important when a generated response can trigger further operations with effects outside the conversation. The interface gives human judgment an explicit place in that continuing process.

Behind the interface, the prompt and policy layer defines the instructions governing the agent’s work. It can combine system instructions with business rules and organizational policies, giving the harness a place to encode behavioral requirements. Those requirements carry into later tool use and execution, where generated decisions can have effects beyond language. The policy layer establishes the conditions under which the rest of the system operates.

Those conditions depend on the information available to the model, so the context manager determines what the model receives and when. It can use compression and other techniques to curate accumulated material for a session, selecting information useful for the current step. Because the model works from the context presented to it, that selection influences its next response. Context management is therefore an active part of coordinating the agent’s work.

With instructions and selected context available, the model interface mediates access to approved LLM services. It sends prompts, context, and parameters to those services and receives their responses. The interface creates a defined boundary between the agent system and the models it may call. Once a response crosses that boundary, the harness can determine which resources are available for the next operation.

The tool registry defines one set of those resources by holding the approved tools and functions the agent can use. Tool access allows a generated decision to lead to an external operation, so developers need to design the available set explicitly. The registry provides that defined set inside the harness. Availability then raises a separate question: whether the agent may perform a particular action.

That question belongs to the permission system, especially when an available operation is sensitive. Permissions specify what the model may do and can impose tighter limits on sensitive actions. Separating available tools from allowed actions lets developers govern behavior where generated output can create an external effect. Human approval through the interface can add another control when an operation warrants intervention.

Once an action is allowed, the execution environment determines where it runs. The agent runtime enters the architecture here by providing an environment in which agent behavior can occur. Coordination and execution are separate engineering jobs, even though they work together during an agent task. Keeping that boundary visible also helps distinguish a harness from a runtime.

Execution may in turn require information or functions outside its immediate environment. The connector layer links the harness to external sources, including retrieval-augmented generation (RAG) repositories and tools enabled through MCP. RAG retrieves relevant information and supplies it for model work, while MCP-enabled tools provide another path to external resources. Connectors extend the systems an agent can reach while keeping that access within the harness’s coordination.

External work also creates information that may need to survive the current interaction. The memory and session store preserves memory between sessions, allowing future work to use retained state. Stored information and model context have different roles: the store can keep material across sessions, while the context manager selects what the model should receive at a particular moment. That separation lets the system retain information without automatically putting every stored item into every model call.

Persistent state and external actions create a further need for accountability. Audit and observability systems maintain records for monitoring, review, and incident response. As an agent gains access to tools, outside systems, and execution environments, engineers may need to inspect what it did after an operation. Observability makes that activity reviewable across a sequence of model calls and actions.

Together, these components describe coordinated responsibilities rather than a mandatory execution sequence. The harness can supply prompts, selected context, and parameters through the model interface, then use returned responses in work involving approved tools, connectors, memory, and an execution environment. Permissions can constrain actions, people can approve or interrupt activity, and audit mechanisms can record events. Feedback loops can repeat parts of this process until the task completes, with the exact path determined by the task and system design.

Context management makes the harness an operational engineering concern

Within that architecture, context management has an unusually direct effect on each model call. Information accumulates as an agent works through successive interactions, and poorly curated material can make a session noisy, expensive, and error prone. Compression and other context-management techniques determine which accumulated information remains useful enough to present to the model. Each selection changes the information available when the next response is generated.

Because each response depends on that selection, context curation becomes a quality decision. Irrelevant or poorly managed material can interfere with useful evidence and instructions, while deliberate selection controls what the model can use for its next step. In an agent operating through loops, the decision recurs as new information enters the system. Context engineering consequently shapes the agent’s continuing judgment throughout the task.

The recurring decision also affects cost. As an unmanaged session grows, sending accumulated material forward can increase expense along with noise. Developers choosing what the model sees are therefore managing both the quality of its working information and the cost of supplying it. The harness gives them an architectural place to make that trade-off throughout an agent’s operation.

Capability needs guardrails: the harness also governs the agent

The operational choices around context sit beside a broader control problem created by agent capabilities. Tools provide callable functions, connectors reach external sources, and memory preserves information across interactions and sessions. Giving an agent those resources also requires decisions about which actions are allowed and when people should be able to inspect, approve, or stop them. The same surrounding layer that coordinates capabilities can enforce those boundaries.

The controls introduced in the architecture divide that responsibility across several mechanisms. The prompt and policy layer carries business rules and organizational policies, while permissions determine which actions the model may perform. Interface controls let people approve or interrupt activity, and audit records support monitoring, review, and incident response. Exposing a powerful tool therefore creates a related engineering task: specifying the controls that govern its use.

Adding another tool or external source increases the operations the system can attempt, which can also create new permission, oversight, and inspection requirements. Harness engineering represents those requirements as software surrounding the model. Safety and operational control become architectural decisions tied to the same mechanisms that enable agent behavior.

Framework, harness, and runtime are different jobs

Those responsibilities make the harness broad, but adjacent infrastructure still has distinct jobs. An agentic framework supports writing and building an agent, an agentic harness supports configuring and running it, and a runtime provides the environment where its behavior takes place. The named examples make the division easier to compare.

Infrastructure role Primary job Examples
Agentic framework Write and build an agent LangChain, OpenAI Agents SDK, LlamaIndex
Agentic harness Configure and run an agent Claude Agent SDK, Deep Agents
Agent runtime Provide the environment where agents perform their behavior LangGraph, Amazon Bedrock AgentCore

The runtime connects back to the harness architecture through the execution environment where actions occur. A harness can coordinate and govern execution while the runtime supplies the environment in which agent behavior takes place. Keeping the roles separate helps engineers locate decisions about control, orchestration, and execution in the appropriate layer.

The framework boundary completes that division of responsibilities. Frameworks support writing and building agents; runtimes host their behavior; harnesses coordinate model access, context, tools, permissions, connectors, memory, human controls, and observability while configuring and running agents. These systems can participate in the same stack, but their engineering roles remain distinct. Using the terms precisely makes it easier to identify where a particular operational decision belongs.

For developers, agent engineering extends beyond the model

Those distinctions expand the practical skills developers need when integrating agentic AI into applications. Pluralsight’s “Integrating AI for Developers” learning path is an on-demand video course covering common frameworks, orchestration strategies, and safety considerations. Its use cases include task automation, assistants, and workflow agents, each of which requires developers to reason about how model behavior fits into a larger application. Pluralsight has a commercial interest in this framing because it sells training for developers who want to build these skills.

That training example connects the architectural distinction to day-to-day development work. Once software can preserve memory, reach external systems, execute actions, and continue through loops, developers have to decide how to orchestrate those abilities and where their safety boundaries sit. Those are application-architecture decisions implemented through the surrounding agent system, and they require skills that extend beyond choosing or prompting an LLM.

Final thoughts

For business leaders, the key distinction is that choosing an LLM is not the same as choosing an agent architecture. The model provides core intelligence, but the harness determines how that intelligence reaches company data, tools, workflows, and execution environments. It also provides the controls that shape what an agent can do, what requires approval, and how its actions can be reviewed.

That makes the harness an important consideration when evaluating agentic AI platforms and investments. Leaders should look beyond model performance to the surrounding architecture for context management, permissions, memory, observability, and human oversight. These capabilities influence operating cost, risk, and how effectively an agent can perform useful work within existing systems.

As organizations move from AI assistants toward agents that can take action, those surrounding systems become more consequential. The strongest agent strategy will account for both sides of the equation: what the model can generate and how the harness turns those outputs into controlled, observable business operations.

Alexander Procter

September 30, 2026

11 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.