Generative AI

From Agent Harnesses to Software Factories

A harness runs the agent, a framework wires the workflow, and a software factory runs many of them against a backlog. We compare the layers, then the factory.

Leonardo Piñeyro

Leonardo Piñeyro

CTO

18 min read
From Agent Harnesses to Software Factories

As of February 2026, Stripe was merging more than 1,300 pull requests a week that contained no human-written code.1 Ramp says the agent it built in-house now writes three of every four pull requests it merges.2 Neither company treats the agent as a chat window. Work goes in as a Slack message or a ticket, and a pull request comes out for a person to review.

The industry calls this a software factory: many agents working a queue of tasks in parallel, with a review gate between them and production.3 It is already a product category: Factory and 8090 alone raised $285 million this year to sell one.45

Here is our bet: in 2027, most teams will build software this way. Not with every engineer steering a coding agent in a terminal, but with work flowing into a system that runs agents in parallel, verifies the results, and hands engineers pull requests to review.

A factory is built from the same layers as any agent, so we start with those layers and the tools we see most in production, then assemble them into a factory.

AI workflows vs. AI agents: why agents need a harness

An AI workflow is a process you design in advance. The steps, their order, and every branch live in your code, and the model fills in a decision or a piece of text at the points you chose: classify this ticket, extract these fields, draft this reply. An AI agent gets a goal instead of a script. At each step it reads the task, the result of its last action, and the tools it can reach, then picks the next action until the job is done.6

The tools and permissions you grant set the limits of that autonomy, and inside them the model chooses the path, sometimes more creatively than you would like. In July 2026, OpenAI disclosed that agents it was testing on ExploitGym, a cybersecurity benchmark, escaped their sandbox and broke into Hugging Face's production servers to steal data for their own evaluation.7 Hopefully your invoice-extraction workflow has never shown that kind of initiative.

That autonomy is why harnesses exist. A workflow needs little more than a function that calls the model. An agent needs somewhere to act: tools, a sandbox to run them in, rules about what it may touch, and a way to keep its context useful across hundreds of steps. The harness supplies all of it.

Harness engineering: why the wrapper matters

LangChain moved its coding agent from 52.8 to 66.5 percent on Terminal Bench 2.0 without changing the model underneath (gpt-5.2-codex). What changed was the harness, the software that runs the loop around the model.8 When frontier models are this close in raw capability, the harness often decides whether an agent finishes a long task or drifts off after 50 steps. That is why the field's shorthand in 2026 became agent = model + harness,9 and why tuning the wrapper got a name of its own from Dex Horthy in November 2025:

Harness engineering is also the right lens for the comparison below: how much of the wrapper you get for free, how much you can change, and how much you will build yourself.

Frameworks, harnesses, and platforms

Search for "agentic AI frameworks" and you get lists that put Claude Code, LangGraph, and CrewAI side by side, as if they did the same job. They don't. They sit at different layers, and mixing them up is how teams end up with an open-ended coding agent where they needed an auditable workflow, or with a hand-drawn graph where a capable agent would have finished the work in a day. Here is how the major labs and analysts define them, from the bottom up.

Framework. A code library of building blocks (model calls, tools, an agent loop, memory, and a way to wire steps into a workflow) that you assemble yourself. LangGraph, the OpenAI Agents SDK, Google ADK, Mastra, and eve are frameworks. Most run on a durable runtime, such as LangGraph's own or Temporal, that checkpoints each step and resumes a run after a crash or a two-day wait for approval.10

Harness. A finished agent: the loop plus a system prompt, tools, context management, permissions, and subagents, with the defaults already tuned. You tailor one instead of assembling it.11 Claude Code, Codex, OpenCode, Pi, and Deep Agents are harnesses.

Platform. Managed infrastructure where agents built with any framework or harness are deployed, run, and governed.12 Amazon Bedrock AgentCore, Microsoft Foundry Agent Service, Gemini Enterprise Agent Platform, and Claude Managed Agents are platforms.

Agent platform
Where agents run as a service. It hosts and governs agents built either way.
Running, for example, three agents on shared services
Support agentHarness
Code review agentHarness
Refund workflowFramework
Runtime & sandboxes
Identity & permissions
Memory
Tool gateway
Observability
Governance
Bedrock AgentCore · Foundry Agent Service · Gemini Enterprise Agent Platform · Claude Managed Agents
Agent harness
A finished agent you tailor. The model picks each next step.
Tools & MCP
Context & compaction
Agent loop
Permissions & hooks
Subagents & skills
Claude Code · Codex · OpenCode · Pi · Deep Agents
Agent framework
Building blocks you assemble. You own the architecture.
Model callsToolsAgent loopMemoryHandoffsGraphsCheckpoints
Assembled, for example, into a path you draw
Router
Agent A
Agent B
Human approval
Done
LangGraph · OpenAI Agents SDK · Google ADK · Mastra · eve

To pick a layer, use the whiteboard test from our recipe book for shipping agents to production: if a competent engineer can draw the decision tree, you have a workflow, not an agent. A path you can draw belongs in plain code, or in a framework once it needs state, checkpoints, approvals, and an audit trail. Work that looks more like a person at a computer (researching across ten sources, refactoring a codebase, reconciling a messy spreadsheet) belongs in a harness, and you drop down to a framework only when you need control over the loop itself.13 Add a platform when either one has to run as a service, with its own identity and permissions. The layers also compose: a harness can call a framework-built workflow as a tool.

Your business logic & evals
Domain rules · guardrails · test suites
Agent platform
Hosts and governs agents: deployment, identity, observability
Tools / Data / MCP / ...
Agent harness
A finished agent: loop, tools, context, permissions
Agent framework
Building blocks: model calls, tools, memory, durable workflows
Models
Anthropic · OpenAI · Google · open-weight

Models sit underneath all of it, and with clean abstractions, swapping providers is a configuration change. On top sits the only layer you own: your business logic and evals, meaning the domain rules, guardrails, and test suites that make an agent useful for your company specifically. Keep them behind your own interfaces, and moving to a different harness or framework becomes a refactor instead of a rewrite.

Agentic AI frameworks compared

Reach for a framework when you want to own the architecture: a path you draw that has outgrown plain code, or an agent whose loop you need to control. Star counts are from GitHub in September 2026, a rough proxy for community size and staying power.

FrameworkLanguageStarsBest fit
LangChainPython, TS~147kIntegrations, retrieval, and a prebuilt agent loop
CrewAIPython~59kRole-based multi-agent teams
LlamaIndexPython, TS~52kRetrieval-heavy agents over documents
LangGraphPython, TS~42kStateful workflows with checkpoints and approvals
OpenAI Agents SDKPython, TS~30kFast start with handoffs, guardrails, tracing
smolagentsPython~29kCode-writing agents on open-weight models
MastraTypeScript~28kAgents, workflows, memory, and evals in TypeScript
Vercel AI SDKTypeScript~27kStreaming chat UIs, provider-agnostic calls
Google ADKPython, Java, Go~22kGemini and Google Cloud teams
Pydantic AIPython~20kType-safe agents with validated outputs
Microsoft Agent FrameworkPython, .NET~14kAzure and Microsoft 365 environments
eveTypeScript~5.4kDurable agents with a default harness and chat channels
AG2Python~5kConversational multi-agent group chat

For regulated, audit-heavy workflows we start with LangGraph, because every node and edge is explicit and checkpointed, with human-in-the-loop interrupts built in. The cost is verbosity. The OpenAI Agents SDK gets a prototype running in an afternoon, but durability is yours to add. CrewAI and AG2 model work as a conversation between specialists, which demos well and gets awkward once your process stops matching the team metaphor. If you are still on AutoGen, Microsoft has moved it to maintenance mode in favor of Microsoft Agent Framework.14

The newest entry, Vercel's eve, ships more of the stack than most: an agent is a directory of Markdown instructions and TypeScript tools on a default harness, with checkpointed runs and deployment to Slack, Teams, web chat, or an API. It is TypeScript-only and still in beta.15

Agent harnesses compared

When the work is open-ended, start from a harness instead of assembling the loop yourself. If you are building a product, watch the last column: it is how your code runs the harness.

HarnessTypeStarsModelsEmbed via
Claude CodeCoding agent~148kClaudeClaude Agent SDK
CursorCoding agentn/aFrontier models, ComposerCLI, TS and Python SDKs, Cloud Agents API
CodexCoding agent~126kOpenAIPython, TS SDKs
OpenCodeCoding agent~210k75+ providersServer + JS SDK
Gemini CLICoding agent~107kGeminiCLI
Qwen CodeCoding agent~28kQwen, others, localSDKs
Hermes AgentGeneral agent~249kAny providerCLI, chat gateway
OpenHandsGeneral agent~89kAnyPython SDK, REST
GooseGeneral agent~55k15+ providersCLI, API
PiMinimal core~109k15+ providersSDK, RPC
fxMinimal core~3.1kAny, incl. locallibfx, WebAssembly
Deep AgentsLibrary~30kAny with tool callingPython, TS
Pydantic AI HarnessLibrary~0.9kAnyPython

Coding agents: Claude Code, Cursor, Codex, OpenCode

Millions of developers push these through long, messy sessions every day, so they are the most battle-tested harnesses available, and all of them now run from your own code. The Claude Agent SDK gives your application the same loop, tools, hooks, and subagents that power Claude Code, and for agents that read documents, run scripts, and produce files, it is the fastest path we know to something capable. The trade-offs are Claude models only, API-key billing under Anthropic's commercial terms, and a heavier runtime, because the SDK drives the bundled Claude Code binary.16

Cursor exposes its agent through TypeScript and Python SDKs, with any model Cursor supports. Codex, the pick for teams on OpenAI, has an open-source Rust core you can read and patch. OpenCode runs as a server your code controls through a typed SDK, a strong base when one backend serves several interfaces or you swap models per task. Gemini CLI now serves only API-key and enterprise accounts,17 and Qwen Code stands out for running local models.

General agents, minimal cores, and libraries

General-purpose agents such as Hermes, OpenHands, Goose, and OpenClaw (the most-starred of the lot, at about 391,000 GitHub stars) reach beyond the terminal, into chat apps and self-hosted fleets.18 They are assistants you run more than components you embed.

Pi is deliberately small: a minimal system prompt, no subagents, plan mode, built-in MCP, or permission gates, and TypeScript extensions for anything you add. You get full control over the context window, and you own whatever Pi leaves out, sandboxing included.

Deep Agents and Pydantic AI Harness are libraries you import into your own service. Deep Agents runs on LangGraph, so you inherit checkpointing, streaming, and LangSmith tracing, along with LangChain's abstractions. Pydantic AI Harness is early, but for a Python service already on Pydantic it is the cleanest option we have seen.

How to choose a framework, harness, or platform

Start with the layer, then the tool. This is the table we use to open the conversation with clients:

If your problem looks likeStart withWhy
A fixed process with a few LLM decisionsNo framework: plain code and structured outputsNothing extra to maintain
A regulated workflow with audit trails and approvalsFramework: LangGraphEvery path explicit and checkpointed
Open-ended work in your product, Claude models are fineHarness: Claude Code via the Agent SDKTools, hooks, and subagents built in
Open-ended work where you need model choice or local modelsHarness: OpenCode, Pi, or Deep AgentsNo provider lock-in
A long-running agent that lives in Slack, Teams, or web chatFramework: eveDefault harness, durable runs, and channels built in
An open-ended agent you would rather not hostPlatform: Claude Managed Agents or the AgentCore harnessHosted harness, sandboxes, and sessions
A voice agentPlatform: ElevenAgentsSpeech and tools without owning the loop
A conversational storefront or support assistantFramework: Vercel AI SDK or Mastra, LangGraph for complex flowsStreaming UI plus explicit business rules
A company standardized on one cloudPlatform: your cloud's agent platform, plus the framework your team prefersIdentity and governance native to your cloud

Once you have two or three candidates, build the same core workflow in each, with evals and tracing from the first commit, and check what will hurt in production:

  • Context control: Can you see and change everything that enters the window? (Our recipe book covers the context engineering behind this.)
  • Sandbox and permissions: Where do tool calls run, and what can the agent touch without asking?
  • Embedding: Can your code start, steer, and stop a session and stream its events?
  • Durability: Does a run survive a crash, a deploy, or a two-day wait for approval?
  • Observability: Can you trace every step and gate releases on evals?
  • Portability: Can you switch model providers, and can your team own the tool for the next two years?

A week of side-by-side prototypes costs far less than six months spent working around the wrong abstraction.

How we choose at Pento

We have shipped agents across hospitality, consumer hardware, facility management, and ecommerce, and in every project we picked the tools last. The problem shape, the client's existing stack, and the team that will own the system come first. Sometimes that means no framework at all.

For a facility management startup, we built pipelines that turn incoming calls, texts, and emails into work orders on nothing more than Next.js, Vercel, Twilio, OpenAI, and integration code. The path was well defined, so it stayed in plain code, and the thin stack let the team switch phone-call providers mid-project (see the case study). At the other end, ml-ralph, our autonomous ML experimentation agent, runs on the Claude Code CLI, because experimentation is open-ended by nature. Running agents like it every day is how we learned where each harness is strong and where it breaks.

When a client already has a cloud and a compliance regime, those pick the stack. The legal and HR copilots we built for a global hospitality services company run inside its Azure environment, with Azure AI Services, LangChain pipelines, and a Microsoft Teams integration, because that is where the security controls already lived. A consumer hardware company's onboarding assistant needed traceable data flows and guardrails, so LangGraph structured the flows and LangSmith monitored them end to end. In conversational commerce, voice runs on ElevenAgents, and the orchestration varies with how much the conversation branches.

From agent platforms to software factories

Everything so far runs one agent at a time. Once agents serve customers, act with their own credentials, or multiply across teams, someone has to host them, give each one an identity and permissions, trace what they do, and govern the fleet. That is the platform layer. Every major cloud sells one, and more platforms now ship with a harness inside (Claude Managed Agents, AgentCore's managed harness, and OpenAI's Agents API, which runs Codex), so the platform question is often whether you run your harness yourself or rent it.1920

Point that same stack at your backlog instead of your users, and you have a software factory.

What a software factory is

In a factory, work enters as a ticket, a spec, or an alert. Agents plan, implement, test, and review it in parallel, rejected work loops back, and what passes reaches a person as a pull request. The harness does the work, a framework wires it into a workflow with retries and checkpoints, and the platform supplies the sandboxes, identity, and audit trail.

What changesCoding agentSoftware factory
Unit of workA sessionA backlog
RunsOn your machine, one task at a timeIn the cloud, many tasks in parallel
The engineer's jobSteer the agent, then read the diffWrite the spec, tune the loop, own the gate
The bottleneckThe engineer's attentionHow fast a change can be verified

Factories sit at the top of the AI development stack:

L4
Delivery and services
Pento AI, and all other consulting/services companies
L3
AI Software factory
Spec-governed: Steer, 8090.ai, SoftwareForge, Maleus
Autonomous loops: Factory.ai, Devin, Warp
L2
Agent workflow tools
Spec toolchains
GitHub Spec Kit, Kiro, Augment Intent
Issue boards
Multica, Vibe Kanban
Parallel runners
Conductor, Emdash, Sculptor
Code review
Copilot code review, Cursor Bugbot, Greptile, Code Rabbit
L1
Coding agents
Editor and CLI
Claude Code, Codex, Cursor, Copilot
Cloud and background
Cursor Cloud Agents, Copilot cloud agent, Agent HQ
App builders
v0, replit, lovable, base44
L0
Models
Anthropic, OpenAI, Google, xAI, Alibaba, Z.ai, DeepSeek, ...

Coding agents (L1) make one engineer faster, workflow tools (L2) help one developer run several agents at once, and factories (L3) carry the whole organization's work from requirement to production.

Who is building factories

The first ones were internal: Stripe's Minions and Ramp's Inspect run in isolated cloud environments and hand back pull requests for review. The products that followed split on the source of truth. Autonomous loops such as Factory and Devin aim to turn a ticket into a pull request with less human time in the loop. Spec-governed factories such as 8090 keep the spec as the record, and every change has to trace back to it. Steer, which we built at Pento and run our own delivery on, is in the second camp: agents execute implementation specs in parallel on whichever harness fits the task, and every pull request carries evidence (tests, screenshots, traces) tied back to the spec.

Notice what these vendors sell: less the agent than the queue, sandboxes, specs, evals, and audit trail around it. That is the platform layer sold as a product, following the path CI/CD took from homegrown scripts to standard infrastructure, and it is why we expect factories to become the default. A coding agent in every terminal makes the pile of diffs taller. A factory puts specs, tests, and automated review in front of the person who merges.

Lights on or lights off

The live argument is about that last step. In a lights-on factory, agents write most of the code and a person still reads the diff before merge. In a lights-off factory, named after plants that run without people on the floor, code ships that no human has read.3

The public lights-off case is StrongDM, where three people set two rules: code must not be written by humans, and code must not be reviewed by humans.21 The public failure belongs to Dex Horthy, who named harness engineering. His team ran a fully automated factory for months, then abandoned it after the codebase degraded,22 a problem now called comprehension debt: code that grows faster than anyone's understanding of it.3

Our read: 2027 belongs to the lit factory. You can only hand a loop as much autonomy as you can cheaply verify,3 and tests catch what they can express, not a design that gets a little worse with every green build. So the gate stays, and it moves upstream, from reading every diff to approving specs, architecture, and the changes that carry real risk.

What this means for your roadmap

The tools in this guide will look different a year from now, as harnesses absorb orchestration, frameworks ship harnesses, and platforms host them. What carries over is the layer you own: your business logic, the evals that define what "correct" means, and the traces that show what each agent did. A factory runs on exactly those.

You don't need a factory on Monday. Start with the layer your constraint points at:

If this is the constraintAdopt first
One person writing code in an editorL1: a coding agent, plus a cloud agent once a task can run unattended
Several agents, and nobody can say which branch is whoseL2: a workflow tool matched to what you already track (specs, issues, or sessions)
Pull requests are the bottleneckAI code review on the agent you already run
Tickets should land as pull requests, with a record of what was decidedL3: a lights-on factory, where a person still merges
The team has not shipped agents and wants the finished systemL4: delivery on your stack

Whatever the row, three moves get you ready for 2027:

  • Get specs, tests, and evals in order. They are the factory's inputs and its gate. If they are weak, a faster generator only grows the review pile.
  • Decide where the gate sits. Write down, per repository, which changes a person must read (authentication, payments, migrations) and which a machine can verify.
  • Open one narrow queue. Bug fixes, dependency upgrades, and flaky tests in a single repository, where a passing test suite is a fair definition of done.

For the production discipline behind all of this, see our recipe book for shipping agents to production. And if you want a second opinion on your stack, or want to see what a factory would do to your own backlog, schedule a call with our agentic AI web development team. We build on your stack and hand over to your engineers.

References

Footnotes

  1. Stripe. Minions: Stripe's one-shot, end-to-end coding agents, Part 2, Stripe Dot Dev Blog, February 2026. See also Part 1. ↩

  2. Linear. The coding agent behind 75% of Ramp's merged PRs, customer story. Ramp describes the build in Why We Built Our Own Background Agent. ↩

  3. Addy Osmani. Software Factories, Light and Dark, July 2026. ↩ ↩2 ↩3 ↩4

  4. Factory. Factory raises $150M Series C, April 16, 2026. ↩

    1. 8090 Raises $135M Series A to Accelerate Their Rollout of Software Factory, Business Wire, June 2026.
    ↩
  5. Anthropic (Erik Schluntz, Barry Zhang). Building effective agents, December 2024. ↩

  6. OpenAI. OpenAI and Hugging Face partner to address security incident during model evaluation, July 2026. Hugging Face published its own technical timeline of the intrusion. ↩

  7. LangChain (Vivek Trivedy). Improving Deep Agents with harness engineering, February 2026. ↩

  8. LangChain (Vivek Trivedy). The Anatomy of an Agent Harness, March 2026. ↩

  9. LangChain (Harrison Chase). Agent Frameworks, Runtimes, and Harnesses- oh my!, October 2025. ↩

  10. Anthropic (Lance Martin). Agent Harness Design: 3 Patterns for Harnessing Claude's Intelligence, April 2026. ↩

  11. AWS Prescriptive Guidance. Agentic AI frameworks, platforms, protocols, and tools on AWS: Platforms ↩

  12. Gartner. Choosing Between an Agent Harness or Framework to Build AI Agents, July 2026. ↩

  13. Microsoft. AutoGen on GitHub (maintenance mode notice) ↩

  14. Vercel. eve: the framework for building agents ↩

  15. Anthropic. Claude Agent SDK overview ↩

  16. Google for Developers. An important update: Transitioning Gemini CLI to Antigravity CLI, May 2026. ↩

  17. OpenClaw. OpenClaw on GitHub ↩

  18. Anthropic. Claude Managed Agents overview ↩

  19. OpenAI. Introducing the Agents API, September 2026. ↩

  20. Simon Willison. How StrongDM's AI team build serious software without even looking at the code, February 2026. ↩

  21. Dex Horthy (HumanLayer). Why Software Factories Fail: Turning the lights back on, X article, 2026. The series is summarized in PostHog's Can software factories actually work? and in Osmani's essay above. ↩

CONTACT US

Schedule an
AI Strategy Session

Work with Pento to turn promising AI experiments into systems that perform reliably in production, with the right architecture, delivery model, and engineering support.