This study is Part 2 of the ace8 Holon Levels series. Part 1, The Architecture of Disintermediation (September 2026), classified the routed model market. This part classifies the agent harness—the software layer that plans, edits, executes, and recovers on a developer’s behalf—together with the models that run inside it. Grades combine publicly available information with sustained hands-on use of each harness across a full product development cycle during 2026.
Executive Summary
Personal agents have come of age faster than the market expected. Claude Cowork redefined the productivity category by carrying the agentic approach of Claude Code into ordinary knowledge work, and adoption now reaches small businesses that historically trailed enterprise technology by years: Anthropic’s own customer examples come from energy and consumer goods, and this study’s fieldwork finds the same pattern in manufacturing and real estate. The buyer’s question has changed from whether a model can build a spreadsheet to whether an agent can run part of the business. This study reads that shift through the holonic framework that Arthur Koestler introduced in The Ghost in the Machine (1967) and that multi-agent research has since applied to LLM systems (arXiv:2501.07992). A holon is a unit that is both a complete whole and a part of a larger system; a harness is exactly that, a complete agent loop to its user and a replaceable component to the router and settlement layers above it.
Anthropic launched Claude for Small Business inside Cowork on May 13, 2026, addressing firms that account for 44% of US GDP and whose AI adoption had lagged larger enterprises (Anthropic). The enterprise AI coding agent market reaches an estimated $14.18 billion in 2026 and $45.83 billion by 2031, a 26.44% CAGR (Mordor Intelligence). Adoption already outruns oversight: 31% of developers used AI agents in 2025 (Stack Overflow), and 49% of companies using agentic AI have yet to extend governance to it (EY).
Success created a cost problem. In this study’s hands-on use, sustained agent work across premium subscriptions and usage credits approached $1,000 a month, a figure that stops a small business. OpenCode and OpenRouter’s Ori answered it by separating the harness, the model, and the payer, so the same workflow runs on whichever model clears the bar at the lowest price. Open harnesses now carry the volume: Hermes, Kilo Code, pi, and Cline together processed about 3 trillion OpenRouter tokens in a day, against about 1.2 trillion for Claude Code and Codex combined (OpenRouter Apps).
The two forces confirm Benedict Evans’s February 2026 observation that half a dozen labs ship frontier models of broadly equivalent capability (Benedict Evans), and they locate where differentiation remains. No model provider holds a durable lead on capability: open-weight models closed the gap to Anthropic’s Fable within weeks, and Qwen3.8 Max now ties Claude Fable 5.1 on the Artificial Analysis Intelligence Index (Part 1). None holds a durable lead on price or scale: DeepSeek led OpenRouter text requests at 25.4% in the week of September 14 and routes from $0.0356 per million input tokens. Providers differentiate on experience, form factor, and speed. OpenAI’s inclusion of Fast mode in its subscriptions (The Deliberate Egalitarianism of Fast Mode), Anthropic’s Cowork, and Meta’s Muse compete on exactly those three, and the grades in this study follow them.
A larger question now sits above all three: whether society trusts agents at all. 57% of US adults rate the societal risks of AI as high, against 25% who rate its benefits as high, and 52% are more concerned than excited, up from 37% in 2021 (Pew Research Center). Trust is becoming the fourth axis of differentiation, and the vendor that has already answered hard public questions about data has an advantage that capability and price cannot buy.
This taxonomy defines six holon levels (H0–H5) for agent harnesses, evaluates six harness vendors across five capability dimensions, and assigns letter grades to each. The models inside the harnesses appear as supporting holons. The study serves technology selection for engineering leaders, investment theses for capital allocators, and roadmap assessment for harness vendors.
Section 1: The Holonic Framework Applied to Agent Harnesses
1.1 Origins of the Holon Model
Koestler described every holon as Janus-faced. Its inward face governs its own internal logic; its outward face interfaces with the system that contains it. A holarchy is a tree of such units, each whole in itself and each a governed part of the level above. Recent systems research applies the model directly to LLM-enhanced architectures, treating each agent as a holon whose boundary determines what it may decide alone (arXiv:2501.07992; arXiv:2505.00368). This study assesses each harness through four graphs: the interior graph (its planning and tool loop), the boundary membrane (what it permits across its edge), the context graph (the models, tools, and data it can reach), and the projection graph (what it presents to the developer and to the systems above it).
1.2 The Harness Gap
The harness now moves results as much as the model does. An analysis of 254 SWE-bench submissions found that the same model varied by up to 29.8 percentage points across nine scaffolds, with a median spread of 15.6 points, while the top 30 submissions spanned only 8.8 points (arXiv:2609.17394). Harness variance can exceed model variance and even reverse model rankings (arXiv:2605.23950), and a recursive harness lifted a Codex baseline from 71.75% to 81.36% on Oolong-Synthetic (arXiv:2606.13643). For vendors, the implication is direct: a harness is a product in its own right, and grading it separately from the model inside it is necessary to read the market correctly.
1.3 Holon Level Definitions
The levels continue the scale defined in Part 1. Each level inherits all capabilities below it.
Section 2: Holon Level Technical Specifications
H0 — Model Snapshot
The fundamental indivisible holon. A snapshot is a fixed set of weights that completes a prompt and retains nothing between calls. Open-weight snapshots now reach harness-grade quality on a single workstation: Alibaba released Qwen3.8-27B, a dense multimodal model, under Apache 2.0 on August 14, 2026 (The Decoder), and DeepSeek V4 Flash activates 13 billion of its 284 billion parameters per token (OpenRouter).
Key Characteristics:
Stateless inference
Local or hosted execution
License governs the membrane
27B dense models runnable on one Apple silicon workstation
H1 — Hosted Endpoint
The first functional holon that sells access. A hosted endpoint wraps a snapshot with tools, context limits, rate limits, and a price. OpenAI launched GPT-6 Sol at $2/$10 per million tokens and GPT-6 Luna at $0.10/$0.50 on September 22, 2026 (TechCrunch); Anthropic released Claude Opus 5.5 at $4/$20 the same day (Anthropic); SpaceXAI prices Grok 4.6 at $2/$6 (Latent.Space).
Key Characteristics:
Per-token pricing with speed tiers
Provider-defined rate limits
Tool and context contracts
Frontier input prices from $0.10 to $10 per million tokens
H2 — Routed Market
The first holon level at which the model becomes interchangeable. A router selects a provider per call by price, latency, and fit. OpenRouter serves 400+ models from 80+ providers (Stripe) and routes DeepSeek V4 Flash from $0.0356 per million input tokens (OpenRouter).
Key Characteristics:
Provider fallback and price routing
One key across many models
OpenAI-compatible request format
10 trillion+ tokens routed per day
H3 — Embeddable Agent Harness
The first compositional holon level and the dominant production pattern. A harness runs a plan-edit-execute-recover loop over a workspace, calling models from H1 or H2. Claude Code, OpenAI Codex, OpenCode, Muse Code, Cline, and Kilo Code operate here.
Key Characteristics:
File, terminal, and browser tools
Subagents and worktree isolation
Session memory and resumption
Scaffold choice moves SWE-bench results by a median 15.6 points (arXiv:2609.17394)
H4 — Meta-Harness
The first federated holon level. A meta-harness launches and supervises several harnesses behind one login, one policy, and one model catalog. OpenRouter’s Ori Harness launched on August 4, 2026 with four harnesses and now supports twelve (OpenRouter Blog; Ori Harness docs).
Key Characteristics:
Harness-agnostic model routing
Single sign-on across harnesses
Shared budgets and guardrails
Twelve supported harnesses as of September 2026
H5 — Settlement and Governance Fabric
The governing holon level. H5 adds no new agent capability; it governs the membrane of every level below through payment, budgets, approvals, and audit. Stripe’s agreement to acquire OpenRouter on August 19, 2026 places settlement directly above the routed market (Stripe). Meta’s Muse applies the same pattern to personal agents, with a separate Sentinel agent approving every internet-bound action (Meta).
Key Characteristics:
Transaction settlement per call
Organization-wide budgets and guardrails
Independent approval of outbound actions
A reported $7–8 billion price for the H2–H5 position
Section 3: Vendor Taxonomy and Grading Framework
3.1 Grading Methodology
Each vendor receives a letter grade from A to D, with plus and minus modifiers, on five equally weighted dimensions. The overall grade combines the five and reflects sustained hands-on use where the dimensions alone understate a result. Part 1 replaced Developer Experience with Price Efficiency because it graded markets; this part restores Developer Experience because it grades the layer a developer works inside
Section 4: Vendor Profiles and Grades
4.1 Personal and Platform Agents (H3–H5 Capable)
Max Holon Level: H5 (Personal Estate) | 1.8 million US and Canada iOS downloads in its first 12 days
Meta launched Muse as a personal agent on September 8, 2026, running on Muse Spark inside a Muse Secure VM that holds both the agent and the user’s data; a separate Sentinel agent approves every internet-bound action (Meta). Users can opt out of training use, Muse data stays out of Meta’s advertising systems, and an end-to-end encrypted Muse Confidential VM is announced. Muse Code, the terminal coding agent, entered beta on August 5, 2026 on Muse Spark 1.2, fans work out to subagents in separate worktrees, and charges $1.25/$4.25 per million tokens (TechCrunch). The Secure VM is the interior graph, Sentinel is the boundary membrane, and the pairing reaches H5 for a single person’s estate. Muse outpaced ChatGPT’s early mobile launch, with 642,000 US mobile daily users against 231,000 over the same window (TechCrunch). Its differentiator is purpose-built simplicity: the work starts immediately, with model hosting, build tooling, and authentication handled inside the membrane. Its key limitation is enterprise reach: Muse launched as a US-only consumer product, priced in individual tiers, and requires a payment card even on the free tier (Tech Insider).
4.2 Meta-Harness and Settlement (H4–H5)
Max Holon Level: H5 | 12 supported harnesses; acquisition by Stripe announced August 19, 2026
Ori shipped as CLI version 0.1.0 on June 27, 2026; version 0.7.1 added Prime Agent, DeepSeek Harness, and Grok Build on August 15, version 0.8.0 made Ori’s own agent loop the default on August 21, and version 0.10.1 followed on August 24 (Ori Changelog). Ori Harness launches Claude Code, Cline, Codex, Grok Build, Hermes, Kilo Code, Muse Code, omp, OpenCode, Pi, Prime Agent, and DeepSeek Harness behind one OAuth login, with organization guardrails and budgets applied to every request (Ori Harness docs). Ori Eval tests a team’s own prompts across providers and identifies the model that meets the required accuracy at the lowest cost. Stripe’s agreement to acquire OpenRouter, at a reported $7–8 billion ($7.5 billion per The New York Times), attaches settlement to the router (Stripe; Slashdot). Ori’s differentiator is the separation of interface, model, and payer into independent variables. Its key limitation is maturity: ten minor releases in two months and an ownership transition leave its interfaces still moving.
4.3 Lab Harnesses (H3)
Max Holon Level: H3 | 20 million active users reported August 21, 2026
Codex is included in every ChatGPT plan: Free, Go at $8, Plus at $20, Pro from $100 at five or twenty times Plus limits, and Business at $20 to $25 per seat, which brings seat administration and enterprise billing to the harness; Fast mode consumes 2.5 times the standard credit rate (ChatGPT pricing). Including Fast mode in the subscription is the pricing decision that separates Codex from its lab peer. This grade reflects sustained use of GPT-5.6 Sol in Fast mode, inside Codex and inside OpenCode. GPT-6 Sol and GPT-6 Luna reached Codex on September 22, 2026 at half the GPT-5.6 rate (TechCrunch); they are one day old at this writing and appear as supporting holons. An Ultrafast preview on Cerebras runs GPT-5.6 Sol at up to 750 tokens per second, currently limited to an API preview (OpenAI). OpenAI’s Codex lead reported 20 million active users on August 21 without defining the activity window (The New Stack). Its differentiator is the model: GPT-5.6 Sol carries more of the result than the harness around it, and it performs at least as well inside OpenCode. Its key limitation is capacity: new $200 Pro sign-ups and upgrades have been paused since September 10, 2026 (OpenAI Help).
Max Holon Level: H3 | 954 billion tokens per day through OpenRouter alone
Claude Code entered research preview on February 24, 2025 alongside Claude 3.7 Sonnet and set the pattern for the H3 harness (Anthropic). Claude Opus 5.5 arrived on September 22, 2026 at $4/$20, 20% below Opus 5, with higher five-hour limits across Pro, Max, Team, and Enterprise plans, the last two carrying the seat administration and billing enterprise buyers require (Anthropic). Max plans run at $100 and $200 (Claude Help Center). Fast mode costs $8/$40 on Opus 5.5 and bills to usage credits only, even when plan usage remains (Claude Code Docs). Anthropic closed subscription access to third-party harnesses, with formal enforcement from April 4, 2026 (Decode the Future). Its differentiator is breadth: subagents, hooks, and skills inside the harness, multimodal input, and the same agentic approach carried into Claude Cowork for non-coding work (Claude Help Center), with Anthropic’s Compliance API covering both products since August 2026 (Claude blog). Its key limitation is weight: in sustained hands-on use it added the most overhead per task of the six harnesses graded, and its closed subscription membrane narrows the models and harnesses a subscriber can combine. Opus 5.5 is one day old at this writing and appears as an ungraded supporting holon.
4.4 Open Harnesses (H3)
Max Holon Level: H3 | 209,650 GitHub stars, MIT license
OpenCode, now maintained at anomalyco/opencode, is an open-source terminal harness that reaches 75+ providers and local models (Morph). ChatGPT Plus and Pro subscribers sign in natively through /connect from version 1.1.11, which lets OpenAI’s subscription models run inside a neutral harness (GitHub). OpenCode is one of Ori’s twelve supported harnesses. Its differentiator is openness: any model, any provider, one interface. Its key limitation is governance: its controls operate at the individual developer, and organization-wide budgets and guardrails come from a meta-harness above it.
4.5 Teammate Agents (H3, Domain-Constrained)
Max Holon Level: H3 (Domain-Constrained) | In beta since August 11, 2026
Grok Bot provides persistent AI teammates that sign into a user’s tools, learn a task by demonstration in up to ten minutes, and share one cloud computer per account (Big Hat Group). It has no standalone price; it ships inside Cursor Ultra at $200, Cursor Teams Premium at $120 per seat, and SuperGrok Heavy, following SpaceX’s $60 billion all-stock acquisition of Cursor, which closed on August 14, 2026 (TechCrunch). Its model, Grok 4.6, launched August 12 at $2/$6, scored 61 on the Intelligence Index, level with GPT-5.6 Sol, and 88.4% on Terminal-Bench v2.1 (Latent.Space). Its differentiator is learning by demonstration. Its key limitation is the membrane: a beta agent that holds sign-ins to a user’s tools on a shared cloud computer carries a governance burden the product has yet to document.
4.6 Supporting Holons (H0–H1)
The models below run inside the graded harnesses and set the price floor beneath them. They carry no grade; models released September 22 had one day of availability at this writing.
Section 5: Consolidated Vendor Scorecard
The two lab harnesses lead on enterprise readiness, Ori and OpenCode lead on openness, and Meta Muse is the fastest riser in the field.
HD = Holon Depth · PR = Production Readiness · GT = Governance & Trust · EO = Ecosystem Openness · DX = Developer Experience. Production Readiness includes enterprise billing, seat administration, and procurement terms. Claude Code and Codex sit above their dimension averages because enterprise buyers weight those terms heavily; Claude Code leads on multimodal input and on its integration with Claude Cowork under one compliance layer. Muse’s grade marks its position two weeks after launch; its Developer Experience and Governance & Trust grades show where it is heading.
Section 6: Structural Market Observations
6.1 The OpenAI-Compatible Request as Boundary Membrane (H1–H3)
The OpenAI-compatible request format has become the standard membrane between harness and model. OpenCode reaches 75+ providers through it, Ori routes twelve harnesses across 400+ models through it, and local runtimes expose it on the developer’s own machine. A harness that speaks the format can swap its model per call, which moves competition from the harness-model pairing to each layer on its own merits.
6.2 Subscription Terms Are the New Ceiling at H3
The sharpest differences between the two lab harnesses come from subscription terms, and the models themselves are closely matched. OpenAI includes Fast mode in ChatGPT plans at a 2.5x credit rate; Anthropic bills Fast mode to usage credits even when plan usage remains and closed its subscription to third-party harnesses. The ceiling on a developer’s day is therefore a design choice by each lab, separate from any capability limit, and developers route around it.
6.3 The Governance Gap Moves Up to H5
49% of companies using agentic AI have yet to extend governance to it, and 39% leave accountability for deployed agents undefined (EY). The vendors closing that gap operate at H5: Ori applies organization budgets and guardrails to every request, and Muse places an independent Sentinel agent on every outbound action. Governance is becoming a layer the buyer adopts once, above every harness, and harness-level controls are becoming secondary.
6.4 Commoditization at H0–H2, Moats at H4–H5
The price floor at H0–H2 keeps falling. GPT-6 Luna lists at $0.10 per million input tokens, DeepSeek V4 Flash routes from $0.0356, and Qwen3.8-27B runs locally under Apache 2.0. Harness volume follows: open harnesses Hermes, Kilo Code, pi, and Cline together processed more OpenRouter tokens in a day than Claude Code and Codex combined (OpenRouter Apps). Durable position is accruing at H4–H5, where Stripe paid a reported $7–8 billion for the router and settlement position and Meta built trust controls into the product itself.
6.5 Concentration Follows Distribution
The strongest new entrants arrived with distribution already in hand. Muse launched to Meta’s user base and passed ChatGPT’s early mobile numbers within 12 days; Grok Bot ships inside Cursor after SpaceX’s $60 billion acquisition; Ori arrives inside Stripe. Harness quality decides which product a developer prefers, and distribution decides which product the developer meets first.
6.6 Societal Risk Makes Trust the Consumer Moat (H5)
Public concern about AI is rising faster than adoption: 57% of US adults rate its societal risks as high, and about two-thirds say it is moving too fast (Pew Research Center). Meta enters personal agents after years of congressional scrutiny, including the January 31, 2024 Senate Judiciary hearing on child safety (Washington Post). Muse reflects that history in its design: a Secure VM that holds the agent and the user’s data, a separate Sentinel agent approving every outbound action, a training opt-out, and separation from Meta’s advertising systems. Meta built its trust controls into the product ahead of the questions a hearing would ask, and that placement at H5 is the clearest reason Muse is rising in this study’s grades.
Section 7: Forward-Looking Considerations
7.1 The Rate Limiter Is Inference Speed Under Subscription
OpenAI’s Ultrafast preview runs GPT-5.6 Sol at up to 750 tokens per second on Cerebras, roughly fourteen times standard speed, and remains an unpriced API preview (Cerebras). Once a speed tier of that class enters a subscription, the harness loop becomes limited by review and approval rather than generation, and H3 differentiation shifts to planning quality and recovery.
7.2 Local and Cloud Become One Session
Apache 2.0 models of 27 billion parameters now run on a single Apple silicon workstation, and the OpenAI-compatible format lets a harness treat a local model and a hosted one as the same endpoint. The next layer to form sits between H1 and H2: a session that routes each step to local or cloud models by sensitivity, cost, and speed, with the locality of every step stated to the user. Cisco’s September 15 release of air-gapped Splunk AI shows enterprise demand for the same pattern at data-center scale.
7.3 Settlement Signals Maturity
Pricing is moving from seats toward settled consumption. Muse tiers weekly token allowances at $0, $20, and $100 (Tech Insider); Codex meters Fast mode in credits; Ori settles every routed call through what will become Stripe. When a harness vendor prices per settled outcome, it has reached H5, and that shift is the clearest signal of level maturity in the next twelve months.
Methodology Note
This study reflects information available as of September 23, 2026. Grades combine vendor documentation, press coverage, and live usage boards with sustained hands-on use of each graded harness during 2026. Token volumes come from OpenRouter’s live application board and cover traffic routed through OpenRouter only; each vendor’s direct traffic is excluded. Market sizing comes from Mordor Intelligence, adoption from the Stack Overflow 2025 Developer Survey (2026 results were unpublished at this writing), and governance data from EY. Where sources disagree, the study reports the range: the Stripe–OpenRouter price appears as $7 billion+ (Bloomberg), $7.5 billion (The New York Times), and $8 billion+ (Semafor).
References
Introducing Muse - Meta’s launch of its personal agent, Secure VM, and Sentinel
Meta launches Muse Code - Muse Code beta, subagents, and pricing
Meta’s Muse is outpacing ChatGPT’s early mobile launch - First 12 days of downloads and daily users
Meta Muse personal AI agent launch - Muse tiers and availability
Stripe agrees to acquire OpenRouter - Acquisition announcement and OpenRouter scale
Stripe buys AI startup OpenRouter for $7.5 billion - Reported deal value
Stripe nears deal to buy OpenRouter for over $7 billion - Earlier reported deal value
Stripe agrees to acquire OpenRouter - Reported value above $8 billion
Ori Harness - Supported harnesses, login, guardrails, and budgets
Ori Changelog - Ori release history
Ori Harness announcement - Launch with four harnesses
OpenRouter Apps - Live daily token volume by application
DeepSeek V4 Flash on OpenRouter - Parameters and routed pricing
GLM-5.3 on OpenRouter - Routed pricing
OpenCode - Repository, stars, and license
Open Source AI Coding Assistants (2026) - OpenCode provider support
OpenCode ChatGPT sign-in - Native ChatGPT Plus and Pro sign-in
ChatGPT pricing - Codex plan inclusion and Fast mode credit rate
OpenAI launches GPT-6 Sol and Luna - Launch and pricing
Previewing Ultrafast - Ultrafast preview on Cerebras
Cerebras powers Ultrafast mode - Throughput figures
Codex user surge - Reported Codex active users
About ChatGPT Pro tiers - Pro tiers and sign-up pause
Claude 3.7 Sonnet and Claude Code - Claude Code research preview
Claude Opus 5.5 - Release, pricing, and limits
Claude Code fast mode - Fast mode pricing and billing
What is the Max plan - Max plan tiers
Anthropic blocks third-party tools - Subscription access enforcement
xAI weekly, August 14, 2026 - Grok Bot beta and bundling
SpaceXAI Grok 4.6 - Grok 4.6 pricing and benchmarks
SpaceX to acquire Cursor - Cursor acquisition
Gemini 3.8 Flash launch pricing - Gemini 3.8 Flash pricing
Qwen3.8 open weights under Apache 2.0 - Qwen3.8-27B release and license
Cisco delivers trusted AI at scale through new Splunk advancements - Air-gapped Splunk AI
Enterprise AI coding agent market - Market size and CAGR
Stack Overflow Developer Survey 2025: AI - Developer agent adoption
EY autonomous AI governance survey - Governance gap
arXiv:2609.17394 - Coding agents and scaffold variance on SWE-bench
arXiv:2605.23950 - Harness disclosure in agent comparisons
arXiv:2606.13643 - Recursive agent harnesses
arXiv:2501.07992 - LLM-enhanced holonic architecture
arXiv:2505.00368 - LLM-enhanced holonic approach to systems of systems
Koestler, A. The Ghost in the Machine (1967) - Origin of the holon model
The Architecture of Disintermediation - Part 1 of this series
Get started with Claude Cowork - Cowork’s shared agentic approach with Claude Code
Compliance API coverage extends to Claude Cowork and Claude Code - Enterprise compliance coverage, August 2026
Introducing Claude for Small Business - Launch inside Cowork, May 13, 2026
How will OpenAI compete? - Benedict Evans on frontier model differentiation, February 19, 2026
Key findings about how Americans view artificial intelligence - Public views of AI risk, March 2026
Meta’s Zuckerberg apologizes to child abuse victims in emotional Senate hearing - Senate Judiciary hearing, January 31, 2024





