The following industry analysis is based on my in-depth daily use of various AI platforms, agentic tools, and other interface layers across a product management lifecycle including product development, minimum viable product assessments, ideal customer profile research, pricing strategy, and general disruptive trends.
The artificial intelligence industry is currently navigating a profound structural reset driven by Stripe and by Open Model companies, one that demands a sober assessment of where value truly accumulates. For the last several years, the market has operated under the assumption that frontier capability required an inescapable, capital-intensive bundle: massive stockpiles of Western hardware, proprietary closed-source models, and vertically integrated agentic workflows.
Today, that assumption is collapsing. We are witnessing a systemic dismantling of the moats, a reality check akin to the one recently delivered by Citadel in the public markets. In July 2026, Leopold Aschenbrenner’s AI-focused hedge fund, Situational Awareness, suffered a catastrophic 67% loss—wiping out roughly $35 billion—due to highly leveraged bets on the exact premise that AI progress required an endless expansion of physical computing infrastructure. When margin calls forced the fund into liquidation, Ken Griffin’s Citadel pragmatically stepped in, spending hours analyzing the risk before purchasing the distressed public equity portfolio. Just as Citadel capitalized on the collapse of an over-leveraged infrastructure thesis, the broader enterprise market is now capitalizing on the collapse of the proprietary AI bundle.
By applying Michael Porter’s Five Forces framework, we can clearly see two distinct disintermediations actively stripping power from frontier labs like OpenAI and Anthropic, shifting leverage permanently toward the enterprise buyer and the sovereign open-weight ecosystem.
Supplier Power and the Threat of Substitution
In Porter’s model, the bargaining power of suppliers—in this case, Western silicon and infrastructure providers—historically dictated the pace of AI advancement. The proprietary frontier labs, mapping their strategies to the legacy hardware curves of Moore and Kurzweil, continue to equate intelligence with sheer physical mass. This capital-intensive doctrine requires constant funding rounds at staggering valuations simply to maintain their hardware stockpiles. Aschenbrenner’s fund was built around this very premise: that the AI race would be won by those with the most physical infrastructure, such as advanced chips and data centers.
But a profound substitution effect was triggered when Meta open-sourced the Llama architecture. Originally deployed as a way to confront the meteoric rise of OpenAI, Llama now provides a foundational architecture that was rapidly absorbed and evolved by Chinese model makers. Now confronting severe geopolitical constraints that eliminated their traditional supplier power, China’s AI industry was forced to adopt a classic Amazon management principle: frugality as the primary spur for innovation.
Treating their lack of hardware as a mandate to rethink computational routing and model architecture, they established an entirely new deflationary curve. Capability was no longer bound to physical mass, but to algorithmic elegance—doing more with less. This constraint-driven ecosystem is now maturing into fully sovereign tooling. With DeepSeek anticipating a new generation of training chips by late 2026, the laboratory dictating the global price floor for agentic compute is poised to operate entirely beyond the reach of Western export controls. They substituted brute-force hardware for superior engineering.
The Bargaining Power of Buyers
However, hyper-efficient sovereign models alone could not disintermediate the enterprise market if buyers remained locked into legacy vendor workflows. The ultimate structural shift—and the dramatic expansion of buyer power—arrived at the software layer with OpenRouter Ori.
To understand why Stripe recently acquired OpenRouter for $7.5 billion, one must look at how the Ori framework systematically destroys vendor lock-in. Traditionally, the threat of substitute products was low because switching costs were prohibitively high; if an engineering team used a premium interface like Anthropic’s Claude Code, they were inextricably tied to Anthropic’s models and billing.
OpenRouter Ori eliminates those switching costs entirely through two mechanical tools: Ori Eval and Ori Harness. Ori Eval acts as an automated engineering suite that scans an enterprise codebase, tests real project prompts against multiple providers, and mathematically proves which model delivers the required accuracy and cost efficiency. Concurrently, Ori Harness intercepts the developer’s command-line interface, allowing them to use their preferred agent workflow but routing the backend requests to any model in OpenRouter’s network.
By separating the user interface, the underlying model, and the payer into completely independent variables, Ori perfectly commoditizes the intelligence layer. An enterprise can utilize a refined Western interface but rely on Ori Eval to prove that a DeepSeek model can execute the task perfectly for thirty-two cents, while Stripe simply clears the network transaction.
The New Reality
When these two forces combine, the proprietary bundle is entirely dismantled. The premium once placed on closed intelligence has collapsed, shifting enterprise value toward agnostic orchestration and cost-efficient execution. Just as the market swiftly punished leveraged bets on endless infrastructure expansion, it will punish enterprises that remain locked into artificially expensive proprietary ecosystems. It is within this unbundled, sovereign, and ruthlessly efficient new reality that we present this Holon review.
How DeepSeek’s Architecture Works
DeepSeek’s models are built to do less computation, hold less memory and move fewer bits for every word they produce. The DeepSeek-V4 technical report describes four parts of the design.
Mixture of experts. The model is divided into many small specialist sub-networks called experts. For each token, a router selects a few of them to run. DeepSeek-V4-Flash holds 284 billion parameters but uses 13 billion for each token, so each token costs about as much as a 13-billion-parameter model while drawing on what the full 284 billion learned. DeepSeek-V4-Pro holds 1.6 trillion parameters and uses 49 billion. DeepSeek’s version uses many fine-grained routed experts plus a few shared experts that every token passes through. The shared experts carry common knowledge, so the routed experts can specialize. Since DeepSeek-V3, the router keeps work spread evenly across experts by adjusting a bias per expert, which avoids the extra training penalty other designs use and protects quality.
Compressed attention. To write each new word, a model looks back at everything it has read, which it keeps in a memory called the key-value (KV) cache. At long context lengths this memory dominates the cost. DeepSeek-V3 shrank it with multi-head latent attention, which stores a compressed summary in place of the full cache. V4 goes further with a hybrid of two methods. Compressed Sparse Attention compresses the cache along the length of the text and then looks only at the positions most likely to matter. Heavily Compressed Attention compresses harder but looks at everything that remains. At a one-million-token context, V4-Flash needs 10% of the computation and 7% of the KV cache that DeepSeek-V3.2 needed.
Low-precision numbers. Most of V4’s weights are stored in FP8, an 8-bit number format, and the routed experts in FP4, a 4-bit format. Fewer bits per number means less memory and less data moved per token. DeepSeek-V3 was one of the first models to show that FP8 training works at very large scale.
Training and post-training. V4-Flash was trained on 32 trillion tokens. DeepSeek then trained separate specialist models for different domains and merged them into one model through distillation, in which the combined model learns from the specialists’ outputs. Reinforcement learning uses Group Relative Policy Optimization (GRPO), which scores a group of answers to the same prompt against each other and so avoids training a separate critic model.
Each of these choices cuts the cost of producing a token. Together they let the cheapest endpoint serve DeepSeek-V4-Flash at $0.036 per million input tokens.
Executive Summary
The market for agentic model capacity is moving from a frontier-priced supply, set by a few closed labs, to an open supply that a marketplace reprices every week. This report applies the holonic systems framework—proposed by Arthur Koestler in The Ghost in the Machine (1967) and formalized in multi-agent research—to classify that supply chain. A holon is simultaneously a self-contained whole and a functional part of a higher-order system. The duality maps directly onto the agent stack of 2026: a model snapshot, a hosted endpoint, a router, an agent harness and a harness of harnesses are each complete products, and each is also a component of the layer above it.
OpenRouter processes more than 10 trillion tokens a day across 400+ models for more than 10 million developers and companies, and announced on August 19, 2026 that it is joining Stripe. Though OpenRouter’s numbers do not include direct calls from Claude Code to Anthropic’s servers or Codex’s calls to OpenAI, the metrics at least provide an understanding of the long tail disruption.
In the week beginning September 14, 2026, DeepSeek led OpenRouter text requests at 25.4%, followed by Google (18.6%), OpenAI (17.0%, down 30% week over week), Z.ai (9.4%, up 31%), Qwen (6.7%) and Tencent (6.4%). Anthropic held 2.7%. Nine of the ten most-used models that week came from outside the US frontier labs. On capability, the gap has closed at the top: Qwen3.8 Max ties Claude Fable 5.1 at 53.4 on the Artificial Analysis Intelligence Index.
Price is only half of the shift. There’s an experiential issue that Claude is currently facing, where their new models are a bit more, studied in their response, yet do not do what you actually ask them to do, and practitioner threads on Reddit since Opus 5’s July 24 release report the same pattern: scope inflation, unrequested features and skipped instructions. Section 6 tests that claim against independent benchmarks.
The pattern—frontier labs set the capability ceiling while open and low-cost labs capture token volume—mirrors the commodity x86 server displacing proprietary systems in the 2000s. Agent harnesses have followed the models into commodity status: OpenRouter now runs twelve of them, from Claude Code to Prime Agent, on any model under one login and one bill.
This taxonomy defines six holon levels (H0–H5) for open models, routers and embeddable agents, evaluates 12 vendors across five capability dimensions, and assigns letter grades to each. The findings are intended to support agent-stack selection, investment thesis development, and the pricing of agent-produced work.
Section 1: The Holonic Framework Applied to the Agent Supply Chain
1.1 Origins of the Holon Model
Koestler described each holon as having two faces. The inward face governs its own internal logic, and the outward face interfaces with the system that contains it. A holarchy is a tree of such units, in which every node is both a complete decision-maker and a governed component of the node above.
The DFKI Research Report Holonic Multi-Agent Systems (1999) applied the model to distributed AI, proposing that agents give up parts of their autonomy to merge into a super-agent. The same structure now describes how agentic work is bought. A model snapshot is merged into an endpoint, the endpoint into a routed market, the market into an agent harness, and the harness into a governed, billed fleet.
The four-graph model of holons supplies the assessment template. In this study the interior graph is the model weights, the boundary membrane is the API and its pricing variant, the context graph is the unit’s position in routing, and the projection graph is what the buyer sees: tokens processed and the bill.
1.2 The Capability–Price Gap
OpenRouter’s June 2026 Fusion study ran 100 DRACO deep-research tasks through frontier and budget models. A fused budget panel of Gemini 3 Flash, Kimi K2.6 and DeepSeek V4 Pro scored 64.7%, against 65.3% for Claude Fable 5 alone, at about 50% of the cost. The same month, OpenRouter’s Subagent server tool let a frontier orchestrator hand self-contained tasks to a cheaper worker model, with up to 10 delegations per request.
Both results point the same way. Capability is increasingly assembled at runtime from cheaper parts, so the price of a finished task is set by routing and orchestration more than by the list price of the strongest model.
1.3 Holon Level Definitions
The taxonomy defines six holon levels (H0–H5) of ascending compositional depth. Each level inherits all capabilities of the levels below it.
Section 2: Holon Level Technical Specifications
H0 — Model Snapshot
The fundamental indivisible holon. An H0 unit is a trained model released as a dated version, with no serving, price or policy of its own. Snapshots now turn over faster than buyers can evaluate them: three DeepSeek V4 Flash versions (0423, 0731 and 4.1) sat in OpenRouter’s weekly top ten at the same time in September 2026. DeepSeek V4 Flash is a 284-billion-parameter mixture-of-experts model that activates 13 billion parameters per token and carries a 1-million-token context window.
Key Characteristics:
Dated version identifier
Parameter count and active parameters per token
Context window and supported reasoning modes
License terms that decide who may serve it
Three concurrent DeepSeek V4 Flash snapshots in one weekly top ten
H1 — Hosted Endpoint
The first holon with a price. An H1 unit is one provider serving one model at a stated price, service tier, latency and region. The same weights can carry very different prices at this level. DeepSeek V4 Flash 0423 is served by 15 providers, with input prices from $0.03556 per million tokens (StreamLake) to $0.21 (Azure), a 5.9-times spread for identical weights. Frontier models sit on fewer endpoints: Claude Opus 5 is hosted by five providers at $5 input and $25 output per million tokens.
Key Characteristics:
Price per million input and output tokens
Service tier (standard or discounted flex)
Availability and latency record (DeepSeek V4 Flash: 99.67% over three days, 0.77 s median latency)
Data retention and region policy
H2 — Routed Market
The first holon that chooses. An H2 unit selects an H1 endpoint for each request according to a rule the buyer sets. OpenRouter’s :floor variant sorts every eligible endpoint by price and admits discounted flex-tier endpoints to the pool, billing at whichever tier serves the request. The ~ prefix and -latest suffix in the run above follow the newest snapshot, so the buyer holds a standing order for the cheapest current version rather than a fixed model. OpenRouter’s Auto router, relaunched on August 10, 2026, classifies the task and ranks candidate models by real-world spend share over the previous seven days, drawn from more than 55 trillion tokens of weekly spend.
Key Characteristics:
Per-request provider selection by price, speed or policy
Variant suffixes (
:floor,:nitro,:free) as routing ordersMarket-driven model choice from trailing seven-day spend share
Price exposure that changes weekly with the market
H3 — Embeddable Agent Harness
The first compositional holon. An H3 unit wraps a model in a plan-act-observe loop with tools, files, memory and subagents, and it can be embedded in an IDE, a terminal, a messaging app or another product. The category is crowded and largely open source. On September 20, 2026, the most-used apps on OpenRouter were Hermes Agent (1.44 trillion tokens), Claude Code (505 billion), Kilo Code (493 billion), Cline (412 billion) and pi (305 billion).
Key Characteristics:
Session persistence and memory across runs
Tool use, shell and file access
Subagent delegation
Model-agnostic configuration through OpenAI-compatible endpoints
One-line install for most harnesses
H4 — Meta-Harness
The first holon that treats harnesses as interchangeable parts. An H4 unit runs any H3 harness on any H2-routed model under a single identity, budget and policy. OpenRouter’s Ori, first released on June 27, 2026, is the only H4 product graded in this study. Ori is a family of tools. Ori Harness launches twelve third-party agents with commands such as ori claude --model <id>. Ori Code is OpenRouter’s own terminal agent, built for working across many models. Ori Eval tests an app’s own prompts across models. Ori Intern deploys Slack agents whose skills work on any model, and Ori Desktop sends tasks to any model from the Mac menu bar. The CLI is live, Intern is in private testing, and Desktop is on a waitlist.
Key Characteristics:
One login (OAuth) in place of per-harness API keys
Harness and model chosen independently per task
A first-party agent (Ori Code) alongside twelve third-party harnesses
Embeddable agents in Slack and on the desktop (Ori Intern, Ori Desktop)
Organization guardrails and budgets applied to every harness
Model selection from evals on the app’s own prompts (Ori Eval)
One consolidated bill
H5 — Settlement and Governance Fabric
The enterprise governance holon. H5 adds no new agent capability. It supplies the membrane that governs every lower level: identity, budgets, guardrails, data residency, audit and payment. OpenRouter added in-region routing for US and EU data on September 9, 2026, a per-agent activity dashboard on August 17, and task classifiers that track what agents do and what it costs on July 24. Its pending combination with Stripe places the settlement layer of agentic work inside a payments company.
Key Characteristics:
Identity federation and workspace permissions
Budgets and spend limits enforced per organization
Data residency controls (US and EU in-region routing)
Per-agent, per-model usage attribution
Payment, fraud control and settlement
Section 3: Vendor Taxonomy and Grading Framework
3.1 Grading Methodology
Each vendor is evaluated across five dimensions, each scored A through D and weighted equally.
This study replaces the Developer Experience dimension of the May 2026 agentic study with Price Efficiency. Most harnesses now install with one command, so developer experience no longer separates vendors, while price per finished task now does.
Section 4: Vendor Profiles and Grades
4.1 Frontier Labs (H0–H3)
Anthropic
Max Holon Level: H3 | OpenRouter request share: 2.7% (week of Sep 14, 2026)
Anthropic sets the capability ceiling. Claude Fable 5.1 shares first place on the Artificial Analysis Intelligence Index at 53.4, and Claude Opus 5, released July 24, 2026, scores 77.0 on the Coding Index and 93.7% on GPQA Diamond. Opus 5 is priced at $5 input and $25 output per million tokens across five hosts, including Amazon Bedrock, Google Vertex and Azure. Claude Code is the second most-used app on OpenRouter at 505 billion tokens a day, and Ori runs it against non-Anthropic models with a single flag. The key limitation is price: Anthropic’s share of OpenRouter requests is one-tenth of DeepSeek’s, and its harness now carries competitor tokens.
OpenAI
Max Holon Level: H3 | OpenRouter request share: 17.0% (down 30% week over week)
OpenAI competes at both ends of the price curve. GPT-5.6 Luna was the most-used model on OpenRouter over the trailing 30 days at 49.4 trillion tokens, up 209%, while GPT-6 Astra scores 52.7 on the Intelligence Index. Codex runs at 150 billion tokens a day and is one of the twelve harnesses Ori supports. The key limitation is volatility: request share fell 30% in a single week as open models gained.
4.2 Open-Weight and Low-Cost Labs (H0–H3)
DeepSeek
Max Holon Level: H3 | OpenRouter request share: 25.4% (first)
DeepSeek is the volume leader. DeepSeek V4 Flash 0423, released April 24, 2026, is a 284B-parameter mixture-of-experts model with 13B active parameters, a 1-million-token context window, and a 56.2 Coding Index score at maximum effort. It is served by 15 providers with 99.67% availability, from $0.03556 input and $0.07112 output per million tokens. Three V4 Flash snapshots ranked in the weekly top ten, with V4.1 Flash up 219% week over week at 15.8 trillion tokens. DeepSeek Harness, its own H3 agent, processes 131 billion tokens a day and runs under Ori. The key limitation is that governance depends on which endpoint serves the request, so data policy is set at H1 or H5 rather than by the lab.
Z.ai
Max Holon Level: H1 | OpenRouter request share: 9.4% (up 31%)
Z.ai’s GLM 5.3 Flash was the single most-used model on OpenRouter on September 20, 2026, at 2.52 trillion tokens, and ranked second for the week at 14.1 trillion. GLM-5.3 at maximum effort scores 44.8 on the Intelligence Index. Z.ai also ships batch and FlashX variants that target the same price-sensitive agent traffic. The key limitation is the absence of a native harness, so Z.ai depends on third-party H3 products for agent distribution.
Qwen (Alibaba)
Max Holon Level: H1 | OpenRouter request share: 6.7% (up 25%)
Qwen holds the strongest capability claim among the non-frontier labs. Qwen3.8 Max ties Claude Fable 5.1 at 53.4 on the Intelligence Index, and a free Qwen3.8 27B variant entered OpenRouter’s trending list at 12 billion tokens in its first week. The key limitation is that Qwen’s volume trails DeepSeek and Z.ai despite its benchmark position, which suggests that price and snapshot cadence decide routed volume more than peak score.
4.3 Routed Market and Settlement (H2–H5)
OpenRouter
Max Holon Level: H5 | Volume: 10+ trillion tokens per day, 400+ models, 10M+ developers
OpenRouter is the market-making layer of this taxonomy. Its variant suffixes turn a model ID into a routing order, and :floor sorts all eligible endpoints by price and admits flex-tier capacity. The Auto router ranks models by seven-day spend share; in OpenRouter’s own test on MMLU Pro it scored 85.2% against 86.6% for the previous router at $140.93 instead of $393.34. Fusion and the Subagent tool assemble frontier-grade answers from cheaper models. H5 functions include in-region routing for the US and EU, per-agent activity reporting, task classifiers, organization guardrails and budgets, and a pending combination with Stripe announced August 19, 2026. The key limitation is concentration: one marketplace now prices, routes and settles a large share of open-model agent traffic.
4.4 Embeddable Agent Harnesses (H3)
Hermes Agent (Nous Research)
Max Holon Level: H3 | Volume: 1.44 trillion tokens per day (first among OpenRouter apps)
Hermes Agent is an open-source agent that runs persistently with memory across sessions, builds reusable skills from experience, and ships with more than 40 tools including web search, browser automation and vision, plus scheduled automations and subagents. It leads OpenRouter app volume by nearly three times the next harness. The key limitation is that governance is left to the operator who hosts it, unless it runs under an H4 or H5 layer.
Kilo Code and Cline
Max Holon Level: H3 | Volume: 493 billion and 412 billion tokens per day
Both are open-source coding agents that run inside the IDE and edit files, run terminal commands and use browser automation. Kilo Code also runs in JetBrains and the CLI. Both are supported by Ori. The key limitation is weak differentiation: each performs the same plan-edit-test loop as Claude Code and Codex, on the same routed models.
OpenClaw
Max Holon Level: H3 | Volume: 166 billion tokens per day
OpenClaw is an open-source agent that connects to messaging apps and takes actions for the user, including running commands, browsing, managing files and sending email. It extends H3 beyond software development into general work. The key limitation is permission scope: an agent that can send email and run commands from a chat thread needs H5 controls that the harness does not supply by itself.
Prime Agent (Prime Intellect)
Max Holon Level: H3 | Funding: $130M Series A at a $1B valuation (July 2026)
Prime Agent is an MIT-licensed harness that adds session persistence, subagents, memory, skills and self-improvement to any model with an OpenAI-compatible endpoint. Prime Intellect reports 95.5% on ARC-AGI-3 when paired with Claude Opus 5; that figure is vendor-reported and not independently verified. Ori added Prime Agent support on September 15, 2026. The key limitation is time in market, measured in weeks.
4.5 Meta-Harness (H4)
Ori (OpenRouter)
Max Holon Level: H4 | Price: free; inference billed through OpenRouter
Ori is OpenRouter’s family of agent tools for the terminal, Slack and the desktop. The CLI shipped as version 0.1.0 on June 27, 2026. Commands for Claude Code, Codex, OpenCode and Hermes followed on July 29, headless ori code with non-Anthropic model support on August 6, Prime Agent, DeepSeek Harness and Grok Build on August 15, and a first-party agent loop as the default on August 21. Ori now launches twelve agents, and one OAuth login replaces the 12+ environment variables Claude Code otherwise needs. Guardrails, allowlists, workspace permissions and budgets set on OpenRouter apply to every request, and all costs land on one bill.
The family extends past the CLI. Ori Eval finds an app’s prompts, builds evals for them and scores models with a judge model. Ori Intern puts agents in Slack that learn skills from ordinary requests; the skills run on any model and can be shared between interns across a team. Ori Desktop sends tasks from the Mac menu bar to any harness and model. Ori separates three decisions that were bundled until this year: which agent does the work, which model thinks, and who pays. The key limitation is maturity: only the CLI is generally available, Intern and Desktop are pre-release, and no service-level commitment is published.
Section 5: Consolidated Vendor Scorecard
The two highest grades go to the layers that sit above the models: the routed market with its settlement fabric, and the meta-harness.
Section 6: Structural Market Observations
6.1 The Shift from Frontier to Open Supply
The volume of enterprise routing offers a clear empirical signal: the monopoly of the proprietary frontier is dissolving. In the week of September 14, 2026, a coalition of open and sovereign models—DeepSeek, Z.ai, Qwen, and Tencent—processed 47.9 percent of all text requests on OpenRouter. On the limited overall traffic to language model providers, the portion routed to OpenRouter towards OpenAI and Anthropic accounted for just 19.7 percent.
While frontier labs still nominally occupy the highest tier of the Intelligence Index, the functional margin of superiority has vanished. Qwen3.8 Max now ties Claude Fable 5.1 at the very top of the board. OpenRouter’s Fusion studies mirror this convergence in real-world application, demonstrating that a routed panel of three budget-tier models can come within one percent of Claude Fable 5’s performance at half the cost. The resulting market dynamic is a strict division of labor. Frontier models are increasingly relegated to narrow orchestrator and advisor roles—functions formalized by OpenRouter’s native tools—while highly efficient, open-weight models carry the vast majority of the token volume.
6.2 Directability as the Second Advantage of Open Models
Beyond sheer cost efficiency, open models possess a secondary, often underappreciated advantage: directability. In practical application, open architectures adhere to developer constraints far more strictly than their proprietary counterparts. Community telemetry reflects this shift; developer forums are inundated with reports of Opus 5 exhibiting severe scope inflation, spawning unrequested subagents, and bypassing explicit boundaries. In one documented instance, a prompt that required 75,000 tokens on Opus 4.6 expanded to consume 150,000 tokens on Opus 5, turning a simple sitemap correction into an unrequested, autonomous site rebuild.
Independent benchmarks quantify the premium exacted by this proprietary friction. Artificial Analysis evaluates both models at maximum reasoning effort, scoring Claude Opus 5 at 51 on its Intelligence Index and DeepSeek V4.1 Flash at 39. However, executing that index cost $7,275 on Opus 5 compared to a mere $477 on Flash. Effectively, each marginal point of capability costs an enterprise 12 times more on Opus 5, which also generates tokens at a quarter of the speed (59 tokens per second versus Flash’s 236). While third-party reviewers credit frontier models with superior coherence across multi-hour autonomous horizons, these benchmarks fail to measure literal adherence to a user’s brief. For the enterprise buyer, directability has become just as critical as peak capability. Practitioner guidance has converged on a clear routing logic: default to the low-cost open model for bounded execution, escalating to the frontier only when correctness-critical or long-horizon orchestration is strictly required.
6.3 Disruptive Pricing in the Routed Market
In this unbundled ecosystem, price is no longer a fixed property of the model; it is a dynamic property of the route. The exact same DeepSeek V4 Flash weights can cost anywhere from $0.03556 to $0.21 per million input tokens, entirely dependent on which endpoint serves the request. By utilizing a dynamic model ID, a buyer essentially holds a standing order that resolves both the model snapshot and the most efficient provider at the exact moment of request.
It is worth noting that the following numbers only apply to the portion of the market that OpenRouter tracks. The market supporting this routed infrastructure is violently dynamic. In a single week, DeepSeek V4.1 Flash volume surged 219 percent, while OpenAI’s request share dropped by 30 percent. Tencent’s Hy4 preview absorbed 47.1 trillion tokens in its inaugural month, and stealth models like Ox Alpha process tens of trillions of tokens in a matter of weeks. Because OpenRouter’s auto-router dynamically ranks models based on a rolling seven-day spend, static monthly budgets are obsolete. Enterprise cost is now best measured per finished run, adapting in real time to the cheapest available compute.
6.4 Embeddable Agents as a Commodity
The interfaces that harness these models have undergone a similar commoditization. On September 20, 2026, all ten of the most utilized applications on OpenRouter were agent harnesses, and nearly half (including Hermes Agent, Kilo Code, Cline, and OpenClaw) were entirely open source. Frontier labs now distribute their own harnesses natively, and Ori allows developers to run over a dozen of them behind a single command-line interface. When a harness is free, open, and swappable with a single flag, it possesses the baseline economics of a raw, hosted endpoint. Strategic differentiation and profit margins no longer reside in the interface; they have migrated up the stack to governance, settlement, and orchestration.
6.5 Ori and the Unbundling of the Agent
It is immediately clear why Stripe acquired OpenRouter at a bargain-basement price of $7 billion—a stark contrast to the trillion-dollar valuations of proprietary labs. The true disruption lies in Ori, a product that fundamentally unbundles the AI agent itself.
Until recently, choosing an agentic harness meant accepting a rigid, vertically integrated bundle: an Anthropic interface required an Anthropic model and an Anthropic invoice. Ori shatters this paradigm by separating the harness, the underlying model, and the payer into completely independent variables. Empowered by rigorous empirical evidence from platforms like Artificial Analysis—which proves the capability delta between proprietary and open-weight models has practically closed—buyers can now route workloads based on pure price efficiency. A user can operate within a premium frontier interface but dynamically swap the intelligence engine to a far cheaper open-weight model, all while Stripe clears the transaction. Absolute pricing and routing power has shifted from the model maker to the buyer.
The economic reality of this unbundling is devastating to traditional organizational structures. For the last decade, the product manager has been the most celebrated and highly compensated player in the tech ecosystem. Yet, a full financial quarter’s worth of that exact cognitive labor—at least at a highly competent first-draft level—can now be replicated for roughly thirty-two cents. If we strip away the navigation of corporate matrix politics, the actual, tangible output of human synthesis has become an absolute commodity. While professionals will inevitably reinvent themselves as strategic editors, the unavoidable math of this transformation means slimming down teams of tens into teams of one.
We are already witnessing the macroeconomic fallout. Gartner, long considered the ultimate market proxy for the implied value of white-collar expertise, has seen its valuation collapse by 71 percent from its peak. The market is aggressively pricing in a new reality: premium cognitive synthesis is no longer scarce.
This structural collapse has triggered a desperate institutional backlash, encapsulated by The New York Times’s ongoing copyright trial against OpenAI and Microsoft. In a fierce attempt to defend the legacy moat of human synthesis, The Times claims that AI models represent a direct substitution that siphons web traffic and belittles the core concept of intellectual property. The cries from legacy institutions to close Pandora’s box are growing deafening. But with the unbundling of the agent complete and premium compute routing seamlessly for pennies, the box is already empty. The era of the expertise premium is over.
Section 7: Forward-Looking Considerations
7.1 Settlement Becomes the Moat
With models and agent harnesses effectively priced down to commodity levels, the only defensible position remains at the top of the Holon architecture (H5). This is the layer where identity, budgets, data residency, and payments are strictly enforced. OpenRouter’s integration with Stripe embeds this governance layer directly inside a global payments network. The vendors that own the financial settlement for agentic work will dictate the terms for every lower level of the AI stack.
7.2 Data Residency at the Route
The introduction of in-region routing for the US and EU in September 2026 shifts the burden of data sovereignty away from the model maker and places it squarely on the router. As open-weight models proliferate across global endpoints, enterprise buyers can simultaneously secure the lowest market price and enforce strict residency policies, provided the routing infrastructure evaluates both variables on every single request.
7.3 Pricing per Outcome
When three months of product management labor consumes a mere 33 cents of inference compute, traditional SaaS pricing models break down. Both seat-based software licenses and token-based API billing lose their intrinsic link to actual business value. The next definitive pricing signal for the enterprise will be per-task or per-outcome billing for agent-produced work—where the cost is inextricably tied to the strategic quality of the brief, the rigorous review of the output, and the human author who ultimately stakes their reputation on the result.
Closing
In September 2026, the theoretical became empirical. We observed a lean, 27-billion-parameter model achieve absolute parity with a 753-billion-parameter on independent indices. We watched a 33-cent agentic session yield production-ready work, while a 90-dollar frontier session generated output that was ultimately discarded. The cheaper model did not just cost less; it executed the brief. With OpenRouter continuously repricing this commoditized capacity, and Ori seamlessly decoupling the agent interface from the underlying model with a single command, the structural unbundling of the enterprise AI stack is now a functional reality. This disintermediation leaves every stakeholder in the ecosystem with a stark, unavoidable mandate:
Engineering leaders must now mathematically justify their routing logic, determining exactly which workloads still demand a proprietary frontier model when that choice carries a 12x cost premium per unit of capability.
Enterprise executives and institutional investors must rigorously stress-test their capital assumptions: how long can trillion-dollar valuations, predicated on the supremacy of Western hardware monopolies, survive in a market where a fiscal quarter of product management executes for thirty-three cents?
Finally, policymakers are forced to confront the limits of their geopolitical leverage, asking what utility export controls truly possess when the most aggressive deflationary curve in the industry has migrated to sovereign silicon.
The architecture of this market has been permanently reset. The next definitive grading point for this taxonomy will arrive the moment the first DeepSeek model is trained entirely on domestic hardware.
Methodology Note
This taxonomy draws on OpenRouter’s public rankings (usage data through September 20, 2026, licensed CC BY 4.0), OpenRouter model pages, the Ori product page, changelog and documentation, Artificial Analysis model comparisons, third-party reviews, and Reddit threads on Claude Opus 5.
References
OpenRouter LLM Rankings - model, app and author market share through Sep 20, 2026
Ori - Use Every Model, Right Where You Are - Ori Harness, Code, Eval, Intern and Desktop
Ori Changelog - CLI releases from 0.1.0 (Jun 27, 2026) to 0.10.1 (Aug 24, 2026)
Ori Harness documentation - twelve supported agents, organization controls
Ori Harness announcement - Aug 4, 2026
OpenRouter is Joining Stripe - Aug 19, 2026; 10T+ tokens a day, 400+ models, 10M+ developers
Floor Variant - Lowest-Cost Inference - price sorting and flex-tier eligibility
Model Routing Powered by Wisdom of the Market - Auto router, Aug 10, 2026
Surpassing Frontier Performance with Fusion - DRACO results, Jun 12, 2026
Subagent: Let Your Model Delegate the Busywork - Jun 16, 2026
OpenRouter Announcements - in-region routing, activity dashboard, classifiers
DeepSeek V4 Flash on OpenRouter - specifications, 15 providers, price range
Claude Opus 5 on OpenRouter - pricing and release date
DeepSeek V4.1 Flash vs Claude Opus 5, Artificial Analysis - Intelligence Index, cost and tokens to run it
Reddit says Opus 5 is a genius that will not shut up - summary of r/ClaudeCode, r/ClaudeAI and r/Anthropic threads with vote counts
The Opus 5 Experience, r/ClaudeCode - 2,743 votes
Switched back to Opus 4.8, r/Anthropic - 592 votes
Is Opus 5 actually that bad or just Reddit hype, r/ClaudeAI - 527 votes, 427 replies
Claude Opus 5 Over-Engineering: Reddit Reaction - Aug 6, 2026 r/ClaudeAI thread and workarounds
DeepSeek V4-Flash vs Claude Opus 5: The 36x Cost Gap - pricing, LiveBench agentic-coding ranking, long-horizon coherence
DeepSeek V4 Flash vs Claude Haiku, Sonnet and Opus: A Coding Review - Terminal-Bench results and routing recommendation
Prime Agent Explained - license, funding, vendor-reported ARC-AGI-3 score
Koestler, A. The Ghost in the Machine (1967)
DFKI Research Report, Holonic Multi-Agent Systems (1999)





