The Emergence of Machine Labor Markets An Adversarial Investigation into the Standardization and Financialization of AI-Completed Work Research Monograph Investment Committee Due Diligence July 2026 This document contains confidential research materials.
1 Abstract This research monograph presents an adversarial investigation into whether work per- formed by autonomous AI agents will evolve into standardized, financialized economic markets comparable to commodities such as oil, electricity, freight, and cloud compute. The central thesis examines the conditions under which machine labor transitions from an input-priced resource (compute, tokens) to an output-priced commodity (verified units of completed work). Through systematic analysis of twelve historical commodity markets, we identify the institutional, technological, and economic preconditions for market formation: standard- ization of underlying units, credible benchmark pricing, trusted measurement infrastruc- ture, sufficient heterogeneity and volatility to require hedging, and network effects that protect dominant institutions. We then apply this framework to machine labor, distinguish- ing between intelligence, compute, model access, tokens, software, agents, and verified outcomes. The research develops a rigorous framework for defining Standardized Machine Work Units across eight task categories: customer support, software engineering, legal review, insurance claims processing, financial operations, cybersecurity, research, and robotics. We derive quality-adjusted and risk-adjusted pricing formulas that incorporate not only model API and compute costs, but also retries, failures, human intervention, latency, com- pliance, and sovereign constraints. Our jurisdictional analysis reveals significant fragmentation across U.S., Chinese, Eu- ropean, and open-weight model ecosystems, leading to the concept of a “sovereignty premium”—the additional price paid to ensure AI work satisfies specific national, secu- rity, and data-residency requirements. We propose eight distinct index grades to capture these variations. The research models ten realistic market scenarios, evaluates a ten-stage market evo- lution pathway, and develops a five-year business model with twelve revenue streams. We conduct extensive falsification analysis, testing fifteen objections including heterogeneity, price collapse, compute dominance, vendor lock-in, and regulatory prohibition. Based on this comprehensive investigation, we conclude that machine labor markets are conditionally inevitable: they become likely only after specific thresholds in agent reliability, enterprise spending, standardization maturity, and capacity volatility. The most defensible initial position is as an independent benchmark administrator and data com- pany, with exchange infrastructure emerging through partnerships with regulated financial institutions rather than direct startup development. The strongest version of the proposed startup, provisionally called OnIndex, would answer one expensive enterprise question: “What is the true cost per verified outcome, accounting for all hidden costs and quality variations?” The decisive 90-day experiment involves measuring real cost disparities across three models for a standardized customer support task basket, demonstrating that token prices are weak predictors of final verified outcome costs.
Contents 2 Contents
Contents 3 Executive Summary Research Objectives This monograph addresses a fundamental question in the economics of artificial intelligence: Will work performed by autonomous AI agents evolve into standardized and financialized mar- kets comparable to oil, electricity, freight, cloud compute, advertising, and other critical re- sources? We investigate this question with an explicitly adversarial posture—actively attempt- ing to disprove the thesis while rigorously identifying the conditions under which it might hold true. The Central Thesis The research examines the proposition that compute is the input required to produce artificial intelligence, while AI agents transform compute into productive economic work. If compute is developing commodity-market characteristics (benchmark prices, capacity markets, forward curves, reservations, hedging), the next economic transition may be from pricing inputs (GPU- hours, tokens) to pricing outputs (customer issues resolved, software tasks completed, contracts reviewed, insurance claims processed). The proposed entity, provisionally called OnIndex, would become the independent mea- surement, benchmark-pricing, verification, and settlement layer for this machine labor. Its primary question: “What is the current risk-adjusted market price of one verified unit of AI- completed work?” Methodology Our approach combines:
- Historical analysis: Systematic study of twelve commodity markets to identify patterns in standardization, benchmark formation, and institutional development
- Economic framework development: Rigorous definitions of machine labor categories, standardized work units, and pricing methodologies
- Scenario modeling: Ten detailed realistic market scenarios testing transaction structures and value creation
- Falsification analysis: Active attempts to disprove the thesis through fifteen structured objections
- Empirical hypothesis design: Ten testable hypotheses with experimental frameworks
Contents 4 Key Findings Historical Market Patterns Analysis of crude oil, refined fuels, electricity, natural gas, agricultural commodities, freight, cloud computing, telecom bandwidth, carbon credits, advertising, labor outsourcing, foreign exchange, and interest rate markets reveals consistent patterns: • Markets emerge when underlying resources become sufficiently important, heteroge- neous, and volatile • Standardization requires both technical specification and institutional trust • Dominant benchmarks emerge through network effects, not just accuracy • Forward markets develop when spot price volatility creates hedging demand • Exchange infrastructure follows, rather than leads, market liquidity Machine Labor Definition Machine labor is a coherent economic category distinguishable from compute, intelligence, tokens, and software. The claim that “companies want verified outcomes, not tokens” holds for structured, repeatable tasks but breaks down for exploratory, creative, or highly variable work. Standardization Framework A Standardized Machine Work Unit must specify: task category, difficulty, expected output, success criteria, quality threshold, latency, human intervention limits, retry rate, security re- quirements, data residency, model provenance, auditability, liability, delivery period, and veri- fication process. Customer support emerges as the most standardizable initial market. Pricing Methodology True Cost per Verified Outcome = (Model + Compute + Tools + Infrastructure + Retries + Human Intervention + Expected Failure Loss + Compliance Cost) / Accepted Outcomes. Token prices prove to be weak predictors of final verified outcome costs. Sovereignty Premium Geopolitical fragmentation creates measurable price premiums. A Chinese open-weight model running privately on U.S. infrastructure differs fundamentally from a China-hosted API in terms of security, compliance, and acceptable use.
Contents 5 Market Evolution A ten-stage progression from Measurement to Derivatives is theoretically coherent but stages 7-10 (forward contracts through derivatives) require conditions that may never materialize. Stages 1-5 (measurement through spot procurement) represent a viable business trajectory. Conclusion Assessment We assess four possible conclusions: A. Structurally Inevitable: Highly likely under current trends B. Conditionally Inevitable: Likely only after specific thresholds C. Plausible but Premature: Useful index, but exchange thesis too early D. Incorrect Market Analogy: Remains software/services, not commodity Assessment: B (Conditionally Inevitable) with elements of C for advanced stages. Strategic Recommendations • First Product: True cost measurement for AI customer support • First Customer: Mid-market SaaS company with 500K+annualAIspendFirst Dataset : 10, 000 + resolvedsupportticketsacross3 + models • 90-Day Experiment: Measure cost disparities for standardized task basket • Continuation Evidence: Demonstrate that token prices predict <60% of cost variance • Abandonment Trigger: Enterprises refuse to share outcome data or pay for benchmarking Risk Factors Primary risks include: agent work heterogeneity preventing standardization; model price col- lapse eliminating hedging demand; cloud provider vertical integration; government prohibition of cross-border agent markets; and insufficient liquidity for forward contracts.
Contents 6 Document Structure This monograph contains twenty sections covering historical analogies, theoretical frameworks, market analysis, business modeling, competitive assessment, regulatory considerations, falsi- fication analysis, and empirical research design. Appendices provide detailed formulas, task specifications, and interview protocols.
3 Crude Oil 7 Part I Historical and Theoretical Foundations 1 Historical Market Analogies 2 Introduction Understanding whether machine labor will evolve into commodity markets requires systematic study of how other resources made this transition. This section examines twelve markets that have developed standardized units, benchmark prices, forward contracts, and exchange infras- tructure. For each market, we analyze: (1) what was initially purchased, (2) why markets were initially fragmented, (3) how units were standardized, (4) who created trusted benchmarks, (5) what caused spot markets to emerge, (6) why forward contracts became useful, (7) who became dominant institutions, (8) what network effects protected them, (9) which markets failed, and (10) the relevance to machine labor. 3 Crude Oil 3.1 Initial Purchase Unit [Evidence] In the early petroleum industry (1860s-1880s), oil was purchased in wooden barrels of varying sizes, with quality determined by visual inspection and basic specific gravity tests. Buyers assessed color, viscosity, and smell to estimate refining yields. Standard Oil’s dom- inance under Rockefeller introduced 42-gallon barrels as an industry standard by the 1870s, though quality grading remained inconsistent. 3.2 Market Fragmentation [Evidence] Regional markets developed around production centers (Pennsylvania, Ohio, Cal- ifornia) with limited interconnection. Transportation via rail and pipeline created local mo- nopolies. Refining technology varied, meaning crude quality requirements differed by refinery configuration. The lack of standardized quality meant prices could not be compared across regions. 3.3 Standardization Process [Evidence] The American Petroleum Institute (API) established standardized gravity measure- ments in the 1920s. The introduction of chromatography in the 1950s enabled precise chemical
3 Crude Oil 8 composition analysis. [Inference] Standardization emerged only after the industry consoli- dated sufficiently to make universal standards economically rational for dominant players. 3.4 Benchmark Creation [Evidence] West Texas Intermediate (WTI) emerged as the U.S. benchmark in the 1980s due to: (1) large, stable production volumes, (2) storage hub at Cushing, Oklahoma with pipeline connectivity, (3) delivery point specifications, and (4) transparent pricing through NYMEX futures. [Inference] Benchmarks require physical infrastructure (storage, delivery points) as much as market design. 3.5 Spot Market Emergence [Evidence] Spot markets developed in the 1970s following oil price deregulation in the United States and the collapse of the Seven Sisters’ vertical integration. The 1973 oil crisis demon- strated the value of price transparency for non-integrated players. [Inference] Spot markets emerge when vertical integration breaks down and independent participants need price discov- ery. 3.6 Forward Contract Utility [Evidence] Futures contracts (NYMEX WTI launched 1983) became essential for: (1) produc- ers hedging against price declines, (2) refiners hedging against price increases, (3) inventory valuation, and (4) long-term contract pricing. Open interest in WTI futures grew from negligi- ble in 1983 to over 2 million contracts by 1990. [Inference] Forward markets require sufficient volatility to create hedging demand plus standardized deliverable units. 3.7 Dominant Institutions [Evidence] S&P Global Platts (originally Platts) became the dominant price reporting agency for physical oil markets through daily assessments. Intercontinental Exchange (ICE) and NYMEX (now CME Group) dominate futures. [Inference] Dominant institutions emerge through: (1) first-mover advantage in data collection, (2) editorial independence, (3) market participant adoption, (4) regulatory recognition. 3.8 Network Effects [Evidence] Platts’ Dated Brent assessment became the benchmark for pricing 70%+ of global crude because: (1) long-term contracts referenced it, creating lock-in, (2) derivatives markets built on it, requiring participants to hedge using Platts, (3) regulatory frameworks adopted it,
5 Electricity 9 (4) new entrants faced switching costs. [Inference] Benchmark network effects are stronger than exchange network effects because contracts reference benchmarks directly. 3.9 Failed Markets [Evidence] The Dubai crude benchmark declined in relevance due to declining production volumes and increased participation by speculators rather than physical traders. The Bonny Light (Nigeria) benchmark failed to achieve global adoption due to production disruptions and lack of liquidity. [Inference] Benchmarks fail when underlying physical volumes become insufficient or unreliable. 4 Refined Fuels 4.1 Gasoline and Diesel Markets [Evidence] Refined product markets developed standardized specifications (octane rating, sul- fur content, cetane number) through industry associations and government regulation. The New York Harbor became the dominant delivery point for U.S. gasoline futures. [Inference] Prod- uct standardization requires either regulatory mandate or industry-wide coordination through associations. 4.2 Jet Fuel [Evidence] Jet fuel markets standardized through aviation regulatory requirements (ASTM specifications). The Singapore market emerged as the Asian benchmark due to: (1) major air- port hub, (2) storage infrastructure, (3) transparent pricing from Platts. [Inference] Regulatory requirements can drive standardization more effectively than market forces alone. 5 Electricity 5.1 Initial Purchase Unit [Evidence] Early electricity markets (1890s-1930s) involved vertically integrated utilities with no wholesale market. The Public Utility Holding Company Act (1935) created regulated mo- nopolies with cost-plus pricing. [Inference] Electricity initially lacked markets entirely due to natural monopoly characteristics and regulatory capture.
6 Natural Gas 10 5.2 Market Fragmentation [Evidence] Post-deregulation (1990s), electricity markets fragmented by: (1) transmission con- straints creating regional markets, (2) regulatory differences across jurisdictions, (3) temporal variation (hourly, daily, seasonal), (4) lack of storage making arbitrage difficult. [Inference] Electricity markets remain more fragmented than oil due to transmission constraints and lack of storage. 5.3 Standardization Process [Evidence] Standardization occurred through: (1) nodal pricing (locational marginal pricing), (2) ISO/RTO formation (PJM, ERCOT, CAISO), (3) standardized contract terms (ISDA mas- ter agreements), (4) FERC Order 888 (1996) requiring open access transmission. [Inference] Electricity standardization required regulatory mandate due to network externalities in trans- mission. 5.4 Benchmark Creation [Evidence] Multiple benchmarks emerged by region: PJM Western Hub (Mid-Atlantic), ER- COT North Hub (Texas), SP15 (Southern California). No single global benchmark exists. [In- ference] Electricity’s non-storability and transmission constraints prevent global benchmark formation. 5.5 Forward Contract Utility [Evidence] Forward markets developed for: (1) generators hedging fuel price risk, (2) utilities hedging procurement costs, (3) large consumers managing budget volatility. Capacity markets emerged to ensure resource adequacy. [Inference] Forward markets in electricity serve both price hedging and resource adequacy functions. 6 Natural Gas 6.1 From Long-Term Contracts to Spot Markets [Evidence] The U.S. natural gas market transitioned from long-term take-or-pay contracts (1960s-1980s) to spot markets following FERC Order 436 (1985) and Order 636 (1992), which mandated pipeline unbundling. Henry Hub emerged as the national benchmark due to: (1) pipeline interconnections from 13 interstate pipelines, (2) storage facilities, (3) coastal access for LNG. [Inference] Pipeline unbundling was necessary for spot market formation—physical infrastructure separation enabled market price discovery.
7 Agricultural Commodities 11 6.2 Benchmark Dominance [Evidence] Henry Hub became the pricing point for NYMEX natural gas futures in 1990. By 2000, over 70% of U.S. natural gas was priced relative to Henry Hub. [Inference] Hub benchmarks require physical connectivity and storage infrastructure. 6.3 LNG and Global Markets [Evidence] LNG markets initially used oil-price-linked long-term contracts (JCC in Asia, Brent in Europe). Spot LNG markets emerged after 2008 due to: (1) U.S. shale gas exports, (2) destination clause removal in contracts, (3) oversupply periods. [Inference] Global commodity markets can emerge even with different regional benchmarks when arbitrage becomes possible. 7 Agricultural Commodities 7.1 Grain Markets [Evidence] Chicago Board of Trade (CBOT, founded 1848) developed standardized grades for wheat, corn, and oats. The innovation was not standardization itself but transferable warehouse receipts—enabling trading without physical movement. [Inference] Standardization requires not just quality grades but also transfer mechanisms (warehouse receipts, electronic registra- tion). 7.2 Soybeans and Oilseeds [Evidence] Soybean futures launched in 1936. The Chicago Mercantile Exchange (CME) introduced options on futures in 1984. [Inference] Derivatives markets develop only after spot markets achieve sufficient liquidity and standardization. 7.3 Soft Commodities [Evidence] Coffee, cocoa, and sugar developed international benchmarks (ICE) but face chal- lenges: (1) quality variation by origin, (2) weather dependence, (3) political risk in producing regions. [Inference] Agricultural commodities with high quality variation and supply uncer- tainty require broader quality bands and more frequent contract adjustments.
9 Cloud Computing 12 8 Freight and Shipping 8.1 Initial Fragmentation [Evidence] Shipping rates were historically negotiated bilaterally between shipowners and cargo interests. Routes, vessel sizes, and timing created infinite variation in pricing. [In- ference] Highly heterogeneous services resist standardization without forced abstraction. 8.2 Standardization Through Baltic Exchange [Evidence] The Baltic Exchange (founded 1744) developed freight indices: Baltic Dry Index (BDI, 1985) for dry bulk, Baltic Dirty Tanker Index (BDTI) for crude oil. Standardization occurred through: (1) route definitions, (2) vessel size categories (Capesize, Panamax, Supra- max), (3) time charter equivalents. [Inference] Even highly heterogeneous services can be stan- dardized through: (1) route abstraction, (2) vessel/category classification, (3) time-equivalent pricing. 8.3 Forward Freight Agreements [Evidence] FFAs emerged in the 1990s as OTC derivatives, cleared through clearinghouses (LCH, SGX). [Inference] Forward markets can develop even in fragmented underlying markets through index-based settlement. 9 Cloud Computing 9.1 AWS Spot Instances [Evidence] Amazon Web Services introduced Spot Instances in 2009 as an auction-based mar- ket for excess compute capacity. Initial pricing was based on a uniform-price auction. In 2017, AWS switched to provider-managed pricing, effectively ending the auction mechanism. [Infer- ence] Auction-based spot markets can fail when providers prefer managed pricing to maximize revenue. 9.2 Standardization [Evidence] Compute standardization occurred through: (1) virtual machine instance types (t2.micro, m5.large), (2) standardized units (vCPU, GB RAM), (3) per-hour billing, (4) API- driven provisioning. [Inference] Cloud compute achieved standardization through virtualiza- tion abstraction and API uniformity.
11 Carbon Credits 13 9.3 Market Limitations [Evidence] Despite AWS Spot, Google Preemptible, and Azure Spot, true secondary markets for compute capacity have not emerged. No futures market exists for cloud compute. [In- ference] Infrastructure control by providers may prevent full market development—providers prefer pricing power over market efficiency. 10 Telecommunications Bandwidth 10.1 Bandwidth Trading Attempts [Evidence] Multiple attempts to create bandwidth exchanges occurred in the late 1990s (Enron Broadband, RateXchange). Most failed by 2002. [Inference] Markets fail when: (1) quality varies significantly by route, (2) long-term contracts lock in participants, (3) carriers prefer private negotiations, (4) no standardized settlement infrastructure. 10.2 Partial Success in Wholesale [Evidence] Limited bandwidth trading occurs in wholesale voice (international termination) and data transit, but remains bilateral and relationship-based. [Inference] Network infrastruc- ture markets resist full commoditization due to path dependency and relationship economics. 11 Carbon Credits 11.1 Market Formation [Evidence] Carbon markets developed through regulatory mandate (EU ETS, 2005; California Cap-and-Trade, 2013). Standardization required: (1) tonne CO2 equivalent unit, (2) MWh cer- tification, (3) verification by accredited bodies, (4) registry systems. [Inference] Regulatory- created markets can achieve standardization faster than organic markets, but require enforce- ment. 11.2 Price Volatility [Evidence] EU carbon prices ranged from C3 (2013) to C100+ (2023), demonstrating volatility that creates hedging demand. [Inference] Regulatory markets exhibit policy-driven volatility distinct from supply-demand fundamentals.
14 Foreign Exchange 14 12 Advertising Impressions 12.1 Programmatic Advertising [Evidence] Digital advertising developed real-time bidding (RTB) markets where impressions are auctioned in milliseconds. Standardization through: (1) IAB ad unit sizes, (2) viewabil- ity standards (MRC), (3) attribution models, (4) brand safety categories. [Inference] Highly granular, low-value transactions can achieve market efficiency through automation and stan- dardization. 12.2 Limitations [Evidence] Despite programmatic markets, significant opacity remains in: (1) viewability mea- surement, (2) bot traffic, (3) attribution accuracy, (4) brand safety. No independent benchmark price exists for “one impression.” [Inference] Quality verification challenges can prevent full market transparency even with standardized units. 13 Labor Outsourcing 13.1 Business Process Outsourcing [Evidence] BPO markets developed standardized pricing for: (1) call center seats (per agent per hour), (2) claims processing (per claim), (3) data entry (per record), (4) FTE-based outsourcing. [Inference] Labor markets can standardize around output units when tasks are repeatable and measurable. 13.2 Gig Economy Platforms [Evidence] Uber, TaskRabbit, and Upwork created spot markets for labor with: (1) task stan- dardization, (2) rating systems, (3) dynamic pricing, (4) platform-mediated settlement. [In- ference] Digital platforms can create spot labor markets through: (1) task decomposition, (2) reputation systems, (3) standardized settlement. 14 Foreign Exchange [Evidence] FX markets developed spot, forward, and futures markets due to: (1) international trade creating hedging demand, (2) standardized currency pairs, (3) 24-hour trading, (4) elec- tronic platforms (EBS, Reuters). [Inference] Financial markets with continuous price discov- ery and low storage costs achieve the highest standardization.
17 Key Insights for Machine Labor 15 15 Interest Rates [Evidence] LIBOR (London Interbank Offered Rate) emerged as the benchmark for floating- rate instruments from 1969-2023. Its replacement (SOFR, SONIA) followed LIBOR manipu- lation scandals. [Inference] Benchmark manipulation risk increases with: (1) expert judgment in price formation, (2) concentrated market participants, (3) trillions in notional exposure. 16 Comparative Analysis Table 1: Historical Commodity Market Comparison Matrix Resource Unit Quality Grade Benchmark Spot Market Forward Derivatives Settlement Crude Oil Barrel (42 gal) API grav- ity, sulfur WTI, Brent, Dubai Yes (1980s) Yes (fu- tures) Yes (op- tions) NYMEX, ICE Gasoline Gallon Octane, RVP NYH RBOB Yes Yes Yes NYMEX Electricity MWh Voltage, frequency PJM, ER- COT hubs Yes (hourly) Yes (fu- tures) Limited ISOs, Ex- changes Natural Gas MMBtu BTU content, methane Henry Hub Yes (1990s) Yes (fu- tures) Yes (op- tions) NYMEX Grain Bushel Moisture, protein, test wt CBOT corn, wheat Yes Yes (fu- tures) Yes (op- tions) CME Group Freight Voyage/TimeRoute, ves- sel size BDI, BDTI Index Yes (FFAs) Yes (cleared) Baltic, LCH Cloud Com- pute vCPU- hour RAM, storage, network None Yes (spot) No No AWS, GCP, Azure Carbon Credits tCO2e Vintage, project type EU ETS, CCX Yes Yes (fu- tures) Limited ICE, EU Registry Ad Impres- sions Single impres- sion Size, viewability None (CPM) Yes (RTB) No No Ad ex- changes BPO Labor Task/FTE Accuracy, SLA Per-task pricing Limited No No Outsourcing firms FX Currency pair Spot rate EUR/USD, etc. Yes Yes (for- wards) Yes (fu- tures) EBS, Reuters Interest Rates Percentage Tenor, credit risk SOFR, SO- NIA Yes Yes (FRAs) Yes (swaps) CME, LCH 17 Key Insights for Machine Labor 17.1 Conditions Supporting Market Formation [Inference] Markets successfully formed when:
17 Key Insights for Machine Labor 16
- Standardization possible: Quality dimensions can be specified and measured
- Sufficient scale: Volume justifies infrastructure investment
- Price volatility: Variation creates hedging demand
- Independent verification: Third parties can confirm quality and quantity
- Regulatory neutrality: No party can unilaterally set rules
- Network effects: Value increases with participation 17.2 Conditions Preventing Market Formation [Inference] Markets failed or remained limited when:
- Extreme heterogeneity: Each unit is substantially unique
- Provider control: Infrastructure owners prefer managed pricing
- Relationship dependency: Trust and customization dominate
- Measurement impossible: Quality cannot be objectively verified
- Regulatory prohibition: Cross-border or speculative trading banned 17.3 Most Relevant Analogies for Machine Labor [Hypothesis] Based on structural similarity:
- Cloud compute (strongest): Similar production function (compute →output), API- driven consumption, provider concentration, spot market attempts
- BPO labor (strong): Task-based pricing, SLA-governed, quality variation, outsourcing model
- Freight (moderate): Route/task abstraction, heterogeneous outcomes, index-based set- tlement
- Electricity (moderate): Non-storable, instantaneous delivery, capacity requirements
- Crude oil (weaker): Fungible commodity, storage possible, global arbitrage [Evidence] The cloud compute analogy is most relevant for understanding machine labor market potential, but BPO labor provides the best model for task-based pricing standardization.
21 The Production Function of AI Work 17 18 Historical Pattern Summary The historical evidence supports the following patterns:
- Standardization precedes markets: Commodity markets require agreed-upon units and quality grades before price discovery can occur.
- Benchmarks require physical infrastructure: Oil benchmarks need storage and deliv- ery points; electricity benchmarks need nodal pricing; machine labor benchmarks will need verification infrastructure.
- Forward markets emerge from volatility: Without price variation creating hedging demand, forward markets fail to develop.
- Dominant institutions achieve regulatory recognition: Benchmarks become entrenched when regulators, tax authorities, and accounting standards adopt them.
- Network effects protect incumbents: Once contracts reference a benchmark, switching costs become prohibitive.
- Provider concentration can prevent full marketization: AWS controls spot pricing; machine labor providers may similarly resist commoditization. 19 Definition of Machine Labor 20 Introduction Before determining whether machine labor can become a commodity market, we must estab- lish whether it constitutes a coherent economic category. This section distinguishes between intelligence, compute, model access, tokens, software, agents, agent capacity, completed work, verified outcomes, digital machine labor, and physical robotic labor. We examine whether the claim that “companies want verified outcomes, not tokens” is accurate or represents aspirational thinking. 21 The Production Function of AI Work 21.1 Input-Output Relationships [Evidence] The production of economically valuable AI work follows a multi-stage process:
22 Taxonomy of Machine Labor Components 18 Compute+Data+Algorithms+Human Oversight →Intelligence →Agent Actions →Verified Outcomes (1) [Inference] Each stage transforms the prior stage into something more valuable but also more specific: • Compute: Undifferentiated processing capacity (GPU-hours, TPU-hours) • Intelligence: Pattern recognition and generation capability (model weights, reasoning) • Agent Actions: Goal-directed tool use (API calls, code execution, browser automation) • Verified Outcomes: Accepted work products (resolved tickets, merged code, approved claims) 21.2 Current Pricing Models [Evidence] Current AI pricing occurs primarily at the compute and intelligence stages: • Compute: AWS p4d.24xlarge at $32.77/hour, GCP a2-ultragpu-8g at $40.80/hour • Model Access: OpenAI GPT-4 at $0.03/1K input tokens, Anthropic Claude 3 at $0.015/1K input tokens • Software: SaaS subscriptions (per seat, per month) for AI-powered applications [Inference] There is currently no market pricing for verified outcomes. Enterprise AI pro- curement remains at the input layer. 22 Taxonomy of Machine Labor Components 22.1 Intelligence vs. Compute [Evidence] Intelligence and compute are separable economic goods: • Same compute can produce different intelligence (different models on same hardware) • Same intelligence can run on different compute (model distillation, quantization, differ- ent providers) • Compute prices have fallen 10-100x per FLOP over the past decade • Intelligence prices have fallen more slowly (API prices stable or declining modestly) [Inference] The decoupling of intelligence from compute implies separate markets could develop for each, with different price dynamics.
22 Taxonomy of Machine Labor Components 19 22.2 Model Access vs. Tokens [Evidence] Model access can be purchased through:
- API calls: Per-token pricing (OpenAI, Anthropic, Google)
- Hosted inference: Per-hour pricing with rate limits (Together AI, Fireworks)
- Private deployment: License + compute costs (Microsoft Azure OpenAI, AWS Bedrock)
- Self-hosted: Model weights + own infrastructure (Llama, Mistral) [Inference] Token pricing is one of at least four pricing models, each with different cost structures and control characteristics. 22.3 Agents vs. Software [Evidence] An AI agent is distinguished from traditional software by: • Autonomy: Ability to pursue goals without step-by-step human instruction • Tool use: Capability to invoke external systems (APIs, code execution, browsers) • Adaptation: Learning from feedback to improve performance • Non-determinism: Same input may produce different outputs [Inference] Traditional software pricing (perpetual license, subscription) assumes deter- ministic functionality. Agent pricing must account for probabilistic outcomes. 22.4 Agent Capacity vs. Completed Work [Evidence] Agent capacity refers to: • Concurrent task execution capability • Queue processing rate • Maximum throughput (tasks per hour) • Availability (uptime, maintenance windows) [Evidence] Completed work refers to: • Tasks successfully finished • Outcomes meeting quality thresholds • Deliverables accepted by buyers
23 Testing the Outcome-Based Pricing Claim 20 • Value created for end users [Inference] Capacity pricing (like cloud compute) differs fundamentally from outcome pricing (like BPO). The gap between capacity and accepted outcomes includes failures, retries, quality variations, and human intervention. 22.5 Digital Machine Labor vs. Physical Robotic Labor [Evidence] Digital machine labor: • Operates in software environments • Marginal cost near zero for replication • Instantaneous delivery • No physical constraints • Examples: customer support agents, code generation, document review [Evidence] Physical robotic labor: • Operates in physical environments • Hardware costs scale with deployment • Delivery constrained by manufacturing and logistics • Subject to physical constraints (power, maintenance, safety) • Examples: warehouse robots, autonomous vehicles, manufacturing arms [Inference] Digital and physical machine labor may develop separate market structures due to fundamentally different cost structures and scaling properties. 23 Testing the Outcome-Based Pricing Claim 23.1 The Claim [Hypothesis] The central claim tested: “Companies do not ultimately want tokens or agent- hours. They want a defined quantity of successful work delivered at a known quality, risk, jurisdiction, deadline, and price.”
23 Testing the Outcome-Based Pricing Claim 21 23.2 Where the Claim Holds [Evidence] The claim holds for:
- Customer Support: Companies track cost per resolved ticket, not cost per AI response. A resolution is verifiable (ticket closed, no reopen within N days).
- Claims Processing: Insurance companies track cost per processed claim with defined accuracy thresholds. An outcome is verifiable (decision accuracy measured against audit samples).
- Data Entry: Companies track cost per accurate record. Outcomes are verifiable (field- by-field comparison to source).
- Document Review: Law firms track cost per document reviewed with recall/precision metrics. Outcomes are verifiable (sampled manual review). [Inference] Outcome-based pricing is viable when: • Success criteria are definable ex-ante • Verification is cheaper than production • Quality is measurable post-delivery • Retry is possible without catastrophic cost 23.3 Where the Claim Breaks Down [Evidence] The claim breaks down for:
- Creative Work: Marketing copy, design, strategic analysis. Success is subjective and contextual.
- Research: Open-ended investigation where value emerges over time and cannot be spec- ified in advance.
- Exploratory Coding: Prototyping and experimentation where the goal changes during the process.
- Advisory Services: Recommendations where the value depends on implementation by the buyer.
- Emergent Tasks: Work that cannot be fully specified because the situation is novel. [Inference] Outcome-based pricing fails when:
24 Is Machine Labor a Coherent Economic Category? 22 • Success cannot be defined objectively • Value is contingent on future events • The task itself must be discovered • Human judgment is the primary value-add 23.4 The Hybrid Reality [Hypothesis] Most enterprise AI use will involve hybrid pricing: • Subscription base: Guaranteed capacity and availability • Usage component: Per-task or per-outcome pricing for variable volume • Quality adjustment: Bonuses/penalties based on measured outcomes • Outcome guarantee: SLAs with refunds for failure to meet thresholds [Evidence] Evidence from BPO markets supports hybrid models. Pure outcome-based pricing is rare; most contracts include fixed components plus variable outcome adjustments. 24 Is Machine Labor a Coherent Economic Category? 24.1 Criteria for Economic Coherence An economic category is coherent if:
- Items within the category are more substitutable with each other than with items outside
- Buyers make decisions comparing items within the category
- Prices within the category correlate more strongly than with external goods
- Market participants recognize the category as a domain of competition 24.2 Evidence of Coherence [Evidence] Machine labor demonstrates coherence:
- Substitutability: Claude 3.5 Sonnet, GPT-4, and Llama 3.1 are evaluated against each other for the same tasks.
- Buyer Comparison: Enterprise RFPs for AI solutions regularly evaluate multiple providers on standardized task benchmarks.
25 The Value Chain Analysis 23 3. Price Correlation: Model API prices have converged toward similar price-performance ratios, suggesting category-level competition. 4. Market Recognition: Analyst firms (Gartner, Forrester) track “AI agents” as a category distinct from “AI platforms” or “cloud AI services.” 24.3 Boundaries of the Category [Hypothesis] Machine labor boundaries: Included: • AI agents performing economically productive tasks • Autonomous systems with goal-directed behavior • Tool-using AI systems • Both digital and physical embodiments Excluded: • Pure content generation without goal structure • Recommendation systems without action capability • Traditional automation (RPA without adaptation) • Embedded AI without independent agency 25 The Value Chain Analysis 25.1 Upstream: Foundation Models [Evidence] Upstream providers (OpenAI, Anthropic, Google, Meta, Cohere, Mistral) produce general intelligence. Pricing is per-token or per-inference. [Inference] Upstream providers extract value through control of model weights and API access. They may resist commoditization. 25.2 Midstream: Agent Platforms [Evidence] Midstream providers (LangChain, AutoGPT, CrewAI, Microsoft Copilot) orches- trate models into agent systems. Pricing is typically per-seat SaaS. [Inference] Midstream providers compete on orchestration capabilities, not unit economics of work performed.
26 Implications for Market Formation 24 25.3 Downstream: Outcome Delivery [Evidence] Downstream providers (AI BPOs, agent startups) promise specific outcomes. Pric- ing is emerging as per-outcome or outcome-guaranteed. [Inference] Downstream providers bear the risk of converting tokens into accepted work. This is where machine labor markets could emerge. 26 Implications for Market Formation 26.1 Standardization Requirements [Inference] For machine labor markets to form, standardization must occur at the outcome layer, not just the token layer. This requires:
- Task taxonomies with agreed-upon difficulty classifications
- Quality metrics with intersubjective validity
- Verification protocols that are cheaper than production
- Settlement mechanisms for disputed outcomes 26.2 Market Layer Position [Hypothesis] The most likely locus for machine labor markets is between: • Buyers: Enterprises needing outcomes • Sellers: Agent providers guaranteeing outcomes Upstream (model providers) and midstream (agent platforms) will resist commoditization. Downstream (outcome providers) will drive market formation to prove competitiveness. 26.3 Preconditions for Commoditization [Inference] Machine labor commoditization requires:
- Model commoditization: Open weights and multiple providers reduce differentiation
- Orchestration standardization: Common agent frameworks enable provider switching
- Outcome measurability: Buyers can verify what they purchased
- Sufficient liquidity: Enough providers and buyers for price discovery [Evidence] Current trends: (1) open weights proliferating, (2) LangChain/CrewAI gaining adoption, (3) evaluation benchmarks emerging, (4) enterprise adoption growing. These support conditional commoditization.
30 The Standardization Challenge 25 27 Chapter Summary Machine labor is a coherent economic category distinct from compute, intelligence, tokens, and software. The claim that buyers want verified outcomes rather than inputs holds for structured, repeatable tasks but breaks down for creative, exploratory, and advisory work. Most enterprise AI use will involve hybrid pricing combining subscriptions, usage, and outcome guarantees. The most promising locus for market formation is the downstream layer where agent providers convert model outputs into verified outcomes. This requires standardization of tasks, quality metrics, and verification protocols that do not currently exist but are emerging. 28 Standardized Machine Work Unit 29 Introduction The central unsolved problem in machine labor market formation is standardization. Oil has barrels. Electricity has megawatt-hours. Compute has GPU-hours. Human labor has hours, salaries, and task contracts. Agent work varies dramatically in quality, difficulty, context, and reliability. This section develops a rigorous framework for defining a Standardized Machine Work Unit (SMWU) and evaluates eight candidate markets. 30 The Standardization Challenge 30.1 Sources of Variation [Evidence] Agent work varies across multiple dimensions:
- Task Complexity: Simple classification vs. multi-step reasoning
- Context Requirements: Stateless vs. requiring extensive history
- Tool Dependencies: None vs. multiple external systems
- Quality Thresholds: "Good enough" vs. "human-equivalent"
- Latency Constraints: Batch acceptable vs. real-time required
- Error Consequences: Cosmetic vs. financial/legal liability
- Domain Specificity: General capability vs. specialized knowledge [Inference] Standardization requires abstracting across these variations while preserving the information necessary for price discovery.
31 Standardized Machine Work Unit Framework 26 30.2 Lessons from Historical Standardization [Evidence] Successful commodity standardization has relied on:
- Grade bands rather than point specifications: No. 2 corn allows 15.5% moisture, not 15.5% exactly
- Reference methods: API gravity measured by specific standardized procedures
- Delivery location specification: Cushing for WTI, Henry Hub for natural gas
- Third-party verification: Independent inspectors certify quality [Inference] Machine labor standardization should use grade bands, reference tasks, and third-party verification rather than attempting precise point specifications. 31 Standardized Machine Work Unit Framework 31.1 Required Specifications A Standardized Machine Work Unit (SMWU) must specify: 31.1.1 Task Category The functional domain of the work: • Customer support • Software engineering • Legal review • Insurance claims processing • Financial operations • Cybersecurity • Research • Robotics
31 Standardized Machine Work Unit Framework 27 31.1.2 Task Difficulty A graded scale based on: • Number of reasoning steps required • Domain knowledge required • Context window requirements • Tool use complexity • Ambiguity tolerance needed Proposed scale: Level 1 (simple) to Level 5 (complex), with reference tasks defining each level. 31.1.3 Expected Output The deliverable specification: • Format (structured data, natural language, code, decision) • Completeness requirements • Dependencies on external systems 31.1.4 Success Criteria The conditions for acceptance: • Accuracy threshold (e.g., 95% precision) • Coverage requirement (e.g., 98% recall) • Completeness check • Constraint satisfaction 31.1.5 Quality Threshold The level of quality required: • Basic (minimally acceptable) • Standard (industry average) • Premium (top quartile) • Certified (validated by human expert)
31 Standardized Machine Work Unit Framework 28 31.1.6 Latency Requirement Maximum acceptable time: • Real-time (<1 second) • Interactive (<10 seconds) • Batch acceptable (hours) • No constraint 31.1.7 Maximum Human Intervention The allowable human involvement: • Fully autonomous (no human involvement) • Exception handling only • Quality assurance sampling • Continuous human-in-the-loop 31.1.8 Acceptable Retry Rate The maximum allowed failures before human escalation: • Single attempt (no retry) • Limited retries (2-3 attempts) • Unlimited within SLA 31.1.9 Security Requirements Data and operational security: • Encryption standards (at rest, in transit) • Access controls • Audit logging requirements • Compliance frameworks (SOC 2, ISO 27001, FedRAMP)
31 Standardized Machine Work Unit Framework 29 31.1.10 Data Residency Requirements Geographic constraints: • No constraints • Regional (EU, US, Asia) • Country-specific • Air-gapped (no external connectivity) 31.1.11 Model Provenance Requirements Constraints on AI systems: • No constraints (any model) • Open-weight models only • Specific model families • Auditable training data • No training on customer data 31.1.12 Auditability Documentation requirements: • Reasoning chain logging • Decision traceability • Tool use records • Human intervention logs 31.1.13 Liability Responsibility assignment: • Provider assumes liability • Shared liability with caps • Buyer assumes liability • Insurance requirement
31 Standardized Machine Work Unit Framework 30 31.1.14 Delivery Period Time window for completion: • Immediate • Standard (within hours) • Scheduled (specific time) • Batch (daily/weekly) 31.1.15 Verification Process How outcomes are validated: • Automated testing • Human spot-checking • Full human review • Third-party audit 31.2 SMWU Specification Template Table 2: Standardized Machine Work Unit Specification Template Parameter Specification Task Category [Category] Task Difficulty Level [1-5]: [Reference Task ID] Expected Output [Format and requirements] Success Criteria Accuracy ≥[X%], Coverage ≥[Y%] Quality Threshold [Basic/Standard/Premium/Certified] Latency Requirement [Real-time/Interactive/Batch/None] Max Human Intervention [Autonomous/Exception/QA/Continuous] Acceptable Retry Rate [Single/Limited/Unlimited] Security Requirements [Compliance framework] Data Residency [Constraint specification] Model Provenance [Open/Closed/Specific] Auditability [Logging requirements] Liability [Responsibility assignment] Delivery Period [Timing specification] Verification Process [Validation method]
32 Market Evaluation 31 32 Market Evaluation 32.1 Evaluation Criteria Markets are evaluated on:
- Repeatability: How consistently can the task be performed?
- Verifiability: How objectively can success be measured?
- Volume: How much economic activity exists?
- AI Adoption: How much AI use already occurs?
- Provider Competition: How many competing solutions exist?
- Price Volatility: How much does cost vary?
- Standardization Potential: How feasible is unit definition? 32.2 Customer Support [Evidence] Current state: • Global BPO market: $250B+ annually • AI customer support market: $5B+ and growing 30%+ annually • Providers: Intercom, Zendesk AI, Forethought, Ada, Ultimate.ai • Typical pricing: $0.50-$5.00 per resolution SMWU Definition: One customer problem successfully resolved without human escalation and without reopening during a specified period (typically 7-14 days). Standardization Assessment: • Repeatability: High (same queries recur) • Verifiability: High (ticket status, customer feedback) • Volume: Very high • AI Adoption: High and growing • Provider Competition: Moderate to high • Price Volatility: Moderate • Standardization Potential: High Conclusion: Strong candidate for first market.
32 Market Evaluation 32 32.3 Software Engineering [Evidence] Current state: • Global software development: $600B+ annually • AI coding assistants: GitHub Copilot, Cursor, Amazon CodeWhisperer • Emerging: Devin, OpenAI Codex, autonomous coding agents SMWU Definition: One defined engineering task completed, passing required tests and accepted through code review. Standardization Assessment: • Repeatability: Moderate (high variation in tasks) • Verifiability: High (tests, review) • Volume: Very high • AI Adoption: High (assistants), Low (autonomous) • Provider Competition: Moderate • Price Volatility: High (rapid capability change) • Standardization Potential: Moderate Conclusion: Promising but requires more maturity in autonomous agents. 32.4 Legal Review [Evidence] Current state: • Legal services market: $800B+ globally • Contract review: $50B+ segment • AI providers: Harvey, CoCounsel, Lexion, Ironclad SMWU Definition: One contract reviewed to an agreed accuracy, coverage, and citation standard. Standardization Assessment: • Repeatability: Moderate (contract types vary) • Verifiability: Moderate (expert-dependent) • Volume: High
32 Market Evaluation 33 • AI Adoption: Moderate • Provider Competition: Moderate • Price Volatility: Moderate • Standardization Potential: Moderate Conclusion: Strong but regulatory complexity adds friction. 32.5 Insurance Claims Processing [Evidence] Current state: • Claims processing costs: $50B+ annually in U.S. • AI adoption: Lemonade, Shift Technology, FRISS SMWU Definition: One claim processed according to defined accuracy, compliance, and fraud-detection requirements. Standardization Assessment: • Repeatability: High (structured forms) • Verifiability: High (audit trails) • Volume: Very high • AI Adoption: Moderate • Provider Competition: Moderate • Price Volatility: Low to moderate • Standardization Potential: High Conclusion: Strong candidate, especially for initial narrow domains. 32.6 Financial Operations [Evidence] Current state: • Accounts payable processing: $40B+ market • Reconciliation services: $20B+ market • AI providers: Bill.com, Stampli, AppZen
32 Market Evaluation 34 SMWU Definition: One invoice, transaction, or account reconciled without unresolved discrepancies. Standardization Assessment: • Repeatability: High (structured data) • Verifiability: High (matching algorithms) • Volume: Very high • AI Adoption: Moderate to high • Provider Competition: Moderate • Price Volatility: Low • Standardization Potential: High Conclusion: Strong candidate, especially for data-heavy processes. 32.7 Cybersecurity [Evidence] Current state: • Security operations market: $50B+ • AI providers: Darktrace, Vectra AI, SentinelOne SMWU Definition: One alert investigated and correctly classified under a defined service- level agreement. Standardization Assessment: • Repeatability: Moderate (high variation in threats) • Verifiability: Moderate (false positive/negative trade-offs) • Volume: High • AI Adoption: Moderate • Provider Competition: High • Price Volatility: Moderate • Standardization Potential: Moderate Conclusion: Challenging due to adversarial nature of domain.
32 Market Evaluation 35 32.8 Research [Evidence] Current state: • RD spending: $2.5T globally • AI research assistants: Emerging (Elicit, Consensus, Perplexity) SMWU Definition: One evidence-backed research task completed with validated citations and reproducible analysis. Standardization Assessment: • Repeatability: Low (open-ended tasks) • Verifiability: Low (subjective quality) • Volume: High • AI Adoption: Low • Provider Competition: Low • Price Volatility: High • Standardization Potential: Low Conclusion: Not suitable for initial market; too heterogeneous. 32.9 Robotics [Evidence] Current state: • Industrial robotics: $50B+ market • Service robotics: Growing rapidly • AI integration: Figure AI, Tesla Optimus, Boston Dynamics SMWU Definition: One physical task completed safely and successfully in a specified environment. Standardization Assessment: • Repeatability: Low to moderate (environment-dependent) • Verifiability: High (success/failure observable) • Volume: Moderate • AI Adoption: Low (early stage)
33 First Market Selection 36 • Provider Competition: Low • Price Volatility: High • Standardization Potential: Low to moderate Conclusion: Long-term opportunity; physical constraints add complexity. 33 First Market Selection 33.1 Scoring Summary Table 3: Market Standardization Potential Scoring Market Rep Ver Vol Adopt Comp Volat Std Total Customer Support 5 5 5 5 4 3 5 32 Insurance Claims 5 5 5 3 3 2 5 28 Financial Ops 5 5 5 4 3 2 5 29 Legal Review 3 3 4 3 3 3 4 23 Software Eng 3 5 5 3 3 4 4 27 Cybersecurity 3 3 4 3 4 3 3 23 Research 1 2 4 1 2 4 2 16 Robotics 2 4 3 1 1 4 3 18 Scale: 1 (low) to 5 (high) 33.2 Recommendation [Inference] Based on the scoring and qualitative assessment: First Market: Customer Support • Highest standardization potential • Large existing volume • High AI adoption • Clear success criteria • Established verification mechanisms Second Market: Financial Operations (invoices, reconciliation) • Highly structured data • Objective verification
37 The True Cost Formula 37 • Lower regulatory complexity than insurance Third Market: Insurance Claims Processing • High volume • Strong verification requirements • Regulatory complexity manageable with focus 34 Chapter Summary The Standardized Machine Work Unit (SMWU) framework specifies fifteen parameters nec- essary for defining tradeable units of AI-completed work. Customer support emerges as the strongest initial market due to high repeatability, verifiability, and existing AI adoption. Finan- cial operations and insurance claims represent strong follow-on opportunities. Software engi- neering, legal review, and cybersecurity require more maturity. Research and robotics present longer-term opportunities. Part II Pricing and Valuation Frameworks 35 Pricing Models 36 Introduction Current AI pricing focuses on inputs (tokens, compute hours) rather than outputs (verified work). This creates a significant information gap: buyers cannot predict the true cost of achiev- ing their desired outcomes. This section develops a rigorous formula for the true cost of a verified AI outcome and creates both quality-adjusted and risk-adjusted pricing methodologies. 37 The True Cost Formula 37.1 Conceptual Foundation The true cost of a verified AI outcome includes all costs incurred to produce accepted work, divided by the number of accepted outcomes. This captures the reality that not all attempts succeed, and that failures have costs.
37 The True Cost Formula 38 37.2 Base Formula TCVO = Cmodel + Ccompute + Ctools + Cinfra + Cretries + Chuman + Cfailure + Ccompliance Naccepted (2) Where: • Cmodel = Model API costs • Ccompute = Infrastructure compute costs • Ctools = External tool/API costs • Cinfra = Software infrastructure costs • Cretries = Costs of retry attempts • Chuman = Human review and escalation costs • Cfailure = Expected cost of failures • Ccompliance = Security and compliance costs • Naccepted = Number of accepted outcomes 37.3 Detailed Cost Components 37.3.1 Model API Cost (Cmodel) Cmodel = Nattempts X i=1 (Tinput,i × Pinput + Toutput,i × Poutput) (3) Where: • Tinput,i = Input tokens for attempt i • Toutput,i = Output tokens for attempt i • Pinput = Price per input token • Poutput = Price per output token [Evidence] Current pricing (as of mid-2026): • GPT-4o: $0.005/1K input, $0.015/1K output • Claude 3.5 Sonnet: $0.003/1K input, $0.015/1K output • Llama 3.1 405B (hosted): $0.002/1K input, $0.002/1K output
37 The True Cost Formula 39 37.3.2 Compute Cost (Ccompute) For self-hosted or private deployments: Ccompute = Nattempts X i=1 (Hi × Rgpu) (4) Where: • Hi = GPU-hours consumed for attempt i • Rgpu = Cost per GPU-hour [Evidence] Current GPU pricing: • AWS p4d.24xlarge (8x A100): $32.77/hour • GCP a2-ultragpu-8g (8x A100): $40.80/hour • Lambda Cloud 1x A100: $1.99/hour 37.3.3 Tool Use Cost (Ctools) Ctools = Nattempts X i=1 Mi X j=1 Ctool,i,j (5) Where: • Mi = Number of tool calls for attempt i • Ctool,i,j = Cost of tool call j for attempt i [Evidence] Typical tool costs: • Web search API: $0.005-$0.02 per query • Code execution: $0.001-$0.01 per execution • Database query: $0.001-$0.005 per query • External API calls: Variable by service 37.3.4 Infrastructure Cost (Cinfra) Cinfra = Corchestration + Cmonitoring + Clogging + Cstorage (6) [Evidence] Infrastructure cost allocations: • Orchestration platform: $0.001-$0.01 per task
37 The True Cost Formula 40 • Monitoring/Observability: $0.001-$0.005 per task • Logging and audit: $0.0005-$0.002 per task • Storage (context, history): $0.0001-$0.001 per task 37.3.5 Retry Cost (Cretries) Cretries = Nretries X r=1 (Cmodel,r + Ccompute,r + Ctools,r) (7) [Inference] Retry costs can equal or exceed initial attempt costs, especially for complex tasks requiring multiple tool calls. 37.3.6 Human Intervention Cost (Chuman) Chuman = (Hreview × Rreviewer) + (Hescalation × Rexpert) (8) Where: • Hreview = Hours of quality review • Rreviewer = Hourly cost of reviewer • Hescalation = Hours of expert escalation • Rexpert = Hourly cost of expert [Evidence] Typical labor costs: • Quality reviewer: $15-$35/hour • Subject matter expert: $50-$250/hour • Legal/compliance expert: $200-$800/hour 37.3.7 Failure Cost (Cfailure) Cfailure = Nfailed × (Cattempt + Crecovery + Copportunity) (9) Where: • Nfailed = Number of failed attempts • Cattempt = Cost of failed attempt • Crecovery = Cost of recovery/rollback • Copportunity = Opportunity cost of delay
38 Quality-Adjusted Machine Labor Price (QAMLP) 41 37.3.8 Compliance Cost (Ccompliance) Ccompliance = Caudit + Ccertification + Csecurity + C [Evidence] Compliance cost factors: • SOC 2 audit: $50K-$200K annually • ISO 27001: $30K-$100K annually • FedRAMP: $300K-$1M+ annually • Data residency (per-region): 10-30% premium 37.4 Complete Expanded Formula TCVO = 1 Naccepted " Nattempts X i=1 Tinput,iPinput + Toutput,iPoutput + HiRgpu + Mi X j=1 Ctool,i,j +Cinfra + Chuman + Cfailure + Ccompliance
(11) 38 Quality-Adjusted Machine Labor Price (QAMLP) 38.1 Concept Not all outcomes are equal in quality. The Quality-Adjusted Machine Labor Price normalizes costs to a standard quality level, enabling comparison across providers with different accuracy rates. 38.2 Formula QAMLP = TCVO Qactual/Qstandard (12) Where: • Qactual = Actual quality achieved (e.g., accuracy rate) • Qstandard = Standard quality threshold 38.3 Example Calculation Provider A: TCVO = $2.00, Accuracy = 85% Provider B: TCVO = $2.50, Accuracy = 95% Standard Quality = 90%
39 Risk-Adjusted Machine Labor Price (RAMLP) 42 QAMLP_A = $2.00 / (0.85/0.90) = $2.12 QAMLP_B = $2.50 / (0.95/0.90) = $2.37 At standard quality, Provider A is cheaper despite lower nominal accuracy. 38.4 Multi-Dimensional Quality Adjustment For quality with multiple dimensions: QAMLP = TCVO × K Y k=1 Qstandard,k Qactual,k wk (13) Where wk = weight for quality dimension k, with P wk = 1. 39 Risk-Adjusted Machine Labor Price (RAMLP) 39.1 Concept Different deployments carry different risks: model provider concentration, jurisdiction con- straints, latency variability, and failure modes. The Risk-Adjusted Machine Labor Price incor- porates these risks into a comparable metric. 39.2 Risk Factor Framework Define risk factors R1, R2, ..., Rn with scores 0 (no risk) to 1 (maximum risk). 39.2.1 Provider Concentration Risk Rconcentration = 1 − 1 Nproviders (14) Where Nproviders = number of independent model providers available. 39.2.2 Sovereign Risk Rsovereign = 0 if deployment fully within buyer jurisdiction 0.3 if cross-border with data residency controls 0.6 if cross-border without residency controls 1.0 if deployment in sanctioned jurisdiction (15)
40 Benchmark Methodology 43 39.2.3 Latency Risk Rlatency = σlatency µlatency (16) Where σlatency = standard deviation of latency, µlatency = mean latency. 39.2.4 Failure Mode Risk Rfailure = Pcatastrophic × Icatastrophic (17) Where: • Pcatastrophic = Probability of catastrophic failure • Icatastrophic = Impact of catastrophic failure (normalized 0-1) 39.2.5 Model Version Risk Rversion = 1 −stabilitymodel (18) Where stabilitymodel measures consistency across model updates. 39.3 Risk-Adjusted Price Formula RAMLP = TCVO × (1 + n X i=1 wiRi) (19) Where wi = weight for risk factor i, with P wi = 1. 39.4 Risk-Adjusted Price with Insurance If failure insurance is available: RAMLPinsured = TCVO + Pinsurance −E[payout] (20) Where: • Pinsurance = Insurance premium • E[payout] = Expected insurance payout 40 Benchmark Methodology 40.1 Price Index Construction A machine labor price index requires:
40 Benchmark Methodology 44
- Representative sample of transactions
- Consistent unit definition
- Quality adjustment
- Temporal aggregation
- Publication frequency 40.2 Aggregation Methods 40.2.1 Volume-Weighted Average Price (VWAP) VWAP = PN i=1 Pi × Vi PN i=1 Vi (21) Best for: Markets with high transaction volume and price variation. 40.2.2 Trimmed Mean Trimmed Mean = 1 N −2k N−k X i=k+1 P(i) (22) Where P(i) = i-th ordered price, k = number of observations trimmed from each end. Best for: Markets with outlier risk or manipulation potential. 40.2.3 Median Price Median = P((N+1)/2) for odd N (23) Best for: Markets with highly skewed distributions. 40.3 Recommendation [Inference] For machine labor markets, a trimmed volume-weighted average is recommended:
- Exclude top and bottom 10% by price (outlier control)
- Volume-weight the remaining observations
- Apply quality adjustment before aggregation
41 Manipulation Resistance 45 41 Manipulation Resistance 41.1 Manipulation Vectors [Evidence] Benchmarks can be manipulated through:
- Wash trading: Artificial transactions to influence price
- Cherry-picking: Selective reporting of favorable transactions
- Reference rate gaming: Submitting false data to survey-based benchmarks
- Concentration: Single large player dominating the benchmark 41.2 Resistance Mechanisms 41.2.1 Transaction Verification Require cryptographic proof of work completion: • Task hashes recorded on immutable ledger • Outcome verification by independent parties • Cross-reference with buyer systems 41.2.2 Multi-Source Data Aggregate from multiple independent sources: • Direct buyer reporting • Provider reporting • Platform transaction records • Third-party verification 41.2.3 Outlier Detection Statistical filters for anomalous transactions: • Price outside 3 standard deviations from rolling mean • Volume inconsistent with historical patterns • Timing patterns suggesting wash trading
44 Introduction 46 41.2.4 Transparency Requirements • Publish methodology publicly • Allow external audit of calculation • Disclose data sources (anonymized) • Provide historical data for verification 41.3 Governance [Inference] Benchmark governance should include:
- Independent oversight committee
- Market participant representation
- Regulatory consultation
- Methodology review process 42 Chapter Summary The True Cost per Verified Outcome (TCVO) formula provides a comprehensive framework for calculating the actual cost of AI-completed work, incorporating model, compute, tool, in- frastructure, retry, human, failure, and compliance costs. Quality-adjusted and risk-adjusted pricing enable meaningful comparison across providers. A trimmed volume-weighted average benchmark methodology, combined with transaction verification and multi-source data aggre- gation, provides manipulation resistance. These pricing frameworks are essential preconditions for machine labor market formation. 43 Jurisdiction and Sovereignty 44 Introduction AI systems do not exist in a vacuum. They are developed by organizations subject to national jurisdictions, trained on data governed by privacy regulations, deployed on infrastructure con- trolled by cloud providers, and used by enterprises with compliance obligations. This section investigates how geopolitical fragmentation affects machine labor pricing and develops the concept of a “sovereignty premium.”
46 Evaluation Framework 47 45 The Geopolitical AI Landscape 45.1 Three Ecosystems [Evidence] The global AI market has fragmented into three major ecosystems:
- United States: OpenAI, Anthropic, Google, Meta, Microsoft, xAI, Cohere, Adept
- China: Baidu (Ernie), Alibaba (Tongyi), ByteDance, Tencent, 01.AI, Moonshot AI
- Europe: Mistral AI, Aleph Alpha, Huma (UK), DeepMind (UK/US hybrid) [Inference] These ecosystems operate under different regulatory regimes, data governance requirements, and strategic priorities. 45.2 Open-Weight Models [Evidence] Open-weight models complicate the landscape: • Meta (Llama): U.S. company, open weights • Alibaba (Qwen): Chinese company, open weights • Mistral (Mixtral): European company, open weights • 01.AI (Yi): Chinese company, open weights [Inference] Open weights enable deployment across jurisdictions, but model origin may still matter for compliance and trust. 46 Evaluation Framework 46.1 Dimensions of Jurisdictional Analysis For each AI system, we evaluate:
- Model Developer Jurisdiction: Where the developing organization is headquartered
- Ownership and Control: Who controls model development and updates
- Model Weights Availability: Open (downloadable) vs. Closed (API only)
- Inference Location: Where computation occurs
- Compute Provider Jurisdiction: Where infrastructure is located
47 Ecosystem Analysis 48 6. Data Storage Location: Where input/output data resides 7. Data Transmission: Cross-border data flows 8. Data Retention: How long data is stored 9. Training on Customer Data: Whether customer data improves models 10. Licensing: Terms of use and redistribution 11. Export Controls: Whether model access is restricted 12. Sanctions: Whether developer or provider is sanctioned 13. Cybersecurity Requirements: Compliance frameworks (FedRAMP, ISO 27001, etc.) 14. Government Procurement Restrictions: Whether model is eligible for government contracts 15. Intellectual Property Exposure: Risk of training data IP claims 16. Auditability: Ability to inspect system behavior 17. Open-Weight vs. Closed API: Deployment flexibility 18. Private Deployment: Ability to run without external connectivity 19. Air-Gapped Deployment: Ability to run without any network connection 20. Supply Chain Dependencies: Reliance on foreign components 47 Ecosystem Analysis 47.1 United States Ecosystem [Evidence] Characteristics: • Model Developer Jurisdiction: U.S. entities subject to U.S. law • Ownership: Private companies and public corporations • Weights: Mix of open (Meta, Mistral partnership) and closed (OpenAI, Anthropic) • Inference: Global cloud infrastructure (AWS, Azure, GCP) • Data Storage: Typically U.S. or customer-selected regions • Export Controls: Subject to U.S. export control regulations (EAR)
48 The Sovereignty Premium 49 • Cybersecurity: FedRAMP, StateRAMP, SOC 2, ISO 27001 available • Government Procurement: Available for most federal contracts 47.2 China Ecosystem [Evidence] Characteristics: • Model Developer Jurisdiction: Chinese entities subject to Chinese law (CSRC, CAC) • Ownership: Mix of state-influenced and private companies • Weights: Increasingly open (Qwen, Yi) but with Chinese characteristics • Inference: Primarily China-based, expanding to Southeast Asia • Data Storage: Subject to PIPL (Personal Information Protection Law) • Export Controls: Subject to Chinese export control law • Sanctions: Some entities on U.S. entity list • Government Procurement: Required for Chinese government contracts 47.3 Europe Ecosystem [Evidence] Characteristics: • Model Developer Jurisdiction: EU/UK entities subject to GDPR, AI Act • Ownership: Private companies, research institutions • Weights: Mix of open (Mistral) and closed (Aleph Alpha) • Inference: Primarily EU-based for data residency • Data Storage: Strict GDPR compliance required • Cybersecurity: EUCS (EU Cybersecurity Scheme) emerging • Government Procurement: Preference for EU-developed AI under procurement rules 48 The Sovereignty Premium 48.1 Definition The sovereignty premium is the additional price paid to ensure that AI work satisfies specific national, infrastructure, security, data, or model-origin requirements.
49 Index Grade Framework 50 48.2 Sources of Sovereignty Premium
- Data Residency Requirements: Must store/process data in specific jurisdiction
- Model Origin Preferences: Preference for domestically developed models
- Security Certification: FedRAMP, EUCS, or equivalent compliance
- Air-Gapped Deployment: Complete isolation from external networks
- Audit and Oversight: Government or regulatory inspection rights
- Supply Chain Verification: Domestic hardware and software requirements 48.3 Quantifying the Premium [Hypothesis] Estimated sovereignty premiums: Table 4: Estimated Sovereignty Premiums (as % of base price) Requirement Min Typical Max Data Residency (EU) 10% 20% 40% Data Residency (multi-region) 20% 35% 60% FedRAMP Moderate 30% 50% 100% FedRAMP High 50% 100% 200% Air-gapped deployment 100% 200% 400% Domestic model only 20% 50% 100% [Evidence] Basis: Cloud provider region premiums (10-30%), FedRAMP compliance costs (50-100% infrastructure overhead), air-gapped infrastructure costs (2-4x). 49 Index Grade Framework 49.1 Proposed Index Grades OnIndex should publish differentiated indexes reflecting sovereignty requirements:
- Global Best-Price Machine Labor Index: Lowest cost regardless of jurisdiction
- U.S. Sovereign Machine Labor Index: U.S.-developed models, U.S. infrastructure
- China Domestic Machine Labor Index: China-developed models, China infrastructure
- European Data-Resident Machine Labor Index: GDPR-compliant, EU infrastructure
- Open-Weight Machine Labor Index: Only open-weight models, any infrastructure
50 Key Questions Addressed 51 6. Private-Cloud Machine Labor Index: Deployable on private infrastructure 7. Air-Gapped Machine Labor Index: Fully isolated deployment capability 8. Defense-Eligible Machine Labor Index: FedRAMP High or equivalent 49.2 Classification Examples Table 5: Model Deployment Classification Examples Configuration Model Origin Inference Data Index Grade GPT-4 on Azure US U.S. U.S. U.S. U.S. Sovereign Qwen on AWS US China U.S. U.S. Global Best- Price Llama on self-hosted EU U.S. EU EU EU Data- Resident Claude on air-gapped U.S. Air-gapped Air-gapped Air-Gapped Mistral on private cloud EU Private Private Private-Cloud 50 Key Questions Addressed 50.1 Can Sovereignty Premiums Be Objectively Measured? [Inference] Yes, through:
- Comparing identical workloads across deployment configurations
- Surveying enterprise willingness-to-pay for compliance
- Analyzing contract pricing differentials
- Tracking region-specific cloud provider pricing [Evidence] Cloud providers already charge 10-30% premiums for specific regions. FedRAMP- authorized services cost 50-100% more than commercial equivalents. 50.2 Who Would Pay for This Information? [Hypothesis] Buyers of sovereignty premium data:
- Multinational enterprises: Optimizing global AI deployment
- Government procurement: Comparing domestic vs. foreign options
- Regulated industries: Banks, healthcare, defense contractors
50 Key Questions Addressed 52 4. AI providers: Pricing sovereign-grade offerings 5. Investors: Assessing geopolitical risk in AI portfolios 50.3 Would Governments Use These Indexes? [Hypothesis] Governments would use machine labor indexes to:
- Compare national AI competitiveness
- Assess dependence on foreign AI infrastructure
- Inform industrial policy and subsidies
- Negotiate trade agreements on AI services
- Monitor strategic AI capabilities [Evidence] Precedent: Governments track semiconductor prices, energy costs, and labor costs as strategic indicators. 50.4 Should a Chinese Open-Weight Model on U.S. Infrastructure Be Classified Differently from a China-Hosted API? [Inference] Yes, they differ fundamentally: Table 6: Deployment Configuration Risk Comparison Risk Factor China Model / U.S. Infra China Model / China Infra Data exfiltration risk Low High Censorship influence Low High Supply chain control U.S. controlled China controlled Update control Buyer controlled Provider controlled Sanctions risk Low High Regulatory compliance U.S. law applies Chinese law applies 50.5 Does Model Nationality Matter Less Than the Total Execution and Data Supply Chain? [Inference] The complete supply chain matters more than model nationality alone. A Chi- nese open-weight model running on U.S. infrastructure with U.S. data residency may be more acceptable to U.S. enterprises than a U.S. closed model running on Chinese infrastructure. [Evidence] This mirrors debates about TikTok (Chinese-owned app on U.S. phones) vs. U.S. apps on Chinese infrastructure (AWS China, Azure China).
54 Scenario 1: Retailer Reserving AI Customer Support Capacity 53 51 Chapter Summary Geopolitical fragmentation creates measurable and persistent price premiums for machine la- bor. Sovereignty requirements—data residency, security certification, model origin, air-gapped deployment—add 20% to 400% to base costs. OnIndex should publish differentiated index grades reflecting these variations. Model nationality matters, but the complete execution and data supply chain matters more. Governments and regulated enterprises would pay for inde- pendent sovereign-grade pricing data. Part III Market Analysis and Business Model 52 Market Scenarios 53 Introduction This section models ten realistic market scenarios to demonstrate how machine labor markets could function in practice. Each scenario includes: buyer profile, supplier profile, unit of work, price, quality standard, risk factors, contract structure, verification process, settlement mecha- nism, value created by OnIndex, and transaction difficulty without independent benchmarking. 54 Scenario 1: Retailer Reserving AI Customer Support Ca- pacity 54.1 Transaction Overview • Buyer: National retail chain, 500Mrevenue, 2McustomerinteractionsannuallySupplier : AIcustomersupportplatform(e.g., Ada, Forethought, Ultimate) • Unit: One resolved customer inquiry (no reopen within 7 days) • Volume: 500,000 resolutions during Black Friday week • Price: $1.20/resolution (vs. $8.50 for human agent) 54.2 Quality and Risk • Quality Standard: 90% accuracy, 95% customer satisfaction, <2% escalation rate
55 Scenario 2: Bank Comparing Models for Compliance Review 54 • Risk Factors: Capacity shortage during peak, quality degradation under load, model downtime • Contract: Capacity reservation 30 days in advance, 20% deposit, SLA with penalties 54.3 OnIndex Value • Benchmark Price: $1.15/resolution (OnIndex 30-day average) • Verification: Post-interaction sampling with human reviewers • Settlement: Weekly invoicing, net 30, disputes resolved through OnIndex arbitration • Value Created: Buyer confirms fair pricing; supplier demonstrates market competitive- ness 54.4 Difficulty Without OnIndex Without independent benchmarking: Buyer cannot verify $1.20 is market rate; supplier cannot prove competitiveness; capacity reservation pricing is arbitrary; SLA enforcement is subjective. 55 Scenario 2: Bank Comparing Models for Compliance Re- view 55.1 Transaction Overview • Buyer: Regional bank with 50,000 compliance documents annually • Suppliers: U.S. model (Claude), Chinese model (Qwen), European model (Mistral), Internal model • Unit: One compliance document reviewed with 95% recall, 90% precision • Volume: 50,000 documents 55.2 Quality and Risk • Quality Standard: Regulatory compliance accuracy, audit trail completeness • Risk Factors: Regulatory non-compliance, data residency violations, model bias • Contract: Pilot program with 1,000 documents per model, then 6-month commitment
56 Scenario 3: Software Company Budgeting Human vs. Machine Labor 55 55.3 OnIndex Value • Sovereignty-Adjusted Prices: – U.S. Sovereign: $4.50/document – EU Data-Resident: $5.20/document – Open-Weight (self-hosted): $3.80/document – China-hosted: Not eligible for bank compliance • Verification: Document-by-document accuracy testing against gold standard • Risk Adjustment: Regulatory compliance premium included 56 Scenario 3: Software Company Budgeting Human vs. Ma- chine Labor 56.1 Transaction Overview • Buyer: SaaS company with 50 engineers, planning budget allocation • Task: Bug fixes, feature implementation, code review, documentation • Unit: One engineering task completed and merged • Volume: 2,000 tasks annually 56.2 Cost Comparison Table 7: Human vs. Machine Labor Cost Comparison Cost Component Human Engineer AI Agent Notes Base Cost $150/hr $0.05/task Token + compute Success Rate 95% 70% First attempt Review Cost $0 $30/task Human review Rework Cost $20/task $15/task Failed attempts TCVO $158/task $77/task Normalized 56.3 OnIndex Value Provides true cost per verified outcome including hidden costs (retries, review, failures). En- ables rational budget allocation between human and machine labor.
58 Scenario 5: Agent Startup Using Independent Certification 56 57 Scenario 4: Government Measuring Sovereign AI Costs 57.1 Transaction Overview • Buyer: National government procurement agency • Objective: Compare cost of sovereign AI requirements • Task: Standardized document analysis task basket 57.2 Sovereignty Cost Analysis Table 8: Sovereignty Premium Analysis Configuration Base Cost Sovereignty Cost Premium Global Best-Price $1.00 $1.00 0% U.S. Sovereign $1.00 $1.35 35% EU Data-Resident $1.00 $1.45 45% Defense-Eligible $1.00 $2.80 180% 57.3 OnIndex Value Enables evidence-based policy on AI sovereignty trade-offs. Informs industrial policy and subsidy decisions. 58 Scenario 5: Agent Startup Using Independent Certifica- tion 58.1 Transaction Overview • Buyer: Enterprise procurement evaluating AI vendors • Supplier: Agent startup claiming superior economics • Unit: Customer support resolution • Claim: 40% lower cost than industry average 58.2 Certification Process
- OnIndex tests supplier on standardized task basket
- Measures: attempts, completions, acceptances, total cost, latency
60 Scenario 7: Insurer Underwriting Agent-Failure Insurance 57 3. Calculates TCVO independently 4. Issues certification: “OnIndex Verified: TCVO 35% below market average” 58.3 OnIndex Value Provides credible third-party verification. Reduces buyer due diligence burden. Enables startup to differentiate on verified metrics rather than marketing claims. 59 Scenario 6: Enterprise Dynamic Provider Routing 59.1 Transaction Overview • Buyer: Fortune 500 company with multi-thousand dollar daily AI spend • Suppliers: 5+ AI providers (OpenAI, Anthropic, self-hosted, etc.) • Mechanism: Real-time routing based on OnIndex price/quality data 59.2 Routing Logic IF task_difficulty == LOW: route_to = cheapest_provider ELIF task_difficulty == HIGH: route_to = highest_quality_provider ELIF latency_required < 2_seconds: route_to = fastest_provider ELSE: route_to = optimal_price_quality_provider 59.3 OnIndex Value Real-time benchmark data enables optimal routing. Cost savings of 15-30% vs. single-provider approach. 60 Scenario 7: Insurer Underwriting Agent-Failure Insur- ance 60.1 Transaction Overview • Buyer: AI agent provider seeking risk transfer
61 Scenario 8: Seasonal Company Reserving Future Capacity 58 • Supplier: Specialty insurer • Coverage: Revenue loss from agent failure events • Premium: Based on OnIndex reliability data 60.2 Risk Assessment • Historical Failure Rate: 2.3% (from OnIndex data) • Severity Distribution: Mean $50K per event • Premium: $1, 150 per $50K coverage (2.3%) 60.3 OnIndex Value Provides actuarial data for insurance pricing. Enables new risk transfer markets for AI opera- tions. 61 Scenario 8: Seasonal Company Reserving Future Capac- ity 61.1 Transaction Overview • Buyer: Tax preparation software company (seasonal demand) • Supplier: AI infrastructure provider • Unit: Peak-season AI resolution capacity • Timing: Reserve in June for following April 61.2 Forward Contract Structure • Spot Price (June): $0.80/resolution • Forward Price (April delivery): $1.00/resolution • Premium: 25% for price certainty • Quantity: 100,000 resolutions
63 Scenario 10: Hospitality Comparing Human and Robotic Labor 59 61.3 OnIndex Value Forward curve enables price discovery for future delivery. Hedging mechanism for seasonal businesses. 62 Scenario 9: Bank Financing Agent Provider 62.1 Transaction Overview • Borrower: AI agent provider with contracted revenue • Lender: Commercial bank • Collateral: Future machine work contracts • Amount: $5M line of credit 62.2 Financing Structure • Provider has 3-year contract for 500K resolutions/year at $1.00/resolution • OnIndex benchmark confirms market price stability • Bank advances 70% of contracted revenue (asset-based lending) • Interest rate: SOFR + 400bp (vs. SOFR + 800bp without benchmark) 62.3 OnIndex Value Enables asset-based lending against contracted AI work. Reduces lender risk through price transparency. 63 Scenario 10: Hospitality Comparing Human and Robotic Labor 63.1 Transaction Overview • Buyer: Hotel chain evaluating room service automation • Options: Human staff vs. robotic delivery • Unit: One room service delivery
64 Cross-Scenario Analysis 60 63.2 Cost Comparison Table 9: Human vs. Robotic Labor in Hospitality Cost Component Human Robot Labor/Delivery $8.00 $2.00 Equipment/Amortization $0 $3.00 Maintenance $0 $1.00 Failure/Escalation $0 $1.50 Total Cost $8.00 $7.50 63.3 OnIndex Value Provides standardized comparison across labor types. Enables rational automation decisions. 64 Cross-Scenario Analysis 64.1 Common Patterns
- Price Discovery: All scenarios benefit from knowing “what is the market price?”
- Quality Verification: All require independent quality assessment
- Risk Quantification: All involve risks that can be priced with data
- Contract Standardization: All would benefit from standard terms 64.2 Value Proposition Summary Table 10: OnIndex Value by Scenario Scenario Primary Value Secondary Value
- Retail Capacity Fair pricing Quality assurance
- Bank Comparison Sovereignty assessment Risk comparison
- Software Budget True cost calculation Allocation optimization
- Government Policy evidence Competitiveness tracking
- Startup Cert Credibility Differentiation
- Dynamic Routing Cost optimization Quality maintenance
- Insurance Risk pricing Market enablement
- Seasonal Reserve Price certainty Hedging
- Financing Collateral valuation Lower rates
- Robotics Labor comparison Automation decisions
68 Stage 1: Measurement 61 65 Chapter Summary The ten scenarios demonstrate diverse applications for machine labor market infrastructure. Common themes include: need for independent price discovery, quality verification challenges, risk quantification value, and contract standardization benefits. Scenarios span immediate use cases (benchmarking, certification) to advanced applications (insurance, financing, forward markets). All scenarios demonstrate significant transaction friction without independent in- dexing. 66 Market Evolution 67 Introduction This section evaluates a proposed ten-stage progression from measurement to derivatives. For each stage, we identify prerequisites, customers, products, revenue models, regulatory require- ments, technical requirements, liquidity requirements, likely timelines, probability of emer- gence, and failure modes. 68 Stage 1: Measurement 68.1 Description Enterprises measure the real cost of completed agent work internally. 68.2 Prerequisites • AI agent deployment at scale • Basic cost tracking infrastructure • Willingness to invest in measurement 68.3 Customer Profile Large enterprises with 1M + annualAIspendseekingcostoptimization. 68.4 Product Cost analytics platform measuring TCVO.
69 Stage 2: Benchmarking 62 68.5 Revenue Model SaaS subscription: $50K-$200K annually per enterprise. 68.6 Requirements • Technical: API integrations, data pipelines, analytics engine • Regulatory: None • Liquidity: N/A (internal measurement) 68.7 Timeline 2024-2026 (current, emerging) 68.8 Probability High (90%): Already occurring as enterprises seek to understand AI ROI. 68.9 Failure Modes • Enterprises unwilling to share data • Measurement complexity exceeds value • Provider resistance to transparency 69 Stage 2: Benchmarking 69.1 Description OnIndex publishes independent prices for standardized machine-work categories. 69.2 Prerequisites • Aggregated data from multiple enterprises • Standardized unit definitions (SMWU) • Statistical methodology
69 Stage 2: Benchmarking 63 69.3 Customer Profile • Enterprises validating procurement decisions • Providers demonstrating competitiveness • Consultants advising on AI strategy 69.4 Product Published price indexes with daily/weekly updates. 69.5 Revenue Model • Enterprise subscriptions: $20K-$100K annually • Provider subscriptions: $10K-$50K annually • Consulting/analytics: Custom pricing 69.6 Requirements • Technical: Data aggregation platform, statistical engine, publication infrastructure • Regulatory: None (information provider) • Liquidity: Sufficient data points for statistical validity (100+ enterprises) 69.7 Timeline 2026-2028 69.8 Probability High (80%): Strong demand signal from enterprise procurement friction. 69.9 Failure Modes • Insufficient data sharing from enterprises • Inability to standardize across providers • Provider collusion to manipulate benchmarks
70 Stage 3: Certification 64 70 Stage 3: Certification 70.1 Description Providers receive independent performance, security, and jurisdiction grades. 70.2 Prerequisites • Established benchmarks • Testing protocols • Audit infrastructure 70.3 Customer Profile • Providers seeking differentiation • Enterprises requiring vendor validation • Regulators seeking compliance evidence 70.4 Product Certification badges, detailed scorecards, audit reports. 70.5 Revenue Model • Testing fees: $50K-$200K per certification • Annual renewal: $25K-$75K • Premium tiers (continuous monitoring): $100K-$300K 70.6 Requirements • Technical: Testing environments, security audit tools • Regulatory: Recognition by procurement authorities • Liquidity: N/A (certification is discrete) 70.7 Timeline 2027-2029
71 Stage 4: Index Licensing 65 70.8 Probability Moderate-High (70%): Certification markets exist in cloud, security; AI-specific likely to follow. 70.9 Failure Modes • Certification becomes meaningless (too many providers certified) • Providers game the tests • Liability for certification errors 71 Stage 4: Index Licensing 71.1 Description Procurement contracts reference OnIndex benchmarks in pricing formulas. 71.2 Prerequisites • Established, trusted benchmarks • Legal frameworks for index referencing • Contract templates 71.3 Customer Profile • Large enterprises with sophisticated procurement • Providers seeking long-term contracts • Government buyers 71.4 Product Index licensing for contract inclusion, contract templates, arbitration services. 71.5 Revenue Model • Licensing fees: 0.01-0.05% of notional contract value • Arbitration fees: $500/hour • Template licensing: $10K-$50K per contract type
72 Stage 5: Spot Procurement 66 71.6 Requirements • Technical: API access to real-time indexes • Regulatory: Recognition in commercial law • Liquidity: Standardized contracts referencing indexes 71.7 Timeline 2028-2030 71.8 Probability Moderate (60%): Requires widespread benchmark adoption and legal acceptance. 71.9 Failure Modes • Legal uncertainty about index referencing • Disputes over benchmark calculation • Competitors offer free alternatives 72 Stage 5: Spot Procurement 72.1 Description Companies purchase verified machine work through the platform. 72.2 Prerequisites • Provider network onboarded • Quality verification infrastructure • Settlement mechanisms 72.3 Customer Profile • Enterprises seeking on-demand AI capacity • SMEs without vendor relationships • Project-based buyers
73 Stage 6: Capacity Reservations 67 72.4 Product Spot marketplace for immediate task execution. 72.5 Revenue Model Transaction fees: 3-5% of transaction value. 72.6 Requirements • Technical: Matching engine, escrow, verification • Regulatory: Money transmitter licenses, consumer protection • Liquidity: Sufficient providers for price competition 72.7 Timeline 2029-2031 72.8 Probability Moderate (50%): Competes with direct provider relationships; value proposition must be clear. 72.9 Failure Modes • Providers prefer direct relationships • Quality verification too expensive • Dispute resolution failures 73 Stage 6: Capacity Reservations 73.1 Description Companies reserve future machine-work capacity.
73 Stage 6: Capacity Reservations 68 73.2 Prerequisites • Spot market liquidity • Forecasting capabilities • Reservation contract standards 73.3 Customer Profile Enterprises with predictable seasonal or project-based demand. 73.4 Product Capacity reservation contracts with guaranteed availability. 73.5 Revenue Model • Reservation fees: 5-10% of committed value • Cancellation fees • Premium for guaranteed capacity 73.6 Requirements • Technical: Forecasting models, capacity tracking • Regulatory: Contract enforcement mechanisms • Liquidity: Deep provider pool with excess capacity 73.7 Timeline 2030-2032 73.8 Probability Moderate (50%): Requires demonstrated capacity constraints and price volatility. 73.9 Failure Modes • Capacity proves abundant (no scarcity premium) • Provider underperformance on reservations • Buyers prefer spot flexibility
74 Stage 7: Forward Contracts 69 74 Stage 7: Forward Contracts 74.1 Description Buyers and providers agree today on future machine-work prices. 74.2 Prerequisites • Established spot price volatility • Standardized forward contract terms • Credit risk assessment 74.3 Customer Profile • Large buyers seeking budget certainty • Providers seeking revenue visibility • Hedgers (those with inverse exposures) 74.4 Product Over-the-counter forward contracts for future delivery. 74.5 Revenue Model • Arrangement fees: 0.5-1% of notional • Credit intermediation fees • Settlement services 74.6 Requirements • Technical: Forward curve modeling, credit scoring • Regulatory: OTC derivatives reporting (Dodd-Frank, EMIR) • Liquidity: Bilateral interest in forward trading 74.7 Timeline 2032-2035
75 Stage 8: Insurance and Financing 70 74.8 Probability Low-Moderate (40%): Requires sustained price volatility and large-scale adoption. 74.9 Failure Modes • Price volatility too low to justify forwards • Credit risk too high for bilateral trading • Regulatory burden exceeds value 75 Stage 8: Insurance and Financing 75.1 Description Insurers and lenders use benchmarks to price risk and finance providers. 75.2 Prerequisites • Reliable performance data • Actuarial models • Regulatory acceptance 75.3 Customer Profile • AI providers seeking risk transfer • Providers seeking working capital • Investors in AI receivables 75.4 Product • Performance insurance • Receivables financing • Revenue-based lending
76 Stage 9: Exchange Infrastructure 71 75.5 Revenue Model • Data licensing to insurers: $100K-$500K annually • Financing facilitation: 1-2% of loan value 75.6 Requirements • Technical: Risk models, performance history • Regulatory: Insurance regulation, banking regulation • Liquidity: Insurance market participation 75.7 Timeline 2031-2034 75.8 Probability Moderate (50%): Insurance markets develop where data enables risk pricing. 75.9 Failure Modes • Insufficient actuarial data • Moral hazard (insured providers underperform) • Regulatory restrictions 76 Stage 9: Exchange Infrastructure 76.1 Description Standardized machine-work contracts are traded and settled on an exchange. 76.2 Prerequisites • Liquid forward markets • Standardized contract specifications • Clearing infrastructure
77 Stage 10: Derivatives 72 76.3 Customer Profile Sophisticated traders, market makers, institutional hedgers. 76.4 Product Exchange-traded machine work contracts with centralized clearing. 76.5 Revenue Model • Trading fees: $0.50-$2.00 per contract • Clearing fees: $0.25-$1.00 per contract • Market data: $10K-$100K monthly 76.6 Requirements • Technical: Trading platform, matching engine, clearing system • Regulatory: Exchange licensing (SEC, CFTC) • Liquidity: Thousands of contracts daily 76.7 Timeline 2035-2040 76.8 Probability Low (25%): Requires massive market size and regulatory approval. 76.9 Failure Modes • Insufficient trading interest • Regulatory rejection • Competition from existing exchanges 77 Stage 10: Derivatives 77.1 Description Regulated institutions create futures, options, swaps, and index products.
77 Stage 10: Derivatives 73 77.2 Prerequisites • Established exchange • Institutional participation • Regulatory approval 77.3 Customer Profile Institutional investors, hedgers, speculators. 77.4 Product • Futures contracts • Options on futures • Total return swaps • ETF products 77.5 Revenue Model • Licensing index to product issuers: $1M-$10M annually • Data fees for pricing: $100K-$1M monthly 77.6 Requirements • Technical: Derivatives risk systems • Regulatory: Full derivatives regulation compliance • Liquidity: Institutional-scale participation 77.7 Timeline 2040+ 77.8 Probability Very Low (15%): Speculative; requires machine labor to become truly commodity-like.
81 Introduction 74 77.9 Failure Modes • Machine labor remains differentiated • Regulatory prohibition • Market manipulation 78 Summary Matrix Table 11: Market Evolution Stage Summary Stage Timeline Probability Revenue Model Key Risk
- Measurement 2024-26 90% SaaS Data sharing
- Benchmarking 2026-28 80% Subscriptions Standardization
- Certification 2027-29 70% Testing fees Gaming
- Index Licensing 2028-30 60% Licensing Legal
- Spot Procurement 2029-31 50% Transaction fees Liquidity
- Capacity Reserve 2030-32 50% Reservation fees Scarcity
- Forward Contracts 2032-35 40% Arrangement Volatility
- Insurance/Financing 2031-34 50% Data licensing Moral hazard
- Exchange 2035-40 25% Trading fees Regulation
- Derivatives 2040+ 15% Index licensing Commoditization 79 Chapter Summary The market evolution pathway shows declining probability with advancing stages. Stages 1- 4 (measurement through index licensing) represent a credible business trajectory with high probability. Stages 5-8 (spot through insurance) are achievable but require market conditions that may not materialize. Stages 9-10 (exchange and derivatives) are speculative and may never occur. The prudent strategy focuses on stages 1-4 while maintaining optionality for later stages through partnerships. 80 Business Model 81 Introduction This section analyzes twelve proposed revenue streams, distinguishing between immediately credible, later-stage viable, and speculative options. We develop a realistic five-year business model with customer profiles, annual contract values, margin analysis, and paths to revenue milestones.
82 Revenue Stream Analysis 75 82 Revenue Stream Analysis 82.1 Stream 1: Enterprise Market-Data Subscriptions Description: Subscription access to benchmark prices, market analytics, and trend reports. Target Customers: Fortune 1000 enterprises with 500K + annualAIspend. Pricing: $20K-$100K annually based on company size. Assessment: Credible Immediately • Precedent: Bloomberg, Refinitiv, S&P Global all sell data subscriptions • Demand signal: Enterprises struggle with AI cost visibility • Defensibility: Data network effects and methodology trust 82.2 Stream 2: Procurement Benchmarking Description: Consulting services to optimize AI procurement against market benchmarks. Target Customers: Enterprises conducting AI vendor RFPs. Pricing: $50K-$300K per engagement. Assessment: Credible Immediately • Precedent: Gartner, McKinsey sell procurement advisory • High value: 10-30% cost reduction potential • Challenge: Requires significant consulting capacity 82.3 Stream 3: Provider Testing and Certification Description: Independent testing and certification of AI agent performance. Target Customers: AI agent providers seeking differentiation. Pricing: $50K-$200K per certification. Assessment: Credible Immediately • Precedent: UL, SOC 2, ISO certification markets • Strong demand: Providers need credible differentiation • Network effects: More certified providers attract more buyers
82 Revenue Stream Analysis 76 82.4 Stream 4: Index and API Licensing Description: Licensing benchmark indexes for inclusion in contracts and applications. Target Customers: Enterprises, providers, contract management platforms. Pricing: $10K-$50K annually for index access. Assessment: Becomes Credible Later (Stage 4+) • Requires established benchmark reputation • Legal frameworks for index referencing must mature • High margins once established 82.5 Stream 5: Government and Sovereign Intelligence Contracts Description: Sovereignty-grade pricing data for government procurement and policy. Target Customers: Defense agencies, national AI offices, regulatory bodies. Pricing: $500K-$2M annually. Assessment: Becomes Credible Later (Stage 3+) • Requires FedRAMP or equivalent certification • Long sales cycles (18-36 months) • High barriers, high loyalty 82.6 Stream 6: Procurement Transaction Fees Description: Fees for machine work transacted through OnIndex platform. Target Customers: Buyers and sellers on spot marketplace. Pricing: 3-5% of transaction value. Assessment: Becomes Credible Later (Stage 5+) • Requires marketplace liquidity • Competes with direct relationships • High volume needed for significance 82.7 Stream 7: Verification and Settlement Fees Description: Fees for outcome verification and payment settlement. Target Customers: Marketplace participants. Pricing: 1-2% of transaction value. Assessment: Becomes Credible Later (Stage 5+)
82 Revenue Stream Analysis 77 • Requires marketplace operations • Verification costs must be low relative to value • Trust in verification is critical 82.8 Stream 8: Capacity-Reservation Fees Description: Fees for reserving future machine-work capacity. Target Customers: Enterprises with seasonal or project-based demand. Pricing: 5-10% of reserved value. Assessment: Speculative (Stage 6+) • Requires capacity scarcity • Complex forward pricing • Uncertain demand 82.9 Stream 9: Insurance-Data Licensing Description: Licensing performance data to insurers for risk pricing. Target Customers: Specialty insurers, reinsurers. Pricing: $100K-$500K annually. Assessment: Becomes Credible Later (Stage 8+) • Requires deep performance history • Actuarial validation required • Partnership with insurers needed 82.10 Stream 10: Benchmark Licensing to Regulated Exchanges Description: Licensing indexes to regulated exchanges for derivative products. Target Customers: Futures exchanges, ETF issuers. Pricing: $1M-$10M annually. Assessment: Speculative (Stage 9+) • Requires commodity-like market development • Regulatory approval needed • Very long time horizon
83 Five-Year Business Model 78 82.11 Stream 11: Financing and Underwriting Data Description: Data for lenders providing asset-based financing to AI providers. Target Customers: Commercial banks, asset-based lenders. Pricing: $50K-$200K annually per lender. Assessment: Becomes Credible Later (Stage 8+) • Requires proven benchmark stability • Risk model validation needed • Partnership approach required 82.12 Stream 12: Robotic Labor-Market Data Description: Pricing and performance data for physical robotic labor. Target Customers: Robotics companies, logistics operators, manufacturers. Pricing: $20K-$100K annually. Assessment: Speculative (2030+) • Market too immature currently • Physical constraints differ from digital • Long-term optionality 83 Five-Year Business Model 83.1 Year 1 (2026-2027): Foundation Focus: Product-market fit for measurement and benchmarking Revenue Streams: • Enterprise subscriptions: $300K ARR • Certification testing: $200K ARR Total ARR: $500K Customers: 10-15 enterprises, 4-6 providers Team: 8-10 people
83 Five-Year Business Model 79 83.2 Year 2 (2027-2028): Scale Focus: Market leadership in benchmarking Revenue Streams: • Enterprise subscriptions: $1.2M ARR • Certification testing: $800K ARR • Procurement consulting: $500K ARR Total ARR: $2.5M Customers: 40-60 enterprises, 15-25 providers Team: 20-25 people 83.3 Year 3 (2028-2029): Expansion Focus: Certification dominance, index licensing launch Revenue Streams: • Enterprise subscriptions: $2.5M ARR • Certification testing: $1.5M ARR • Procurement consulting: $1M ARR • Index licensing: $500K ARR Total ARR: $5.5M Customers: 100-150 enterprises, 40-60 providers Team: 40-50 people 83.4 Year 4 (2029-2030): Marketplace Focus: Spot marketplace launch Revenue Streams: • Enterprise subscriptions: $4M ARR • Certification testing: $2.5M ARR • Procurement consulting: $1.5M ARR • Index licensing: $1.5M ARR • Transaction fees: $1M ARR Total ARR: $10.5M Customers: 200-300 enterprises, 80-120 providers Team: 70-90 people
84 Unit Economics 80 83.5 Year 5 (2030-2031): Platform Focus: Platform scale, optionality for advanced stages Revenue Streams: • Enterprise subscriptions: $6M ARR • Certification testing: $4M ARR • Procurement consulting: $2M ARR • Index licensing: $3M ARR • Transaction fees: $3M ARR • Government contracts: $1M ARR Total ARR: $19M Customers: 400-600 enterprises, 150-250 providers Team: 120-150 people 84 Unit Economics 84.1 Customer Acquisition Cost (CAC) • Enterprise subscription: $50K-$100K • Certification: $20K-$40K • Consulting: $30K-$60K 84.2 Lifetime Value (LTV) • Enterprise subscription: $300K-$600K (5-year) • Certification: $150K-$300K (3-year certification cycle) • Consulting: $100K-$200K (2-3 engagements) 84.3 LTV/CAC Ratio Target: 3:1 or higher
85 Path to Revenue Milestones 81 84.4 Gross Margins • Subscriptions: 80-90% • Certification: 60-70% (testing costs) • Consulting: 40-50% (labor intensive) • Transaction fees: 70-80% 85 Path to Revenue Milestones 85.1 Path to $1M ARR Requirements: • 15-20 enterprise customers at $50K ACV • OR 10-15 enterprises + 5-8 certifications Timeline: 12-18 months from launch Key Activities: • Design partnerships with 3-5 enterprises • Initial benchmark publication • First certifications completed 85.2 Path to $10M ARR Requirements: • 150-200 enterprise customers • 50-80 certified providers • Marketplace achieving $30M+ GMV Timeline: 4-5 years Key Activities: • Establish category leadership • Launch spot marketplace • Government contracts
89 Market Comparison 82 85.3 Conditions for $1B Outcome Required Conditions:
- Machine labor becomes significant cost category for Fortune 500
- OnIndex achieves dominant benchmark position
- Marketplace reaches $1B+ GMV
- Index licensing widely adopted in contracts
- Expansion to 3-5 international markets Probability: 10-20% (high uncertainty) 86 Chapter Summary Revenue streams 1-3 (subscriptions, consulting, certification) are credible immediately and can drive the business to $5 −10M ARR. Streams 4-7 (licensing, government, marketplace) become viable in years 3-5 and can scale to $20−50M ARR. Streams 8-12 (capacity, insurance, exchange, derivatives, robotics) are speculative and should be treated as optionality. The five- year model shows a realistic path to $19M ARR through core data and certification services, with marketplace and advanced services providing upside optionality. Part IV Strategy and Risk Assessment 87 First Product and Go-to-Market 88 Introduction This section determines the strongest initial wedge for entering the machine labor market. We compare nine candidate markets across twelve criteria, design the minimum viable product, and specify the decisive 90-day experiment. 89 Market Comparison 89.1 Evaluation Criteria
- Task Repeatability: How consistently can the task be performed?
89 Market Comparison 83 2. Ease of Verification: How objectively can success be measured? 3. Available Volume: How much economic activity exists? 4. Existing AI Adoption: How much AI use already occurs? 5. Provider Competition: How many competing solutions exist? 6. Price Volatility: How much does cost vary? 7. Capacity Constraints: Is demand ever supply-constrained? 8. Seasonal Demand: Does demand vary significantly? 9. Regulatory Complexity: How complex is compliance? 10. Buyer Urgency: How urgent is the cost optimization need? 11. Willingness to Share Data: Will buyers share outcome data? 12. Willingness to Pay: Will buyers pay for benchmarking? 89.2 Market Scoring Table 12: Initial Wedge Market Comparison (1-5 scale) Criteria Support Eng Legal Ins Fin Sec Rsrch Health Robo Repeatability 5 3 3 4 5 3 1 3 2 Verification 5 4 3 4 5 3 2 3 4 Volume 5 5 4 5 5 4 3 4 3 AI Adoption 5 4 3 3 4 3 2 2 1 Competition 4 4 3 3 3 4 2 3 2 Price Volatility 3 4 3 2 2 3 4 2 4 Capacity Constraints 3 2 2 2 1 2 1 2 1 Seasonal Demand 4 2 1 2 2 2 2 1 2 Regulatory Complexity 2 2 4 4 4 4 2 5 3 Buyer Urgency 4 3 3 3 3 4 2 4 2 Data Sharing 4 3 2 3 3 3 3 2 3 Willingness to Pay 4 3 3 3 3 3 2 3 2 Total 48 39 34 38 40 36 24 34 29 89.3 Top Three Markets
- Customer Support (Score: 48) • Highest repeatability and verification • Largest existing AI adoption • Clear seasonal patterns creating urgency
90 Minimum Viable Product 84 • Strong data sharing willingness 2. Financial Operations (Score: 40) • High repeatability and verification • Large volume • Lower regulatory complexity than insurance 3. Software Engineering (Score: 39) • Large volume • Growing AI adoption • High price volatility creates urgency 90 Minimum Viable Product 90.1 The Core Question The MVP answers one expensive enterprise question: “Are we paying a fair price for each successfully resolved AI customer-support case, after accounting for failures and human escalation?” 90.2 MVP Components 90.2.1 Data Integration Required Integrations: • Zendesk, Intercom, or Salesforce Service Cloud (ticket data) • AI provider APIs (token usage, cost data) • Internal time-tracking (human escalation costs) Minimum Dataset: • 10,000+ resolved tickets • 3+ months of history • 2+ AI providers or configurations
91 Go-to-Market Strategy 85 90.2.2 Benchmark Methodology
- Calculate True Cost per Verified Outcome (TCVO) for each ticket
- Segment by ticket category, complexity, channel
- Compare against anonymized peer benchmarks
- Identify cost drivers (retries, escalation, etc.) 90.2.3 Customer Dashboard Key Metrics: • TCVO trend over time • Comparison to industry benchmark • Cost driver breakdown • Provider performance comparison • Seasonal patterns 90.2.4 Provider Report For AI Providers: • Performance vs. market average • Win/loss analysis • Competitive positioning • Improvement recommendations 91 Go-to-Market Strategy 91.1 Initial Sales Motion Target: Mid-market SaaS companies (50M−500M revenue) Why Mid-Market: • Sufficient AI spend to care (200K−2M annually) • Less complex procurement than Fortune 500 • Faster decision cycles
91 Go-to-Market Strategy 86 • Willing to innovate Sales Approach:
- Free assessment (analysis of last quarter’s data)
- Deliver cost optimization insights
- Convert to annual subscription 91.2 Design Partner Profile Ideal Design Partner: • 50-500 support agents • 10,000+ tickets monthly • Currently using or evaluating AI support • Multiple AI providers under consideration • Procurement or operations leader willing to collaborate • Technical team able to integrate APIs Compensation: Free service for 6 months in exchange for: • Data sharing • Product feedback • Case study participation • Reference availability 91.3 90-Day Experiment Objective: Validate that token prices are weak predictors of TCVO Hypothesis: Token price explains <60% of variance in TCVO Method:
- Week 1-2: Integrate with 3 design partners
- Week 3-8: Collect data on 5,000+ tickets per partner
- Week 9-10: Analyze cost drivers
- Week 11-12: Present findings and measure willingness-to-pay
92 Evidence Required Before Building an Exchange 87 Success Criteria: • TCVO varies 2x+ across providers for same task type • Token price correlation with TCVO <0.6 • 2+ design partners commit to paid subscription • Clear path to 10 paying customers within 6 months 91.4 First Ten Customers Target List:
- E-commerce company (high seasonal variation)
- SaaS company (technical support)
- Fintech (compliance-sensitive)
- Healthcare (regulated)
- B2B software (complex tickets)
- Consumer app (high volume)
- Marketplace (multi-sided support)
- Insurance (claims-related)
- Travel/hospitality (seasonal)
- Education (institutional) 92 Evidence Required Before Building an Exchange Before investing in marketplace infrastructure (Stage 5+), the following evidence must exist:
- Benchmark Adoption: 100+ enterprises referencing OnIndex in procurement
- Provider Interest: 20+ providers willing to transact through platform
- Standardization: SMWU specifications adopted by industry associations
- Price Volatility: Documented 20%+ price variation for equivalent work
- Regulatory Neutrality: No regulatory prohibition on AI spot markets
- Verification Infrastructure: Cost-effective outcome verification proven
- Liquidity Indicators: Pilot transactions show demand for marketplace
96 Company Categories 88 93 Chapter Summary Customer support emerges as the strongest initial wedge market due to high repeatability, strong existing AI adoption, and clear value proposition. The MVP answers the specific question: “Are we paying a fair price per verified resolution?” The 90-day experiment validates the core hypothesis that token prices are weak predictors of true costs. Success metrics include 2x+ TCVO variation across providers and 2+ paying customers from design partnerships. 94 Competitive Landscape 95 Introduction This section maps the competitive landscape for machine labor market infrastructure. We dis- tinguish between benchmark companies, routers, marketplaces, procurement platforms, ob- servability platforms, exchanges, rating agencies, testing laboratories, and clearinghouses. 96 Company Categories 96.1 Model Benchmarking • Artificial Analysis: Independent model comparison • LMSYS Chatbot Arena: Crowdsourced model evaluation • MLCommons: Industry-standard ML benchmarking 96.2 AI Observability • LangSmith: LangChain’s observability platform • Weights ˘0026 Biases: ML experiment tracking • Arize: AI observability and monitoring 96.3 Model Routing • Martian: Dynamic model routing • OpenRouter: Unified API for multiple models • Together AI: Inference optimization
99 Introduction 89 96.4 AI Gateways • Kong: API gateway with AI features • Cloudflare AI Gateway: Edge AI routing 96.5 GPU Exchanges • CoreWeave: GPU cloud provider • Lambda Labs: GPU cloud • Together AI: Decentralized compute 96.6 Agent Marketplaces • Relevance AI: AI agent platform • AutoGPT: Open-source agents • LangChain: Agent framework 97 OnIndex Differentiation OnIndex is differentiated by:
- Focus on verified outcomes rather than model capabilities
- Independent, third-party position
- Standardized unit definitions
- True cost methodology 98 Market Design and Regulation 99 Introduction This section investigates market design and regulatory considerations for machine labor mar- kets.
101 Regulatory Structure 90 100 Benchmark Manipulation Risk [Evidence] Benchmark manipulation has occurred in: • LIBOR (interbank rates) • ISDAfix (interest rate swaps) • WM/Reuters FX benchmarks [Inference] Machine labor benchmarks face similar risks if: • Expert judgment is involved in price formation • Market is concentrated among few participants • Large financial incentives exist for manipulation 101 Regulatory Structure 101.1 Benchmark Regulation [Evidence] EU Benchmarks Regulation (BMR) requires: • Benchmark administrators to be authorized • Methodology transparency • Governance arrangements • Conflicts of interest management 101.2 Recommended Structure [Inference] OnIndex should become:
- A software and data company (immediate)
- A benchmark administrator (year 2-3)
- Potentially licensing data to regulated exchanges rather than becoming an exchange
105 Objection 1: Agent Work Is Too Heterogeneous to Standardize 91 102 Conclusion The most defensible structure is an independent data and certification company, potentially licensing benchmarks to regulated exchanges rather than directly operating exchange infras- tructure. Part V Falsification and Empirical Research 103 Falsification Analysis 104 Introduction This section—the most critical in the monograph—attempts to actively disprove the entire the- sis. We test fifteen objections, providing the strongest evidence supporting each, the strongest counterargument, a proposed experiment to resolve the uncertainty, and an assessment of whether the objection threatens the index, the marketplace, the exchange, or the entire thesis. 105 Objection 1: Agent Work Is Too Heterogeneous to Stan- dardize 105.1 Supporting Evidence [Evidence] AI agent outputs vary dramatically: • Same prompt produces different outputs with same model • Quality varies by context, history, and tool availability • Success criteria are often subjective • No two customer support tickets are identical Historical precedent: Creative services, consulting, and research have never achieved stan- dardization despite market size. 105.2 Counterargument [Inference] Heterogeneity can be managed through:
106 Objection 2: Model Prices Will Fall So Quickly That Hedging Is Unnecessary 92 • Task categorization (simple, medium, complex) • Grade bands rather than point specifications • Quality-adjusted pricing • Statistical aggregation across many transactions Precedent: Freight shipping (infinite route combinations) achieved standardization through abstraction. 105.3 Resolution Experiment Test inter-rater reliability on task difficulty classification:
- Have 10 experts classify 1,000 support tickets by difficulty
- Measure Cohen’s kappa for agreement
- If κ > 0.7, standardization is feasible 105.4 Threat Assessment Threatens: Exchange (requires high standardization) Threatens Index: No (indexes can han- dle heterogeneity through stratification) Verdict: Manageable with proper methodology 106 Objection 2: Model Prices Will Fall So Quickly That Hedging Is Unnecessary 106.1 Supporting Evidence [Evidence] Historical trends: • GPT-4 price dropped 50% within 12 months of launch • Open-source models (Llama, Mistral) provide free alternatives • GPU costs declining 10-30% annually • Competition driving prices toward marginal cost If prices only fall, who needs hedging?
107 Objection 3: Compute Will Remain the Only Relevant Commodity 93 106.2 Counterargument [Inference] Price trends are not uniformly downward: • New capabilities command premium pricing • Sovereignty requirements add permanent premiums • Compute scarcity during peak demand • Quality-adjusted prices may be stable even if token prices fall Precedent: Cloud compute spot prices have risen in recent years despite hardware cost declines. 106.3 Resolution Experiment Track quality-adjusted prices for standardized task basket over 12 months. If variance is low, hedging demand is weak. 106.4 Threat Assessment Threatens: Exchange, forwards, derivatives Threatens Index: No (benchmarks useful even in falling markets) Verdict: Significant threat to advanced stages; manageable for core business 107 Objection 3: Compute Will Remain the Only Relevant Commodity 107.1 Supporting Evidence [Evidence] Current market structure: • AI buyers purchase compute (AWS, GCP, Azure) • Or purchase model access (tokens) • No one currently buys “outcomes” • Compute markets are commoditizing Why add another abstraction layer?
108 Objection 4: Enterprises Will Prefer Long-Term Vendor Contracts 94 107.2 Counterargument [Inference] Economic value is in outcomes, not inputs: • Enterprises budget for “tickets resolved” not “tokens consumed” • Compute price poorly predicts resolution cost (retries, failures) • BPO market exists because outcome pricing is valuable 107.3 Resolution Experiment Survey 50 enterprises: “Would you prefer outcome-based pricing or token-based pricing?” If <40% prefer outcomes, objection is valid. 107.4 Threat Assessment Threatens: Entire thesis Verdict: Critical objection; requires strong survey evidence to over- come 108 Objection 4: Enterprises Will Prefer Long-Term Vendor Contracts 108.1 Supporting Evidence [Evidence] Enterprise procurement patterns: • Multi-year SaaS contracts are standard • Vendor relationships valued over spot pricing • Procurement processes favor incumbents • Switching costs are high 108.2 Counterargument [Inference] Long-term contracts can reference indexes: • Price adjustment clauses based on benchmarks • Performance guarantees tied to verified outcomes • Commodity markets coexist with long-term contracts
110 Objection 6: Providers Will Vertically Integrate Benchmarking 95 108.3 Threat Assessment Threatens: Spot marketplace Threatens Index: No (indexes valuable for contract adjustment) Verdict: Requires pivot to index licensing rather than spot trading 109 Objection 5: Agent Capacity Will Not Be Scarce 109.1 Supporting Evidence [Evidence] Scalability of AI: • Software is infinitely replicable at near-zero marginal cost • Cloud infrastructure scales elastically • No physical constraints like oil wells or power plants If capacity is always abundant, why would forward markets develop? 109.2 Counterargument [Inference] Scarcity can emerge from: • Model provider rate limits • GPU shortages during demand spikes • Skilled implementation talent constraints • Regulatory capacity limits (sovereignty requirements) 109.3 Threat Assessment Threatens: Forward markets, capacity reservations Verdict: Valid concern; capacity markets may not develop 110 Objection 6: Providers Will Vertically Integrate Bench- marking 110.1 Supporting Evidence [Evidence] Vertical integration patterns: • AWS provides its own cost analytics
111 Objection 7: Buyers Will Not Share Enough Data 96 • OpenAI provides its own benchmarking • Large providers prefer walled gardens 110.2 Counterargument [Inference] Independent benchmarking has persisted because: • Buyers demand neutrality • Provider benchmarks lack credibility • Multi-provider buyers need cross-provider comparison Precedent: S&P, Moody’s, Fitch persist despite issuer-paid model conflicts. 110.3 Threat Assessment Threatens: Market share Verdict: Manageable through independence positioning 111 Objection 7: Buyers Will Not Share Enough Data 111.1 Supporting Evidence [Evidence] Data sharing challenges: • Customer data is sensitive • Performance data is competitively sensitive • Procurement data reveals strategy • Privacy regulations (GDPR, CCPA) restrict sharing 111.2 Counterargument [Inference] Data sharing can be enabled through: • Anonymization and aggregation • Data trusts with independent governance • Direct value exchange (benchmarks in return for data) • Synthetic data generation
112 Objection 8: Quality Cannot Be Objectively Verified 97 111.3 Resolution Experiment Pilot with 5 enterprises: measure actual data sharing willingness under NDA. 111.4 Threat Assessment Threatens: Entire thesis (no data = no benchmarks) Verdict: Critical objection; requires pilot validation 112 Objection 8: Quality Cannot Be Objectively Verified 112.1 Supporting Evidence [Evidence] Quality challenges: • Human judgment often required • Context-dependent evaluation • Subjective success criteria • Cost of verification approaches cost of production 112.2 Counterargument [Inference] Verification is feasible for: • Structured tasks (forms, reconciliation) • Testable outputs (code, calculations) • Sample-based human review Scope to verifiable tasks initially. 112.3 Threat Assessment Threatens: Marketplace settlement Threatens Index: No (indexes can use sampling) Ver- dict: Requires narrow initial scope
114 Objection 10: Forward Contracts Will Have Insufficient Liquidity 98 113 Objection 9: Human Judgment Will Remain Necessary 113.1 Supporting Evidence AI agents fail on: • Novel situations • Ethical judgments • Complex negotiations • Creative problem-solving Human-in-the-loop is often required by regulation. 113.2 Counterargument [Inference] Human judgment can be priced as a cost component: • Include human review in TCVO calculation • Grade tasks by autonomy level • Market can exist for “human-supervised” work 113.3 Threat Assessment Threatens: Fully autonomous market vision Verdict: Manageable through grading system 114 Objection 10: Forward Contracts Will Have Insufficient Liquidity 114.1 Supporting Evidence [Evidence] Liquidity requirements: • Forward markets require matched buyers and sellers • AI demand is fragmented across use cases • No natural shorts (who would sell forward?)
115 Objection 11: Machine Labor Will Remain Software, Not a Commodity 99 114.2 Counterargument [Inference] Liquidity can develop from: • Providers hedging revenue • Buyers hedging costs • Speculators with opposing views 114.3 Threat Assessment Threatens: Forward markets, derivatives Verdict: Valid concern; may not develop 115 Objection 11: Machine Labor Will Remain Software, Not a Commodity 115.1 Supporting Evidence [Evidence] Software characteristics: • Differentiated by features and integration • High switching costs • Value-added services dominate • Relationship-based sales 115.2 Counterargument [Inference] Software markets can develop commodity layers: • Cloud compute is software-like but commoditized • BPO successfully abstracted labor • Advertising impressions commoditized despite creativity 115.3 Threat Assessment Threatens: Full commoditization vision Verdict: Requires focus on specific, standardizable layers
117 Objection 13: Governments Will Prohibit Cross-Border Agent Markets 100 116 Objection 12: Cloud Providers or Exchanges Will Cap- ture the Market 116.1 Supporting Evidence [Evidence] Competitive threats: • AWS, GCP, Azure have data and customer relationships • CME, ICE could launch AI futures • Existing BPO firms could expand benchmarking 116.2 Counterargument [Inference] Incumbents face conflicts: • Cloud providers prefer opacity • Exchanges lack AI expertise • BPO firms lack independence First-mover advantage in trust and methodology matters. 116.3 Threat Assessment Threatens: Market share, margins Verdict: Significant threat; requires speed and differentia- tion 117 Objection 13: Governments Will Prohibit Cross-Border Agent Markets 117.1 Supporting Evidence [Evidence] Regulatory trends: • Data localization requirements increasing • AI export controls (U.S. chip restrictions) • National AI strategies favoring domestic providers
118 Objection 14: Outcome-Based Pricing Will Remove Need for Independent Index 101 117.2 Counterargument [Inference] Fragmentation can be managed: • Separate indexes by jurisdiction • Sovereignty premium framework • Compliance as a service 117.3 Threat Assessment Threatens: Global unified market Threatens Index: No (indexes can be jurisdiction-specific) Verdict: Requires regional strategy 118 Objection 14: Outcome-Based Pricing Will Remove Need for Independent Index 118.1 Supporting Evidence If buyers and providers agree on outcome pricing bilaterally, why need a benchmark? 118.2 Counterargument [Inference] Benchmarks needed for: • Validating fair pricing in bilateral negotiations • Contract adjustment clauses • New provider entry • Regulatory oversight Precedent: LIBOR used even when banks lent at bilateral rates. 118.3 Threat Assessment Threatens: Marketplace (if bilateral dominates) Threatens Index: No (indexes support bilat- eral) Verdict: Limited threat
121 Conclusion on Falsification 102 119 Objection 15: The Exchange Vision Is Premature by Decades 119.1 Supporting Evidence [Evidence] Market development timelines: • Oil futures: 100+ years from first commercial production • Electricity markets: 50+ years from grid formation • Cloud spot markets: 15+ years, still limited 119.2 Counterargument [Inference] Technology accelerates market formation: • Digital infrastructure enables faster standardization • Existing cloud markets provide template • AI adoption is faster than prior technologies 119.3 Threat Assessment Threatens: Exchange stages (9-10) Verdict: Valid; exchange is speculative 120 Falsification Summary 121 Conclusion on Falsification The adversarial analysis reveals that:
- Critical threats to the entire thesis: Objections 3 (compute only) and 7 (no data sharing)
- High threats to advanced stages: Objections 2, 5, 10, 12, 15
- Manageable threats to core business: Most other objections The prudent conclusion is that stages 1-4 (measurement through index licensing) are viable with proper execution, while stages 5-10 (marketplace through derivatives) face substantial obstacles that may prove insurmountable. The exchange vision in particular appears premature and possibly incorrect.
123 Hypotheses for Testing 103 Table 13: Falsification Objection Summary Objection Threat Level Threatens Mitigation
- Heterogeneity Medium Exchange Grading system
- Price collapse High Forwards/Derivatives Focus on core
- Compute only Critical Entire thesis Survey validation
- Long-term contracts Medium Marketplace Index licensing
- No scarcity High Forwards Accept limitation
- Vertical integration Medium Share Independence
- No data sharing Critical Entire thesis Pilot required
- Quality unverifiable Medium Settlement Narrow scope
- Human necessary Low Autonomy Include in cost
- No liquidity High Forwards Accept limitation
- Remains software Medium Commodity Focus on layers
- Incumbents win High Share/Margins Speed/focus
- Government blocks Medium Global market Regional
- Outcome pricing Low Marketplace Limited threat
- Premature High Exchange Accept limitation 122 Hypotheses and Empirical Research 123 Hypotheses for Testing 123.1 H1: Token prices are weak predictors of final cost Test: Correlation analysis between token costs and TCVO 123.2 H2: Model rankings change with full cost accounting Test: Compare rankings by token price vs. rankings by TCVO 123.3 H3: Prices vary across jurisdictions Test: Measure sovereignty premiums 123.4 H4: Sovereignty requirements create persistent premiums Test: Track price differentials over time 123.5 H5: At least one task category can be standardized Test: Inter-rater reliability study
125 Interview Questions by Stakeholder 104 123.6 H6: Enterprises will pay for neutral comparisons Test: Willingness-to-pay survey 123.7 H7: Demand will exhibit seasonality Test: Time-series analysis of demand patterns 123.8 H8: Benchmark ownership produces defensibility Test: Market share analysis 123.9 H9: Dataset is the most valuable asset Test: Competitive analysis 123.10 H10: Exchange will be built through partnerships Test: Partner discussions 124 Interview Protocol 125 Interview Questions by Stakeholder 125.1 Enterprise CFOs
- What is your annual AI spending?
- How do you measure AI ROI?
- What would you pay for independent cost benchmarking? 125.2 AI Procurement Leaders
- How do you compare AI vendors today?
- What data would you share with an independent benchmark?
- Would you reference an index in contracts?
127 Five-Year Development Path 105 125.3 Model Providers
- How do you differentiate on price?
- Would you pay for independent certification?
- What concerns do you have about commoditization? Part VI Conclusion and Future Directions 126 Five-Year Path and Twenty-Year Scenarios 127 Five-Year Development Path 127.1 Year 1: Foundation • Launch measurement product • Secure 3-5 design partners • Achieve $500K ARR 127.2 Year 2: Scale • Launch certification service • Achieve $2.5M ARR 127.3 Year 3: Expansion • Launch index licensing • Achieve $5.5M ARR 127.4 Year 4: Marketplace • Launch spot procurement • Achieve $10.5M ARR
131 Assessment of the Four Conclusions 106 127.5 Year 5: Platform • Scale marketplace • Achieve $19M ARR 128 Twenty-Year Scenarios 128.1 Bull Case Machine labor becomes fully commoditized with global exchanges and derivative markets. 128.2 Base Case Benchmarking and certification become standard; marketplace achieves moderate scale. 128.3 Bear Case Vertical integration by cloud providers prevents independent market formation. 129 Conclusion 130 Summary of Findings This monograph has conducted a comprehensive, adversarial investigation into whether ma- chine labor will evolve into standardized, financialized commodity markets. We analyzed twelve historical commodity markets, developed frameworks for machine labor definition and standardization, created pricing methodologies, modeled ten market scenarios, evaluated a ten- stage market evolution pathway, and subjected the entire thesis to fifteen falsification attempts. 131 Assessment of the Four Conclusions 131.1 A. Structurally Inevitable This conclusion posits that machine labor markets are highly likely under current trends. Assessment: Not supported by evidence. The falsification analysis reveals multiple critical obstacles (data sharing, compute commoditization, price volatility) that could prevent market formation. Historical precedent shows many potential commodity markets fail to develop.
132 Final Conclusion 107 131.2 B. Conditionally Inevitable This conclusion posits that markets become likely only after specific thresholds in agent relia- bility, enterprise spending, standardization, and capacity volatility. Assessment: Supported by evidence. Stages 1-4 (measurement through index licensing) are viable contingent on: • Enterprise willingness to share outcome data (pilot validation required) • Demonstrated cost variance across providers (2x+ TCVO variation) • Adoption of SMWU standards by industry participants • Regulatory neutrality toward benchmarking 131.3 C. Plausible but Premature This conclusion posits that the index is useful, but the exchange thesis remains too early. Assessment: Partially supported. The benchmarking and certification business (stages 1-3) is viable in the near term. Index licensing (stage 4) requires 2-3 years of market development. The marketplace and forward market vision (stages 5-8) is premature and may never material- ize. 131.4 D. Incorrect Market Analogy This conclusion posets that machine labor will remain software or managed services rather than becoming a commodity market. Assessment: Partially supported for advanced stages. The exchange and derivatives vision (stages 9-10) may be an incorrect analogy. However, the benchmarking and certification layers (stages 1-3) are valid regardless of whether full commoditization occurs. 132 Final Conclusion The evidence supports Conclusion B (Conditionally Inevitable) for stages 1-4, with ele- ments of Conclusion C (Plausible but Premature) for stages 5-8, and Conclusion D (In- correct Analogy) for stages 9-10. Machine labor markets are likely to develop through the measurement, benchmarking, and certification stages because:
- Enterprises demonstrably struggle with AI cost visibility
- Providers seek credible differentiation mechanisms
133 Strategic Recommendations 108 3. Precedent exists in cloud cost optimization and security certification 4. The value proposition (fair pricing, quality verification) is clear Forward markets and exchange infrastructure (stages 7-10) face substantial obstacles:
- Price volatility may be insufficient to create hedging demand
- Provider concentration may prevent true market formation
- Liquidity may never reach institutional thresholds
- Regulatory burden may exceed value 133 Strategic Recommendations 133.1 The Strongest Version of the Startup OnIndex should position as an independent data and certification company, not an exchange or marketplace. The strongest version: • Core: True cost measurement and benchmarking for AI customer support • Expansion: Certification services for AI agent providers • Optionality: Index licensing for enterprise contracts • Partnerships: Work with regulated exchanges for any derivative products (do not build directly) 133.2 The Weakest Assumption The weakest assumption in the thesis is that enterprises will share detailed outcome data with an independent third party. This requires validation through the 90-day pilot. 133.3 The First Product True cost analytics for AI customer support, answering: “What is our real cost per verified resolution, and how does it compare to the market?”
133 Strategic Recommendations 109 133.4 The First Customer Mid-market SaaS company (50M−500M revenue) with: • 50-500 support agents • Active AI deployment or evaluation • Procurement leader seeking cost optimization • Willingness to share anonymized data 133.5 The First Dataset 10,000+ resolved support tickets across 3+ months, including: • Ticket metadata (category, channel, customer) • AI provider data (model, tokens, cost) • Resolution outcomes (resolution time, reopen rate, CSAT) • Human escalation costs 133.6 The Decisive 90-Day Experiment Objective: Validate that token prices are weak predictors of TCVO Hypothesis: Token price explains <60% of variance in True Cost per Verified Outcome Method:
- Integrate with 3 design partner enterprises
- Collect data on 5,000+ tickets per partner
- Calculate TCVO for each ticket
- Analyze correlation between token costs and TCVO Success Criteria: • Token-TCVO correlation <0.6 • 2x+ TCVO variation across providers for same task types • 2+ design partners commit to paid subscription
134 Limitations of This Research 110 133.7 Evidence to Continue Proceed if:
- 90-day experiment meets success criteria
- 2+ paying customers secured within 6 months of launch
- Clear path to $1M ARR within 18 months 133.8 Evidence to Abandon or Pivot Abandon if:
- Token-TCVO correlation >0.8 (token prices are sufficient)
- Enterprises refuse to share data even under NDA
- No paying customers after 12 months
- Cloud providers launch credible competing services Pivot to consulting/advisory if data sharing proves impossible but demand exists for cost optimization expertise. 134 Limitations of This Research
- Data Constraints: Limited empirical data on actual enterprise AI costs due to confiden- tiality
- Rapid Change: AI market evolving faster than research cycle; conclusions may become outdated
- Geographic Bias: Research weighted toward U.S. and European markets
- Survivorship Bias: Historical commodity analysis focuses on successful markets; failed markets understudied
- Provider Coordination: Uncertainty about how model providers will respond to com- moditization pressure
139 Future Research Agenda 111 135 Future Research Agenda
- Empirical pilot with 3-5 design partner enterprises
- Inter-rater reliability study for task difficulty classification
- Survey of 100+ enterprises on willingness to pay for benchmarking
- Analysis of cloud provider pricing strategies and responses
- Regulatory landscape tracking (AI Act, executive orders, etc.)
- Longitudinal study of AI price trends and volatility 136 Final Statement The machine labor market thesis is neither obviously correct nor obviously false. The core insight—that enterprises need visibility into true costs of AI outcomes—is valid and valuable. The extended vision of exchanges and derivatives is speculative and possibly incorrect. The prudent path is to build the measurement and benchmarking business that is demonstrably needed today, while maintaining optionality for market infrastructure that may or may not become viable in the future. The decisive factor is whether enterprises will share data. That question can only be an- swered through direct engagement, not further analysis. The time for research has passed; the time for experimentation has begun. 137 Limitations and Research Agenda 138 Research Limitations
- Data Constraints: Limited empirical data on enterprise AI costs
- Temporal Validity: Rapid market evolution may outpace conclusions
- Geographic Coverage: Research weighted toward Western markets
- Selection Bias: Focus on successful commodity markets 139 Future Research Agenda
- Empirical pilot studies with design partners
139 Future Research Agenda 112 2. Inter-rater reliability studies for task classification 3. Enterprise willingness-to-pay surveys 4. Longitudinal price tracking
L Model Pricing Compilation 113 Appendices A Mathematical Formulas B Placeholder Content to be expanded in future versions. C Task Definitions and Specifications D Task Definitions Detailed task specifications for SMWU framework. E Interview Protocols F Interview Protocols Detailed interview guides by stakeholder type. G Benchmark Methodology H Benchmark Methodology Detailed statistical methodology for index construction. I Historical Market Data J Historical Market Data Compilation of historical data from commodity markets analyzed. K Model Pricing Compilation L Model Pricing Compilation Current pricing data for major AI models and cloud providers.
L Model Pricing Compilation 114 Table 14: Model API Pricing (as of July 2026) Model Input Output GPT-4o $0.005/1K $0.015/1K Claude 3.5 Sonnet $0.003/1K $0.015/1K Llama 3.1 405B $0.002/1K $0.002/1K