Business Viability, Defensibility, and AI Economics
How to test AI product defensibility, model costs and review load, choose pricing, and prove the economics of a production workflow.

On this page
TL;DR
- Model prices and capabilities change quickly, but reliability, human review, integration, and support still determine commercial viability.
- Per-seat pricing can create a cannibalisation paradox when successful agents reduce the number of human users. Pricing should follow the value unit the customer receives.
- AI assistance adds variable cost. It becomes a margin trap when that cost grows without enough workflow value, willingness to pay, or removed operating cost.
- Access to a capable model is rarely defensible. Advantage compounds through distribution, workflow position, lawful learning, trust, ecosystems, and operating depth.
The AI economics shift
Inference is only one part of the AI business case.
Provider prices, caching discounts, open-weight capabilities, and fine-tuning tools move too quickly for a durable handbook to freeze one comparison. Use current prices and measured workload traces. Include model calls, retries, tools, storage, infrastructure, human review, support, and governance.
Low prototype cost can hide expensive production behaviour. A feature that works 92% of the time sounds impressive until you calculate the consequence of the other 8% at scale. The viable AI product handles failure economically and safely, rather than optimising only the token bill.
This changes the viability question. Market sizing still matters, but the AI-specific questions centre on data, defensibility, margins, pricing, and the operating system around the job.
Data feasibility and the data flywheel
For AI products, feasibility includes whether the team has lawful, representative, maintainable data and enough engineering depth to operate the system. Neither should be assumed.
The data feasibility report
Before an AI bet proceeds, document answers to four questions:
Acquisition and quality. Do you have this data? Is it sufficient in volume and representative (unbiased)? What's the plan to clean, label, and maintain quality over time?
Provenance and legality. Where did this data come from? Do you have the legal rights, licences, and permissions to use it for training a commercial model?
Privacy and ethics. Does this data contain PII or other sensitive information? What's the plan for anonymisation, de-identification, and aggregation? The AI governance chapter covers PII handling, data residency, and compliance requirements in depth.
Bias audit. What inherent biases exist in this data? What's the mitigation plan to ensure the model doesn't amplify them?
This report is the central artefact that unlocks the rest of the business case. You cannot choose a pricing model or build strategy until you've de-risked data acquisition.
The data flywheel
Data feasibility answers whether you can build the product. The data flywheel determines whether you can defend it.
Each user interaction can generate signal: corrections, selections, feedback, and usage patterns. Useful signal can improve retrieval, evals, workflow design, or models. A better product can attract more use and produce more evidence.
A functioning data flywheel can improve a product in ways that access to the same base model cannot. It is only defensible when the team has permission to use the signal, the signal predicts quality, and the learning reaches the product faster than competitors can reproduce the workflow.
When assessing viability, ask two flywheel questions:
- Does usage generate lawful signal that improves the product? Static data does not create a learning loop, but products can still differentiate through workflow, distribution, trust, or service.
- Is the improvement measurable and compounding? Track product quality against the evidence entering the loop. If usage grows while quality stays flat, there is no data advantage.
Build a defensibility stack
Model access and generated features diffuse quickly. Defensibility comes from assets and relationships that become harder to reproduce as the product operates.
Assess six layers:
| Layer | Defensibility test |
|---|---|
| Distribution | Does the product have a repeatable path to the right customers through existing users, workflow exposure, channels, community, brand, or partners? |
| Workflow position | Does the product hold context, approvals, collaboration, or system-of-record responsibility that makes replacement costly for a legitimate reason? |
| Learning loop | Does permitted production evidence improve quality, evaluation, or decisions faster than a competitor can reproduce the loop? |
| Trust | Has the product earned confidence through reliability, assurance, domain credibility, recovery, and honest boundaries? |
| Ecosystem | Do customers, developers, suppliers, or partners create complementary value that grows with participation? |
| Operating depth | Has the company built implementation, service, security, compliance, physical operations, or domain capability that software alone cannot copy? |
Score each layer with evidence, not aspiration. A proprietary dataset has weak value when it does not predict the outcome. Workflow lock-in built through data captivity can increase switching cost while destroying trust. A partner programme with no active exchange is a logo page.
Look for combinations. Distribution can bring more of the right work into a lawful learning loop. Better outcomes can earn trust and deeper workflow responsibility. Ecosystem participation can expand distribution and make the product more useful without the company building every complement.
Defensibility also creates obligations. Holding more customer context raises privacy and security exposure. Operating a marketplace creates governance work. Hardware and service depth add capital, supply, and support cost. Include those burdens in the business case.
Use the stack to guide investment:
- Which layer already has evidence?
- Which layer could compound from normal product use?
- What must the company operate to sustain it?
- Can a platform provider or well-funded competitor reproduce it?
- Does the advantage improve customer value, or only make exit painful?
The AI go-to-market chapter turns distribution and trust into a route from attention to repeat value.
The margin trap
Most established software companies are adding AI the wrong way. They bolt a "copilot" sidebar onto a platform architected a decade ago. Every product demo has a chatbot. Every roadmap has an "AI Assistant" workstream. Every earnings call mentions generative AI fifteen times.
The problem: they've added inference cost without removing workflow cost. The user still navigates the same screens, fills out the same forms, follows the same multi-step processes. The copilot auto-fills a few fields. Margins shrink. The feature looks modern. The P&L looks worse.
This is the margin trap. If the AI is optional (a sidebar the user can ignore), you've added COGS to your platform without changing the value equation. The user experience is incrementally better. The unit economics are materially worse.
The distinction that matters:
| Approach | Value proposition | Cost impact | Margin effect |
|---|---|---|---|
| Copilot (assistance) | Do work faster | Inference added on top of existing platform cost | Margin compression |
| Agent (replacement) | Do work for you | Inference replaces platform + human cost | Margin expansion |
A copilot may help someone fill out a form faster. An agent may remove the form, but it also introduces operating and review costs. Model the net value and cost for the actual workflow. Neither interaction pattern guarantees a margin outcome.
The litmus test: remove the AI from your product. Does the product still work? If yes, the AI is a feature, not a strategy. Now imagine the inverse: build the product where the AI is the product, where removing it means the product doesn't function. That's the difference between a margin trap and a viable AI business.
The cannibalisation paradox
Per-seat pricing powered a generation of billion-dollar SaaS companies. The logic was elegant: customer hires more humans, you sell more seats, revenue grows. Headcount and ARR moved in lockstep.
Agentic AI inverts the logic. You build autonomous agents. The customer needs fewer humans. You sell fewer seats. Revenue shrinks. The more successful your AI product, the more it cannibalises your seat-based revenue.
This is the Cannibalisation Paradox. Every efficiency gain you ship is a seat your customer no longer needs. The better your product gets, the less they pay you. No product leader wants to present that slide at a board meeting, but if you're building agentic capabilities on a per-seat model, that's the trajectory.
The fix: stop selling access and start selling outcomes. Shift from Software-as-a-Service to Service-as-a-Software.
When you sell a tool (Salesforce, Jira, Figma), you charge for the login. When you sell a result (the work itself, completed autonomously), you charge for the completed task. Tickets resolved. Contracts reviewed. Reports generated. Leads qualified. The pricing unit should be the smallest meaningful outcome your agent delivers, something the customer already understands and already values.
Price anchor: what did this task cost when a human did it? Your price should be meaningfully less than human cost, meaningfully more than inference cost. The spread is your margin, and unlike seat-based margin, it scales with volume.
Pricing for AI products
The COGS formula

Before setting any price, model the cost per query. Prompt caching changes the math significantly for products with repeated context (system prompts, document templates, recurring workflows).
COGS per Query = (Input Tokens × Cost/Token × Cache Miss Rate) + (Input Tokens × Cached Cost/Token × Cache Hit Rate) + (Output Tokens × Cost/Token) + Infrastructure Overhead
Run the formula with current provider pricing and the cache hit rate observed on representative traffic. Caching can materially reduce repeated-input cost, but the saving depends on provider rules, prefix stability, expiry, and workload shape. Do not put a generic discount into the business case without measuring it.
AI pricing models
| Model | How it works | When it fits | Risk |
|---|---|---|---|
| Pure usage-based | Charge per unit (API call, token, query) | Developer tools, infrastructure products | Customer anxiety over unpredictable bills; revenue volatility |
| Outcome-based (Service-as-a-Software) | Charge per successful result (ticket resolved, report generated) | Agentic products replacing human work | Defining "success" is hard; you absorb the cost of failures |
| Stand-alone add-on | AI features sold as a separate subscription tier | Quick monetisation of AI on an existing platform | Creates adoption friction; risks becoming the "optional copilot" |
| Hybrid: platform fee + metered outcomes | Flat base fee for access, metered charge for autonomous work done | Most agentic products transitioning from SaaS | More complex to build and communicate |
| Hybrid: seat + credit pool | Each seat contributes to a shared pool of AI usage credits | Teams transitioning gradually from per-seat | Power users exhaust the pool; doesn't solve the cannibalisation paradox long-term |
A hybrid platform fee + metered outcome model is one useful starting point when the platform creates baseline value and autonomous work has a measurable unit. Other products fit usage, subscription, add-on, or outcome pricing better. Test customer value, predictability, and measurement integrity before choosing.
This model solves the cannibalisation paradox because revenue grows with agent output, not headcount. It also creates natural expansion revenue: as agents prove themselves, customers route more volume through them without a sales conversation.
The internal alignment problem
Pricing changes fail when internal incentives don't follow. If engineering builds features that reduce human workload while sales is incentivised to increase seat count, your company is at war with itself.
Realign three things simultaneously:
- Sales incentives. Commission on consumption revenue and platform expansion, not seat count.
- Customer success metrics. Measure outcomes delivered, not daily active users and logins.
- Product metrics. Track work completed autonomously. A user who spends less time in your product because the agent handled everything is a success, not a churn risk.
The audit tax
Multi-agent architectures (where a manager model audits worker model outputs) introduce a cost multiplier that most teams don't model until it's too late. The agentic AI patterns chapter covers the architectural side; here, the focus is on the economics.
The math
Calculate audit cost from the actual trace:
Audit cost per task = Worker cost + (Review rate × Manager cost) + Deterministic checks + Human review cost
The manager cost must include its own input and output tokens, the worker traces it consumes, retries, and any tools it calls. If a larger model reviews every small-model result, oversight may cost far more than execution. The multiplier changes with provider pricing and workflow shape, so measure it rather than publishing one universal percentage.
The spot-check architecture
Route outputs through deterministic validation where possible. Use calibrated confidence, known-risk rules, and statistical sampling to decide which cases need model or human review. High-consequence cases may still require full approval regardless of confidence.
Four approaches to confidence scoring (combine them):
- Model-native confidence. Ask the worker to rate uncertainty, or generate multiple candidates and measure agreement.
- Rule-based validation. For structured outputs, validate against known constraints. Nearly free.
- Historical calibration. Track actual accuracy against confidence scores over time. Adjust thresholds based on observed performance.
- Domain heuristics. Route known hard inputs (long documents, ambiguous language) to the manager proactively.
Pricing reliability as a tier
Customers with stricter quality or assurance requirements create higher evaluation, audit, and support costs. Where the value unit is clear, reliability can become part of packaging. Keep safety and regulatory minimums outside optional pricing. Customers cannot buy a lower control standard where the organisation remains obligated to provide it.
Buy, build, or route
The buy, build, and partner decision now includes a fourth option: route tasks across managed and self-hosted models. Routing is useful only when the workload benefits enough to justify it.
What changed
Open-weight models widened deployment choice. For some tasks they provide acceptable quality with greater control over hosting, adaptation, or cost. Evaluate them against the same workload and operating standard as managed APIs.
Prompt caching changes some workloads. Repeated stable context may receive substantial discounts, while highly variable prompts receive little benefit. Model the observed hit rate.
Fine-tuning became more accessible. Running a job is easier than proving that the resulting model is lawful, better, maintainable, and worth operating. Treat adaptation as a measured product investment.
The routing layer emerged. Instead of choosing one model, production systems route different tasks to different models based on complexity, cost, and latency requirements. Simple classification goes to a small, fast model. Complex reasoning goes to a frontier model. The routing layer is the architectural decision, not the model selection. I cover this in detail in the chapter on multi-model orchestration.
The decision matrix
| Factor | API (proprietary) | Open-weight (self-hosted) | Routed (multi-model) |
|---|---|---|---|
| Time to market | Usually fastest | Includes infrastructure and operations | Includes routing and evaluation logic |
| Cost profile | Variable usage price | Infrastructure plus operations | Mixed, with routing overhead |
| Data control | Depends on provider and contract | Greater deployment control | Configurable by route |
| Customisation | Prompting and provider-supported adaptation | Full deployment and adaptation control | Varies by model in the stack |
| Migration effort | High if tightly coupled | High if infrastructure-specific | Lower only when interfaces and evals are clean |
| Strategic value | Speed and managed capability | Control and workload-specific economics | Task-level optimisation when evidence supports it |
Practical starting point: Use the simplest deployment that can validate the use case within its privacy and risk constraints. Put model access behind a stable interface and maintain task-level evals. Add routing, self-hosting, or fine-tuning when measured quality, cost, control, or resilience justifies the extra operations.
The strategic value rarely comes from the model alone. It comes from the product's workflow position, context, evaluation data, distribution, trust, and ability to improve the job over time.
What commercially rigorous AI PMs look like
| Behaviour | What it looks like in practice |
|---|---|
| Models COGS before features | Runs the cost-per-query formula before writing the PRD, not after launch |
| Treats reliability as a pricing lever | Offers tiered audit rates rather than promising blanket accuracy |
| Monitors the flywheel | Tracks model accuracy against usage volume to prove the compounding loop |
| Stress-tests the cannibalisation math | Models what happens to revenue when the agent handles 50%, then 80% of the workload |
| Prices on outcomes, not access | Defines the work unit the customer values and builds pricing around it |
| Builds for replacement, not assistance | Asks "should this workflow exist?" before asking "how do we add AI to this workflow?" |
| Preserves model options | Keeps model access testable and adds routing only when measured value exceeds its operating cost |
| Builds a defensibility stack | Invests in distribution, workflow, learning, trust, ecosystem, or operating depth that compounds with customer value |
The anti-pattern: viability theatre
The PM who runs a market analysis, picks a pricing model from a textbook, labels the dataset a moat, and calls the business case "validated." No COGS modelling. No cannibalisation analysis. No proof that usage improves the product or that distribution can reach the market.
Viability theatre produces a confident investment case without a measured production trace. After launch, review load, retries, integrations, or inference volume may push the workflow beyond its margin or capacity ceiling. The failure began before pricing: the team never analysed the complete system.
Do the math first. Then build.
v3.1 · Updated July 2026