← Research

Strategy

AI Evaluation & Decision-Making

An evaluation-first approach means scoring AI vulnerability, sizing opportunities with real data, and making build-vs-buy decisions before investing. Most teams skip evaluation and pay for it in wasted builds.

Core thesis

An evaluation-first approach means scoring AI vulnerability, sizing opportunities with real data, and making build-vs-buy decisions before investing. Most teams skip evaluation and pay for it in wasted builds. The LangChain State of Agent Engineering survey, published in early 2026, found that 57.3 percent of organizations have AI agents in production, but the number with measured ROI is far lower. Production without evaluation is just activity.

Why evaluation goes first

The natural instinct is to build. AI makes building faster than ever. A team can ship an AI feature in a week that would have taken a quarter two years ago. The problem is that building fast does not mean building the right thing. I have watched teams ship six AI features in six months, only to discover that none of them moved the number that mattered. They built fast. They evaluated never. The evaluation-first approach inverts the sequence. Before you build, you answer three questions. Is this the right problem? Is AI the right solution? Will the investment return more than it costs? Answering these questions before building saves months of wasted effort. The teams that win are not the fastest builders. They are the fastest evaluators.

The evaluation stack: three levels

Each level gates the next

Do not skip levels. A great opportunity sizing on a problem you should not solve is wasted work. A detailed build-vs-buy analysis on an opportunity too small to matter is analysis theater.

  • Level 1: AI vulnerability scoring. Does AI threaten or enable your business? Score your product, channel, and business model against AI risk factors.
  • Level 2: Opportunity sizing. Which AI opportunities generate the highest return for the lowest risk? Size each opportunity with real data, not guesswork.
  • Level 3: Build-vs-buy decision framework. For each sized opportunity, should you build, buy, partner, or wait? The decision changes based on capability maturity and market timing.

Level 1: AI vulnerability scoring

AI vulnerability scoring is a structured assessment of how exposed your business is to AI disruption. Score across three dimensions. Product vulnerability: can AI replicate your core value proposition? Score from one (no AI can touch it) to ten (AI can already do it better). Channel vulnerability: can AI saturate or bypass your distribution channels? Score from one (channel requires relationships AI cannot replicate) to ten (channel is fully automatable). Model vulnerability: can AI disrupt your business model by changing cost structures or customer willingness to pay? Score from one (model is AI-immune) to ten (model is obsolete if AI costs drop 50 percent). A combined score above 21 out of 30 means AI is an existential threat. Below 12 means AI is more opportunity than threat. Between 12 and 21 means selective vulnerability that requires targeted response.

Level 2: Opportunity sizing with real data

Opportunity sizing is where most teams fail. They estimate market size from Gartner reports and call it done. Real opportunity sizing requires three data points. First, the addressable problem: how many customers have this problem and how much does it cost them? Second, the AI advantage: how much better does AI solve this problem than the current alternative? Third, the capture rate: what percentage of the problem can you realistically capture given your channel and competitive position? Multiply those three numbers. That is your sized opportunity. Writer's 2026 Enterprise AI Adoption survey found that 54 percent of C-suite leaders cite integration difficulty as the primary AI adoption blocker. An opportunity that requires zero integration is 10x more capturable than one that requires deep integration. Size the capturable opportunity, not the theoretical one.

How to estimate the AI advantage

The AI advantage is the hardest number to estimate because the technology is moving. A good rule is to estimate the advantage at current capability and then discount it by 30 percent for the first year. AI always takes longer to operationalize than the demo suggests. The LangChain survey found that 32 percent of teams cite output quality as the top barrier. Quality problems erode the theoretical advantage. A feature that looks 10x better in a demo may deliver 3x better in production. Estimate conservatively. If the opportunity does not pencil out at conservative estimates, it probably does not pencil out. The teams I work with that get this wrong consistently overestimate the AI advantage and underestimate the operational cost.

Level 3: Build-vs-buy decision framework

Five factors that determine the decision

No single factor decides build vs. buy. The decision is a weighted score across all five. Weight depends on your strategy and risk tolerance.

  • Strategic importance: is this capability core to your differentiation or a commodity? Core capabilities are build candidates. Commodities are buy candidates.
  • Time to market: can you build faster than the market window or is buying the only way to ship in time? AI market windows are measured in months, not years.
  • Capability maturity: is the AI capability mature enough to buy or so nascent that building gives you an advantage? Buying immature AI creates integration debt.
  • Total cost of ownership: what is the three-year TCO of building vs. buying, including maintenance, model updates, and team cost? Build costs are always higher than initial estimates.
  • Switching cost: how hard is it to switch from a bought solution to a built one later? High switching costs favor building early. Low switching costs favor buying now and building later.

The risk assessment methodology

Every AI investment carries four categories of risk. Technical risk: will the AI work reliably in production? Score from one (proven technology, well-understood integration) to ten (research-grade AI with unknown reliability). Adoption risk: will customers actually use it? Score from one (customers are demanding this feature) to ten (customers have not asked for AI). Economic risk: will the unit economics work? Score from one (clear cost savings or revenue upside) to ten (unknown cost to serve, unknown willingness to pay). Competitive risk: what happens if a competitor ships first? Score from one (competitors are far behind) to ten (competitors are actively building the same thing). Score each risk. Multiply by the opportunity size. The result is a risk-adjusted opportunity score. Prioritize the highest risk-adjusted scores.

Total cost of ownership for AI investments

AI TCO is different from traditional software TCO. Three costs that most teams miss. First, model cost volatility. Frontier model prices may stay constant, but usage per task doubles every six months. A feature that costs $1,000 per month in API calls today may cost $8,000 per month in 18 months, even if per-token prices do not change. Second, evaluation cost. Every AI feature requires an evaluation pipeline that checks output quality. Building and maintaining that pipeline costs 20 to 40 percent of the feature build cost on an ongoing basis. Third, iteration cost. AI features require continuous tuning as models improve and customer expectations shift. Budget 30 percent of the initial build cost per year for ongoing iteration. A $100,000 AI feature actually costs $150,000 to $180,000 per year when you include all three hidden costs.

The evaluation timeline: when to evaluate vs. when to move

Evaluation has a shelf life. An evaluation done six months ago is stale. AI capabilities change. Competitive landscape shifts. Customer expectations reset. The evaluation timeline works on three horizons. Horizon one, zero to three months: evaluation is active and actionable. Decisions made on this evaluation are current. Horizon two, three to six months: evaluation is aging but still directionally useful. Decisions should be reviewed before execution. Horizon three, six-plus months: evaluation is expired. Re-evaluate before acting. The teams that struggle most are the ones that did a thorough evaluation 18 months ago and are still executing against it. The world moved. They did not re-evaluate. Build re-evaluation into the operating rhythm.

Real company examples: good and bad evaluations

A B2B SaaS company I worked with evaluated 12 AI opportunities. They scored each on vulnerability, opportunity size, and build-vs-buy. The top opportunity was an AI-powered onboarding assistant. They estimated a 40 percent reduction in time-to-value for new customers. They built it in eight weeks. Time-to-value dropped 35 percent. Retention at week four improved 22 percent. The evaluation was accurate because they used real customer data, not market reports. A different company skipped evaluation entirely. They built an AI feature because a board member asked for it. The feature launched. Nobody used it. Six months of engineering time. Zero adoption. The post-mortem revealed that no customer had ever asked for the feature. Evaluation would have caught that in week one.

The evaluation team: who needs to be in the room

AI evaluation requires a cross-functional team. Product owns the problem definition and customer context. Engineering owns the technical feasibility and build cost estimates. Data science or AI engineering owns the model capability assessment and quality risk. Finance owns the TCO model and ROI calculation. Sales or customer success owns the adoption risk assessment and customer willingness-to-pay signal. Legal or compliance owns the regulatory risk assessment, especially for regulated industries. The evaluation fails when any of these functions is absent. Product evaluates a feature engineering cannot build. Engineering builds a feature customers will not adopt. Finance approves an investment with hidden model costs. Get all six functions in the room for the evaluation. It takes longer. It produces better decisions.

Common evaluation failures

The five most expensive evaluation mistakes

Every one of these failures is preventable. I have seen all five cost teams six figures or more in wasted investment.

  • Evaluating AI opportunities in isolation instead of as a portfolio. The second-best opportunity may enable the first. The best opportunity may cannibalize the third. Evaluate as a set.
  • Overestimating AI capability maturity. The demo looks perfect. Production is messier. Assume the AI will be 70 percent as good as the demo for the first six months.
  • Underestimating evaluation pipeline cost. Building the AI feature is 60 percent of the work. Building the system that checks whether the AI is working correctly is the other 40 percent.
  • Ignoring adoption risk. Customers do not care about AI. They care about outcomes. If the AI feature delivers a clear outcome, they adopt. If it delivers AI for the sake of AI, they ignore it.
  • Confusing activity with outcomes. Shipping an AI feature is not success. Moving a business metric is success. Evaluate opportunities against metrics, not milestones.

The evaluation-to-execution handoff

The most dangerous moment in the evaluation process is the handoff to execution. The evaluation team makes a decision. The execution team receives a project. Information gets lost. Assumptions get buried. The fix is a one-page evaluation summary that travels with the project. It contains the opportunity size estimate, the key assumptions, the risk scores, the build-vs-buy rationale, and the success criteria. When the execution team hits a challenge, they consult the summary. When the assumptions prove wrong, they update the summary. The summary is a living document, not a one-time artifact. The teams that skip this step rediscover the evaluation's assumptions through expensive mistakes.

Getting started: your first AI evaluation

Start with one AI opportunity that matters to your business. Do not evaluate 20 opportunities. Evaluate one well. Score its vulnerability, size it with real data, and make a build-vs-buy decision. Write the one-page summary. Set a re-evaluation date 90 days out. If the evaluation recommends building, build the smallest version that tests the riskiest assumption. If the evaluation recommends buying, run a 30-day trial and measure adoption. If the evaluation recommends waiting, document the trigger that would change the recommendation. The evaluation-first approach is not about avoiding action. It is about taking action that is informed by evidence instead of enthusiasm. Start this week.


Enjoyed this framework? Get more research and practical notes in your inbox.

Subscribe to Build Notes →