Engineer monitoring server infrastructure on a laptop in a data center

How to Build a Decision Model

October 11, 2026 · 12 min read · By Rafael

Key Takeaways:

  • AI-related incidents now account for more than one in 10 reported technology outages, roughly six times the 2023 rate, according to StackGen’s analysis of nearly 178,000 public incidents.
  • About 70% of enterprise AI initiatives fail, and 54% of C-suite executives say adopting AI is not working, per HPCwire’s September 2026 analysis.
  • The average number of AI agents in production per organization grew from five in February 2025 to 13 by April 2026, according to Salesforce’s Agentic Enterprise Index.
  • 47% of organizations with formal AI governance policies admit they have bypassed those policies for urgent deployments, per EY’s survey of 202 executives at $1B+ companies.
  • A decision model is only as good as its inputs; the failure mode is usually process and governance, not the model itself.

AI-related incidents now account for more than one in 10 reported technology outages, roughly six times the 2023 rate, according to StackGen’s analysis of nearly 178,000 public technology incidents. When a tenth of your outages come from the systems making decisions, the criteria those systems use become an operational risk, not just a modeling preference.

Off-the-shelf tools are not inadequate. The decisions that matter most to a specific organization (such as which vendor to sign with, which workload to move to which cloud, or which alert to escalate) depend on weights and thresholds no vendor can provide. This article explains how to build a decision model, where the build-versus-buy calculation stands, and the governance layer that prevents it from drifting.

Why Enterprises Are Building Their Own Decision Models

About 70% of enterprise AI initiatives fail, and 54% of C-suite executives say adopting AI is not just ineffective but actively disappointing, according to HPCwire’s September 2026 analysis. The model itself is rarely the problem. The issue lies in the surrounding decision process: who decides, based on what evidence, and with what threshold for moving forward.

Defining Criteria Before Choosing a Tool

Healthcare illustrates this clearly. Health systems run many successful AI pilots; what they lack is a repeatable method to turn those pilots into enterprise value, as Becker’s Hospital Review explains in its coverage of the AI value gap. A pilot that works in one department rarely applies elsewhere because the decision criteria that made it successful (the specific workflow, data, and accountability chain) were never documented.

At the same time, deployment scale keeps increasing. The average number of AI agents activated per organization nearly tripled, from five in February 2025 to 13 by April 2026, and the time to create a new agent dropped 53% to an average of 1.9 days, according to Salesforce’s Agentic Enterprise Index. More agents making more decisions faster create conditions where an undocumented decision model becomes a liability.

Defining Criteria Before Choosing a Tool

A decision model is a scoring function. You identify the factors, assign each a weight, specify the evidence that satisfies each factor, and set a threshold above which you proceed. Determining those factors is organizational work, not technical.

Four categories cover most enterprise technology decisions. Cost includes total cost of ownership across the full lifecycle, not just the license fee. Control covers customization, intellectual property, and vendor lock-in. Compliance covers data residency, sovereignty, and auditability. Time-to-value covers how quickly the option produces a measurable result. Each category receives a weight reflecting your organization’s actual priorities, which vary significantly by industry. For example, a regulated bank will assign much higher weight to compliance than a consumer retailer.

The build-versus-buy question itself shows why generic answers fail. An Omdia report on build-versus-buy dynamics for enterprise AI found that 95% of 376 technical and business stakeholders surveyed agreed that building AI offers greater customization and control, while 91% acknowledged the speed and benefits of prebuilt platforms, as reported by TechTarget’s LLM decision framework. Both numbers are high because both are true. A decision model makes the trade-off explicit instead of leaving it to whoever argues loudest.

A systematic literature review of 1,697 records found that generative AI consistently supports design ideation and trade-off analysis, rapid artifact creation, and architectural decision support, while introducing risks including opacity, contextually incorrect outputs that cause rework, and privacy and compliance concerns, according to the 2025 review on arXiv. The tool helps you analyze trade-offs. It does not decide which trade-offs matter.

A Six-Step Process for Building the Model

Teams that fail often skip steps two and five: they collect inputs without verifying them, and they never set a threshold, so every option gets a “yes, but let’s pilot it.”

A Six-Step Process for Building the Model
A Six-Step Process for Building the Model, architecture diagram

1. Define decision criteria. For each decision type, specify the owner, the threshold, and the evidence required. For example, a cloud migration decision might require a documented TCO comparison, a data residency assessment, and a named accountable executive before it can be scored.

2. Collect verified inputs. Every input needs a source. If your cost figure comes from a vendor’s marketing page, label it as a vendor claim. If your uptime figure comes from a status page, note the measurement window. Unverified inputs are the most common reason a model produces confident errors.

3. Set weights. Score each factor from 1 to 5 and normalize. Do this with the people who will live with the consequences, because the weights encode organizational priorities.

4. Score options. Apply the weights to each candidate. Keep the raw inputs visible so a reviewer can challenge a single number without redoing the entire analysis.

5. Set the go/no-go line. Define the minimum score required to proceed. Without this, the model becomes a discussion aid rather than a decision tool.

6. Run the review loop. Re-score quarterly. Log the outcome of each decision so the model can be calibrated against actual results rather than intuition.

The pre-execution layer is as important as the scoring. Research on the “Right-to-Act” protocol proposes a deterministic, non-compensatory decision layer that evaluates whether an AI-generated decision is allowed to be executed, halting or deferring execution if any required condition is unmet, rather than letting a high-confidence signal override a failed condition, as described in the 2026 arXiv paper. Some conditions are gates, not points. For example, a data residency violation should not be offset by a favorable cost score.

Build vs. Buy: Where the Numbers Land

The choice between building a decision model and buying one depends on how much of your decision logic is proprietary. The table below matches common situations to the approach that fits, with the trade-offs to manage.

Situation Approach Why it fits Main trade-off
Commodity decisions with public benchmarks (e.g., standard cloud pricing comparisons) Buy or use a vendor tool Inputs are public and the criteria are stable across organizations Little differentiation; the tool cannot encode your specific risk tolerance
Decisions that turn on proprietary operational data Build a custom model Your data is the source of the advantage; a generic tool cannot see it Requires ongoing data pipeline and maintenance ownership
Regulated decisions with audit requirements Build with a documented, explainable layer Every weight and threshold must be defensible to a regulator Slower to change; governance overhead on every update
Fast-moving pilot decisions Start with a lightweight model, formalize later Speed matters more than precision early Risk of the criteria never being documented

Two data points illustrate the build side. Building an enterprise-class LLM can take six months to two years to deploy, and cloud infrastructure for deployment can cost from $1,000 to more than $50,000 per month across the lifecycle depending on sophistication, according to TechTarget’s build-versus-buy framework. Those figures apply to model building; a decision model that scores options rather than generating text costs far less, but the maintenance effort is similar.

Sovereignty changes the calculation for some buyers. Cohere and PwC announced a global alliance in October 2026 pairing Cohere’s enterprise AI platform with PwC’s transformation practice, with support for private-cloud, on-premises, and air-gapped deployments that give organizations more control over sensitive data and infrastructure, per the joint announcement. That is the vendor’s own description, and the trade-off is real: a fully controlled deployment costs more to operate and depends on one vendor’s platform.

Rows of server racks in a modern enterprise data center
Where the model runs is itself a decision the model should score: cost, control, and residency rarely lead to the same answer.

Governance and the Pre-Execution Gate

Governance is where the 70% failure rate occurs. EY’s survey of 202 senior AI executives at organizations generating at least $1 billion in annual revenue found that 98% report having formal AI governance policies in place, yet 47% admit their organization has previously bypassed its AI governance process for urgent deployments, according to the EY AI Risk and Governance Survey.

The same survey found that 36% of organizations have experienced an AI incident or failure causing materially negative impact, including data loss, financial damage, brand damage, and operational disruption. Among organizations using agentic AI, 26% cannot detect unauthorized AI agents operating internally, and 49% say their governance framework has not been updated to include agentic AI requirements and risks.

These figures support treating the decision model itself as a governed artifact. Every weight, threshold, and input source should be versioned and reviewable. When a decision goes wrong, the first question is whether the model was followed, and the second is whether the model was correct. Without versioning, you cannot answer either.

Among those conducting formal AI assurance reviews, 64% significantly modified a quarter or more of their AI systems, 29% paused a quarter or more, and 25% fully stopped a quarter or more. The most common issues found were data quality problems (57%), AI model drift (48%), and shadow AI (39%). A decision model with a quarterly review loop detects drift before an assurance review becomes necessary.

Scaling Without Losing the Thread

Scaling a decision model means applying it to more decisions without weakening the criteria that made it effective. The Salesforce index provides a benchmark: agents now handle four unique actions on average, up from two in 2025, and the share of customer service sessions handled autonomously by AI agents has reached seven in ten, with over five million conversations at help.salesforce.com handled by AI agents versus 2.4 million by humans as of August 2026.

These are vendor-reported figures from Salesforce’s own platform data, so treat them as indicative rather than universal. The pattern they describe (agents taking on more scope over time) makes a written decision model necessary. When an agent handles a decision that used to require a human, the criteria that human used must be encoded, or decision quality silently declines.

Personal AI points to the next layer. Liquid AI’s approach to on-device agents focuses on keeping context within fixed hardware limits, using device signals to build understanding of the user and continuous improvement loops that refine both model and harness over time, according to SiliconANGLE’s coverage of the Fully Connected event. The company’s COO, Jeffrey Li, explained the constraint clearly: “The problem with devices is that you have fixed compute. You have to fit within zero-sum compute.” Decision models face the same constraint in a different form. Every factor you add demands attention, and every threshold you relax reduces precision.

For teams already running observability and security tools, the decision model connects directly to existing work. Our analysis of OpenTelemetry-based observability explains how telemetry feeds operational decisions, and the SIEM setup guide walks through alert-tuning metrics such as detection rate, false positive rate, time to respond, and dwell time that a decision model can use as inputs. The GPU pricing forecast shows how fast the cost inputs to a compute decision change, which is why the review loop matters more than the initial scoring.

Engineer monitoring server infrastructure on a laptop in a data center
A decision model without a review loop drifts out of alignment with the environment it was built to describe.

Frequently Asked Questions

What is a decision model in an enterprise AI context?

A decision model is a documented scoring function that identifies the factors relevant to a decision, assigns each a weight, specifies the evidence required, and sets a threshold for proceeding. It turns a debate into a repeatable process. It differs from the AI model itself; the AI model generates options or predictions, while the decision model determines which options to act on.

How much does it cost to build a custom decision model?

Costs vary widely by scope. Building an enterprise-class LLM can take six months to two years, and cloud infrastructure for deployment can cost from $1,000 to more than $50,000 per month across the lifecycle, according to TechTarget’s build-versus-buy framework. A decision model that scores options rather than generating text costs far less to build, but the ongoing maintenance and governance effort is similar.

Should I build or buy a decision model?

Build when the decision depends on proprietary operational data or when a regulator requires an explainable, auditable trail. Buy when the inputs are public and the criteria are stable across organizations. Many teams use both: a purchased tool for commodity decisions and a custom model for the decisions that carry the most risk.

How often should a decision model be reviewed?

Quarterly is a reasonable baseline, with immediate re-scoring triggered by significant changes in inputs such as a major price shift, a new regulatory requirement, or an AI incident. Organizations conducting formal AI assurance reviews often significantly modify a quarter or more of their AI systems, which suggests annual review is too infrequent.

What is the biggest mistake teams make when building one?

Skipping the threshold. Without a defined go/no-go line, the model becomes a discussion aid and every option gets a pilot. The second most common mistake is using unverified inputs, where a cost or uptime figure has no traceable source.

It addresses one specific failure mode: decisions made without documented criteria. AI-related incidents now account for more than one in 10 reported outages, and 36% of organizations report an AI incident with materially negative impact. A decision model with a review loop detects drift and data quality problems before they compound, but it does not prevent model failures or security incidents on its own.

For related reading, see our guide to capex and opex for infrastructure decisions and our analysis of hyperscaler AI infrastructure spending, both of which feed directly into the cost and control factors a decision model scores.

More in-depth coverage from this blog on closely related topics:

Sources and References

Sources cited while researching and writing this article:

Rafael

Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...