We’re putting $5.5M behind research and development of smaller, specialized AI models for real-world deployment — Read our manifesto →
    All posts

    Build vs Buy Is the Wrong Question for Enterprise AI

    The evidence from Morgan Stanley, Uber, Walmart, Intuit, JPMorganChase and BBVA points to a better strategy: buy commodity capability, then commission the intelligence your business cannot buy off the shelf.

    Conscious Engines

    The evidence from Morgan Stanley, Uber, Walmart, Intuit, JPMorganChase and BBVA points to a better strategy: buy commodity capability, then commission the intelligence your business cannot buy off the shelf.

    Executive answer

    Most enterprise AI committees are given two choices:

    1. Buy an AI product from a software vendor.
    2. Build an AI system with an internal team.

    That framing is incomplete. A third option is often better for high-value workflows: commission a specialist to build a bespoke system that the enterprise controls.

    The enterprise does not need to train a frontier foundation model from zero. It can buy model access, use an open-weight model, or combine several models. What it should selectively own is the layer that converts general intelligence into an operating capability:

    • the enterprise data model and permission logic;
    • the retrieval system and source hierarchy;
    • the task-specific model, adapter or classifier;
    • the speech and language layer for domain vocabulary;
    • the agent's tools, authority and approval rules;
    • the evaluation set that defines acceptable performance;
    • the integrations with systems of record;
    • the telemetry, audit trail, fallback and model-switching layer.

    This is not theoretical. Morgan Stanley combined a third-party foundation model with private retrieval, evaluations and adviser workflows. Uber built a multi-model gateway around vendor and internally hosted models. Walmart developed retail-specific models using decades of proprietary data. Intuit built financial models and a runtime that selects models and data sources. JPMorganChase put leading models inside a controlled enterprise layer. BBVA bought a horizontal platform, then created thousands of workflow-specific assistants.

    The pattern is consistent: buying a model is not the same as buying an outcome.

    The three decisions that executives should separate

    An enterprise AI program actually contains three different sourcing decisions.

    DecisionTypical optionsWhat matters
    Foundation intelligenceCommercial API, managed cloud model, open-weight model, internally trained base modelGeneral reasoning quality, modality, context, price, latency, deployment rights
    Enterprise intelligencePrompting, RAG, fine-tuning, task-specific model, speech model, policy engine, model routerProprietary language, current knowledge, accuracy, control, cost per accepted outcome
    Production systemPackaged SaaS, internal application, bespoke partner-built systemWorkflow integration, permissions, adoption, auditability, monitoring, support and ownership

    A company can buy the first layer, commission the second and third layers, and still own the resulting operating capability. That is the strategy this article calls selective ownership.

    Selective ownership avoids two expensive extremes:

    • The black-box extreme: purchase a generic application and accept its workflow, data model, roadmap, model choices and control boundaries.
    • The research-lab extreme: recruit a large permanent team to reproduce foundation-model infrastructure that specialist vendors and open-source communities already provide.

    The objective is not maximum ownership. It is ownership of the parts that determine business value, risk and switching power.

    What the market data actually says

    The build-versus-buy market has moved quickly enough that one year's conclusion can look like the opposite of the next year's conclusion.

    Menlo Ventures' 2024 enterprise report found a near-even split: 47% of surveyed solutions were developed in-house and 53% were sourced from vendors. More important than the split was the buying criterion. Respondents prioritized measurable value and industry-specific customization. Only 1% named price as a primary selection concern. Menlo also reported implementation costs in 26% of failed pilots, data privacy in 21%, disappointing ROI in 18% and hallucinations in 15%.

    Those percentages are survey responses, not audited failure rates. They still expose a useful problem: the purchase price visible in procurement is not the same as the cost of making AI work.

    Andreessen Horowitz saw enterprises building more applications in 2024, then reported a shift toward buying mature third-party applications in 2025. In its survey of 100 CIOs across 15 industries, 37% used five or more models, up from 29% in the prior survey. More than 90% were testing third-party customer-support applications. The report also noted that regulated and high-risk sectors such as healthcare remained an exception to the broad shift toward packaged applications.

    This is not a contradiction. It is market segmentation:

    • Buy when the workflow is common and a mature product is already better than what an internal team can economically maintain.
    • Customize when a horizontal platform can be configured without compromising the workflow.
    • Commission bespoke development when the use case depends on private context, unusual inputs, regulated decisions, proprietary process or deep integration.
    • Build internally when AI itself is a permanent core product capability and the enterprise can support the full lifecycle.

    The strongest evidence against a one-size-fits-all purchase comes from McKinsey's 2025 global AI survey. It found that 88% of respondents said their organizations regularly used AI, but only about one-third said they had begun scaling AI across the enterprise. AI high performers were nearly three times as likely to redesign workflows. Access to a model was common. Operational redesign was scarce.

    That is precisely where bespoke engineering earns its place.

    Six public company cases

    The following cases are not presented as controlled experiments. They are public, company-reported deployments and should be read as evidence of feasibility and architecture, not as guaranteed outcomes. Together, however, they show what serious enterprise AI programs build after they obtain model access.

    1. Morgan Stanley: buy the foundation, build trust around the task

    Morgan Stanley did not attempt to recreate GPT-4. It worked with OpenAI to embed the model into a wealth-management system designed around the firm's intellectual capital, controls and adviser workflow.

    According to the public Morgan Stanley case study, the system included:

    • retrieval across a corpus that grew to 100,000 documents;
    • evaluation datasets graded by advisers and prompt engineers;
    • daily regression testing;
    • retrieval-method refinement;
    • human review for meeting notes and follow-up drafts;
    • zero-data-retention arrangements;
    • CRM integration for adviser debrief outputs.

    The reported results were unusually concrete:

    • more than 98% of adviser teams actively used the Assistant;
    • access to documents rose from 20% to 80%;
    • the retrievable knowledge scope moved from answering about 7,000 questions to operating over 100,000 documents;
    • follow-ups that had taken days could happen within hours.

    Jeff McMillan, Morgan Stanley's Head of Firmwide AI, described the strategic edge as “new products and services that only people close to the problem can imagine.”

    The lesson is not that every enterprise should buy GPT-4. The lesson is that the valuable asset was the system around the model: proprietary content, expert evaluations, compliant retrieval, workflow integration and user trust.

    A generic adviser chatbot could reproduce the interface. It could not reproduce Morgan Stanley's permissioned knowledge, acceptance criteria, adviser feedback loop or operational integration.

    2. Uber: one control plane, many models

    Uber's public architecture rejects the assumption that an enterprise should choose one winning model vendor.

    Uber reported more than 60 generative AI use cases. It built a GenAI Gateway that provides a consistent interface to external models, including commercial providers, and Uber-hosted models. The gateway added controls that a raw model endpoint did not provide:

    • authentication and authorization;
    • personally identifiable information redaction and restoration;
    • safety and policy guardrails;
    • cost attribution and usage alerts;
    • audit logging;
    • quality evaluation;
    • a common interface across programming languages and model providers.

    In July 2024, Uber reported that the gateway served close to 30 internal teams, handled 16 million queries per month and reached a peak of 25 queries per second.

    Uber also explained why it wanted both model classes. External models were useful for broad knowledge and difficult reasoning. Internally hosted open models could be fine-tuned on proprietary data for Uber-specific tasks “at a fraction of the cost and lower latency.” Its wider Michelangelo platform was already managing about 400 active ML projects, more than 20,000 training jobs per month and more than 5,000 production models, according to Uber's engineering account.

    Few enterprises need Uber's internal platform scale. Many need the same design principle: a modular control layer that can route work, protect data, record cost and replace models without rewriting every application.

    This is an important distinction for procurement. The enterprise can buy inference from several vendors without letting any one vendor become the architecture.

    3. Walmart: proprietary retail context is the product

    Walmart announced Wallaby, a family of retail-specific language models trained with decades of Walmart data. The company said it would combine Wallaby with other LLMs rather than treat its proprietary model as a universal replacement.

    The Walmart announcement described customer-facing use cases and a support assistant that could recognize the customer, understand intent and take actions such as finding orders and managing returns. Walmart also described a common global platform strategy spanning Walmart US, Sam's Club and international operations.

    Suresh Kumar, then Walmart's global CTO and chief development officer, framed the model as common capabilities “built once and deployed across Walmart U.S., Sam's Club and Walmart International.”

    Why not buy a standard retail chatbot? Because the key capability is not fluent conversation. It is the combination of:

    • a product and catalog ontology;
    • inventory and fulfillment state;
    • customer identity and permissions;
    • returns and service policies;
    • merchandising logic;
    • Walmart-specific language and values;
    • action APIs connected to real orders.

    The public evidence does not disclose audited ROI for Wallaby. It does show that one of the world's largest retailers considered its data and retail context important enough to create a specialized model family while retaining access to general models.

    4. Intuit: custom models plus a model-selection runtime

    Intuit's GenOS is another rejection of the single-model thesis. The company announced custom-trained financial LLMs for tax, accounting, cash flow, marketing and personal finance. It also built GenRuntime, a layer that chooses the model and data access points for each request.

    At the time of its GenOS announcement, Intuit reported:

    • 400,000 customer and financial attributes per small business;
    • 55,000 tax and financial attributes per consumer;
    • connections to more than 24,000 financial institutions;
    • more than 730 million AI-driven customer interactions per year;
    • 58 billion machine-learning predictions per day;
    • more than 100 million customers served across its products.

    These are company-reported platform-scale figures, not measures of incremental GenOS ROI. They explain why bespoke intelligence can matter. The company's data structure, financial attributes, expert network and customer context are not available inside a generic model.

    CEO Sasan Goodarzi stated the thesis directly: “The depth of our customer data, along with our proprietary GenOS platform, create a competitive advantage.”

    For most enterprises, the correct imitation is not to train multiple foundation models. It is to make their own data, policies and expertise usable by the best-fit models through a controlled runtime.

    5. JPMorganChase: a controlled enterprise layer at workforce scale

    JPMorganChase disclosed in its 2024 annual-report material that it deployed LLM Suite to more than 200,000 colleagues. It described the system as a controlled desktop environment for accessing leading generative AI capabilities while protecting company and customer data.

    The bank also reported more than 60,000 technologists, more than 6,000 applications and nearly an exabyte of data. Those numbers make a fully internal platform plausible for JPMorganChase. They do not make it the default choice for an enterprise with a 50-person technology organization.

    The transferable lesson is narrower: once AI reaches sensitive work and thousands of employees, the organization needs shared controls and reusable capabilities. It cannot safely manage each use case as an employee's direct account with a model vendor.

    For enterprises without JPMorganChase's engineering scale, commissioning that controlled layer can be more rational than trying to hire and retain every required specialist.

    6. BBVA: buying the platform did not eliminate building

    BBVA is a strong example of buying a horizontal enterprise product and then creating task-specific systems on top.

    According to BBVA's public case study, the bank reported:

    • about three hours saved per employee per week;
    • 83% weekly active usage;
    • more than 20,000 custom GPTs created;
    • around 4,000 custom GPTs used frequently;
    • efficiency improvements above 80% in selected workflow tests.

    One assistant used by more than 3,000 employees in Peru reportedly reduced average query handling time from about 7.5 minutes to about one minute. The rollout also included security, legal and compliance participation and training for 250 senior leaders.

    Antonio Bravo, BBVA's Global Head of Data and AI, said AI had to become “part of our business strategy, not a tech effort that sits on the side.”

    The case supports buying secure horizontal access for general employee productivity. It also demonstrates that value quickly fragments into thousands of specialized workflows. A license created the safe experimentation surface. Business-specific assistants created repeatable operating value.

    What all six companies built that a model license did not provide

    Across the cases, the custom layer repeatedly contained the same assets.

    Bespoke assetMorgan StanleyUberWalmartIntuitJPMorganChaseBBVA
    Private knowledge or proprietary dataYesYesYesYesYesYes
    Enterprise permissions and controlsYesYesImpliedYesYesYes
    Workflow-specific interface or actionsYesYesYesYesYesYes
    Evaluation or monitoringYesYesNot publicly detailedKnowledge checksNot publicly detailedWorkflow tests
    Multi-model or model-routing designNot publicly detailedYesYesYesLeading modelsPlatform plus custom GPTs
    Domain-specific modelsRetrieval specializationUber-hosted fine-tunesRetail LLMsFinancial LLMsNot publicly detailedCustom assistants

    The strongest common denominator is not a proprietary base model. It is an enterprise-specific system.

    That is the category Conscious Engines is positioned to build.

    When buying is the correct decision

    A persuasive bespoke-AI company should be willing to say when a customer should not commission custom development.

    Buy a mature product when most of the following are true:

    • The workflow is nearly identical across companies.
    • The required integrations already exist and are maintained by the vendor.
    • The product has credible references in the same regulatory environment.
    • The enterprise's data adds limited differentiation.
    • Standard permissions and audit features satisfy policy.
    • Users can adopt the vendor's workflow without costly process change.
    • The application is not a material source of competitive advantage.
    • Switching costs are understood and acceptable.
    • The vendor exposes enough outcome telemetry to evaluate value.

    Common examples include general meeting transcription, basic copy drafting, standard coding assistance, commodity OCR and broad employee chat. A custom build in these categories can become an expensive copy of a product that a dedicated vendor improves every week.

    Buying is also appropriate as a discovery tool. Enterprise licenses can reveal where employees repeatedly create custom instructions, upload the same documents or perform the same manual checks. Those patterns become evidence for later bespoke workflow investment.

    When a bespoke system is economically stronger

    When a bespoke system is economically stronger

    The process is a source of differentiation

    If the workflow encodes how the company prices risk, schedules operations, serves clients, maintains equipment, reviews claims or allocates capital, forcing it into a generic...

    Proprietary data changes the answer

    Private documents alone do not justify a custom model.

    The input is operationally difficult

    Generic systems are often optimized for clean typed text.

    The AI must take controlled action

    An agent that only drafts text is relatively easy to buy.

    Quality must be defined locally

    Generic benchmarks do not measure whether a claims assistant applies the correct endorsement, whether a maintenance copilot retrieves the current lockout procedure or whether...

    Volume or latency changes the model economics

    A large model may be inexpensive during a 100-user pilot and uneconomic across millions of narrow transactions.

    Commission a bespoke enterprise AI system when several of these conditions are present.

    1. The process is a source of differentiation

    If the workflow encodes how the company prices risk, schedules operations, serves clients, maintains equipment, reviews claims or allocates capital, forcing it into a generic product can erase the difference the business is trying to preserve.

    2. Proprietary data changes the answer

    Private documents alone do not justify a custom model. The stronger case appears when private data has a unique structure, terminology, hierarchy, temporal meaning or permission model. Examples include clinical abbreviations, asset telemetry, policy endorsements, legal matter boundaries, fuel ledgers and maintenance histories.

    3. The input is operationally difficult

    Generic systems are often optimized for clean typed text. Enterprise reality includes accented speech, code-switching, radio audio, machine noise, scanned forms, tables, handwriting, sensor events and incomplete identifiers. Custom speech-to-text, document processing and entity resolution may determine whether the language model receives usable evidence at all.

    4. The AI must take controlled action

    An agent that only drafts text is relatively easy to buy. An agent that updates a work order, books a shipment, changes an appointment, initiates a payment or writes to a clinical system needs least-privilege tools, deterministic checks, approvals and rollback.

    OWASP's guidance on excessive agency identifies excessive functionality, permissions and autonomy as core causes of damaging agent behavior. Those controls must be designed around the enterprise's systems and authority model. They cannot be inherited from a general chatbot.

    5. Quality must be defined locally

    Generic benchmarks do not measure whether a claims assistant applies the correct endorsement, whether a maintenance copilot retrieves the current lockout procedure or whether an ASR model recognizes a site-specific asset code.

    The enterprise needs its own evaluation set. That set becomes the specification, procurement test, regression suite and feedback asset. Read the related guide on why an enterprise evaluation set becomes an AI moat.

    6. Volume or latency changes the model economics

    A large model may be inexpensive during a 100-user pilot and uneconomic across millions of narrow transactions. A smaller task-specific model, deterministic classifier or retrieval-first pipeline can reduce latency and cost if it meets the same acceptance threshold.

    The key metric is not price per token. It is cost per accepted outcome:

    cost per accepted outcome = total run cost / outputs accepted without rework or harm

    A cheap model with low acceptance can be more expensive than a premium model. A bespoke small model with high acceptance on one stable task can be far cheaper than routing every request to a frontier model. See the evidence behind why the right model is rarely the biggest.

    7. Portability has strategic value

    IBM's 2026 survey of 1,000 senior executives found that 71% said switching their primary AI model or vendor would be difficult. Only 9% reported an excellent understanding of their AI dependencies. Among those that had switched or attempted to switch, 75% described difficulty related to portability, revalidation, compliance and technical lock-in. These are vendor-sponsored survey findings, but they identify real contract and architecture questions. IBM calls the response selective AI sovereignty: control where risk and economics justify it, not everywhere.

    A bespoke system can reduce lock-in if prompts, retrieval, tool contracts, evaluations and observability are separated from the model provider. Custom does not automatically mean portable. Portability must be designed and contracted.

    When an internal build is justified

    Building internally is strongest when the enterprise has:

    • a permanent AI product mandate rather than a small set of projects;
    • enough recurring demand to keep a platform and model team fully utilized;
    • strong data engineering, ML engineering, security and product management;
    • the ability to recruit and retain specialist talent;
    • a mature software-delivery and on-call culture;
    • a reason to own the complete implementation rather than its outputs and artifacts;
    • a multiyear budget for evaluation, maintenance, model migrations and support.

    Uber and JPMorganChase are useful reference points because their public disclosures show the scale behind internal platforms. Uber reported thousands of production models. JPMorganChase reported 60,000 technologists. An enterprise should not copy the organizational form without the workload that makes it efficient.

    Commissioning a specialist fills the gap between packaged software and a permanent internal AI platform. The enterprise keeps domain ownership and governance. The specialist supplies concentrated technical execution.

    A weighted decision scorecard

    Use the following scorecard at the use-case level. Score each factor from 0 to 3, where 0 means absent and 3 means critical. Multiply by the suggested weight.

    FactorWeightFavors packaged buyFavors bespoke commissionFavors internal build
    Workflow differentiation5LowHighVery high and strategic
    Proprietary data advantage5LowHighVery high and reusable
    Regulatory or safety exposure5Standard controls fitCustom controls neededCore risk capability
    Integration complexity4Standard connectorsSeveral custom systemsPlatform-wide estate
    Task volume4Low to moderateModerate to highVery high across many teams
    Latency or edge requirement4Standard cloud worksSpecial deployment neededEnterprise platform need
    Model and vendor portability3Low concernContractual and architectural needStrategic platform mandate
    Time to first value4ImmediateWeeks to monthsMonths to years
    Internal AI operating maturity5LowLow to moderateHigh
    Product maturity in market5Strong existing categoryCategory gapCategory gap plus permanent team

    Do not total the columns mechanically and call the highest number truth. Use the scores to expose disagreements. A compliance leader may rate risk as critical while the project sponsor sees a simple chatbot. That disagreement must be resolved before procurement.

    The procurement test most enterprises miss

    Run a competitive bake-off against the enterprise task, not a vendor demonstration.

    Step 1: define the operating outcome

    Examples:

    • reduce average documentation time without increasing correction rate;
    • resolve shipment exceptions while respecting financial approval limits;
    • retrieve the correct controlled procedure with source and effective date;
    • classify claims at a target recall while preserving manual review for severe errors.

    Step 2: create the evaluation set

    Build a representative, permission-cleared and time-separated sample. Include ordinary cases, difficult cases and severe failure cases. Define acceptance before seeing model output.

    Step 3: compare three real alternatives

    • the incumbent manual or rules-based process;
    • the best credible packaged product;
    • a bespoke prototype using the simplest viable architecture.

    If the internal team is a serious option, include it under the same timeline and acceptance criteria.

    Step 4: measure the full system

    Evaluate input quality, retrieval, answer faithfulness, tool selection, action correctness, latency, human review, uptime and cost. A model benchmark cannot substitute for a workflow benchmark.

    Step 5: calculate quality-adjusted economics

    Include license or inference, integration, data preparation, review labor, error remediation, security, support, model migration and change management. The companion article provides a detailed bespoke enterprise AI business-case model.

    Step 6: negotiate ownership and exit before scaling

    The contract should say who owns:

    • cleaned and labeled datasets;
    • evaluation cases and expert rubrics;
    • prompts, retrieval configuration and policy rules;
    • fine-tuned adapters or task-specific weights;
    • connectors and tool schemas;
    • monitoring data and incident records;
    • deployment scripts and documentation;
    • rights to move the system to another model or environment.

    Paying for bespoke development without securing the operational assets can create a custom form of lock-in.

    Common objections

    “Foundation models improve so quickly that customization will become obsolete.”

    Better models reduce the amount of customization needed for some tasks. They do not create the enterprise's permissions, current source hierarchy, tool contracts, action approvals, local vocabulary or acceptance criteria. A modular bespoke system benefits from model improvement because its model layer can be upgraded.

    “Our SaaS vendors will add AI to every system.”

    They will, and enterprises should use those features when the workflow is native to that product. Cross-system workflows remain difficult. A fuel exception may span procurement, telematics, finance and site operations. A patient conversation may span audio, clinical vocabulary, consent, documentation and an electronic record. No single incumbent necessarily owns the end-to-end outcome.

    “Custom development is too expensive.”

    It can be. The correct comparison is not custom cost versus a small monthly license. It is total cost at the required quality and scale. A16z's 2024 enterprise analysis reported one executive estimating that LLMs were roughly one-quarter of the cost of building use cases. Data, implementation and operations dominate many projects regardless of whether the starting model is bought.

    The bespoke case is strongest when it removes material labor, loss, delay or risk and becomes reusable. It is weak when it recreates commodity software.

    “We do not want a services project.”

    Neither should the goal be an endless services project. A responsible bespoke engagement should produce a versioned product, measurable acceptance criteria, reusable components, operating documentation and a transition path. The work is custom because the system fits the enterprise, not because delivery remains permanently manual.

    A practical commissioning model

    The highest-confidence approach has four stages.

    Stage 1: workflow and evidence discovery

    Map the current process, systems, permissions, error costs and baseline metrics. Identify where intelligence is actually needed and where deterministic software is safer.

    Stage 2: architecture bake-off

    Compare prompting, RAG, fine-tuning, task-specific models and frontier models against the same evaluation set. The technical options are explained in RAG, fine-tuning or a bespoke model.

    Stage 3: controlled production pilot

    Integrate with real systems under constrained permissions. Shadow traffic, measure human acceptance and rehearse fallback. A strong pilot proves the operating loop, not only a demo response.

    Stage 4: scale and transfer

    Harden the system, document it, train users, define support, establish model-refresh policy and transfer agreed artifacts. The enterprise should leave the engagement with more internal capability and more control than it had at the start.

    The strategic conclusion

    The build-versus-buy question treats AI as a single item. Enterprise AI is a stack of decisions.

    The public evidence from Morgan Stanley, Uber, Walmart, Intuit, JPMorganChase and BBVA shows that serious deployments rarely stop at buying access to a model or application. The companies created proprietary retrieval, evaluations, controls, domain models, model routers, shared platforms and workflow integrations.

    Most enterprises should not attempt to become foundation-model laboratories. They should become intelligent owners of their business-specific AI layer.

    That leads to a simple allocation rule:

    • Buy the commodity.
    • Commission the differentiator.
    • Own the evidence, controls and operating assets.

    Research note

    This article prioritizes first-party company engineering posts, corporate disclosures and attributed customer cases, supplemented by enterprise surveys. Company metrics are self-reported unless stated otherwise. Survey categories use different definitions of build and buy. No public case establishes a universal ROI or break-even point. Evidence reviewed through September 5, 2026.

    Building a Production-Ready System

    Conscious Engines builds bespoke AI systems for enterprise workflows that do not fit a generic product. The model stack can combine domain speech-to-text, text-to-speech, voice agents, enterprise RAG, task-specific small language models and controlled agents. We benchmark the simplest viable architecture against the enterprise's own evaluation set, integrate it with the systems where work happens, and design the control and portability layer around the customer's risk.

    The enterprise does not pay us to build “an AI.” It commissions a measurable operating capability that it cannot buy off the shelf.