We’re putting $5.5M behind research and development of smaller, specialized AI models for real-world deployment — Read our manifesto →
    All posts

    Sovereign by Design: The Enterprise AI Model Stack for Government

    A practical guide to multilingual speech, citizen-service agents, policy RAG, casework automation, document intelligence, and controllable models for public institutions.

    Conscious Engines

    Government AI must work under a higher standard than ordinary productivity software. A wrong answer may affect a benefit, license, legal right, safety decision, or public record. The system must also serve people across languages, channels, abilities, and levels of digital access.

    The right strategy is not one national chatbot. It is a governed portfolio of task-specific models connected to authoritative public information and accountable workflows.

    Adoption is real, but evidence is uneven

    The OECD reviewed 200 real-world government AI examples across 11 core government functions. It found prominent activity in public services, civic participation, and justice. In the United States, the Government Accountability Office found that reported AI use cases at selected federal agencies increased from 571 in 2023 to 1,110 in 2024. Reported generative-AI cases rose from 32 to 282, or roughly ninefold.

    GAO also found that 10 of 12 selected agencies cited policy or privacy obstacles. Scale has arrived before institutional readiness is complete.

    Productivity evidence should be read carefully. A UK cross-government trial covering about 20,000 users reported 26 minutes saved per person per day, equivalent to nearly two weeks a year. A later Department for Work and Pensions evaluation of 3,549 licensed staff estimated 19 minutes saved per day, a 0.56-point increase in job satisfaction, and a 0.49-point increase in perceived work quality on seven-point scales.

    The DWP study used comparison groups and regression, but licenses were not randomized, there was no pretrial baseline, and outcomes were self-reported. The honest conclusion is promising perceived productivity, not proven cash savings.

    The public-sector model stack

    The public-sector model stack

    Multilingual speech-to-text

    Government receives information by phone, at counters, in hearings, through field visits, and in public meetings.

    Text-to-speech and citizen voice agents

    A citizen should be able to ask about eligibility, application status, required documents, appointments, or service locations by phone.

    Policy and procedure RAG

    RAG can help public servants retrieve current law, policy, procedure, guidance, forms, and precedent.

    Casework copilots

    A casework model can summarize a file, identify missing evidence, retrieve the controlling rule, prepare a chronology, and draft a communication.

    Document intelligence

    Government processes forms, certificates, permits, invoices, evidence, correspondence, and scans. Document models can classify, extract, validate, and route them.

    Forecasting and resource allocation

    Models can forecast service demand, inspections, call volume, maintenance, emergency needs, and staffing.

    Multilingual speech-to-text

    Government receives information by phone, at counters, in hearings, through field visits, and in public meetings. Domain ASR can transcribe, translate, classify, and structure these interactions.

    Use cases include contact-center calls, grievance intake, inspection notes, council meetings, court administration, emergency communications, and accessibility captions. Measure entity accuracy for names, addresses, identifiers, dates, and legal or program terms, not only overall word error.

    Text-to-speech and citizen voice agents

    A citizen should be able to ask about eligibility, application status, required documents, appointments, or service locations by phone. The agent must answer from approved information, authenticate when accessing a case, and transfer high-risk or ambiguous matters.

    India's Bhashini program reported support for more than 36 text languages and 22 voice languages, integration across more than 500 websites, and over 100 live use cases in a January 2026 government release. Earlier releases used different model-count definitions, so language coverage and deployed use cases are more stable comparison points than raw model totals.

    Policy and procedure RAG

    RAG can help public servants retrieve current law, policy, procedure, guidance, forms, and precedent. It must preserve jurisdiction, effective date, authority, and security classification.

    The UK's Redbox is described as a protected generative-AI tool used across parts of government, including for documents up to OFFICIAL SENSITIVE under its controls. Cabinet Office Assist uses RAG to support government communicators and states that its AWS Bedrock configuration does not make customer data available to model providers for training.

    Casework copilots

    A casework model can summarize a file, identify missing evidence, retrieve the controlling rule, prepare a chronology, and draft a communication. It should not make a final eligibility, enforcement, or adjudicative decision unless the law, policy, and governance explicitly permit automation.

    Document intelligence

    Government processes forms, certificates, permits, invoices, evidence, correspondence, and scans. Document models can classify, extract, validate, and route them. A small language model can explain missing fields in plain language.

    Measure field precision and recall, straight-through rate, accessibility, review time, and incorrect rejection rate. Performance must be checked across languages, document quality, and citizen groups.

    Forecasting and resource allocation

    Models can forecast service demand, inspections, call volume, maintenance, emergency needs, and staffing. Optimization can schedule appointments, vehicles, officers, or field crews under policy and fairness constraints.

    Internal productivity models

    Drafting, meeting summaries, research, and search are lower-risk entry points. They still need protected environments, classification controls, quality review, and measured task outcomes.

    A sovereign architecture

    Sovereignty is not identical to on-premise hosting. It is operational control over data, identity, keys, model versions, dependencies, logs, and the ability to continue or exit.

    1. Identity and citizen boundary: public, authenticated, employee, privileged, and classified contexts.
    2. Authoritative source registry: owner, legal authority, jurisdiction, effective date, language, and status.
    3. Model router: deterministic rule, SLM, specialist model, or larger model based on task and risk.
    4. Data controls: residency, encryption, retention, redaction, and provider-training restrictions.
    5. Workflow controls: approval, appeal, escalation, and service-system integration.
    6. Evidence record: query, sources, model and prompt version, output, edits, and final action.
    7. Portability: exportable corpus, evaluation set, logs, and model alternatives.

    Risk tiers

    TierExampleDefault control
    1public FAQ grounded in published guidanceautomated answer with source and feedback
    2employee draft or meeting summaryuser review before use
    3authenticated case explanationpermission check and trained officer review
    4eligibility, enforcement, or legal effectformal human decision, reason, and appeal path
    5safety, defense, or critical infrastructure controlspecialized assurance and separate authority

    Evaluation beyond accuracy

    Government must test:

    • factual and citation correctness
    • completeness and appropriate abstention
    • current-versus-obsolete source selection
    • access-control leakage
    • language and accent parity
    • accessibility
    • inconsistent outcomes across protected or vulnerable groups
    • time to service completion
    • escalation and appeal quality
    • cost per successfully completed case

    Publish an algorithmic transparency record where appropriate. The UK records for Redbox and Assist show a practical pattern: state purpose, data, model, risks, oversight, and contact.

    A 90-day first deployment

    Choose an internal search or public information workflow with authoritative content and no automated legal decision. Assemble 200 real questions, including adversarial and multilingual cases. Define what the system may answer and when it must abstain. Test permissions and outdated documents.

    Run staff-only first. Compare time to a verified answer, citation correctness, escalation, and user correction. Expand to citizens only after language, accessibility, privacy, and complaint routes are ready.

    The strategic conclusion

    Government should not pursue the largest model it can access. It should pursue the smallest controllable system that meets the public task.

    That means multilingual models for access, RAG for authoritative knowledge, document models for case intake, forecasting for resources, and explicit human authority for consequential decisions. Sovereignty is the ability to inspect, govern, replace, and defend the whole system.

    Research note

    Research is current through September 5, 2026. Government inventory counts depend on agency reporting and definitions. Productivity trials relied materially on self-reported outcomes. Public-sector deployments require jurisdiction-specific legal, procurement, accessibility, privacy, and records review.

    Continue the research

    Building a Production-Ready System

    Conscious Engines builds sovereign AI for government with deployment, data, language, permissions, and audit designed around the institution. Specialized speech, small language models, private RAG, and constrained agents can improve public services without surrendering control of sensitive data or official decisions.