Close Menu
    Facebook X (Twitter) Instagram
    KSA News TodayKSA News Today
    Facebook X (Twitter) Instagram
    • KSA
    • Business
    • Technology
    • Sports
    • Lifestyle
    KSA News TodayKSA News Today
    • KSA
    • Business
    • Technology
    • Sports
    • Lifestyle
    • Contact us
    Technology

    Core42 advances secure and scalable AI deployment for UAE government services

    Editorial TeamBy Editorial TeamAugust 18, 2026
    Share Facebook Twitter Pinterest Copy Link Telegram LinkedIn Tumblr Email
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Raghu Chakravarthi, Chief Product and Technology Officer, Core42.

    Raghu Chakravarthi outlines how workload-aware orchestration, diverse accelerator architectures and embedded governance can support secure, scalable and cost-efficient AI deployment across UAE government services.

    The UAE’s ambition to transition 50 per cent of government services and operations to Agentic AI will require infrastructure capable of delivering intelligence securely, economically and at production scale. Sovereignty, compute flexibility, workload orchestration, observability and financial governance will become increasingly important as autonomous agents generate multiple inference events to complete complex tasks.

    In an interview with Tahawultech.com, Raghu Chakravarthi, Chief Product and Technology Officer at Core42, explains how integrated compute, cloud, data and governance capabilities can move Agentic AI beyond isolated pilots. He also discusses how routing workloads across different models, accelerators and deployment environments can improve performance, control costs and protect sensitive government data.

    Interview excerpts

    How critical will sovereign, scalable AI infrastructure be to achieving the UAE’s goal of transitioning 50 per cent of government services and operations to Agentic AI?
    Sovereign, scalable AI infrastructure is foundational to achieving this ambition. Moving a significant share of government services to Agentic AI is ultimately a question of whether intelligence can be delivered repeatedly, economically and under control at production scale. Agentic systems raise the stakes because they change the unit economics of AI. A single user request can become many inference events behind the scenes as the system reasons, calls tools, and completes multi-step tasks on its own. Consumption stops scaling with headcount and starts scaling with autonomy, which is exactly the dynamic that can produce bills far larger than anticipated at the pilot stage. 

    What is distinctive about the Gulf is that sovereignty and scale are being pursued together from the outset, rather than sovereignty being retrofitted onto infrastructure built elsewhere. In this region, sovereign control is a starting requirement, which, combined with significant national investment in compute, has produced infrastructure that is locally governed, globally competitive, and built to support AI adoption at scale from day one.

    For government, the priority is therefore to embed scalability, sovereignty, observability and governance into the inference layer so these capabilities can expand alongside Agentic AI rather than being added later.

    What compute, cloud, data and governance capabilities must be established to move Agentic AI from isolated pilots into secure government-wide deployment?
    Government-wide deployment requires compute, cloud, data and governance to operate as one integrated system rather than as separate layers. At the compute level, AI workloads place very different demands on infrastructure. Accelerator architectures vary in areas such as time to first token, sustained throughput, memory capacity, utilisation, power efficiency and cost. The optimal choice also changes according to the model being used, the length of the context, the number of concurrent users and whether the workload is interactive or asynchronous. This makes access to diverse silicon and the ability to right-size models and accelerators increasingly important as workloads scale. Rather than forcing every workload onto one architecture, infrastructure should span diverse silicon, including NVIDIA, AMD, Qualcomm and Cerebras, matching each workload to the most appropriate accelerator based on latency sensitivity, throughput requirements, utilisation and cost.

    When every workload is locked to a single accelerator ecosystem, organisations risk running tasks on more expensive silicon than necessary and paying for idle or over-provisioned capacity across the estate.

    The deployment architecture must also provide flexibility across real-time and batch processing, cloud and on-premises environments, and shared, private or sovereign infrastructure. Routing should be treated as a first-class capability across three dimensions simultaneously – models, accelerators and deployment paths. To make that routing automatic rather than manual, Compass is developing a workload classifier that inspects each incoming request and determines the right target for it – the most suitable silicon, whether NVIDIA, AMD, Qualcomm or Cerebras; the right model, whether a proprietary or open-source option; and the location in which it should run. Classifying the request at the point of entry is what allows the platform to align performance, cost and sovereignty for every workload without a person making that trade-off by hand.

    On data and governance, sovereignty needs to be embedded directly into the inference layer. In-country data residency, encryption, access controls, guardrails and zero logging of customer data should be applied as part of the deployment path rather than managed through separate processes, supported by more than 170 security policies and SOC 2 Type II certification.

    Finally, observability and financial controls are essential. Organisations need centralised visibility into usage, cost and activity, together with metering, budget controls and guardrails, so they can understand what an agent consumed, which models it used and whether its actions remained within approved governance boundaries. For government, that means spending can be attributed by department, project, application or model, with budgets, quotas and alerts active before broad deployment rather than introduced after consumption has already scaled.

    Together, these capabilities allow a government entity to move a workload from an experimentation environment to a production-grade deployment without rebuilding its stack.

    How can government entities control inference costs as autonomous agents generate multiple model calls and AI consumption increases at scale?
    The starting point is to manage the economics of the complete agentic workflow rather than focusing only on the number of users, prompts or the quoted price of an individual token. Government entities need to understand how many inference events a task generates, how context grows across each step, which stages require low latency, which can be processed in batches and how frequently an agent may repeat or redirect an action. This becomes particularly important as Agentic AI is deployed across high-volume government services.

    Cost controls also need to be active before consumption scales. Spend caps, budgets, alerts and routing logic should be built in from the outset rather than added once usage has already expanded. The most common planning mistake is basing capacity on user counts rather than inference events per task. Agentic workloads therefore require observability, governance and workload-aware routing to be engineered in from day one.

    The next lever is workload placement, and this is where a workload classifier becomes central. By classifying each incoming request as it arrives, the platform can direct it to the right model, silicon and location automatically, so a routine classification or extraction request is not sent to the largest available model, while a high-volume summarisation workload is processed through a high-throughput batch path rather than occupying capacity designed for real-time interaction.

    The objective is to improve tokens per second per dollar by delivering the required outcome at the right speed and cost, rather than optimising for token price alone.

    Centralised visibility, real-time budgets and alerts, access controls and guardrails then allow government entities to see not only whether an agent completed a task, but how many resources it consumed, which models it used and whether every action remained within approved data and governance boundaries.

    How can workload-aware orchestration across different models, accelerators and deployment environments improve the performance, sovereignty and economics of government AI services?
    Workload-aware orchestration improves all three by turning infrastructure from a fixed architectural choice into a workload-placement decision. Multi-silicon routing allows infrastructure to become a workload-placement decision, matching each task to the model, accelerator and deployment path that can meet its quality and latency requirements at the best economics. Efficiency in the inference era comes from treating that diversity as an economic asset rather than a procurement inconvenience. Compass applies that principle through a workload classifier that evaluates each incoming request and routes it across three dimensions. Across models, so each task runs on a right-sized model from our library of more than 60 open and proprietary models rather than the biggest one available.

    Across accelerators, NVIDIA, AMD, Qualcomm, and Cerebras, so latency-sensitive and throughput-heavy workloads land on the hardware suited to each. And across deployment paths, balancing real-time against batch, cloud against on-premises, and shared against private or sovereign, including the location in which a request is processed. The routing criteria then combine latency sensitivity, throughput requirements, utilisation and cost with the governance question of where a workload is permitted to run and how its data must be handled. For government services, that can mean faster responses for citizen-facing applications, better utilisation for high-volume workloads and stronger sovereign control over sensitive data, without forcing every workload onto the most expensive model or infrastructure.

     


    Source: Tahawul Tech

    Previous ArticleTwo Holy Mosques Authority Invites Makkah Visitors to Grand Mosque Exhibition
    Next Article Prime Inspections & Snagging to launch prime inspections, its own snagging and booking app

    Related Posts

    OpenAI targets AI-native enterprise growth in UAE with local data, inference residency

    August 18, 2026

    Logitech advances human-centric hybrid workplaces across IMEA through AI

    August 18, 2026

    Apple collaborates with Alibaba on Chinese AI model

    August 17, 2026
    Latest Posts

    ‘Set an example at home’: UAE doctors share ways to reverse teens’ insulin resistance

    Saudi Arabia warns repeated Iranian attacks signal dangerous escalation

    Aramco, Maaden form mining JV to explore nearly 10% of Saudi Arabia’s land

    Prime Inspections & Snagging to launch prime inspections, its own snagging and booking app

    Latest News

    OpenAI targets AI-native enterprise growth in UAE with local data, inference residency

    August 18, 2026

    Logitech advances human-centric hybrid workplaces across IMEA through AI

    August 18, 2026

    Apple collaborates with Alibaba on Chinese AI model

    August 17, 2026
    Facebook X (Twitter) Instagram Pinterest
    • KSA
    • Business
    • Technology
    • Sports
    • Lifestyle
    • Contact us
    2026. All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.