Running the Model Mix: Why MODaaS Is the Next Telco AI Priority 

Contributing experts

Telco AI investment has not stalled for lack of ambition or budget. As HTEC’s research makes clear, the constraint is organizational absorption: the gap between knowing what AI should do and getting it to actually run in production, across fragmented systems, stretched teams, and infrastructure that was never designed for this pace of change.

Most operators have cleared the first hurdle. Pilots are running, use cases are identified, and vendor relationships are in place. The problem showing up now is subtler and more expensive: the model portfolio underlying those pilots lacks a governance layer. Every request, whether a simple ticket classification or a multi-step RAN configuration task, routes to whatever frontier model the developer reached for first. The bill grows. The cost per outcome stays opaque. And when anyone asks which model is doing what, the honest answer is that nobody knows.

TMForum recognized this problem at an industry scale. GB1085 defines MODaaS as the governance and routing layer that should sit between AI consumers and the model portfolio, organizing access around nine operational domains: Unified Access, Operational Control, FinOps, Observability, Security and Governance, Discovery and Registry, Predictive Resilience, Quality Evaluation and Drift, and Supply-Chain Integrity.

AT&T’s production deployment illustrates the potential of intelligent model routing. Across 3,601 real-world tasks, 88% achieved equivalent results with either a lower-cost or premium model, while the quality gap was only three percentage points, and the cost gap reached 67% per million output tokens. Yet developers chose the premium model 85% of the time, even though only 16% of tasks required it. Implementing a Smart LLM Router based on the GB1085 MODaaS pattern reduced premium-model usage to 27%, maintained 99.5% of output quality, and cut costs by approximately 29%.

Why telco is the right vertical for this

HTEC’s earlier analysis of implementation-layer stalling identified something specific to telco: the communication model is shifting from people talking to people toward agents talking to agents. That shift changes what the network carries, but it also changes what the AI stack needs to do. An agentic workflow running network configuration, fault triage, or customer journey automation across dozens of turns is a different engineering problem from a single-prompt query. Simple per-prompt routing breaks in those sessions because switching models mid-session resets the prompt cache, and the cost compounds exponentially as context grows.

Cache-aware routing holds model selection stable across a session to exploit cached prefix tokens. For a telco running long-horizon agentic workflows across OSS, BSS, and customer systems, that capability is not optional optimization: the cost difference between a naive router and a cache-aware one grows with every turn of every session running in production.

Telcos also carry a regulatory dimension, such as data residency constraints, EU AI Act configurability, and per-jurisdiction governance. GB1085‘s governance domains address them directly, which is one reason the TM Forum specified MODaaS as a standard rather than leaving it to individual operators to solve on their own.

Where HTEC plays

OneLoopAI gives telecom operators the technical foundation they need to build MODaaS (Model Operations as a Service) without having to assemble and integrate dozens of AI components themselves.

HTEC OneLoopAi adds real-time measurement of AI usage, costs, and realized value, giving operators the transparency that makes the MODaaS cost case visible and the FinOps governance domain operational from day one.

The operators pulling ahead on AI are those treating it as an operational redesign. MODaaS provides a governed routing layer that makes the model mix visible, auditable, and cost-accountable.

FAQ

MODaaS stands for Model-as-a-Service. The term and its formal definition come from TM Forum’s GB1085 specification, which established MODaaS as the governance and routing architecture for managing AI model portfolios in telecommunications and enterprise environments.

Tokenmaxxing describes the organizational pattern of routing every AI request to the most expensive available model, regardless of task complexity. The term captures the waste that occurs when premium model access is the default rather than the exception, driving costs well above what the actual task mix requires.

A hyperscaler gateway (AWS Bedrock, Azure AI Studio, and equivalents) solves the connection problem: making models accessible through a single API surface. MODaaS solves the control problem: governing which model handles which request, enforcing data residency, attributing cost by agent or cost center, supporting open and self-hosted models, and providing the audit trail that compliance and finance teams require.

A Smart LLM Router intercepts each incoming prompt, scores it for complexity, and dispatches it to the least expensive model capable of producing an adequate result. The scoring model is trained on a corpus of real tasks, learning which request types require premium capability and which can be handled by cheaper alternatives.

Cache-aware routing holds the model selection steady across all turns of a multi-turn session rather than re-evaluating it on every prompt. When a model switch occurs mid-session, the prompt cache resets and the operator pays to re-process the full context. On long agentic runs (network configuration workflows, customer journey automation, and similar), that cost compounds per turn. Cache-aware routing prevents that by locking in the model for the session once it is selected.

Telcos run long-horizon agentic workflows across OSS, BSS, and customer systems where cache-aware routing produces the largest cost savings. They operate under strict data residency and jurisdictional governance requirements that MODaaS addresses natively. They also face the agent-to-agent communication shift that requires the underlying AI stack to be governed at a more substantive level.

The lowest-risk entry point combines the three native-ready GB1085 domains (Unified Access, Operational Control, and FinOps) with an observability integration that makes the current model spend visible by workload. That combination produces the baseline measurement needed to build the routing model and justify the full MODaaS build.

Explore more

Most popular articles