Vietnam crab exporterVietnamese mud crab export
🏈 Draft. Trade. Win. Explore Marvel comics The Fall Reset 🛍️ Check home prices 🏠

"We're at an Inflection Point": Telnyx CEO Says AI Infrastructure Is Becoming the Industry's Next Bottleneck

Adobe Stock
Lyssanoel Frater
Contributor
July 22, 2026, 3:03 p.m. ET

GitHub recently revealed it needs to expand its infrastructure 30 times over to keep pace with AI-driven development.  The announcement has become a reference point for infrastructure providers who argue the industry may be underestimating the scale required to support widespread agent adoption. One AI infrastructure CEO is speaking out about the need to scale. 

As AI agent deployments move from pilot programs to enterprise scale, a growing body of research and a string of real-world outages suggest the communications and computing infrastructure underneath those systems must be scaled for what is coming.

Several companies like UnitedHealth, Salesforce, and OpenAI have each announced major agent initiatives, promising software systems that can handle complex tasks, make decisions, and interact with customers with growing levels of autonomy. But while much of the conversation focuses on what AI agents can do, far less attention is being paid to the infrastructure that makes them possible.

"Everyone is focused on building smarter agents. That's the right problem to solve," said David Casem, CEO of Telnyx, a voice AI infrastructure platform. "What we're not talking about is what happens when you go from running 1,000 agents to running 10 million of them simultaneously. That's a fundamentally different engineering problem, and most of the industry isn't ready for it."

The concern is not theoretical. As organizations move from pilot programs to enterprise-wide deployments, researchers and engineers are beginning to grapple with what some are calling the agent scaling problem. According to a study by researchers at the University of Hong Kong, current GPU schedulers inflate end-to-end AI agent latency by three to eight times because they treat each model call as an independent request rather than part of a continuous workflow. 

Separately, researchers at UC Berkeley and Google DeepMind, presenting at the USENIX Networked Systems symposium in May 2026, found that existing AI serving systems ignore dependencies between agent program calls, causing cascading wait times that grow more severe as deployments scale.

For companies deploying voice-enabled agents, the stakes are higher. According to research on human conversational turn-taking from the University of Victoria, people naturally respond within 200 to 500 milliseconds. A chatbot user might tolerate a brief, visually signaled delay of a couple of seconds, according to a 2025 study in the International Journal of Human-Computer Interaction. A customer on a phone call is much less tolerant of unexplained silence.

"At Telnyx, we understand voice is where the infrastructure problem becomes impossible to hide," Casem said. "When you're running real-time conversations at scale, latency isn't a performance metric. It's the product. If the network can't deliver, the agent fails, no matter how good the model is."

Telnyx, which was ranked the sixth fastest-growing software vendor in Brex's Spring 2026 benchmark, has built its business around that gap. Telnyx operates its own global communications network and has expanded into AI inference infrastructure and AI voice orchestration, allowing communications workloads and AI workloads to run closer together. The company argues that reducing network hops becomes increasingly important as real-time AI deployments scale.

Casem believes that shift reflects a broader change occurring across the AI industry. "Intelligence is becoming abundant," he said. "Infrastructure is becoming scarce."

The company has put their theory of the future of the industry to the test internally. According to Casem, Telnyx now runs 1,400 autonomous bots alongside its 150-person engineering team, with every engineer supervising AI agents rather than writing code line by line.

"If you're building agent infrastructure for the next decade, you have to engineer it like telecommunications. Reliability, latency, and uptime aren't features. They're table stakes," Casem said.

According to GitHub's own public blog, the software development platform acknowledged it initially planned for a tenfold capacity increase to handle surging AI-assisted development, only to discover within months it needed to engineer for 30 times today's scale. 

Meanwhile, researchers at the University of Edinburgh and Tencent published findings this year showing that without purpose-built scheduling systems, multi-agent pipelines systematically underutilize GPU resources while accumulating bottlenecks across every stage of execution.

"We're at an inflection point," Casem said. "The demand for frontier and near-frontier intelligence is becoming insatiable. Infrastructure is becoming scarce. The companies that solve the infrastructure layer, networks, inference, and real-time communications, are going to have enormous influence over how AI actually gets deployed in the world."

More from Contributor Content