Artificial intelligence is having its broadband moment.
For years, AI was something most organizations discussed in innovation labs, strategy meetings, and carefully contained proof-of-concept environments. Then generative AI arrived in a form that practically anyone could use. Suddenly, employees were creating content, writing code, summarizing documents, analyzing data, and imagining entirely new business models from a browser window.
That shift is exciting. It is also misleading.
The chatbot is the visible part of the experience, but the real transformation is happening behind it. Every useful enterprise AI service depends on an enormous stack of computing, networking, storage, data engineering, security, governance, and operational discipline. The interface may feel simple. The infrastructure absolutely is not.
That is why Greg Diamos’ perspective on the future of AI deserves the attention of every CIO, CTO, infrastructure leader, and technology executive. Diamos has worked across several important layers of modern machine learning, including large-scale AI systems, GPU architecture, model benchmarking, and enterprise large language model development. His work points toward a reality that IT leaders can no longer treat as a side conversation: smarter AI will require dramatically more capable infrastructure.
After more than 20 years leading technology organizations, I find this part of the AI revolution especially interesting. The algorithms are remarkable, but turning them into dependable business services is an infrastructure and operating-model challenge. The demo is the easy part. Keeping AI fast, secure, governed, available, and financially sustainable on a Tuesday afternoon is where enterprise IT earns its paycheck.
The Real AI Race Is an Infrastructure Race
Most public discussion focuses on which company has the smartest model. Enterprises need to ask a different question: Which organization can consistently deliver useful AI at scale?
Model quality matters, but it is only one component. A production AI platform also needs access to trusted data, sufficient compute capacity, high-throughput storage, low-latency networking, effective orchestration, cybersecurity controls, observability, cost governance, and a skilled operational team.
Experienced IT leaders have seen this pattern with virtualization, cloud computing, big data, and Kubernetes. A powerful technology appears, early adopters build custom environments, demand accelerates, and infrastructure teams must industrialize what began as experimentation. AI is following the same pattern, only faster and with a much larger appetite for resources.
Stanford’s 2023 AI Index showed that industry had moved decisively ahead of academia in producing significant machine-learning models, largely because advanced AI increasingly requires levels of data, compute, and capital that commercial organizations are better positioned to assemble. That trend makes the infrastructure barrier impossible to ignore.
Scaling Intelligence Means Scaling Compute
One of the most important ideas behind the current AI boom is that model performance can improve predictably as developers increase model size, training data, and computing resources. Research on neural scaling laws established measurable relationships between performance and the amount of compute used to train language models.
For an enterprise executive, the takeaway is not that every company should train a frontier model. Most should not.
The takeaway is that AI capability has a direct infrastructure cost. Even when an organization uses an existing foundation model, it still must support inference, fine-tuning, retrieval-augmented generation, vector databases, data pipelines, application integration, identity management, monitoring, and model governance.
Inference can become particularly expensive because it happens whenever an employee, customer, application, device, or automated process asks the model to do something. A successful AI service can create its own infrastructure problem simply by becoming popular.
That changes capacity planning. AI workloads can be bursty, GPU-intensive, data-hungry, and highly sensitive to latency. Sizing infrastructure once and revisiting it during the next budget cycle will not survive contact with enterprise AI.
The GPU Shortage Is Only the Beginning
Much of the current infrastructure conversation revolves around GPUs, and for good reason. High-performance accelerators are essential for training and serving many modern AI models. In 2023, cloud and hardware providers were already racing to bring new GPU platforms to market because businesses lacked sufficient access to the infrastructure required for large generative AI workloads.
However, buying GPUs does not create an AI platform any more than buying servers automatically creates a private cloud.
Organizations also need to solve for:
- High-bandwidth networking between accelerators, storage, and data services
- Storage capable of feeding large datasets without starving expensive compute
- Scheduling and orchestration that prevent scarce GPUs from sitting idle
- Data pipelines that deliver governed, current, and properly classified information
- Security controls for prompts, models, APIs, identities, and sensitive enterprise data
- Observability across infrastructure, model performance, application behavior, and cost
- Power, cooling, rack density, lifecycle management, and capacity expansion
This is a full-stack architecture problem. A poorly designed environment can spend premium money on accelerators and still deliver disappointing performance because the surrounding network, storage, software, or operational processes cannot keep up.
The winning architecture will not necessarily have the most GPUs. It will produce the greatest amount of trusted business value from every unit of compute.
Hybrid Multicloud Becomes the Practical AI Architecture
The cloud remains essential to AI because it offers immediate access to models, managed services, and specialized infrastructure. It is often the fastest place to begin an experiment. Yet public cloud cannot be the only answer for every workload, especially when GPU availability, data-transfer costs, persistent inference demand, regulatory requirements, and data gravity enter the equation.
This is where hybrid multicloud becomes a practical operating strategy, not just a marketing phrase.
An enterprise might use public cloud services to test several models, burst for temporary training capacity, or consume a managed AI API. The same company may run production inference on-premises because its data is local, its usage is predictable, or its privacy requirements are strict. Edge infrastructure may handle manufacturing, retail, healthcare, logistics, or remote-site use cases where latency and connectivity matter.
The goal is not to place everything in one location. The goal is to place each AI workload where it delivers the best combination of performance, economics, security, compliance, and operational control.
That requires portability. IT leaders should favor platforms that allow teams to move models, containers, data services, and application components without rewriting the entire environment. Lock-in is inconvenient in traditional infrastructure. In AI, where models, accelerators, frameworks, and commercial options are evolving at remarkable speed, lock-in can become a strategic liability.

Enterprise Data Is the Competitive Advantage
Foundation models are impressive because they understand broad patterns. Enterprise AI becomes valuable when it understands the organization.
That means the competitive advantage is not merely access to a public model. Most competitors can obtain access to similar models. The real differentiator is the ability to connect AI safely to proprietary data, operational history, customer context, industry expertise, workflows, policies, and intellectual property.
This is why data architecture is inseparable from AI infrastructure.
Organizations need clean data, useful metadata, access controls, lineage, classification, and a repeatable method for grounding model responses in authoritative sources. Retrieval-augmented generation can use internal knowledge without retraining an entire foundation model, but it still depends on reliable ingestion, indexing, permissions, and lifecycle management.
AI will expose every weakness in an organization’s data strategy. Duplicate records, conflicting definitions, inaccessible repositories, and poorly governed file shares do not magically improve when connected to a language model. They simply become faster ways to produce inconsistent answers.
AI Infrastructure Is Also Security Infrastructure
The rush to deploy generative AI creates a familiar temptation: Build first and govern later. That approach is risky.
AI applications introduce new attack surfaces and amplify existing ones. Sensitive information can leak through prompts. Models can produce inaccurate output. Access controls can be applied inconsistently. APIs can be abused. Unapproved tools can create shadow AI across the business. Model dependencies and training data can introduce supply-chain concerns.
McKinsey’s 2023 survey found that generative AI adoption was already widespread, but many organizations had not established policies or mature controls for its risks. Only a minority reported actively mitigating inaccuracy, cybersecurity, and related concerns.
The NIST AI Risk Management Framework, released in January 2023, gives technology leaders a useful structure for incorporating trustworthiness into the design, deployment, and operation of AI systems. It reinforces an important point: Governance is not paperwork added after deployment. It is an architectural requirement.
Security, privacy, explainability, resilience, auditability, and human oversight should be designed into the platform from the beginning. The organizations that do this well will move faster because they will not need to stop every successful pilot for a six-month compliance rescue mission.
The Business Case Is Too Large to Ignore
The infrastructure challenge is substantial, but so is the opportunity.
McKinsey estimated in 2023 that generative AI use cases could create between $2.6 trillion and $4.4 trillion in annual economic value across industries. It also found significant potential to automate or augment activities that consume a large portion of employee time, particularly in knowledge-intensive work.
This is why AI infrastructure should not be treated purely as a technology expense. It is a business capability.
In healthcare, AI can help interpret complex information and identify patterns across large datasets. In manufacturing, it can accelerate quality analysis, maintenance planning, and access to operating procedures. In software development, it can assist with code, testing, documentation, and troubleshooting. In financial services, it can improve research, service, fraud analysis, and regulatory workflows.
The executive question is not whether AI will affect the business. It is whether the company will build the capabilities to capture that value responsibly.
What IT Leaders Should Do in 2024
I would not recommend beginning with a massive AI data center purchase. I would begin with a disciplined enterprise AI roadmap.
First, identify a small portfolio of high-value use cases tied to revenue, customer experience, employee productivity, risk reduction, or operational efficiency. Second, classify the data those use cases require. Third, determine the performance, security, compliance, and availability requirements. Fourth, select the model and deployment approach only after the business and data requirements are understood.
From there, build a reusable platform rather than a collection of isolated pilots.
That platform should include approved models, secure access to enterprise data, identity-based controls, standardized APIs, infrastructure automation, cost reporting, model monitoring, and a clear path from development to production. It should support multiple deployment locations and avoid assuming that one model, cloud, or hardware vendor will remain the best answer indefinitely.
Most importantly, IT must partner with the business. AI cannot succeed as a science project owned by one technical team. It needs executive sponsorship, domain expertise, legal and security participation, data ownership, workforce enablement, and measurable outcomes.
The Next AI Breakthrough May Be Operational
The future of AI is often described in terms of smarter models and increasingly human-like capabilities. Those advances will be fascinating, but the next major enterprise breakthrough may be far less theatrical.
It may be the moment when organizations can deploy AI as reliably as they deploy databases, virtual machines, and cloud applications.
That means predictable performance, governed data, repeatable automation, transparent costs, resilient operations, secure access, and enough capacity for a pilot to become a platform used by thousands of people.
Greg Diamos’ infrastructure-focused view of AI is compelling because it brings the conversation back to reality. Intelligence does not float in the cloud as an abstract idea. It runs on physical systems, consumes power, moves data, depends on networks, and requires people who know how to operate complex technology.
For IT leaders, that is not bad news. It is an invitation.
The AI era will not reduce the importance of infrastructure. It will elevate infrastructure from a supporting function to one of the most strategic capabilities in the enterprise. The organizations that understand this early will not just adopt AI faster. They will build AI that is useful, defensible, scalable, and ready to create measurable business value.
References Published Before January 2024
- Stanford Institute for Human-Centered Artificial Intelligence, The 2023 AI Index Report — 2023
- McKinsey & Company, The Economic Potential of Generative AI: The Next Productivity Frontier — June 2023
- McKinsey & Company, The State of AI in 2023: Generative AI’s Breakout Year — August 2023
- Google Cloud and NVIDIA, Next-Generation AI Infrastructure and Software for Large-Scale Models — March 2023
- National Institute of Standards and Technology, AI Risk Management Framework 1.0 — January 2023
- OpenAI researchers, Scaling Laws for Neural Language Models — January 2020
- NVIDIA, Global Data Center System Manufacturers to Supercharge Generative AI and Industrial Digitalization — August 2023




