At the Hot Interconnects 2026 conference, NVIDIA Senior Vice President of Networking Gilad Shainer outlined a fundamental shift in how large-scale AI computing infrastructure must be conceived, moving from a collection of discrete servers to a single, co-designed unit of computing. His keynote, titled "Networking Innovations for Gigascale AI Systems," presented an architecture where the network is the central nervous system of what he termed an "AI factory."
As reported by TechCrunch, this approach argues that the "harness"—the surrounding infrastructure of networking, storage, and software—is now the critical differentiator for performance, not just the raw power of the AI models themselves. Shainer's blueprint divides the system into purpose-built infrastructure domains: NVLink for tightly coupling GPUs within a server (Scale Up), Spectrum-X Ethernet or InfiniBand for connecting server racks (Scale Out), Spectrum-XGS Ethernet for data center-wide connectivity (Scale Across), and BlueField-4 data processing units for infrastructure services.
A key update from Shainer was the announcement that NVIDIA's co-packaged optics (CPO) technology, which integrates silicon photonics directly with switch silicon to drastically reduce power consumption and increase bandwidth density, is now in production. This move from roadmap to product underscores the intense focus on maximizing useful AI token throughput within strict power and facility constraints, a core tenet of the new architecture which converges around what NVIDIA calls the "Vera Rubin" generation of systems.
Concurrently, NVIDIA released a suite of developer tools targeting specific, persistent engineering challenges in applied AI. According to a separate report, these releases span generative recommender systems, on-device robot control, AI agent skill evaluation, and full-stack observability for AI infrastructure. The company highlighted that recommender systems, one of the most widespread machine learning applications, present unique scaling difficulties due to their complex mix of categorical and continuous feature data, which the new tools aim to address.
Perhaps the most symbolically significant release is TensorRT Model Connect (TRTMC), now in public preview. NVIDIA announced that the project was built entirely using OpenAI Codex agents under human direction and review. The agents handled model implementations, performance tuning, tests, integrations, and documentation. This development directly addresses friction points in AI deployment workflows and serves as a real-world case study in using AI agents for complex software development.
The dual announcements—one focused on the physical and network architecture of AI factories, the other on the software tools that run on them—signal NVIDIA's expanding focus beyond just manufacturing accelerators. The company is now engineering the full stack, from the photonics in the network switches to the AI agents that help write its software libraries. This holistic approach reflects the industry's maturation, where the bottleneck for large-scale AI has shifted from merely acquiring GPUs to efficiently orchestrating thousands of them as a single, power-aware system and providing the tools to leverage that system effectively.
Shainer's presentation frames the AI data center not as a passive housing for computers but as a factory where the network is the assembly line, dictating the pace and efficiency of production (in this case, tokens of AI output). The new developer tools, particularly the Codex-agent-built TensorRT Model Connect, represent the specialized tooling and automation required to operate such a factory. Together, they illustrate a broader industry trend where foundational AI companies are vertically integrating, providing not just the core silicon but the entire ecosystem necessary to deploy AI at gigascale, turning raw computational power into reliable, usable intelligence.








