Nvidia started shipping Vera Rubin this month. Most edge sites can't physically host it. Here's the part that never makes the roadmap slide. Every Nvidia generation makes the rack hotter. A single B300 pulls 1,400 watts. A GB300 NVL72 rack pulls 132 to 140 kW today, liquid cooled, no exceptions. Rubin climbs from there. And inference is the workload driving the volume now. It doesn't want to sit in three mega-campuses. It wants to live near the data and the users, which pushes it out to regional and edge sites. So watch the collision. The silicon is getting denser and hotter. The workload is moving to smaller, distributed locations. And most of the buildings in those locations were spec'd for 5 to 10 kW racks, back when that was a lot. Air cooling quits above 40 kW per rack. That's not a vendor preference. It's physics. You do not turn a 2015 shell into a 130 kW liquid-cooled hall by adding fans. This is the whole case for factory-built modular. You build the power and the cooling for the density up front, in a plant, then ship it to the site. Edge inference at 40 kW per rack and up, cooled properly, live in months instead of years. The roadmap isn't slowing down. Either your real estate catches up to the silicon, or the silicon sits in a crate.
Yuri Milyutin’s Post
More Relevant Posts
-
Nvidia started shipping Vera Rubin this month. Most edge sites can't physically host it. Here's the part that never makes the roadmap slide. Every Nvidia generation makes the rack hotter. A single B300 pulls 1,400 watts. A GB300 NVL72 rack pulls 132 to 140 kW today, liquid cooled, no exceptions. Rubin climbs from there. And inference is the workload driving the volume now. It doesn't want to sit in three mega-campuses. It wants to live near the data and the users, which pushes it out to regional and edge sites. So watch the collision. The silicon is getting denser and hotter. The workload is moving to smaller, distributed locations. And most of the buildings in those locations were spec'd for 5 to 10 kW racks, back when that was a lot. Air cooling quits above 40 kW per rack. That's not a vendor preference. It's physics. You do not turn a 2015 shell into a 130 kW liquid-cooled hall by adding fans. This is the whole case for factory-built modular. You build the power and the cooling for the density up front, in a plant, then ship it to the site. Edge inference at 40 kW per rack and up, cooled properly, live in months instead of years. The roadmap isn't slowing down. Either your real estate catches up to the silicon, or the silicon sits in a crate.
To view or add a comment, sign in
-
Most GPUs in the world are waiting. Waiting on data. Waiting on an agent, a loop, a graph. Waiting in a queue behind a job that shouldn't be running. The industry's answer has been to buy more GPUs. It is the most expensive way to solve an idle-GPU problem ever created. We measure this across more than a million GPUs. The gap between what an AI factory can do and what it actually does is 10x to 100x more. Same silicon. Same power. Same capital. Nobody publishes their utilization. That should tell you something. The next gigawatt will not fix what the last gigawatt is wasting. And no single layer fixes it, because the chip, the data path, the scheduler and the workload all have to agree. We've worked alongside NVIDIA on this for a decade. If you're building at the model layer, the cloud layer, or the energy layer and you think utilization is the real bottleneck, I want to talk. Comment and I'll reach out.
To view or add a comment, sign in
-
-
A standard enterprise rack was built for 8 to 12 kilowatts. A single NVIDIA GB200 NVL72 rack is specified at 120. That comparison explains more about this market than any model benchmark does. At 8 to 12 you can air-cool it, you can retrofit, and you can put AI workloads into space you already have. At 120 none of that survives. You're into liquid cooling, redesigned power distribution, different floor loading, and a facility planned around the density rather than adapted to it. NVIDIA's 2027 roadmap already shows a 600 kW rack. This is why "we'll just run it in our existing footprint" stalls so often. Budget and intent are usually fine. Physics isn't, and the building was finished before the requirement existed. It's also why purpose-built AI infrastructure became a category instead of a feature. You don't upgrade your way to that density. Before anything else gets committed, the number to establish is the actual per-rack ceiling in the facility you're planning to run this in. #AIInfrastructure #DataCenter #GPUCompute
To view or add a comment, sign in
-
-
We shipped two firsts last month, and they tell one story. First came the numbers 📝 We published the first measured silicon results for NVIDIA Vera Rubin NVL72: at similar interactivity, 10x more tokens per second per megawatt than Blackwell, with NVIDIA TensorRT-LLM and NVIDIA Dynamo optimizations enabled. Those are REAL results from live hardware, not projections. Then came the fabric underneath them. We deployed the NVIDIA Spectrum-X SN6600-LD, the industry's first fully liquid-cooled Ethernet switch, and packed 1.64 Pb/s of fully non-blocking capacity into a single rack. Those aren't two separate wins. Reasoning models push enormous token traffic across the network, and every watt the fabric saves is a watt the GPUs get back. You only see numbers like these when compute, networking, and cooling are engineered as one system, which is exactly how we build. More info: https://t.co/qwQHD9cSn2
To view or add a comment, sign in
-
A second real measurement, this time isolating precision instead of hardware. Same NVIDIA A100, same Qwen3-8B, same everything — except one run served in fp16, the other in bf16. No hardware changed at all. Result: 29.5% of prompts produced different output text. That's the same divergence rate we measured moving to a completely different vendor (AMD MI300X). Precision alone accounts for about as much "noise" as switching silicon vendors entirely. Full data: https://ruitong.io — 第二组真实测量,这次隔离的变量是精度,而非硬件。 同一块 NVIDIA A100、同一个 Qwen3-8B、其余条件完全一致——唯一区别是一次以 fp16 提供服务,另一次以 bf16。硬件本身没有任何变化。 结果:61 条提示词中有 29.5% 输出文本不同——这与换成完全不同厂商(AMD MI300X)时测得的分叉率几乎一致。仅精度差异产生的"噪声",就已经接近换一个芯片厂商所产生的幅度。 完整数据:https://ruitong.io
To view or add a comment, sign in
-
https://lnkd.in/gf7tYGxT "Nvidia says it will be able to produce up to 1,000 Vera Rubin racks ~PER DAY~. This would be 72000 Rubin GPUs per day. This would be over 2 million Rubin GPUs per month.It represents future peak capability once TSMC N3 wafers, CoWoS advanced packaging, HBM4 supply, power delivery, and data center infrastructure all align at scale. When this happens Nvidia would generate over $630 Billion in revenue per quarter for Nvidia & manufacturing partners." And, as these are deployed, the implications for our world is huge.
To view or add a comment, sign in
-
Frore LiquidJet Claim Targets Rubin GPU Cooling Efficiency Tom's Hardware reported that Frore Systems says its LiquidJet coldplate could lower Nvidia Rubin GPU junction temperatures by 6C to 12C in an analytical thermal model. Frore Systems' white paper claims a 10C reduction would raise token generation effic... Full report: https://lnkd.in/gsq7Jqia #CloudComputing #AIInfrastructure #EnterpriseTech #Semiconductors
To view or add a comment, sign in
-
NVIDIA is not building faster GPUs. It is building an AI token factory, and every architectural decision flows from one objective: produce tokens at the lowest possible cost per unit. From $1.95 per million tokens on Hopper to a projected $0.10 on Feynman. How? smashing the Von Neumann memory bottleneck, offloading 30% of infrastructure overhead via BlueField DPUs, mandating liquid cooling for 200kW+ racks, and transitioning to 800V DC power to eliminate conversion losses. At CloudSyntrix, we're already seeing hardware conversations shift toward this reality.
To view or add a comment, sign in
-
-
⚡ NVIDIA's Nemotron 3.5 Lightning dropped yesterday. Had it running on my Dell GB10 the same day, curious to see how it'd actually perform. First impression: it's fast. ~115 tokens/sec decode (single-stream) on real diagnostic tasks, thanks to DSpark — NVIDIA's speculative decoder, hand-built for this exact chip. 💡But instead of "how fast is it", I wanted to test something more meaningful: Can it actually solve a realistic ML debugging problem, and how efficiently does it reason through it? 🧩 The task: Loss stuck at ln(1000) for 200K iterations. The agent has tools to inspect the loss curve, gradients, LR schedule, and config. 🕵️ The real bug: scheduler.step() called every micro-batch instead of every optimizer step with 8× gradient accumulation, the LR silently decayed to zero. Subtle, not a toy question — the kind of bug that hides behind a normal-looking config ⚔️ Nemotron vs Qwen3.6 35B-A3B NVIDIA's own comparator, same ~3B active-parameter MoE class. Both models correctly diagnosed the bug. But: 🧠 Nemotron: 571 tokens 🧠 Qwen: 1,839 tokens ~3.2× fewer tokens for the same correct answer. 🎯 Takeaway: Accuracy alone isn't enough for agentic workloads. At production scale, token efficiency is what drives inference cost, throughput, and scalability. Not just "did it solve it", "how efficiently did it get there?" #NVIDIA #Nemotron #AIAgents #DellProPrecision #DellTech #DellProMax
To view or add a comment, sign in
-
-
NVIDIA's open-source Nemotron 3.5 Lightning delivers 4x faster throughput while its Switchyard router cuts agent costs to a third "That is the power of a system of models, matching the right model to each step of the workflow," Kari Ann Briski said.
To view or add a comment, sign in
Explore content categories
- Career
- Productivity
- Finance
- Soft Skills & Emotional Intelligence
- Project Management
- Education
- Technology
- Leadership
- Ecommerce
- User Experience
- Recruitment & HR
- Customer Experience
- Real Estate
- Marketing
- Sales
- Retail & Merchandising
- Science
- Supply Chain Management
- Future Of Work
- Consulting
- Writing
- Economics
- Artificial Intelligence
- Employee Experience
- Workplace Trends
- Fundraising
- Networking
- Corporate Social Responsibility
- Negotiation
- Communication
- Engineering
- Hospitality & Tourism
- Business Strategy
- Change Management
- Organizational Culture
- Design
- Innovation
- Event Planning
- Training & Development