By the spring of 2022, a strange thing had happened to the computer industry. The most important model in artificial intelligence was no longer a secret research idea — it was a shape. The Transformer, introduced in a 2017 Google paper titled “Attention Is All You Need,” had become the beating heart of nearly every ambitious AI system: language models, translation, recommender engines, protein folding. And the workloads that shape produced — enormous stacks of matrix multiplications, repeated billions of times — had started to strain even NVIDIA’s formidable Ampere GPUs.
So NVIDIA did something it had never done quite so explicitly before. It built a chip for the Transformer. On March 22, 2022, at its GTC conference, the company unveiled the Hopper architecture and its first product, the H100. It was named for Grace Hopper, the U.S. Navy rear admiral and computing pioneer who helped invent the compiler — a fitting namesake for a machine designed to translate mathematics into intelligence at industrial scale.
A chip that changes its own math
The headline feature was the Transformer Engine, and to understand why it mattered you have to understand a trade-off that had haunted deep learning for years: precision versus speed. Represent numbers with more bits and you get accuracy but pay in memory and time; use fewer bits and everything runs faster but risks falling apart. Hopper’s answer was to refuse to choose.
The Transformer Engine introduced support for an 8-bit floating-point format, FP8, and paired it with logic that watches the statistics flowing through each layer of a network. The heavy lifting — the giant matrix multiplications inside attention and the feed-forward blocks — could run in fast, cheap FP8. The delicate operations, where rounding errors compound, stayed in 16-bit precision. The engine tuned this balance automatically, layer by layer, so that developers got the speed of low precision without watching their model’s accuracy quietly rot. NVIDIA claimed the result could accelerate Transformer training and inference by up to six times over the previous generation, without loss of accuracy.
Eighty billion transistors of ambition
The silicon underneath was staggering. The H100 packed 80 billion transistors onto a die built using a custom TSMC 4N process — up from Ampere’s 54 billion. Jensen Huang, NVIDIA’s founder and CEO, framed it in language that would define the company’s next chapter. “Data centers are becoming AI factories,” he said, “processing and refining mountains of data to produce intelligence. NVIDIA H100 is the engine of the world’s AI infrastructure.”
Nearly every subsystem was rebuilt to feed that engine. The H100 was the first GPU to use HBM3 memory, delivering three terabytes per second of bandwidth, and the first to support PCIe Gen5. A fourth-generation NVLink interconnect moved 900 gigabytes per second between chips, and a new external NVLink Switch could stitch together as many as 256 H100s into a single high-speed domain — effectively turning a rack of GPUs into one gigantic accelerator. There was more: a Tensor Memory Accelerator to shuttle data without burning compute, thread-block clusters for finer-grained parallelism, second-generation Multi-Instance GPU to safely slice one card into seven, and even confidential-computing features to protect models and data while they were being processed.
The engine of the boom
Timing, as ever, was everything. Hopper was revealed in March 2022 and the H100 shipped later that year. Just weeks after it reached customers in volume, a chatbot called ChatGPT arrived and the world discovered, all at once, what large language models could do. Suddenly every cloud provider, startup, and research lab on earth wanted as many H100s as they could buy — and NVIDIA had built, almost prophetically, the exact machine the moment demanded.
The H100 would become one of the most sought-after pieces of hardware on the planet, its lead times stretching into quarters, its presence in a company’s data center a signal of serious AI intent. NVIDIA had spent years betting that the Transformer would matter. With Hopper, that bet stopped being a gamble and started being an empire.
Next: the moment the numbers went vertical — when NVIDIA’s data-center business exploded and the company raced toward a valuation few chipmakers had ever imagined.
Comments
Post a Comment