Blast from the Future: Tensordyne’s Wild Mathematical Strategy
The AI infrastructure startup Tensordyne is on an ambitious journey, and their inaugural commercial accelerator is prepared to launch, developed using TSMC’s advanced 3nm technology. They’ve partnered with Juniper Networks and Broadcom, guaranteeing that their systems will provide enhanced throughput while consuming less power than standard GPUs. By leveraging logarithms – those tricky concepts from math class – they believe they’ve simplified the handling of AI workloads. Logs transform multiplication into a straightforward addition problem. Instead of a*b, it’s log(a) + log(b). The challenge lies in converting those numbers into logs and back without causing chaos.
Logarithmic Fun: The Mitchell Approximation
Here’s the twist: opting for a lookup table (LUT) would have been the easy way out, but Gilles Backhus, one of the key innovators behind Tensordyne, informs GadgetLad that this was ruled out due to size constraints. Instead, they’ve fully committed to a heuristic known as the Mitchell approximation, which approximates the log and antilog for each value. Indeed, it’s merely an approximation and can introduce some errors. To mitigate this, they’ve incorporated a section-wise correction system in the hardware which provides accuracy comparable to FP16. Napier will also explore FP8 and 4-bit block floating data types, enabling the MAC unit to function cleverly without performing traditional multiplication.
Napier: The Powerhouse Under the Hood
Tensordyne’s inaugural commercial chip, Napier, showcases specifications one would expect from a high-end GPU from the past. It boasts a 300-watt TDP, 144 GB of HBM3e distributed across four stacks, 4.7 TB/s memory bandwidth, and has the capability to produce up to 2.1 petaFLOPS of dense FP8 performance. In terms of power consumption, it devours nearly 60 percent less than the Nvidia H200 accelerators. But bear in mind, maximum FLOPS often don’t equate to peak FLOPS, so take this with a pinch of skepticism. We will truly assess Napier in comparison to Nvidia or AMD’s latest GPU offerings next year.
The TDN72: A Powerful System
The TDN72 system stands strong with eight air-cooled compute blades, a 10-core Intel Xeon-D host CPU, and nine Napier accelerators. These chips connect through a high-speed interconnect fabric akin to Nvidia’s GB200 NVL72 systems. It’s more compact, avoids the complexities of liquid cooling, and is designed to fit seamlessly in older datacenters. Up to four 30 kW TDN72 systems can be housed in a 52U rack, delivering 608 petaFLOPS within a 120 kW footprint. That translates to approximately 1.68x more dense FP8 compute per rack than Nvidia’s GB200 NVL72. However, resist the temptation to link peak FLOPS too closely.
Software Smart: Making It User-Friendly
From its inception, Tensordyne has gone above and beyond to ensure its software is accessible for clients. The prototype initially lacked the error correction found in the Napier chips and wasn’t well-suited for users targeting trillion-parameter models. Now, the software has evolved, allowing the compiler to seamlessly convert existing models for the new hardware. Tensordyne has also provided a proprietary serving platform and a runtime environment for utilizing popular inference servers like vLLM. PyTorch compatibility is also on the horizon. Ambitious claims are being made, with expectations of processing more than 1,000 tokens per second. Neocloud providers such as Cirrascale and BlueSky Compute have already expressed interest. However, as AMD can confirm, software can significantly influence success.
Summary: Logging In and Preparing for Launch
“Logs: Not Exclusively for Lumberjacks”: Here we are, as Tensordyne emphasizes its ability to convert traditional arithmetic into a computing transformation. If they succeed, Nvidia will need to be cautious. The chips are projected for release in mid-2027, so it’s a matter of staying calm and anticipating the silicon upheaval.