AWS Hints at Cost-Saving Technology – Yell, Dude!

Glancing Behind the AWS Veil

Last month, I found myself in a networking lab alongside AWS VP of Global Network Engineering Matt Rehder, and I made a quip about cage nuts tearing up your hands. I wasn’t ready for the blank look that met my joke, but perhaps I should have been. Not because it lacked humor (my jokes are comedic gold), but due to the fact that the context of the joke has faded away for them. Racks arrive pre-assembled, and evidently, no one is bolting hardware into place within an AWS datacenter anymore. It was a rare glimpse into a realm that drives all our cloud activities, yet astonishingly few are aware of its existence. My fellow GadgetLad colleague Thomas Claburn previously toured the lab and provided us with a thorough examination of AWS’s documentation on this. To summarize, “a flat, single-tier network arranged in a purposefully quasi-random fashion results in a significantly more resilient network at a much lower operational cost.” Now you’re informed. The element I believed warranted more focus (which is why I invited myself to take a tour of an AWS facility, much to AWS’s astonishment) is the economic aspect and its effects on AWS clientele. In essence, their novel networking strategy can achieve up to a 40 percent increase in energy efficiency, and it’s the standard for most new datacenter constructions (and dear reader, they are currently engaged in numerous builds), while remaining completely unnoticed both inside and outside of the grandest bookstore on the planet. The real savings and cost benefits are substantial — so where do they go?

Abundant Cost Savings

Seventeen years prior, Amazon SVP James Hamilton took to the stage and remarked [PDF] that network vendors operated businesses with margins akin to those of mainframe vendors, who in turn were similar to starving pigs at a trough (a vivid metaphor courtesy of me, but also… quite accurate). His primary concern distilled down to “they’re obstructing my path,” prompting him and his team to take action. It turns out that “their course of action” involved nothing short of entirely reconstructing the networking stack on commodity hardware.

Slashing Networking Expenditures

This brings us to just a few years back. Their networking cost efficiencies had already plunged well below what they would have been if reliant on traditional networking providers, and then their Resilient Network Graph shufflebox initiative further reduced networking costs. So… where have those savings gone over recent years? I directly inquired if it was reasonable to claim that they retained the margin improvement instead of returning it to customers. Their response was clear: “From a cost perspective? Effectively, yes.”

The Economics of AWS Cloud

It’s tempting to view that response as disheartening, but I suggest that if that’s your initial thought, you may not have been following broader industry movements these past few years. Unlike its rivals, AWS has not increased prices on existing SKUs. (Yes, the costs for GPU capacity blocks have escalated quarterly, but that’s how those were designed to function, resembling a sluggish version of Spot Instances. And yes, a few years back, they began charging per public IPv4 address.) You can create the same instance with 64GB of RAM that you could in 2020 and pay the same amount today; that’s not even raising a price to keep pace with inflation! While there is a 9 percent increase when transitioning from a Graviton4-based c8g.2xlarge to its c9g.2xlarge Graviton5 counterpart, nothing compels you to upgrade. Quite frankly, given the rising costs of components, I find it astonishing that they managed to limit the increase to just 9 percent.

The Intricacies of AWS Billing

Moreover, despite the bewildering level of granularity present in an AWS bill with its multitude of SKUs, it’s essential to recognize that each SKU conceals an overwhelming amount of complexity. An EC2 instance incurs charges per hour (yes, it’s metered by the second; please don’t email me about it), but that single charge encompasses the CPU, the RAM, the backplane, the electricity, the building’s physical security, the IAM security features that prevent it from resembling a public computer or an Azure instance, the personnel required to build and maintain these systems, and an impressive amount of networking wizardry.

A Fundamental Shift in Strategy

One of the facets of networking wizardry is that while the cost to transfer data between Availability Zones can be notably steep (in major regions, each gigabyte has a list price of one cent in and one cent out, totaling “two cents per gigabyte,” which adds up), transferring data within an AZ remains free. This is remarkable when you consider how many distinct datacenter facilities can constitute a single AZ. It’s even more impressive upon realizing that AZs are enlarging, data volumes are surging, and that AWS has thus far chosen to absorb the associated costs instead of passing them on. Moving data out of AWS is substantially pricier. Their now-obsolete Snowball devices allowed you to transport data back and forth to AWS. Sending one in incurred a set charge, while sending one out included that charge plus a per-GB fee. Annoying, wouldn’t you agree?

The Emergence of AWS Interconnect

That situation transformed over the past year with the introduction of AWS Interconnect – multicloud, which charges nothing per GB, just a per-port hourly rate. There’s a free tier granting one 500 Mbps port per provider, effectively eliminating the AWS-side cost for transmitting that traffic to GCP, Oracle, and soon Azure. For perspective, it would roughly cost $12,000 to transmit a month’s worth of usage for that port across the public internet. To put it plainly, this is an unprecedented move on the AWS side, and it’s quite intriguing. Although it isn’t the primary focus for those I consulted, I still posed a question about it for an official response: “We understand customers prefer flat-rate network pricing because it simplifies cost predictability, which is why we’re shifting toward flat-rate pricing for new network offerings.”

How Difficult Could It Be? Extremely Difficult

The challenge lies in how much of what AWS does goes unnoticed. For instance, why am I sharing this information instead of AWS? AWS admits they are typically poor at conveying their own narrative, yet they clearly wish to share it. Historically, they’ve deferred to customers to voice their own experiences, believing that the advantages will accrue to them.

Hidden Networking Mysteries

Here’s my preferred illustration: the challenging aspect of a quasi-random network lies not in the cabling but in the routing. Every protocol you’ve heard of (BGP, OSPF, some Cisco-specific protocols that sane individuals don’t utilize, etc.) calculates shortest paths, and “shortest path” loses its meaning when there are nearly thousands of equivalent routes between any two points. AWS addressed that with a protocol named SIDR (Scalable Intent-Driven Routing, pronounced “cider,” similar to the CIDR method for IP address allocation, showcasing their flair for naming things) which governs the control plane, while Spraypoint manages the actual forwarding path. They unveiled SIDR publicly at re:Invent in 2023, followed up with it at the “Monday Night Live” re:Invent keynote in 2024. Then, they released the RNG paper, which doesn’t mention SIDR whatsoever. To my knowledge, no one outside the company has ever connected the protocol to the network it enables. I had to inquire to ensure I wasn’t overlooking something. I’m correct. It’s not a secret! They simply haven’t mentioned its vital role in the functioning of the entire structure. Unfortunately, their storytelling shortcomings regarding these aspects are not serving them well.

The Market’s Perspective

The perception skews the other way: the market thinks that the neoclouds and Nvidia reference architectures hold an advantage. They do not. You can construct a neocloud with minimal understanding of networking, and certain entities clearly have. Numerous datacenter networks are managed by utter incompetents. I’m aware of this; I used to belong to that group before I had an epiphany / realized I could engage a cloud provider to handle these challenges, making them Matt Rehder’s responsibility. AWS has become remarkably adept at rendering network complexities invisible, to the extent that we seldom pause to appreciate how extraordinary it is that you can achieve nearly full line-rate transfers between any two points in AWS, and it simply functions. In the datacenters we might establish, you confront amusing issues like “the switch at the top of the rack can only handle a limited amount of traffic, so not every node can communicate at full speed consistently.” I have never encountered a real-world scenario of