Just Give Me A Second … (We’re Just Floppin’ Around)

A second. Seems harmless enough. Just a little time to get things right. Blink of an eye. Snap of a finger. Seconds seem precious some of the time (we still measure our best athletes in seconds, breaking records past generations thought were impossible). We measure our work sometimes in second/minutes/hours – always trying to shave a little time off to increase efficiency and performance. So, saying to the grandkids, “Sure, just give me a few seconds; today is an eternity.) The other day I was reading an article about FLOPS – A “flop” (or FLOPS) isn’t a physical size but a measure of computer speed. FLOPS stands for Floating Point Operations Per Second, indicating how many complex decimal calculations a computer processor can do in one second. And it blew my mind.  I remember when a “flop” was when a Broadway show or product or service or movie that just fell flat – (history is full of them, like clear Pepsi, Cheetos Lip Balm, Solyndra, DeLorean cars, early Google Glass, Carrie, the musical, CATS the movie (along with CATS the musical (Not a flop, just awful!), Harley Davidson cologne, and most of McDonald’s efforts in new products … (what actually is in a McRib – although very weird, still weirdly good??). Nowadays, FLOPS are super cool, and super-fast. I mean incomprehensibly fast!! I did some digging and learned a whole bunch of cool things about how computational power is just going nuts (Remember, I am a child of AOL CD’s and dial phone up log in’s – for my older readers, we can still remember the sounds: “beep, buzz, zing, aaarrrrr, zap” and then the connection) … so stepping into this arena is quite a rush. 

What Are FLOPS – FLOPS stands for Floating-Point Operations Per Second. While the term sounds technical, the idea behind it is simple: FLOPS measure how many calculations a computer can perform every second. It’s a way of describing raw computing power—how much mathematical work a machine can get done in a very small amount of time.

Why FLOPS Became a Useful Measure – In the earliest days of computing, speed was easy to describe. A computer might perform thousands or millions of calculations per second, and that alone was impressive. But as computers began tackling scientific and engineering challenges, those rough descriptions stopped being useful.

Researchers needed a consistent way to compare machines across different designs and purposes. FLOPS became that yardstick. While clock speed tells you how fast a processor ticks, FLOPS tell you how much meaningful work it accomplishes with each tick. Two computers may run at the same clock speed but deliver very different FLOPS depending on how efficiently they handle math.

A Brief History of Exploding Computing Speed – In the 1970s and early 1980s, powerful computers measured performance in thousands or millions of FLOPS. These machines were rare, expensive, and often filled entire rooms. Access to that level of computing power was limited to governments, universities, and large research labs. Computers began with the abacus back in 1100 BCE.

By the 1990s, gigaflops—billions of calculations per second—became achievable. This marked a major leap forward and enabled breakthroughs in graphics, engineering, and scientific modeling that had previously been out of reach.

The early 2000s brought teraflops, or trillions of calculations per second. At the time, this level of performance felt almost unreal. It powered detailed climate models, advanced physics research, and increasingly realistic digital simulations.

In the 2010s, supercomputers crossed into petaflops, performing quadrillions of calculations every second. Today, the most advanced systems and AI-focused chips operate at the exaflop scale—quintillions of calculations per second, with yetaflops just around the corner. At that point, the numbers stop feeling like quantities and start feeling like abstractions. Here’s a Breakdown: (buckle up, ‘cause we’re going on a speedy ride … )

MegaFLOPS (MFLOPS):  Millions (10^6) of operations per second.

GigaFLOPS (GFLOPS):  Billions (10^9) of operations per second (e.g., older gaming PCs).

TeraFLOPS (TFLOPS):  Trillions (10^12) of operations per second (e.g., modern gaming GPUs, game consoles like PS4).

PetaFLOPS (PFLOPS):  Quadrillions (10^15) of operations per second (e.g., supercomputers).

ExaFLOPS (EFLOPS):  Quintillions (10^18) (e.g., the fastest supercomputers today).

YetaFLOPS… oh, just stop! 1million to the 24th power

Why My Brain Struggles to Comprehend These Speeds – One reason FLOPS are so hard to grasp is that human intuition and reasoning were never designed for numbers this large. A simple analogy helps illustrate the gap:

  • If one calculation took one second …
  • A million calculations would take about 11 days
  • A billion would take roughly 32 years
  • A trillion would stretch to about 32,000 years
  • A quintillion would take more than 30 billion years—longer than the age of the universe

Modern computers perform all of that work in a single second! (Go ahead and blink – and you did a calculation of quintillion – WHAT??? )The scale is so extreme that it breaks our natural sense of time and effort. (“Jackie, hang on, I’ll be there in a second…”).

Why FLOPS Matter So Much for AI – Artificial intelligence is, at its core, math performed at enormous scale. Training an AI model involves adjusting billions or even trillions of numerical values, each requiring repeated floating-point calculations. The faster a system can perform those operations, the more complex and capable the model can become.

This is why computing power has become such a central topic in discussions about AI. Higher FLOPS don’t just make AI systems faster; they make entirely new applications possible. Problems that once took months or years to compute can now be solved in hours…or minutes… or seconds!

Modern Chips: Speed Through Parallelism – Today’s computing leap isn’t just about making a single processor run faster. Modern chips achieve their performance by doing many things at once. Thousands of small processing units work in parallel, each handling a portion of the overall workload.

These chips are also designed specifically for the kinds of math AI requires, making them far more efficient than general-purpose processors. The result is not just speed, but a new way of computing—less like one brain thinking faster and more like thousands of brains working together in perfect coordination. AI doesn’t “think” like a human – it calculates. The faster it can perform math, the faster it can recognize patterns, generate language, or analyze images. AI progress isn’t just smarter algorithms; it’s raw computational speed, measured in FLOPS. More FLOPS means bigger models, faster answers, and more capable systems.

Hard to imagine, but one chip can only go so fast- a single computer chip, no matter how advanced, has physical limits. Heat, power, and size constrain how fast it can run. Even the most powerful AI chips can only deliver so many FLOPS on their own. To go faster, the industry stopped trying to build one super-chip and started connecting many chips together.

Parallel Processing: Many Chips, One Problem – Modern AI systems split huge problems into smaller pieces and run them in parallel. Thousands of chips work at the same time, each handling part of the math. Think of it like 10,000 people solving different rows of the same spreadsheet simultaneously – the work finishes far faster than one person doing it alone. 

Now, on to Scale – When chips are connected correctly, their performance stacks.
10 chips at 100 TFLOPS ≈ 1 PFLOP, 10,000 chips ≈ 1 EFLOP, 100,00 chips – YIKES!
This is how AI systems reach mind-bending speeds — not by one miracle chip, but by orchestrating armies of processors. That’s why places that house them are so massive.

Ok, so let’s put them in Buildings – AI chips are mounted inside servers. Servers are stacked into racks. Racks fill rooms. Rooms fill buildings. At the largest scale, multiple buildings operate as a single system. Each layer exists to move data fast enough so the chips don’t sit idle waiting for information. But raw FLOPS mean nothing if chips can’t talk to each other quickly. That’s why AI data centers invest heavily in ultra-fast networking – miles of fiber-optic connections – so results from one chip reach another in microseconds. Much of the cost isn’t computing power but trying to move data at extreme speed.

Why These Data Centers Are So Big – AI chips consume enormous power and generate intense heat. Cooling systems, power distribution, backup generators, and networking equipment often take up as much space as the computers themselves. The buildings exist as much to support the chips as to house them. At the highest level, AI computing happens across campuses of data centers (we have a campus here at KHT – but our cooling is simply walking outside and water breaks!!) Workloads are distributed across buildings, cities, and even regions. From the outside it looks like real estate. Inside, it functions like one massive, synchronized computer. Think of it as your nervous system in your body, all talking at once.

Now Mega-Factories – These facilities are the factories of the AI age. Instead of stamping steel or assembling cars, they manufacture computational output – FLOPS at ridiculous scale. The companies that control the fastest, largest computer systems gain enormous advantages in AI capability, speed, and cost. So, when you hear about trillion-dollar investments in AI infrastructure, you’re not hearing about software hype – you’re hearing about a global race to build machines capable of performing unimaginable amounts of math, every second, by making thousands of chips behave like one mind.

OK, Now Blow My Mind … Here’s an expanded breakdown of estimated computing power (FLOPS) for the biggest AI compute “factories” today. Keep in mind, these are broad estimates … (can you imagine ten years from now??)

  1. xAI Colossus (Memphis, Tennessee, USA)

Compute Setup & Scale

  • Initially launched with ~100,000 Nvidia GPUs, expanded to ~200,000 and continuing to scale toward ~1 million GPUs. 
  • xAI’s long-term vision includes up to ~50 million H100-equivalent GPUs industry-wide by 2030 (50 ExaFLOPS of FP16/BF16 training compute. (wait – 50 quintillion …what?)

FLOPS Estimate

  • single Nvidia H100 GPU delivers ~1 petaFLOPS (1×10¹⁵ FLOPS) peak in FP16/BF16 AI workloads. 
  • Current 200k GPU cluster (Colossus):
    ~200,000 GPUs × ~1 PFLOPS ≈ 200 petaFLOPS (0.2 exaFLOPS) theoretical peak.
    (Real-world sustained training performance will be lower.)
  • If Colossus scales toward ~1 million GPUs:
    ~1,000,000 GPUs × ~1 PFLOPS ≈ ~1 exaFLOPS theoretical peak (mixed precision).
  • Long-term industry goal (all sites and next-gen chips): 50 exaFLOPS of AI training compute by ~2030.  (wonder how many extension cords they’re gonna need??)
  1. Microsoft Fairwater AI Datacenter (Wisconsin, USA)

Compute Scale

  • Expected to house hundreds of thousands to millions of next-gen Nvidia GPUs (GB200/GB300 and future Blackwell/Rubin units). 
  • Predicted to reach ~3.3 GW of total electrical capacity at full buildout by late 2027 — suggesting scale beyond 300,000–500,000 GPUs and potentially into millions. 

FLOPS Estimate

  • GPU capabilities vary by generation (e.g., GB200/GB300 and soon Rubin/Rubin Ultra), but many current high-end GPU accelerators exceed ~2–4 petaFLOPS of mixed-precision AI throughput per unit. 
  • Conservative cluster estimate: ~500,000 GPUs × ~2 PFLOPS ≈ ~1 exaFLOPS theoretical peak or higher.
  • With future GPUs and denser scaling (e.g., millions of units), theoretical capabilities could push toward several exaFLOPS of aggregate AI throughput.

Why This Matters: Microsoft itself claims Fairwater will deliver 10× the performance of the fastest current supercomputers — which today operate in the low exaFLOPS range, implying a cluster performance well into the exaFLOPS tier or beyond. 

  1. Meta Hyperion AI Campus (Louisiana, USA) – estimated

Planned Scale

  • Projected to become the largest AI compute campus on Earth with ~5 GW of power capacity once fully built (2030+). 
  • Expected footprint: ~2,250 acres of datacenter facilities (that’s over 1700 football fields…including end zones of course!!)

FLOPS Estimate (Future)

  • With projected GPU counts possibly in the millions, and next-gen AI accelerators (Rubin, Rubin Ultra, etc.) delivering ~5–10+ petaFLOPS each:
    • 2 million GPUs × ~5 PFLOPS ≈ ~10 exaFLOPS theoretical peak
    • 5 million GPUs × ~5 PFLOPS ≈ ~25 exaFLOPS theoretical peak

These are high-level illustrative estimates — actual aggregate FLOPS depends on the mix of chips, utilization, and precision (FP8/FP16 vs. FP32). Nonetheless, the projected compute capacity aligns with a facility that could easily exceed tens of exaFLOPS at peak theoretical scale. Now we know why other countries are also racing to get chips and stay in the game.

Now, consider this … AI clusters’ FLOPS growth is exponential: performance roughly doubles every ~9 months as hardware and network efficiency improve.  By the time the buildings are built and operational, the hardware is somewhat obsolete inside … YIKES! – and when a massive structure is “slow”, what becomes of it??

For me, all this is amazingly exciting along with being a little bit unnerving or flat out scary!! 

I’m exhausted and need to lie down for a few seconds.

How did you do on our last logo contest?

Check out our logo guide for the “Mixed Results” post here!

0 replies

Leave a Reply

Want to join the discussion?
Feel free to contribute!

Leave a Reply

Your email address will not be published. Required fields are marked *

Please prove you aren't a robot: * Time limit is exhausted. Please reload CAPTCHA.