In short
The same number of operations can execute at very different speeds, and on new hardware the relationship turned out to be even less predictable. The breakdown shows why AI efficiency estimates need real measurements rather than a look at FLOPs alone.
If two models perform the same number of operations, that does not yet mean they run in the same time. A new replication of the study on the α-FLOPs formula confirms the main thesis: raw FLOPs are poorly suited to estimating real execution speed.
The reason is that operations parallelise differently. The spatial dimensions of a computation usually scale more easily than kernel dimensions, so the same formal amount of work can take different times on one and the same accelerator.
But the most interesting part begins on new hardware. The replication's authors found jumps and oscillations in execution time — not the smooth dependency convenient to describe with a single formula. α-FLOPs generally underestimates real time, turning from a useful correction into an unreliable forecast.
The practical conclusion is unwelcome but important: FLOPs can serve as a rough guide, yet model efficiency has to be compared by measurements on specific hardware. Especially when it comes to choosing an architecture for production, where what matters is not abstract operations but latency, accelerator utilisation and the cost of running.
There is a separate lesson for academic work too. In reproducing the study the authors lacked detail about dependencies and transparency of the data for the regression — gaps of exactly that kind make verifying results markedly harder. For their own implementation they prepared a full replication package, but the story itself shows that without code, environment and source data an efficiency result stays tied to one particular system and is poorly verifiable.
Are you ready to trust FLOPs when choosing a model if the real time on your hardware can behave in jumps? Source: cs.AI updates on arXiv.org