Disk fast for a few seconds then a cliff? Detecting I/O cache fraud
Many “NVMe” plans ace short benches and fail under real writes. Classic trick: scores ride page/device cache, then throughput and latency fall off a cliff.
Why short tests lie
Multi-second fio runs often hit page cache or write cache — you measure cache speed, not sustained media capability.
On oversold hosts, shared IO makes the cliff sharper: full speed briefly, then await spikes and IOPS collapse.
Measure sustained, not the opening seconds
At minimum:
- Run sequential/random writes for tens of seconds to minutes
- Keep a time series: does throughput drop after cache exhaustion?
- Compare claimed media (NVMe/SSD/HDD) to measured reality
- Test peak and off-peak to filter noise
Read it with steal time
Disk fraud and CPU overselling often co-exist. CPU steal can worsen IO; IO stalls can look like “app slowness.”
Full picture = steal time + sustained disk curve + hardware generation + price.
How CloudWorth exposes it
The probe’s disk bench and timed sampling target cache cliffs and fold them into an audit report you can show vendors.
Pair with steal and FinOps modules for one exportable conclusion — start at getcloudworth.com.