Home / Gotcha guides / Cloud Server Memory Bandwidth Detection: Why Does Speed Drop Even When Capacity Meets Spec?

Cloud Server Memory Bandwidth Detection: Why Does Speed Drop Even When Capacity Meets Spec?

Enough capacity does not mean enough bandwidth; only three-dimensional forensics can reveal the truth.

Updated 2026-10-07 · CloudWorth

Memory BandwidthSTREAMNUMAOverselling DetectionFinOpsCloudWorthSteal TimeVPS oversellingVPS benchmark

Cloud Server Memory Bandwidth Detection: Why Does Speed Drop Even When Capacity Meets Spec?

Capacity meeting spec does not equal bandwidth meeting spec. Use STREAM plus Steal Time and NUMA pinning for cross-forensics, then compare the premium rate to decide whether to upgrade or request a refund.

Test Capacity Before Testing Bandwidth

When you buy a cloud server, almost everyone's first glance goes to memory capacity: 8GB, 16GB, 32GB. Capacity is written on the order page—visible, verifiable—so it gets assumed to mean "the memory is fine." But capacity is just the warehouse floor space, while bandwidth is the forklift speed—an 8GB instance can have capacity showing all green while memory throughput is only half the nominal figure for its generation.

This is the classic scenario of meeting capacity specs while dropping speed: workloads like AI inference, vector search, large Redis values, and JVM heap copies are far more sensitive to GB/s than to GB, and having enough capacity actually means hitting the bandwidth ceiling first.

Start by separating the two steps:

  1. Capacity side: use free -h to check total/available, then compare against the order specs to confirm nothing has been reclaimed by a hidden balloon driver. Under KVM you can check dmesg | grep -i balloon.
  2. Bandwidth side: matching capacity does not mean matching bandwidth—you need a dedicated memory throughput test, covered in the next section.

The order must be capacity first, bandwidth second. Capacity falling short is false advertising, so go straight to a refund; capacity matching while bandwidth drops is the more insidious problem of overselling / throttling, and it requires the evidence chain below.

A useful rule of thumb: if you bought an instance with "the same capacity but noticeably a tier cheaper," assume bandwidth is a suspect by default rather than assuming you got a bargain. This kind of cheapness usually comes from CPU oversubscription, cross-node NUMA placement, or downclocked memory frequency, and it ultimately bites back in the form of a price-performance premium rate—you paid for the capacity, but not the throughput.

After recording the raw data (instance specs, region, image, billing type, order time), move on to actual testing. The record itself is evidence for later support tickets and downsizing decisions; a ticket that only says "the memory is very slow" will almost never get an effective response.

STREAM and sysbench Hands-On Testing

Memory bandwidth testing has two complementary tools: STREAM measures sustainable vector throughput (Copy/Scale/Add/Triad), while sysbench memory measures a scenario closer to “small-block read/write.” The debate over sysbench vs STREAM memory bandwidth test for cloud servers is pointless; you need to run both, because their distortion directions differ: STREAM is too friendly to the L3 cache of small instances, and sysbench is too sensitive to instruction overhead. Only by cross-checking them will you avoid being fooled by a single number.

STREAM usage (you can compile it into your home directory without root):

sudo apt install -y gcc gfortran make
# 下载 stream.c 后:
gcc -O3 -fopenmp -DSTREAM_ARRAY_SIZE=200000000 stream.c -o stream
OMP_NUM_THREADS=$(nproc) ./stream

Set the array size to 1/4–1/2 of memory. If it is too small, everything will land in cache and you will get “inflated bandwidth”; otherwise you are measuring L3, not DRAM.

sysbench side:

sysbench memory --memory-block-size=1M --memory-total-size=20G --memory-oper=read run
sysbench memory --memory-block-size=1M --memory-total-size=20G --memory-oper=write run

There are three key criteria:

  • Single-core vs full-core: run taskset -c 0 and all cores respectively, and record GB/s. If full-core barely increases, bandwidth is already capped, not cores insufficient.
  • Compare against nominal: the generational gap between DDR4 and DDR5 is roughly 1.5–2x. Among cloud instances of the same generation and frequency, there should not be a two- or threefold difference; if there is, something is wrong.
  • Convert to per-vCPU bandwidth: total bandwidth / vCPU count is the most practical metric for judging memory bandwidth oversubscription. The per-vCPU bandwidth of shared instances is often only half that of dedicated instances, which is one of the real sources of cloud server premium rate.

Run three times and take the median, avoiding the top-of-the-hour neighbor peak. To automatically convert these numbers into per-vCPU bandwidth and premium rate, you can create a template in /app, enter STREAM, sysbench, and instance price together, and generate a comparable price-performance profile. The next section uses Steal Time and NUMA to attribute “slowdown” to specific causes.

Steal Time and NUMA Forensics

The previous section can only prove that “speed has dropped”; attribution still depends on two other metrics: Steal Time and NUMA topology. This is also the layer most easily skipped in cloud server memory bandwidth testing—capacity is sufficient and a single run meets the target, yet the bottleneck hides in scheduling and memory affinity. Spec sheets never list these two items, and the gap in real benchmark scores often comes from exactly here.

First, look at st:

vmstat 1 10        # watch the st column (Steal Time)
mpstat -P ALL 1    # per-core %steal

If st remains >3% over time and moves in sync with bandwidth decline, it means vCPU time is being borrowed by neighbors. On shared KVM, CPU steal correlates very strongly with memory bandwidth; what drops is the entire path, not just compute capacity.

Next, look at NUMA:

lscpu | grep -i numa
numactl --hardware
numactl --cpunodebind=0 --membind=0 \
  sysbench memory --memory-block-size=1M --memory-total-size=20G --memory-oper=read run

If bandwidth rises noticeably after binding the process to the local node, it is cross-node penalty (commonly 20%–40%), not overselling; if it still hugs the per-user bandwidth ceiling after binding, then it is genuine throttling.

Take three screenshots—st, NUMA topology, and bandwidth before and after binding—and open a ticket; it is far more effective than saying “it got slower” with no evidence. If you want to turn this into a reusable forensics template, you can put these three numbers together with the unit price you pay in /app, calculate the premium rate, and then decide whether to upgrade, switch availability zones, or go through the refund process.

Disk Cache Cliff Verification

After NUMA binding, bandwidth is still pinned to the ceiling, so one final proof is missing: use a cache cliff to separate “slow memory” from “slow disk.” The approach is to first do a hot read, then a cold read of the same file, and see at which working-set size bandwidth drops.

# 1) 热路径:让 page cache 命中,测的是缓存速度
fio --name=hot --filename=/data/blob --rw=read --bs=1M \
    --size=4G --direct=0 --runtime=30 --time_based

# 2) 冷路径:清缓存后测真实内存→IO 通路
sync && echo 3 > /proc/sys/vm/drop_caches
fio --name=cold --filename=/data/blob --rw=read --bs=1M \
    --size=4G --direct=1 --runtime=30 --time_based

Interpretation points:

  • Cold-read bandwidth shows a smooth decline as the working set grows from 256M to 4G, which is normal memory-hierarchy and prefetch behavior;
  • If it falls off a cliff (for example, a direct halving around 2G), and the break point is far smaller than the instance's nominal L3/cache expectations, it usually means memory bandwidth is being throttled, rather than the disk giving out first;
  • Also check bi/bo and si/so in vmstat 1; if IO is not high but throughput drops sharply, the blame is basically on the memory side.

This is also where “real benchmarks vs. advertised specs” is most likely to go wrong: an 8GB instance claiming DDR4/DDR5 and capacity is not wrong, but once the working set of an AI inference KV cache or vector search is pushed out of cache, bandwidth drops first, and latency jitters immediately after. Put the cliff point, bandwidth before and after NUMA binding, and %steal together with the unit price you actually pay to calculate the premium rate, then decide whether to upgrade, switch availability zones, or go for a refund—far more effective than repeatedly rebooting. If you need a ready-made evidence template, you can use it directly in /app.

In one sentence: meeting capacity is just the ticket to entry; cloud server memory bandwidth detection is about where the slowdown inflection point is, and what price you're still paying after that inflection point.

Benchmark Comparison and Premium Decision

Earlier, we already obtained three sets of hard data: STREAM's Copy/Triad, the bandwidth difference before and after NUMA binding, and the %steal curve. Now only the last step remains—align them with the bill. The method is simple: write the 'nominal vs measured' values for the same instance in two columns:

# 实测带宽(GB/s)
sysbench memory --memory-block-size=1M --memory-total-size=10G run | grep transferred
# 单价(元/GB 内存/月)= 月费 / 标称容量
# 性价比 = 实测带宽 / 单价

The interpretation thresholds can be a bit rough, but they must be quantified:

  • Measured bandwidth reaches more than 70% of the published value for same-generation DDR4/DDR5, and %steal is normally below 2%: normal sharing, keep using it;
  • Bandwidth is only 50%–70% of nominal expectations, and NUMA binding can recover 15%+: this is a scheduling issue, and changing availability zone or specifying vCPU affinity is often cheaper than upgrading;
  • Bandwidth is below 50%, %steal stays above 5% long-term, and the Cache cliff appears early: this is a typical VPS overselling signal, hard evidence for detecting cloud server overselling.

Calculate the price-performance ratio before discussing action. Suppose an 8GB instance costs 120 yuan per month, and measured bandwidth is only 60% of a same-priced AMD EPYC instance; then your premium rate is actually 40%—at this point, upgrading to 16GB often just amplifies 'more expensive per GB'; switching instance type or pursuing a refund is more cost-effective. Calculate the premium rate before upgrading, and preserve evidence before requesting a refund: submit the raw output of three benchmark runs, %steal from vmstat, and a screenshot of numactl --hardware together; in the support ticket, write only data and no adjectives, and your chances of getting the decision reversed are much higher. Run the same tests before renewal as well, to prevent a sneaky downgrade at renewal. The complete evidence-collection table and ticket template are at /guides/memory-bandwidth-checklist; you can copy them directly.

Remember the conclusion: capacity meeting spec does not mean bandwidth meeting spec. Use STREAM plus Steal Time and NUMA binding for cross-verification, then compare against the premium rate to decide whether to upgrade or request a refund.

FAQ

How to investigate when memory capacity meets spec but bandwidth drops?

First use STREAM to measure actual bandwidth, then compare with the advertised value; a difference of more than 30% is abnormal.

How to run STREAM accurately?

Pin to NUMA nodes, use taskset for core pinning, run multithreaded tests 3 times and take the median, avoid neighbor peak hours.

What level of Steal Time is considered abnormal?

Sustained >5% indicates overselling; check st with vmstat; if over 10%, gather evidence and request a refund.

How to determine if it's a cache cliff rather than slow memory?

Use dd to test the disk; if bandwidth drops sharply and iowait spikes, it indicates Cache throttling, not a memory issue.

How to decide using benchmark scores versus premium rate?

Calculate the price per GB of bandwidth; if the premium is >30% and speed drop is >20%, downgrade or request a refund.

Capacity meeting spec does not equal bandwidth meeting spec. Use STREAM plus Steal Time and NUMA pinning for cross-forensics, then compare the premium rate to decide whether to upgrade or request a refund.

Start free detection →