
Proxmox ARM64 Was Crawling at 6 GB/s on the ASUS GX10. The Bottleneck Was Hiding in Plain Sight.
On one ASUS GX10 running Proxmox VE 9.2 ARM64, kernel NVMe/TCP read throughput increased from 6.3 GB/s to 28.1 GB/s after IOMMU, PCIe, NIC and queue tuning. These are host-level read benchmarks on a specific test setup, not a claim about general VM or AI-inference speed.
The published Blockbridge test explains the progression: iommu.strict=0 took the kernel initiator to 24.4 GB/s, pci=pcie_bus_perf took it to 26.5 GB/s, and additional network tuning produced the 28.1 GB/s result. The team also measured 28.9 GB/s using SPDK.
Why was ASUS GX10 NVMe/TCP stuck near 6.3 GB/s?
The published ASUS GX10 result attributes the low default read rate primarily to strict IOMMU invalidation in the tested host configuration.
The GX10 has the kind of specification sheet that makes a small box feel improbable: twenty Arm CPU cores, 128 GB of memory, a GPU, and a dual-port 200 Gb/s NVIDIA ConnectX-7 network adapter. NVMe over TCP provides a way to access remote NVMe storage over standard Ethernet networks, making it a useful test of the whole receiving path rather than a neat demonstration of one chip's headline capability. Bytes have to arrive at the NIC, move over PCIe, pass through IOMMU handling, and reach host memory. A bottleneck anywhere along that journey shows up in the final throughput.
For background on how KVM and virtual devices fit together, see the QEMU versus KVM explanation.
The engineer disclosed working at Blockbridge, the company supplying the NVMe/TCP target used in testing. The technical write-up gives the detailed configuration and test commands. The researchers changed host settings in stages, compared the kernel initiator with SPDK, and checked their results against the calculated capacity of the PCIe links.
At the starting point, the contrast was stark. The kernel initiator managed 6.3 GB/s, while SPDK already achieved 21.3 GB/s on the same default host configuration. Profiling supplied the clue: arm_smmu_cmdq_issue_cmdlist reportedly consumed 57% of cycles on the larger CPU cores. The machine was spending an extraordinary amount of effort on IOMMU command handling. The bottleneck was IOMMU overhead.
Which kernel parameters improved GX10 throughput?
iommu.strict=0 and pci=pcie_bus_perf improved different bottlenecks; the 28.1 GB/s result also required further NIC tuning.
The first change was iommu.strict=0. On the tested ARM64 kernel, strict IOMMU invalidation caused every DMA unmap to wait for the relevant hardware translation-cache invalidation. On a system whose SMMU has a single command queue, those waits became expensive. Switching to lazy invalidation allowed operations to be batched, taking the kernel initiator from 6.3 to 24.4 GB/s. SPDK moved from 21.3 to 24.7 GB/s. That is a striking difference in how the two paths responded.
Next came pci=pcie_bus_perf. The GX10 firmware had initialized the NIC's PCIe root ports with a 128-byte maximum payload size, although the hardware supported 512 bytes. Linux preserved the smaller setting, leaving useful link capacity on the table. The performance-oriented PCIe parameter allowed the larger payload, moving the kernel initiator to 26.5 GB/s and SPDK to 28.8 GB/s. The team later enabled large receive offload, reduced the number of outstanding reads, and pinned NIC interrupts. Those additional changes took the averages to 28.1 and 28.9 GB/s, respectively. In other words, the memorable two-flag story is true as a description of the breakthrough, but the final kernel number also reflects subsequent tuning. Treating every gain as the work of two magic lines would obscure what the testing actually found.
The reported physical limit explains why the last gigabytes were difficult to extract. Behind the ConnectX-7 sit two separate PCIe Gen5 x4 links, not a single x16 connection. The team calculated a combined NVMe payload ceiling of approximately 29.9 GB/s. Reaching 28.9 GB/s is therefore impressive relative to this particular topology, but it is not evidence of unlimited bandwidth or a universal ARM64 performance record. One favorable SPDK run reached roughly 29.3 GB/s of payload, and the engineer cautioned that such runs were not fully repeatable because TCP connections could land unevenly across NIC receive queues. Even near the hardware ceiling, packet placement and scheduling still matter.
How did ConnectX-7 topology limit the GX10?
The ConnectX-7 exposes paths across two PCIe Gen5 x4 links, so traffic has to use the right physical functions to exploit both.
A dual-port NIC usually sounds like two obvious places to send traffic. On this GX10, that assumption was misleading. Each physical network port exposed a physical function on each of the two PCIe halves, leaving four PFs visible in Linux. A given PF could use only its own x4 link; two PFs attached to the same half still competed for the same roughly 14.8 GB/s ceiling. Meanwhile, traffic confined to one 200 Gb/s physical port faced that port's approximate 24.7 GB/s limit. To reach the reported 28.9 GB/s, the experiment needed all four PFs so both physical ports and both PCIe halves carried traffic. The relevant limit was determined by the route each packet actually took, not by the impressive bandwidth printed on the adapter's specification sheet.
Storage tuning only helps when the storage design fits the workload, a distinction covered in the Proxmox storage comparison.
Then there was an almost comically quiet network problem. Two PFs sharing a physical port could each answer ARP requests for the other's address under the default Linux behavior. Whichever response reached the remote target first could determine where traffic flowed. If the wrong PF won, one path sat idle. No obvious error message appeared, and the headline link speed still looked fine. The only reliable clue was the per-PF rx_bytes counter. Using arp_ignore=1 and arp_announce=2 addressed that behavior in the test setup.
A question about the screenshot's monitoring interface revealed another grounded detail. The colorful dashboard was a small Python program built with Rich, drawing data from /proc, /sys, and ethtool. It displayed network, PCIe, and NVMe throughput together, along with per-core activity, so bottlenecks became easier to distinguish. The author cautioned that it was tailored to this machine's unusual layout and probably wasn't broadly reusable. Custom instrumentation can be exactly what a strange hardware topology requires.
Should you use lazy IOMMU invalidation on Proxmox?
iommu.strict=0 can improve throughput, but the Linux kernel explicitly documents reduced device isolation as a tradeoff.
The clearest pushback focused on iommu.strict=0. One participant saw the result as the familiar bargain of reduced security for increased speed. The engineer did not dismiss the concern. Instead, they explained that lazy invalidation does not disable IOMMU translation. Mappings remain in use; the change is when the system flushes stale translation entries after a DMA unmap. Batching improves performance, but it can briefly leave a device able to access memory that the kernel has already unmapped. The Linux kernel's parameter documentation explicitly describes that as a trade-off between throughput and device isolation. For machines with untrusted devices or PCIe hardware passed into untrusted guests, the original author recommended retaining strict mode and accepting the performance penalty.
Another participant wondered whether Intel and AMD systems might benefit from the same settings. The response was a useful check on enthusiasm: the author said x86 commonly already uses lazy IOMMU behavior, so adding iommu.strict=0 would often change nothing. Likewise, the unusual 128-byte PCIe payload setting was tied to the GB10 firmware behavior seen here; other systems may initialize their links differently. The team also tested IOMMU passthrough and reported no improvement over lazy mode on the GX10, making a more aggressive reduction in isolation hard to justify for this result. A benchmark's best configuration is not automatically a deployment's best security policy.
That leaves a surprisingly grounded outcome. The researchers began with boot faults, unexplained resets, a power limitation, and a networking number that made capable hardware look mediocre. They ended with a Proxmox ARM64 host operating close to the measured topology's calculated read ceiling, using an ordinary kernel and standard drivers. The result makes ARM64 virtualization look promising on the GX10, but it does not prove production readiness for every guest, GPU workload, or security model. The lesson is more useful than that. When powerful infrastructure disappoints, the fastest route to an answer may be to stop looking at the advertised bandwidth and trace the actual path of the bytes.
The jump from 6.3 to 28.1 GB/s wasn't a mystery solved by throwing more hardware at the box. It came from finding where the host was waiting, where PCIe had been unnecessarily constrained, and where packets were choosing the wrong route. Anyone tempted to repeat it should start with the same discipline: measure first, change one thing at a time, and know what each performance setting gives up.
I would reproduce the result only on a controlled test host and measure each setting individually. On a multi-tenant or isolation-sensitive server, I would keep strict invalidation unless a documented risk assessment justifies the performance gain.
Frequently Asked Questions
How fast was NVMe/TCP on the ASUS GX10 under Proxmox ARM64?
Blockbridge reported 6.3 GB/s with the default kernel NVMe/TCP initiator and 28.1 GB/s after tuning. The measurement was made on the host, not from inside a virtual machine.
What does iommu.strict=0 do?
It enables deferred IOMMU invalidation where supported, reducing overhead from DMA unmapping. Linux documents higher throughput at the cost of reduced device isolation.
Do two GX10 kernel parameters alone deliver 28.1 GB/s?
No. The published test reached 26.5 GB/s after the two parameter changes and used additional NIC and queue adjustments to achieve the final 28.1 GB/s result.