Prime Intellect, 8× H100 SXM5
This is a repeated single-node acceptance study of one Prime Intellect 8× H100 SXM5 spot rental. The qualification suite ran on 21 and 22 June 2026. Persistent counters, the gateway address, and nearly identical measurements indicate that both rentals returned the same physical machine, so the evidence contains two repeated measurements but only one independent node.
Both archived runs returned an overall FAIL under the suite version used at the time. Peak compute, sustained CUDA stress, NVLink, numerical integrity, and storage passed. The repeated concerns were the host-to-device path under load and cooldown-model instability. A legacy training probe also reported the same GPU 0 disparity twice, but later validation identified weaknesses in that probe's workload and timing protocol. The prior 17% throughput and 5× masking projections are therefore withdrawn pending remeasurement.
Executive assessment
| Finding | Repeated evidence | Assessment |
|---|---|---|
| GPU compute and sustained load | 719–789 TFLOPS per GPU across the two runs; 300 s burns peaked at 70°C with 0/240 throttle samples | Passed for the tested synthetic workloads |
| GPU interconnect | Minimum measured NVLink pair bandwidth 359.6 and 360.0 GB/s; NCCL all-reduce about 295 GB/s with zero wrong results | Functional and repeatable on this node |
| Host transfer path | GPU 3 measured 12.644 GB/s H2D and 13.329 GB/s D2H in both archived runs | Repeatable concern, but measured by a legacy single-shot probe without an independent cross-check |
| Cooldown behavior | Cooldown fit failed both runs; worst fitted time constant was 69.1 s in run 1 and 109.5 s in run 2, on different GPUs | Evidence of unstable or nonuniform decay; root cause not identified |
| Training-step disparity | Legacy probe reported GPU 0 at 30.81 and 31.44 ms while the fastest ranks were 5.41 and 5.37 ms | Quantitative claim withdrawn until the corrected probe is rerun |
Rental and test configuration
The console exposed the 8× H100 SXM5 SKU at $11.92 per hour on spot and $14.40 per hour on demand during the test window. These are checkout observations from June 2026, not current-price claims. The rental was delivered as a container on a shared host and was accessed over SSH after attaching a customer key. Launch latency was not instrumented tightly enough to report.
Each run requested eight NVIDIA H100 80GB HBM3 devices and executed sixtytwo test --full inside the container. The detected H100 SXM profile used a 480 TFLOPS diagnostic reference target, 3,350 GB/s memory-bandwidth reference, and 450 GB/s NVLink reference. These are suite baselines, not vendor-peak certifications.
The first run lasted 13 minutes 14 seconds and produced 44 PASS, 11 WARN, 6 SKIP, and 3 FAIL results across 64 checks. The second lasted 13 minutes 11 seconds and produced 43 PASS, 12 WARN, 6 SKIP, and 3 FAIL results. The repeated run was intended to test measurement stability, not to estimate pool prevalence.
Measurement results
| Measurement | Run 1, 21 June | Run 2, 22 June | Interpretation |
|---|---|---|---|
| Per-GPU matrix multiply | 718.8–788.9 TFLOPS | 741.7–760.3 TFLOPS | Passed the 480 TFLOPS diagnostic target |
| Isolated-kernel dispersion | 3.4% | 3.4% | No isolated compute straggler observed |
| Minimum GPU memory bandwidth | 2,427.5 GB/s | 2,425.2 GB/s | 72% of the suite's 3,350 GB/s reference; warning |
| Minimum NVLink pair bandwidth | 359.58 GB/s | 360.02 GB/s | Stable across 56 measured directed pairs |
| NCCL all-reduce bus bandwidth | 295.03 GB/s | 294.73 GB/s | Completed with zero wrong results |
| NCCL all-gather bus bandwidth | 248.62 GB/s | 247.73 GB/s | Below the suite's H100 healthy band; warning |
| Worst PCIe transfer under load | 12.644 / 13.329 GB/s H2D / D2H | 12.644 / 13.329 GB/s H2D / D2H | Same GPU and values in both runs; legacy probe |
| Sustained 300 s burn | 70°C peak, 0/240 throttle samples, 4% droop | 70°C peak, 0/240 throttle samples, 4% droop | Passed |
| Cooldown fit | median tau 23.9 s, worst tau 69.1 s, minimum r² 0.55 | median tau 22.5 s, worst tau 109.5 s, minimum r² 0.53 | Failed fit-quality and outlier rules |
| GPU ECC counters | 0 volatile corrected, 0 volatile uncorrectable, 2 aggregate historical uncorrectable | Same | No active event observed; historical counter requires context |
| Numerical reference | maximum absolute error 1.19×10⁻⁷, bit-identical repeat | Same | Passed the tested reference workload |
Peak matrix multiply and the isolated kernel show that the arithmetic units were not uniformly slow. The 300-second burns also completed without thermal throttling, and NVLink was stable. Those results bound the finding: this was not a general loss of GPU compute capability during the tests.
The transfer result was localized. GPU 3 was the minimum in both runs at 12.644 GB/s host-to-device and 13.329 GB/s device-to-host. GPUs 4 through 7 showed approximately 49 GB/s H2D but about 19 GB/s D2H, while other GPUs were higher in both directions. The exact pattern repeated a day later. That repeatability argues against an isolated transient, but the archived probe used the pre-July single-shot timing method and no NVIDIA nvbandwidth cross-check was captured. The defensible conclusion is that the legacy method found a repeatable asymmetry that requires current-protocol confirmation; it does not identify whether the cause was physical PCIe, host NUMA placement, virtualization, or probe interaction.
Cooldown behavior was also abnormal under the suite's exponential-decay model, but its localization did not reproduce. Run 1's maximum fitted time constant was 69.1 seconds on GPU 2; run 2's was 109.5 seconds on GPU 1. Several GPUs in both runs also had poor exponential fit. This is consistent with unstable or nonuniform cooling behavior. It is not sufficient to diagnose a thermal interface, airflow obstruction, or ambient-control fault without host telemetry and a controlled remeasurement.
Correction to the training-step claim
The archived short training probe reported GPU 0 at 30.81 and 31.44 milliseconds per step, compared with 5.41 and 5.37 milliseconds for the fastest GPU. The same GPU-level ordering across both runs is noteworthy. However, the probe reused a small fixed batch, reached zero final loss, performed host-synchronizing scalar reads inside the timed region, and predated the fixed-warmup, five-trial, median-of-trials protocol later adopted by the suite.
Those defects prevent the old step times from supporting precise throughput or cost projections. In particular, synchronized throughput cannot be asserted to equal 17% of a clean node, and masking GPU 0 cannot be asserted to yield a 5× improvement, without rerunning the corrected probe on this hardware. The original measurements remain in the evidence record for auditability; their economic interpretation is withdrawn.
Cost exposure and acceptance testing
At the observed $11.92 hourly spot price, one 13.2-minute suite run corresponds to approximately $2.62 of GPU rental before setup time, billing granularity, storage, or network charges. Seven uninterrupted days at the same rate would be $2,002.56. This arithmetic illustrates why a bounded acceptance test is inexpensive relative to a long reservation; it does not estimate savings, because the current protocol has not been rerun on this node and the tested synthetic checks may not predict every workload.
For a new rental, the acceptance decision should use the current timing protocol and a job-relevant workload. At minimum, remeasure bidirectional host transfer with both the Sixtytwo probe and an independent tool, run the corrected multi-trial training probe, confirm volatile ECC and Xid counters, repeat sustained load and cooldown, and execute the actual collective sizes and precision used by the intended job.
Scope and limitations
- Both runs almost certainly sampled the same physical machine. The effective independent node count is one.
- The study estimates neither Prime Intellect fleet prevalence nor a service-level failure probability.
- Container access limits attribution: host NUMA policy, BIOS, BMC, fan control, and physical PCIe topology were not independently observable.
- Six checks were skipped, largely because host-level facilities do not apply or were unavailable inside the container. The DCGM diagnostic did not complete.
- The PCIe and training measurements predate the suite's July timing-standard revision. The PCIe pattern reproduced but lacks an independent archived cross-check; the training-derived economic claims are withdrawn.
- Cooldown-model failure is an observation, not a root-cause diagnosis, and the worst GPU changed between runs.
- The 300-second burn and short numerical probes cannot establish a rare-fault rate for a multi-day job.
- Prices describe the June 2026 checkout and exclude ancillary charges and spot-preemption risk.
Reproducibility and evidence
The method is described in How we benchmark providers, and the harness is distributed as sixtytwo-cli. The evidence bundle contains the complete JSON reports for both runs, including per-GPU rows, timestamps, check statuses, and raw summaries. Reproduction should use the current harness because the legacy timing results are not numerically interchangeable with the revised protocol.