Two weeks on one DGX Spark — the power cut out, the clock got stuck, and we gave up on the 70B
Notes · Operations
Two weeks on one DGX Spark — the power cut out, the clock got stuck, and we gave up on the 70B
Two power cuts under a 70B, the clock stuck at 507 MHz and the cold drain that clears it, the decode probe that tells, and the rules we adopted.
2026-10-05
Every benchmark we publish ran on one DGX Spark (GB10, 128 GB unified memory, DGX OS 7.2.3, driver 580.173.02, vLLM 0.27.1 in a container). Eight models of about 31B ran for two weeks without trouble; the moment we loaded a 70B, the power cut out twice. Here is only what may help someone on the same machine.
What happened
- Two power cuts, both under the 70B. On the evening of 3 October, about 45 minutes into a public-benchmark run with llama-3.3-70b (FP8, vLLM), the machine alone went dark. The next afternoon, six minutes after loading the same model again, it went dark again. Both times the system journal ends mid-line with no shutdown sequence and no kernel panic. The second time the box would not power on until it had cooled (an ice pack, ten minutes). The GPU die was at 56–77 °C in the power log we had kept on disk, so the cut sits elsewhere: the power stage (board VRM) or the USB-C PD adapter, tripping on heat. Nothing like it in two weeks on the 31B models.
- After a cut, the clock gets stuck. Power back on, model loaded, and the GPU sits at 507–513 MHz, 8–10 W, P0, with no throttle reason shown, at 95 % utilisation. Decode falls from the usual 6.4 tok/s to 1.7–4.3, and every turn takes twice as long. We first met this on 19 September (the documented USB-PD controller wedge); it came back after the cuts. Two warm reboots do nothing.
The fix — a full drain
Power off, unplug the adapter at the wall and at the unit, unplug every USB-C peripheral, hold the power button for 30 seconds, wait at least 60 seconds, then power on. Our first attempt on 4 October skipped the 60-second wait and the box stayed stuck (4.3 tok/s, 507 MHz); the second, done properly with the box off for about 15 minutes, cleared it (6.4 tok/s, 2528 MHz, 35 W). In September one drain was enough.
The check — do not trust idle
A stuck box looks normal at idle: clocks and temperature read fine. A light load (300 embeddings) only reached 858 MHz, too little to tell. Only a real decode load decides it. We run npm run preflight, a 200-token decode probe, and pass it at 5 tok/s or more and an SM clock of 1400 MHz or more (31B FP8). Forty minutes before any demo.
| State | Decode (31B FP8, 200 tokens) | SM clock under load | Power |
|---|---|---|---|
| Healthy | 6.3 to 6.4 tok/s | 2528 MHz | 35 to 37 W |
| Stuck (after the first drain, 4 Oct) | 4.3 tok/s | 507 MHz | — |
| Stuck (right after the cut, 70B, 4 Oct) | 1.7 tok/s | 507 MHz | 10 W |
The rules we adopted
- No 70B on this unit. The row resumes on a second machine, which will also tell whether the fault is this unit’s power stage or the adapter.
- A power log on disk for every heavy run (
nvidia-smi --query-gpu=timestamp,power.draw,temperature.gpu --format=csv -l 30). Never under/tmp: it is tmpfs and gone after a reboot, which is how we lost that night’s logs. - When Ollama (embeddings) shares the GPU, vLLM runs with
--gpu-memory-utilization0.55–0.6. At 0.7 the 70B produced 23 memory warnings while loading (not the cause of the cut). - NVIDIA’s field diagnostics (
dgx-spark-fieldiag) needs Secure Boot off andinit 3, takes 30 minutes, and an abort needs a power cycle. We run it on an empty day.
What we do not know
Whether the thermal cut-off is this unit’s power stage or the adapter and cable. The second machine and the field diagnostics will say, and we will add it here. How much speculative decoding (Google’s Gemma 4 drafter) speeds up decode is being measured and goes up as soon as it is.
Related notes: Choosing a local writer · Tool-call reliability
