NVIDIA DGX Spark instances with 128GB unified memory and Blackwell GB10 architecture. Bare-metal SSH access, Docker, NVMe storage. Built for AI researchers and engineers who need GPU power without enterprise overhead.
Starting at $0.75/hour — View pricing
Run multi-agent systems or 200B parameter models without sharding. Load a 120B reasoning model, a 6.7B coder, and a 4B embedding model simultaneously — all in memory without swapping.
Combine NVFP4 with Speculative Decoding for 3x faster inference while maintaining accuracy. Both draft and target models fit in 128GB unified memory.
Develop on the same architecture you'll deploy to production. Code written for Blackwell uses SM 10.0 features — 5th-gen Tensor Cores and improved sparsity support — that don't exist on Hopper's SM 9.0.
Second-generation attention acceleration delivers 2x attention speedup. Measure actual inference throughput on your models before committing to hardware.
| GPU | VRAM | Hourly |
|---|---|---|
| NVIDIA DGX Spark | 128GB | $0.75 |
| NVIDIA H100 | 80GB | $3.85 |
| NVIDIA H200 | 141GB | $4.50 |
| NVIDIA B200 | 192GB | $7.15 |
Pay-per-hour, no commitment. Same features on every configuration — pick the one that fits your workload.
Single node, 128 GB unified memory.
Dual node interconnect, 256 GB unified memory.
Quad node interconnect, 512 GB unified memory.
DGX Spark Cloud gives you remote SSH access to a dedicated NVIDIA DGX Spark workstation — featuring the GB10 Blackwell GPU with 128GB unified memory. You get bare-metal performance with Docker support, NVMe storage, and full CUDA 13 compatibility, without buying the hardware.
Pricing is pay-per-hour with no commitment. A single DGX Spark is $0.75/hour, a 2× DGX Spark (dual node, 256 GB unified memory) is $1.65/hour, and a 4× DGX Spark (quad node, 512 GB unified memory) is $4.50/hour. All tiers include full root access, SSH, Docker, NVMe storage, and founder support during beta — only the GPU configuration differs.
DGX Spark uses the Blackwell GB10 GPU (SM 10.0) with 128GB unified memory at 273 GB/s bandwidth. The H100 uses Hopper architecture (SM 9.0) with 80GB HBM3. DGX Spark offers 60% more memory, native NVFP4 quantization, and 2nd-generation Transformer Engine — at a fraction of the hourly cost ($0.75/hr vs $3.85+/hr for H100).
AI researchers, ML engineers, and teams who need to train models, fine-tune LLMs, build multi-agent systems, or benchmark on Blackwell architecture — without committing to enterprise hardware purchases or long-term cloud contracts.
You get direct SSH access to your dedicated DGX Spark instance. Docker is pre-installed, CUDA 13 and the full NVIDIA AI stack are ready to use. Connect from any terminal.
Yes. With 128GB unified memory, DGX Spark can run models up to 200B parameters without sharding. You can load multiple models simultaneously for multi-agent workflows.
No. NVIDIA DGX Cloud is an enterprise platform with multi-node clusters for large-scale training. DGX Spark Cloud by Enverge gives you a single dedicated DGX Spark workstation — ideal for individual researchers and small teams at a fraction of the cost.
Each instance comes with Ubuntu, CUDA 13, cuDNN, NVIDIA driver 535+, Docker, Python 3, PyTorch, and the full NVIDIA AI Enterprise stack. You have root access (Unlimited plan) to install anything else.