Understanding GPU Cloud Costs
GPU Cloud Pricing Reference (On-Demand, 2024)
| GPU |
VRAM |
AWS (p-series) |
GCP (a-series) |
Lambda |
Notes |
| NVIDIA T4 | 16 GB | ~$0.53/hr | ~$0.35/hr | ~$0.50/hr | Inference, light training |
| NVIDIA A10G | 24 GB | ~$1.01/hr | ~$0.90/hr | ~$0.75/hr | Mid-range training |
| NVIDIA A100 (40GB) | 40 GB | ~$3.20/hr | ~$2.93/hr | ~$1.10/hr | Production AI training |
| NVIDIA A100 (80GB) | 80 GB | ~$4.10/hr | ~$3.67/hr | ~$1.50/hr | Large model training |
| NVIDIA H100 (SXM) | 80 GB | ~$6.00/hr+ | ~$5.18/hr | ~$2.49/hr | Flagship; LLM training |
| NVIDIA L4 | 24 GB | ~$0.71/hr | ~$0.70/hr | — | Efficient inference |
GPU Cloud Provider Comparison
| Provider |
Strengths |
Weaknesses |
Best For |
| AWS (EC2, SageMaker) | Largest ecosystem, many integrations | Expensive, complex billing | Enterprise, existing AWS users |
| Google Cloud (GCP) | TPUs available, good Vertex AI | Less competitive GPU pricing | ML teams using Google stack |
| Azure | Deep Microsoft/OpenAI integration | Complex pricing | Enterprise, Azure/M365 shops |
| Lambda Labs | Lowest GPU pricing | Fewer services, less reliability | Researchers, cost-sensitive |
| Vast.ai | Cheapest (spot pricing) | Unreliable, setup overhead | Dev/test, budget projects |
| CoreWeave | H100 focus, competitive | Less ecosystem | High-end AI training |
| RunPod | Easy UI, competitive | Smaller platform | Individual researchers, ML |
Typical GPU Training Job Cost Examples
| Workload |
GPU |
Hours |
Approx. Cost |
Notes |
| Fine-tune BERT (small) | T4 | 2–4 hrs | $1–3 | Classification task |
| Fine-tune LLaMA-7B | A10G | 8–16 hrs | $8–16 | With PEFT/LoRA |
| Train ResNet-50 (ImageNet) | A100 40GB | 12–24 hrs | $40–80 | Standard benchmark |
| Fine-tune GPT-2 (large) | A100 80GB | 4–8 hrs | $20–40 | With full fine-tuning |
| Pre-train small LLM (1B) | H100 ×8 | 100–500 hrs | $1,000–5,000 | Distributed training |
| Pre-train large LLM (70B) | H100 ×64 | 10,000+ hrs | $500,000+ | Research/enterprise |
GPU cloud computing has become essential for machine learning, AI development, rendering, and scientific computing. Choosing the right cloud provider and GPU type can significantly impact your project's budget. This comprehensive guide helps you understand GPU cloud pricing across major providers and make informed decisions for your workloads.
Cloud GPU prices vary widely between providers, GPU types, and pricing models. A single H100 GPU can cost over $30 per hour on-demand, while a T4 might cost less than $0.50 per hour. Understanding these differences is crucial for optimizing your cloud spending.
Major Cloud GPU Providers
Amazon Web Services (AWS)
AWS offers GPU instances through EC2, with options ranging from the budget-friendly T4 to the powerful H100. Key instance families include:
- P5 instances: Latest H100 GPUs for demanding AI workloads
- P4d/P4de instances: A100 GPUs for training and inference
- G5 instances: A10G GPUs for graphics and ML inference
- G4dn instances: T4 GPUs for cost-effective inference
Google Cloud Platform (GCP)
GCP provides flexible GPU attachment to VMs and offers competitive preemptible pricing:
- A3 VMs: H100 GPUs for cutting-edge AI
- A2 VMs: A100 GPUs with various configurations
- G2 VMs: L4 GPUs for inference workloads
- N1 with GPUs: Flexible V100, T4 attachment
Microsoft Azure
Azure offers GPU VMs optimized for different workloads:
- ND H100 v5: H100 GPUs for large-scale training
- ND A100 v4: A100 GPUs for demanding workloads
- NC A100 v4: A100 for cost-effective training
- NCas T4 v3: T4 GPUs for inference
Lambda Labs
Lambda Labs offers competitive pricing focused on ML workloads with simplified pricing and no hidden fees. They're known for excellent GPU availability and straightforward billing.
RunPod
RunPod provides flexible, pay-as-you-go GPU cloud with some of the most competitive spot pricing in the market. Ideal for development, testing, and burst workloads.
Pricing Models Explained
On-Demand Pricing
Pay for compute capacity by the hour with no long-term commitments. Best for variable workloads, development, and testing. Highest flexibility but also highest cost.
Spot/Preemptible Pricing
Access unused cloud capacity at steep discounts (50-90% off on-demand). Instances can be interrupted with short notice. Best for fault-tolerant workloads with checkpointing.
Reserved Instances
Commit to 1-3 year terms for significant discounts (30-70% off on-demand). Best for stable, predictable workloads. Requires upfront planning and commitment.
GPU Comparison
| GPU |
Memory |
FP16 TFLOPS |
Best For |
| H100 SXM |
80GB HBM3 |
1,979 |
Large LLM training, fastest inference |
| A100 80GB |
80GB HBM2e |
312 |
LLM training, large model inference |
| A100 40GB |
40GB HBM2e |
312 |
General training, medium models |
| V100 |
32GB HBM2 |
125 |
Training, good price/performance |
| A10G |
24GB GDDR6 |
125 |
Inference, graphics, rendering |
| L4 |
24GB GDDR6 |
121 |
Inference, video processing |
| T4 |
16GB GDDR6 |
65 |
Budget inference, development |
Tips for Reducing GPU Cloud Costs
1. Right-Size Your GPU Selection
Don't pay for more GPU power than you need. Profile your workload to determine minimum GPU requirements. A T4 may be sufficient for inference that doesn't require an A100.
2. Leverage Spot Instances
For training workloads with checkpointing, spot instances can reduce costs by 60-90%. Implement robust checkpoint saving and loading to handle interruptions.
3. Consider Alternative Providers
Lambda Labs and RunPod often offer lower prices than major cloud providers. They may have better GPU availability for high-demand models like H100.
4. Use Reserved Capacity for Steady Workloads
If you have predictable GPU usage, reserved instances can provide significant savings over on-demand pricing.
5. Optimize Training Efficiency
Mixed-precision training, gradient checkpointing, and efficient data loading can reduce training time and therefore costs.
Conclusion
GPU cloud costs vary significantly across providers and configurations. Use our GPU Cloud Cost Comparison Calculator to find the most cost-effective option for your specific needs. Consider factors beyond just price, including GPU availability, support quality, and ecosystem integration.
Frequently Asked Questions
How accurate are the results?
The GPU Cloud Cost Comparison applies a standard formula to your inputs — accuracy depends on how precisely you measure those inputs. For planning and estimation, results are reliable. For high-stakes or professional decisions, cross-check the output with a domain expert or primary source.
Can I use this on mobile?
Yes — the calculator is designed to work on any device. For complex multi-input calculations on small screens, landscape orientation gives more room to see all fields and results simultaneously.
How much does it cost to run a GPU in the cloud?
GPU cloud pricing varies widely depending on the GPU type, provider, and pricing model (on-demand vs. spot vs. reserved). Representative on-demand prices (2024): Entry-level (T4, 16GB VRAM): $0.35–$0.53/hour (AWS, GCP, Lambda). Mid-range (A10G, 24GB): $0.75–$1.01/hour. Production AI (A100 40GB): $1.10–$3.20/hour. Flagship (H100, 80GB): $2.49–$6.00+/hour. Cheapest options: Vast.ai (community GPU marketplace): $0.10–$0.50/hour for T4/A10G (variable availability). Lambda Labs: Often 50–60% cheaper than AWS/GCP for the same GPU. RunPod: Competitive spot pricing, easy to use. Cost calculation: Total cost = hourly rate × GPU hours. 10 A100 hours on Lambda: 10 × $1.10 = $11. Same on AWS: 10 × $3.20 = $32. Pro tip — spot/preemptible instances: AWS Spot: 60–90% discount vs. on-demand, but can be interrupted. GCP Preemptible: ~80% discount. Best for: Training jobs you can checkpoint and resume. Not suitable for: Real-time inference, jobs that can't be interrupted. Reserved instances: 1-year or 3-year commitment saves 30–60% vs. on-demand. Best for: Production inference endpoints with predictable usage.
How do I choose the right GPU for machine learning?
GPU selection depends on your model size, task type, and budget. Key factors: VRAM (video RAM): Most important factor for training large models. Model won't fit in VRAM → training fails. Rule of thumb (mixed precision training): Model parameters × 2 bytes/param (FP16) + optimizer states (~8 bytes/param for Adam). LLM fine-tuning: LLaMA-7B: ~14GB VRAM minimum (FP16). LLaMA-13B: ~26GB VRAM minimum. LLaMA-70B: Requires multi-GPU or quantization. GPU recommendations by task: Inference only (small models): T4 (16GB, cheapest). Computer vision training: A10G or older V100. LLM fine-tuning (7–13B): A100 40GB. LLM fine-tuning (30–70B): A100 80GB or multi-GPU A100s. LLM pre-training: H100 (best performance/watt). VRAM-saving techniques: Quantization: Load model in 4-bit or 8-bit (reduces VRAM 4–8×). LoRA/QLoRA: Fine-tune only a small fraction of weights. Gradient checkpointing: Trade compute for memory. Mixed precision (FP16/BF16): Half the memory of FP32. Compute vs. memory bandwidth: For inference, memory bandwidth matters more than FLOPS (most inference is memory-bandwidth bound). H100 vs. A100 for inference: H100 has 2× memory bandwidth → roughly 2× inference throughput. Multi-GPU scaling: Linear scaling only works when the bottleneck is compute, not memory/communication. Typical efficiency: 70–85% for data parallelism, 60–75% for model parallelism.
What is the difference between on-demand, spot, and reserved GPU instances?
Cloud GPU instances come in three main pricing models with very different cost/reliability tradeoffs. On-demand: Pricing: Standard list price, highest cost. Availability: Immediate (when available). Interruption: Never interrupted by the provider. Best for: Short jobs, inference, testing, anything where interruption would be costly. Cost: 100% of list price. Spot / Preemptible: Pricing: 60–90% discount vs. on-demand. Availability: Uses spare capacity — can run out. Interruption: Provider can reclaim instance with 2-minute warning (AWS) or 30-second (GCP). Best for: Training jobs with checkpointing, batch processing, experiments. Required: Implement checkpointing so training can resume after interruption. Cost saving: AWS Spot A100: $3.20 → ~$0.48–0.96/hr. GCP Preemptible A100: $2.93 → ~$0.59/hr. Reserved instances: Pricing: 30–60% discount vs. on-demand. Commitment: 1-year or 3-year purchase, paid upfront or monthly. Availability: Guaranteed capacity (capacity reservation). Best for: Production inference endpoints, predictable usage. Risk: Paying for capacity even when unused. Hybrid strategy: Use reserved instances for baseline load, spot for burst training, on-demand for critical jobs. Many teams use on-demand for dev/test, spot for experiments, reserved for production serving.
Is buying a GPU better than renting cloud GPUs?
This is the classic "buy vs. rent" tradeoff. The answer depends heavily on utilization rate and time horizon. When buying wins: Continuous high utilization (>50–60% of hours): A used A100 80GB: ~$8,000–10,000. Cloud cost at $1.50/hr: $1.50 × 24 × 365 = $13,140/year. ROI < 1 year at full utilization. You prefer total control over hardware and data. Long-term research projects (2–5+ years). When cloud wins: Sporadic or unpredictable workload: Cloud scales to zero; owned hardware costs electricity even idle. Need the latest GPU hardware (H100, B200): Owned hardware depreciates; cloud offers access. No upfront capital available. Short projects: 6–12 months rarely justify hardware purchase. Hardware comparison (2024 approximate): RTX 4090 (24GB): ~$1,600, ~$0.25–0.35/hr equivalent. A100 80GB (new): ~$25,000. A100 80GB (used): ~$8,000–12,000. H100 SXM5 (new): ~$30,000–40,000. Hidden costs of ownership: Electricity: A100 draws ~250–400W = $0.02–0.04/hr at $0.10/kWh. Cooling and maintenance. Depreciation (GPUs depreciate ~20–30% per year). Support and redundancy. Colocation (colocating your GPU in a data center): Middle ground — own the hardware, pay $0.10–0.30/hr for rack space + power. Works well for teams with 3+ A100s and steady utilization.