产品官方一手精选

AWS SageMaker AI评测小型LLM推理性能

Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6

精选理由

AWS实测了两种30B模型在四种GPU上的表现,帮你算清楚哪种实例性价比最高。

AWS在SageMaker AI上评测了Qwen3-Coder-30B和NVIDIA Nemotron-3-Nano-30B两个30B专家混合模型。测试对比了G5、G6、G6e和G7 GPU实例的吞吐量、延迟和每token成本。G7搭载的NVIDIA Blackwell GPU在实时LLM推理中展现出明显的价格性能优势。

图片来源 · AWS Machine Learning Blog
原文 · AWS Machine Learning Blog

Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6

Benchmark two 30B Mixture-of-Experts models, Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B, across G5, G6, G6e, and G7 GPU instances on Amazon SageMaker AI. Compare throughput, latency, and cost-per-token, and see how G7's NVIDIA Blackwell GPUs deliver measurable price-performance gains for real-time LLM inference.