NVIDIA · Ada Lovelace Architecture

Rent NVIDIA L40 in the Cloud

VRAM 48 GB GDDR6
Bandwidth 864 GB/s
FP16 181.0 TFLOPS
FP32 90.5 TFLOPS
TDP 300W
Architecture Ada Lovelace

No pricing data available yet for this GPU model. Check back soon.

NVIDIA L40 Technical Specifications

Manufacturer NVIDIA
Architecture Ada Lovelace
VRAM 48 GB GDDR6
Memory Bandwidth 864 GB/s
FP16 (Tensor) 181.0 TFLOPS
FP32 90.5 TFLOPS
TDP 300W
Release Year 2023
Segment Data center
Memory Type GDDR6

Best For

Inference video processing rendering

Frequently Asked Questions

Is NVIDIA L40 memory bandwidth enough for LLM production inference?

Short version of the NVIDIA L40 spec sheet: 48 GB GDDR6, 864 GB/s, 181 FP16 TFLOPS, 90.5 FP32 TFLOPS, Ada Lovelace (2023), 300W.

Long version: the card is tuned for mixed-precision matrix multiplication on large tensors, which is exactly what transformer training and production inference demand. Bandwidth is generous enough to avoid stalling on attention operations, and VRAM capacity covers modern model sizes without requiring offloading to CPU memory.

Full specs, benchmarks, and comparisons are on the NVIDIA L40 page.

NVIDIA L40 memory-bound vs compute-bound workloads

NVIDIA L40 performance headline: 181 FP16 TFLOPS, 90.5 FP32 TFLOPS, 864 GB/s bandwidth, 48 GB VRAM.

Converted into practical benchmarks: model training a 7B-parameter LLM in FP16 with reasonable batch sizes typically saturates compute before bandwidth; real-time serving on the same model is usually bandwidth-bound and tracks the 864 GB/s figure. Diffusion image generation benchmarks sit between the two — compute-heavy steps utilise tensor cores well, while attention blocks still touch bandwidth.

Check the NVIDIA L40 page for complete specifications and related GPU matchups.

NVIDIA L40 alternatives — what else should I consider?

NVIDIA L40 is best for workloads where its 48 GB VRAM and Ada Lovelace tensor cores are well-matched: Inference, video processing, rendering.

If your workload needs significantly more memory (e.g., training frontier-scale models from scratch), NVIDIA L40 is undersized and you'd want an H100/H200/B200 class card. If your workload needs less (e.g., small-scale serving on 7B-parameter models), cheaper cards like L4 or RTX 4090 may be more cost-efficient. For the middle band, NVIDIA L40 is usually the sensible pick.

Full specs, benchmarks, and comparisons are on the NVIDIA L40 page.

Compare with Other GPUs

See how NVIDIA L40 stacks up against other popular cloud GPUs in specs, pricing, and availability.