Modal Pricing (2026): Per-Second GPU Costs and Free Credits
Modal charges per second for the GPU, CPU and memory your code uses, and nothing while it is scaled to zero. An H100 costs $0.001097 a second (about $3.95 an hour), an A100 80 GB about $2.50 and an L4 about $0.80; CPU costs $0.0000131 per core per second and memory $0.00000222 per GiB per second. The Starter plan is free and includes $30 of compute every month; Team costs $250 a month plus usage, with $100 of monthly credit and higher limits; Enterprise is custom.
Prices are Modal's list prices in US dollars, checked on 25 September 2026. Pinning a region costs 1.15 to 1.75 times base prices.
Prices checked 25 September 2026 on Modal's official pricing page. Prices are in US dollars unless stated and can change at any time.
Modal plans and prices
Starter
- $30 of free compute every month
- 3 workspace seats
- 100 containers and 10 GPUs at once
- Real-time metrics and logs; region selection
- Scheduled and web functions limited (5 deployed crons, 200 deployed apps)
- 1-day log retention
- Community Slack support only
Team
- $100 of free compute every month
- Unlimited seats
- 5,000 containers and 50 GPUs at once
- Unlimited crons, custom domains, static IP proxy, deployment rollbacks, environment budgets
- 30-day log retention
- Base fee is charged even in quiet months
Enterprise
- Volume-based discounts
- Higher GPU concurrency
- Audit logs, SAML SSO and HIPAA
- Private Slack support and embedded ML engineering services
- Pricing through sales
GPU compute (all plans)
- B300 about $7.10, B200 $6.25, H200 $4.54, RTX PRO 6000 $3.03, A100 80 GB $2.50 per hour
- A100 40 GB $2.10, L40S $1.95, A10 $1.10, L4 $0.80 per hour
- Up to 8 GPUs per container
- GPU functions are always preemptible
- Region pinning 1.15-1.75x base price
CPU and memory (all plans)
- About $0.047 per physical core-hour (2 vCPUs)
- About $0.008 per GiB-hour of memory
- Minimum 0.125 cores per container
- Non-preemptible CPU execution costs 3x base prices
Sandboxes and Notebooks
- Isolated containers created on demand
- Not preempted unless a GPU is attached
- CPU and memory cost 3x the standard function rates
Volumes (storage)
- First 1 TiB each month free
- Charged beyond 1 TiB
Startup and academic credits
- Startup compute credits
- Up to $10,000 of credits for graduate students, labs and researchers
- Eligibility decided by Modal
Is there a free plan?
Modal's Starter plan costs $0 a month and includes $30 of compute every month, enough for roughly 7.5 hours of H100 GPU time (before CPU and memory) or about 37 hours on an L4. You get 3 seats, up to 100 containers and 10 GPUs at once, 5 deployed cron jobs and 1 day of logs. There is no need to buy credits in advance: usage above $30 is billed per second. Startups and academic researchers can also apply for credit grants, up to $10,000 for academics.
What you will actually pay
Worked examples using the published prices above.
| Scenario | Cost | How we got there |
|---|---|---|
| Hobby image API on an L4, busy 1 hour a day for 30 days (1 core, 16 GiB) | About $29.23/month, covered by the $30 Starter credit | 108,000 seconds. GPU: 108,000 x $0.000222 = $23.98. CPU: 108,000 x $0.0000131 = $1.41. Memory: 16 x 108,000 x $0.00000222 = $3.84. Total $29.23, so $0 after the Starter credit (plus a little idle time per burst). |
| LLM endpoint on one H100, busy 2 hours a day for 30 days (4 cores, 32 GiB), Starter plan | About $233.62/month after the $30 credit | 216,000 seconds. GPU: 216,000 x $0.001097 = $236.95. CPU: 4 x 216,000 x $0.0000131 = $11.32. Memory: 32 x 216,000 x $0.00000222 = $15.34. Total $263.62 minus $30 = $233.62, plus 60 seconds of idle billing after each burst. |
| Same H100 endpoint running 24/7 for 30 days | About $3,163.38/month | 2,592,000 seconds. GPU $2,843.42 + CPU $135.82 + memory $184.14 = $3,163.38. A RunPod Secure Cloud H100 pod left on for 720 hours costs 720 x $3.49 = $2,512.80, so always-on work is cheaper on a rented GPU. |
| Startup spending $1,000 of compute a month | $970 on Starter or $1,150 on Team | Starter: $1,000 - $30 credit = $970, but capped at 3 seats and 10 concurrent GPUs. Team: $250 + ($1,000 - $100) = $1,150, with unlimited seats, 50 GPUs and 30-day logs. |
Modal pricing compared with rivals
| Tool | Paid plans from | Free plan | Note |
|---|---|---|---|
| Modal | Team $250/month + usage ($100 credit); Enterprise custom | Starter: $0/month with $30 of free compute every month | The tool on this page |
| RunPod | H100 pod $2.69/hour (Community); Serverless H100 $4.79/hour | No | Cheaper rented GPUs for steady work; Serverless flex workers also scale to zero. |
| Replicate | H100 $0.001525/sec ($5.49/hour) | No free tier listed | Public models bill only processing time; private models also bill setup and idle time. |
| Together AI | H100 $3.99/GPU-hour (on-demand cluster) | No | GPU clusters plus pay-per-token inference and fine-tuning APIs. |
| Lambda | H100 SXM $4.29/hour (1x); $3.99 per GPU on 8x | No | Simple self-serve instances; no serverless product. Plus sales tax or VAT. |
| Nebius | H100 $3.85/GPU-hour ($4.50 from 1 Oct 2026) | No | Platinum-rated reliability; preemptible H100 from $0.79. |
| CoreWeave | $49.24/hour per 8x H100 node ($6.16/GPU-hour) | No | Large reliable clusters for big training; sales-approved accounts; up to 60% off with commitments. |
Modal is excellent value for bursty work. Because you pay only while containers run, an API that is busy two hours a day on an H100 costs about $264 a month before credits, against about $2,500 for an always-on rented H100. The $30 monthly credit makes small projects free, and there is no base fee on Starter.
It is poor value for steady, round-the-clock GPU use. At about $3.95 an hour plus CPU and memory, a 24/7 H100 on Modal costs roughly $650 more a month than a RunPod Secure Cloud pod. Upgrade to Team ($250) only when you need more than 3 seats or 10 concurrent GPUs. Compare options on our Modal alternatives page.
Modal GPU prices per second and per hour (September 2026)
GPU only; CPU and memory are billed separately. Hourly figures are the per-second price x 3,600, rounded.
| GPU | Per second | About per hour |
|---|---|---|
| B300 | $0.001972 | $7.10 |
| B200 | $0.001736 | $6.25 |
| H200 SXM | $0.001261 | $4.54 |
| H100 SXM5 | $0.001097 | $3.95 |
| RTX PRO 6000 | $0.000842 | $3.03 |
| A100 80 GB | $0.000694 | $2.50 |
| A100 40 GB | $0.000583 | $2.10 |
| L40S | $0.000542 | $1.95 |
| A10 | $0.000306 | $1.10 |
| L4 | $0.000222 | $0.80 |
| T4 | $0.000164 | $0.59 |
| Other resource | Price |
|---|---|
| CPU (physical core = 2 vCPUs) | $0.0000131 per core per second (about $0.047 an hour) |
| Memory | $0.00000222 per GiB per second (about $0.008 an hour) |
| Sandbox and Notebook CPU | $0.00003942 per core per second |
| Sandbox and Notebook memory | $0.00000667 per GiB per second |
| Volumes | $0.09 per GiB a month, first 1 TiB free |
Price multipliers to know
- Region selection: 1.15 to 1.75 times base prices on every plan. An H100 pinned to a region costs about $4.54 to $6.91 an hour instead of $3.95.
- Non-preemptible execution: 3 times base prices for CPU and memory. It is not available for GPU functions, which are always preemptible.
- Sandboxes and Notebooks: CPU and memory cost 3 times the standard function rates.
- H100 to H200 upgrades: Modal may run an H100 request on an H200 at no extra cost; request
H100!if you need exactly an H100.
Paying with cloud commitments or credits
You can arrange to buy Modal through the AWS and GCP marketplaces, so the spend counts against existing cloud commitments; this goes through Modal's sales team. Early-stage startups can apply for free compute credits, and graduate students, labs and researchers can get up to $10,000 in credits through Modal's academic programme.
- Stay on Starter until you actually need more than 3 seats or 10 concurrent GPUs; Team's $250 fee only returns $100 in credit.
- Right-size CPU and memory requests, since you pay for the higher of request or use on every GPU second.
- Use a smaller GPU when the model fits: an L4 costs about $0.80 an hour against $3.95 for an H100.
- Cache model weights in a Volume (first 1 TiB free) so cold starts spend less billed time downloading.
- If a workload is busy most of the day, price it on a rented GPU such as RunPod before scaling it up on Modal.
Jargon explained
- Serverless
- A model where the platform starts computers only when your code needs them and stops them afterwards, so you pay for use rather than for idle machines.
- Cold start
- The delay while a new container starts and loads a model before it can answer its first request.
- Preemption
- When the platform stops your running job to reclaim the machine and restarts it elsewhere. Saving checkpoints lets a job resume without losing work.
- Decorator
- A line starting with @ placed above a Python function that changes how it runs. On Modal it tells the platform to run that function in the cloud.
- Scale to zero
- Shutting down every container when there is no work, so the bill drops to nothing between requests.
Frequently asked questions
How much does Modal cost?
The Starter plan is $0 a month with $30 of free compute; Team is $250 a month with $100 of credit; Enterprise is custom. On top, you pay per second for GPU, CPU and memory, for example about $3.95 an hour for an H100.
Does Modal have a free tier?
Yes. Starter includes $30 of compute every month with no base fee, 3 seats and up to 10 GPUs at once. Usage above $30 is billed per second.
Is Modal billed per second?
Yes. GPU, CPU and memory are all billed per second while containers run, including the default 60-second idle window before a container shuts down.
Why is my Modal bill higher than the GPU price?
Because CPU and memory are billed separately, at the higher of what you requested or used, plus idle time before scale-down, start-up time and any region multiplier.
Is Modal cheaper than RunPod?
For bursty work, usually yes, because Modal scales to zero. For always-on GPUs, no: a RunPod H100 pod costs $2.69 to $3.49 an hour against about $3.95 plus CPU and memory on Modal.
Can I use AWS or Google Cloud credits on Modal?
Modal can be bought through the AWS and GCP marketplaces, so the spend counts against committed cloud spend; you arrange this with Modal's sales team. Azure is not listed.
Sources
Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.
- Modal pricing (Modal)
- Preemption (Modal docs) (Modal)
- Reserving CPU and memory (Modal docs) (Modal)
- Cold start performance (Modal docs) (Modal)
- GPU acceleration (Modal docs) (Modal)
- Multi-node clusters (Modal docs) (Modal)
- How we achieved truly serverless GPUs (Modal)
- About Modal (Modal)
- Serverless AI infrastructure startup Modal Labs seals $355M funding round (SiliconANGLE)
- Modal reviews (Product Hunt)
- ClusterMAX 3.0: the industry standard GPU cloud rating system (SemiAnalysis)
- Runpod GPU cloud pricing (Runpod)
- Replicate pricing (Replicate)
- Together AI pricing (Together AI)
- Lambda pricing (Lambda)
- Nebius AI Cloud prices (Nebius)
- CoreWeave cloud pricing (CoreWeave)