Hosting cost
Serverless GPU, scale-to-zero — you pay only for seconds used. Figures verified mid-2026. Per-song = one ~3.5-min track through the full pipeline (~90–180 GPU-seconds for Demucs + transcription; lighter steps run on cheap CPU).
Serverless GPU price ($/hour)
| Provider | T4 | L4 | A10 | L40S | A100-80GB | H100 | Cold start |
|---|
| Modal | $0.59 | $0.80 | $1.10 | $1.95 | $2.50 | $3.95 | <5 s |
| RunPod (Flex) | — | ~$0.39 | ~$0.44 | ~$0.86 | ~$1.39 | ~$4.18 | 5–20 s |
| Replicate | $0.81 | — | — | $3.51 | $5.04 | $5.49 | ~11 s |
| fal.ai | — | — | — | — | $1.08 | $1.80 | ~2–5 s |
Estimated cost per 3.5-min song
| Provider / GPU | $ per song | Notes |
|---|
| RunPod Flex · L4 | ~$0.01–0.02 | cheapest; more DevOps (Docker images) |
| Modal · L4 (recommended) | ~$0.02–0.04 | best ergonomics for a custom pipeline; <5 s cold start |
| Modal · T4 | ~$0.015–0.03 | budget; slower, fine for v1 |
| fal.ai · A100 | ~$0.03–0.05 | faster wall-clock |
Monthly projection (Modal · L4, ~$0.03/song)
| Volume / month | Compute | + Vercel/Supabase |
|---|
| 100 songs | ~$3 | free tiers |
| 1,000 songs | ~$30 | mostly free tiers |
| 10,000 songs | ~$300 | ~$25–50 storage/egress |
Third-party stem APIs (no hosting, for comparison)
| Service | $/min audio | $/3.5-min song |
|---|
| AudioShake (bulk) | $0.01–0.05 | $0.04–0.18 |
| Music.ai (5-stem) | $0.07 | ~$0.25 |
| LALAL.AI | $0.06 | ~$0.21 |
| Moises Pro | ~$0.10 | ~$0.35 |
Recommendation: build on Modal + L4 for v1 (scale-to-zero, sub-5 s cold start, ~$0.02–0.04/song). Port to RunPod Flex if pure cost at scale becomes the priority. Self-hosting also gets you transcription + notation, which the stem APIs don't provide. Full sourcing in docs/hosting-cost-chart.md.