Hosting cost

Serverless GPU, scale-to-zero — you pay only for seconds used. Figures verified mid-2026. Per-song = one ~3.5-min track through the full pipeline (~90–180 GPU-seconds for Demucs + transcription; lighter steps run on cheap CPU).

Serverless GPU price ($/hour)

ProviderT4L4A10L40SA100-80GBH100Cold start
Modal$0.59$0.80$1.10$1.95$2.50$3.95<5 s
RunPod (Flex)~$0.39~$0.44~$0.86~$1.39~$4.185–20 s
Replicate$0.81$3.51$5.04$5.49~11 s
fal.ai$1.08$1.80~2–5 s

Estimated cost per 3.5-min song

Provider / GPU$ per songNotes
RunPod Flex · L4~$0.01–0.02cheapest; more DevOps (Docker images)
Modal · L4 (recommended)~$0.02–0.04best ergonomics for a custom pipeline; <5 s cold start
Modal · T4~$0.015–0.03budget; slower, fine for v1
fal.ai · A100~$0.03–0.05faster wall-clock

Monthly projection (Modal · L4, ~$0.03/song)

Volume / monthCompute+ Vercel/Supabase
100 songs~$3free tiers
1,000 songs~$30mostly free tiers
10,000 songs~$300~$25–50 storage/egress

Third-party stem APIs (no hosting, for comparison)

Service$/min audio$/3.5-min song
AudioShake (bulk)$0.01–0.05$0.04–0.18
Music.ai (5-stem)$0.07~$0.25
LALAL.AI$0.06~$0.21
Moises Pro~$0.10~$0.35

Recommendation: build on Modal + L4 for v1 (scale-to-zero, sub-5 s cold start, ~$0.02–0.04/song). Port to RunPod Flex if pure cost at scale becomes the priority. Self-hosting also gets you transcription + notation, which the stem APIs don't provide. Full sourcing in docs/hosting-cost-chart.md.