GPU batch on Modal architecture

Bursty GPU work triggered from a queue, with inputs and outputs in cheap object storage.

PaaSmodalgpuinferencebatch
Submit API1 GB RAM · light vCPUJob queueFixed 250 MB planGPU workers~50 h A10G at $1.10/hInputs500 GB · no egress feeOutputs500 GB · no egress fee
5 nodes. Drawn by the same layout the product uses.
What is in it

Every resource, and what it costs.

Projections from August 2026 list prices for always-on resources. Connect an account and these become the figures your provider actually bills.

Resources in the GPU batch on Modal template with projected monthly cost.
NodeTypeWhat it isProjected
Submit APICompute1 GB RAM · light vCPUfrom $12/mo
Job queueDatabaseFixed 250 MB planfrom $10/mo
GPU workersCompute~50 h A10G at $1.10/hfrom $55/mo
InputsStorage500 GB · no egress fee$8.00/mo
OutputsStorage500 GB · no egress fee$8.00/mo

Agent steps are priced from provider-reported token usage once the agent runs, not estimated. Guardrails cost nothing and are the reason a runaway agent cannot. 3 rows show from because those services bill by usage, so the total is a scenario at the stated volumes, not a quote.

Open GPU batch on Modal on the canvas.

It loads as an editable graph. Connect an account or instrument an agent and the projected figures above become measured ones.