Skip to content

RunPod

Phase 3.

Pods

  • gateway runs as a small CPU pod; TLS comes from RunPod's proxy URL, so no domain needed
  • GPU pods (secure cloud by default; community cloud is opt-in with a warning) run the engine images directly
  • spot pods for capacity: spot

Serverless

capacity: serverless is the one exception to the tunnel model:

  • the gateway proxies to RunPod's load-balancing endpoint with a shared secret
  • RunPod scales the workers; the gateway sets worker min/max and sets max=0 when the budget cap trips
  • the agent runs in "direct" mode inside the worker container