The deploy file¶
deploy.yaml is the only artifact you keep. It holds desired state plus the two fields apply writes back (gateway.url, gateway.admin_token). Any value can be a literal or ${ENV_VAR} (with :-default fallbacks); the CLI resolves env vars locally before pushing.
Unknown fields are a hard error, so typos fail loudly.
Full schema¶
version: 1 # must be 1
name: acme-ai # [a-z0-9-]{3,32}, used in resource names and tags
provider:
type: aws # aws | gcp | azure | runpod | lambda | hyperbolic | ssh | fake
region: us-east-1 # not used for ssh
capacity: spot-fallback # spot | on-demand | spot-fallback | serverless (runpod only)
subnet: null # optional aws/gcp/azure: use instead of the default network
hosts: # ssh only: your machines
- {addr: 10.0.0.5, user: ubuntu, gpus: 2, cost_hr: 0.30}
budget:
monthly_usd: 500 # hard cap; null/absent = unlimited (ssh defaults to unlimited)
alerts:
webhook: https://ntfy.sh/... # optional; receives generic JSON posts at 80/90/95%
gateway:
domain: ai.acme.com # optional; default <ip-dashed>.sslip.io
host: null # optional ssh://user@host; required for lambda and hyperbolic
url: https://... # written by apply
admin_token: gr_admin_... # written by apply
models: # ids must exist in the catalog
- id: llama-3.1-8b-instruct
variant: awq # optional; one of the catalog's tested variants
replicas: {min: 0, max: 4} # defaults 0 / 1
idle_timeout: 15m # scale to zero after this idle; default 15m
keys:
- name: alice
key: ${ALICE_KEY} # the value users present as Bearer tokens
models: ["*"] # which catalog ids this key may use
limits:
rpm: 60 # token bucket
tokens_per_day: 200000 # from engine usage (streaming included)
stt_mb_per_day: 200 # request body size
tts_chars_per_day: 100000 # input length
secrets:
hf_token: ${HF_TOKEN} # needed for gated models; passed to nodes
Semantics¶
- capacity —
on-demandis what people usually call "dedicated".spot-fallbacktries spot first and falls back to on-demand at launch time and on interruption.serverless(RunPod only) delegates scaling to the provider. - ssh pools are fixed —
maxis bounded by the GPUs that physically exist. Spend iscost_hrx uptime; the budget is ignored. - api model names — the
modelfield in requests is the catalogid, e.g."llama-3.1-8b-instruct". - secrets — the file is plaintext by design and belongs to you; use env references for hygiene. The gateway stores the resolved config in its data dir.
Validation¶
spec.Validate runs in the CLI before any push, and again in the gateway on receipt; a rejected push keeps the previous config running.