Contributing recipes¶
Recipes live in catalog/*.yaml in the main repo. One file per model.
Rules¶
- You must have actually run it. Fill the
tested:block with the date, provider, gpu and engine version. CI rejects recipes without it. - Prefer tested quantized variants — they are the cheap path for most operators (awq 8B fits an L4).
- Order
gpusby preference — the planner ranks the listed classes by price, but only the classes you list. - Set a conservative
concurrency— it is the autoscaler target per replica. - Only LLM, STT and TTS. Only
vllm,llamacpp,speachesengines for v1.
PR process¶
- CI validates schema and the
tested:block - a maintainer may ask for the exact flags you used; put anything non-obvious in
args: - gpu smoke tests are manual for now: run the variant on the listed gpu class, hit
/health, run one chat/audio request, then submit
Example¶
id: qwen2.5-7b-instruct
type: llm
license: apache-2.0
gated: false
default_variant: awq
variants:
- name: awq
engine: vllm
hf_repo: Qwen/Qwen2.5-7B-Instruct-AWQ
gpus: [L4, A10G, L40S]
args: [--max-model-len, "16384", --gpu-memory-utilization, "0.92"]
capacity: {concurrency: 32}
tested: {date: "2026-10-01", provider: aws, gpu: L4, engine_version: "0.11.0"}