Recommended configuration values
Recommended values for a set of vLLM models and the Parakeet model.
CosmicAC can serve any model that vLLM supports. The following tables list recommended values for a set of models. The field names match the form that opens when you go to Models > Recommended configurations > Add new in the web interface. For what each vLLM environment variable does, see vLLM serving options.
| Field | Value |
|---|
| Model name | Qwen/Qwen3-VL-235B-A22B-Thinking-FP8 |
| Job type | Managed Inference |
| CUDA version | CUDA 13.0 |
| Runtime image | vllm/vllm-openai:v0.15.1 |
| Data type | Auto |
| Quantisation | None |
| Reasoning parser | deepseek_r1 |
| Replicas | 1 |
| GPUs per replica | 8 |
| GPU memory utilisation | 0.9 |
| Max model length | 131072 |
| Max concurrent sequences | 64 |
| Root disk (GB) | 500 |
| Video and image input | On |
Environment variables
| Name | Value |
|---|
TRUST_REMOTE_CODE | true |
SWAP_SPACE | 0 |
ENABLE_EXPERT_PARALLEL | true |
ENFORCE_EAGER | false |
| Field | Value |
|---|
| Model name | Qwen/Qwen3.5-122B-A10B |
| Job type | Managed Inference |
| CUDA version | CUDA 13.0 |
| Runtime image | vllm/vllm-openai:v0.17.1 |
| Data type | Auto |
| Quantisation | None |
| Reasoning parser | deepseek_r1 |
| Replicas | 1 |
| GPUs per replica | 8 |
| GPU memory utilisation | 0.9 |
| Max model length | 32768 |
| Max concurrent sequences | 32 |
| Root disk (GB) | 500 |
| Video and image input | On |
Environment variables
| Name | Value |
|---|
TRUST_REMOTE_CODE | true |
SWAP_SPACE | 0 |
ENABLE_EXPERT_PARALLEL | true |
ENFORCE_EAGER | true |
| Field | Value |
|---|
| Model name | MiniMaxAI/MiniMax-M2.5 |
| Job type | Managed Inference |
| CUDA version | CUDA 13.0 |
| Runtime image | vllm/vllm-openai:v0.15.1 |
| Data type | Auto |
| Quantisation | None |
| Reasoning parser | deepseek_r1 |
| Replicas | 1 |
| GPUs per replica | 4 |
| GPU memory utilisation | 0.85 |
| Max model length | 131072 |
| Max concurrent sequences | 32 |
| Root disk (GB) | 500 |
| Video and image input | On |
Environment variables
| Name | Value |
|---|
TRUST_REMOTE_CODE | true |
ENABLE_EXPERT_PARALLEL | true |
ENFORCE_EAGER | false |
| Field | Value |
|---|
| Model name | Qwen/Qwen2-VL-2B-Instruct |
| Job type | Managed Inference |
| CUDA version | CUDA 13.0 |
| Runtime image | vllm/vllm-openai:v0.15.1 |
| Data type | Auto |
| Quantisation | None |
| Reasoning parser | default |
| Replicas | 1 |
| GPUs per replica | 1 |
| GPU memory utilisation | 0.9 |
| Max model length | 32768 |
| Max concurrent sequences | 64 |
| Root disk (GB) | 150 |
| Video and image input | On |
Environment variables
| Name | Value |
|---|
TRUST_REMOTE_CODE | true |
SWAP_SPACE | 0 |
ENFORCE_EAGER | true |
| Field | Value |
|---|
| Model name | nvidia/parakeet-tdt-0.6b-v3 |
| Job type | Parakeet Inference |
| Chunk duration (seconds) | 10 |
| Chunk overlap (seconds) | 5 |
| Maximum file size (MB) | 1024 |
| Require authentication header | Off |