Register a vLLM model
Register a vLLM model on the Models page to create its recommended model configuration, with values prefilled from vLLM Recipes.
Register a vLLM model to create its recommended model configuration. If the model has a published vLLM recipe, CosmicAC can prefill the values from it.
Prerequisites
Before you start, make sure that you have the following.
- A running CosmicAC deployment. See Set up CosmicAC.
- The platform administrator role. See Teams and roles.
- The model's Hugging Face ID, in the format
namespace/name.
Steps
Open the Register model page
In the left navigation, click Models, and then click Register model.
Select the model type
Select Deployable · vLLM.
Enter the model ID
In Hugging Face model ID, enter the model's ID, such as Qwen/Qwen3-VL.
To fill the values from the model's vLLM recipe, click Prefill configuration. CosmicAC uses the matching configuration from vLLM Recipes.
Set the recommended values
Under Recommended values, set the following fields.
- Recommended GPU count: the number of GPUs for one replica.
- Runtime image: the vLLM serving image.
- Data type: the numeric precision that the model runs at.
- Quantisation: the method that compresses the model weights.
- GPU memory utilisation: the fraction of GPU memory to use.
- Max model length: the maximum context length.
- Max concurrent sequences: the maximum number of requests handled at the same time.
- Root disk (GB): the root disk size, in GB.
Set the deviation directions
Under Deviation directions, select when CosmicAC warns about a new job for GPU count, GPU memory utilisation, Max concurrent sequences, and Max model length. By default, GPU count is set to Below unsafe, and the other three are set to Above unsafe.
- Below unsafe: CosmicAC warns when the job's value is lower than the recommended value.
- Above unsafe: CosmicAC warns when the job's value is higher than the recommended value.
- Exact: CosmicAC warns when the job's value differs from the recommended value.
Register the model
Click Register. CosmicAC returns you to the Models page. To see the model's recommended configuration, click Recommended configurations.