CosmicAC Logo
Installation

Set up recommended model configurations

Add a recommended model configuration for each model you plan to serve. CosmicAC uses these values to prefill new Managed Inference Jobs.

Add a recommended model configuration for each model you plan to serve. When you create a job, CosmicAC lists only models that have a recommended configuration. For more information about recommended model configurations and how jobs use them, see the overview.

Prerequisites

Before you start, make sure that you have the following.

  • A running CosmicAC deployment. See Set up CosmicAC.
  • Access to the web interface as a platform administrator. On a default deployment, authentication is off, so every user has platform administrator access.
  • The recommended values for each model you plan to serve. See Recommended configuration values.

Steps

In the web interface, go to Models > Recommended configurations, and then click Add new.

Enter the model name and job type

Under Default parameters, enter the Model name.

For Job type, select Managed Inference for a vLLM model or Parakeet Inference for a Parakeet model.

Enter the model configuration

If the model is listed in Recommended configuration values, enter the values listed for the model.

For a Managed Inference model, set the following fields.

  • CUDA version: the CUDA driver version.
  • Runtime image: the vLLM serving image.
  • Data type: the numeric precision the model runs at.
  • Quantisation: the method that compresses the model weights.
  • Reasoning parser: the parser that separates thinking tokens from the final response.
  • Replicas: the number of model copies to serve.
  • GPUs per replica: the number of GPUs each replica uses.
  • GPU memory utilisation: the fraction of GPU memory to use.
  • Max model length: the maximum context length.
  • Max concurrent sequences: the maximum number of requests handled at the same time.
  • Root disk (GB): the root disk size in GB.
  • Environment variables: additional vllm serve options, as name and value pairs.
  • Video and image input: whether the model accepts multimodal input. This setting applies to vision-language models only.

To add an environment variable, click Add variable, and then enter its name and value.

Set the job creation warnings

Under Job creation warnings, select when CosmicAC warns about a job that differs from a recommended value. The person who creates the job can acknowledge the warning and continue.

For a Managed Inference model, select a setting for each of the following fields.

  • GPU memory utilisation
  • Max model length
  • Max concurrent sequences
  • Replicas
  • Root disk

Each field offers the following settings. Root disk offers only No warning and Warn above.

  • No warning: CosmicAC never warns about this field.
  • Warn below: CosmicAC warns when the job's value is lower than the recommended value.
  • Warn above: CosmicAC warns when the job's value is higher than the recommended value.
  • Warn when different: CosmicAC warns when the job's value is lower or higher than the recommended value.

Click Save. The recommended configuration appears on the Recommended configurations page.

Add a configuration for each model

Repeat the previous steps for each model you plan to serve.

Verify the configurations

On the Recommended configurations page, confirm that each model appears in the list.

Next steps

On this page