Create a Parakeet Managed Inference Job with the CLI
Create a speech-to-text Parakeet Managed Inference Job with the CLI and verify that it's running.
Create a Parakeet Managed Inference Job with the CosmicAC CLI to serve a speech-to-text model behind an OpenAI-compatible transcription endpoint.
Prerequisites
Before you start, make sure that you have the following.
- A running CosmicAC deployment. See Set up CosmicAC.
- The CosmicAC CLI installed and configured. See Install the CLI.
- A recommended configuration for the Parakeet model you want to serve. See Set up recommended model configurations.
Steps
Create the job
Create the job interactively by answering prompts, or pass the job configuration as flags.
Start the interactive job setup.
cosmicac jobs createSelect Managed Inference (Parakeet) as the job type, and then answer the prompts.
Configure the following fields.
- Job name: a name that identifies the job.
- Tags: comma-separated labels for the job.
- Location: the region where the job runs.
- GPU type: the GPU to use. The CLI lists each GPU type with the number of free GPUs on one node and in the selected location.
- GPU count: the number of GPUs for one replica. One of 1, 2, 4, or 8.
- Model: the Parakeet model to serve, such as
nvidia/parakeet-tdt-0.6b-v3. The CLI lists the Parakeet models that have a recommended configuration. - Chunk duration: the audio chunk length in seconds. Minimum
10. - Chunk overlap: the overlap between chunks in seconds. Must be less than the chunk duration. Minimum
5. - Max file size: the maximum audio upload size in MB. Minimum
1024. - Endpoint name: the name used in the endpoint URL. Use lowercase letters, numbers, and hyphens only. Interactive mode checks that the name is available and shows the endpoint URL.
- Replicas: the number of model copies to run. CosmicAC doesn't autoscale replicas.
- Require Authorization header: whether callers must send an API key. See Create an API key.
- Notifications: the job lifecycle events this job reports. All four are on by default, and interactive mode prompts for them with a checkbox.
An event reaches your webhook only if it's also turned on in Settings > Notifications. See What controls delivery.
In interactive mode, the CLI shows a job summary and asks you to confirm the job. If a team is active, the prompt shows the team ID. Enter yes to create the job. The CLI then shows the job ID.
For every field and its CLI flag, see Parakeet Managed Inference Job configuration.
Verify the job
List your jobs.
cosmicac jobs listCheck that the new job appears in the list with its ID, name, tags, and status. Wait for the job to reach running. The endpoint accepts requests after the job reaches this status.
Help and troubleshooting
Job stuck in Creating or Starting
If a job stays in Creating or Starting, check the status of its KubeVirt virtual machine instance (VMI).
-
Find the job's container ID.
cosmicac jobs detail <jobId>The output lists the Container ID for each container.
-
Find the VMI for the container.
CosmicAC creates one VMI for each container and names it
<container-id>-n0. A multi-node job has one VMI per node.From a machine with
kubectlaccess to your Kubernetes cluster, run the following command.kubectl get vmi -n <namespace>Replace
<namespace>with the namespace configured inK8S_NAMESPACE. -
Check the VMI status.
-
If the VMI is not Running, inspect its events.
kubectl describe vmi <container-id>-n0 -n <namespace> -
If the VMI is Running but the job stays in Creating or Starting, cosmicac-wrk-agent-inference cannot reach cosmicac-wrk-server-k8s-nvidia. These two components connect directly, and some cluster network configurations can block the connection.
To route the connection through a relay, see Set up a relay for CosmicAC.
-