Connect to a GPU Container Job
Open an interactive shell in a running GPU Container Job with the CLI.
Open a shell in a running GPU Container Job to run commands directly on the container.
Prerequisites
Before you start, make sure that you have the following.
- The CosmicAC CLI installed and configured. See Install the CLI.
- A running GPU Container Job. See Create a GPU Container Job with the CLI or Create a GPU Container Job in the web interface.
Steps
Find the container index
View the job details. Replace <jobId> with the job ID.
cosmicac jobs detail <jobId>In the Containers section, each container has an index that starts at 0, such as Container 0. Note the index of the container you want to connect to.
Open the container shell
Make sure that the job's status is running. If the job is still starting, wait until it's running.
Open a shell in the container. Replace <jobId> with the job ID and <containerId> with the container index.
cosmicac jobs shell <jobId> <containerId>The shell opens as appuser. This user can run commands as root with sudo. To close the shell, run exit.
Help and troubleshooting
nvidia-smi fails with Failed to initialize NVML: Unknown Error
If nvidia-smi displays Failed to initialize NVML: Unknown Error, the NVIDIA device files under /dev and /proc/driver/nvidia/version are present, but the GPU is not available to nvidia-smi.
-
Restart the container from the shell.
kill 1If you need root access, run
sudo kill 1instead. -
Reconnect to the container with
cosmicac jobs shelland runnvidia-smiagain.
The shell doesn't open when the job is Running
cosmicac-cli connects to cosmicac-wrk-agent-instance over hyperswarm-ssh. The connection then goes directly to the job's virtual machine.
Some cluster network configurations can block this connection, so the job can stay Running while the shell connection fails.
To route the connection through a relay, see Set up a relay for CosmicAC.