Set up a relay for CosmicAC
Route CosmicAC job-agent connections through a relay when your cluster network blocks direct connections.
Set up a relay when your cluster network blocks direct connections between CosmicAC job agents and cosmicac-wrk-server-k8s-nvidia.
CosmicAC components normally connect directly over the Holepunch peer-to-peer stack. Use a relay when either of the following problems occurs.
- Communication failure: cosmicac-wrk-agent-inference or cosmicac-wrk-agent-instance can't reach cosmicac-wrk-server-k8s-nvidia. Job statuses stop updating, and the last communication can be minutes old. A job can also remain in Creating or Starting after CosmicAC creates it.
- Access issues: users can't open a shell in the GPU Container Jobs on a cluster.
Prerequisites
Before you start, make sure that you have the following.
- A running CosmicAC deployment. See Set up CosmicAC.
- A relay host machine that runs Node.js and npm. The deployment host machine and the job VMIs on your cluster must both be able to reach it.
- An open UDP port on the relay host that no firewall blocks. The relay listens on port
49737by default. - Enough free memory on the relay host machine. The relay uses up to 200 MB of RAM and little CPU.
Steps
Install the relay service
On the relay host machine, install blind-relay-service.
npm install -g blind-relay-serviceStart the relay
Start the relay with a persistent storage directory.
blind-relay --storage /var/lib/blind-relayUse the same storage directory each time you restart the relay. The relay creates its key pair from the Corestore in this directory. If you start the relay with a different storage directory, it can generate a different key. The key in CosmicAC then no longer matches, and jobs can't reach the relay.
If you don't specify --storage, the relay uses ./corestore in the directory where you run the command.
To use a different UDP port, add --port <port>.
The relay prints its storage path and key.
Using corestore storage at /var/lib/blind-relay
Server listening on <relay-key>Copy the key from the Server listening on line. You need it in the next step.
Configure the relay key
On the deployment host machine, add the relay key to .env, so that cosmicac-wrk-server-k8s-nvidia passes it to the jobs that it creates.
K8S_COMMON_CONFIG__interconnect__relays='["<relay-key>"]'The value is a JSON array. For the variable format, see Config overrides.
Apply the configuration and restart the worker
Apply the updated configuration.
task apply-wrk-server-k8s-nvidia-common-configRestart the worker.
task restart SERVICES="cosmicac-wrk-server-k8s-nvidia"Existing jobs don't receive the relay key
Only jobs created after the worker restarts use the relay. The worker passes the relay key to a job when it creates the job.
To recreate a stuck job, delete it, and then create it again.
cosmicac jobs delete <job-id>Verify the relay
Create a new job, and then list the jobs.
cosmicac jobs listThe new job leaves Creating, and its status keeps updating.