File size: 1,654 Bytes
dfb775d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
# Ops

Container, compose, Kubernetes, vmm, and Gensyn manifests for mindxtrain. Per
mindxtrain2.md Β§Part 4 (`ops/`).

## Layout

```
ops/
β”œβ”€β”€ containerfiles/
β”‚   β”œβ”€β”€ containerfile_train     # FROM rocm/primus:v26.2          (training image)
β”‚   β”œβ”€β”€ containerfile_serve     # FROM rocm/vllm-dev:rocm7.2.1    (inference image)
β”‚   └── digest.lock             # SHA256 pins, populated post-pull
β”œβ”€β”€ compose/
β”‚   └── compose_dev.yaml        # full stack: vLLM + operator FastAPI for the live demo
β”œβ”€β”€ k8s/
β”‚   └── train_job.yaml          # single-MI300X training Job
β”œβ”€β”€ vmm/                        # OpenBSD vmm vm definitions (post-hackathon)
└── gensyn/                     # Gensyn distributed-training configs (post-hackathon)
```

## Build the train image (on the MI300X droplet)

```bash
podman build -f ops/containerfiles/containerfile_train -t mindxtrain/train:latest .
podman inspect --format '{{index .RepoDigests 0}}' mindxtrain/train:latest \
  | tee -a ops/containerfiles/digest.lock
```

## Build the serve image

```bash
podman build -f ops/containerfiles/containerfile_serve -t mindxtrain/serve:latest .
```

## Run the demo stack

```bash
podman-compose -f ops/compose/compose_dev.yaml up -d
# vLLM-ROCm at :8000, mindxtrain operator FastAPI at :8080
```

## Submit the K8s training job

```bash
kubectl create configmap mindxtrain-run-config --from-file=run.yaml=run.yaml
kubectl apply -f ops/k8s/train_job.yaml
kubectl logs -f job/mindxtrain-job
```

The Job spec assumes a node labeled `accelerator: mi300x` and the AMD ROCm
device plugin (`amd.com/gpu` resource).