Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
116 changes: 116 additions & 0 deletions docs/user/component-catalog.md
Original file line number Diff line number Diff line change
Expand Up @@ -1555,6 +1555,122 @@ also installs `computedomains` from its chart. This pairing has not yet been
verified on a live OpenShift cluster; see
[NVIDIA/aicr#2969](https://github.com/NVIDIA/aicr/issues/2969).

### `mariadb-operator`: `26.6.0` (or earlier) to `26.10.1`

`26.10.0` changes the replication configuration rendered by the MariaDB init
container and the replication liveness probe served by the agent, so the data
plane has to move with the operator, and `26.10.1` keeps that requirement.
`updateStrategy.autoUpdateDataPlane` defaults to `false`, so an operator
upgraded without setting it first runs `26.10.1` against init and agent
containers left at the prior version.

The init container and agent are the HA data plane, so they exist only where
Galera or replication is enabled. AICR's accounting database ships as a single
non-HA instance (`galera.enabled: false`, `replicas: 1`), so steps 1 and 5
below are inert for it: the field is accepted and does nothing. They matter for
any HA `MariaDB` the same operator manages. List what you have with:

```bash
kubectl get mariadb -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,IMAGE:.spec.image,GALERA:.spec.galera.enabled,REPLICATION:.spec.replication.enabled
```

Fresh installs are unaffected, and so is a cluster already on `26.10.0`:
upstream's `26.10.1` guide applies only when coming from before `26.10.x`, and
`26.10.1` adds only optional CRD fields. To migrate an existing cluster from
`26.6.0` or earlier, set the flag **before** the operator moves, then upgrade
CRDs first and the operator second:

1. Enable data-plane auto-update on every **HA** MariaDB the operator manages.
AICR's own accounting database is `mariadb` in namespace `slurm` and is not
HA, so this is a no-op for it:
```bash
kubectl patch mariadb mariadb -n slurm --type merge \
-p '{"spec":{"updateStrategy":{"autoUpdateDataPlane":true}}}'
```
With Argo CD or Flux, do not patch the live resource. Set
`spec.updateStrategy.autoUpdateDataPlane: true` in the desired configuration
the application syncs, commit it, and sync it before step 2, then confirm
the live value with
`kubectl get mariadb <name> -n <namespace> -o jsonpath='{.spec.updateStrategy.autoUpdateDataPlane}'`.
A pruning or self-healing sync restores whatever git says, and if git still
says `false` when the upgraded operator first reconciles an HA `MariaDB`,
the operator's Galera and replication defaulting keep the old init and
agent images: the data-plane upgrade is skipped and the waits in step 4
time out. Keep `true` in git until step 4 passes.
2. Upgrade `mariadb-operator-crds` to `26.10.1` **in place**. Confirm which
namespace the existing release is in first, because Helm scopes a release
by namespace: without `--namespace` the request lands in whatever namespace
the kubeconfig context points at, `--install` does not find the existing
release, and Helm installs a second one that then fights the first for
ownership of the cluster-scoped CRDs. `mariadb-system` is the registry
default and no overlay overrides it, but an inherited bundle may differ.
```bash
helm list -A | grep mariadb-operator-crds
helm upgrade --install mariadb-operator-crds \
oci://ghcr.io/mariadb-operator/charts/mariadb-operator-crds \
--version 26.10.1 --namespace mariadb-system
```
Never `helm uninstall` the CRD chart: that deletes the CRDs and
cascade-deletes every `MariaDB`, `User`, `Database` and `Grant` with them.
3. Upgrade `mariadb-operator` to `26.10.1` (re-run `install.sh`, `helmfile
apply`, or sync the release).
4. For each HA MariaDB patched in step 1, wait until the `26.10.1` data plane
is running before continuing. The `Updated` and `Ready` conditions cannot
show this: the operator computes both against whatever StatefulSet exists,
so they are already `True` before `26.10.1` first reconciles the resource,
and a flag reverted at that point leaves `26.10.1` keeping the old init and
agent images. The waits below cannot pass on that old state. Substitute the
name, namespace and pod names (`<name>-0` up to `<name>-<replicas-1>`), and
set `3` to the replica count:
```bash
IMG=ghcr.io/mariadb-operator/mariadb-operator:26.10.1
# The StatefulSet renders the new data plane.
kubectl wait sts <name> -n <namespace> --timeout=10m \
--for=jsonpath='{.spec.template.spec.initContainers[?(@.name=="init")].image}'=$IMG
kubectl wait sts <name> -n <namespace> --timeout=10m \
--for=jsonpath='{.spec.template.spec.containers[?(@.name=="agent")].image}'=$IMG
# Every pod runs it.
kubectl wait pod <name>-0 <name>-1 <name>-2 -n <namespace> --timeout=15m \
--for=jsonpath='{.status.initContainerStatuses[?(@.name=="init")].image}'=$IMG
kubectl wait pod <name>-0 <name>-1 <name>-2 -n <namespace> --timeout=15m \
--for=jsonpath='{.status.containerStatuses[?(@.name=="agent")].image}'=$IMG
# The roll has finished.
kubectl wait sts <name> -n <namespace> --timeout=15m \
--for=jsonpath='{.status.updatedReplicas}'=3
kubectl wait sts <name> -n <namespace> --timeout=15m \
--for=jsonpath='{.status.readyReplicas}'=3
```
The operator writes the new images into the `MariaDB` spec before it
renders the StatefulSet, so once the template carries them step 5 can no
longer take them back; the pod and replica waits confirm the roll itself.
A resource that names its own init or agent image, such as a mirror, keeps
that repository and takes only the `26.10.1` tag, so set `IMG` to match.
Name the pods rather than selecting them with `-l`: a label selector
resolves the pod list once, so a pod being recreated when the command
starts is never waited on, while a missing named pod fails the command and
you rerun it. The replica counts mean something only after the image waits,
because before the new template exists they already equal the replica
count for the old revision. Skip non-HA instances: they have no init or
agent container, so these waits run to their timeout.
5. Return the flag to `false` so a later operator bump does not update the data
plane unattended. With Argo CD or Flux, set it back to `false` in the
desired configuration, commit it, and sync, only after step 4 passes.

`26.10.0` also changes the operator's default server image to
`mariadb:12.3.3`, and `26.10.1` keeps it. AICR pins `mariadb:11.8.8` in
`recipes/components/slurm-accounting-mariadb/values.yaml`, so a cluster bundled
from this recipe stays on `11.8.8`; the pin also puts the server image into the
rendered `MariaDB` resource, where the BOM records it. Any `MariaDB` that omits
`spec.image` takes the operator default and will move to a new MariaDB major
version on its next reconcile. Check with:

```bash
kubectl get mariadb -A \
-o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,IMAGE:.spec.image
```

A blank `IMAGE` column means that cluster takes the default.

### `agentgateway`: upgrading across breaking releases

AICR pins the `agentgateway` and `agentgateway-crds` charts in the component
Expand Down
52 changes: 26 additions & 26 deletions docs/user/container-images.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ A machine-readable **CycloneDX 1.6 JSON** companion to this page is produced by
## Summary

- Components: **49**
- Unique images: **113**
- Unique images: **114**
- Distinct registries: **11**

Registries: `602401143452.dkr.ecr.us-west-2.amazonaws.com`, `cr.agentgateway.dev`, `docker.io`, `gcr.io`, `ghcr.io`, `gke.gcr.io`, `nvcr.io`, `public.ecr.aws`, `quay.io`, `registry.k8s.io`, `us-docker.pkg.dev`
Expand Down Expand Up @@ -53,12 +53,12 @@ _Rendering fidelity:_ `catalog-parity: charts are rendered with the shared recip
| k8s-ephemeral-storage-metrics | helm | k8s-ephemeral-storage-metrics/k8s-ephemeral-storage-metrics | 1.19.2 | 1 |
| k8s-nim-operator | helm | k8s-nim-operator | 3.1.0 | 1 |
| k8s-nim-operator-ocp | helm | k8s-nim-operator | 3.1.0 | 1 |
| kai-scheduler | helm | kai-scheduler | v0.16.9 | 12 |
| kai-scheduler | helm | kai-scheduler | v0.17.2 | 12 |
| kube-prometheus-stack | helm | prometheus-community/kube-prometheus-stack | 84.4.0 | 8 |
| kubeflow-trainer | helm | kubeflow-trainer | 2.2.0 | 4 |
| kueue | helm | kueue | 0.19.3 | 1 |
| mariadb-operator | helm | mariadb-operator | 26.6.0 | 1 |
| mariadb-operator-crds | helm | mariadb-operator-crds | 26.6.0 | 0 |
| kueue | helm | kueue | 0.19.6 | 1 |
| mariadb-operator | helm | mariadb-operator | 26.10.1 | 1 |
| mariadb-operator-crds | helm | mariadb-operator-crds | 26.10.1 | 0 |
| network-operator | helm | nvidia/network-operator | 26.4.1 | 12 |
| network-operator-ocp | manifest | — | — | 0 |
| network-operator-ocp-olm | manifest | — | — | 0 |
Expand All @@ -75,11 +75,11 @@ _Rendering fidelity:_ `catalog-parity: charts are rendered with the shared recip
| prometheus-adapter | helm | prometheus-community/prometheus-adapter | 5.3.0 | 1 |
| prometheus-adapter-ocp | helm | prometheus-community/prometheus-adapter | 5.3.0 | 1 |
| prometheus-operator-crds | helm | prometheus-community/prometheus-operator-crds | 28.0.1 | 0 |
| slinky-slurm | helm | slurm | 1.2.0 | 5 |
| slinky-slurm-operator | helm | slurm-operator | 1.2.0 | 2 |
| slinky-slurm-operator-crds | helm | slurm-operator-crds | 1.2.0 | 0 |
| slinky-slurm | helm | slurm | 1.2.2 | 5 |
| slinky-slurm-operator | helm | slurm-operator | 1.2.2 | 2 |
| slinky-slurm-operator-crds | helm | slurm-operator-crds | 1.2.2 | 0 |
| slinky-topograph | helm | topograph/topograph | 1.0.0 | 1 |
| slurm-accounting-mariadb | helm | mariadb-cluster | 26.6.0 | 0 |
| slurm-accounting-mariadb | helm | mariadb-cluster | 26.10.1 | 1 |

## Version variants

Expand Down Expand Up @@ -212,18 +212,18 @@ _No images extracted._

### kai-scheduler

- `ghcr.io/kai-scheduler/kai-scheduler/admission:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/binder:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/crd-upgrader:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/nodescaleadjuster:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/numa-placement-exporter:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/operator:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/podgroupcontroller:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/podgrouper:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/queuecontroller:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/resourcereservation:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/scalingpod:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/scheduler:v0.16.9`
- `ghcr.io/kai-scheduler/kai-scheduler/admission:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/binder:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/crd-upgrader:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/nodescaleadjuster:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/numa-placement-exporter:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/operator:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/podgroupcontroller:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/podgrouper:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/queuecontroller:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/resourcereservation:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/scalingpod:v0.17.2`
- `ghcr.io/kai-scheduler/kai-scheduler/scheduler:v0.17.2`

### kube-prometheus-stack

Expand All @@ -245,11 +245,11 @@ _No images extracted._

### kueue

- `registry.k8s.io/kueue/kueue:v0.19.3`
- `registry.k8s.io/kueue/kueue:v0.19.6`

### mariadb-operator

- `ghcr.io/mariadb-operator/mariadb-operator:26.6.0`
- `ghcr.io/mariadb-operator/mariadb-operator:26.10.1`

### mariadb-operator-crds

Expand Down Expand Up @@ -352,8 +352,8 @@ _No images extracted._

### slinky-slurm-operator

- `ghcr.io/slinkyproject/slurm-operator-webhook:1.2.0`
- `ghcr.io/slinkyproject/slurm-operator:1.2.0`
- `ghcr.io/slinkyproject/slurm-operator-webhook:1.2.2`
- `ghcr.io/slinkyproject/slurm-operator:1.2.2`

### slinky-slurm-operator-crds

Expand All @@ -365,7 +365,7 @@ _No images extracted._

### slurm-accounting-mariadb

_No images extracted._
- `mariadb:11.8.8`

### kube-prometheus-stack@83.7.0 (variant)

Expand Down
Loading
Loading