Install OpenShift 4 on cloudscale.ch
Steps to install an OpenShift 4 cluster without Puppet-managed LBs on cloudscale.ch.
These steps follow the Installing a cluster on bare metal docs to set up a user provisioned installation (UPI). Terraform is used to provision the cloud infrastructure.
|
The commands are idempotent and can be retried if any of the steps fail. The certificates created during bootstrap are only valid for 24h. So make sure you complete these steps within 24h. |
|
If you’re setting up a cluster that has more demanding routing requirements, consider setting up a cluster with Puppet LBs. See "Install OpenShift 4 with Puppet-managed LBs on cloudscale.ch" for details. |
Starting situation
-
You already have a Tenant and its git repository
-
You have a CCSP Red Hat login and are logged into Red Hat Openshift Cluster Manager
Don’t use your personal account to login to the cluster manager for installation. -
You want to register a new cluster in Lieutenant and are about to install Openshift 4 on cloudscale.ch
Guided Setup
This how-to guide is exported from the Guided Setup automation tool. It’s highly recommended to run these instructions using said tool, as opposed to running them manually.
Guided Setup can be run easily in docker using the following aliases:
|
These shell scripts depend on |
guided-setup-base() {
local OPTIND opt extra_env extra_volume extra_groups docker_group docker_path sshop_volumes ecr_volume
extra_env=()
extra_volume=()
sshop_volumes=()
ecr_volume=()
while getopts 'h?v:e:' opt; do
case "$opt" in
h|\?)
echo "usage: $0 [-y] [-v EXTRA_VOLUME_MOUNT] [-e EXTRA_ENV_VAR]"
;;
e)
extra_env+=(--env "$OPTARG")
;;
v)
extra_volume+=(--volume "$OPTARG")
;;
esac
done
shift $((OPTIND-1))
local pubring="${HOME}/.gnupg/pubring.kbx"
if command -v gpgconf &>/dev/null && test -f "${pubring}"; then
gpg_opts=(--volume "${pubring}:/app/.gnupg/pubring.kbx:ro" --volume "$(gpgconf --list-dir agent-extra-socket):/app/.gnupg/S.gpg-agent:ro")
else
gpg_opts=()
fi
readonly GANDALF_CONFIG="${XDG_CONFIG_HOME:-~/.config}/gandalf"
if [[ -f "${GANDALF_CONFIG}/env" ]]
then
extra_env+=(--env-file "${GANDALF_CONFIG}/env")
fi
docker_group="$( getent group docker | cut --delimiter ':' --fields 3 )"
docker_path=${DOCKER_HOST:-/var/run/docker.sock}
if [[ -n "$docker_group" ]]; then
extra_groups=("--group-add=$( getent group docker | cut --delimiter ':' --fields 3 )")
fi
if [[ "$OSTYPE" == "linux-gnu"* ]]; then
open="xdg-open"
elif [[ "$OSTYPE" == "darwin"* ]]; then
open="open"
fi
if [[ -e "${HOME}/.ssh/sshop_config" ]]
then
sshop_volumes=("--volume" "${HOME}/.ssh/sshop_config:/app/.ssh/sshop_config" "--volume" "${HOME}/.ssh/sshop_known_hosts:/app/.ssh/sshop_known_hosts")
fi
if [[ -d "${HOME}/.config/emergency-credentials-receive/" ]]
then
ecr_volume=("--volume" "${HOME}/.config/emergency-credentials-receive/:/app/.config/emergency-credentials-receive:ro")
fi
socat \
tcp-listen:8105,fork,reuseaddr,bind=127.0.0.1 \
system:"xargs $open" &
# NOTE(aa): Host network is required for the Vault OIDC callback, since Vault only binds the callback handler to 127.0.0.1
# cf. https://github.com/hashicorp/vault/issues/29064
docker run \
--interactive=true \
--tty \
--rm \
--user="$(id -u)" \
"${extra_groups[@]}" \
--env SSH_AUTH_SOCK=/tmp/ssh_agent.sock \
--env GLAB_CONFIG_DIR="${PWD}" \
--network host \
--volume "${SSH_AUTH_SOCK}:/tmp/ssh_agent.sock" \
--volume "${HOME}/.ssh/config:/app/.ssh/config:ro" \
--volume "${HOME}/.ssh/known_hosts:/app/.ssh/known_hosts:ro" \
--volume "${HOME}/.gitconfig:/app/.gitconfig:ro" \
--volume "${HOME}/.cache:/app/.cache" \
--volume "${XDG_CONFIG_HOME:-~/.config}/gandalf:/app/.config/gandalf" \
--volume "${XDG_CONFIG_HOME:-~/.config}/io.vshn.kharon:/app/.config/io.vshn.kharon" \
--volume "${docker_path}":/var/run/docker.sock \
"${sshop_volumes[@]}" \
"${ecr_volume[@]}" \
"${extra_volume[@]}" \
"${extra_env[@]}" \
"${gpg_opts[@]}" \
--volume "${PWD}:${PWD}" \
--workdir "${PWD}" \
ghcr.io/appuio/guided-setup:latest \
"${@}"
kill %socat
}
guided-setup() {
local OPTIND opt workflow_dir
workflow_dir=
while getopts 'h?w:' opt; do
case "$opt" in
h|\?)
echo "usage: $0 [-y] [-w LOCAL_WORKFLOW_DIR]"
;;
w)
workflow_dir="$OPTARG"
;;
esac
done
shift $((OPTIND-1))
if [[ -z $workflow_dir ]]
then
guided-setup-base "${1}" "/workflows/${2}.workflow" "/workflows/${2}/*.yml" "/workflows/shared/*.yml" "${@:3}"
else
guided-setup-base -v "$workflow_dir":/workflows "${1}" "/workflows/${2}.workflow" "/workflows/${2%%-*}/*.yml" "/workflows/shared/*.yml" "${@:3}"
fi
}
To run the workflows detailed below, simply source the above aliases and run:
guided-setup run cloudscale
|
It’s recommended to run |
Prerequisites
-
jq -
yqyq YAML processor (version 4 or higher - use the go version by mikefarah, not the jq wrapper by kislyuk) -
vaultVault CLI -
curl -
emergency-credentials-receiveInstall instructions -
commodore, see Installing Commodore -
kapitan(should automatically be available in$PATHifcommodoreis installed withuvas described in the link above) -
gzip -
docker -
mc>=RELEASE.2024-01-18T07-03-39ZMinio client (aliased tomcif necessary) -
awsCLI Official install instructions. You can also install the Python package with your favorite package manager (we recommenduv:uv tool install awscli).
Workflow
Given I have all prerequisites installed
This step checks if all necessary prerequisites are installed on your system, including 'yq' (version 4 or higher, by Mike Farah) and 'oc' (OpenShift CLI).
Script
OUTPUT=$(mktemp)
set -euo pipefail
echo "Checking prerequisites..."
if which yq >/dev/null 2>&1 ; then { echo "✅ yq is installed."; } ; else { echo "❌ yq is not installed. Please install yq to proceed."; exit 1; } ; fi
if yq --version | grep -E 'version v[4-9]\.' | grep 'mikefarah' >/dev/null 2>&1 ; then { echo "✅ yq by mikefarah version 4 or higher is installed."; } ; else { echo "❌ yq version 4 or higher is required. Please upgrade yq to proceed."; exit 1; } ; fi
if which jq >/dev/null 2>&1 ; then { echo "✅ jq is installed."; } ; else { echo "❌ jq is not installed. Please install jq to proceed."; exit 1; } ; fi
if which oc >/dev/null 2>&1 ; then { echo "✅ oc (OpenShift CLI) is installed."; } ; else { echo "❌ oc (OpenShift CLI) is not installed. Please install oc to proceed."; exit 1; } ; fi
if which vault >/dev/null 2>&1 ; then { echo "✅ vault (HashiCorp Vault) is installed."; } ; else { echo "❌ vault (HashiCorp Vault) is not installed. Please install vault to proceed."; exit 1; } ; fi
if which curl >/dev/null 2>&1 ; then { echo "✅ curl is installed."; } ; else { echo "❌ curl is not installed. Please install curl to proceed."; exit 1; } ; fi
if which docker >/dev/null 2>&1 ; then { echo "✅ docker is installed."; } ; else { echo "❌ docker is not installed. Please install docker to proceed."; exit 1; } ; fi
if which glab >/dev/null 2>&1 ; then { echo "✅ glab (GitLab CLI) is installed."; } ; else { echo "❌ glab (GitLab CLI) is not installed. Please install glab to proceed."; exit 1; } ; fi
if which dig >/dev/null 2>&1 ; then { echo "✅ dig (DNS lookup utility) is installed."; } ; else { echo "❌ dig (DNS lookup utility) is not installed. Please install dig to proceed."; exit 1; } ; fi
if which mc >/dev/null 2>&1 ; then { echo "✅ mc (MinIO Client) is installed."; } ; else { echo "❌ mc (MinIO Client) is not installed. Please install mc >= RELEASE.2024-01-18T07-03-39Z to proceed."; exit 1; } ; fi
mc_version=$(mc --version | grep -Eo 'RELEASE[^ ]+')
if echo "$mc_version" | grep -E 'RELEASE\.202[4-9]-' >/dev/null 2>&1 ; then { echo "✅ mc version ${mc_version} is sufficient."; } ; else { echo "❌ mc version ${mc_version} is insufficient. Please upgrade mc to >= RELEASE.2024-01-18T07-03-39Z to proceed."; exit 1; } ; fi
if which aws >/dev/null 2>&1 ; then { echo "✅ aws (AWS CLI) is installed."; } ; else { echo "❌ aws (AWS CLI) is not installed. Please install aws to proceed. Our recommended installer is uv: 'uv tool install awscli'"; exit 1; } ; fi
if which restic >/dev/null 2>&1 ; then { echo "✅ restic (Backup CLI) is installed."; } ; else { echo "❌ restic (Backup CLI) is not installed. Please install restic to proceed."; exit 1; } ; fi
if which kharon >/dev/null 2>&1 ; then { echo "✅ kharon (Cluster access helper) is installed."; } ; else { echo "❌ kharon is not installed. Please install it from https://github.com/vshn/kharon ."; exit 1; } ; fi
if which commodore >/dev/null 2>&1 ; then { echo "✅ commodore (Project Syn) is installed."; } ; else { echo "❌ commodore (Project Syn) is not installed. Please install it with 'uv tool install syn-commodore && commodore tool install --missing' ."; exit 1; } ; fi
echo "✅ All prerequisites are met."
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I download the openshift-install binary for version "4.21"
This step downloads the openshift-install binary for the specified OpenShift version.
Script
OUTPUT=$(mktemp)
set -euo pipefail
openshift-install() {
./openshift-install "${@}"
}
if [[ -f openshift-install ]]
then
echo "Found existing openshift-install binary, checking version ..."
INSTALLED_VERSION=$(openshift-install version | sed -E -n '/openshift-install/s/^[^ ]+ v?//;s/\.[0-9]{1,2}$//p')
if [ "$INSTALLED_VERSION" = "$MATCH_ocp_version" ]; then
echo "✅ openshift-install version ${MATCH_ocp_version}.XX is present."
env -i "openshift_install_bin=$( pwd )/openshift-install" >> "$OUTPUT"
exit 0
else
echo "⚠️ openshift-install version $INSTALLED_VERSION is present, but version $MATCH_ocp_version is required. Deleting local binary and installing $MATCH_ocp_version"
rm openshift-install
fi
fi
rm -f openshift-install-linux.tar.gz
echo -n "Determining latest patch version for ${MATCH_ocp_version} ... "
# NOTE(sg): The OpenShift channel graph API returns a JSON object with the
# following shape:
#
# {
# nodes: [
# {
# version: "MAJOR.MINOR.PATCH",
# ...
# },
# ...
# ],
# edges: [
# [from_index, to_index],
# ...
# ],
# ...
# }
#
# where from_index and to_index for each edges entry are indices into the
# `nodes` list. Each entry of `edges` represents a possible upgrade path.
ocp_graph=$(curl -fsSL "https://api.openshift.com/api/upgrades_info/v1/graph?channel=stable-${MATCH_ocp_version}&arch=amd64")
# This jq query filters the nodes list by our target minor version, splits
# the version field (so it will be sorted numerically for `max`) and joins
# the result of max with dots to return the latest patch version that's
# available on the stable channel.
latest_version=$(jq -n -r --argjson g "$ocp_graph" --arg minor "${MATCH_ocp_version}" \
'[ $g.nodes[] | select(.version | test($minor)) | .version ]
| map(split(".") | map(tonumber))
| max | join(".")')
echo "${latest_version}"
echo -n "Determining latest patch version which can be upgraded to ${latest_version} ... "
# This jq query takes the latest patch version that's available on the
# stable channel and finds the index of that patch version's entry in the
# nodes list. Afterwards the query finds all edges which point to the
# latest patch version, extracts the version of edge's origin, and returns
# the highest patch version that has an upgrade edge to the latest patch
# version.
prev_version=$(jq -n -r --argjson g "$ocp_graph" --arg latest "$latest_version" \
'$g.nodes
| to_entries[]
| select(.value.version==$latest)
| .key as $curr_index
| [
$g.edges[]
| select(.[1]==$curr_index)
| $g.nodes[.[0]].version
| split(".") | map(tonumber)
] | max | join(".")')
echo "${prev_version}"
echo "Downloading openshift-install for version ${prev_version} ..."
curl -LO "https://mirror.openshift.com/pub/openshift-v4/clients/ocp/${prev_version}/openshift-install-linux.tar.gz"
tar -xf openshift-install-linux.tar.gz openshift-install
./openshift-install version
env -i "openshift_install_bin=$( pwd )/openshift-install" >> "$OUTPUT"
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And a lieutenant cluster
This step retrieves the Commodore tenant ID associated with the given lieutenant cluster ID.
Use api.syn.vshn.net as the Commodore API URL for production clusters. You might use the WebUI at control.vshn.net/syn/lieutenantapiendpoints to create and manage your clusters.
For customer clusters ensure the following facts are set:
-
sales_order: Name of the sales order to which the cluster is billed, such as S10000
-
service_level: Name of the service level agreement for this cluster, such as guaranteed-availability
-
access_policy: Access-Policy of the cluster, such as regular or swissonly
-
release_channel: Name of the syn component release channel to use, such as stable
-
maintenance_window: Pick the appropriate upgrade schedule, such as monday-1400 for test clusters, tuesday-1000 for prod or custom to not (yet) enable maintenance
-
cilium_addons: Comma-separated list of cilium addons the customer gets billed for, such as advanced_networking or tetragon. Set to NONE if no addons should be billed.
This step checks that you have access to the Commodore API and the cluster ID is valid.
Inputs
-
commodore_api_url: URL of the Commodore API to use for retrieving cluster information.
Use api.syn.vshn.net as the Commodore API URL for production clusters. Use api-int.syn.vshn.net for test clusters.
You might use the WebUI at control.vshn.net/syn/lieutenantapiendpoints to create and manage your clusters.
-
commodore_cluster_id: Project Syn cluster ID for the cluster to be set up.
In the form of c-example-infra-prod1.
You might use the WebUI at control.vshn.net/syn/lieutenantapiendpoints to create and manage your clusters.
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_api_url=
# export INPUT_commodore_cluster_id=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
echo "Retrieving Commodore tenant ID for cluster ID '$INPUT_commodore_cluster_id' from API at '$INPUT_commodore_api_url'..."
tenant_id=$(curl -sH "Authorization: Bearer $(commodore fetch-token)" ${COMMODORE_API_URL}/clusters/${INPUT_commodore_cluster_id} | jq -r .tenant)
if echo "$tenant_id" | grep 't-' >/dev/null 2>&1 ; then { echo "✅ Retrieved tenant ID '$tenant_id' for cluster ID '$INPUT_commodore_cluster_id'."; } else { echo "❌ Failed to retrieve valid tenant ID for cluster ID '$INPUT_commodore_cluster_id'. Got '$tenant_id'. Please check your Commodore API access and cluster ID."; exit 1; } ; fi
env -i "commodore_tenant_id=$tenant_id" >> "$OUTPUT"
region=$(curl -sH "Authorization: Bearer $(commodore fetch-token)" ${COMMODORE_API_URL}/clusters/${INPUT_commodore_cluster_id} | jq -r .facts.region)
if test -z "$region" || test "$region" == "null" ; then { echo "❌ Failed to retrieve CSP region for cluster ID '$INPUT_commodore_cluster_id'."; exit 1; } ; else { echo "✅ Retrieved CSP region '$region' for cluster ID '$INPUT_commodore_cluster_id'."; } ; fi
env -i "csp_region=$region" >> "$OUTPUT"
echo "Retrieving Vault address and login method..."
vault_addr=$(curl -s "${COMMODORE_API_URL}" | jq -r '.vault.addr')
if test -z "$vault_addr" || test "$vault_addr" == "null"; then echo "❌ Failed to retrieve Vault address from Lieutenant at '${COMMODORE_API_URL}'."; exit 1; else echo "✅ Retrieved Vault address '${vault_addr}' for Lieutenant at '${COMMODORE_API_URL}'."; fi
env -i "vault_address=${vault_addr}" >> "$OUTPUT"
vault_login_method=$(curl -s "${COMMODORE_API_URL}" | jq -r '.vault.loginMethod')
if test -z "$vault_login_method" || test "$vault_login_method" == "null"; then echo "❌ Failed to retrieve Vault login method from Lieutenant at '${COMMODORE_API_URL}'."; exit 1; else echo "✅ Retrieved Vault login method '${vault_login_method}' for Lieutenant at '${COMMODORE_API_URL}'."; fi
env -i "vault_login_method=${vault_login_method}" >> "$OUTPUT"
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And a Keycloak service
In this step, you have to create a Keycloak service for the new cluster via the VSHN Control Web UI at control.vshn.net/vshn/services/_create
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_cluster_id=
echo '#########################################################'
echo '# #'
echo "# Please create a Keycloak service with the cluster's #"
echo '# ID as Service Name via the VSHN Control Web UI. #'
echo '# #'
echo '#########################################################'
echo
echo "The name and ID of the service should be ${INPUT_commodore_cluster_id}."
echo "You can go to https://control.vshn.net/vshn/services/_create"
sleep 2
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And a cloudscale API token
Create a new cloudscale API token with read+write permissions and name it {{ .commodore_cluster_id }} on control.cloudscale.ch/service/<your-project>/api-token.
This step currently does not validate whether the token has write permission.
Inputs
-
cloudscale_token: Cloudscale API token with read+write permissions.
Used for setting up the cluster and for the machine api provider.
Outputs
-
cloudscale_token_floaty: Dummy output so other steps don’t unnecessarily ask for this as an input. Some steps also use this output to identify when we provision a cluster without Puppet LBs.
Script
OUTPUT=$(mktemp)
# export INPUT_cloudscale_token=
set -euo pipefail
if [[ $( curl -sH "Authorization: Bearer ${INPUT_cloudscale_token}" https://api.cloudscale.ch/v1/flavors -o /dev/null -w"%{http_code}" ) != 200 ]]
then
echo "Cloudscale token not valid!"
fi
env -i "cloudscale_token_floaty=none" >> "$OUTPUT"
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And a personal VSHN GitLab access token
This step ensures that you have provided a personal access token for VSHN GitLab.
Create the token at git.vshn.net/-/user_settings/personal_access_tokens with the "api" scope.
This step currently does not validate the token’s scope.
Inputs
-
gitlab_api_token: Personal access token for VSHN GitLab with the "api" scope.
Create the token at git.vshn.net/-/user_settings/personal_access_tokens with the "api" scope.
Script
OUTPUT=$(mktemp)
# export INPUT_gitlab_api_token=
set -euo pipefail
user="$( curl -sH "Authorization: Bearer ${INPUT_gitlab_api_token}" "https://git.vshn.net/api/v4/user" | jq -r .username )"
if [[ "$user" == "null" ]]
then
echo "Error validating GitLab token. Are you sure it is valid?"
exit 1
fi
env -i "gitlab_user_name=$user" >> "$OUTPUT"
echo "Token is valid."
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And basic cluster information
This step collects two essential pieces of information required for cluster setup: the base domain and the Red Hat pull secret.
See kb.vshn.ch/oc4/explanations/dns_scheme.html for more information about the base domain. Get a pull secret from cloud.redhat.com/openshift/install/pull-secret.
Inputs
-
base_domain: The base domain for the cluster without the cluster ID prefix and the last dot.
Example: appuio-beta.ch
See kb.vshn.ch/oc4/explanations/dns_scheme.html for more information about the base domain.
-
redhat_pull_secret: Red Hat pull secret for accessing Red Hat container images.
Get a pull secret from cloud.redhat.com/openshift/install/pull-secret.
Then I download the OpenShift image for version "4.21.0"
This step downloads the OpenShift image for the version specified by in the step.
If the image already exists locally, it skips the download.
Script
OUTPUT=$(mktemp)
set -euo pipefail
. "$GANDALF_SPELLBOOK_DIR"/scripts/semver.sh
MAJOR=0
MINOR=0
PATCH=0
SPECIAL=""
semverParseInto "$MATCH_image_name" MAJOR MINOR PATCH SPECIAL
image_path="rhcos-$MAJOR.$MINOR.qcow2"
env -i "image_major=$MAJOR" >> "$OUTPUT"
env -i "image_minor=$MINOR" >> "$OUTPUT"
env -i "image_patch=$PATCH" >> "$OUTPUT"
echo "Image is $image_path"
if [ -f "$image_path" ]; then
echo "Image $image_path already exists, skipping download."
env -i "image_path=$image_path" >> "$OUTPUT"
exit 0
fi
echo Downloading OpenShift image "$MATCH_image_name" to "$image_path"
curl -L "https://mirror.openshift.com/pub/openshift-v4/dependencies/rhcos/${MAJOR}.${MINOR}/${MATCH_image_name}/rhcos-${MATCH_image_name}-x86_64-openstack.x86_64.qcow2.gz" | gzip -d > "$image_path"
env -i "image_path=$image_path" >> "$OUTPUT"
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I set up required S3 buckets
This step sets up the required S3 buckets for the OpenShift cluster installation.
It uses the MinIO Client (mc) to create the necessary buckets if they do not already exist.
Script
OUTPUT=$(mktemp)
# export INPUT_cloudscale_token=
# export INPUT_commodore_cluster_id=
# export INPUT_csp_region=
set -euo pipefail
response=$(curl -sH "Authorization: Bearer ${INPUT_cloudscale_token}" \
https://api.cloudscale.ch/v1/objects-users | \
jq -e ".[] | select(.display_name == \"${INPUT_commodore_cluster_id}\")" ||:)
if [ -z "$response" ]; then
echo "Creating Cloudscale S3 user for cluster ID '${INPUT_commodore_cluster_id}'..."
response=$(curl -sH "Authorization: Bearer ${INPUT_cloudscale_token}" \
-F display_name=${INPUT_commodore_cluster_id} \
https://api.cloudscale.ch/v1/objects-users)
echo "Created user with id $(echo "$response" | jq -r .id)"
else
echo "Cloudscale S3 user for cluster ID '${INPUT_commodore_cluster_id}' already exists. id: $(echo "$response" | jq -r .id)"
fi
echo -n "Waiting for S3 credentials to become available ..."
until mc alias set \
"${INPUT_commodore_cluster_id}" "https://objects.${INPUT_csp_region}.cloudscale.ch" \
"$(echo "$response" | jq -r '.keys[0].access_key')" \
"$(echo "$response" | jq -r '.keys[0].secret_key')"
do
echo -n .
sleep 5
done
echo "OK"
mc mb --ignore-existing \
"${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-bootstrap-ignition"
mc mb --ignore-existing \
"${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-image-registry"
mc mb --ignore-existing \
"${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-logstore"
keyid=$(mc alias list ${INPUT_commodore_cluster_id} -json | jq -r .accessKey)
export AWS_ACCESS_KEY_ID="${keyid}"
secretkey=$(mc alias list ${INPUT_commodore_cluster_id} -json | jq -r .secretKey)
export AWS_SECRET_ACCESS_KEY="${secretkey}"
echo "Configuring S3 bucket policies..."
aws s3api put-public-access-block \
--endpoint-url "https://objects.${INPUT_csp_region}.cloudscale.ch" \
--bucket "${INPUT_commodore_cluster_id}-image-registry" \
--public-access-block-configuration BlockPublicAcls=false
aws s3api put-bucket-lifecycle-configuration \
--endpoint-url "https://objects.${INPUT_csp_region}.cloudscale.ch" \
--bucket "${INPUT_commodore_cluster_id}-image-registry" \
--lifecycle-configuration '{
"Rules": [
{
"ID": "cleanup-incomplete-multipart-registry-uploads",
"Prefix": "",
"Status": "Enabled",
"AbortIncompleteMultipartUpload": {
"DaysAfterInitiation": 1
}
}
]
}'
echo "S3 buckets are set up."
env -i "bucket_user=$(echo "$response" | jq -c .)" >> "$OUTPUT"
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I import the image in Cloudscale
This step uploads the Red Hat CoreOS image to the S3 bucket for the image registry.
It then imports the image into Cloudscale as a custom image.
It uses the MinIO Client (mc) to perform the upload.
Inputs
-
image_path -
commodore_cluster_id -
csp_region -
bucket_user -
image_major -
image_minor -
cloudscale_token
Script
OUTPUT=$(mktemp)
# export INPUT_image_path=
# export INPUT_commodore_cluster_id=
# export INPUT_csp_region=
# export INPUT_bucket_user=
# export INPUT_image_major=
# export INPUT_image_minor=
# export INPUT_cloudscale_token=
set -euo pipefail
auth_header="Authorization: Bearer ${INPUT_cloudscale_token}"
image_slug="rhcos-${INPUT_image_major}.${INPUT_image_minor}"
target_zones=$(jq -n -r --arg zone "${INPUT_csp_region}1" '[$zone]')
# NOTE(sg): The value of zones is a space-separated list of zones where
# the image with the given slug exists or the empty string if the image
# doesn't exist in any zone.
images_resp=$(curl -fsSL -H "$auth_header" https://api.cloudscale.ch/v1/custom-images)
zones=$(jq -n -r --argjson resp "$images_resp" --arg slug "$image_slug" \
'$resp[] | select(.slug == $slug) | [.zones[]|.slug]|join(" ")')
old_image_href=
if [ -n "$zones" ] && [[ "$zones" == *"${INPUT_csp_region}1"* ]]; then
echo "✅ Image '$image_slug' already exists in cloudscale zone ${INPUT_csp_region}, skipping upload."
exit 0
elif [ -n "$zones" ]; then
echo "Image '$image_slug' already exists in cloudscale zone(s) '${zones}', configuring multi-zone image upload"
old_image_href=$(jq -n -r --argjson resp "$images_resp" --arg slug "$image_slug" \
'$resp[] |select(.slug == $slug) | .href')
target_zones=$(jq -n -r --argjson zones "$target_zones" \
--arg zone "$INPUT_csp_region}1" '$zones + [$zone]')
fi
mc alias set \
"${INPUT_commodore_cluster_id}" "https://objects.${INPUT_csp_region}.cloudscale.ch" \
"$(echo "$INPUT_bucket_user" | jq -r '.keys[0].access_key')" \
"$(echo "$INPUT_bucket_user" | jq -r '.keys[0].secret_key')"
echo "Uploading Red Hat CoreOS image '$INPUT_image_path' to S3 bucket '${INPUT_commodore_cluster_id}-image-registry'..."
mc cp "rhcos-${INPUT_image_major}.${INPUT_image_minor}.qcow2" "${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-bootstrap-ignition/"
echo "Upload completed."
mc anonymous set download "${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-bootstrap-ignition/rhcos-${INPUT_image_major}.${INPUT_image_minor}.qcow2"
echo "Importing image into Cloudscale..."
import_req=$(jq -n -r \
--argjson zones "$target_zones" \
--argjson image_download "$(mc share download --json "${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-bootstrap-ignition/rhcos-${INPUT_image_major}.${INPUT_image_minor}.qcow2")" \
--arg name "RHCOS ${INPUT_image_major}.${INPUT_image_minor}" \
--arg image_slug "$image_slug" \
'{
"slug": $image_slug,
"url": $image_download.url,
"name": $name,
"zones": $zones,
"source_format": "raw",
"user_data_handling": "pass-through"
}')
import_resp=$(curl -fsSL -H "$auth_header" \
https://api.cloudscale.ch/v1/custom-images/import \
--json "$import_req")
echo "Waiting for import to complete..."
import_id=$(jq -n -r --argjson import "$import_resp" '$import.uuid')
import_done=false
status="initializing"
step_exit=0
while ! $import_done; do
import_status_resp=$(curl -fsSL -H "$auth_header" \
"https://api.cloudscale.ch/v1/custom-images/import/${import_id}")
status=$(jq -n -r --argjson status "$import_status_resp" '$status.status')
case "$status" in
"started"|"in_progress")
echo -n "."
import_done=false
import_error=
;;
"failed")
import_done=true
import_error="❌ $(jq -n -r --argjson status "$import_status_resp" '$status.error_message')"
step_exit=1
break
;;
"success")
import_done=true
import_error=
break
;;
*)
import_done=true
import_error="⚠️ Unknown import status '${status}', canceling wait for completion, please verify the import status manually!"
step_exit=0
break
;;
esac
sleep 5
done
echo ""
echo "Import completed with status ${status}"
if [ -n "$import_error" ]; then
echo -e "\n$import_error"
exit "$step_exit"
else
echo -e "\n✅ Successfully imported image"
fi
sleep 0.1
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
Then I set secrets in Vault
This step stores the collected secrets and tokens in the ProjectSyn Vault.
Inputs
-
vault_address: Address of the Vault server associated with the Lieutenant API to store cluster secrets. -
vault_login_method -
commodore_cluster_id -
commodore_tenant_id -
bucket_user -
cloudscale_token -
cloudscale_token_floaty
Script
OUTPUT=$(mktemp)
# export INPUT_vault_address=
# export INPUT_vault_login_method=
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_tenant_id=
# export INPUT_bucket_user=
# export INPUT_cloudscale_token=
# export INPUT_cloudscale_token_floaty=
set -euo pipefail
export VAULT_ADDR=${INPUT_vault_address}
vault login -method=${INPUT_vault_login_method}
# Set the cloudscale.ch access secrets
vault kv put clusters/kv/${INPUT_commodore_tenant_id}/${INPUT_commodore_cluster_id}/cloudscale \
token=${INPUT_cloudscale_token} \
s3_access_key="$(echo "${INPUT_bucket_user}" | jq -r '.keys[0].access_key')" \
s3_secret_key="$(echo "${INPUT_bucket_user}" | jq -r '.keys[0].secret_key')"
# Put LB API key in Vault
if [ "${INPUT_cloudscale_token_floaty}" != "none" ]; then
vault kv put clusters/kv/${INPUT_commodore_tenant_id}/${INPUT_commodore_cluster_id}/floaty \
iam_secret="${INPUT_cloudscale_token_floaty}"
fi
# Generate an HTTP secret for the registry
vault kv put clusters/kv/${INPUT_commodore_tenant_id}/${INPUT_commodore_cluster_id}/registry \
httpSecret="$(LC_ALL=C tr -cd "A-Za-z0-9" </dev/urandom | head -c 128)"
# Generate a master password for K8up backups
vault kv put clusters/kv/${INPUT_commodore_tenant_id}/${INPUT_commodore_cluster_id}/global-backup \
password="$(LC_ALL=C tr -cd "A-Za-z0-9" </dev/urandom | head -c 32)"
# Generate a password for the cluster object backups
vault kv put clusters/kv/${INPUT_commodore_tenant_id}/${INPUT_commodore_cluster_id}/cluster-backup \
password="$(LC_ALL=C tr -cd "A-Za-z0-9" </dev/urandom | head -c 32)"
hieradata_repo_secret=$(vault kv get \
-format=json "clusters/kv/lbaas/hieradata_repo_token" | jq '.data.data')
env -i "hieradata_repo_user=$(echo "${hieradata_repo_secret}" | jq -r '.user')" >> "$OUTPUT"
env -i "hieradata_repo_token=$(echo "${hieradata_repo_secret}" | jq -r '.token')" >> "$OUTPUT"
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I check the cluster domain
Please verify that the base domain generated is correct for your setup.
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_cluster_id=
# export INPUT_base_domain=
set -euo pipefail
cluster_domain="${INPUT_commodore_cluster_id}.${INPUT_base_domain}"
echo "Cluster domain is set to '$cluster_domain'"
echo "cluster_domain=$cluster_domain" >> "$OUTPUT"
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I prepare the cluster repository
This step prepares the local cluster repository by cloning the Commodore hieradata repository and setting up the necessary configuration for the specified cluster.
Inputs
-
commodore_api_url -
commodore_cluster_id -
commodore_tenant_id -
cluster_domain -
image_major -
image_minor
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_api_url=
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_tenant_id=
# export INPUT_cluster_domain=
# export INPUT_image_major=
# export INPUT_image_minor=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
rm -rf inventory/classes/
mkdir -p inventory/classes/
git clone "$(curl -sH"Authorization: Bearer $(commodore fetch-token)" "${INPUT_commodore_api_url}/tenants/${INPUT_commodore_tenant_id}" | jq -r '.gitRepo.url')" inventory/classes/${INPUT_commodore_tenant_id}
pushd "inventory/classes/${INPUT_commodore_tenant_id}/"
yq eval -i ".parameters.openshift.baseDomain = \"${INPUT_cluster_domain}\"" \
${INPUT_commodore_cluster_id}.yml
git diff --exit-code --quiet || git commit -a -m "Configure cluster domain for ${INPUT_commodore_cluster_id}"
if ls openshift4.y*ml 1>/dev/null 2>&1; then
yq eval -i '.classes += ".openshift4"' ${INPUT_commodore_cluster_id}.yml;
git diff --exit-code --quiet || git commit -a -m "Include openshift4 class for ${INPUT_commodore_cluster_id}"
fi
yq eval -i '.parameters.openshift.cloudscale.subnet_uuid = "TO_BE_DEFINED"' ${INPUT_commodore_cluster_id}.yml
yq eval -i '.parameters.openshift.cloudscale.rhcos_image_slug = "rhcos-'"${INPUT_image_major}.${INPUT_image_minor}"'"' \
${INPUT_commodore_cluster_id}.yml
yq eval -i ".parameters.openshift4_terraform.terraform_variables.ignition_ca = \"TO_BE_DEFINED\"" \
${INPUT_commodore_cluster_id}.yml
git diff --exit-code --quiet || git commit -a -m "Configure Cloudscale metaparameters on ${INPUT_commodore_cluster_id}"
yq eval -i '.applications += ["cloudscale-loadbalancer-controller"]' ${INPUT_commodore_cluster_id}.yml
yq eval -i '.applications = (.applications | unique)' ${INPUT_commodore_cluster_id}.yml
cat ${INPUT_commodore_cluster_id}.yml
git diff --exit-code --quiet || git commit -a -m "Enable cloudscale loadbalancer controller for ${INPUT_commodore_cluster_id}"
yq eval -i '.applications += ["cilium"]' ${INPUT_commodore_cluster_id}.yml
yq eval -i '.applications = (.applications | unique)' ${INPUT_commodore_cluster_id}.yml
yq eval -i '.parameters.networkpolicy.networkPlugin = "cilium"' ${INPUT_commodore_cluster_id}.yml
yq eval -i '.parameters.openshift.infraID = "TO_BE_DEFINED"' ${INPUT_commodore_cluster_id}.yml
yq eval -i '.parameters.openshift.clusterID = "TO_BE_DEFINED"' ${INPUT_commodore_cluster_id}.yml
yq eval -i '.parameters.cilium.olm.generate_olm_deployment = true' ${INPUT_commodore_cluster_id}.yml
git diff --exit-code --quiet || git commit -a -m "Add Cilium addon to ${INPUT_commodore_cluster_id}"
git push
popd
commodore catalog compile ${INPUT_commodore_cluster_id} --push \
--dynamic-fact kubernetesVersion.major=1 \
--dynamic-fact kubernetesVersion.minor="$((INPUT_image_minor+13))" \
--dynamic-fact openshiftVersion.Major=${INPUT_image_major} \
--dynamic-fact openshiftVersion.Minor=${INPUT_image_minor}
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
Then I configure the OpenShift installer
This step configures the OpenShift installer for the Cloudscale cluster by generating the necessary installation files using Commodore.
Inputs
-
commodore_cluster_id -
commodore_tenant_id -
base_domain -
cluster_domain -
vault_address -
vault_login_method -
redhat_pull_secret -
csp_region -
bucket_user -
cloudscale_token -
openshift_install_bin
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_tenant_id=
# export INPUT_base_domain=
# export INPUT_cluster_domain=
# export INPUT_vault_address=
# export INPUT_vault_login_method=
# export INPUT_redhat_pull_secret=
# export INPUT_csp_region=
# export INPUT_bucket_user=
# export INPUT_cloudscale_token=
# export INPUT_openshift_install_bin=
set -euo pipefail
openshift-install() {
"${INPUT_openshift_install_bin}" "${@}"
}
export VAULT_ADDR="${INPUT_vault_address}"
vault login -method="${INPUT_vault_login_method}"
ssh_private_key="$(pwd)/ssh_${INPUT_commodore_cluster_id}"
ssh_public_key="${ssh_private_key}.pub"
env -i "ssh_public_key_path=$ssh_public_key" >> "$OUTPUT"
if vault kv get -format=json clusters/kv/${INPUT_commodore_tenant_id}/${INPUT_commodore_cluster_id}/cloudscale/ssh >/dev/null 2>&1; then
echo "SSH keypair for cluster ${INPUT_commodore_cluster_id} already exists in Vault, skipping generation."
vault kv get -format=json clusters/kv/${INPUT_commodore_tenant_id}/${INPUT_commodore_cluster_id}/cloudscale/ssh | \
jq -r '.data.data.private_key|@base64d' > "${ssh_private_key}"
chmod 600 "${ssh_private_key}"
ssh-keygen -f "${ssh_private_key}" -y > "${ssh_public_key}"
else
echo "Generating new SSH keypair for cluster ${INPUT_commodore_cluster_id}."
ssh-keygen -C "vault@${INPUT_commodore_cluster_id}" -t ed25519 -f "$ssh_private_key" -N ''
base64_no_wrap='base64'
if [[ "$OSTYPE" == "linux"* ]]; then
base64_no_wrap='base64 --wrap 0'
fi
vault kv put clusters/kv/${INPUT_commodore_tenant_id}/${INPUT_commodore_cluster_id}/cloudscale/ssh \
private_key="$(cat "$ssh_private_key" | eval "$base64_no_wrap")"
fi
echo Adding SSH private key to ssh-agent...
echo You might need to start the ssh-agent first using: eval "\$(ssh-agent)"
echo ssh-add "$ssh_private_key"
ssh-add "$ssh_private_key"
installer_dir="$(pwd)/target"
rm -rf "${installer_dir}"
mkdir -p "${installer_dir}"
cat > "${installer_dir}/install-config.yaml" <<EOF
apiVersion: v1
metadata:
name: ${INPUT_commodore_cluster_id}
baseDomain: ${INPUT_base_domain}
platform:
external:
platformName: cloudscale
cloudControllerManager: External
networking:
networkType: Cilium
pullSecret: |
${INPUT_redhat_pull_secret}
sshKey: "$(cat "$ssh_public_key")"
EOF
echo Running OpenShift installer to create manifests...
openshift-install --dir "${installer_dir}" create manifests
echo Copying machineconfigs...
machineconfigs=catalog/manifests/openshift4-nodes/10_machineconfigs.yaml
if [ -f $machineconfigs ]; then
yq --no-doc -s \
"\"${installer_dir}/openshift/99x_openshift-machineconfig_\" + .metadata.name" \
$machineconfigs
fi
echo Copying Cloudscale CCM manifests...
for f in catalog/manifests/cloudscale-cloud-controller-manager/*; do
cp "$f" "${installer_dir}/manifests/cloudscale_ccm_$(basename "$f")"
done
yq -i e ".stringData.access-token=\"${INPUT_cloudscale_token}\"" \
"${installer_dir}/manifests/cloudscale_ccm_01_secret.yaml"
echo Copying Cilium OLM manifests...
for f in catalog/manifests/cilium/olm/[a-z]*; do
cp "$f" "${installer_dir}/manifests/cilium_$(basename "$f")"
done
# shellcheck disable=2016
# We don't want the shell to execute network.operator.openshift.io as a
# command, so we need single quotes here.
echo 'Generating initial `network.operator.openshift.io` resource...'
yq '{
"apiVersion": "operator.openshift.io/v1",
"kind": "Network",
"metadata": {
"name": "cluster"
},
"spec": {
"deployKubeProxy": false,
"clusterNetwork": .spec.clusterNetwork,
"externalIP": {
"policy": {}
},
"networkType": "Cilium",
"serviceNetwork": .spec.serviceNetwork
}}' "${installer_dir}/manifests/cluster-network-02-config.yml" \
> "${installer_dir}/manifests/cilium_cluster-network-operator.yaml"
gen_cluster_domain=$(yq e '.spec.baseDomain' \
"${installer_dir}/manifests/cluster-dns-02-config.yml")
if [ "$gen_cluster_domain" != "$INPUT_cluster_domain" ]; then
echo -e "\033[0;31mGenerated cluster domain doesn't match expected cluster domain: Got '$gen_cluster_domain', want '$INPUT_cluster_domain'\033[0;0m"
exit 1
else
echo -e "\033[0;32mGenerated cluster domain matches expected cluster domain.\033[0;0m"
fi
echo Running OpenShift installer to create ignition configs...
openshift-install --dir "${installer_dir}" \
create ignition-configs
mc alias set \
"${INPUT_commodore_cluster_id}" "https://objects.${INPUT_csp_region}.cloudscale.ch" \
"$(echo "$INPUT_bucket_user" | jq -r '.keys[0].access_key')" \
"$(echo "$INPUT_bucket_user" | jq -r '.keys[0].secret_key')"
mc cp "${installer_dir}/bootstrap.ign" "${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-bootstrap-ignition/"
ignition_bootstrap=$(mc share download \
--json --expire=24h \
"${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-bootstrap-ignition/bootstrap.ign" | jq -r '.share')
env -i "ignition_bootstrap=$ignition_bootstrap" >> "$OUTPUT"
echo "✅ OpenShift installer configured successfully."
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I configure Terraform for team "aldebaran"
This step configures Terraform the Commodore rendered terraform configuration.
Inputs
-
commodore_api_url -
commodore_cluster_id -
commodore_tenant_id -
ssh_public_key_path -
hieradata_repo_user -
base_domain -
image_major -
image_minor
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_api_url=
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_tenant_id=
# export INPUT_ssh_public_key_path=
# export INPUT_hieradata_repo_user=
# export INPUT_base_domain=
# export INPUT_image_major=
# export INPUT_image_minor=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
installer_dir="$(pwd)/target"
pushd "inventory/classes/${INPUT_commodore_tenant_id}/"
yq eval -i '.classes += ["global.distribution.openshift4.no-opsgenie"]' ${INPUT_commodore_cluster_id}.yml;
yq eval -i '.classes = (.classes | unique)' ${INPUT_commodore_cluster_id}.yml
yq eval -i ".parameters.openshift.infraID = \"$(jq -r .infraID "${installer_dir}/metadata.json")\"" \
${INPUT_commodore_cluster_id}.yml
yq eval -i ".parameters.openshift.clusterID = \"$(jq -r .clusterID "${installer_dir}/metadata.json")\"" \
${INPUT_commodore_cluster_id}.yml
yq eval -i 'del(.parameters.cilium.olm.generate_olm_deployment)' \
${INPUT_commodore_cluster_id}.yml
yq eval -i ".parameters.openshift.ssh_key = \"$(cat ${INPUT_ssh_public_key_path})\"" \
${INPUT_commodore_cluster_id}.yml
ca_cert=$(jq -r '.ignition.security.tls.certificateAuthorities[0].source' \
"${installer_dir}/master.ign" | \
awk -F ',' '{ print $2 }' | \
base64 --decode)
yq eval -i ".parameters.openshift4_terraform.terraform_variables.base_domain = \"${INPUT_base_domain}\"" \
${INPUT_commodore_cluster_id}.yml
yq eval -i ".parameters.openshift4_terraform.terraform_variables.ignition_ca = \"${ca_cert}\"" \
${INPUT_commodore_cluster_id}.yml
yq eval -i ".parameters.openshift4_terraform.terraform_variables.team = \"${MATCH_team_name}\"" \
${INPUT_commodore_cluster_id}.yml
yq eval -i ".parameters.openshift4_terraform.terraform_variables.hieradata_repo_user = \"${INPUT_hieradata_repo_user}\"" \
${INPUT_commodore_cluster_id}.yml
git commit -a -m "Setup cluster ${INPUT_commodore_cluster_id}"
git push
popd
commodore catalog compile ${INPUT_commodore_cluster_id} --push \
--dynamic-fact kubernetesVersion.major=1 \
--dynamic-fact kubernetesVersion.minor="$((INPUT_image_minor+13))" \
--dynamic-fact openshiftVersion.Major=${INPUT_image_major} \
--dynamic-fact openshiftVersion.Minor=${INPUT_image_minor}
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I configure Terraform for Cloudscale
Review the file ./inventory/classes/{{ .commodore_tenant_id }}/{{ .commodore_cluster_id }}.yml. Override default parameters or add more component configurations as required for your cluster.
Then press enter to commit and compile.
Inputs
-
commodore_api_url -
commodore_cluster_id -
commodore_tenant_id -
image_major -
image_minor -
ssh_public_key_path -
cloudscale_token_floaty
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_api_url=
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_tenant_id=
# export INPUT_image_major=
# export INPUT_image_minor=
# export INPUT_ssh_public_key_path=
# export INPUT_cloudscale_token_floaty=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
export CLUSTER_ID="${INPUT_commodore_cluster_id}"
export TENANT_ID="${INPUT_commodore_tenant_id}"
pushd "inventory/classes/${TENANT_ID}/"
yq eval -i ".parameters.openshift4_terraform.terraform_variables.allocate_router_vip_for_lb_controller = true" \
${INPUT_commodore_cluster_id}.yml
if [ "$INPUT_cloudscale_token_floaty" == "none" ]; then
echo "Configuring cloudscale cluster without Puppet LBs"
yq eval -i ".parameters.openshift4_terraform.terraform_variables.enable_api_lbaas = true" \
${INPUT_commodore_cluster_id}.yml
yq eval -i ".parameters.openshift4_terraform.terraform_variables.enable_cloudscale_router = true" \
${INPUT_commodore_cluster_id}.yml
fi
yq eval -i ".parameters.openshift4_terraform.terraform_variables.ssh_keys = [\"$(cat ${INPUT_ssh_public_key_path})\"]" \
${INPUT_commodore_cluster_id}.yml
echo "Committing changes ..."
if git commit -a -m "Set cloudscale specific values in cluster ${CLUSTER_ID}"; then
git push
else
echo "No changes."
fi
popd
echo "Compile and push cluster catalog"
commodore catalog compile ${CLUSTER_ID} --push \
--dynamic-fact kubernetesVersion.major=1 \
--dynamic-fact kubernetesVersion.minor="$((INPUT_image_minor+13))" \
--dynamic-fact openshiftVersion.Major=${INPUT_image_major} \
--dynamic-fact openshiftVersion.Minor=${INPUT_image_minor}
echo "✅ Changes committed and cluster catalog compiled successfully."
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
Then I provision the cloudscale LBs and router
This step provisions the cloudscale network infrastructure for the OpenShift cluster using Terraform.
Inputs
-
cloudscale_token -
ignition_bootstrap -
gitlab_user_name -
gitlab_api_token -
commodore_cluster_id -
commodore_api_url -
cluster_domain
Outputs
-
lb_fqdn_1: Dummy output so we can reuse later steps which need to interact with the Puppet LBs only if we provision them -
lb_fqdn_2: Dummy output so we can reuse later steps which need to interact with the Puppet LBs only if we provision them -
control_vshn_api_token: Dummy output so we can reuse later steps which need to interact with control.vshn.net only if we provision Puppet LBs
Script
OUTPUT=$(mktemp)
# export INPUT_cloudscale_token=
# export INPUT_ignition_bootstrap=
# export INPUT_gitlab_user_name=
# export INPUT_gitlab_api_token=
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_api_url=
# export INPUT_cluster_domain=
set -euo pipefail
# set dummy outputs
env -i "lb_fqdn_1=none" >> "$OUTPUT"
env -i "lb_fqdn_2=none" >> "$OUTPUT"
env -i "control_vshn_api_token=none" >> "$OUTPUT"
# provision cloudscale infra
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
cat <<EOF > ./terraform.env
CLOUDSCALE_API_TOKEN=${INPUT_cloudscale_token}
TF_VAR_ignition_bootstrap=${INPUT_ignition_bootstrap}
TF_VAR_lb_cloudscale_api_secret=none
TF_VAR_control_vshn_net_token=none
GIT_AUTHOR_NAME=$(git config --global user.name)
GIT_AUTHOR_EMAIL=$(git config --global user.email)
HIERADATA_REPO_TOKEN=none
EOF
tf_image=$(\
yq eval ".parameters.openshift4_terraform.images.terraform.image" \
dependencies/openshift4-terraform/class/defaults.yml)
tf_tag=$(\
yq eval ".parameters.openshift4_terraform.images.terraform.tag" \
dependencies/openshift4-terraform/class/defaults.yml)
echo "Using Terraform image: ${tf_image}:${tf_tag}"
base_dir=$(pwd)
terraform() {
touch .terraformrc
docker run --rm -e REAL_UID="$(id -u)" -e TF_CLI_CONFIG_FILE=/tf/.terraformrc --env-file "${base_dir}/terraform.env" -w /tf -v "$(pwd):/tf" --ulimit memlock=-1 "${tf_image}:${tf_tag}" /tf/terraform.sh "${@}"
}
gitlab_repository_url=$(curl -sH "Authorization: Bearer $(commodore fetch-token)" ${INPUT_commodore_api_url}/clusters/${INPUT_commodore_cluster_id} | jq -r '.gitRepo.url' | sed 's|ssh://||; s|/|:|')
gitlab_repository_name=${gitlab_repository_url##*/}
gitlab_catalog_project_id=$(curl -sH "Authorization: Bearer ${INPUT_gitlab_api_token}" "https://git.vshn.net/api/v4/projects?simple=true&search=${gitlab_repository_name/.git}" | jq -r ".[] | select(.ssh_url_to_repo == \"${gitlab_repository_url}\") | .id")
gitlab_state_url="https://git.vshn.net/api/v4/projects/${gitlab_catalog_project_id}/terraform/state/cluster"
pushd catalog/manifests/openshift4-terraform/
terraform init \
"-backend-config=address=${gitlab_state_url}" \
"-backend-config=lock_address=${gitlab_state_url}/lock" \
"-backend-config=unlock_address=${gitlab_state_url}/lock" \
"-backend-config=username=${INPUT_gitlab_user_name}" \
"-backend-config=password=${INPUT_gitlab_api_token}" \
"-backend-config=lock_method=POST" \
"-backend-config=unlock_method=DELETE" \
"-backend-config=retry_wait_min=5"
cat > override.tf <<EOF
module "cluster" {
bootstrap_count = 0
master_count = 0
}
EOF
terraform apply -auto-approve
terraform output -raw cluster_dns > ../../../dns.txt
echo "@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@"
echo "@ @"
echo "@ Please add the DNS records shown in the Terraform output to your DNS provider. @"
echo "@ Most probably in https://git.vshn.net/vshn/vshn_zonefiles @"
echo "@ @"
echo "@ If terminal selection does not work the entries can also be copied from @"
echo "@ dns.txt @"
echo "@ @"
echo "@ Waiting for record to propagate... @"
echo "@ @"
echo "@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@"
# NOTE(sg): We first query the SOA record for the cluster zone name. This
# shouldn't significantly pollute our cache, since that record should
# exist and is unlikely to change. Afterwards, we directly query the
# authoritative nameserver for the api record, which doesn't pollute any
# recursive DNS servers' caches.
auth_ns=$(dig +noall +authority SOA "${INPUT_cluster_domain}" | tr "\t" " " | tr -s " " | cut -d " " -f5)
while [ "$(dig +short A "@${auth_ns}" "api.${INPUT_cluster_domain}")" == "" ]
do
echo -n "."
sleep 15
done
echo "✅ API record present"
rm ../../../dns.txt
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I provision the bootstrap node and control plane
This step provisions the bootstrap node and control plane for the cloudscale OpenShift cluster using Terraform.
Inputs
-
cloudscale_token -
ignition_bootstrap -
gitlab_user_name -
gitlab_api_token -
commodore_cluster_id -
commodore_api_url
Script
OUTPUT=$(mktemp)
# export INPUT_cloudscale_token=
# export INPUT_ignition_bootstrap=
# export INPUT_gitlab_user_name=
# export INPUT_gitlab_api_token=
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_api_url=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
installer_dir="$(pwd)/target"
cat <<EOF > ./terraform.env
CLOUDSCALE_API_TOKEN=${INPUT_cloudscale_token}
TF_VAR_ignition_bootstrap=${INPUT_ignition_bootstrap}
TF_VAR_lb_cloudscale_api_secret=none
TF_VAR_control_vshn_net_token=none
GIT_AUTHOR_NAME=$(git config --global user.name)
GIT_AUTHOR_EMAIL=$(git config --global user.email)
HIERADATA_REPO_TOKEN=none
EOF
tf_image=$(\
yq eval ".parameters.openshift4_terraform.images.terraform.image" \
dependencies/openshift4-terraform/class/defaults.yml)
tf_tag=$(\
yq eval ".parameters.openshift4_terraform.images.terraform.tag" \
dependencies/openshift4-terraform/class/defaults.yml)
echo "Using Terraform image: ${tf_image}:${tf_tag}"
base_dir=$(pwd)
terraform() {
touch .terraformrc
docker run --rm -e REAL_UID="$(id -u)" -e TF_CLI_CONFIG_FILE=/tf/.terraformrc --env-file "${base_dir}/terraform.env" -w /tf -v "$(pwd):/tf" --ulimit memlock=-1 "${tf_image}:${tf_tag}" /tf/terraform.sh "${@}"
}
gitlab_repository_url=$(curl -sH "Authorization: Bearer $(commodore fetch-token)" ${INPUT_commodore_api_url}/clusters/${INPUT_commodore_cluster_id} | jq -r '.gitRepo.url' | sed 's|ssh://||; s|/|:|')
gitlab_repository_name=${gitlab_repository_url##*/}
gitlab_catalog_project_id=$(curl -sH "Authorization: Bearer ${INPUT_gitlab_api_token}" "https://git.vshn.net/api/v4/projects?simple=true&search=${gitlab_repository_name/.git}" | jq -r ".[] | select(.ssh_url_to_repo == \"${gitlab_repository_url}\") | .id")
gitlab_state_url="https://git.vshn.net/api/v4/projects/${gitlab_catalog_project_id}/terraform/state/cluster"
pushd catalog/manifests/openshift4-terraform/
terraform init \
"-backend-config=address=${gitlab_state_url}" \
"-backend-config=lock_address=${gitlab_state_url}/lock" \
"-backend-config=unlock_address=${gitlab_state_url}/lock" \
"-backend-config=username=${INPUT_gitlab_user_name}" \
"-backend-config=password=${INPUT_gitlab_api_token}" \
"-backend-config=lock_method=POST" \
"-backend-config=unlock_method=DELETE" \
"-backend-config=retry_wait_min=5"
cat > override.tf <<EOF
module "cluster" {
bootstrap_count = 1
}
EOF
terraform apply -auto-approve
popd
echo -n "Waiting for Kubernetes API to become available .."
API_URL=$(yq e '.clusters[0].cluster.server' "${installer_dir}/auth/kubeconfig")
while ! curl --connect-timeout 1 "${API_URL}/healthz" -k &>/dev/null; do
echo -n "."
sleep 5
done && echo " ✅ API is up"
export KUBECONFIG="${installer_dir}/auth/kubeconfig"
echo "Waiting for masters to be created ..."
kubectl wait --for create --timeout=600s node -l node-role.kubernetes.io/master
echo "Waiting for masters to become ready ..."
kubectl wait --for condition=ready --timeout=600s node -l node-role.kubernetes.io/master
env -i "kubeconfig_path=${installer_dir}/auth/kubeconfig" >> "$OUTPUT"
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I store the subnet ID and floating IP in the Syn hierarchy
This step retrieves the subnet ID and ingress floating IP from Terraform and stores them in the Syn hierarchy.
Inputs
-
cloudscale_token -
cloudscale_token_floaty -
control_vshn_api_token -
ignition_bootstrap -
gitlab_user_name -
gitlab_api_token -
commodore_cluster_id -
commodore_tenant_id -
commodore_api_url -
image_major -
image_minor
Script
OUTPUT=$(mktemp)
# export INPUT_cloudscale_token=
# export INPUT_cloudscale_token_floaty=
# export INPUT_control_vshn_api_token=
# export INPUT_ignition_bootstrap=
# export INPUT_gitlab_user_name=
# export INPUT_gitlab_api_token=
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_tenant_id=
# export INPUT_commodore_api_url=
# export INPUT_image_major=
# export INPUT_image_minor=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
cat <<EOF > ./terraform.env
CLOUDSCALE_API_TOKEN=${INPUT_cloudscale_token}
TF_VAR_ignition_bootstrap=${INPUT_ignition_bootstrap}
TF_VAR_lb_cloudscale_api_secret=${INPUT_cloudscale_token_floaty}
TF_VAR_control_vshn_net_token=${INPUT_control_vshn_api_token}
GIT_AUTHOR_NAME=$(git config --global user.name)
GIT_AUTHOR_EMAIL=$(git config --global user.email)
HIERADATA_REPO_TOKEN=${INPUT_gitlab_api_token}
EOF
tf_image=$(\
yq eval ".parameters.openshift4_terraform.images.terraform.image" \
dependencies/openshift4-terraform/class/defaults.yml)
tf_tag=$(\
yq eval ".parameters.openshift4_terraform.images.terraform.tag" \
dependencies/openshift4-terraform/class/defaults.yml)
echo "Using Terraform image: ${tf_image}:${tf_tag}"
base_dir=$(pwd)
terraform() {
touch .terraformrc
docker run --rm -e REAL_UID="$(id -u)" -e TF_CLI_CONFIG_FILE=/tf/.terraformrc --env-file "${base_dir}/terraform.env" -w /tf -v "$(pwd):/tf" --ulimit memlock=-1 "${tf_image}:${tf_tag}" /tf/terraform.sh "${@}"
}
gitlab_repository_url=$(curl -sH "Authorization: Bearer $(commodore fetch-token)" ${INPUT_commodore_api_url}/clusters/${INPUT_commodore_cluster_id} | jq -r '.gitRepo.url' | sed 's|ssh://||; s|/|:|')
gitlab_repository_name=${gitlab_repository_url##*/}
gitlab_catalog_project_id=$(curl -sH "Authorization: Bearer ${INPUT_gitlab_api_token}" "https://git.vshn.net/api/v4/projects?simple=true&search=${gitlab_repository_name/.git}" | jq -r ".[] | select(.ssh_url_to_repo == \"${gitlab_repository_url}\") | .id")
gitlab_state_url="https://git.vshn.net/api/v4/projects/${gitlab_catalog_project_id}/terraform/state/cluster"
pushd catalog/manifests/openshift4-terraform/
terraform init \
"-backend-config=address=${gitlab_state_url}" \
"-backend-config=lock_address=${gitlab_state_url}/lock" \
"-backend-config=unlock_address=${gitlab_state_url}/lock" \
"-backend-config=username=${INPUT_gitlab_user_name}" \
"-backend-config=password=${INPUT_gitlab_api_token}" \
"-backend-config=lock_method=POST" \
"-backend-config=unlock_method=DELETE" \
"-backend-config=retry_wait_min=5"
SUBNET_UUID="$(terraform output -raw subnet_uuid)"
INGRESS_FLOATING_IP_V4="$(terraform output -raw router_vip)"
INGRESS_FLOATING_IP_V6="$(terraform output -raw router_vip_v6)"
pushd ../../../inventory/classes/${INPUT_commodore_tenant_id}
yq eval -i '.parameters.openshift.cloudscale.subnet_uuid = "'"$SUBNET_UUID"'"' \
${INPUT_commodore_cluster_id}.yml
yq eval -i '.parameters.openshift.cloudscale.ingress_floating_ip_v4 = "'"$INGRESS_FLOATING_IP_V4"'"' \
${INPUT_commodore_cluster_id}.yml
yq eval -i '.parameters.openshift.cloudscale.ingress_floating_ip_v6 = "'"$INGRESS_FLOATING_IP_V6"'"' \
${INPUT_commodore_cluster_id}.yml
if git diff-index --quiet HEAD
then
echo "No changes, skipping commit"
else
git commit -am "Configure cloudscale subnet UUID and ingress floating IP for ${INPUT_commodore_cluster_id}"
git push origin master
fi || true
popd
popd # yes, twice.
# Recompile the catalog
commodore catalog compile ${INPUT_commodore_cluster_id} --push \
--dynamic-fact kubernetesVersion.major=1 \
--dynamic-fact kubernetesVersion.minor="$((INPUT_image_minor+13))" \
--dynamic-fact openshiftVersion.Major=${INPUT_image_major} \
--dynamic-fact openshiftVersion.Minor=${INPUT_image_minor}
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
Then I deploy initial manifests
This step deploys some manifests required during bootstrap, including cert-manager, machine-api-provider, machinesets, loadbalancer controller, and ingress loadbalancer.
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_api_url=
# export INPUT_vault_address=
# export INPUT_vault_login_method=
# export INPUT_kubeconfig_path=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
export KUBECONFIG="${INPUT_kubeconfig_path}"
export VAULT_ADDR=${INPUT_vault_address}
vault login -method=${INPUT_vault_login_method}
echo '# Applying cert-manager ... #'
kubectl apply -f catalog/manifests/cert-manager/00_namespace.yaml
kubectl apply -Rf catalog/manifests/cert-manager/10_cert_manager
# shellcheck disable=2046
# we need word splitting here
kubectl -n syn-cert-manager patch --type=merge \
$(kubectl -n syn-cert-manager get deploy -oname) \
-p '{"spec":{"template":{"spec":{"tolerations":[{"operator":"Exists"}]}}}}'
echo '# Waiting for cert-manager to become available ... #'
kubectl -n syn-cert-manager wait --for=condition=available --timeout=300s \
deploy/cert-manager deploy/cert-manager-webhook
echo '# Applied cert-manager. #'
echo
echo '# Applying machine-api-provider ... #'
VAULT_TOKEN=$(vault token lookup -format=json | jq -r .data.id)
export VAULT_TOKEN
kapitan refs --reveal --refs-path catalog/refs -f catalog/manifests/machine-api-provider-cloudscale/00_secrets.yaml | kubectl apply -f -
kubectl apply -f catalog/manifests/machine-api-provider-cloudscale/10_clusterRoleBinding.yaml
kubectl apply -f catalog/manifests/machine-api-provider-cloudscale/10_serviceAccount.yaml
kubectl apply -f catalog/manifests/machine-api-provider-cloudscale/11_deployment.yaml
echo '# Applied machine-api-provider. #'
echo
echo '# Applying machinesets ... #'
for f in catalog/manifests/openshift4-nodes/machineset-*.yaml;
do kubectl apply -f "$f";
done
echo '# Applied machinesets. #'
echo
echo '# Applying loadbalancer controller ... #'
kubectl apply -f catalog/manifests/cloudscale-loadbalancer-controller/00_namespace.yaml
kapitan refs --reveal --refs-path catalog/refs -f catalog/manifests/cloudscale-loadbalancer-controller/10_secrets.yaml | kubectl apply -f -
# TODO(aa): This fails on the first attempt because likely some of the previous resources need time to come online; figure out what to wait for
# TODO(sg): This still fails sometimes, even when waiting for cert-manager
# to become available. From the errors I've observed, there's some case
# where deploying the cert-manager resources runs into a timeout related
# to the validatingwebhookconfig for cert-manager.
until kubectl apply -Rf catalog/manifests/cloudscale-loadbalancer-controller/10_kustomize
do
echo "Manifests didn't apply, waiting a moment to try again ..."
sleep 20
done
echo "Waiting for load balancer controller to become available ..."
# NOTE(sg): We need to wait for at least one app node to become ready
# before the cloudscale LB controller can be scheduled. This can take 5+
# minutes, depending on how quickly the worker nodes are becoming ready.
kubectl -n appuio-cloudscale-loadbalancer-controller \
wait --for condition=available --timeout 10m \
deploy cloudscale-loadbalancer-controller-controller-manager
echo '# Applied loadbalancer controller. #'
echo
echo '# Applying ingress loadbalancer ... #'
kubectl apply -f catalog/manifests/cloudscale-loadbalancer-controller/20_loadbalancers.yaml
echo '# Applied ingress loadbalancer. #'
echo
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I wait for bootstrap to complete
This step waits for OpenShift bootstrap to complete successfully.
Script
OUTPUT=$(mktemp)
# export INPUT_openshift_install_bin=
set -euo pipefail
openshift-install() {
"${INPUT_openshift_install_bin}" "${@}"
}
installer_dir="$(pwd)/target"
openshift-install --dir "${installer_dir}" \
wait-for bootstrap-complete --log-level debug
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
Then I remove the bootstrap node
After successful bootstrapping, this step removes the bootstrap node again.
Inputs
-
cloudscale_token -
cloudscale_token_floaty -
control_vshn_api_token -
ignition_bootstrap -
gitlab_user_name -
gitlab_api_token -
commodore_cluster_id -
commodore_api_url -
lb_fqdn_1 -
lb_fqdn_2
Script
OUTPUT=$(mktemp)
# export INPUT_cloudscale_token=
# export INPUT_cloudscale_token_floaty=
# export INPUT_control_vshn_api_token=
# export INPUT_ignition_bootstrap=
# export INPUT_gitlab_user_name=
# export INPUT_gitlab_api_token=
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_api_url=
# export INPUT_lb_fqdn_1=
# export INPUT_lb_fqdn_2=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
cat <<EOF > ./terraform.env
CLOUDSCALE_API_TOKEN=${INPUT_cloudscale_token}
TF_VAR_ignition_bootstrap=${INPUT_ignition_bootstrap}
TF_VAR_lb_cloudscale_api_secret=${INPUT_cloudscale_token_floaty}
TF_VAR_control_vshn_net_token=${INPUT_control_vshn_api_token}
GIT_AUTHOR_NAME=$(git config --global user.name)
GIT_AUTHOR_EMAIL=$(git config --global user.email)
HIERADATA_REPO_TOKEN=${INPUT_gitlab_api_token}
EOF
tf_image=$(\
yq eval ".parameters.openshift4_terraform.images.terraform.image" \
dependencies/openshift4-terraform/class/defaults.yml)
tf_tag=$(\
yq eval ".parameters.openshift4_terraform.images.terraform.tag" \
dependencies/openshift4-terraform/class/defaults.yml)
echo "Using Terraform image: ${tf_image}:${tf_tag}"
base_dir=$(pwd)
terraform() {
touch .terraformrc
docker run --rm -e REAL_UID="$(id -u)" -e TF_CLI_CONFIG_FILE=/tf/.terraformrc --env-file "${base_dir}/terraform.env" -w /tf -v "$(pwd):/tf" --ulimit memlock=-1 "${tf_image}:${tf_tag}" /tf/terraform.sh "${@}"
}
gitlab_repository_url=$(curl -sH "Authorization: Bearer $(commodore fetch-token)" ${INPUT_commodore_api_url}/clusters/${INPUT_commodore_cluster_id} | jq -r '.gitRepo.url' | sed 's|ssh://||; s|/|:|')
gitlab_repository_name=${gitlab_repository_url##*/}
gitlab_catalog_project_id=$(curl -sH "Authorization: Bearer ${INPUT_gitlab_api_token}" "https://git.vshn.net/api/v4/projects?simple=true&search=${gitlab_repository_name/.git}" | jq -r ".[] | select(.ssh_url_to_repo == \"${gitlab_repository_url}\") | .id")
gitlab_state_url="https://git.vshn.net/api/v4/projects/${gitlab_catalog_project_id}/terraform/state/cluster"
pushd catalog/manifests/openshift4-terraform/
terraform init \
"-backend-config=address=${gitlab_state_url}" \
"-backend-config=lock_address=${gitlab_state_url}/lock" \
"-backend-config=unlock_address=${gitlab_state_url}/lock" \
"-backend-config=username=${INPUT_gitlab_user_name}" \
"-backend-config=password=${INPUT_gitlab_api_token}" \
"-backend-config=lock_method=POST" \
"-backend-config=unlock_method=DELETE" \
"-backend-config=retry_wait_min=5"
# NOTE(sg): override.tf doesn't exist anymore at this point if we're not
# provisioning Puppet LBs.
rm override.tf || true
terraform apply --auto-approve
if [ "$INPUT_lb_fqdn_1" != "none" ]; then
echo "@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@"
echo "@ @"
echo "@ Please review and merge the LB hieradata MR listed in Terraform output hieradata_mr. @"
echo "@ @"
echo "@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@"
while (GITLAB_HOST=git.vshn.net GITLAB_TOKEN="${INPUT_gitlab_api_token}" glab mr list -R=appuio/appuio_hieradata | grep "${INPUT_commodore_cluster_id}")
do
sleep 10
done
echo PR merged, waiting for CI to finish...
sleep 10
while (GITLAB_HOST=git.vshn.net GITLAB_TOKEN="${INPUT_gitlab_api_token}" glab mr list -R=appuio/appuio_hieradata | grep "running")
do
sleep 10
done
ssh "${INPUT_lb_fqdn_1}" sudo puppetctl run
ssh "${INPUT_lb_fqdn_2}" sudo puppetctl run
fi
popd
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I configure initial deployments
This step configures some deployments that require manual changes after cluster bootstrap, such as enabling proxy protocol on the Ingress controller, and scheduling the ingress controller on the infrastructure nodes.
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_api_url=
# export INPUT_kubeconfig_path=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
export KUBECONFIG="${INPUT_kubeconfig_path}"
echo '# Enabling proxy protocol ... #'
kubectl -n openshift-ingress-operator patch ingresscontroller default --type=json \
-p '[{
"op":"replace",
"path":"/spec/endpointPublishingStrategy",
"value": {"type": "HostNetwork", "hostNetwork": {"protocol": "PROXY"}}
}]'
echo '# Enabled proxy protocol. #'
echo
distribution="$(curl -sH "Authorization: Bearer $(commodore fetch-token)" ${COMMODORE_API_URL}/clusters/${INPUT_commodore_cluster_id} | jq -r .facts.distribution)"
if [[ "$distribution" != "oke" ]]
then
echo '# Scheduling ingress controller on infra nodes ... #'
kubectl -n openshift-ingress-operator patch ingresscontroller default --type=json \
-p '[{
"op":"replace",
"path":"/spec/nodePlacement",
"value":{"nodeSelector":{"matchLabels":{"node-role.kubernetes.io/infra":""}}}
}]'
echo '# Scheduled ingress controller on infra nodes. #'
echo
fi
echo '# Removing temporary cert-manager tolerations ... #'
# shellcheck disable=2046
# we need word splitting here
kubectl -n syn-cert-manager patch --type=json \
$(kubectl -n syn-cert-manager get deploy -oname) \
-p '[{"op":"remove","path":"/spec/template/spec/tolerations"}]'
echo '# Removed temporary cert-manager tolerations. #'
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I wait for installation to complete
This step waits for OpenShift installation to complete successfully.
Script
OUTPUT=$(mktemp)
# export INPUT_openshift_install_bin=
set -euo pipefail
openshift-install() {
"${INPUT_openshift_install_bin}" "${@}"
}
installer_dir="$(pwd)/target"
openshift-install --dir "${installer_dir}" \
wait-for install-complete --log-level debug
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
Then I synthesize the cluster
This step enables Project Syn on the cluster.
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_api_url=
# export INPUT_commodore_cluster_id=
# export INPUT_kubeconfig_path=
set -euo pipefail
export COMMODORE_API_URL="${INPUT_commodore_api_url}"
export KUBECONFIG="${INPUT_kubeconfig_path}"
LIEUTENANT_AUTH="Authorization:Bearer $(commodore fetch-token)"
if ! kubectl get deploy -n syn steward > /dev/null; then
INSTALL_URL=$(curl -H "${LIEUTENANT_AUTH}" "${COMMODORE_API_URL}/clusters/${INPUT_commodore_cluster_id}" | jq -r ".installURL")
if [[ $INSTALL_URL == "null" ]]
# TODO(aa): consider doing this programmatically - especially if, at a later point, we add the lieutenant kubeconfig to the inputs anyway
then
echo '###################################################################################'
echo '# #'
echo '# Could not fetch install URL! Please reset the bootstrap token and try again. #'
echo '# #'
echo '###################################################################################'
echo
echo 'See https://kb.vshn.ch/corp-tech/projectsyn/explanation/bootstrap-token.html#_resetting_the_bootstrap_token'
sleep 0.1
exit 1
fi
echo "# Deploying steward ..."
kubectl create -f "$INSTALL_URL"
fi
echo "# Waiting for ArgoCD resource to exist ..."
kubectl wait --for=create crds/argocds.argoproj.io --timeout=5m
echo "# Waiting for ArgoCD resource to become ready ..."
sleep 1
kubectl wait --for=condition=NamesAccepted=True crd argocds.argoproj.io --timeout=60s
kubectl wait --for=condition=Established=True crd argocds.argoproj.io --timeout=60s
echo "# Refreshing local Kubernetes API cache ..."
kubectl api-resources --api-group=argoproj.io
echo "# Waiting for ArgoCD instance to exist ..."
kubectl wait --for=create argocd/syn-argocd -nsyn --timeout=90s
echo "# Waiting for ArgoCD instance to be ready ..."
kubectl wait --for=jsonpath='{.status.phase}'=Available argocd/syn-argocd -nsyn --timeout=5m
echo "Done."
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
Then I set acme-dns CNAME records
This step ensures CNAME records exist for ACME challenges once cert-manager is properly deployed.
Script
OUTPUT=$(mktemp)
# export INPUT_kubeconfig_path=
# export INPUT_cluster_domain=
set -euo pipefail
export KUBECONFIG="${INPUT_kubeconfig_path}"
echo '# Waiting for cert-manager namespace ...'
kubectl wait --for=create ns/syn-cert-manager --timeout 3m
echo '# Waiting for cert-manager secret ...'
kubectl wait --for=create secret/acme-dns-client --timeout 10m -nsyn-cert-manager
fulldomain=""
while [[ -z "$fulldomain" ]]
do
fulldomain=$(kubectl -n syn-cert-manager \
get secret acme-dns-client \
-o jsonpath='{.data.acmedns\.json}' | \
base64 -d | \
jq -r '[.[]][0].fulldomain')
echo "$fulldomain"
done
echo "_acme-challenge.api IN CNAME $fulldomain." > dns.txt
echo "_acme-challenge.apps IN CNAME $fulldomain." >> dns.txt
echo "@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@"
echo "@ @"
echo "@ Please add the acme DNS records below to your DNS provider. @"
echo "@ Most probably in https://git.vshn.net/vshn/vshn_zonefiles @"
echo "@ @"
echo "@ If terminal selection does not work the entries can also be copied from @"
echo "@ dns.txt @"
echo "@ @"
echo "@ Waiting for record to propagate... @"
echo "@ @"
echo "@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@"
echo
echo "The following entry must be created in the same origin as the api record:"
echo "_acme-challenge.api IN CNAME $fulldomain."
echo "The following entry must be created in the same origin as the apps record:"
echo "_acme-challenge.apps IN CNAME $fulldomain."
echo
# back up kubeconfig, just in case.
# NOTE(sg): Don't overwrite backed up kubeconfig when the step is retried.
if [ ! -f "${INPUT_kubeconfig_path}.bak" ]; then
cp "${INPUT_kubeconfig_path}" "${INPUT_kubeconfig_path}.bak"
fi
# remove cluster CA cert and disable certificate verification in kubeconfig
echo "Removing cluster CA and temporarily disabling TLS verification in admin kubeconfig ..."
yq -i e 'del(.clusters[0].cluster.certificate-authority-data)' "${INPUT_kubeconfig_path}"
# Set insecure-skip-tls-verify=true in kubeconfig so we can use it to wait
# for the kube-apiserver CO
yq -i e '.clusters[0].cluster.insecure-skip-tls-verify = true' "${INPUT_kubeconfig_path}"
echo "Waiting for cluster certificate to be issued ..."
# NOTE(sg): We query the cluster zone's SOA record and extract the
# negative TTL from it. We then configure the wait for the cert-manager
# certificate to become ready to 2*negative TTL. This should be
# sufficient to ensure that any possibly cached NXDOMAIN responses for the
# acme-challenge records have expired.
zone_nx_ttl_secs=$(dig +noall +authority SOA "${INPUT_cluster_domain}" \
| tr "\t" " " | tr -s " " | cut -d " " -f11)
kubectl wait -n openshift-config --for condition=ready \
--timeout="$(( zone_nx_ttl_secs * 2 ))s" \
certificate/api-server-cluster-certificate-default
echo "Waiting for all kube-apiserver instances to be updated ..."
kubectl wait --for condition=progressing=false --timeout=15m co kube-apiserver
echo "Enabling TLS verification in admin kubeconfig again ..."
yq -i e 'del(.clusters[0].cluster.insecure-skip-tls-verify)' "${INPUT_kubeconfig_path}"
kubectl get nodes
rm dns.txt
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I verify emergency access
This step ensures the emergency credentials for the cluster can be retrieved.
Inputs
-
kubeconfig_path -
commodore_cluster_id -
commodore_api_url -
passbolt_passphrase: Your password for Passbolt.
This is required to access the encrypted emergency credentials.
Script
OUTPUT=$(mktemp)
# export INPUT_kubeconfig_path=
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_api_url=
# export INPUT_passbolt_passphrase=
set -euo pipefail
export KUBECONFIG="${INPUT_kubeconfig_path}"
echo '# Waiting for emergency-credentials-controller namespace ...'
kubectl wait --for=create ns/appuio-emergency-credentials-controller
echo '# Waiting for emergency-credentials-controller ...'
kubectl wait --for=create secret/acme-dns-client -nsyn-cert-manager
echo '# Waiting for emergency credential tokens ...'
until kubectl -n appuio-emergency-credentials-controller get emergencyaccounts.cluster.appuio.io -o=jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.lastTokenCreationTimestamp}{"\n"}{end}' | grep "$( date '+%Y' )" >/dev/null
do
echo -n .
done
export KHARON_PASSBOLT_PASSPHRASE="${INPUT_passbolt_passphrase}"
export KUBECONFIG="em-${INPUT_commodore_cluster_id}"
kharon update --lieutenant-url "${INPUT_commodore_api_url}" --inventory-file kharon-inventory.json --mapping-file kharon-mapping.json
kharon emergency-credentials "${INPUT_commodore_cluster_id}" --inventory-file kharon-inventory.json
yq -i e '.clusters[0].cluster.insecure-skip-tls-verify = true' "em-${INPUT_commodore_cluster_id}"
kubectl get nodes
oc whoami | grep system:serviceaccount:appuio-emergency-credentials-controller: || exit 1
env -i "kubeconfig_path=$(pwd)/em-${INPUT_commodore_cluster_id}" >> "$OUTPUT"
echo "# Invalidating 10-year admin kubeconfig ..."
kubectl -n openshift-config patch cm admin-kubeconfig-client-ca --type=merge -p '{"data": {"ca-bundle.crt": ""}}'
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I configure the cluster alerts
This step configures monitoring alerts on the cluster.
Script
OUTPUT=$(mktemp)
# export INPUT_kubeconfig_path=
set -euo pipefail
export KUBECONFIG="${INPUT_kubeconfig_path}"
kubectl wait --for create --timeout=10m -n openshift-monitoring cronjob silence
echo '# Installing default alert silence ...'
oc --as=system:admin -n openshift-monitoring get job silence-manual &>/dev/null || \
oc --as=system:admin -n openshift-monitoring create job --from=cronjob/silence silence-manual
oc wait -n openshift-monitoring --for=condition=complete job/silence-manual
oc --as=system:admin -n openshift-monitoring delete job/silence-manual
echo '# Retrieving active alerts ...'
kubectl --as=system:admin -n openshift-monitoring exec sts/alertmanager-main -- \
amtool --alertmanager.url=http://localhost:9093 alert --active
echo
echo '#######################################################'
echo '# #'
echo '# Please review the list of open alerts above, #'
echo '# address any that require action before proceeding. #'
echo '# #'
echo '#######################################################'
sleep 2
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I enable Opsgenie alerting
This step enables Opsgenie alerting for the cluster via Project Syn.
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_cluster_id=
# export INPUT_commodore_tenant_id=
set -euo pipefail
pushd "inventory/classes/${INPUT_commodore_tenant_id}/"
yq eval -i 'del(.classes[] | select(. == "*.no-opsgenie"))' ${INPUT_commodore_cluster_id}.yml
git commit -a -m "Enable opsgenie alerting on cluster ${INPUT_commodore_cluster_id}"
git push
popd
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I verify the image registry config
This step verifies that the image registry config has bootstrapped correctly.
Script
OUTPUT=$(mktemp)
# export INPUT_kubeconfig_path=
set -euo pipefail
export KUBECONFIG="${INPUT_kubeconfig_path}"
echo '# Checking image registry status conditions ...'
status="$( kubectl get config.imageregistry/cluster -oyaml --as system:admin | yq '.status.conditions[] | select(.type == "Available").status' )"
if [[ $status != "True" ]]
then
kubectl get config.imageregistry/cluster -oyaml --as system:admin | yq '.status.conditions'
echo
echo ERROR: image registry is not available.
echo Please review the status reports above and manually fix the registry.
echo
echo > kubectl get config.imageregistry/cluster
exit 1
fi
echo '# Checking image registry pods ...'
numpods="$( kubectl -n openshift-image-registry get pods -l docker-registry=default --field-selector=status.phase==Running -oyaml | yq '.items | length' )"
if (( numpods != 2 ))
then
kubectl -n openshift-image-registry get pods -l docker-registry=default
echo
echo ERROR: unexpected number of registry pods
echo Please review the running pods above and ensure the 2 registry pods are running.
echo
echo > kubectl -n openshift-image-registry get pods -l docker-registry=default
exit 1
fi
echo '# Ensuring openshift-samples operator is enabled ...'
mgstate="$( kubectl get config.samples cluster -ojsonpath='{.spec.managementState}' )"
if [[ $mgstate != "Managed" ]]
then
kubectl patch config.samples cluster -p '{"spec":{"managementState":"Managed"}}'
fi
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I wait until all ArgoCD apps are synced and healthy
Wait until all ArgoCD applications are synced and healthy.
Script
OUTPUT=$(mktemp)
# export INPUT_kubeconfig_path=
set -euo pipefail
export KUBECONFIG="${INPUT_kubeconfig_path}"
sync_timeout=600s
healthy_timeout=300s
echo "Wait until all ArgoCD apps are synced... (timeout=${sync_timeout})"
if ! kubectl -n syn wait --for jsonpath='{.status.sync.status}="Synced"' \
applications.argoproj.io --all --timeout="${sync_timeout}"; then
echo "❌ Wait for ArgoCD app sync timed out, check status manually!"
else
echo "✅ All ArgoCD apps synced"
fi
echo "Wait until all ArgoCD apps are healthy... (timeout=${healthy_timeout})"
if ! kubectl -n syn wait --for jsonpath='{.status.health.status}="Healthy"' \
applications.argoproj.io --all --timeout="${healthy_timeout}"; then
echo "❌ Wait for ArgoCD app health timed out, check status manually!"
else
echo "✅ Cluster synthesis complete. All ArgoCD apps healthy."
fi
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I schedule the first maintenance
This step verifies that the UpgradeConfig object is present on the cluster, and schedules a first maintenance.
Script
OUTPUT=$(mktemp)
# export INPUT_kubeconfig_path=
set -euo pipefail
export KUBECONFIG="${INPUT_kubeconfig_path}"
numconfs="$( kubectl -n appuio-openshift-upgrade-controller get upgradeconfig -oyaml | yq '.items | length' )"
if (( numconfs < 1 ))
then
kubectl -n appuio-openshift-upgrade-controller get upgradeconfig
echo
echo ERROR: did not find an upgradeconfig
echo Please review the output above and ensure an upgradeconfig is present.
echo
echo "Double check the cluster's maintenance_window fact."
exit 1
fi
echo '# Scheduling a first maintenance ...'
uc="$(yq .parameters.facts.maintenance_window inventory/classes/params/cluster.yml)"
kubectl -n appuio-openshift-upgrade-controller get upgradeconfig "$uc" -oyaml | \
yq '
.metadata.name = "first",
.metadata.labels = {},
.spec.jobTemplate.metadata.labels.upgradeconfig/name = "first",
.spec.schedule.cron = ((now+"1m")| tz("Europe/Zurich") | format_datetime("4 15")) + " * * *",
.spec.pinVersionWindow = "0m"
' | \
kubectl create -f - --as=system:admin
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I remove the bootstrap bucket
This step deletes the S3 bucket with the bootstrap ignition config.
Script
OUTPUT=$(mktemp)
# export INPUT_commodore_cluster_id=
# export INPUT_csp_region=
# export INPUT_bucket_user=
set -euo pipefail
mc alias set \
"${INPUT_commodore_cluster_id}" "https://objects.${INPUT_csp_region}.cloudscale.ch" \
"$(echo "$INPUT_bucket_user" | jq -r '.keys[0].access_key')" \
"$(echo "$INPUT_bucket_user" | jq -r '.keys[0].secret_key')"
mc rm -r --force "${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-bootstrap-ignition"
mc rb "${INPUT_commodore_cluster_id}/${INPUT_commodore_cluster_id}-bootstrap-ignition"
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"
And I wait for maintenance to complete
This step waits for the first maintenance to complete, and then removes the initial UpgradeConfig.
Script
OUTPUT=$(mktemp)
# export INPUT_kubeconfig_path=
set -euo pipefail
export KUBECONFIG="${INPUT_kubeconfig_path}"
echo "# Waiting for initial maintenance to complete ..."
oc get clusterversion
until kubectl wait --for=condition=Succeeded upgradejob -l "upgradeconfig/name=first" -n appuio-openshift-upgrade-controller 2>/dev/null
do
oc get clusterversion | grep -v NAME
sleep 30
done
echo "# Deleting initial UpgradeConfig ..."
kubectl --as=system:admin -n appuio-openshift-upgrade-controller \
delete upgradeconfig first
# echo "# Outputs"
# cat "$OUTPUT"
# rm -f "$OUTPUT"