Files
DodoSSH/.github/workflows/ci.yml
T
jaap-jan 1bcf422bbe Build the phone once per run, and stop rebuilding the toolchain image
The android job compiled everything twice. The plain `dotnet build` before the
packaging step looked like a cheap check ahead of an expensive one and was
neither: SignAndroidPackage depends on Build, so the packaging line compiles
everything anyway — and the build above it ran with no -p:DodoChannel, which
means it ran as the *release* channel. Different application id, different
version, different assembly metadata; MSBuild treats a different set of global
properties as a different project instance, so not one output was reused. It was
a full second compile of the reference closure, producing an APK for the one
channel this job must never build, thrown away unread.

Measured in the toolchain image with a warm package volume, same commit:

  two builds   2m52 + 2m26   5m21 total
  one build            3m01   3m04 total

Byte for byte the same artefact out of both — versionCode 203, versionName
0.0.0-alpha.0.136, dev.dodotech.dodossh.nightly.

The toolchain image is now built once per Dockerfile rather than once per run.
The tag is the Dockerfile's own digest and docker applies a tag only on success,
so an existing tag is by construction the right image and `docker image inspect`
is a sound check rather than a guess. What that trades away is the JDK from apt
drifting; everything that decides what is in the image is pinned in the
Dockerfile, so anything that matters changes the digest. It is a build tool, not
something shipped — the image job takes the opposite trade with --pull, because
what it builds is what users run.

And the build job now says whether its package cache did anything. setup-dotnet's
cache: true is actions/cache underneath, which needs a cache server act_runner
ships and can have turned off — and when it is off it does nothing and says
nothing about it. A cold restore and a perfect cache look identical from outside:
both are green, and the difference is minutes. One `find` before anything writes
to the folder turns that from a belief into a line in the log. It does not fail
the build, because a runner without a cache server is slow rather than wrong.

What is deliberately not cached: the apt installs in each job's preamble, which
need the runner's image fixed rather than a workflow change and already say so;
the Testcontainers pulls and the API image's layers, which the daemon already
caches on a persistent runner; and the android obj/bin, which would not help —
the source arrives by `docker cp` with fresh timestamps, so MSBuild rebuilds it
whatever is in there.
2026-08-05 12:40:28 +02:00

911 lines
47 KiB
YAML

name: ci
on:
push:
branches: [main]
# Release tags run the whole workflow, not only the image job that gates on it. A tag
# is the one build nobody is watching, so it is the last place to take the tests on
# trust.
tags: ['v*']
pull_request:
branches: [main]
# Actions are pinned to commit SHAs, not tags: a tag can be moved to point at new code,
# which would let a compromised action run with this workflow's permissions.
permissions:
contents: read
concurrency:
group: ci-${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
env:
DOTNET_NOLOGO: true
DOTNET_CLI_TELEMETRY_OPTOUT: true
DOTNET_SKIP_FIRST_TIME_EXPERIENCE: true
CI: true
jobs:
build:
name: build and test
runs-on: [linux]
steps:
# act_runner runs every `uses:` action with node inside the job container, and this
# runner's image has none — the run died on the first line of actions/checkout with
# "Cannot find: node in PATH". A `run:` step is shell rather than node, so this one
# can go first and unblock the rest.
#
# This is a workaround and the real fix is one line of the runner's own config.yaml:
# point container.image at an image that ships node, the way Gitea's default
# catthehacker/ubuntu:act-latest does. Kept anyway, because a pipeline that depends
# on a runner being configured correctly somewhere else fails confusingly when it is
# not, and because it costs nothing on a runner that is.
#
# git as well as node, and said in the step name rather than smuggled in: checkout
# shells out to git the moment node has loaded it, so an image thin enough to lack
# one usually lacks the other, and learning that costs a whole second CI round trip.
#
# Repeated verbatim in all three jobs, which is not laziness. It cannot be a local
# composite action — that would need the checkout it exists to unblock — and YAML
# anchors, which would deduplicate it, are rejected by GitHub's parser and would make
# this file portable to nothing. Change one copy, change all three.
- name: ensure node and git
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
missing=""
command -v node >/dev/null 2>&1 || missing="$missing nodejs"
command -v git >/dev/null 2>&1 || missing="$missing git"
if [ -z "$missing" ]; then
echo "node $(node --version), git $(git --version)"
exit 0
fi
echo "Installing:$missing"
if command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends $missing
elif command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache $missing
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y $missing
else
echo "No apt-get, apk or dnf here, so node cannot be installed from inside the" >&2
echo "job. Point the runner's container.image at something that ships node." >&2
exit 1
fi
echo "node $(node --version), git $(git --version)"
# Warned about rather than failed on. Distributions pin their nodejs package to
# the release they shipped with — Ubuntu 24.04 still serves 18, which is past end
# of life and older than the runtime these actions declare. It generally runs
# them anyway, since act_runner uses whichever node is on PATH regardless of what
# the action asked for, so this is a note for when one of them misbehaves in a
# way that makes no sense, not a reason to stop a build that is probably fine.
major="$(node --version | sed 's/^v//; s/\..*//')"
if [ "$major" -lt 20 ]; then
echo "::warning::node $major is older than the runtime these actions target;" \
"give the runner an image with node 20 or newer if actions misbehave."
fi
# fetch-depth 0, and it is load-bearing rather than tidy. MinVer derives the version from the
# nearest v* tag, and checkout's default shallow clone has no tags at all — so it would not fail,
# it would quietly answer 0.0.0-alpha.0.N and every build would ship that. Velopack decides
# whether an installed client is out of date by comparing versions, which makes a plausible wrong
# answer here a client that never updates.
#
# Repeated in all three jobs, like the node preamble above and for the same reason. Change one
# copy, change all three.
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 0
# cache: true is actions/cache underneath, keyed on the lock files. That needs a cache server,
# which act_runner ships and can be turned off — and when it is off, this does nothing at all and
# says nothing about it. See the step below, which is how you find out.
- uses: actions/setup-dotnet@a98b56852c35b8e3190ac28c8c2271da59106c68 # v6.0.0
with:
global-json-file: global.json
cache: true
cache-dependency-path: '**/packages.lock.json'
# ◆ ONE LINE, AND IT EXISTS BECAUSE THE ALTERNATIVE IS ASSUMING.
#
# A restore that is cold every single run and a cache working perfectly look identical from
# outside: both end in a green build, and the only difference is minutes. This job's whole
# package cache is the step above, so whether that step works is worth knowing as a fact rather
# than as a belief — and it is one `find` before anything has written to the folder.
#
# It does not fail the build. A runner with no cache server is slow rather than wrong, and a
# pipeline that refused to run on one would be a worse trade than the minutes.
- name: what the package cache brought
run: |
set -eu
root="${NUGET_PACKAGES:-$HOME/.nuget/packages}"
if [ -d "$root" ]; then
echo "$root holds $(find "$root" -mindepth 1 -maxdepth 1 -type d | wc -l) package(s) already."
else
echo "$root is not there, so this restore is cold."
echo "If every run says that, the cache above is doing nothing: check that this runner's"
echo "config.yaml has the cache server enabled, and that ACTIONS_CACHE_URL reaches it."
fi
# Locked mode fails if packages.lock.json does not match the project files, so a
# dependency cannot change without the lock file change being reviewed.
- name: restore
run: dotnet restore DodoSSH.slnx --locked-mode
# The version comes from the tag, so on a tag build there are two ways to say the same number
# and they can disagree — a moved tag, a tag on the wrong commit, or a checkout that somehow
# still lost its history. What makes that worth a step of its own is that the disagreement is
# silent everywhere else: MinVer answers 0.0.0-alpha.0.N rather than failing, the build goes
# green, the package is cut, and the symptom arrives weeks later as clients that never update.
#
# -getProperty evaluates without building, so this costs a second and runs before the build.
- name: the tag and the version agree
if: startsWith(github.ref, 'refs/tags/v')
run: |
set -euo pipefail
tag="${GITHUB_REF#refs/tags/v}"
declared="$(dotnet msbuild src/DodoSSH.Client.App/DodoSSH.Client.App.csproj \
-getProperty:Version -nologo | tr -d '[:space:]')"
if [ "$tag" != "$declared" ]; then
echo "The tag says v$tag and MinVer computed $declared." >&2
echo >&2
echo "These come from the same place, so a mismatch means the checkout did not see the" >&2
echo "tag it is building — most likely fetch-depth, which must stay 0 in every job here." >&2
exit 1
fi
echo "v$declared"
# No `dotnet format --verify-no-changes` step. It re-analysed the whole solution before
# the build did, for minutes, to check something the build already checks: IDE0055 is an
# error in .editorconfig and TreatWarningsAsErrors is on, so a misformatted file fails
# the build step below on its own. What the separate step added was the ability to say
# so a few minutes earlier, and it cost more than that on every run.
# Avalonia's headless renderer is still Skia, and libSkiaSharp.so — which the layout
# test project copies into its own output — links against libfontconfig. Without that
# one library every test in DodoSSH.Client.App.Layout.Tests dies inside
# HeadlessUnitTestSession before it measures anything, and 69 tests fail for a reason
# none of their names or assertions mention.
#
# The library, and not fonts. Verified in a container where fc-list returns zero and
# the suite passes anyway: the application carries Inter itself, so nothing here needs
# a typeface installed — only the thing that would have gone looking for one.
- name: ensure skia's native dependency
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
if ldconfig -p 2>/dev/null | grep -q 'libfontconfig\.so\.1'; then
echo "libfontconfig present"
exit 0
fi
echo "Installing fontconfig"
if command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends libfontconfig1
elif command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache fontconfig
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y fontconfig
else
echo "No package manager here, so Skia cannot be given its dependency and the" >&2
echo "layout suite will fail to start. Add fontconfig to the runner's image." >&2
exit 1
fi
- name: build
run: dotnet build DodoSSH.slnx --no-restore --configuration Release
- name: test
run: dotnet test DodoSSH.slnx --no-build --configuration Release
# The one build shape nothing else here exercises: a self-contained RID-specific publish. Its
# failure mode is a restore graph or a native asset that resolves for net10.0 and not for
# net10.0/win-x64, which nobody would see until a person was halfway through cutting a release
# on a Windows machine. vpk can pack a Windows package from Linux; only signing needs Windows,
# and this repository signs nothing yet, so proving the publish here is worth the minutes.
#
# RestoreLockedMode=false for this command only, and it is not a loosened gate. The committed
# lock files are deliberately RID-free: declaring win-x64 on the desktop head writes a
# net10.0/win-x64 target into every project it references transitively, which includes
# DodoSSH.Contracts and DodoSSH.Crypto — and the API's Dockerfile restores those with no RID
# under locked mode, so the image job would fail NU1004. The gate is the locked solution
# restore at the top of this job, which is unchanged.
#
# It rewrites the lock files as it goes; nothing after this step reads them, and the runner's
# checkout is thrown away. The release script does the same thing and puts them back, because
# there the tree is somebody's working copy.
#
# After the tests rather than before them, so a red suite does not first spend a hundred
# megabytes pulling a win-x64 runtime pack. main and tags only, for the same reason: a break
# found by the person about to release is found early enough.
- name: the windows publish still resolves
if: github.event_name != 'pull_request'
run: >
dotnet publish src/DodoSSH.Client.App/DodoSSH.Client.App.csproj
--configuration Release --runtime win-x64 --self-contained true
-p:RestoreLockedMode=false
--output "$RUNNER_TEMP/win-x64-check"
# This includes the end-to-end suite, which starts PostgreSQL, Keycloak and an OpenSSH
# server through Testcontainers and runs the API as a child process — so it needs a
# Docker daemon and gets one here. That is why the tests run on ubuntu rather than
# macOS, whose runners have no daemon at all. Expect the Keycloak image pull to
# dominate a cold run.
# A failing run says only that tests failed and names a log file on a machine nobody
# has a shell on. Every diagnostic thing — exception type, message, stack — is inside
# that file, so a red build was a filename and a guess. This prints it.
#
# head rather than tail, and that is the whole trick: when a suite fails wholesale it
# writes one stack per test and they are all the same stack. The first is the one that
# explains it, and the last two hundred lines are the same sentence repeated.
- name: what actually failed
if: failure()
run: |
set +e
echo "=== distro ==="
cat /etc/os-release 2>/dev/null | head -3
id
echo "=== what Skia needs, and whether it is here ==="
# ldd against the copy the test project carries. Its unresolved rows are the
# answer whenever the layout suite dies in HeadlessUnitTestSession, and asking
# here beats inferring it from a managed TypeInitializationException.
skia="$(find tests -name 'libSkiaSharp.so' 2>/dev/null | head -1)"
if [ -n "$skia" ]; then
echo "$skia"
ldd "$skia" 2>&1 | grep -Ei 'not found|fontconfig|freetype' || echo " all resolved"
else
echo " libSkiaSharp.so was not in the test output at all"
fi
ldconfig -p 2>/dev/null | grep -ci fontconfig | sed 's/^/fontconfig entries in ldconfig: /'
echo "=== docker, for the Testcontainers suites ==="
docker version --format '{{.Server.Version}}' 2>&1 | head -2
echo "=== test logs ==="
find tests -path '*/TestResults/*.log' 2>/dev/null | while read -r log; do
echo "----- $log"
head -n 120 "$log"
done
exit 0
android:
name: android head
runs-on: [linux]
# Writes the nightly release at the end of the job; see that step for what the capability is and why
# it is acceptable here and nowhere else. Job-scoped, so nothing else in this file gains it.
permissions:
contents: write
steps:
# Duplicated from the build job; see the comment there for why it cannot be factored
# out. Any change here has to be made in all three.
- name: ensure node and git
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
missing=""
command -v node >/dev/null 2>&1 || missing="$missing nodejs"
command -v git >/dev/null 2>&1 || missing="$missing git"
if [ -z "$missing" ]; then
echo "node $(node --version), git $(git --version)"
exit 0
fi
echo "Installing:$missing"
if command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends $missing
elif command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache $missing
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y $missing
else
echo "No apt-get, apk or dnf here, so node cannot be installed from inside the" >&2
echo "job. Point the runner's container.image at something that ships node." >&2
exit 1
fi
echo "node $(node --version), git $(git --version)"
# Warned about rather than failed on. Distributions pin their nodejs package to
# the release they shipped with — Ubuntu 24.04 still serves 18, which is past end
# of life and older than the runtime these actions declare. It generally runs
# them anyway, since act_runner uses whichever node is on PATH regardless of what
# the action asked for, so this is a note for when one of them misbehaves in a
# way that makes no sense, not a reason to stop a build that is probably fine.
major="$(node --version | sed 's/^v//; s/\..*//')"
if [ "$major" -lt 20 ]; then
echo "::warning::node $major is older than the runtime these actions target;" \
"give the runner an image with node 20 or newer if actions misbehave."
fi
# fetch-depth 0, and it is load-bearing rather than tidy. MinVer derives the version from the
# nearest v* tag, and checkout's default shallow clone has no tags at all — so it would not fail,
# it would quietly answer 0.0.0-alpha.0.N and every build would ship that. Velopack decides
# whether an installed client is out of date by comparing versions, which makes a plausible wrong
# answer here a client that never updates.
#
# Repeated in all three jobs, like the node preamble above and for the same reason. Change one
# copy, change all three.
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 0
# A job of its own, because DodoSSH.Client.Android is deliberately not in DodoSSH.slnx.
# Adding it there would make the android workload and a full Android SDK a prerequisite of
# `dotnet build DodoSSH.slnx` for everyone — including the build job above, which needs
# neither and would grow several minutes for a head it does not compile.
#
# The cost of keeping it out is that nothing in the main job would notice this head
# breaking, which for a project sharing view models with the desktop one is a matter of
# when rather than whether. This job is that notice.
#
# ============ ◆ WHY THIS HEAD IS BUILT IN A CONTAINER AND THE OTHER TWO ARE NOT ============
#
# .NET for Android cannot build on this runner. Not "needs setting up" — cannot. The runner is
# Alpine, and the SDK's own MSBuild tasks load glibc shared libraries out of the workload pack
# into the build process, which a musl-linked dotnet will not do:
#
# error XARLP7000: Error relocating …/libZipSharpNative-3-3.so: __snprintf_chk: symbol not found
#
# There is no musl variant of that pack and no shim for a library loaded in-process — gcompat
# gets a glibc *executable* started and is no help whatsoever here. Everything this job used to
# do on the host to make Android work was therefore treatment of symptoms: an aapt2 reported
# missing while being present, an "unsupported version" of a binary that had never run. All of
# them were this one sentence in a different accent, and docs/platform-flags.md has the history
# so the next reader does not repeat the three rounds it took to hear it.
#
# So the toolchain moves into an image that is glibc, and the job keeps on the host only what the
# host is good at: checkout, git, and publishing. The daemon needs no arranging — the image job
# builds with it and every Testcontainers suite in the build job reaches it over the socket.
#
# There is no setup-dotnet here any more, and its absence is the point: nothing outside the
# container compiles anything, so a dotnet on the host would be a second toolchain with nothing
# to do and a version nobody checks.
- name: ensure the docker cli
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
# Only the client. The socket is already there and a daemon is already answering on it, so
# installing an engine would start a second one beside the one in use. The image job carries
# the same step and the same reasoning; change one, look at the other.
if ! command -v docker >/dev/null 2>&1; then
echo "Installing the docker cli"
if command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache docker-cli
elif command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends docker.io
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y docker-cli
else
echo "No apt-get, apk or dnf here, so the docker client cannot be installed from" >&2
echo "inside the job. Add it to the runner's image." >&2
exit 1
fi
fi
docker --version
# Tagged by the digest of the Dockerfile that made it, so the tag moves exactly when the toolchain
# does and never otherwise. This runner is persistent, so the first build pays for a JDK, an
# Android SDK and the android workload — several minutes — and every run after it is a cache hit
# that prints two lines. A change to that Dockerfile is what buys a new one.
# ◆ BUILT ONCE PER DOCKERFILE, NOT ONCE PER RUN.
#
# The tag is the Dockerfile's own digest, so a tag that exists is by construction an image built
# from exactly this Dockerfile — and docker applies a tag only when the build succeeded, so it
# cannot be a half-built one. That makes `docker image inspect` a sound cache check rather than a
# guess, and it skips the context upload and the layer walk that a no-op `docker build` still does.
#
# The staleness this trades away is real and bounded. Everything that decides what is *in* the
# image — the command-line tools zip, the build-tools version, the workload — is pinned in the
# Dockerfile, so changing any of them changes the digest and rebuilds. What drifts is the JDK from
# apt. This is a build tool rather than something shipped, so that is the right way round; the
# image job takes the opposite trade with --pull, because what it builds is what users run.
#
# Deleting the image, or editing the Dockerfile, is how you force it.
- name: the android toolchain image
run: |
set -eu
tag="dodossh-android-build:$(sha256sum build/android-build.Dockerfile | cut -c1-16)"
if docker image inspect "$tag" >/dev/null 2>&1; then
echo "$tag is already on this runner"
else
docker build -t "$tag" -f build/android-build.Dockerfile build
fi
echo "ANDROID_BUILD_IMAGE=$tag" >> "$GITHUB_ENV"
echo "$tag"
# ◆ COPIED IN AND COPIED OUT, NOT BIND-MOUNTED, AND THAT IS NOT A PREFERENCE.
#
# `-v "$PWD:/build"` fails here, and it fails quietly: this runner is itself a container with the
# host's Docker socket handed to it, so the workspace path it reports — /root/.cache/act/<hash>/…
# — is a path inside the runner, not on the daemon's host. Docker resolves a bind source on the
# daemon side, finds nothing there, helpfully creates an empty directory and mounts that. The
# container then starts perfectly and answers:
#
# MSBUILD : error MSB1009: Project file does not exist.
#
# Nothing in that says "empty mount", and nothing earlier in the job would have caught it: the
# image job's `docker build` sends its context over the API and Testcontainers mounts nothing, so
# neither of them proves a bind mount would work.
#
# `docker cp` goes over the same API and so does not care where the daemon lives. The repository
# is well under a megabyte packed, so copying the whole checkout in — .git included, because
# MinVer and the versionCode both read it — costs a moment.
#
# The NuGet cache is a named volume for the same reason. It lives on the daemon, needs no path
# either side can agree on, and turns the restore into a cache hit on every run after the first.
- name: build and package the nightly
id: nightly
run: |
set -eu
staged=artifacts/android-nightly
rm -rf "$staged"
# Created rather than run, because the copy has to happen between creating and starting: the
# script being executed arrives with it.
cid="$(docker create \
-e CI=true \
-v dodossh-nuget:/root/.nuget/packages \
"$ANDROID_BUILD_IMAGE" bash /build/scripts/ci-android.sh)"
trap 'docker rm -f "$cid" >/dev/null 2>&1 || true' EXIT
# ./. rather than . — with the trailing dot it is the directory's *contents* that land in
# /build, and without it the whole directory lands as /build/<name> and nothing is where the
# script expects it.
docker cp ./. "$cid:/build"
# --attach streams the build log and exits with the container's own status, so a failure in
# there is a failure here.
docker start --attach "$cid"
mkdir -p artifacts
docker cp "$cid:/build/$staged" artifacts/
# Read out of the manifest the container wrote rather than out of anything it printed. A
# container's stdout is the build log as well as its result, so a value parsed from it is one
# stray MSBuild line away from being wrong.
name="$(sed -n 's/.*"versionName":"\([^"]*\)".*/\1/p' "$staged/android-nightly.json")"
if [ -z "$name" ]; then
echo "The container produced no usable manifest." >&2
exit 1
fi
echo "staged=$staged" >> "$GITHUB_OUTPUT"
echo "version=$name" >> "$GITHUB_OUTPUT"
ls -la "$staged"
# ◆ PUBLISHING IT, AND WHAT THAT CAPABILITY IS.
#
# Whoever can write a release here can put a build on every nightly phone, because the client fetches
# from this feed and Android's only check is that the signature matches — and this channel's key is
# in the repository for everybody. That is the same capability as the signing key, reached through a
# different door, and it is exactly what ADR 0011 rule 1 keeps off runners.
#
# It is acceptable here for one reason: this is not that channel. A nightly is signed by a key with
# no secrecy to lose, installs under its own package id, and cannot update the application anybody
# is trusting with their credentials. The release channel has none of this — no job, no token, no
# key on a runner — and the two are separate applications so that no mistake here can reach it.
#
# main only. A tag build must not touch this: a v* tag is the release channel's, and cutting it is a
# person's job.
- name: publish the nightly
if: github.ref == 'refs/heads/main'
env:
FORGE: https://git.dodotech.cloud
REPO: DodoTech/DodoSSH
TOKEN: ${{ secrets.GITHUB_TOKEN }}
STAGED: ${{ steps.nightly.outputs.staged }}
VERSION: ${{ steps.nightly.outputs.version }}
run: |
set -euo pipefail
if [ -z "${TOKEN:-}" ]; then
echo "No token, so the nightly was built and not published." >&2
echo "GITHUB_TOKEN is provided by the runner; an empty one means Actions is configured" >&2
echo "without it, and the job's contents: write permission is what asks for it." >&2
exit 1
fi
api="$FORGE/api/v1/repos/$REPO"
auth="Authorization: token $TOKEN"
# ◆ THE FIRST "id" IN THE DOCUMENT, AND NOT THE LAST ONE A GREEDY MATCH LANDS ON.
#
# This was `sed -n 's/.*"id":\([0-9]*\).*/\1/p'`, and the leading .* is greedy, so it walked past
# the release's own id to the last "id": in the response — which belongs to the embedded author
# object and is -2. [0-9]* then matched no digits at all and the answer was the empty string. A
# release had been created and the step failed saying none had.
#
# It broke both readings, and the other one silently: the existing-release lookup below got the
# same empty id, so the delete never fired, so the second run would have failed to create a
# release for a tag that already had one. A rolling channel that works exactly once.
#
# grep matches left to right and [0-9]\+ needs a digit, so the author's -2 cannot match at all
# and head -1 takes the release's. jq would say this in four characters and is not on this
# runner; anything parsing JSON with a regex should say which assumption it is making, and this
# one is: the release id is the first "id": with digits after it.
release_id() {
grep -oE '"id":[0-9]+' | head -1 | cut -d: -f2
}
# Deleted and recreated rather than updated in place. A rolling tag has to move, and moving one
# through this API is two calls with no atomic form either way — so the shape with the fewest
# states is to remove both and make them again. The window where no nightly exists is a few
# seconds and the client's answer to it is the same as to an unreachable forge: try later.
existing="$(curl -fsS -H "$auth" "$api/releases/tags/nightly" 2>/dev/null || true)"
if [ -n "$existing" ]; then
id="$(printf '%s' "$existing" | release_id)"
if [ -n "$id" ]; then
echo "Removing the previous nightly release $id"
curl -fsS -X DELETE -H "$auth" "$api/releases/$id" >/dev/null || true
else
echo "A nightly release exists and its id could not be read:" >&2
printf '%s\n' "$existing" >&2
exit 1
fi
fi
curl -fsS -X DELETE -H "$auth" "$api/tags/nightly" >/dev/null 2>&1 || true
created="$(curl -fsS -X POST -H "$auth" -H 'Content-Type: application/json' \
-d "$(printf '{"tag_name":"nightly","target_commitish":"%s","name":"Nightly %s","prerelease":true,"body":"Built from %s by CI, signed with the public nightly key. Installs beside the release build, never over it. See docs/adr/0014-android-updates.md."}' \
"$GITHUB_SHA" "$VERSION" "$GITHUB_SHA")" \
"$api/releases")"
release="$(printf '%s' "$created" | release_id)"
if [ -z "$release" ]; then
echo "Gitea accepted the release call and returned no id:" >&2
printf '%s\n' "$created" >&2
exit 1
fi
# The APK first and the manifest last, which is the ordering the client depends on: it reads the
# manifest and then fetches what the manifest names, so a manifest published before its APK is a
# few seconds in which every phone is told to download something that is not there yet.
names=""
for file in "$STAGED"/*.apk "$STAGED"/android-nightly.json; do
name="$(basename "$file")"
echo "Uploading $name ($(wc -c < "$file") bytes)"
curl -fsS -X POST -H "$auth" \
-F "attachment=@$file" \
"$api/releases/$release/assets?name=$name" >/dev/null
names="$names $name"
done
# Asked for back rather than assumed, because the failure this whole channel is exposed to is a
# release that exists and carries nothing — which is what a phone sees as a feed it can read and
# never update from, and what a person sees as a release page offering source tarballs. That is
# exactly what the run before this one left behind, and it reported success for every upload it
# never made.
published="$(curl -fsS -H "$auth" "$api/releases/$release")"
for name in $names; do
if ! printf '%s' "$published" | grep -qF "\"name\":\"$name\""; then
echo "The release was created but $name is not on it:" >&2
printf '%s\n' "$published" >&2
exit 1
fi
done
echo "Published nightly $VERSION with$names"
image:
name: api image
# Gated on the tests rather than parallel with them, which costs a few minutes of wall
# clock on every main commit and buys the thing worth having: no image reaches the
# registry from a commit whose tests were red. An image is not a build artefact anyone
# inspects — it is the thing that gets deployed.
needs: [build]
runs-on: [linux]
steps:
# Duplicated from the build job; see the comment there for why it cannot be factored
# out. Any change here has to be made in all three.
- name: ensure node and git
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
missing=""
command -v node >/dev/null 2>&1 || missing="$missing nodejs"
command -v git >/dev/null 2>&1 || missing="$missing git"
if [ -z "$missing" ]; then
echo "node $(node --version), git $(git --version)"
exit 0
fi
echo "Installing:$missing"
if command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends $missing
elif command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache $missing
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y $missing
else
echo "No apt-get, apk or dnf here, so node cannot be installed from inside the" >&2
echo "job. Point the runner's container.image at something that ships node." >&2
exit 1
fi
echo "node $(node --version), git $(git --version)"
# Warned about rather than failed on. Distributions pin their nodejs package to
# the release they shipped with — Ubuntu 24.04 still serves 18, which is past end
# of life and older than the runtime these actions declare. It generally runs
# them anyway, since act_runner uses whichever node is on PATH regardless of what
# the action asked for, so this is a note for when one of them misbehaves in a
# way that makes no sense, not a reason to stop a build that is probably fine.
major="$(node --version | sed 's/^v//; s/\..*//')"
if [ "$major" -lt 20 ]; then
echo "::warning::node $major is older than the runtime these actions target;" \
"give the runner an image with node 20 or newer if actions misbehave."
fi
# fetch-depth 0, and it is load-bearing rather than tidy. MinVer derives the version from the
# nearest v* tag, and checkout's default shallow clone has no tags at all — so it would not fail,
# it would quietly answer 0.0.0-alpha.0.N and every build would ship that. Velopack decides
# whether an installed client is out of date by comparing versions, which makes a plausible wrong
# answer here a client that never updates.
#
# Repeated in all three jobs, like the node preamble above and for the same reason. Change one
# copy, change all three.
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 0
# The daemon and the client are two different things to have, and this runner had only
# one of them. Testcontainers reaches Docker straight over /var/run/docker.sock from a
# .NET library, so every integration suite in the build job passed while `docker` was
# not a command here at all — which surfaced as exit 127 from the build step below,
# after the tags had been worked out and everything looked healthy.
#
# Only the CLI. The socket is already there and a daemon is already answering on it;
# installing an engine would start a second one beside the one being used.
- name: ensure the docker cli
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
if ! command -v docker >/dev/null 2>&1; then
echo "Installing the docker cli"
if command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache docker-cli
elif command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends docker.io
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y docker-cli
else
echo "No apt-get, apk or dnf here, so the docker client cannot be installed from" >&2
echo "inside the job. Add it to the runner's image." >&2
exit 1
fi
fi
docker --version
# buildx after the client, and wanted rather than required. Without the plugin
# `docker build` falls back to the legacy builder, which still produces the image
# and says on every run that it will not do so forever; with it the same command
# routes through BuildKit and the Dockerfile's independent stages stop being
# serialised. Alpine's docker-cli package does not carry it, which is why a job
# that had just been given a working client still built the deprecated way.
#
# A distribution with no package for it should get a warning and an image, not a
# failed release — so every branch here ends in `|| true` and the check below
# reports rather than exits.
if ! docker buildx version >/dev/null 2>&1; then
echo "Installing buildx"
if command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache docker-cli-buildx || true
elif command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq || true
$SUDO apt-get install -y --no-install-recommends docker-buildx || true
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y docker-buildx || true
fi
fi
if docker buildx version >/dev/null 2>&1; then
docker buildx version
else
echo "::warning::buildx is unavailable, so this image was built by the legacy" \
"builder Docker has deprecated. Add a buildx package to the runner image."
fi
# No docker/* actions here, deliberately. The build is single-architecture, so it
# wants the daemon this runner already has, a client, and BuildKit — all of which the
# step above arranges with two packages. What it does not want is QEMU, a builder
# instance to create and tear down, or a third-party action whose SHA has to be
# audited and re-pinned on a schedule. Adding linux/arm64 later is where that trade
# changes, and where setup-buildx-action starts earning its place.
- name: work out the tags
id: tags
env:
REGISTRY: registry-docker.dodotech.cloud
IMAGE: dodotech/dodossh-api
run: |
set -euo pipefail
repo="$REGISTRY/$IMAGE"
short="$(git rev-parse --short HEAD)"
# sha- prefixed, because a bare hex tag is ambiguous with a digest to both a human
# and a fair amount of tooling. This one is on every build and never moves, which
# makes it the only tag safe to pin a deployment to.
tags="$repo:sha-$short"
version="$short"
case "$GITHUB_REF" in
refs/tags/v*)
v="${GITHUB_REF#refs/tags/v}"
version="$v"
tags="$tags $repo:$v"
# The moving major.minor tag and :latest, but only for a release proper.
# v1.3.0-rc1 sorts after v1.2.9 and would otherwise take :latest with it,
# which is how a release candidate ends up on somebody's server.
case "$v" in
*-*) ;;
*)
tags="$tags $repo:${v%.*}"
tags="$tags $repo:latest"
;;
esac
;;
refs/heads/main)
tags="$tags $repo:main"
version="main-$short"
;;
esac
# The version MSBuild is allowed to see, which is not the same string as the one above.
# `version` is a docker tag and is `main-<short sha>` on a main build; handing that to
# -p:Version fails the publish with NETSDK1018. So this is set only when it is a real
# version, and the Dockerfile leaves the SDK default alone when it is empty.
assembly_version=""
case "$GITHUB_REF" in
refs/tags/v*) assembly_version="${GITHUB_REF#refs/tags/v}" ;;
esac
echo "tags=$tags" >> "$GITHUB_OUTPUT"
echo "version=$version" >> "$GITHUB_OUTPUT"
echo "assemblyVersion=$assembly_version" >> "$GITHUB_OUTPUT"
echo "created=$(date -u +%Y-%m-%dT%H:%M:%SZ)" >> "$GITHUB_OUTPUT"
echo "Tagging: $tags"
# Every value from a step output or the event goes through env rather than being
# interpolated into the script text. A git tag may contain a semicolon, and
# `${{ }}` is a textual substitution performed before the shell ever sees the line —
# so an interpolated tag name is a command the workflow agreed to run.
- name: build the image
env:
TAGS: ${{ steps.tags.outputs.tags }}
VERSION: ${{ steps.tags.outputs.version }}
ASSEMBLY_VERSION: ${{ steps.tags.outputs.assemblyVersion }}
REVISION: ${{ github.sha }}
CREATED: ${{ steps.tags.outputs.created }}
run: |
set -euo pipefail
args=()
for tag in $TAGS; do
args+=(--tag "$tag")
done
# --pull rather than whatever the runner happens to have cached: the base images
# are floating tags, and a runner that has held aspnet:10.0-noble-chiseled for a
# month is a month of unapplied CVE fixes shipping in every image built on it.
docker build \
--pull \
--file src/DodoSSH.Api/Dockerfile \
--build-arg "VERSION=$VERSION" \
--build-arg "ASSEMBLY_VERSION=$ASSEMBLY_VERSION" \
--build-arg "REVISION=$REVISION" \
--build-arg "CREATED=$CREATED" \
"${args[@]}" \
.
# Everything above runs on a pull request too. Building a fork's Dockerfile is safe —
# nothing is pushed and no credential is in scope — and it means a change that breaks
# the image fails on the PR rather than on main. Only these last two steps are held
# back, and the condition is on the event rather than on the branch so that a PR
# targeting main cannot reach them.
- name: log in to the registry
if: github.event_name != 'pull_request'
env:
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: |
set -eu
# Checked before use, because an unset secret is not an error anywhere upstream of
# here: an expression that resolves to nothing renders as the empty string, so
# docker is handed --username "" and answers with something about credentials,
# which sends people to the registry to debug a value that never left the
# settings page.
#
# Note for anyone editing this comment: an expression delimiter written literally
# here is interpolated even though this is a shell comment. The runner substitutes
# the whole script before any shell sees it, so an empty one fails the step with a
# parse error and no line number — which is how this very block broke the release
# it was added to protect.
#
# Reported by length, and never by value. Gitea masks known secret values in logs,
# but a mask is only as good as the runner's bookkeeping and a length answers the
# only question being asked: did anything arrive.
missing=""
[ -n "${REGISTRY_USERNAME:-}" ] || missing="$missing REGISTRY_USERNAME"
[ -n "${REGISTRY_PASSWORD:-}" ] || missing="$missing REGISTRY_PASSWORD"
if [ -n "$missing" ]; then
echo "Empty or unset:$missing" >&2
echo >&2
echo "Both come from repository secrets, which in Gitea are at" >&2
echo " Settings -> Actions -> Secrets" >&2
echo "and are a different page from Settings -> Actions -> Variables. A value" >&2
echo "added as a variable is invisible to the secrets context and arrives here" >&2
echo "as an empty string, which is exactly what this message means." >&2
exit 1
fi
echo "username: ${#REGISTRY_USERNAME} characters; password: set"
printf '%s' "$REGISTRY_PASSWORD" \
| docker login registry-docker.dodotech.cloud \
--username "$REGISTRY_USERNAME" --password-stdin
- name: push
if: github.event_name != 'pull_request'
env:
TAGS: ${{ steps.tags.outputs.tags }}
run: |
set -euo pipefail
for tag in $TAGS; do
docker push "$tag"
done
# The daemon is shared with every other job on this runner, and a credential left in
# ~/.docker/config.json outlives the job that created it. always(), so a failed push
# does not leave it behind.
- name: log out
if: always() && github.event_name != 'pull_request'
run: docker logout registry-docker.dodotech.cloud
# There is no job here that publishes the desktop client, and there is not going to be one. Two
# independent reasons, and both need saying because someone will fix one and think they are done.
#
# The smaller one is mechanical: vpk stamps and embeds the Setup.exe and Update.exe stubs with Windows
# tooling, and every job in this file is runs-on: [linux]. A Windows runner would answer that.
#
# The larger one is that a Windows runner would not answer the other. Velopack clients fetch from the
# release feed and do not verify a package signature when they apply it, so whoever can write a release
# on this repository can publish an update that every installed client downloads and runs. That is the
# same capability as the signing key, reached through a different door — and docs/adr/0011 rule 1 puts
# that capability on a machine which is not a runner, because a workflow secret is held by everyone who
# can change a workflow file. See docs/adr/0013-desktop-distribution-and-updates.md.
#
# What cuts a release is scripts/release-windows.ps1, run by a person. What this file does is prove the
# thing still builds and packages, which is the same division of labour the android job above already
# has: it packages an APK nobody installs, so that a link-time break fails here rather than later.