Files
DodoSSH/.github/workflows/ci.yml
T
jaap-jan 30a3edb1d4
ci / build and test (push) Successful in 1m54s
ci / android head (push) Failing after 3m54s
ci / api image (push) Successful in 44s
Build the phone in a container, because this runner cannot build it at all
The runner is Alpine, and .NET for Android does not work on musl. Not "needs setting up" — the SDK's
own MSBuild tasks pull glibc shared objects out of the workload pack into the build process, and a
musl-linked dotnet will not load one:

  error XARLP7000: Error relocating .../libZipSharpNative-3-3.so: __snprintf_chk: symbol not found

That is a glibc fortify symbol musl does not implement, reached through a DllImport rather than an
exec, so gcompat is no help: it gets a glibc *executable* started, which is a different problem. There
is no musl variant of the pack.

Everything this job did on the host to make Android work was therefore treatment of symptoms, mine
included. The missing aapt2 was present. The "unsupported version" was of a binary that had never run.
Both were this one sentence in a different accent, and the loader was the accent, not the sentence.

So the toolchain moves into build/android-build.Dockerfile — Microsoft's own sdk:10.0-noble plus a JDK,
the Android SDK and the workload — and the job keeps on the host only what the host is good at:
checkout, git, publishing. The daemon needed no arranging, since the image job already builds with it
and every Testcontainers suite reaches it over the socket. The image is tagged by the digest of the
Dockerfile that made it, so on a persistent runner every run after the first is a cache hit, and a
change to the toolchain is the only thing that buys a new one.

Built rather than pulled: a community image with the Android SDK already in it would put a stranger in
the path of a package this project signs and publishes. Eleven lines of apt and sdkmanager is the
cheaper trade.

Verified end to end in that image against a real clone rather than reasoned about, which after three
rounds of reasoning seemed the least I could do. Restore under locked mode, Release build, then
SignAndroidPackage:

  package: name='dev.dodotech.dodossh.nightly' versionCode='195' versionName='0.0.0-alpha.0.128'
  Signer #1 certificate SHA-256 digest: a9f067877724ddb0fdc04b637fbd5bfb97df753976616f100b48b522e132ba22

which is the keystore in build/. The versionName carries MinVer's height, so the csproj's target fires
in the container too, and the manifest the feed publishes parses back on the host.

Staging moves from RUNNER_TEMP to artifacts/, which is forced rather than preferred: the package is
made inside a container and read outside one, so it has to land under the bind-mounted checkout.
2026-08-04 23:28:47 +02:00

885 lines
45 KiB
YAML

name: ci
on:
push:
branches: [main]
# Release tags run the whole workflow, not only the image job that gates on it. A tag
# is the one build nobody is watching, so it is the last place to take the tests on
# trust.
tags: ['v*']
pull_request:
branches: [main]
# Actions are pinned to commit SHAs, not tags: a tag can be moved to point at new code,
# which would let a compromised action run with this workflow's permissions.
permissions:
contents: read
concurrency:
group: ci-${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
env:
DOTNET_NOLOGO: true
DOTNET_CLI_TELEMETRY_OPTOUT: true
DOTNET_SKIP_FIRST_TIME_EXPERIENCE: true
CI: true
jobs:
build:
name: build and test
runs-on: [linux]
steps:
# act_runner runs every `uses:` action with node inside the job container, and this
# runner's image has none — the run died on the first line of actions/checkout with
# "Cannot find: node in PATH". A `run:` step is shell rather than node, so this one
# can go first and unblock the rest.
#
# This is a workaround and the real fix is one line of the runner's own config.yaml:
# point container.image at an image that ships node, the way Gitea's default
# catthehacker/ubuntu:act-latest does. Kept anyway, because a pipeline that depends
# on a runner being configured correctly somewhere else fails confusingly when it is
# not, and because it costs nothing on a runner that is.
#
# git as well as node, and said in the step name rather than smuggled in: checkout
# shells out to git the moment node has loaded it, so an image thin enough to lack
# one usually lacks the other, and learning that costs a whole second CI round trip.
#
# Repeated verbatim in all three jobs, which is not laziness. It cannot be a local
# composite action — that would need the checkout it exists to unblock — and YAML
# anchors, which would deduplicate it, are rejected by GitHub's parser and would make
# this file portable to nothing. Change one copy, change all three.
- name: ensure node and git
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
missing=""
command -v node >/dev/null 2>&1 || missing="$missing nodejs"
command -v git >/dev/null 2>&1 || missing="$missing git"
if [ -z "$missing" ]; then
echo "node $(node --version), git $(git --version)"
exit 0
fi
echo "Installing:$missing"
if command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends $missing
elif command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache $missing
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y $missing
else
echo "No apt-get, apk or dnf here, so node cannot be installed from inside the" >&2
echo "job. Point the runner's container.image at something that ships node." >&2
exit 1
fi
echo "node $(node --version), git $(git --version)"
# Warned about rather than failed on. Distributions pin their nodejs package to
# the release they shipped with — Ubuntu 24.04 still serves 18, which is past end
# of life and older than the runtime these actions declare. It generally runs
# them anyway, since act_runner uses whichever node is on PATH regardless of what
# the action asked for, so this is a note for when one of them misbehaves in a
# way that makes no sense, not a reason to stop a build that is probably fine.
major="$(node --version | sed 's/^v//; s/\..*//')"
if [ "$major" -lt 20 ]; then
echo "::warning::node $major is older than the runtime these actions target;" \
"give the runner an image with node 20 or newer if actions misbehave."
fi
# fetch-depth 0, and it is load-bearing rather than tidy. MinVer derives the version from the
# nearest v* tag, and checkout's default shallow clone has no tags at all — so it would not fail,
# it would quietly answer 0.0.0-alpha.0.N and every build would ship that. Velopack decides
# whether an installed client is out of date by comparing versions, which makes a plausible wrong
# answer here a client that never updates.
#
# Repeated in all three jobs, like the node preamble above and for the same reason. Change one
# copy, change all three.
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 0
- uses: actions/setup-dotnet@a98b56852c35b8e3190ac28c8c2271da59106c68 # v6.0.0
with:
global-json-file: global.json
cache: true
cache-dependency-path: '**/packages.lock.json'
# Locked mode fails if packages.lock.json does not match the project files, so a
# dependency cannot change without the lock file change being reviewed.
- name: restore
run: dotnet restore DodoSSH.slnx --locked-mode
# The version comes from the tag, so on a tag build there are two ways to say the same number
# and they can disagree — a moved tag, a tag on the wrong commit, or a checkout that somehow
# still lost its history. What makes that worth a step of its own is that the disagreement is
# silent everywhere else: MinVer answers 0.0.0-alpha.0.N rather than failing, the build goes
# green, the package is cut, and the symptom arrives weeks later as clients that never update.
#
# -getProperty evaluates without building, so this costs a second and runs before the build.
- name: the tag and the version agree
if: startsWith(github.ref, 'refs/tags/v')
run: |
set -euo pipefail
tag="${GITHUB_REF#refs/tags/v}"
declared="$(dotnet msbuild src/DodoSSH.Client.App/DodoSSH.Client.App.csproj \
-getProperty:Version -nologo | tr -d '[:space:]')"
if [ "$tag" != "$declared" ]; then
echo "The tag says v$tag and MinVer computed $declared." >&2
echo >&2
echo "These come from the same place, so a mismatch means the checkout did not see the" >&2
echo "tag it is building — most likely fetch-depth, which must stay 0 in every job here." >&2
exit 1
fi
echo "v$declared"
# No `dotnet format --verify-no-changes` step. It re-analysed the whole solution before
# the build did, for minutes, to check something the build already checks: IDE0055 is an
# error in .editorconfig and TreatWarningsAsErrors is on, so a misformatted file fails
# the build step below on its own. What the separate step added was the ability to say
# so a few minutes earlier, and it cost more than that on every run.
# Avalonia's headless renderer is still Skia, and libSkiaSharp.so — which the layout
# test project copies into its own output — links against libfontconfig. Without that
# one library every test in DodoSSH.Client.App.Layout.Tests dies inside
# HeadlessUnitTestSession before it measures anything, and 69 tests fail for a reason
# none of their names or assertions mention.
#
# The library, and not fonts. Verified in a container where fc-list returns zero and
# the suite passes anyway: the application carries Inter itself, so nothing here needs
# a typeface installed — only the thing that would have gone looking for one.
- name: ensure skia's native dependency
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
if ldconfig -p 2>/dev/null | grep -q 'libfontconfig\.so\.1'; then
echo "libfontconfig present"
exit 0
fi
echo "Installing fontconfig"
if command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends libfontconfig1
elif command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache fontconfig
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y fontconfig
else
echo "No package manager here, so Skia cannot be given its dependency and the" >&2
echo "layout suite will fail to start. Add fontconfig to the runner's image." >&2
exit 1
fi
- name: build
run: dotnet build DodoSSH.slnx --no-restore --configuration Release
- name: test
run: dotnet test DodoSSH.slnx --no-build --configuration Release
# The one build shape nothing else here exercises: a self-contained RID-specific publish. Its
# failure mode is a restore graph or a native asset that resolves for net10.0 and not for
# net10.0/win-x64, which nobody would see until a person was halfway through cutting a release
# on a Windows machine. vpk can pack a Windows package from Linux; only signing needs Windows,
# and this repository signs nothing yet, so proving the publish here is worth the minutes.
#
# RestoreLockedMode=false for this command only, and it is not a loosened gate. The committed
# lock files are deliberately RID-free: declaring win-x64 on the desktop head writes a
# net10.0/win-x64 target into every project it references transitively, which includes
# DodoSSH.Contracts and DodoSSH.Crypto — and the API's Dockerfile restores those with no RID
# under locked mode, so the image job would fail NU1004. The gate is the locked solution
# restore at the top of this job, which is unchanged.
#
# It rewrites the lock files as it goes; nothing after this step reads them, and the runner's
# checkout is thrown away. The release script does the same thing and puts them back, because
# there the tree is somebody's working copy.
#
# After the tests rather than before them, so a red suite does not first spend a hundred
# megabytes pulling a win-x64 runtime pack. main and tags only, for the same reason: a break
# found by the person about to release is found early enough.
- name: the windows publish still resolves
if: github.event_name != 'pull_request'
run: >
dotnet publish src/DodoSSH.Client.App/DodoSSH.Client.App.csproj
--configuration Release --runtime win-x64 --self-contained true
-p:RestoreLockedMode=false
--output "$RUNNER_TEMP/win-x64-check"
# This includes the end-to-end suite, which starts PostgreSQL, Keycloak and an OpenSSH
# server through Testcontainers and runs the API as a child process — so it needs a
# Docker daemon and gets one here. That is why the tests run on ubuntu rather than
# macOS, whose runners have no daemon at all. Expect the Keycloak image pull to
# dominate a cold run.
# A failing run says only that tests failed and names a log file on a machine nobody
# has a shell on. Every diagnostic thing — exception type, message, stack — is inside
# that file, so a red build was a filename and a guess. This prints it.
#
# head rather than tail, and that is the whole trick: when a suite fails wholesale it
# writes one stack per test and they are all the same stack. The first is the one that
# explains it, and the last two hundred lines are the same sentence repeated.
- name: what actually failed
if: failure()
run: |
set +e
echo "=== distro ==="
cat /etc/os-release 2>/dev/null | head -3
id
echo "=== what Skia needs, and whether it is here ==="
# ldd against the copy the test project carries. Its unresolved rows are the
# answer whenever the layout suite dies in HeadlessUnitTestSession, and asking
# here beats inferring it from a managed TypeInitializationException.
skia="$(find tests -name 'libSkiaSharp.so' 2>/dev/null | head -1)"
if [ -n "$skia" ]; then
echo "$skia"
ldd "$skia" 2>&1 | grep -Ei 'not found|fontconfig|freetype' || echo " all resolved"
else
echo " libSkiaSharp.so was not in the test output at all"
fi
ldconfig -p 2>/dev/null | grep -ci fontconfig | sed 's/^/fontconfig entries in ldconfig: /'
echo "=== docker, for the Testcontainers suites ==="
docker version --format '{{.Server.Version}}' 2>&1 | head -2
echo "=== test logs ==="
find tests -path '*/TestResults/*.log' 2>/dev/null | while read -r log; do
echo "----- $log"
head -n 120 "$log"
done
exit 0
android:
name: android head
runs-on: [linux]
# Writes the nightly release at the end of the job; see that step for what the capability is and why
# it is acceptable here and nowhere else. Job-scoped, so nothing else in this file gains it.
permissions:
contents: write
steps:
# Duplicated from the build job; see the comment there for why it cannot be factored
# out. Any change here has to be made in all three.
- name: ensure node and git
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
missing=""
command -v node >/dev/null 2>&1 || missing="$missing nodejs"
command -v git >/dev/null 2>&1 || missing="$missing git"
if [ -z "$missing" ]; then
echo "node $(node --version), git $(git --version)"
exit 0
fi
echo "Installing:$missing"
if command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends $missing
elif command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache $missing
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y $missing
else
echo "No apt-get, apk or dnf here, so node cannot be installed from inside the" >&2
echo "job. Point the runner's container.image at something that ships node." >&2
exit 1
fi
echo "node $(node --version), git $(git --version)"
# Warned about rather than failed on. Distributions pin their nodejs package to
# the release they shipped with — Ubuntu 24.04 still serves 18, which is past end
# of life and older than the runtime these actions declare. It generally runs
# them anyway, since act_runner uses whichever node is on PATH regardless of what
# the action asked for, so this is a note for when one of them misbehaves in a
# way that makes no sense, not a reason to stop a build that is probably fine.
major="$(node --version | sed 's/^v//; s/\..*//')"
if [ "$major" -lt 20 ]; then
echo "::warning::node $major is older than the runtime these actions target;" \
"give the runner an image with node 20 or newer if actions misbehave."
fi
# fetch-depth 0, and it is load-bearing rather than tidy. MinVer derives the version from the
# nearest v* tag, and checkout's default shallow clone has no tags at all — so it would not fail,
# it would quietly answer 0.0.0-alpha.0.N and every build would ship that. Velopack decides
# whether an installed client is out of date by comparing versions, which makes a plausible wrong
# answer here a client that never updates.
#
# Repeated in all three jobs, like the node preamble above and for the same reason. Change one
# copy, change all three.
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 0
# A job of its own, because DodoSSH.Client.Android is deliberately not in DodoSSH.slnx.
# Adding it there would make the android workload and a full Android SDK a prerequisite of
# `dotnet build DodoSSH.slnx` for everyone — including the build job above, which needs
# neither and would grow several minutes for a head it does not compile.
#
# The cost of keeping it out is that nothing in the main job would notice this head
# breaking, which for a project sharing view models with the desktop one is a matter of
# when rather than whether. This job is that notice.
#
# ============ ◆ WHY THIS HEAD IS BUILT IN A CONTAINER AND THE OTHER TWO ARE NOT ============
#
# .NET for Android cannot build on this runner. Not "needs setting up" — cannot. The runner is
# Alpine, and the SDK's own MSBuild tasks load glibc shared libraries out of the workload pack
# into the build process, which a musl-linked dotnet will not do:
#
# error XARLP7000: Error relocating …/libZipSharpNative-3-3.so: __snprintf_chk: symbol not found
#
# There is no musl variant of that pack and no shim for a library loaded in-process — gcompat
# gets a glibc *executable* started and is no help whatsoever here. Everything this job used to
# do on the host to make Android work was therefore treatment of symptoms: an aapt2 reported
# missing while being present, an "unsupported version" of a binary that had never run. All of
# them were this one sentence in a different accent, and docs/platform-flags.md has the history
# so the next reader does not repeat the three rounds it took to hear it.
#
# So the toolchain moves into an image that is glibc, and the job keeps on the host only what the
# host is good at: checkout, git, and publishing. The daemon needs no arranging — the image job
# builds with it and every Testcontainers suite in the build job reaches it over the socket.
#
# There is no setup-dotnet here any more, and its absence is the point: nothing outside the
# container compiles anything, so a dotnet on the host would be a second toolchain with nothing
# to do and a version nobody checks.
- name: ensure the docker cli
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
# Only the client. The socket is already there and a daemon is already answering on it, so
# installing an engine would start a second one beside the one in use. The image job carries
# the same step and the same reasoning; change one, look at the other.
if ! command -v docker >/dev/null 2>&1; then
echo "Installing the docker cli"
if command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache docker-cli
elif command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends docker.io
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y docker-cli
else
echo "No apt-get, apk or dnf here, so the docker client cannot be installed from" >&2
echo "inside the job. Add it to the runner's image." >&2
exit 1
fi
fi
docker --version
# Tagged by the digest of the Dockerfile that made it, so the tag moves exactly when the toolchain
# does and never otherwise. This runner is persistent, so the first build pays for a JDK, an
# Android SDK and the android workload — several minutes — and every run after it is a cache hit
# that prints two lines. A change to that Dockerfile is what buys a new one.
- name: the android toolchain image
run: |
set -eu
tag="dodossh-android-build:$(sha256sum build/android-build.Dockerfile | cut -c1-16)"
docker build -t "$tag" -f build/android-build.Dockerfile build
echo "ANDROID_BUILD_IMAGE=$tag" >> "$GITHUB_ENV"
echo "$tag"
# The checkout is bind-mounted rather than copied in, so what the container writes — obj, bin, the
# signed package — is on the host the moment it exits and the publishing step can read it. The
# NuGet cache is mounted for the mirror-image reason: a container that started empty would fetch
# every package again on every run.
#
# The script arrives on stdin under `bash -s` rather than as `bash -c '…'`, which is not a style
# choice: the packaging step below has to match a versionName out of aapt2 with sed, and nesting
# that quoting inside a shell string is how a working command becomes a silently empty variable.
# A quoted heredoc passes the script through untouched, so what is written here is what runs.
#
# CI=true is handed over rather than assumed. Directory.Build.props turns ContinuousIntegrationBuild
# on when it is set — which is what normalises the source paths baked into the PDBs — and a
# container inherits none of the runner's environment unless it is given it.
- name: restore and build
run: |
set -eu
docker run --rm -i \
-v "$PWD:/build" \
-v "$HOME/.nuget/packages:/root/.nuget/packages" \
-e CI=true \
"$ANDROID_BUILD_IMAGE" bash -s <<'IN_CONTAINER'
set -euo pipefail
dotnet restore src/DodoSSH.Client.Android/DodoSSH.Client.Android.csproj --locked-mode
# Aapt2ToolPath, so the build takes aapt2 from the Android SDK in this image rather than the
# copy inside the workload pack. Both work here; using the SDK's makes it the same binary the
# packaging step reads the versionName back with, so the manifest the feed publishes is read
# by the thing that wrote it.
dotnet build src/DodoSSH.Client.Android/DodoSSH.Client.Android.csproj \
--no-restore --configuration Release \
-p:Aapt2ToolPath="$ANDROID_HOME/build-tools/$ANDROID_BUILD_TOOLS"
IN_CONTAINER
# Packaging rather than only compiling, because the two failures this head is most exposed to are
# both link-time: a native library with no android ABI, and a managed assembly that resolves for
# net10.0 but has nothing to dex. Neither shows up in a compile.
#
# ◆ THE NIGHTLY CHANNEL, WHICH IS AN INSTALLABLE APPLICATION AND NOT THE ONE. It has its own package
# id and is signed by a keystore committed to this repository in the open, so it can neither replace
# nor be replaced by the release channel — see the csproj, and ADR 0014. The APK this produces is
# meant to be installed; the release APK is cut from a v* tag by a person running
# scripts/release-android.ps1, on a machine that holds the key ADR 0011 rule 1 keeps off runners.
#
# No RuntimeIdentifier, where this step used to pin android-arm64. That produced the smallest
# possible build check and the least installable artefact: an arm64-only APK will not run on an
# x86_64 emulator, which is what most people testing a nightly actually have. Every supported ABI
# costs size on a package nobody ships to users.
#
# versionCode is the commit count, which is monotonic by construction and needs nobody to remember
# anything. It is not a version and is never displayed; versionName carries MinVer's full answer
# including the height, which is what tells two nightlies apart.
- name: package the nightly
id: nightly
run: |
set -euo pipefail
# Under artifacts/ rather than RUNNER_TEMP, and that is forced rather than preferred: the
# package is produced inside a container and read outside one, so it has to land somewhere
# both can see, which means somewhere under the bind-mounted checkout. .gitignore already
# excludes artifacts/* with two named exceptions, neither of which is this.
staged="artifacts/android-nightly"
rm -rf "$staged"
docker run --rm -i \
-v "$PWD:/build" \
-v "$HOME/.nuget/packages:/root/.nuget/packages" \
-e CI=true \
"$ANDROID_BUILD_IMAGE" bash -s <<'IN_CONTAINER'
set -euo pipefail
staged="artifacts/android-nightly"
mkdir -p "$staged"
code="$(git rev-list --count HEAD)"
dotnet build src/DodoSSH.Client.Android/DodoSSH.Client.Android.csproj \
--no-restore --configuration Release \
-t:SignAndroidPackage \
-p:DodoChannel=nightly \
-p:DodoNightlyVersionCode="$code" \
-p:Aapt2ToolPath="$ANDROID_HOME/build-tools/$ANDROID_BUILD_TOOLS"
apk="$(find src/DodoSSH.Client.Android/bin/Release -name '*-Signed.apk' | head -1)"
if [ -z "$apk" ]; then
echo "The package step produced no signed APK." >&2
exit 1
fi
# Read back out of the APK rather than recomputed, so what the feed advertises is what the
# bytes say. A versionName derived a second time in shell is a second implementation of the
# csproj's target, and the two would drift on the first change to either.
badging="$("$ANDROID_HOME/build-tools/$ANDROID_BUILD_TOOLS/aapt2" dump badging "$apk")"
name="$(printf '%s' "$badging" | sed -n "s/.*versionName='\([^']*\)'.*/\1/p" | head -1)"
if [ -z "$name" ]; then
echo "aapt2 reported no versionName for $apk." >&2
exit 1
fi
cp "$apk" "$staged/DodoSSH-nightly-$name.apk"
# The channel manifest, which is what the client reads and the whole reason the feed is
# machine-readable at all. versionCode is the comparison — it is the number Android itself uses
# to accept or refuse an install, so comparing anything else would let the client offer an
# update the platform then rejects. versionName is for the person reading the banner.
printf '{"versionCode":%s,"versionName":"%s","apk":"DodoSSH-nightly-%s.apk"}' \
"$code" "$name" "$name" > "$staged/android-nightly.json"
IN_CONTAINER
# Read back out of the manifest the container just wrote, rather than passed out of it. A
# container's stdout is the build log as well as its return value, so anything parsed from it
# is one stray MSBuild line away from being wrong.
name="$(sed -n 's/.*"versionName":"\([^"]*\)".*/\1/p' "$staged/android-nightly.json")"
if [ -z "$name" ]; then
echo "The container produced no usable manifest." >&2
exit 1
fi
echo "staged=$staged" >> "$GITHUB_OUTPUT"
echo "version=$name" >> "$GITHUB_OUTPUT"
ls -la "$staged"
# ◆ PUBLISHING IT, AND WHAT THAT CAPABILITY IS.
#
# Whoever can write a release here can put a build on every nightly phone, because the client fetches
# from this feed and Android's only check is that the signature matches — and this channel's key is
# in the repository for everybody. That is the same capability as the signing key, reached through a
# different door, and it is exactly what ADR 0011 rule 1 keeps off runners.
#
# It is acceptable here for one reason: this is not that channel. A nightly is signed by a key with
# no secrecy to lose, installs under its own package id, and cannot update the application anybody
# is trusting with their credentials. The release channel has none of this — no job, no token, no
# key on a runner — and the two are separate applications so that no mistake here can reach it.
#
# main only. A tag build must not touch this: a v* tag is the release channel's, and cutting it is a
# person's job.
- name: publish the nightly
if: github.ref == 'refs/heads/main'
env:
FORGE: https://git.dodotech.cloud
REPO: DodoTech/DodoSSH
TOKEN: ${{ secrets.GITHUB_TOKEN }}
STAGED: ${{ steps.nightly.outputs.staged }}
VERSION: ${{ steps.nightly.outputs.version }}
run: |
set -euo pipefail
if [ -z "${TOKEN:-}" ]; then
echo "No token, so the nightly was built and not published." >&2
echo "GITHUB_TOKEN is provided by the runner; an empty one means Actions is configured" >&2
echo "without it, and the job's contents: write permission is what asks for it." >&2
exit 1
fi
api="$FORGE/api/v1/repos/$REPO"
auth="Authorization: token $TOKEN"
# Deleted and recreated rather than updated in place. A rolling tag has to move, and moving one
# through this API is two calls with no atomic form either way — so the shape with the fewest
# states is to remove both and make them again. The window where no nightly exists is a few
# seconds and the client's answer to it is the same as to an unreachable forge: try later.
existing="$(curl -fsS -H "$auth" "$api/releases/tags/nightly" 2>/dev/null || true)"
if [ -n "$existing" ]; then
id="$(printf '%s' "$existing" | sed -n 's/.*"id":\([0-9]*\).*/\1/p' | head -1)"
[ -n "$id" ] && curl -fsS -X DELETE -H "$auth" "$api/releases/$id" >/dev/null || true
fi
curl -fsS -X DELETE -H "$auth" "$api/tags/nightly" >/dev/null 2>&1 || true
created="$(curl -fsS -X POST -H "$auth" -H 'Content-Type: application/json' \
-d "$(printf '{"tag_name":"nightly","target_commitish":"%s","name":"Nightly %s","prerelease":true,"body":"Built from %s by CI, signed with the public nightly key. Installs beside the release build, never over it. See docs/adr/0014-android-updates.md."}' \
"$GITHUB_SHA" "$VERSION" "$GITHUB_SHA")" \
"$api/releases")"
release="$(printf '%s' "$created" | sed -n 's/.*"id":\([0-9]*\).*/\1/p' | head -1)"
if [ -z "$release" ]; then
echo "Gitea accepted the release call and returned no id:" >&2
printf '%s\n' "$created" >&2
exit 1
fi
# The APK first and the manifest last, which is the ordering the client depends on: it reads the
# manifest and then fetches what the manifest names, so a manifest published before its APK is a
# few seconds in which every phone is told to download something that is not there yet.
for file in "$STAGED"/*.apk "$STAGED"/android-nightly.json; do
echo "Uploading $(basename "$file")"
curl -fsS -X POST -H "$auth" \
-F "attachment=@$file" \
"$api/releases/$release/assets?name=$(basename "$file")" >/dev/null
done
echo "Published nightly $VERSION"
image:
name: api image
# Gated on the tests rather than parallel with them, which costs a few minutes of wall
# clock on every main commit and buys the thing worth having: no image reaches the
# registry from a commit whose tests were red. An image is not a build artefact anyone
# inspects — it is the thing that gets deployed.
needs: [build]
runs-on: [linux]
steps:
# Duplicated from the build job; see the comment there for why it cannot be factored
# out. Any change here has to be made in all three.
- name: ensure node and git
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
missing=""
command -v node >/dev/null 2>&1 || missing="$missing nodejs"
command -v git >/dev/null 2>&1 || missing="$missing git"
if [ -z "$missing" ]; then
echo "node $(node --version), git $(git --version)"
exit 0
fi
echo "Installing:$missing"
if command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends $missing
elif command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache $missing
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y $missing
else
echo "No apt-get, apk or dnf here, so node cannot be installed from inside the" >&2
echo "job. Point the runner's container.image at something that ships node." >&2
exit 1
fi
echo "node $(node --version), git $(git --version)"
# Warned about rather than failed on. Distributions pin their nodejs package to
# the release they shipped with — Ubuntu 24.04 still serves 18, which is past end
# of life and older than the runtime these actions declare. It generally runs
# them anyway, since act_runner uses whichever node is on PATH regardless of what
# the action asked for, so this is a note for when one of them misbehaves in a
# way that makes no sense, not a reason to stop a build that is probably fine.
major="$(node --version | sed 's/^v//; s/\..*//')"
if [ "$major" -lt 20 ]; then
echo "::warning::node $major is older than the runtime these actions target;" \
"give the runner an image with node 20 or newer if actions misbehave."
fi
# fetch-depth 0, and it is load-bearing rather than tidy. MinVer derives the version from the
# nearest v* tag, and checkout's default shallow clone has no tags at all — so it would not fail,
# it would quietly answer 0.0.0-alpha.0.N and every build would ship that. Velopack decides
# whether an installed client is out of date by comparing versions, which makes a plausible wrong
# answer here a client that never updates.
#
# Repeated in all three jobs, like the node preamble above and for the same reason. Change one
# copy, change all three.
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 0
# The daemon and the client are two different things to have, and this runner had only
# one of them. Testcontainers reaches Docker straight over /var/run/docker.sock from a
# .NET library, so every integration suite in the build job passed while `docker` was
# not a command here at all — which surfaced as exit 127 from the build step below,
# after the tags had been worked out and everything looked healthy.
#
# Only the CLI. The socket is already there and a daemon is already answering on it;
# installing an engine would start a second one beside the one being used.
- name: ensure the docker cli
run: |
set -eu
SUDO=""
[ "$(id -u)" -eq 0 ] || SUDO="sudo"
if ! command -v docker >/dev/null 2>&1; then
echo "Installing the docker cli"
if command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache docker-cli
elif command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y --no-install-recommends docker.io
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y docker-cli
else
echo "No apt-get, apk or dnf here, so the docker client cannot be installed from" >&2
echo "inside the job. Add it to the runner's image." >&2
exit 1
fi
fi
docker --version
# buildx after the client, and wanted rather than required. Without the plugin
# `docker build` falls back to the legacy builder, which still produces the image
# and says on every run that it will not do so forever; with it the same command
# routes through BuildKit and the Dockerfile's independent stages stop being
# serialised. Alpine's docker-cli package does not carry it, which is why a job
# that had just been given a working client still built the deprecated way.
#
# A distribution with no package for it should get a warning and an image, not a
# failed release — so every branch here ends in `|| true` and the check below
# reports rather than exits.
if ! docker buildx version >/dev/null 2>&1; then
echo "Installing buildx"
if command -v apk >/dev/null 2>&1; then
$SUDO apk add --no-cache docker-cli-buildx || true
elif command -v apt-get >/dev/null 2>&1; then
$SUDO apt-get update -qq || true
$SUDO apt-get install -y --no-install-recommends docker-buildx || true
elif command -v dnf >/dev/null 2>&1; then
$SUDO dnf install -y docker-buildx || true
fi
fi
if docker buildx version >/dev/null 2>&1; then
docker buildx version
else
echo "::warning::buildx is unavailable, so this image was built by the legacy" \
"builder Docker has deprecated. Add a buildx package to the runner image."
fi
# No docker/* actions here, deliberately. The build is single-architecture, so it
# wants the daemon this runner already has, a client, and BuildKit — all of which the
# step above arranges with two packages. What it does not want is QEMU, a builder
# instance to create and tear down, or a third-party action whose SHA has to be
# audited and re-pinned on a schedule. Adding linux/arm64 later is where that trade
# changes, and where setup-buildx-action starts earning its place.
- name: work out the tags
id: tags
env:
REGISTRY: registry-docker.dodotech.cloud
IMAGE: dodotech/dodossh-api
run: |
set -euo pipefail
repo="$REGISTRY/$IMAGE"
short="$(git rev-parse --short HEAD)"
# sha- prefixed, because a bare hex tag is ambiguous with a digest to both a human
# and a fair amount of tooling. This one is on every build and never moves, which
# makes it the only tag safe to pin a deployment to.
tags="$repo:sha-$short"
version="$short"
case "$GITHUB_REF" in
refs/tags/v*)
v="${GITHUB_REF#refs/tags/v}"
version="$v"
tags="$tags $repo:$v"
# The moving major.minor tag and :latest, but only for a release proper.
# v1.3.0-rc1 sorts after v1.2.9 and would otherwise take :latest with it,
# which is how a release candidate ends up on somebody's server.
case "$v" in
*-*) ;;
*)
tags="$tags $repo:${v%.*}"
tags="$tags $repo:latest"
;;
esac
;;
refs/heads/main)
tags="$tags $repo:main"
version="main-$short"
;;
esac
# The version MSBuild is allowed to see, which is not the same string as the one above.
# `version` is a docker tag and is `main-<short sha>` on a main build; handing that to
# -p:Version fails the publish with NETSDK1018. So this is set only when it is a real
# version, and the Dockerfile leaves the SDK default alone when it is empty.
assembly_version=""
case "$GITHUB_REF" in
refs/tags/v*) assembly_version="${GITHUB_REF#refs/tags/v}" ;;
esac
echo "tags=$tags" >> "$GITHUB_OUTPUT"
echo "version=$version" >> "$GITHUB_OUTPUT"
echo "assemblyVersion=$assembly_version" >> "$GITHUB_OUTPUT"
echo "created=$(date -u +%Y-%m-%dT%H:%M:%SZ)" >> "$GITHUB_OUTPUT"
echo "Tagging: $tags"
# Every value from a step output or the event goes through env rather than being
# interpolated into the script text. A git tag may contain a semicolon, and
# `${{ }}` is a textual substitution performed before the shell ever sees the line —
# so an interpolated tag name is a command the workflow agreed to run.
- name: build the image
env:
TAGS: ${{ steps.tags.outputs.tags }}
VERSION: ${{ steps.tags.outputs.version }}
ASSEMBLY_VERSION: ${{ steps.tags.outputs.assemblyVersion }}
REVISION: ${{ github.sha }}
CREATED: ${{ steps.tags.outputs.created }}
run: |
set -euo pipefail
args=()
for tag in $TAGS; do
args+=(--tag "$tag")
done
# --pull rather than whatever the runner happens to have cached: the base images
# are floating tags, and a runner that has held aspnet:10.0-noble-chiseled for a
# month is a month of unapplied CVE fixes shipping in every image built on it.
docker build \
--pull \
--file src/DodoSSH.Api/Dockerfile \
--build-arg "VERSION=$VERSION" \
--build-arg "ASSEMBLY_VERSION=$ASSEMBLY_VERSION" \
--build-arg "REVISION=$REVISION" \
--build-arg "CREATED=$CREATED" \
"${args[@]}" \
.
# Everything above runs on a pull request too. Building a fork's Dockerfile is safe —
# nothing is pushed and no credential is in scope — and it means a change that breaks
# the image fails on the PR rather than on main. Only these last two steps are held
# back, and the condition is on the event rather than on the branch so that a PR
# targeting main cannot reach them.
- name: log in to the registry
if: github.event_name != 'pull_request'
env:
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: |
set -eu
# Checked before use, because an unset secret is not an error anywhere upstream of
# here: an expression that resolves to nothing renders as the empty string, so
# docker is handed --username "" and answers with something about credentials,
# which sends people to the registry to debug a value that never left the
# settings page.
#
# Note for anyone editing this comment: an expression delimiter written literally
# here is interpolated even though this is a shell comment. The runner substitutes
# the whole script before any shell sees it, so an empty one fails the step with a
# parse error and no line number — which is how this very block broke the release
# it was added to protect.
#
# Reported by length, and never by value. Gitea masks known secret values in logs,
# but a mask is only as good as the runner's bookkeeping and a length answers the
# only question being asked: did anything arrive.
missing=""
[ -n "${REGISTRY_USERNAME:-}" ] || missing="$missing REGISTRY_USERNAME"
[ -n "${REGISTRY_PASSWORD:-}" ] || missing="$missing REGISTRY_PASSWORD"
if [ -n "$missing" ]; then
echo "Empty or unset:$missing" >&2
echo >&2
echo "Both come from repository secrets, which in Gitea are at" >&2
echo " Settings -> Actions -> Secrets" >&2
echo "and are a different page from Settings -> Actions -> Variables. A value" >&2
echo "added as a variable is invisible to the secrets context and arrives here" >&2
echo "as an empty string, which is exactly what this message means." >&2
exit 1
fi
echo "username: ${#REGISTRY_USERNAME} characters; password: set"
printf '%s' "$REGISTRY_PASSWORD" \
| docker login registry-docker.dodotech.cloud \
--username "$REGISTRY_USERNAME" --password-stdin
- name: push
if: github.event_name != 'pull_request'
env:
TAGS: ${{ steps.tags.outputs.tags }}
run: |
set -euo pipefail
for tag in $TAGS; do
docker push "$tag"
done
# The daemon is shared with every other job on this runner, and a credential left in
# ~/.docker/config.json outlives the job that created it. always(), so a failed push
# does not leave it behind.
- name: log out
if: always() && github.event_name != 'pull_request'
run: docker logout registry-docker.dodotech.cloud
# There is no job here that publishes the desktop client, and there is not going to be one. Two
# independent reasons, and both need saying because someone will fix one and think they are done.
#
# The smaller one is mechanical: vpk stamps and embeds the Setup.exe and Update.exe stubs with Windows
# tooling, and every job in this file is runs-on: [linux]. A Windows runner would answer that.
#
# The larger one is that a Windows runner would not answer the other. Velopack clients fetch from the
# release feed and do not verify a package signature when they apply it, so whoever can write a release
# on this repository can publish an update that every installed client downloads and runs. That is the
# same capability as the signing key, reached through a different door — and docs/adr/0011 rule 1 puts
# that capability on a machine which is not a runner, because a workflow secret is held by everyone who
# can change a workflow file. See docs/adr/0013-desktop-distribution-and-updates.md.
#
# What cuts a release is scripts/release-windows.ps1, run by a person. What this file does is prove the
# thing still builds and packages, which is the same division of labour the android job above already
# has: it packages an APK nobody installs, so that a link-time break fails here rather than later.