The same application, the same Velopack and the same two-phase person-run release as Windows, with four things forced to differ. Signing is a precondition rather than an improvement: Gatekeeper refuses an un-notarized download outright instead of warning about it, so there was never the "unsigned for now" that ADR 0013 decision 8 argues for on Windows, and release-macos.sh refuses to start without the identities. The packaging split is narrower than it first looked, and the old claim at the foot of ci.yml is why it was worth checking rather than assuming. vpk cross-compiles when told to: 'vpk [osx] bundle' builds a real .app on any platform, and CI now publishes osx-arm64 and bundles it on every main and tag build, which is what catches a restore graph with no macOS native asset. There is no '[osx] pack' off a Mac, and that part is correct — pack drives codesign, notarytool and stapler, which exist nowhere else. The dylib signing loop in the script looks redundant beside vpk's own pass and is not. vpk signs with 'codesign --deep', which is the shape Apple documents as wrong for nested code, and platform-flags has recorded a notarization rejection that names no file since before any of this existed. Signing each native binary inside-out first leaves that pass nothing to get wrong. MacDeviceKeyStore reaches ADR 0007's conclusion through different hardware: a P-256 key in the Secure Enclave under an access control requiring user presence, so the platform enforces the gate rather than this process — which is the whole point of that ADR's amendment. The enclave holds no other kind of key, hence ECIES where Windows uses RSA-OAEP, and the shape that falls out is better than the Windows one: sealing needs only the public half and is silent, so only unlock prompts. IsSupported probes rather than infers, because three ordinary Macs answer no — an Intel machine without a T2, one with no login password, and every unsigned development build, since enclave keys need a signing identity. Two decisions worth stating because they are reversible. arm64 only: a second channel is small work and nobody here has an Intel Mac to walk Phase 18 on, and an x64 package would be the only artefact in this repository reaching users unverified. And the pack id stays DodoSSH.Desktop even though vpk names the bundle after it, so /Applications holds DodoSSH.Desktop.app: decision 2's reasoning binds harder here, because a pack id of DodoSSH would put Velopack's install root on top of ClientPaths.DataDirectory and let an uninstall take the user's un-synced outbox with it. CFBundleDisplayName puts the product name back in front of a person. Measured rather than assumed, since none of it is obvious: the publish and the bundle were both run, LSMinimumSystemVersion is 12.0 because that is the minos in the apphost's own LC_BUILD_VERSION, and vpk copies a custom Info.plist verbatim with no substitution at all — which is why the plist is a template the script renders and not a committed file. What is not done is the half that needs the hardware. There is no macOS runner, so nothing past "it bundles" has ever run. Phase 18 is the whole of the verification, and the two checks most likely to fail are the terminal against WKWebView and the enclave interop, neither of which has executed once.
66 KiB
Platform flags
Things known or suspected to behave differently outside Windows, plus deployment gotchas that have already cost time once. Development is Windows-first, but the full test suite now runs on Linux in CI on every change, so a Linux claim here is usually a measurement now rather than a suspicion. Anything marked unverified has not run on the platform in question and must not be assumed to work.
macOS now builds and packages, and has still never run. The distinction matters more here than
anywhere else on this page, because the two halves are verified in completely different places. The
build is measured on every main and tag build: CI publishes osx-arm64 and runs vpk [osx] bundle
on a Linux runner, which is enough to catch a restore graph with no macOS native asset and an .app
that will not compose. Everything past that — whether the window draws, whether the terminal's
loopback WebSocket reaches WKWebView, whether the Secure Enclave holds a device key — is verified
only by a person walking Phase 18 of manual-checks.md on a Mac, because there
is no macOS runner in CI. Treat every macOS runtime claim below as unverified unless it says
otherwise.
Each entry says what the risk is, why it matters, and what to do about it. Delete an entry when it has been verified or made moot — not when it merely stops being convenient.
Cryptography
ChaCha20Poly1305.IsSupported is false on macOS, and on Windows builds before 10.0.20142.
This is why the client uses NSec (libsodium) rather than the BCL for content encryption; see
docs/crypto.md §1. Already mitigated — but if a BCL AEAD path is ever added as a fallback it
must gate on IsSupported rather than assuming availability, or the client will fail to open
any vault on macOS.
The Secure Enclave holds P-256 keys and nothing else, which is why MacDeviceKeyStore wraps the
device key with ECIES rather than with the RSA-OAEP the Windows store uses. It will not hold an RSA
key at any size, so this is not a preference. The useful consequence is that the macOS shape is
better than the Windows one: SecKeyCopyPublicKey works on an enclave key without prompting, so
registering a device is silent and only unlock asks — where Windows raises a dialog at key creation
too. Unverified: no enclave call in this repository has ever run.
Three ordinary Macs have no usable enclave, and IsSupported probes rather than infers for that
reason: an Intel machine without a T2, a machine with no login password set, and — the one that
surprises people — any build that is not code signed, because enclave key creation needs a
signing identity. So dotnet run correctly offers no device key at all. Do not "fix" this by
checking the OS instead; the offer would then put a wrap on the server that nothing can ever open.
Argon2id timings are measured on one Windows machine only. 256 MiB with t=4 took 323 ms here.
The floor and ceiling in EnrollmentLimits were chosen against that number. Unverified
elsewhere: recalibrate on the slowest target platform before recommending a default profile,
because a cost that is comfortable on a desktop can make unlock unusable on a low-power laptop —
and the parameters are stored per user at enrollment, so a bad default is a per-user migration.
libsodium ships native binaries per RID. This complicates single-file and AOT publishing, and
on macOS every native library (libsodium, libSkiaSharp, libHarfBuzzSharp, libe_sqlite3)
must be signed individually with --options runtime --timestamp before the bundle is signed,
or notarization fails with an error that does not name the offending file. Mitigated in
scripts/release-macos.sh, which signs every .dylib and createdump in a loop before vpk touches
anything — vpk's own pass uses codesign --deep, which is the shape Apple documents as wrong for
nested code and is the likeliest source of that unnamed rejection. The loop looks redundant next to
--deep and is not; do not delete it because a release once succeeded without it.
Desktop client
The local pane's roots bar is built differently per platform, and has to be. On Windows it is the
ready drives, from DriveInfo.GetDrives. On Unix that same call answers with every mount the kernel
holds — around forty on an ordinary laptop, counting /proc, /sys/fs/bpf, one per installed snap and
/run/user/1000/doc — and the transfers screen draws a button per root, so the bar ran to roughly five
thousand pixels inside an eight-hundred pixel window. Fixed in LocalDirectory.Roots, which on Unix
returns the root, the user's home, and whatever is mounted under /run/media/<user>, /media, /mnt
or /Volumes. Do not try to filter GetDrives instead: DriveType reports Fixed for / and /home
but also for every squashfs snap, for efivarfs and for tracefs, while /boot/efi comes back
Removable, and DriveFormat would need a hand-kept list of every virtual filesystem Linux may grow.
Found by the layout suite on its first Linux run, which is the argument for that suite existing.
The WebView runs on Windows. Avalonia.Controls.WebView 12.0.1 (MIT, no licence key) hosts the
terminal page: WebView2 launches, navigates to the loopback page, runs its JavaScript and completes the
WebSocket handshake. Verified by observing an established TCP connection from msedgewebview2 to the data
plane port.
Note precisely what that evidence covers, because it was once stretched to cover more: every clause above is about the process and the socket. It says nothing about how the control composites with Avalonia-drawn content, which is the axis on which it does not behave like an ordinary control — see the next entry.
A native child window cannot be covered by Avalonia content, on any platform that hosts it windowed.
NativeWebView attaches a real Win32 child HWND through NativeControlHost — on Windows the backend
creates a WS_CHILD holder window and SetParents WebView2's HWND into it — and a child window paints
above everything its parent draws, whatever the visual tree's z-order says. This is by design and
acknowledged upstream: "NativeControlHost places native controls over Avalonia content just like WPF one
does. So it suffers from the same airspace problem" (Avalonia's maintainer,
#6605, still open). Reproduced in a 60-line standalone
app with no DodoSSH code: a 340,* grid, a NativeWebView in column 1 and an opaque Border as a later
Panel sibling renders the overlay sliced dead on x=340.
Layering a screen over the terminal therefore does nothing: the WebView's rectangle stays on top. In this shell that sliced the setup and unlock cards at the terminal column's left edge, put every one of their buttons inside the WebView's rectangle at the window's default width — so the flow could only be completed by keyboard — and handed Win32 focus to WebView2 on any click in that region, which makes a text box stop accepting keystrokes with no visible cause. That last symptom is the focus asymmetry documented further down, not a separate fault: focus crosses into the WebView readily and does not come back on its own.
The fix is to collapse the control, not to cover it: IsVisible="{Binding IsUnlocked}" on the
NativeWebView. That is safe, and this is the part worth recording, because the opposite was asserted here
for a while:
-
NativeControlHostcreates the native attachment from attach to the visual tree, not from layout and not from visibility. ItsUpdateHostnever readsIsEffectivelyVisible; onlyTryUpdateNativeControlPositiondoes, choosingHideWithSizeoverShowInBounds. -
NativeWebViewstashes aSourceassigned before its adapter exists and replays it once created, so navigation is never lost to ordering. The shell already depends on that replay. -
So a collapsed WebView still starts WebView2, still loads the page and still lets the renderer attach its socket. Measured on Windows in a harness mirroring the data plane's handshake, with
IsVisible=falseset before the window was ever shown: adapter created,GET /, then the WebSocket 101 sent — the momentRendererAttachedfires — followed by frames arriving over the socket, all while hidden. A cold WebView2 profile behaves the same. Revealing it recomputes bounds within about 7 ms, on oneResizeObservercallback, over the same socket.It must be
IsVisible, not removal from the tree. Detaching runsDestroyNativeControland takes the whole WebView2 process tree with it, so conditional content or a template swap would pay a cold start on every unlock. Hiding merely doesSetWindowPos(holder, …, SWP_HIDEWINDOW). NegativeMarginalso works as a runtime toggle;RenderTransformdoes not, becauseNativeControlHostnever watches it.Note the earlier version of this bullet cited "35
msedgewebview2processes" as the confirmation. A process count cannot show that a socket was accepted — it is the same shape of mistake as the one described below, one level down.
The previous version of this entry claimed the reverse — that hiding it would mean never realising it — and
cited the msedgewebview2 connection as verification. That observation was made while the overlay was
showing but, because of the airspace behaviour above, the WebView was in fact uncovered and in plain view.
It confirmed only that a visible WebView is realised, which nobody disputed, and could not discriminate
the case it was attached to. A process-level check cannot verify a rendering claim; that needs a
screenshot, and this defect shipped because one was never taken.
What the first connection after unlocking actually depends on is the await workspace.WaitForRendererAsync(cancellationToken) in VaultViewModel.ConnectAsync, because
TerminalDataPlane.SendAsync drops frames when no renderer is attached rather than queueing them. That
await is the invariant; the control's visibility is not.
It is now bounded — TerminalWorkspaceOptions.RendererTimeout, 15 s, plus the command's own token —
because whether the renderer attaches at all depends on a runtime this application does not install. A
missing or policy-blocked Evergreen runtime, or an AppContainer that cannot reach loopback, previously
left Connect waiting forever with IsBusy stuck and nothing on screen to explain it. The gate is
unchanged; only the wait is. Why 15 s and not less: attaching is near-instant in the normal case (the page
attaches while the unlock screen is still up), but a cold WebView2 profile creates a user-data directory
and starts its process tree first, and reporting a broken runtime to someone whose runtime was merely slow
is the worse error. The timeout is caught in VaultViewModel and reported as a message naming WebView2,
because TimeoutException.Message is "The operation has timed out" and names nothing.
Hiding the WebView does not pause it. With the holder window hidden, the page keeps
visibilityState: "visible" and requestAnimationFrame keeps firing at roughly 115/s — Chromium does not
treat a hidden child HWND as a hidden page. That is why the handshake completes while collapsed, so it is
load-bearing rather than merely wasteful, but it means a locked DodoSSH is still animating a full-size
off-screen page. Worth revisiting if idle power ever matters.
A degenerate pane size reaches the remote pty. The fit addon floors its proposal at 2 columns by 1 row
rather than refusing, so any path that fits a terminal with almost no viewport sends window-change for a
2x1 window and permanently mangles the wrapped scrollback. Reachable today by minimising, and — once splits
land — by dragging a splitter to the edge. terminal.js now skips the fit below 40 px in either axis.
Related, and now fixed: the conflict log was an ItemsControl with no ScrollViewer and no MaxHeight on
an Auto row, so enough conflicts squeezed the row below it toward nothing. It survived that long because
it lived in MainWindow.axaml, which no test can lay out. Moving it into HostsScreen.axaml — a
UserControl, and therefore measurable — is what surfaced it; it now has both, and
TheHostsScreenFitsWithAConflictLogTooLongToShow fails without them.
Keyboard focus crosses into the WebView by itself and does not come back. This is the asymmetry to
know; the connect-focus bug that led here was only its first symptom. Measured on Windows with a harness
that reports GetFocus(), the class name of the window holding it, and the page's own
document.hasFocus() at each step.
- Into the page: nothing custom is needed.
NativeWebViewoverridesFocusableto true and itsOnGotFocuscalls the adapter'sFocus(), which on Windows isICoreWebView2Controller::MoveFocus(PROGRAMMATIC). A plain AvaloniaTerminal.Focus()therefore moves real Win32 focus to theChrome_WidgetWin_1child and the page reportshasFocus: true. NoSetFocusP/Invoke and no COM work — the package version of this entry that assumed otherwise was wrong. The control also replays aFocus()that arrived before its adapter existed, and re-asserts itself: while it holds Win32 focus itsGotFocushandler pulls Avalonia's logical focus back onto the control. Worth stating positively, because the reasonable guess before measuring — that crossing into a child HWND must needSetFocus— is the wrong way round: it is the return trip that needs it. - Out of the page: the package does nothing at all.
OnLostFocuscalls the adapter'sResignFocus(), and on Windows that method is empty. SosomeTextBox.Focus()moves Avalonia's focused element while Win32 focus stays on WebView2: a text box with a caret that silently receives nothing.Window.Activate()andWindow.Focus()were both measured and neither recovers it. The hand-back has to beSetFocus(topLevelHwnd)— seeViews/NativeKeyboardFocus.cs. A real mouse click does recover it, because Avalonia's window sets focus on pointer input, which is exactly why this is invisible to anyone who clicks before typing. - Collapsing the control does not release the keyboard. With
IsVisible=falsethe holder window is hidden but Win32 focus stays on it — measured as focus held by a window reportingvisible=False, with Avalonia's focused element becoming(none). So locking the vault after touching the terminal left the unlock passphrase box eating keystrokes. The lock path now hands the keyboard back and focuses that box. Focus()on a collapsed control is a no-op and is not replayed on reveal. Order matters: reveal, then focus. Focus does survive a lock/unlock cycle when done that way.- There is no Tab-out. The package subscribes
ICoreWebView2Controller::add_MoveFocusRequestedand its handler body is empty, so WebView2's request to move focus off itself is discarded; xterm eats Tab anyway. The way out isCtrl+Shift+F6, intercepted interminal.jsand sent to the host as a web message — measured arriving verbatim inWebMessageReceivedEventArgs.Body. It has to be handled in the page, because once the child window owns Win32 focus Avalonia sees no key events and noKeyBindingcould fire. Not Escape, and not a bare F6: both are keys a TUI legitimately binds, and Ctrl+Shift is the range terminal emulators conventionally keep for themselves.
On the phone the same asymmetry arrives through a button, and the fix is one property. The terminal's
accessory row — Ctrl, Esc, Tab, the arrows, and the two text-size keys — is a set of ordinary Avalonia
buttons over a NativeWebView. An ordinary button takes focus on tap, which takes it off the WebView, and
OnLostFocus then calls the adapter's ResignFocus(). So pressing Tab handed the terminal one byte and
took the keyboard away from it: everything typed afterwards on a hardware keyboard went nowhere.
The symptom is what makes it worth an entry. The row goes on working — its keys are pressed rather than
typed into — so what a user sees is a terminal that answers the buttons and ignores the keyboard, which
reads as the session having died rather than as a focus problem. Focusable = false is what a toolbar
button is, and it means the focused element never changes, so nothing resigns and nothing has to be handed
back. Every button on that row carries it; check 11.10 is the measurement.
None of this is covered by a test, and cannot be here: headless Avalonia has no native window, so a headless test renders and focuses correctly and would confirm the wrong belief. What the suite covers is the plumbing that drives it — that connecting asks for focus once per session, that a failed connect does not, and that locking stops the forwarding.
The lock/unlock cycle does not resize the pane at all, and the 40 px guard is not what makes it safe.
Measured on Windows with a live shell, against a real sshd in a container, in a harness mirroring
MainWindow.axaml's 340,* grid: with the NativeWebView collapsed by IsVisible=false, the page still
reports paneWidth: 840, paneHeight: 760, unchanged cols/rows, and visibilityState: "visible".
Hiding is SetWindowPos(holder, …, SWP_HIDEWINDOW), which does not resize the holder, so no
ResizeObserver callback fires, no fit runs, and no window-change reaches the remote — before,
during or after the cycle. stty size on the remote answered 50 118 both before locking and after
unlocking, and the renderer's own buffer came back byte for byte, wrapped lines included.
The guard's irrelevance here was established rather than assumed: the same run with
MINIMUM_FITTABLE_PIXELS patched to 0 — the guard fully disabled — produced an identical clean result.
So the guard is still worth keeping for the paths it was written for, minimising and a splitter dragged to
the edge, but it is not on the lock path and must not be cited as the reason locking is safe. It was
described that way when it landed.
Two further results from the same harness, both about the deliberate decision that shells outlive a lock
(README, MainWindowViewModel.LockAsync):
- A collapsed WebView is not typed into. With the harness confirmed as the foreground window and all
twelve injected
SendInputevents accepted, not one character of the probe reached the remote pty, and aCtrl-Uafterwards answeredBEL— nothing was sitting in the remote's line editor either. So the lock screen is a real input barrier even though the session behind it is live, and that is what makes surviving the lock defensible rather than merely convenient. The mechanism is not what this run concluded: it read the result as a hiddenWS_CHILDwindow being ineligible for keyboard focus, but the focus entry above measured Win32 focus still held by the hidden holder window, and the lock path now moves the keyboard off it deliberately. Take the barrier as measured here and the reason from there — which also means the barrier is something the lock path maintains, not something the platform guarantees. - The session survives the cycle in the real control, not only in tests.
LiveSessionCountwas 1 before, during and after, and the shell accepted a command again immediately on unlock.
Suspected, seen once, not reproduced: on the first run — before the harness learned to wait for the
window's scale to settle — the window opened at 2558x1367 px and the page reported a 2202x1328 pane
(312x88 characters) for a window 1180 logical units wide, which looks like physical pixels arriving where
CSS pixels were expected. A later re-push to 1177x672 then reflowed the wrapped line and split it in two.
Both events straddled a DPI settle rather than the lock, and three later runs at RenderScaling 1.00
never showed it. If a user reports mangled scrollback after moving the window between displays of
different scale, start here.
WebView2 will not initialise when the host executable sits under a very long path.
CreateCoreWebView2Environment fails with COMException 0x80080005 CO_E_SERVER_EXEC_FAILURE ("Server
execution failed") and the terminal never appears. Hit while building the harness above: the same binary
that failed from a ~230-character directory ran first time from %TEMP%\h. The exact threshold was not
established and the mechanism is unconfirmed — the user data folder is created beside the executable by
default and the browser process is launched with paths derived from it, so MAX_PATH is the obvious
suspect.
Measured for the packaged layout, so this stops being a worry and becomes a number. Velopack installs to
%LOCALAPPDATA%\DodoSSH.Desktop\current\, and …\AppData\Local\DodoSSH.Desktop\current\DodoSSH.exe is
64 characters against the ~230 that reproduced the failure — about 180 characters of headroom, and a
40-character corporate username adds 35 of them back. The shipped installer is not at risk. Two things
would reopen it and neither is in the plan: a self-extracting single-file publish, whose native libraries
land under a hashed temp path, and %LOCALAPPDATA% folder-redirected to a deep UNC path in a domain.
Manual check 16.2 measures it on the real machine rather than trusting this paragraph.
WebView2's user data folder must be kept out of the install directory. It defaults to a directory
beside the host executable, which under Velopack is inside current\ — and current\ is replaced by
every update. Left alone, the browser profile would be destroyed on each one, so the first connect after
every update would pay a cold WebView2 start: a fresh user-data directory and a new process tree, which is
the slow path RendererTimeout's fifteen seconds was sized for, arriving at the exact moment somebody is
most ready to believe the update broke the terminal. Program.Main sets WEBVIEW2_USER_DATA_FOLDER to
%LOCALAPPDATA%\DodoSSH\WebView2 — under the profile directory, which Velopack never touches. Check 16.8
is what would notice it regressing, and it is worth having because the symptom is "slow but working", which
gets dismissed as a fluke.
The install root and the profile directory must not be the same folder. Velopack removes
%LOCALAPPDATA%\<packId> entirely on uninstall, and ClientPaths puts cache.db (plus -wal and -shm),
settings.json and device.key in %LOCALAPPDATA%\DodoSSH. So a pack id of DodoSSH — the obvious
choice — would have made the uninstaller delete the vault cache and the outbox of changes not yet pushed,
silently, which is the thing the application will not do without a counted confirmation. The pack id is
DodoSSH.Desktop for that reason and no other; --packTitle supplies the name people see, so nothing is
lost. Do not "tidy" it. See ADR 0013 and check 16.9.
The Windows app manifest must declare a supportedOS list. Without it the process reports a
downlevel Windows version and Avalonia's native control host fails outright — "Unable to create child
window for native control host" — so the WebView, and therefore the terminal, does not start at all.
[STAThread] on Main is equally mandatory: WebView2 checks the apartment state and refuses to
initialise on an MTA thread.
WebView2 spawns a process tree, not a process. Around 35 processes were observed for one embedded view. That is the concrete reason the design uses one WebView hosting N terminals rather than one per tab: twenty tabs would mean twenty of those trees.
The Linux WebView needs ExperimentalOffscreen, or the terminal is blank. This entry used to predict
that WPE (libwpewebkit-2.0) would be the Linux backend and be too rarely installed; the prediction was
wrong in its details and right about the outcome. Measured on Fedora 44 with
Avalonia.Controls.WebView 12.0.1, by a spike that hosts a NativeWebView and reads AdapterInfo:
- The adapter is WebKitGTK 2.52.5, not WPE, and it reports
IsSupported = True. Fedora packages no WPE WebKit at all —dnf search wpereturns a computer-algebra package and nothing else — so the WPE path is not merely rare there, it is unavailable. - In its default mode that adapter reports
SupportedScenarios = NativeDialog: a window of its own, and nothing that can be hosted in place. The same under X11 and under Wayland, so this is the adapter's answer rather than a session-type problem. - Setting
ExperimentalOffscreenonGtkWebViewEnvironmentRequestedEventArgschanges the same adapter's answer toOffscreenRenderer— the compositor-drawn mode, which is whatNativeWebViewCompositorHost(mentioned below as an unknown) exists to host.MainWindowsets it; seeOnTerminalEnvironmentRequested, which is a no-op on Windows and macOS by type rather than by OS check. - In both modes the page loads and
InvokeScriptanswers, which is the trap: the failure has no diagnostics. Everything except the pixels works, so it reads as the terminal being broken rather than as the host having nowhere to draw.AdapterInfo.SupportedScenariosis the thing to look at first.
Still unverified: whether the offscreen mode actually paints, and how it behaves for input, IME and
resizing. The spike could not answer it — an XWayland root capture is black under a Wayland compositor
and RenderTargetBitmap does not capture a compositor surface — so it needs eyes on a running client.
macOS 15 remains untested entirely.
ITerminalHost was supposed to be the seam that keeps a backend swap cheap, and it is declared but not
implemented — nothing in the application uses it, and the view navigates NativeWebView.Source directly.
Swapping backends today means editing MainWindow.axaml and its code-behind. That is a small job, but do
not plan around a seam that is currently only a file.
Avalonia.Diagnostics has no 12.x release (latest is 11.3.18), so the developer tools overlay is
unavailable on Avalonia 12. Development-only, so nothing ships differently — but debugging a layout
problem currently means reasoning rather than inspecting.
The xterm bundles are vendored, not built. @xterm/xterm 6.0.0 with the fit and webgl addons, all
MIT, committed as UMD bundles under WebAssets/vendor and embedded as Avalonia resources. No npm or
esbuild step, so a clean clone builds with the .NET SDK alone. The cost is that upgrades are a manual
re-download; the licence and versions are recorded here so that stays visible.
Bracketed paste is the renderer's to decide, and it is why snippets go through a frame. xterm tracks
\e[?2004h from the remote's own output and Terminal.paste(text) wraps the text in paste markers only
when the mode is on — which is what makes a shell treat embedded newlines as text rather than as "run
this". The host process cannot make that decision: TerminalDataPlane moves opaque bytes and never parses
output, so writing a snippet straight into the pump would mean guessing, and guessing wrong executes every
line of a multi-line command. Hence TerminalServerOpcode.Paste. Two consequences worth keeping:
- The Enter for a snippet marked as running goes through
Terminal.input('\r'), outside the wrapper. A\rappended to the pasted text is bracketed with it and arrives as a literal character, so nothing runs. - Against a remote with bracketed paste off — a raw
sh, or a session inside an editor — a multi-line snippet does run line by line, and nothing can prevent that. It is a property of the terminal protocol, not of this client, which is why the screen says "types this into whatever is there".
Both methods were confirmed present on the public API of the vendored @xterm/xterm 6.0.0 bundle before
being written against; neither is reachable from any test in this repository, so they are in
docs/manual-checks.md as checks 3.6–3.8.
SSH.NET's window-change is verified working as of 2025.1.0 — resolved, not a flag.
ShellStream.ChangeWindowSize(columns, rows, width, height) exists and the remote genuinely
observes it: PtyAndResizeSpikeTests reads stty size back from a real sshd after resizing, and
repeated resizes each take effect. The IChannelSession fallback is not needed. That suite stays
in place as a regression guard, because an upgrade that silently stopped sending the request would
present as wrapped output only after a resize — easy to misattribute to the terminal emulator.
ShellStream.Write buffers and requires an explicit Flush. Without one a keystroke is accepted,
reported as written, and never reaches the remote — the terminal displays output perfectly and simply
stops responding to input. SSH.NET's own WriteLine flushes, which is why a spike that used it never
hit this. SshNetShellSession.WriteAsync now flushes per write; batching would be wrong anyway, since
a terminal has to put a keystroke on the wire immediately.
ShellStream does not override ReadAsync. The base Stream implementation therefore runs
the blocking Read on a thread-pool thread, so every open session parks one thread for as long as
it is idle. Fine for the handful of tabs M1 targets; revisit before advertising many concurrent
sessions, since the fix is either an upstream change or driving IChannelSession directly.
A passphrase supplied for an unprotected private key is silently ignored, not refused.
PrivateKeyFile(stream, passphrase) on an unencrypted PKCS#1 RSA key loads it and the connection
authenticates exactly as if no passphrase had been given — measured against a real sshd in
KeyAuthenticationTests.APassphraseOnAnUnprotectedKey_IsIgnoredRatherThanRefused, which was written
expecting the opposite and corrected to match. Two consequences, and the second is the one that bites: a
stray passphrase does no harm, so nothing downstream needs to defend against it; but equally nothing
downstream will report one, so if a user swears they set a passphrase and the key opens without it, no
error will ever say so. Only established for that armour and that algorithm; whether the OpenSSH format's
none cipher path behaves the same way is untested. SshKeySecret.Passphrase still normalises an empty
string to null, for the reasons stated there — one representation of one state — and not for this.
SSH.NET cannot share one connection between SshClient and SftpClient. A shell plus SFTP to
the same host means two TCP connections, two authentications and — later — two relay sockets.
Connect SFTP lazily and reuse the cached decrypted credential so the user is not prompted twice.
Agent forwarding is de-scoped from v1. It needs an upstream SSH.NET change. A vault-backed agent of our own plus ProxyJump covers the real use cases.
The SSH suite pulls linuxserver/openssh-server from Docker Hub, which is rate-limited for
unauthenticated pulls. If CI starts failing on image pulls rather than on tests, that is why.
MSIX packaging is ruled out, not merely deprioritised. A packaged app runs WebView2 in an
AppContainer where loopback connections are blocked without a CheckNetIsolation exemption. The
terminal data plane is a loopback WebSocket, so MSIX would break the product outright. Velopack
for Windows/macOS/AppImage; Flatpak and deb/rpm defer updates to the package manager.
The App Sandbox is ruled out on macOS for the same reason, and the entitlements say so. A
sandboxed process cannot listen on loopback without com.apple.security.network.server, and the
terminal is that listener. Developer ID distribution outside the App Store does not require the
sandbox, so this costs nothing today — but it does mean the Mac App Store is closed to this
application without solving the data plane differently first. See
build/macos/DodoSSH.entitlements.
The hardened runtime is not optional and .NET needs four holes punched in it. Notarization
refuses a Developer ID submission without it, and CoreCLR will not start under it without
allow-jit and allow-unsigned-executable-memory — both, not either, because the runtime allocates
executable memory outside the MAP_JIT path as well. disable-library-validation and
allow-dyld-environment-variables are needed for Velopack's updater rather than for the runtime.
Each is argued individually in the entitlements file; the failure mode for a missing one is a
process that dies during runtime initialisation, before anything exists that could report it.
vpk cross-compiles to macOS only as far as the bundle. vpk [osx] bundle runs anywhere and
produces a real .app; there is no [osx] pack off a Mac, because pack drives codesign,
notarytool and stapler. So CI can prove the bundle builds and only a Mac can produce something
installable. Note this is the opposite of the Windows story, where vpk [win] pack builds the
whole installer on Linux — the asymmetry is Apple tooling, not a Velopack limitation.
A custom Info.plist is copied verbatim by vpk, with no substitution whatsoever. That is why
--plist and --bundleId are mutually exclusive, and why build/macos/Info.plist.template is a
template the release script renders rather than a committed file. A committed plist would carry one
version into every release afterwards, and the symptom is silent: Velopack's index would still be
right, the updater would still work, and only Get Info and any crash report would disagree.
macOS app icons live on an 824-in-1024 grid. An icon that bleeds to the edge of its canvas is
not bolder, it is the one icon in the Dock that is too big. dodossh-icon.ps1 draws the .icns at
that fraction and the .ico at full bleed, from one geometry.
Checked rather than assumed, now that Velopack is actually wired up: its Windows path does not
reintroduce the thing MSIX was ruled out for. Setup.exe is an ordinary Win32 executable that unpacks a
directory under %LOCALAPPDATA% and creates shortcuts — there is no AppxManifest, no package identity,
no runFullTrust, no elevation and no execution alias, so the process stays an ordinary desktop process
and WebView2 stays out of an AppContainer. That is reasoning, not measurement; manual check 16.4 is the
measurement, because if a package identity ever did appear the symptom would be the terminal hanging and
then reporting the WebView2 message after fifteen seconds, which reads like a broken runtime rather than
like packaging.
Linux ships AppImage and Flatpak first, specifically so the WebKit runtime is bundled rather than assumed present on the user's machine.
A field generated for x:Name is null in every view in this repository, and using one is a crash rather
than a mistake you can see. The Avalonia name generator declares a field per x:Name and assigns it
inside the InitializeComponent it also generates. No view here calls that method — every one of them
loads its XAML directly with AvaloniaXamlLoader.Load(this), on both heads. So the field exists, compiles,
resolves in the editor, and is null at run time.
What that costs depends on where it is touched. In a constructor it is a NullReferenceException while the
control is being built, and a control being built by PhoneShell on the way up takes the whole launch with
it: the application dies before the first frame, with a stack that names the view rather than the name.
Nightly 0.0.0-alpha.0.133 shipped exactly that, from one line in HostsScreen's constructor.
this.FindControl<T>("Name")! is the idiom, and it is what PhoneShell, TerminalScreen and HostsScreen
all use. Two of those three carried a <remarks> warning about it before the third did it anyway, which is
why it is written here as well: a comment on the control that already got it right is not where somebody
writing a new one is looking.
Opening the system browser depends on the platform handler. SystemBrowserLauncher uses
UseShellExecute, which delegates to ShellExecute on Windows, open on macOS and xdg-open on
Linux. Unverified off Windows: xdg-open comes from xdg-utils, which is not guaranteed on a
minimal desktop or inside a Flatpak sandbox — where the portal is the correct route instead. If
sign-in silently does nothing on Linux, this is the first thing to check. IBrowserLauncher exists
so a platform-specific opener can be substituted without touching the flow.
Identity provider
A loopback redirect URI must be registered without a port, not with a wildcard port. Keycloak — and
providers implementing RFC 8252 §7.3 generally — ignores the port when the registered redirect URI's host
is a loopback literal, which is what lets a native client bind an ephemeral port. Registering
http://127.0.0.1:*/callback looks more explicit and is broken: the * is parsed as a literal port and
every real authorization request comes back 400 Invalid parameter: redirect_uri. Keycloak's wildcard
support is trailing-only, so a * in the middle of a URI never means what it looks like.
Register http://127.0.0.1/callback. Keep the path — it is the part that stops another process on the
machine having an authorization code delivered to a different endpoint. Oidc:LoopbackRedirectPattern,
which the server advertises through /.well-known/dodossh-configuration, says the same thing so an
operator configuring a different provider copies something that works.
Found by running the sign-in against a real Keycloak; every test until then used a stub that accepted whatever it was given.
Keycloak marks its session cookies Secure even over plain HTTP, because SameSite=None is only
legal alongside Secure. A spec-conformant HTTP client therefore refuses to store them from an http://
origin — .NET's CookieContainer drops every one silently — and the login form POST then comes back
400 with no explanation at all. Browsers complete the flow because they treat loopback as a trustworthy
origin and make the exception.
This does not affect the product: the client uses the system browser, which makes that exception. It does
affect any non-browser automation against a development Keycloak, which has to carry the cookies by hand
(see ScriptedBrowser) or be given HTTPS. Two hours of "the credentials must be wrong".
A user declared in a realm import gets no roles unless realmRoles says so — not even the realm's own
default-roles-<realm> composite, which Keycloak grants automatically to a user created through the admin
API or the registration form. The realm file's alice and bob therefore had no role mappings at all, and
because offline_access lives inside that composite and the desktop client requests that scope, the very
first sign-in died at the token exchange with 400 Offline tokens not allowed for the user or client. The
authorization succeeds and the failure lands one step later, which makes it read like a client bug.
Add "realmRoles": ["default-roles-dodossh"] to every user the file declares. And note the asymmetry,
because it is what let this ship: DodoSSH.SystemTests used to create its own account through the admin
API, so it exercised a provisioning path no real user takes and passed while the documented alice could
not sign in at all. The suite now signs in as the realm's own account, and removing these roles fails it.
Keycloak rejects unknown fields in a realm file. RealmRepresentation deserialises with
FAIL_ON_UNKNOWN_PROPERTIES enabled, so a "_comment" key — the usual way to annotate JSON that has no
comment syntax — does not merely get ignored: the import throws
Unrecognized field ... not marked as ignorable and the container refuses to start at all. Explanations
about the realm belong here or in the compose file, never in the realm JSON.
--import-realm skips a realm that already exists. Editing deploy/keycloak/realm-dodossh.json and
running docker compose restart keycloak therefore changes nothing, and the stale configuration keeps
being served — which reads exactly like the edit being wrong. start-dev keeps its state in an H2
database inside the container, so the realm has to be recreated along with it:
docker compose rm -sf keycloak && docker compose up -d keycloak. Cost an otherwise inexplicable
debugging detour.
DodoSSH.SystemTests is immune to this by construction — its Keycloak is created and destroyed per run —
which is a second reason the end-to-end suite starts its own containers rather than reusing the developer's
stack. Editing the realm file and rerunning the suite always tests the edit.
Local cache
The cache location is per-OS and must stay non-roaming. ClientPaths chooses it:
%LOCALAPPDATA%\DodoSSH on Windows, ~/Library/Application Support/DodoSSH on macOS,
$XDG_DATA_HOME/dodossh or ~/.local/share/dodossh on Linux. It must not land anywhere that syncs
to a cloud drive or roams: two machines writing one SQLite file through a file-sync client corrupts it,
and the whole point of the outbox is that each machine has its own. That is also why Windows uses
%LOCALAPPDATA% and not %APPDATA%, which roams in a domain environment.
The platform branches are explicit rather than delegating to
Environment.SpecialFolder.LocalApplicationData everywhere, because on macOS the runtime maps that to
~/.local/share rather than to ~/Library/Application Support. Verified on Windows only — the client
created %LOCALAPPDATA%\DodoSSH\cache.db and migrated it on first launch. The macOS and Linux branches
are reasoned, not run.
SQLite timestamps are stored as integers, deliberately. EF's default DateTimeOffset mapping for
SQLite is a text form it then refuses to order or compare, so any query that sorts or filters by time
throws at execution rather than at model build. UnixMillisecondsConverter is applied as a convention
so a timestamp added later cannot be the one left unconverted. This is provider behaviour, not
platform behaviour, but it cost a debugging session and will again if the converter is removed.
The cache is three files, not one. EF Core's SQLite provider puts the database in WAL mode, which is
the right mode here — a background sync pass writes while the interface reads, and under the default
rollback journal those reads would fail busy — but it means cache.db is accompanied by cache.db-wal
and cache.db-shm. Any backup, export or uninstall routine that touches only cache.db is wrong.
Verified by launching the client and reading PRAGMA journal_mode, after a comment in the code claimed
the opposite.
Pooled SQLite connections keep the file open after the last context is disposed. On Windows that
means locked, so the application cannot delete or replace its own cache and a test cannot clean up after
itself. ClientCacheFactory.Dispose clears the pool for exactly this reason; removing that line makes
the failure appear only on Windows.
No SQLCipher, on any platform. The rows are already ciphertext from the server, so an encrypted
database file would protect bytes that are protected already at the cost of a native dependency and a
licence obligation — and bundle_e_sqlcipher was deprecated in SQLitePCLRaw 3.0. The consequence to
be honest about: the cache offers no protection against another process running as the same user. See
LocalCacheProtector for what it does and does not defend against.
Build and CI
The layout suite needs two unrelated things on a bare image, and each hides the other. Both were found the slow way, one per CI run, because the first masks the second entirely.
The first is libfontconfig. Avalonia's headless renderer is Skia, and the libSkiaSharp.so the test
project copies into its own output links against it; without it the native library never loads and all
69 tests fail inside HeadlessUnitTestSession with a TypeInitializationException on
SkiaSharp.SKImageInfo naming none of their actual subjects. The CI job installs the package.
The second only becomes visible once the first is fixed, and is not about a package at all. Avalonia
takes its default font family from the platform, and on an image with no fonts installed there is no
answer — FontManager throws "Default font family name can't be null or empty" during
AppBuilder.SetupUnsafe, again before any test body runs and again for all 69. WithInterFont does not
help by itself: it registers a collection without naming a default. Both Program.BuildAvaloniaApp and
the layout suite's HeadlessApp now set FontManagerOptions.DefaultFamilyName to
avares://Avalonia.Fonts.Inter/Assets#Inter explicitly, which owes the host nothing because the font
travels in the package.
Pinning it is worth more than the CI fix. A suite that measures text was taking its metrics from
whatever the machine happened to have — Segoe UI on Windows, DejaVu on Linux — and reporting both as
one number. Verified on Alpine musl with fc-list returning zero and on a Fedora desktop with 595
fonts, which now agree. Unverified on Windows: the pinning changed the metrics there too, so a
tight assertion could conceivably have moved.
The end-to-end suite must state its plaintext exemption rather than inherit it. ServerConnection
allows an http OIDC authority only when it is loopback, which is a sound rule the suite cannot lean
on: Testcontainers reports the host the container is actually reachable at, so running the tests
directly gives localhost and passes, while running them inside a container — which is what a
containerised CI runner does — gives the bridge gateway 172.17.0.1 and is refused. That refusal is
the product being correct; a client that quietly accepted plaintext metadata from a routable address
would be a real weakness. M1VerticalSliceTests therefore passes configureOidc to set
RequireHttpsMetadata = false for the throwaway Keycloak it starts itself, and the rule stays as strict
as it was for everyone else.
Integration tests need a Docker daemon (Testcontainers). They run on ubuntu-latest in CI.
macOS runners have no Docker daemon, and the Windows CI job is deliberately build-only. So
anything proved by an integration test is proved on Linux only — which is the right place for
server code, and no coverage at all for client platform behaviour.
The end-to-end suite launches the API's own launcher executable, falling back to dotnet exec on the
assembly. The fallback exists for one reason: a checkout or artefact copy that lost the execute bit
produces a Win32Exception on Linux and nothing whatsoever on Windows. Verified on Windows only — the
launcher path is what runs here, so the fallback itself is reasoned rather than exercised. If the suite
fails in CI with a permission error before any container work, that is the path to look at.
It also depends on Server:PublicBaseUrl being knowable before startup. The port is chosen by binding
a loopback socket and releasing it, because the API reads that URL at startup and advertises it to clients,
so it cannot be discovered from Kestrel afterwards. The window for another process to take the port is a
few milliseconds; if the suite ever fails with an address-in-use, this is why, and a retry is the fix
rather than a redesign.
A RID must never reach the committed lock files, and the obvious fix for a RID-specific publish puts
one there. dotnet publish -r win-x64 resolves a graph the committed packages.lock.json files do not
describe — they carry a net10.0 target and nothing else — so under locked mode it fails NU1004. The
obvious answer is <RuntimeIdentifiers>win-x64</RuntimeIdentifiers> on the desktop head plus a
--force-evaluate to regenerate. That is wrong here, and it was tried and reverted.
A RID declared on one project flows to every project it references transitively while restoring, so the
regenerated lock files for DodoSSH.Contracts and DodoSSH.Crypto grew a net10.0/win-x64 target as
well — and those two are built by the server. The API's Dockerfile restores them with no RID and
--locked-mode, so it failed:
error NU1004: The project's runtime identifiers have changed from.
Project's runtime identifiers: , lock file's runtime identifiers win-x64.
Packaging the desktop client had broken the server's image build, and nothing but the image job would
have caught it. Found by running docker build locally rather than by reading the lock files.
So the RID stays out of the committed state, and the two commands that need one — the release script's
publish and the publish the windows desktop client step in ci.yml — pass -p:RestoreLockedMode=false
for themselves alone. That restore rewrites the lock files as a side effect, which does not matter on a runner
whose checkout is discarded and does matter on a developer's machine, so the release script runs
git checkout -- '*packages.lock.json' afterwards. -p:RestorePackagesWithLockFile=false is not an
alternative: it fails NU1005 whenever a lock file already exists.
vpk picks its target from the host, and cross-compiling is a bracketed directive rather than a flag.
The CI job packages the Windows desktop client on a Linux runner, and plain vpk pack --runtime win-x64
there refuses outright:
To build packages for Linux, the target rid must be Linux (actually was Windows). If your real intention
was to cross-compile a release for Windows then you should provide an OS directive: eg. 'vpk [win] pack ...'
The directive goes before the verb — dotnet vpk '[win]' pack … — and must be quoted in a POSIX shell,
where [win] is a glob matching any one of w, i and n. With it, a Linux runner logs
Directive enabled for cross-compiling from Linux (current os) to Windows and writes
DodoSSH.Desktop-win-Setup.exe, the portable zip, the .nupkg and releases.win.json — the same set a
Windows machine produces. --runtime win-x64 is still required: the directive says which OS is being
targeted, not which RID. Only signing needs Windows tooling, which is why ci.yml can package and
scripts/release-windows.ps1 will still be the thing that signs when there is a certificate to sign with.
dotnet msbuild -getProperty:Version answers 1.0.0, and MinVer is not to blame. -getProperty
without a target evaluates the project and runs nothing, while MinVer computes the version inside a
target — so the read comes back as the SDK default on a full checkout with every tag present, which looks
exactly like a version that was never configured. -t:MinVer makes -getProperty report the value after
that target has run, and every reader of it — the tag check in ci.yml, the desktop nightly job, and
scripts/release-windows.ps1 — passes it. Neither of the first two had, and neither had ever run: the CI
check is if: a tag ref and there are no tags yet, so the first release would have been refused by its own
guard, which would have blamed fetch-depth.
And naming that target requires a restore first, which is a second failure wearing a very different
face. MinVer arrives as a package, so its target is imported from obj/*.nuget.g.targets and does not
exist at all on a clean checkout:
error MSB4057: The target "MinVer" does not exist in the project.
That reads like a typo in the workflow rather than like a missing restore, and it does not reproduce on any machine that has built the project before — which is every developer machine and no fresh runner. The build job's tag check is safe because it runs after that job's own restore; the desktop nightly job and the release script each restore before reading, deliberately and with a comment saying why.
The Android head's lock file is outside the solution, so nothing checks it until the android job runs
— and the android job was broken for an unrelated reason for the whole of the release that went stale.
DodoSSH.Client.Android is deliberately not in DodoSSH.slnx (it needs a workload the other two jobs
have no reason to install), so dotnet restore DodoSSH.slnx --locked-mode — the gate that keeps every
other lock file honest — has never seen it. Its only gate is the android job's own restore, and that job
could not reach the restore step at all while the runner had no JDK and no SDK.
The result: MinVer was added to Directory.Build.props for the desktop updater and reached fourteen
lock files. The fifteenth was not restorable on a runner, so it silently kept a graph from three releases
earlier, and the first thing the repaired job did was fail:
error NU1004: The project's runtime identifiers have changed from.
Project's runtime identifiers: android-arm64;android-x64, lock file's runtime identifiers android-arm64.
Two changes at once, which is why the message names both: the missing MinVer, and an ABI set that grew
when the -r android-arm64 pin came off the packaging step. --force-evaluate on that project alone is
the fix, and — unlike the win-x64 case above — it is safe to commit: the RIDs land in the Android
project's own lock file and the fourteen it references are rewritten byte-identically, because an
android-* RID is not a graph any of them has a package for. Check that with git status rather than
believing it; it is the same mechanism that broke the server's image build, and it happens to land
harmlessly here rather than by design.
The lasting hazard is the first paragraph and not the fix. Any change to a shared props file is a change to a lock file this repository cannot verify from a machine without the Android workload, and it will go on being noticed later than every other one.
.NET for Android cannot be built on a musl host, and this project's runner is Alpine. Every message the toolchain produces on the way to saying so names a missing file that is present. Three CI rounds went into this and the first two fixed symptoms, so the messages are worth reading in the order they arrive.
It fails first inside .NET for Android's tooling resolution:
warning : An error occurred trying to start process '…/packs/Microsoft.Android.Sdk.Linux/36.1.69/tools/Linux/aapt2' … No such file or directory
error XA0111: Unsupported version of AAPT2 found at path '…/tools/Linux'
XA0111 names the wrong problem. Nothing was found, so nothing had a version, and it sends you to an
Aapt2ToolPath in the project file that has never been set. The warning above it is the real message and
it is only a warning. Pointing the build at Google's aapt2 from build-tools instead moves the failure
one step earlier and says it more plainly:
…/build-tools/36.0.0/aapt2: cannot execute: required file not found
That is bash's wording for ENOENT out of execve, which for a file that exists means the ELF
interpreter is missing — not the binary. aapt2 names /lib64/ld-linux-x86-64.so.2, glibc's loader,
which musl does not have. The runner reports linux-musl-x64. Both copies of aapt2 fail for this one
reason and the workload pack was never incomplete.
gcompat and libstdc++ fix that much — measured on alpine:latest, where aapt2 and zipalign will
not start bare and both answer their version once those are installed. And it is not enough, because
the next thing to fail is not a program the build runs but a library the build loads:
error XARLP7000: Error relocating …/tools/libZipSharpNative-3-3.so: __snprintf_chk: symbol not found
__snprintf_chk is a glibc fortify symbol musl does not implement, and this is a DllImport from an
MSBuild task — a glibc shared object being pulled into a musl-linked dotnet process. gcompat supplies
a loader for glibc executables, which is a different problem; there is no shim for this and no musl
variant of the pack. This is the end of the road on Alpine, not a harder step along it.
So the android job builds in a container instead. build/android-build.Dockerfile is Microsoft's own
sdk:10.0-noble plus a JDK, the Android SDK and the workload; scripts/ci-android.sh is everything that
has to happen inside it; and the job keeps on the host only what the host is good at — checkout, git,
publishing. The daemon needed no arranging: the image job already builds with it and every
Testcontainers suite reaches it over the socket. The image is tagged by the digest of the Dockerfile that
made it, so on a persistent runner every run after the first is a cache hit.
A bind mount into that container does not work, and it does not fail either. This runner is itself a
container holding the host's Docker socket, so the workspace path it reports — /root/.cache/act/<hash>/…
— exists in the runner and not on the daemon's host, which is where Docker resolves a bind source. It
finds nothing, creates an empty directory and mounts that. The container then starts perfectly and says:
MSBUILD : error MSB1009: Project file does not exist.
Nothing in that names an empty mount, and nothing earlier in the job would have caught it: docker build
sends its context over the API and Testcontainers mounts nothing, so neither of the two places this
repository was already using Docker proves that a bind mount would work. A socket is not a shared
filesystem, and every check that looked like it said otherwise was answering a different question.
docker cp goes over the same API and therefore does not care where the daemon lives, which is what the
job does now — in with the whole checkout including .git, since MinVer and the versionCode both read
it, and out with the staged package. The repository is well under a megabyte packed, so it costs a
moment. The NuGet cache is a named volume for the same reason: it lives on the daemon and needs no path
either side has to agree on.
Three smaller things worth keeping. The image is built rather than pulled, because a community image
with the Android SDK already in it would put a stranger in the path of a package this project signs and
publishes. Aapt2ToolPath points at the SDK's build-tools even though the workload's own copy works
inside the container — it makes the build use the same binary the packaging step reads the versionName
back with, so the manifest the feed publishes is read by the thing that wrote it. And the pack's layout
is host-shaped, which matters to anything that goes looking: the Linux pack keeps host binaries under
tools/Linux/, the Windows one puts aapt2.exe straight in tools/.
The thing to carry forward is that nothing in this toolchain will tell you the C library is wrong.
Java tools run, dotnet runs, sdkmanager installs, restore succeeds — and then one native thing fails
with a sentence about a file that is plainly on disk. file on the binary, or the name of the missing
symbol, answers in one step what the build's own diagnostics will not.
A Docker ARG named VERSION silently sets MSBuild's Version. An ARG is an environment variable
for the rest of the stage, MSBuild reads environment variables as global properties, and MSBuild property
names are case-insensitive — so ARG VERSION in a build stage sets Version for every project built in
it, with no line anywhere saying so. The workflow passes main-<short sha> on a main build, which is a
fine docker tag and not a version, and the publish died with NETSDK1018: Invalid NuGet version string
pointing at DodoSSH.Contracts — a project nobody had touched. The build stage's argument is therefore
ASSEMBLY_VERSION, passed empty except on a tag build; the VERSION arg in the final stage is only ever
an OCI label and never meets MSBuild. Renaming is the entire fix, and the reason it is written down is that
the symptom names the wrong project and the cause is invisible.
System.Text.Json's source generator does not honour property initializers on a record. Defaults for a
ClientSettings-style record must live on the constructor parameters, not on property initializers,
and getting it wrong fails silently in the worst direction. The generator emits an
ObjectWithParameterizedConstructorCreator — it treats the init-only properties as constructor arguments
and builds new ClientSettings() { A = (T)args[0], … }, so the initializer runs and is then overwritten by
args, which for a member absent from the JSON is the CLR default. Measured: a settings.json of {}
read back TerminalFontSize 0 (clamped up to the 8px floor, not the 13px the renderer draws at) and, once
it existed, AutomaticUpdateChecks false. Reflection-based deserialisation of the same JSON answers 13
and true, which is what makes it so easy to miss — every way of checking it by hand is right except the
one that ships. JsonSourceGenerationMode.Metadata does not help; it was tried. It stayed invisible while
there was one setting, because that setting was written on every save and so was never absent; it went live
the moment a second one was added, since every existing profile lacks the new key.
ASettingAbsentFromTheFile_ComesBackAsItsDeclaredDefault fails without the fix.
[CallerFilePath] is rewritten to /_/... under ContinuousIntegrationBuild. Any test that
locates a fixture by source path passes locally and fails in CI. Copy fixtures to the output
directory and read them via AppContext.BaseDirectory instead; GoldenVectorTests shows the
pattern.
Formatting fails the build rather than a separate step. IDE0055 is an error in
.editorconfig and EnforceCodeStyleInBuild is on, so dotnet build reports misformatted code
the way it reports a type error. CI used to run dotnet format --verify-no-changes as well; it
was removed for spending minutes to reach a verdict the build reaches anyway. dotnet format is
still how to fix what the build complains about — it just no longer gates anything itself.
Deployment
PostgreSQL 18 moved its data directory to /var/lib/postgresql, not /var/lib/postgresql/data
as in 17 and earlier. A compose file carried over from an older version silently gets an empty
volume — the database appears to work and loses everything on restart. Relevant to any compose
file other than deploy/docker-compose.dev.yml, which is already correct.
Keycloak in the dev stack listens on host port 18080, not 8080. On this machine an unrelated
Apache Tomcat holds 127.0.0.1:8080, and a loopback-specific bind wins over Docker's 0.0.0.0
publish when resolving localhost — so every realm request returned 404 while the container
looked healthy. If discovery fails against a locally-published container, check for another
process bound specifically to loopback before suspecting the container.
A path prefix in the server URL is silently discarded. The client uses the typed address only as
HttpClient.BaseAddress and every request path is root-absolute (/api/v1/meta,
/.well-known/dodossh-configuration, …), so https://example.test/dodossh reaches
https://example.test/api/v1/... and the prefix is dropped without a word. That rules out hosting DodoSSH
under a sub-path — which is exactly what a reverse proxy in front of several services usually does. Nothing
trims or normalises the typed URL either, and it is the raw string, not the parsed form, that becomes the
local cache's identity. The server already publishes a canonical apiBaseUrl in its discovery document
that the client could normalise against and currently ignores.
Sync:CursorSigningKey generates an ephemeral per-process key when unset. Fine for a single
node; on a multi-node deployment cursors issued by one node are rejected by another, so clients
resync from the beginning repeatedly. Must be configured explicitly before running more than one
instance. WarnOnRiskyConfiguration logs this at startup.
Rate limiting is not implemented yet (M2). POST /api/v1/me/enrollment and the sync endpoints
are reachable by any authenticated caller at any rate. Enrollment requires a valid access token
and is idempotent, so the exposure is resource consumption rather than a credential-guessing
surface — but it is still an unmetered write path.
/api/v1/me does not update last_seen_at_utc. Deliberate: a GET that writes on every call is
a smell, and nothing depends on the value yet. Revisit when device management lands, since that is
the first feature that needs it.
ApplicationDisplayVersion cannot be set from a target, so the Android head shipped 1.0.0.
Xamarin.Android.Common.targets reads it in a plain top-level PropertyGroup —
<_AndroidVersionName>$(ApplicationDisplayVersion)</_AndroidVersionName> — which is evaluation, not a
target. Every project property is already final before any target runs, so the MinVer-derived value set in
UseTheDerivedVersionForAndroid was assigned after the only thing that reads it had finished. MinVer cannot
run at evaluation time, so no arrangement of the public property works. The tell is that
-getProperty:ApplicationDisplayVersion answers correctly while aapt2 dump badging on the packaged APK
says versionName='1.0.0' — measured, and the reason a -getProperty check cannot catch this class of bug.
The fix is to assign _AndroidVersionName from a target hooked BeforeTargets="_GenerateJavaStubs"; an
internal name, taken deliberately over passing -p:ApplicationDisplayVersion from every caller and leaving
an ordinary dotnet build lying about its version.
Nothing found so far varies the Android launcher name per build. Four mechanisms were tried and all
four produce the same label. AndroidManifestPlaceholders is wired to the manifest task and does not reach
android:label — measured on a build whose placeholder property evaluated to appLabel=DodoSSH nightly
and whose APK reported DodoSSH. ApplicationTitle, the documented property, feeds an ApplicationLabel
task parameter that the label already on the application element wins against. A second resource directory
under Resources/ is picked up by the SDK's own glob as a qualifier and fails the build outright with
APT2142: invalid configuration 'nightly'. And an AndroidResource Remove/Include swap does nothing
from the project body — the SDK's glob is added by Sdk.targets, imported below it, so the removal runs
before the item exists — while the same swap inside a target that runs before UpdateAndroidResources also
had no effect. Underneath all of it: the launcher shows the activity's label, and that one is a string in
a C# attribute. The two channels ADR 0014 defines therefore share a launcher name, and are told apart by
the package name in Android's app info, by the version, and by the channel the application names on its own
preferences screen.