Files
DodoSSH/docs/platform-flags.md
T
jaap-jan d07b336868 Free the terminal from the Hosts screen, and fill the room it left
The WebView sat inside the Hosts grid, so navigating to Files or the keychain
hid every open terminal and the strip that named them. A connection you had
opened was invisible from four of the five screens. The window now has two
surfaces rather than one: a nav rail that says which page you are on, and a
terminal strip that is always there and switches the whole content area to a
shell. Screen keeps meaning "which page" and never becomes a sixth kind of
page, which is why this is two properties instead of one enum with a terminal
member in it.

Every screen lives inside one wrapper panel that collapses when a terminal is
showing. That is not tidiness — the WebView hosts a Win32 child window that
composites above everything Avalonia draws, so a screen left visible over its
rectangle is a screen sliced in half, and this window has shipped that defect
once already. One decision point, IsTerminalShowing, and a nested panel rather
than five compound bindings nobody would remember to extend.

The focus choreography is the part no test in this repo can see. Every reveal
path now focuses in the same turn the WebView appeared, so all three of them
post at DispatcherPriority.Loaded and let the native control re-push its bounds
first. Going the other way had a real bug: the screen-changed branch called a
bare Focus() where it had to release the keyboard from the native child, so
switching from a terminal to Files silently ate the first keystrokes. Rare
before this commit and the primary gesture after it.

The tab strip grew a cross inside each tab, a plus that opens the quick-connect
palette, and middle-click close. Nested buttons are correct here: Avalonia
handles a left press on the cross and deliberately does not handle other
buttons, which is exactly what lets middle-click bubble up from the cross as
well as the tab. The test is PointerUpdateKind rather than
IsMiddleButtonPressed, because the latter reports button state and is also true
for a left press made while the middle button happens to be held. The handler
is on the tab and not the strip, so the background closes nothing by
construction. Plus opens the palette rather than a flyout, since a menu
dropping into the WebView's rectangle may or may not composite above a child
HWND and this repo does not make rendering claims it has not photographed.

Everything a user reads now says keychain. The wire, the database and the
cryptographic spec still say vault, deliberately: renaming those is a migration
and a protocol change for a word. That split is written down rather than left
to be rediscovered as an inconsistency.

Four things that were squeezed into the keychain's category rail, or into
nothing at all, now have screens. Pinned host keys get one, with fingerprints
never truncated and a filter that matches them, because comparing what you have
against what the operator published is the whole workflow; the approved date is
read out of the item's UUIDv7 rather than added as a column, and says so, since
it means first approval and not last use. Keys can be generated in the client,
which needed the openssh-key-v1 container written by hand — there is no BCL or
NSec helper, and the PKCS#8 route is unverified in the SSH library this uses.
The armour carries no passphrase: encrypting it needs bcrypt_pbkdf, which is
Blowfish with a swizzle, in a project whose crypto is otherwise entirely
libsodium, for a protection the key's own remarks argue is redundant inside a
vault. Generation fills the existing editor and stops, so SAVE stays the one
thing that writes. ~/.ssh/config can be imported behind a preview that is
ticked per row and writes nothing until the button; IdentityFile records the
path and imports the key material only on an explicit opt-in, because reading
somebody's private key into a vault is precisely the act this product exists to
make deliberate. Match blocks and ProxyJump are reported rather than obeyed —
one cannot be evaluated statically and the other has nothing behind it to route
with, and a preview that implied otherwise would be worse than one that admits
it.

Files can be dragged in all four directions that are honestly available. Remote
to Explorer does not ship and is not pretended to: the shell wants the bytes
during the drop, which needs a virtual file and a native COM data object,
outside what Avalonia offers. Note for the next person that Avalonia 12
replaced the drag model outright — DataObject and DataFormats are no-op stubs
and IDataObject is not in the reference assembly, so every tutorial written for
11 does not compile here.

Hosts can be grouped, flat and never nested. A parent id merged as a scalar
lets two offline clients each re-parent A under B and B under A, producing a
cycle inside an encrypted payload that no server can police and every reader
would have to detect for ever. Membership lives in that payload rather than in
the one plaintext concession ADR 0001 allows, whose test is that the relay
cannot function without it — nothing on the server reads a group, so what
plaintext would hand over is a clustering of the estate for nothing. The
plaintext column reserved for it is dropped, provably always null, and the
server now refuses a client that sends one; it was never populated, was copied
on apply, and was not cleared on delete, so a group id would have outlived the
host it described.

Snippets insert through xterm rather than through the pump, because xterm is
the only thing that knows whether the remote has bracketed paste on, and that
is what makes a shell treat embedded newlines as text instead of as execute.
The host process moves opaque bytes and never parses output, so it would have
to guess, and guessing wrong runs every line. Running is off by default and the
copy says the text goes into whatever is there — the terminal has no notion of
being at a prompt, and may be in vi or at a password prompt with echo off, so
the Enter the user presses themselves is the entire safety property.

Connections and keychain changes are recorded as synced encrypted items, which
is what makes them auditable by a team later and costs the server knowledge of
connection rate and timing from row counts alone. ADR 0001 already concedes it
cannot hide that class of metadata; the trade is now written into it rather
than left implicit. A connection entry is written once, at close, which is what
makes a synced log tractable: nothing to merge, one outbox row, no chance of
colliding with itself. Live sessions come from memory, not from the log. The
write is void by contract and posts to a bounded channel, because putting an
encrypt-and-write on the teardown path of every session is how closing the
application comes to take four seconds. A ticket opened before a lock still
closes afterwards, since a shell outlives the vault. The activity log hooks the
one generic repository every kind writes through, so it cannot miss a caller —
which is also why the log kinds themselves declare they are not audited, or the
first entry would write an entry about writing an entry. It records the names
of the fields that changed and never their values; a log with an old password
in it would be a plaintext credential store with no vault around it. Retention
is 90 days or 5,000 entries, whichever bites first, pruned on the sync loop
rather than on a second timer.

That log traffic then broke the status line, which is worth recording because
the fix is a shape and not a patch: background sync counted its own log rows as
pushed items, so the quiet rule stopped being quiet and every action's message
was overwritten a second later by a sync report. The report now separates log
rows from user items and the rule reads the latter.

S3 buckets appear as a remote in the file browser, behind the same interface an
SFTP session implements, so the queue and both panes did not have to learn what
they are talking to. Uploads go through a pipe, because the queue wants to
write and the SDK wants to read; memory is then bounded by the part size
instead of buffering a file to disk twice.

Finally, the Windows device key store moved out of the session project, which
was the one thing keeping it from being portable — everything else in it is
platform-neutral, and a Windows CNG dependency in the middle of the vault code
meant a second head could not reference it without dragging Windows along. The
seam that made the move free was already there. docs/android-port.md is the
audit behind that: what ports, what does not, in order of cost, the four
decisions taken, and an inventory of every screen and state the interface has
to carry, written so a design can be made from it directly.

dotnet build, dotnet test and dotnet format --verify-no-changes are all clean:
1240 tests at zero warnings, including the end-to-end suite against real
containers. The manual checks that headless Avalonia cannot make — the drag
from Explorer, a generated key against a real host, twelve tabs at the minimum
window width — are listed in docs/manual-checks.md and are still outstanding.
2026-07-31 20:30:05 +02:00

35 KiB
Raw Blame History

Platform flags

Things known or suspected to behave differently outside Windows, plus deployment gotchas that have already cost time once. Development and testing are currently Windows-only, so anything here marked unverified has not run on the platform in question and must not be assumed to work.

Each entry says what the risk is, why it matters, and what to do about it. Delete an entry when it has been verified or made moot — not when it merely stops being convenient.

Cryptography

ChaCha20Poly1305.IsSupported is false on macOS, and on Windows builds before 10.0.20142. This is why the client uses NSec (libsodium) rather than the BCL for content encryption; see docs/crypto.md §1. Already mitigated — but if a BCL AEAD path is ever added as a fallback it must gate on IsSupported rather than assuming availability, or the client will fail to open any vault on macOS.

Argon2id timings are measured on one Windows machine only. 256 MiB with t=4 took 323 ms here. The floor and ceiling in EnrollmentLimits were chosen against that number. Unverified elsewhere: recalibrate on the slowest target platform before recommending a default profile, because a cost that is comfortable on a desktop can make unlock unusable on a low-power laptop — and the parameters are stored per user at enrollment, so a bad default is a per-user migration.

libsodium ships native binaries per RID. This complicates single-file and AOT publishing, and on macOS every native library (libsodium, libSkiaSharp, libHarfBuzzSharp, libe_sqlite3) must be signed individually with --options runtime --timestamp before the bundle is signed, or notarization fails with an error that does not name the offending file.

Desktop client

The WebView runs on Windows. Avalonia.Controls.WebView 12.0.1 (MIT, no licence key) hosts the terminal page: WebView2 launches, navigates to the loopback page, runs its JavaScript and completes the WebSocket handshake. Verified by observing an established TCP connection from msedgewebview2 to the data plane port.

Note precisely what that evidence covers, because it was once stretched to cover more: every clause above is about the process and the socket. It says nothing about how the control composites with Avalonia-drawn content, which is the axis on which it does not behave like an ordinary control — see the next entry.

A native child window cannot be covered by Avalonia content, on any platform that hosts it windowed. NativeWebView attaches a real Win32 child HWND through NativeControlHost — on Windows the backend creates a WS_CHILD holder window and SetParents WebView2's HWND into it — and a child window paints above everything its parent draws, whatever the visual tree's z-order says. This is by design and acknowledged upstream: "NativeControlHost places native controls over Avalonia content just like WPF one does. So it suffers from the same airspace problem" (Avalonia's maintainer, #6605, still open). Reproduced in a 60-line standalone app with no DodoSSH code: a 340,* grid, a NativeWebView in column 1 and an opaque Border as a later Panel sibling renders the overlay sliced dead on x=340.

Layering a screen over the terminal therefore does nothing: the WebView's rectangle stays on top. In this shell that sliced the setup and unlock cards at the terminal column's left edge, put every one of their buttons inside the WebView's rectangle at the window's default width — so the flow could only be completed by keyboard — and handed Win32 focus to WebView2 on any click in that region, which makes a text box stop accepting keystrokes with no visible cause. That last symptom is the focus asymmetry documented further down, not a separate fault: focus crosses into the WebView readily and does not come back on its own.

The fix is to collapse the control, not to cover it: IsVisible="{Binding IsUnlocked}" on the NativeWebView. That is safe, and this is the part worth recording, because the opposite was asserted here for a while:

  • NativeControlHost creates the native attachment from attach to the visual tree, not from layout and not from visibility. Its UpdateHost never reads IsEffectivelyVisible; only TryUpdateNativeControlPosition does, choosing HideWithSize over ShowInBounds.

  • NativeWebView stashes a Source assigned before its adapter exists and replays it once created, so navigation is never lost to ordering. The shell already depends on that replay.

  • So a collapsed WebView still starts WebView2, still loads the page and still lets the renderer attach its socket. Measured on Windows in a harness mirroring the data plane's handshake, with IsVisible=false set before the window was ever shown: adapter created, GET /, then the WebSocket 101 sent — the moment RendererAttached fires — followed by frames arriving over the socket, all while hidden. A cold WebView2 profile behaves the same. Revealing it recomputes bounds within about 7 ms, on one ResizeObserver callback, over the same socket.

    It must be IsVisible, not removal from the tree. Detaching runs DestroyNativeControl and takes the whole WebView2 process tree with it, so conditional content or a template swap would pay a cold start on every unlock. Hiding merely does SetWindowPos(holder, …, SWP_HIDEWINDOW). Negative Margin also works as a runtime toggle; RenderTransform does not, because NativeControlHost never watches it.

    Note the earlier version of this bullet cited "35 msedgewebview2 processes" as the confirmation. A process count cannot show that a socket was accepted — it is the same shape of mistake as the one described below, one level down.

The previous version of this entry claimed the reverse — that hiding it would mean never realising it — and cited the msedgewebview2 connection as verification. That observation was made while the overlay was showing but, because of the airspace behaviour above, the WebView was in fact uncovered and in plain view. It confirmed only that a visible WebView is realised, which nobody disputed, and could not discriminate the case it was attached to. A process-level check cannot verify a rendering claim; that needs a screenshot, and this defect shipped because one was never taken.

What the first connection after unlocking actually depends on is the await workspace.WaitForRendererAsync(cancellationToken) in VaultViewModel.ConnectAsync, because TerminalDataPlane.SendAsync drops frames when no renderer is attached rather than queueing them. That await is the invariant; the control's visibility is not.

It is now bounded — TerminalWorkspaceOptions.RendererTimeout, 15 s, plus the command's own token — because whether the renderer attaches at all depends on a runtime this application does not install. A missing or policy-blocked Evergreen runtime, or an AppContainer that cannot reach loopback, previously left Connect waiting forever with IsBusy stuck and nothing on screen to explain it. The gate is unchanged; only the wait is. Why 15 s and not less: attaching is near-instant in the normal case (the page attaches while the unlock screen is still up), but a cold WebView2 profile creates a user-data directory and starts its process tree first, and reporting a broken runtime to someone whose runtime was merely slow is the worse error. The timeout is caught in VaultViewModel and reported as a message naming WebView2, because TimeoutException.Message is "The operation has timed out" and names nothing.

Hiding the WebView does not pause it. With the holder window hidden, the page keeps visibilityState: "visible" and requestAnimationFrame keeps firing at roughly 115/s — Chromium does not treat a hidden child HWND as a hidden page. That is why the handshake completes while collapsed, so it is load-bearing rather than merely wasteful, but it means a locked DodoSSH is still animating a full-size off-screen page. Worth revisiting if idle power ever matters.

A degenerate pane size reaches the remote pty. The fit addon floors its proposal at 2 columns by 1 row rather than refusing, so any path that fits a terminal with almost no viewport sends window-change for a 2x1 window and permanently mangles the wrapped scrollback. Reachable today by minimising, and — once splits land — by dragging a splitter to the edge. terminal.js now skips the fit below 40 px in either axis. Related, and now fixed: the conflict log was an ItemsControl with no ScrollViewer and no MaxHeight on an Auto row, so enough conflicts squeezed the row below it toward nothing. It survived that long because it lived in MainWindow.axaml, which no test can lay out. Moving it into HostsScreen.axaml — a UserControl, and therefore measurable — is what surfaced it; it now has both, and TheHostsScreenFitsWithAConflictLogTooLongToShow fails without them.

Keyboard focus crosses into the WebView by itself and does not come back. This is the asymmetry to know; the connect-focus bug that led here was only its first symptom. Measured on Windows with a harness that reports GetFocus(), the class name of the window holding it, and the page's own document.hasFocus() at each step.

  • Into the page: nothing custom is needed. NativeWebView overrides Focusable to true and its OnGotFocus calls the adapter's Focus(), which on Windows is ICoreWebView2Controller::MoveFocus(PROGRAMMATIC). A plain Avalonia Terminal.Focus() therefore moves real Win32 focus to the Chrome_WidgetWin_1 child and the page reports hasFocus: true. No SetFocus P/Invoke and no COM work — the package version of this entry that assumed otherwise was wrong. The control also replays a Focus() that arrived before its adapter existed, and re-asserts itself: while it holds Win32 focus its GotFocus handler pulls Avalonia's logical focus back onto the control. Worth stating positively, because the reasonable guess before measuring — that crossing into a child HWND must need SetFocus — is the wrong way round: it is the return trip that needs it.
  • Out of the page: the package does nothing at all. OnLostFocus calls the adapter's ResignFocus(), and on Windows that method is empty. So someTextBox.Focus() moves Avalonia's focused element while Win32 focus stays on WebView2: a text box with a caret that silently receives nothing. Window.Activate() and Window.Focus() were both measured and neither recovers it. The hand-back has to be SetFocus(topLevelHwnd) — see Views/NativeKeyboardFocus.cs. A real mouse click does recover it, because Avalonia's window sets focus on pointer input, which is exactly why this is invisible to anyone who clicks before typing.
  • Collapsing the control does not release the keyboard. With IsVisible=false the holder window is hidden but Win32 focus stays on it — measured as focus held by a window reporting visible=False, with Avalonia's focused element becoming (none). So locking the vault after touching the terminal left the unlock passphrase box eating keystrokes. The lock path now hands the keyboard back and focuses that box.
  • Focus() on a collapsed control is a no-op and is not replayed on reveal. Order matters: reveal, then focus. Focus does survive a lock/unlock cycle when done that way.
  • There is no Tab-out. The package subscribes ICoreWebView2Controller::add_MoveFocusRequested and its handler body is empty, so WebView2's request to move focus off itself is discarded; xterm eats Tab anyway. The way out is Ctrl+Shift+F6, intercepted in terminal.js and sent to the host as a web message — measured arriving verbatim in WebMessageReceivedEventArgs.Body. It has to be handled in the page, because once the child window owns Win32 focus Avalonia sees no key events and no KeyBinding could fire. Not Escape, and not a bare F6: both are keys a TUI legitimately binds, and Ctrl+Shift is the range terminal emulators conventionally keep for themselves.

None of this is covered by a test, and cannot be here: headless Avalonia has no native window, so a headless test renders and focuses correctly and would confirm the wrong belief. What the suite covers is the plumbing that drives it — that connecting asks for focus once per session, that a failed connect does not, and that locking stops the forwarding.

The lock/unlock cycle does not resize the pane at all, and the 40 px guard is not what makes it safe. Measured on Windows with a live shell, against a real sshd in a container, in a harness mirroring MainWindow.axaml's 340,* grid: with the NativeWebView collapsed by IsVisible=false, the page still reports paneWidth: 840, paneHeight: 760, unchanged cols/rows, and visibilityState: "visible". Hiding is SetWindowPos(holder, …, SWP_HIDEWINDOW), which does not resize the holder, so no ResizeObserver callback fires, no fit runs, and no window-change reaches the remote — before, during or after the cycle. stty size on the remote answered 50 118 both before locking and after unlocking, and the renderer's own buffer came back byte for byte, wrapped lines included.

The guard's irrelevance here was established rather than assumed: the same run with MINIMUM_FITTABLE_PIXELS patched to 0 — the guard fully disabled — produced an identical clean result. So the guard is still worth keeping for the paths it was written for, minimising and a splitter dragged to the edge, but it is not on the lock path and must not be cited as the reason locking is safe. It was described that way when it landed.

Two further results from the same harness, both about the deliberate decision that shells outlive a lock (README, MainWindowViewModel.LockAsync):

  • A collapsed WebView is not typed into. With the harness confirmed as the foreground window and all twelve injected SendInput events accepted, not one character of the probe reached the remote pty, and a Ctrl-U afterwards answered BEL — nothing was sitting in the remote's line editor either. So the lock screen is a real input barrier even though the session behind it is live, and that is what makes surviving the lock defensible rather than merely convenient. The mechanism is not what this run concluded: it read the result as a hidden WS_CHILD window being ineligible for keyboard focus, but the focus entry above measured Win32 focus still held by the hidden holder window, and the lock path now moves the keyboard off it deliberately. Take the barrier as measured here and the reason from there — which also means the barrier is something the lock path maintains, not something the platform guarantees.
  • The session survives the cycle in the real control, not only in tests. LiveSessionCount was 1 before, during and after, and the shell accepted a command again immediately on unlock.

Suspected, seen once, not reproduced: on the first run — before the harness learned to wait for the window's scale to settle — the window opened at 2558x1367 px and the page reported a 2202x1328 pane (312x88 characters) for a window 1180 logical units wide, which looks like physical pixels arriving where CSS pixels were expected. A later re-push to 1177x672 then reflowed the wrapped line and split it in two. Both events straddled a DPI settle rather than the lock, and three later runs at RenderScaling 1.00 never showed it. If a user reports mangled scrollback after moving the window between displays of different scale, start here.

WebView2 will not initialise when the host executable sits under a very long path. CreateCoreWebView2Environment fails with COMException 0x80080005 CO_E_SERVER_EXEC_FAILURE ("Server execution failed") and the terminal never appears. Hit while building the harness above: the same binary that failed from a ~230-character directory ran first time from %TEMP%\h. The exact threshold was not established and the mechanism is unconfirmed — the user data folder is created beside the executable by default and the browser process is launched with paths derived from it, so MAX_PATH is the obvious suspect. Relevant to packaging: an installer that lands under a deep per-user path would break the terminal with an error that names nothing.

The Windows app manifest must declare a supportedOS list. Without it the process reports a downlevel Windows version and Avalonia's native control host fails outright — "Unable to create child window for native control host" — so the WebView, and therefore the terminal, does not start at all. [STAThread] on Main is equally mandatory: WebView2 checks the apartment state and refuses to initialise on an MTA thread.

WebView2 spawns a process tree, not a process. Around 35 processes were observed for one embedded view. That is the concrete reason the design uses one WebView hosting N terminals rather than one per tab: twenty tabs would mean twenty of those trees.

The Avalonia WebView on Linux remains unproven, and is still the largest risk in the plan. The package's own release notes say NativeWebView gained Linux support via a WPE backend (libwpewebkit-2.0), which is much less widely installed than WebKitGTK — and it ships a separate NativeWebDialog described as "particularly useful for platforms like Linux where embedded WebView controls might not be available", which is the vendor confirming the concern. Unverified: a spike must cover Ubuntu on both Wayland and X11, Fedora KDE, and macOS 15.

ITerminalHost was supposed to be the seam that keeps a backend swap cheap, and it is declared but not implemented — nothing in the application uses it, and the view navigates NativeWebView.Source directly. Swapping backends today means editing MainWindow.axaml and its code-behind. That is a small job, but do not plan around a seam that is currently only a file.

One more reason the Linux picture may be better than this entry assumes: the package also ships NativeWebViewCompositorHost, a non-windowed host drawn through Avalonia's compositor. A compositor host would not have the airspace problem described below at all. Whether it can be selected deliberately is unknown and worth establishing during the spike, because it would change how overlays can be built.

Avalonia.Diagnostics has no 12.x release (latest is 11.3.18), so the developer tools overlay is unavailable on Avalonia 12. Development-only, so nothing ships differently — but debugging a layout problem currently means reasoning rather than inspecting.

The xterm bundles are vendored, not built. @xterm/xterm 6.0.0 with the fit and webgl addons, all MIT, committed as UMD bundles under WebAssets/vendor and embedded as Avalonia resources. No npm or esbuild step, so a clean clone builds with the .NET SDK alone. The cost is that upgrades are a manual re-download; the licence and versions are recorded here so that stays visible.

Bracketed paste is the renderer's to decide, and it is why snippets go through a frame. xterm tracks \e[?2004h from the remote's own output and Terminal.paste(text) wraps the text in paste markers only when the mode is on — which is what makes a shell treat embedded newlines as text rather than as "run this". The host process cannot make that decision: TerminalDataPlane moves opaque bytes and never parses output, so writing a snippet straight into the pump would mean guessing, and guessing wrong executes every line of a multi-line command. Hence TerminalServerOpcode.Paste. Two consequences worth keeping:

  • The Enter for a snippet marked as running goes through Terminal.input('\r'), outside the wrapper. A \r appended to the pasted text is bracketed with it and arrives as a literal character, so nothing runs.
  • Against a remote with bracketed paste off — a raw sh, or a session inside an editor — a multi-line snippet does run line by line, and nothing can prevent that. It is a property of the terminal protocol, not of this client, which is why the screen says "types this into whatever is there".

Both methods were confirmed present on the public API of the vendored @xterm/xterm 6.0.0 bundle before being written against; neither is reachable from any test in this repository, so they are in docs/manual-checks.md as checks 3.63.8.

SSH.NET's window-change is verified working as of 2025.1.0 — resolved, not a flag. ShellStream.ChangeWindowSize(columns, rows, width, height) exists and the remote genuinely observes it: PtyAndResizeSpikeTests reads stty size back from a real sshd after resizing, and repeated resizes each take effect. The IChannelSession fallback is not needed. That suite stays in place as a regression guard, because an upgrade that silently stopped sending the request would present as wrapped output only after a resize — easy to misattribute to the terminal emulator.

ShellStream.Write buffers and requires an explicit Flush. Without one a keystroke is accepted, reported as written, and never reaches the remote — the terminal displays output perfectly and simply stops responding to input. SSH.NET's own WriteLine flushes, which is why a spike that used it never hit this. SshNetShellSession.WriteAsync now flushes per write; batching would be wrong anyway, since a terminal has to put a keystroke on the wire immediately.

ShellStream does not override ReadAsync. The base Stream implementation therefore runs the blocking Read on a thread-pool thread, so every open session parks one thread for as long as it is idle. Fine for the handful of tabs M1 targets; revisit before advertising many concurrent sessions, since the fix is either an upstream change or driving IChannelSession directly.

A passphrase supplied for an unprotected private key is silently ignored, not refused. PrivateKeyFile(stream, passphrase) on an unencrypted PKCS#1 RSA key loads it and the connection authenticates exactly as if no passphrase had been given — measured against a real sshd in KeyAuthenticationTests.APassphraseOnAnUnprotectedKey_IsIgnoredRatherThanRefused, which was written expecting the opposite and corrected to match. Two consequences, and the second is the one that bites: a stray passphrase does no harm, so nothing downstream needs to defend against it; but equally nothing downstream will report one, so if a user swears they set a passphrase and the key opens without it, no error will ever say so. Only established for that armour and that algorithm; whether the OpenSSH format's none cipher path behaves the same way is untested. SshKeySecret.Passphrase still normalises an empty string to null, for the reasons stated there — one representation of one state — and not for this.

SSH.NET cannot share one connection between SshClient and SftpClient. A shell plus SFTP to the same host means two TCP connections, two authentications and — later — two relay sockets. Connect SFTP lazily and reuse the cached decrypted credential so the user is not prompted twice.

Agent forwarding is de-scoped from v1. It needs an upstream SSH.NET change. A vault-backed agent of our own plus ProxyJump covers the real use cases.

The SSH suite pulls linuxserver/openssh-server from Docker Hub, which is rate-limited for unauthenticated pulls. If CI starts failing on image pulls rather than on tests, that is why.

MSIX packaging is ruled out, not merely deprioritised. A packaged app runs WebView2 in an AppContainer where loopback connections are blocked without a CheckNetIsolation exemption. The terminal data plane is a loopback WebSocket, so MSIX would break the product outright. Velopack for Windows/macOS/AppImage; Flatpak and deb/rpm defer updates to the package manager.

Linux ships AppImage and Flatpak first, specifically so the WebKit runtime is bundled rather than assumed present on the user's machine.

Opening the system browser depends on the platform handler. SystemBrowserLauncher uses UseShellExecute, which delegates to ShellExecute on Windows, open on macOS and xdg-open on Linux. Unverified off Windows: xdg-open comes from xdg-utils, which is not guaranteed on a minimal desktop or inside a Flatpak sandbox — where the portal is the correct route instead. If sign-in silently does nothing on Linux, this is the first thing to check. IBrowserLauncher exists so a platform-specific opener can be substituted without touching the flow.

Identity provider

A loopback redirect URI must be registered without a port, not with a wildcard port. Keycloak — and providers implementing RFC 8252 §7.3 generally — ignores the port when the registered redirect URI's host is a loopback literal, which is what lets a native client bind an ephemeral port. Registering http://127.0.0.1:*/callback looks more explicit and is broken: the * is parsed as a literal port and every real authorization request comes back 400 Invalid parameter: redirect_uri. Keycloak's wildcard support is trailing-only, so a * in the middle of a URI never means what it looks like.

Register http://127.0.0.1/callback. Keep the path — it is the part that stops another process on the machine having an authorization code delivered to a different endpoint. Oidc:LoopbackRedirectPattern, which the server advertises through /.well-known/dodossh-configuration, says the same thing so an operator configuring a different provider copies something that works.

Found by running the sign-in against a real Keycloak; every test until then used a stub that accepted whatever it was given.

Keycloak marks its session cookies Secure even over plain HTTP, because SameSite=None is only legal alongside Secure. A spec-conformant HTTP client therefore refuses to store them from an http:// origin — .NET's CookieContainer drops every one silently — and the login form POST then comes back 400 with no explanation at all. Browsers complete the flow because they treat loopback as a trustworthy origin and make the exception.

This does not affect the product: the client uses the system browser, which makes that exception. It does affect any non-browser automation against a development Keycloak, which has to carry the cookies by hand (see ScriptedBrowser) or be given HTTPS. Two hours of "the credentials must be wrong".

A user declared in a realm import gets no roles unless realmRoles says so — not even the realm's own default-roles-<realm> composite, which Keycloak grants automatically to a user created through the admin API or the registration form. The realm file's alice and bob therefore had no role mappings at all, and because offline_access lives inside that composite and the desktop client requests that scope, the very first sign-in died at the token exchange with 400 Offline tokens not allowed for the user or client. The authorization succeeds and the failure lands one step later, which makes it read like a client bug.

Add "realmRoles": ["default-roles-dodossh"] to every user the file declares. And note the asymmetry, because it is what let this ship: DodoSSH.SystemTests used to create its own account through the admin API, so it exercised a provisioning path no real user takes and passed while the documented alice could not sign in at all. The suite now signs in as the realm's own account, and removing these roles fails it.

Keycloak rejects unknown fields in a realm file. RealmRepresentation deserialises with FAIL_ON_UNKNOWN_PROPERTIES enabled, so a "_comment" key — the usual way to annotate JSON that has no comment syntax — does not merely get ignored: the import throws Unrecognized field ... not marked as ignorable and the container refuses to start at all. Explanations about the realm belong here or in the compose file, never in the realm JSON.

--import-realm skips a realm that already exists. Editing deploy/keycloak/realm-dodossh.json and running docker compose restart keycloak therefore changes nothing, and the stale configuration keeps being served — which reads exactly like the edit being wrong. start-dev keeps its state in an H2 database inside the container, so the realm has to be recreated along with it: docker compose rm -sf keycloak && docker compose up -d keycloak. Cost an otherwise inexplicable debugging detour.

DodoSSH.SystemTests is immune to this by construction — its Keycloak is created and destroyed per run — which is a second reason the end-to-end suite starts its own containers rather than reusing the developer's stack. Editing the realm file and rerunning the suite always tests the edit.

Local cache

The cache location is per-OS and must stay non-roaming. ClientPaths chooses it: %LOCALAPPDATA%\DodoSSH on Windows, ~/Library/Application Support/DodoSSH on macOS, $XDG_DATA_HOME/dodossh or ~/.local/share/dodossh on Linux. It must not land anywhere that syncs to a cloud drive or roams: two machines writing one SQLite file through a file-sync client corrupts it, and the whole point of the outbox is that each machine has its own. That is also why Windows uses %LOCALAPPDATA% and not %APPDATA%, which roams in a domain environment.

The platform branches are explicit rather than delegating to Environment.SpecialFolder.LocalApplicationData everywhere, because on macOS the runtime maps that to ~/.local/share rather than to ~/Library/Application Support. Verified on Windows only — the client created %LOCALAPPDATA%\DodoSSH\cache.db and migrated it on first launch. The macOS and Linux branches are reasoned, not run.

SQLite timestamps are stored as integers, deliberately. EF's default DateTimeOffset mapping for SQLite is a text form it then refuses to order or compare, so any query that sorts or filters by time throws at execution rather than at model build. UnixMillisecondsConverter is applied as a convention so a timestamp added later cannot be the one left unconverted. This is provider behaviour, not platform behaviour, but it cost a debugging session and will again if the converter is removed.

The cache is three files, not one. EF Core's SQLite provider puts the database in WAL mode, which is the right mode here — a background sync pass writes while the interface reads, and under the default rollback journal those reads would fail busy — but it means cache.db is accompanied by cache.db-wal and cache.db-shm. Any backup, export or uninstall routine that touches only cache.db is wrong. Verified by launching the client and reading PRAGMA journal_mode, after a comment in the code claimed the opposite.

Pooled SQLite connections keep the file open after the last context is disposed. On Windows that means locked, so the application cannot delete or replace its own cache and a test cannot clean up after itself. ClientCacheFactory.Dispose clears the pool for exactly this reason; removing that line makes the failure appear only on Windows.

No SQLCipher, on any platform. The rows are already ciphertext from the server, so an encrypted database file would protect bytes that are protected already at the cost of a native dependency and a licence obligation — and bundle_e_sqlcipher was deprecated in SQLitePCLRaw 3.0. The consequence to be honest about: the cache offers no protection against another process running as the same user. See LocalCacheProtector for what it does and does not defend against.

Build and CI

Integration tests need a Docker daemon (Testcontainers). They run on ubuntu-latest in CI. macOS runners have no Docker daemon, and the Windows CI job is deliberately build-only. So anything proved by an integration test is proved on Linux only — which is the right place for server code, and no coverage at all for client platform behaviour.

The end-to-end suite launches the API's own launcher executable, falling back to dotnet exec on the assembly. The fallback exists for one reason: a checkout or artefact copy that lost the execute bit produces a Win32Exception on Linux and nothing whatsoever on Windows. Verified on Windows only — the launcher path is what runs here, so the fallback itself is reasoned rather than exercised. If the suite fails in CI with a permission error before any container work, that is the path to look at.

It also depends on Server:PublicBaseUrl being knowable before startup. The port is chosen by binding a loopback socket and releasing it, because the API reads that URL at startup and advertises it to clients, so it cannot be discovered from Kestrel afterwards. The window for another process to take the port is a few milliseconds; if the suite ever fails with an address-in-use, this is why, and a retry is the fix rather than a redesign.

[CallerFilePath] is rewritten to /_/... under ContinuousIntegrationBuild. Any test that locates a fixture by source path passes locally and fails in CI. Copy fixtures to the output directory and read them via AppContext.BaseDirectory instead; GoldenVectorTests shows the pattern.

dotnet format --verify-no-changes is part of the CI gate and exits non-zero on style warnings, not just whitespace. Run it before pushing; a build with zero warnings can still fail that step.

Deployment

PostgreSQL 18 moved its data directory to /var/lib/postgresql, not /var/lib/postgresql/data as in 17 and earlier. A compose file carried over from an older version silently gets an empty volume — the database appears to work and loses everything on restart. Relevant to any compose file other than deploy/docker-compose.dev.yml, which is already correct.

Keycloak in the dev stack listens on host port 18080, not 8080. On this machine an unrelated Apache Tomcat holds 127.0.0.1:8080, and a loopback-specific bind wins over Docker's 0.0.0.0 publish when resolving localhost — so every realm request returned 404 while the container looked healthy. If discovery fails against a locally-published container, check for another process bound specifically to loopback before suspecting the container.

A path prefix in the server URL is silently discarded. The client uses the typed address only as HttpClient.BaseAddress and every request path is root-absolute (/api/v1/meta, /.well-known/dodossh-configuration, …), so https://example.test/dodossh reaches https://example.test/api/v1/... and the prefix is dropped without a word. That rules out hosting DodoSSH under a sub-path — which is exactly what a reverse proxy in front of several services usually does. Nothing trims or normalises the typed URL either, and it is the raw string, not the parsed form, that becomes the local cache's identity. The server already publishes a canonical apiBaseUrl in its discovery document that the client could normalise against and currently ignores.

Sync:CursorSigningKey generates an ephemeral per-process key when unset. Fine for a single node; on a multi-node deployment cursors issued by one node are rejected by another, so clients resync from the beginning repeatedly. Must be configured explicitly before running more than one instance. WarnOnRiskyConfiguration logs this at startup.

Rate limiting is not implemented yet (M2). POST /api/v1/me/enrollment and the sync endpoints are reachable by any authenticated caller at any rate. Enrollment requires a valid access token and is idempotent, so the exposure is resource consumption rather than a credential-guessing surface — but it is still an unmetered write path.

/api/v1/me does not update last_seen_at_utc. Deliberate: a GET that writes on every call is a smell, and nothing depends on the value yet. Revisit when device management lands, since that is the first feature that needs it.