Public Access
Found while sweeping for the stale default. The client uses the typed address only as HttpClient.BaseAddress and every request path is root-absolute, so https://example.test/dodossh reaches https://example.test/api/v1/... with the prefix discarded and no error — which rules out hosting under a sub-path, the usual arrangement behind a proxy fronting several services. The server already publishes a canonical apiBaseUrl the client could normalise against and ignores. Recorded rather than fixed: it is a deployment-shape decision, not a bug in the screen that prompted this.
295 lines
20 KiB
Markdown
295 lines
20 KiB
Markdown
# Platform flags
|
|
|
|
Things known or suspected to behave differently outside Windows, plus deployment gotchas that
|
|
have already cost time once. **Development and testing are currently Windows-only**, so anything
|
|
here marked *unverified* has not run on the platform in question and must not be assumed to work.
|
|
|
|
Each entry says what the risk is, why it matters, and what to do about it. Delete an entry when it
|
|
has been verified or made moot — not when it merely stops being convenient.
|
|
|
|
## Cryptography
|
|
|
|
**`ChaCha20Poly1305.IsSupported` is false on macOS**, and on Windows builds before 10.0.20142.
|
|
This is why the client uses NSec (libsodium) rather than the BCL for content encryption; see
|
|
docs/crypto.md §1. *Already mitigated* — but if a BCL AEAD path is ever added as a fallback it
|
|
**must** gate on `IsSupported` rather than assuming availability, or the client will fail to open
|
|
any vault on macOS.
|
|
|
|
**Argon2id timings are measured on one Windows machine only.** 256 MiB with t=4 took 323 ms here.
|
|
The floor and ceiling in `EnrollmentLimits` were chosen against that number. *Unverified
|
|
elsewhere:* recalibrate on the slowest target platform before recommending a default profile,
|
|
because a cost that is comfortable on a desktop can make unlock unusable on a low-power laptop —
|
|
and the parameters are stored per user at enrollment, so a bad default is a per-user migration.
|
|
|
|
**libsodium ships native binaries per RID.** This complicates single-file and AOT publishing, and
|
|
on macOS every native library (`libsodium`, `libSkiaSharp`, `libHarfBuzzSharp`, `libe_sqlite3`)
|
|
must be signed **individually** with `--options runtime --timestamp` before the bundle is signed,
|
|
or notarization fails with an error that does not name the offending file.
|
|
|
|
## Desktop client
|
|
|
|
**The WebView runs on Windows.** `Avalonia.Controls.WebView` 12.0.1 (MIT, no licence key) hosts the
|
|
terminal page: WebView2 launches, navigates to the loopback page, runs its JavaScript and completes the
|
|
WebSocket handshake. Verified by observing an established TCP connection from `msedgewebview2` to the data
|
|
plane port.
|
|
|
|
Note precisely what that evidence covers, because it was once stretched to cover more: every clause above
|
|
is about the process and the socket. It says nothing about how the control **composites** with
|
|
Avalonia-drawn content, which is the axis on which it does not behave like an ordinary control — see the
|
|
next entry.
|
|
|
|
**A native child window cannot be covered by Avalonia content, on any platform that hosts it windowed.**
|
|
`NativeWebView` attaches a real Win32 child HWND through `NativeControlHost`, and a child window paints
|
|
above everything its parent draws, whatever the visual tree's z-order says. Layering a screen over the
|
|
terminal therefore does nothing: the WebView's rectangle stays on top. In this shell that sliced the setup
|
|
and unlock cards at the terminal column's left edge, put every one of their buttons inside the WebView's
|
|
rectangle at the window's default width — so the flow could only be completed by keyboard — and handed
|
|
Win32 focus to WebView2 on any click in that region, which makes a text box stop accepting keystrokes with
|
|
no visible cause.
|
|
|
|
The fix is to collapse the control, not to cover it: `IsVisible="{Binding IsUnlocked}"` on the
|
|
`NativeWebView`. That is safe, and this is the part worth recording, because the opposite was asserted here
|
|
for a while:
|
|
|
|
- `NativeControlHost` creates the native attachment from **attach to the visual tree**, not from layout and
|
|
not from visibility. Its `UpdateHost` never reads `IsEffectivelyVisible`; only
|
|
`TryUpdateNativeControlPosition` does, choosing `HideWithSize` over `ShowInBounds`.
|
|
- `NativeWebView` stashes a `Source` assigned before its adapter exists and replays it once created, so
|
|
navigation is never lost to ordering. The shell already depends on that replay.
|
|
- So a collapsed WebView still starts WebView2, still loads the page and still lets the renderer attach its
|
|
socket. Confirmed on Windows: 35 `msedgewebview2` processes with the control collapsed behind the setup
|
|
screen.
|
|
|
|
The previous version of this entry claimed the reverse — that hiding it would mean never realising it — and
|
|
cited the `msedgewebview2` connection as verification. That observation was made while the overlay was
|
|
showing but, because of the airspace behaviour above, the WebView was in fact uncovered and in plain view.
|
|
It confirmed only that a *visible* WebView is realised, which nobody disputed, and could not discriminate
|
|
the case it was attached to. A process-level check cannot verify a rendering claim; that needs a
|
|
screenshot, and this defect shipped because one was never taken.
|
|
|
|
**What the first connection after unlocking actually depends on** is the `await
|
|
workspace.WaitForRendererAsync()` in `VaultViewModel.ConnectAsync`, because `TerminalDataPlane.SendAsync`
|
|
drops frames when no renderer is attached rather than queueing them. That await is the invariant; the
|
|
control's visibility is not. It currently has no timeout, so a WebView2 that fails to initialise hangs
|
|
Connect with the busy flag stuck — worth fixing on its own merits.
|
|
|
|
**The Windows app manifest must declare a `supportedOS` list.** Without it the process reports a
|
|
downlevel Windows version and Avalonia's native control host fails outright — *"Unable to create child
|
|
window for native control host"* — so the WebView, and therefore the terminal, does not start at all.
|
|
`[STAThread]` on `Main` is equally mandatory: WebView2 checks the apartment state and refuses to
|
|
initialise on an MTA thread.
|
|
|
|
**WebView2 spawns a process tree, not a process.** Around 35 processes were observed for one embedded
|
|
view. That is the concrete reason the design uses one WebView hosting N terminals rather than one per
|
|
tab: twenty tabs would mean twenty of those trees.
|
|
|
|
**The Avalonia WebView on Linux remains unproven, and is still the largest risk in the plan.** The
|
|
package's own release notes say `NativeWebView` gained Linux support via a **WPE** backend
|
|
(`libwpewebkit-2.0`), which is much less widely installed than WebKitGTK — and it ships a separate
|
|
`NativeWebDialog` described as *"particularly useful for platforms like Linux where embedded WebView
|
|
controls might not be available"*, which is the vendor confirming the concern. *Unverified:* a spike
|
|
must cover Ubuntu on both Wayland and X11, Fedora KDE, and macOS 15.
|
|
|
|
`ITerminalHost` was supposed to be the seam that keeps a backend swap cheap, and it is **declared but not
|
|
implemented** — nothing in the application uses it, and the view navigates `NativeWebView.Source` directly.
|
|
Swapping backends today means editing `MainWindow.axaml` and its code-behind. That is a small job, but do
|
|
not plan around a seam that is currently only a file.
|
|
|
|
One more reason the Linux picture may be better than this entry assumes: the package also ships
|
|
`NativeWebViewCompositorHost`, a non-windowed host drawn through Avalonia's compositor. A compositor host
|
|
would not have the airspace problem described below at all. Whether it can be selected deliberately is
|
|
unknown and worth establishing during the spike, because it would change how overlays can be built.
|
|
|
|
**`Avalonia.Diagnostics` has no 12.x release** (latest is 11.3.18), so the developer tools overlay is
|
|
unavailable on Avalonia 12. Development-only, so nothing ships differently — but debugging a layout
|
|
problem currently means reasoning rather than inspecting.
|
|
|
|
**The xterm bundles are vendored, not built.** `@xterm/xterm` 6.0.0 with the fit and webgl addons, all
|
|
MIT, committed as UMD bundles under `WebAssets/vendor` and embedded as Avalonia resources. No npm or
|
|
esbuild step, so a clean clone builds with the .NET SDK alone. The cost is that upgrades are a manual
|
|
re-download; the licence and versions are recorded here so that stays visible.
|
|
|
|
**SSH.NET's `window-change` is verified working** as of 2025.1.0 — resolved, not a flag.
|
|
`ShellStream.ChangeWindowSize(columns, rows, width, height)` exists and the remote genuinely
|
|
observes it: `PtyAndResizeSpikeTests` reads `stty size` back from a real sshd after resizing, and
|
|
repeated resizes each take effect. The `IChannelSession` fallback is not needed. That suite stays
|
|
in place as a regression guard, because an upgrade that silently stopped sending the request would
|
|
present as wrapped output only after a resize — easy to misattribute to the terminal emulator.
|
|
|
|
**`ShellStream.Write` buffers and requires an explicit `Flush`.** Without one a keystroke is accepted,
|
|
reported as written, and never reaches the remote — the terminal displays output perfectly and simply
|
|
stops responding to input. SSH.NET's own `WriteLine` flushes, which is why a spike that used it never
|
|
hit this. `SshNetShellSession.WriteAsync` now flushes per write; batching would be wrong anyway, since
|
|
a terminal has to put a keystroke on the wire immediately.
|
|
|
|
**`ShellStream` does not override `ReadAsync`.** The base `Stream` implementation therefore runs
|
|
the blocking `Read` on a thread-pool thread, so every open session parks one thread for as long as
|
|
it is idle. Fine for the handful of tabs M1 targets; revisit before advertising many concurrent
|
|
sessions, since the fix is either an upstream change or driving `IChannelSession` directly.
|
|
|
|
**SSH.NET cannot share one connection between `SshClient` and `SftpClient`.** A shell plus SFTP to
|
|
the same host means two TCP connections, two authentications and — later — two relay sockets.
|
|
Connect SFTP lazily and reuse the cached decrypted credential so the user is not prompted twice.
|
|
|
|
**Agent forwarding is de-scoped from v1.** It needs an upstream SSH.NET change. A vault-backed
|
|
agent of our own plus ProxyJump covers the real use cases.
|
|
|
|
**The SSH suite pulls `linuxserver/openssh-server` from Docker Hub**, which is rate-limited for
|
|
unauthenticated pulls. If CI starts failing on image pulls rather than on tests, that is why.
|
|
|
|
**MSIX packaging is ruled out, not merely deprioritised.** A packaged app runs WebView2 in an
|
|
AppContainer where loopback connections are blocked without a `CheckNetIsolation` exemption. The
|
|
terminal data plane *is* a loopback WebSocket, so MSIX would break the product outright. Velopack
|
|
for Windows/macOS/AppImage; Flatpak and deb/rpm defer updates to the package manager.
|
|
|
|
**Linux ships AppImage and Flatpak first**, specifically so the WebKit runtime is bundled rather
|
|
than assumed present on the user's machine.
|
|
|
|
**Opening the system browser depends on the platform handler.** `SystemBrowserLauncher` uses
|
|
`UseShellExecute`, which delegates to `ShellExecute` on Windows, `open` on macOS and `xdg-open` on
|
|
Linux. *Unverified off Windows:* `xdg-open` comes from `xdg-utils`, which is not guaranteed on a
|
|
minimal desktop or inside a Flatpak sandbox — where the portal is the correct route instead. If
|
|
sign-in silently does nothing on Linux, this is the first thing to check. `IBrowserLauncher` exists
|
|
so a platform-specific opener can be substituted without touching the flow.
|
|
|
|
## Identity provider
|
|
|
|
**A loopback redirect URI must be registered without a port, not with a wildcard port.** Keycloak — and
|
|
providers implementing RFC 8252 §7.3 generally — ignores the port when the registered redirect URI's host
|
|
is a loopback literal, which is what lets a native client bind an ephemeral port. Registering
|
|
`http://127.0.0.1:*/callback` looks more explicit and is *broken*: the `*` is parsed as a literal port and
|
|
every real authorization request comes back `400 Invalid parameter: redirect_uri`. Keycloak's wildcard
|
|
support is trailing-only, so a `*` in the middle of a URI never means what it looks like.
|
|
|
|
Register `http://127.0.0.1/callback`. Keep the path — it is the part that stops another process on the
|
|
machine having an authorization code delivered to a different endpoint. `Oidc:LoopbackRedirectPattern`,
|
|
which the server advertises through `/.well-known/dodossh-configuration`, says the same thing so an
|
|
operator configuring a different provider copies something that works.
|
|
|
|
Found by running the sign-in against a real Keycloak; every test until then used a stub that accepted
|
|
whatever it was given.
|
|
|
|
**Keycloak marks its session cookies `Secure` even over plain HTTP**, because `SameSite=None` is only
|
|
legal alongside `Secure`. A spec-conformant HTTP client therefore refuses to store them from an `http://`
|
|
origin — .NET's `CookieContainer` drops every one silently — and the login form POST then comes back
|
|
`400` with no explanation at all. Browsers complete the flow because they treat loopback as a trustworthy
|
|
origin and make the exception.
|
|
|
|
This does not affect the product: the client uses the system browser, which makes that exception. It does
|
|
affect any non-browser automation against a development Keycloak, which has to carry the cookies by hand
|
|
(see `ScriptedBrowser`) or be given HTTPS. Two hours of "the credentials must be wrong".
|
|
|
|
**`--import-realm` skips a realm that already exists.** Editing `deploy/keycloak/realm-dodossh.json` and
|
|
running `docker compose restart keycloak` therefore changes nothing, and the stale configuration keeps
|
|
being served — which reads exactly like the edit being wrong. `start-dev` keeps its state in an H2
|
|
database inside the container, so the realm has to be recreated along with it:
|
|
`docker compose rm -sf keycloak && docker compose up -d keycloak`. Cost an otherwise inexplicable
|
|
debugging detour.
|
|
|
|
`DodoSSH.SystemTests` is immune to this by construction — its Keycloak is created and destroyed per run —
|
|
which is a second reason the end-to-end suite starts its own containers rather than reusing the developer's
|
|
stack. Editing the realm file and rerunning the suite always tests the edit.
|
|
|
|
## Local cache
|
|
|
|
**The cache location is per-OS and must stay non-roaming.** `ClientPaths` chooses it:
|
|
`%LOCALAPPDATA%\DodoSSH` on Windows, `~/Library/Application Support/DodoSSH` on macOS,
|
|
`$XDG_DATA_HOME/dodossh` or `~/.local/share/dodossh` on Linux. It must **not** land anywhere that syncs
|
|
to a cloud drive or roams: two machines writing one SQLite file through a file-sync client corrupts it,
|
|
and the whole point of the outbox is that each machine has its own. That is also why Windows uses
|
|
`%LOCALAPPDATA%` and not `%APPDATA%`, which roams in a domain environment.
|
|
|
|
The platform branches are explicit rather than delegating to
|
|
`Environment.SpecialFolder.LocalApplicationData` everywhere, because on macOS the runtime maps that to
|
|
`~/.local/share` rather than to `~/Library/Application Support`. *Verified on Windows only* — the client
|
|
created `%LOCALAPPDATA%\DodoSSH\cache.db` and migrated it on first launch. The macOS and Linux branches
|
|
are reasoned, not run.
|
|
|
|
**SQLite timestamps are stored as integers, deliberately.** EF's default `DateTimeOffset` mapping for
|
|
SQLite is a text form it then refuses to order or compare, so any query that sorts or filters by time
|
|
throws at execution rather than at model build. `UnixMillisecondsConverter` is applied as a convention
|
|
so a timestamp added later cannot be the one left unconverted. This is provider behaviour, not
|
|
platform behaviour, but it cost a debugging session and will again if the converter is removed.
|
|
|
|
**The cache is three files, not one.** EF Core's SQLite provider puts the database in WAL mode, which is
|
|
the right mode here — a background sync pass writes while the interface reads, and under the default
|
|
rollback journal those reads would fail busy — but it means `cache.db` is accompanied by `cache.db-wal`
|
|
and `cache.db-shm`. Any backup, export or uninstall routine that touches only `cache.db` is wrong.
|
|
Verified by launching the client and reading `PRAGMA journal_mode`, after a comment in the code claimed
|
|
the opposite.
|
|
|
|
**Pooled SQLite connections keep the file open after the last context is disposed.** On Windows that
|
|
means locked, so the application cannot delete or replace its own cache and a test cannot clean up after
|
|
itself. `ClientCacheFactory.Dispose` clears the pool for exactly this reason; removing that line makes
|
|
the failure appear only on Windows.
|
|
|
|
**No SQLCipher, on any platform.** The rows are already ciphertext from the server, so an encrypted
|
|
database file would protect bytes that are protected already at the cost of a native dependency and a
|
|
licence obligation — and `bundle_e_sqlcipher` was deprecated in SQLitePCLRaw 3.0. The consequence to
|
|
be honest about: the cache offers no protection against another process running as the same user. See
|
|
`LocalCacheProtector` for what it does and does not defend against.
|
|
|
|
## Build and CI
|
|
|
|
**Integration tests need a Docker daemon** (Testcontainers). They run on `ubuntu-latest` in CI.
|
|
macOS runners have no Docker daemon, and the Windows CI job is deliberately build-only. So
|
|
anything proved by an integration test is proved on Linux only — which is the right place for
|
|
server code, and no coverage at all for client platform behaviour.
|
|
|
|
**The end-to-end suite launches the API's own launcher executable**, falling back to `dotnet exec` on the
|
|
assembly. The fallback exists for one reason: a checkout or artefact copy that lost the execute bit
|
|
produces a `Win32Exception` on Linux and nothing whatsoever on Windows. *Verified on Windows only* — the
|
|
launcher path is what runs here, so the fallback itself is reasoned rather than exercised. If the suite
|
|
fails in CI with a permission error before any container work, that is the path to look at.
|
|
|
|
**It also depends on `Server:PublicBaseUrl` being knowable before startup.** The port is chosen by binding
|
|
a loopback socket and releasing it, because the API reads that URL at startup and advertises it to clients,
|
|
so it cannot be discovered from Kestrel afterwards. The window for another process to take the port is a
|
|
few milliseconds; if the suite ever fails with an address-in-use, this is why, and a retry is the fix
|
|
rather than a redesign.
|
|
|
|
**`[CallerFilePath]` is rewritten to `/_/...` under `ContinuousIntegrationBuild`.** Any test that
|
|
locates a fixture by source path passes locally and fails in CI. Copy fixtures to the output
|
|
directory and read them via `AppContext.BaseDirectory` instead; `GoldenVectorTests` shows the
|
|
pattern.
|
|
|
|
**`dotnet format --verify-no-changes` is part of the CI gate** and exits non-zero on style
|
|
warnings, not just whitespace. Run it before pushing; a build with zero warnings can still fail
|
|
that step.
|
|
|
|
## Deployment
|
|
|
|
**PostgreSQL 18 moved its data directory** to `/var/lib/postgresql`, not `/var/lib/postgresql/data`
|
|
as in 17 and earlier. A compose file carried over from an older version silently gets an empty
|
|
volume — the database appears to work and loses everything on restart. Relevant to any compose
|
|
file other than `deploy/docker-compose.dev.yml`, which is already correct.
|
|
|
|
**Keycloak in the dev stack listens on host port 18080, not 8080.** On this machine an unrelated
|
|
Apache Tomcat holds `127.0.0.1:8080`, and a loopback-specific bind wins over Docker's `0.0.0.0`
|
|
publish when resolving `localhost` — so every realm request returned 404 while the container
|
|
looked healthy. If discovery fails against a locally-published container, check for another
|
|
process bound specifically to loopback before suspecting the container.
|
|
|
|
**A path prefix in the server URL is silently discarded.** The client uses the typed address only as
|
|
`HttpClient.BaseAddress` and every request path is root-absolute (`/api/v1/meta`,
|
|
`/.well-known/dodossh-configuration`, …), so `https://example.test/dodossh` reaches
|
|
`https://example.test/api/v1/...` and the prefix is dropped without a word. That rules out hosting DodoSSH
|
|
under a sub-path — which is exactly what a reverse proxy in front of several services usually does. Nothing
|
|
trims or normalises the typed URL either, and it is the raw string, not the parsed form, that becomes the
|
|
local cache's identity. The server already publishes a canonical `apiBaseUrl` in its discovery document
|
|
that the client could normalise against and currently ignores.
|
|
|
|
**`Sync:CursorSigningKey` generates an ephemeral per-process key when unset.** Fine for a single
|
|
node; on a multi-node deployment cursors issued by one node are rejected by another, so clients
|
|
resync from the beginning repeatedly. Must be configured explicitly before running more than one
|
|
instance. `WarnOnRiskyConfiguration` logs this at startup.
|
|
|
|
**Rate limiting is not implemented yet** (M2). `POST /api/v1/me/enrollment` and the sync endpoints
|
|
are reachable by any authenticated caller at any rate. Enrollment requires a valid access token
|
|
and is idempotent, so the exposure is resource consumption rather than a credential-guessing
|
|
surface — but it is still an unmetered write path.
|
|
|
|
**`/api/v1/me` does not update `last_seen_at_utc`.** Deliberate: a GET that writes on every call is
|
|
a smell, and nothing depends on the value yet. Revisit when device management lands, since that is
|
|
the first feature that needs it.
|