Merge branch 'main' into the desktop updater, and give way on two numbers

Main landed a realtime push feature while this branch was building the updater,
and the two collided in three places. Every one of them resolves the same way:
main got there first, so this branch moves.

**Two ADRs were both numbered 0012.** Main's is realtime push; this one is now
[ADR 0013](docs/adr/0013-desktop-distribution-and-updates.md). Git did not call
this a conflict — the filenames differ — so it would have merged quietly and left
the directory with two 0012s and every cross-reference ambiguous. Renumbered here
along with the nine places that point at it.

**Two manual-check phases were both numbered 15**, and that one git did catch.
Main's "Changes that arrive without a timer" keeps 15; installing and updating
the desktop client becomes Phase 16, with its checks and every reference to them
renumbered. The file's own rule is that a number is for life, which is exactly
why the one that had not been pushed is the one that gives way.

**The merge rewrote several files with CRLF**, and `.editorconfig` asks for LF on
everything except `*.ps1`. That is not cosmetic here: IDE0055 is an error and
`EnforceCodeStyleInBuild` is on, so it failed the build on three lines of
App.axaml.cs whose only change in this branch was an ADR number in a comment.
Forty-six files normalised back to LF; the release script keeps CRLF, which is
what `.gitattributes` and `.editorconfig` both already say for a PowerShell file.

Nothing else conflicted. The updater does not touch the sync loop or the event
stream, and the one file both sides edited heavily — MainWindowViewModel — merged
without a hunk in common.

Verified after merging: the solution restores locked and builds clean, and 304
shell, 100 layout, 54 session, 28 client-api and 25 contracts tests pass. The
first two counts are higher than before the merge because main's own tests came
with it and pass alongside these.
This commit is contained in:
2026-08-04 17:52:57 +02:00
52 changed files with 6205 additions and 341 deletions
+163 -25
View File
@@ -324,7 +324,8 @@ phone's, whose list has no room for a row of group cards and draws the whole tre
host filed, the grid says so in a sentence rather than sitting empty.
**Then press a group card once.** It is marked as chosen and **nothing else happens** — the grid is still the
level it was, and EDIT and DELETE now aim at that group. **Then double-press it.** The group opens: its hosts
level it was, and no buttons appear beside the GROUPS heading: editing and deleting a group are on the card's
own right-click menu, which is 7.9. **Then double-press it.** The group opens: its hosts
are the grid, the trail above the cards reads `ALL HOSTS <name> `, each card carrying the group's name as
an accent chip, and the card grid shows what is *inside* that group rather than every group in the keychain.
Pressing ALL HOSTS goes back to the outermost level.
@@ -347,9 +348,9 @@ Make two groups and file one under the other with the parent picker in the group
**Pass:** only the outer group has a card to start with. Double-press it and the inner one is the only card
shown, with the trail reading `ALL HOSTS <outer> `. Double-press that, and the cards disappear entirely —
it has nothing inside it — while the trail, EDIT and DELETE stay: with no card selected the two buttons act
on the group the trail ends with, so a group with nothing in it can still be renamed after being opened.
Pressing the **middle** crumb goes back one level rather than all the way out.
it has nothing inside it — while the trail stays. Pressing the **middle** crumb goes back one level rather
than all the way out, which is also how a group with nothing inside it is renamed: back out to the level
where it has a card, and right-click that.
**Failure means:** cards for groups that are not at this level is `VisibleGroups` having been bound past —
the flat `Groups` is the phone's and the lookups'. A group that cannot be reached at all is worse and is the
@@ -370,13 +371,35 @@ takes only the group that is open.
### 3.3 Deleting a group with hosts in it
Select a group with hosts and press DELETE.
Right-click a group with hosts in it and choose **Delete…**. Do it twice: once leaving the tick alone, and
once — on another group — ticking it.
**Pass:** the question names how many hosts are filed under it and says they stay. Agreeing removes the
group; the hosts lose their chip and are otherwise unchanged.
**Pass:** the question names how many hosts are filed under it, says they stay and move to UNGROUPED, and
offers a tick that would delete them as well. The tick starts clear, and it starts clear again on the next
group even if it was set on the last one. Left clear, agreeing removes the group and the hosts stay, without
a chip and otherwise unchanged. Ticked, the hosts go with it — and only the hosts that were filed under that
group. A group with nothing under it is asked no second question and shows no tick.
**Failure means:** if the hosts vanish, the delete is rewriting host payloads, which it must not — see
`HostGroupRepository`.
**Failure means:** a tick that carries from one question to the next is the reset in `OnPendingDeletionChanged`
having gone, and it deletes machines on the strength of a decision about a different group. Hosts that keep
the chip after an unticked delete are the unfiling not happening: they still name a group that is gone, which
is what this used to do on purpose and no longer should.
### 3.3a Moving a group to another vault · **needs a second vault**
Build `outer inner` with a host in `inner`, all in your personal vault, then right-click **outer** and
choose **Move to another vault…**. Pick the shared vault and press MOVE.
**Pass:** the panel says what travels and what does not before you press anything. Afterwards all three items
carry the destination's badge, `inner` is still inside `outer` and the host is still inside `inner` — every
one of them under an id it did not have a moment ago. The sentence names the vault, the counts, and the fact
that the group now sits at the top level if it was nested. Nothing is left behind in the vault it came from.
**Failure means:** a host under UNGROUPED in the destination is the group id having been carried across
rather than remapped — the ids are the destination's making, so every reference has to be rewritten as its
target lands. Anything still in the source vault is a partial move, which is survivable by design but should
not happen with the network up: the groups are written top-down and the hosts last, so an interruption leaves
hosts behind and never a shelf with nothing on it.
### 3.4 A group deleted on another machine · **needs two machines**
@@ -795,6 +818,17 @@ With host A selected, right-click host B and choose Delete.
wrong machine. `HostGridTests` covers both halves headlessly, so this is a confirmation that a real popup
behaves as the headless one did.
**And the same on the group cards above.** Open a group, then right-click a card inside it and choose
Delete.
**Pass:** the question names the **card**, not the group that is open — and Open on that menu goes into the
card, rather than back out to ALL HOSTS. Right-clicking the space around the group cards opens no menu.
**Failure means:** the menu is reading `GroupTarget`'s fallback, which is the group whose contents are on
screen rather than the card the pointer is on. This menu is the only way to edit or delete a group on the
desktop — there are no buttons beside the GROUPS heading any more — so a menu aimed wrongly is the whole of
the mistake.
### 7.10 Clicking a host in the palette connects
Ctrl+K, then click a result with the mouse rather than pressing Enter.
@@ -1279,6 +1313,31 @@ rather than a broken role.
under the people who share it is an administrative act reached without the role for it. The server refuses
it too — this is the interface not offering what the server would turn down.
### 12.10 A group made in a shared vault arrives as a group, not as a heap · **needs two accounts**
1. As Alice, on HOSTS, press + NEW GROUP, choose the shared vault in the editor's VAULT picker, name it
`production`, and give it a default port and username.
2. Add two hosts to the same shared vault and file them under it.
3. Make a second group called `production` in the **personal** vault.
4. Sync, then look at Bob's machine after his own sync.
**Pass on Alice's:** the two cards are told apart by the vault name printed under each — same name, two
folders — and on the phone the two headings carry the same badge. Opening either shows only its own hosts.
Dragging one of the shared vault's host cards onto the personal `production` card is **refused with a
sentence naming both vaults**, and the host stays where it was.
**Pass on Bob's:** the group is a card and a heading on his machine too, with the hosts inside it, and the
port and username they dial are the ones Alice typed into the group rather than 22 and his own account. He
can rename it, and the rename comes back to Alice rather than arriving as a second group in his personal
vault.
**Failure means:** a group that reaches Bob as UNGROUPED hosts is the resolution map having gone narrow
again — cosmetic on its own, except that the port and the username go with it, so his terminal dials the
wrong place. A rename of his that turns up as a new group in his own vault is the editor writing to the
active vault rather than to the row's, which forks the shelf and leaves Alice's untouched. Two identical
cards with no vault under them means one of them is a folder somebody outside the team can read, and
nothing on screen says which.
---
## Phase 13 — Unlocking the phone with a fingerprint
@@ -1466,9 +1525,88 @@ delivery survives the failure because it is held against the transfer rather tha
dropped on the failure. An error saying the staged file is missing is the copy having been deleted at the
stop, which is what `QueueDeliveredDownload` documents it does not do.
## Phase 15 — Changes that arrive without a timer
The socket is covered by tests on both sides: the endpoint suite opens a real one against a real
`TestServer` and proves a push produces a notice, that another account's push does not, and that a frame
carries no ciphertext; the shell suite proves a notice wakes the synchronisation loop long before the
minute. What none of that can reach is **the network in between**, and that is where this feature is most
likely to fail: a reverse proxy that will not upgrade, one that drops an idle socket without telling either
end, a corporate middlebox, a phone moving between Wi-Fi and mobile data. Every one of those looks the same
from inside a test host, which has no proxy and no radio.
The pass condition throughout is *two* things, and the second matters as much as the first: it arrives
quickly, **and** it still arrives when the socket is gone. A build where the timer had stopped working would
pass every "it was fast" check here and fail nobody until somebody's proxy changed.
### 15.1 A colleague's edit appears while you are looking at it
Two accounts sharing a vault, both unlocked, both on the Hosts screen. On the first machine, rename a host
in the shared vault and save.
**Pass:** the second machine's list shows the new name within a second or two, with nothing pressed and no
screen flicker — the row updates, the selection does not move, and the status line is not repainted with a
sync report.
**Failure means:** nothing within a minute, then the new name, is the socket not being established at all —
that is the timer doing its job, which is the correct fallback and not the feature. Check `/api/v1/meta`
lists `events`, then whether the proxy in front of the API forwards `Upgrade` and `Connection`. A list that
never updates at all is a synchronisation failure and has nothing to do with this phase.
### 15.2 A vault shared with you turns up as it is shared
The second account signed in and unlocked, sitting on the VAULTS screen. From the first, add them to a team
and press SHARE KEY.
**Pass:** the vault appears in their list within a second or two of the key being wrapped, and reads as
waiting for a key until the share, then as readable.
**Failure means:** the vault appearing only on the minute is the `vaults.changed` notice not being published
or not being followed. Both the membership add and the grant publish one; if the membership arrives promptly
and the key does not, the grant path is the one to look at.
### 15.3 It still works with the socket taken away
On the second machine, block the WebSocket — the simplest way is a proxy rule rejecting the upgrade, or
setting `Events:Enabled` to `false` on the server and restarting it.
**Pass:** everything above still happens, within the minute rather than within seconds. Nothing on the
screen says anything is wrong, because nothing is: no error, no OFFLINE badge, no repeated status message.
The Sync button still works and still reports.
**Failure means:** an error message, a titlebar claiming to be offline, or a status line that repaints with
a socket failure is the client treating an absent push channel as a fault. It is not one — the timer is the
guarantee and the socket is the optimisation, and a user with a strict proxy must never be told their
keychain is broken.
### 15.4 A laptop that slept comes back on its own
With the second machine idle and connected, close the lid for a few minutes — or disable Wi-Fi for two
minutes and re-enable it. Then make a change on the first machine.
**Pass:** the change arrives quickly again, without the vault having been locked or the application
restarted. The reconnection is invisible.
**Failure means:** changes that arrive only on the timer from then on are the stream having given up after
its socket died — the reconnection loop is what should make that impossible, and a client that reconnects
once and not twice is the specific defect its tests exist to catch. Changes that never arrive again, timer
included, are a different and worse bug in the synchronisation loop rather than in the socket.
### 15.5 An expiring token does not end the push
This one needs a short access-token lifetime in the identity provider — the dev realm's Keycloak client can
be set to a couple of minutes. Leave a machine unlocked and idle for longer than that, then make a change
elsewhere.
**Pass:** the change still arrives quickly. The socket is closed by the server at the token's expiry and the
client reconnects with a fresh one, which should be invisible.
**Failure means:** notices stopping at roughly the token's lifetime is the reconnection not asking for a new
token — it would be dialling with the spent one and being closed again immediately. A burst of reconnection
attempts in the server log is the same defect seen from the other end.
---
## Phase 15 — Installing the desktop client, and being updated by it
## Phase 16 — Installing the desktop client, and being updated by it
Nothing in this phase is reachable by a test, and not for the usual reason. There is no installed
application in CI, no `%LOCALAPPDATA%` worth inspecting, and the update path only exists across two builds
@@ -1478,12 +1616,12 @@ of the view model against a fake channel, and pins the one promise that matters
(`AReadyUpdate_IsNeverAppliedOnItsOwn`); `UpdateBannerTests` measures the banner at the window's minimum
width. Neither can install anything.
Walk it once per release, and in order — 15.6 onwards needs 15.1 to have happened.
Walk it once per release, and in order — 16.6 onwards needs 16.1 to have happened.
Run `pwsh -File scripts/release-windows.ps1` first. It stops after packing, on purpose, so that everything
below happens before anything reaches a user.
### 15.1 The installer needs no administrator, and lands beside the vault rather than on it · **the one that would destroy data**
### 16.1 The installer needs no administrator, and lands beside the vault rather than on it · **the one that would destroy data**
Run `Releases\DodoSSH.Desktop-win-Setup.exe` from an ordinary account. Then look at `%LOCALAPPDATA%`.
@@ -1493,9 +1631,9 @@ reading **DodoSSH**; and `%LOCALAPPDATA%\DodoSSH` either absent (a fresh machine
**Failure means:** a UAC prompt is a per-machine install, which is not what was designed. Anything written
into `%LOCALAPPDATA%\DodoSSH` is the pack id having drifted back to `DodoSSH`, and that is the serious one —
the uninstaller removes its whole install root, so it would take the vault cache and the outbox with it.
See [ADR 0012](adr/0012-desktop-distribution-and-updates.md) decision 2.
See [ADR 0013](adr/0013-desktop-distribution-and-updates.md) decision 2.
### 15.2 The installed path is short, measured rather than assumed
### 16.2 The installed path is short, measured rather than assumed
```powershell
"$env:LOCALAPPDATA\DodoSSH.Desktop\current\DodoSSH.exe".Length
@@ -1507,7 +1645,7 @@ See [ADR 0012](adr/0012-desktop-distribution-and-updates.md) decision 2.
`CO_E_SERVER_EXEC_FAILURE` and name nothing — see the long-path entry in `platform-flags.md`, which is the
reason this check is a number rather than a shrug.
### 15.3 The window opens and carries its own icon
### 16.3 The window opens and carries its own icon
**Pass:** the Start-menu shortcut launches it, and the taskbar and Alt-Tab show the dodo mark rather than a
generic icon.
@@ -1515,7 +1653,7 @@ generic icon.
**Failure means:** `--icon` or `ApplicationIcon` did not survive packaging. Cosmetic, and the first thing
anybody notices.
### 15.4 A terminal connects from the installed build · **the one that would catch an AppContainer**
### 16.4 A terminal connects from the installed build · **the one that would catch an AppContainer**
Sign in, unlock, open a shell against a real host, and type.
@@ -1527,7 +1665,7 @@ WebView2 in an AppContainer where the loopback data plane cannot connect. That i
out for, and this is the check that would find it in the Velopack path. `RendererTimeout` is where the
fifteen seconds comes from.
### 15.5 The version on screen is the version that was built
### 16.5 The version on screen is the version that was built
Right-click `DodoSSH.exe` → Properties → Details, and open PREFERENCES → UPDATES.
@@ -1539,7 +1677,7 @@ dropped from a checkout. `1.0.0.0` is somebody having wired the app manifest's i
version to the real one. A version on screen that differs from the file properties means the two are being
read from different places, which is the thing having one number was for.
### 15.6 A second release produces a delta, not only a full package
### 16.6 A second release produces a delta, not only a full package
Tag `v0.1.1` and run the script again.
@@ -1551,7 +1689,7 @@ previous release did not come down, so every user is about to fetch a ~60 MB ful
change. The
script warns rather than failing when that is legitimate, which is the first release only.
### 15.7 The update arrives, and the restart lands in it · **the whole point of the work**
### 16.7 The update arrives, and the restart lands in it · **the whole point of the work**
With v0.1.0 installed and running, a vault unlocked, a host change made, and **a terminal open**, publish
v0.1.1 (`-Upload`). Then press CHECK NOW on PREFERENCES rather than waiting six hours.
@@ -1566,11 +1704,11 @@ intact.
and reports the client up to date forever. A banner sliced at the terminal's left edge is the occlusion rule
having been broken, and the fallback is to move the offer into the titlebar instead. Coming back as 0.1.0 is
the swap having been blocked, usually by a process still holding a file under `current\`. Being asked to
enrol again means the profile directory did not survive, which is 15.1's failure arriving late.
enrol again means the profile directory did not survive, which is 16.1's failure arriving late.
### 15.8 The first connect after an update is not a cold start
### 16.8 The first connect after an update is not a cold start
Immediately after 15.7, connect to a host.
Immediately after 16.7, connect to a host.
**Pass:** the terminal appears about as quickly as it did before the update.
@@ -1578,7 +1716,7 @@ Immediately after 15.7, connect to a host.
see the entry in `platform-flags.md`. Slow but working, so it gets dismissed as a fluke unless somebody is
looking for it, which is why it is a numbered check rather than a note.
### 15.9 Uninstalling removes the application and leaves the vault · **the data-loss check**
### 16.9 Uninstalling removes the application and leaves the vault · **the data-loss check**
Settings → Apps → DodoSSH → Uninstall.
@@ -1587,5 +1725,5 @@ Settings → Apps → DodoSSH → Uninstall.
is wrong. Reinstalling then asks for the passphrase rather than for a server.
**Failure means:** the cache going with the application is the pack-id collision, and whoever ran this has
lost their offline unlock and any change that was still in the outbox. That is the failure ADR 0012
decision 2 exists to prevent, and it is why 15.1 checks the same thing from the other end.
lost their offline unlock and any change that was still in the outbox. That is the failure ADR 0013
decision 2 exists to prevent, and it is why 16.1 checks the same thing from the other end.