Merge branch 'main' into the group's move and its deletion question

Main took the group's EDIT and DELETE off the GROUPS heading while this branch
was adding a MOVE beside them, so the conflict was about the same six pixels
from both directions. Main's answer wins outright, and it is the better one for
the reason its own message gives: a button beside a heading has no card under a
pointer to mean, and had to work its subject out from the selection or from the
trail. Moving a group had that problem worst of all — the thing it takes with it
is everything on the shelf, and "which shelf" is not a question a button there
could answer plainly.

So the MOVE button is gone and the menu entry it was drawn beside is the whole
of it. That entry was already in this branch, above the separator DELETE sits
below, and it needed no change: the card menu selects whatever was right-clicked
before it runs anything, which is exactly the aiming a group move wants.

Three things went with the button. ShowsGroupActions, which main deleted because
hiding buttons was all it did, and which this branch had extended to hide them
for the move panel as well. CanMoveGroupTarget, which existed to answer whether
that button was worth drawing — CanMoveSelectedHost stays, because the phone
really does leave the host's MOVE out rather than offer a refusal, and a menu
whose entries came and went would be a menu whose items move. And the two test
assertions that read them, which were describing the button rather than the
behaviour; what they were guarding is that the two panels never share the
moment, and IsConfirmingGroupDeletion says that directly.

The move panel and the deletion question both keep their place under the
heading, which is where the buttons were and is now simply where that section
puts things. They still exclude each other, by disarming rather than by a
visibility flag: MoveGroup clears a pending deletion and DeleteGroup folds the
move panel away.

Manual checks 3.3 was rewritten by main for the menu and by this branch for the
tick, and now says both; 3.3a is new and walks a two-level shelf across a vault
boundary, which is the half of this feature no headless test can watch land.
This commit is contained in:
2026-08-04 17:11:06 +02:00
38 changed files with 3416 additions and 165 deletions
+5 -1
View File
@@ -58,7 +58,11 @@ the Npgsql connection string** — the default; do not enable multiplexing.
- One place enforces revision, change-log and ACL invariants. That halves both the endpoint
count and the authorization surface, which is the main reason for the single write path.
- Delta pull makes frequent polling cheap, so multi-device feels live; push notification over
SSE or the existing WebSocket can layer on with polling as the fallback.
SSE or the existing WebSocket can layer on with polling as the fallback. **That has since been
built — see [ADR 0012](0012-realtime-push.md)** — and nothing in this ADR changed to accommodate
it. The socket carries a notice naming a vault and a sequence, whose answer is the delta pull
above, so there is still exactly one path that applies a change; and polling is still what
guarantees a pass rather than a legacy route kept for old clients.
- Conflict resolution is entirely client-side. The client retains a `BaseCiphertext` common
ancestor and performs a field-level three-way merge for structured items, or creates a
visible conflicted copy for opaque ones. **It must never silently drop a key or a host.**
+167
View File
@@ -0,0 +1,167 @@
# ADR 0012 — A WebSocket that carries notices, not data
- Status: accepted
- Date: 2026-08-04
## Context
[ADR 0003](0003-sync-protocol.md) built a delta pull that is cheap enough to run on a timer, and
the client does: one pass a minute. That is the difference between a colleague's change appearing
"soon" and appearing *now*, and it shows up in three places that are not equally forgivable.
- **A vault shared with you** arrives on the next pass. `AdmitNewVaultsAsync` says so in its own
remarks — "the recipient is handed nothing — there is no push channel" — and the README repeats
it. Sharing works and looks broken.
- **Two people editing one keychain** see each other up to a minute late, which is long enough to
make the same edit twice and produce a conflict that nobody needed to have.
- **A revoked grant** keeps serving a client that has not noticed yet, for up to a pass.
Shortening the interval is the obvious answer and the wrong one: it costs a request per client per
interval whether or not anything happened, and it does not converge on *immediate* — it converges
on a busier server that is still late.
There is also a second thing coming that this decision has to not preclude. The intended feature is
a **shared terminal session** — one person's shell, watched or driven by another, TeamViewer-shaped.
That is bidirectional, continuous, and latency-sensitive in a way a keychain notice is not.
## Decision
### One WebSocket per signed-in client, at `GET /api/v1/events`
Subprotocol `dodossh.events.v1`. The client opens it after unlock and keeps it open; the server
sends a notice whenever something the client can read has changed.
**Not SSE.** Server-sent events would carry today's notices perfectly well and would be less code.
It is one-directional, so the shared-session feature would need a second mechanism next to it, and
then two transports would need reconnection, authorization and lifetime rules that agree. The cost
of a WebSocket over SSE is small; the cost of two transports is not.
**Not SignalR.** It brings hub protocol negotiation, its own serialisation and transport fallbacks,
none of which are wanted here: `DodoSSH.Contracts` and its source-generated serialiser are "the
actual contract between the two sides", and a second wire format alongside it is exactly the silent
drift `Setup/Json.cs` records having already cost this project once.
### The notice carries no ciphertext
A `vault.changed` frame is `{ kind, vaultId, sequence }` and nothing else. The client's answer to it
is the pull it would have done on the timer anyway.
This is the load-bearing decision, and it is worth being explicit about why the tempting alternative
is refused. Pushing the changed items themselves would save a round trip and would fork the code
path that applies a change into two — one that arrives by pull and one that arrives by socket — with
the cursor, the merge and the tombstone rules duplicated across both. ADR 0003 put every mutation
through one write path for exactly that reason; this keeps every *read* on one path for the same
one. The socket decides *when* to sync. It never decides *what* a vault contains.
It also means a dropped notice is harmless, which is what lets everything below be simple.
### Polling stays, and is the fallback rather than a legacy path
The one-minute pass is unchanged. The socket makes it *early*; it does not make it *necessary*. A
client on a network that eats WebSockets, an older client, a server that has the feature off, a
notice dropped under backpressure, a second API replica that did not see the write — every one of
those degrades to what the product does today, which is correct and up to a minute late.
Nothing may be reachable only by socket. That is a rule about future features, not an observation
about this one.
### Authorization: the bearer token on the upgrade, not a ticket
[ADR 0004](0004-relay-authorization.md) gives the relay a two-step ticket so its WebSocket carries
no API authority. This one goes the other way and takes the ordinary bearer JWT on the upgrade
request, which is an ordinary authenticated HTTP request. The difference is not inconsistency:
- The relay's socket is a **byte pipe to a third party**, and its whole authorization decision —
which host, which IPs, which port — is made *before* the socket opens and never revisited. It is
also the extraction seam for a standalone relay process that must not hold ACL code.
- This socket is a **view of the caller's own vault list**, and it has to keep answering "what may
this account read" for as long as it is open. It needs the full ACL context, in-process, for the
life of the connection. A ticket would carry that context in a token instead, and it would be
wrong the moment the account's access changed.
A long-lived connection authorised by a short-lived token is the problem this creates, and it is met
head-on rather than ignored:
1. **The socket does not outlive the token.** The `exp` claim is read at accept, and the connection
is closed with `4401` when it passes. The client reconnects with a fresh token; that is a
sub-second gap in a channel whose failure mode is already "poll instead".
2. **The vault set is re-resolved periodically** (`Events:AccessRefreshInterval`, default five
minutes) as well as on the changes that are known to affect it. A withdrawn grant therefore stops
producing notices within that window at the latest, and immediately in the ordinary case.
Both are bounds on **metadata** — the fact that a vault changed and roughly when — because that is
all a notice contains. Nobody's ciphertext is behind this socket, and a client that stayed subscribed
one interval too long could still not read a byte of it: reading requires a vault key grant, which
this server has never held.
### The frames
Text frames, JSON, `DodoSshJsonContext`. Server to client:
| kind | meaning |
| --- | --- |
| `hello` | accepted; carries the heartbeat interval and the vault count subscribed |
| `vault.changed` | `vaultId` moved to `sequence`; pull it |
| `vaults.changed` | the set of vaults this account can reach is different; re-read it |
| `ping` | heartbeat; the client answers `pong` |
Client to server: `ping`, answered with `pong`. Nothing else — subscription is decided by the server
from the caller's access, not asked for by the client, because a client that could ask to subscribe
to a vault id is a client that can probe for vault ids.
`kind` is a **string**, not an enum, and that is deliberate. `UseStringEnumConverter` throws on a
value it does not know, so a newer server sending a kind an older client has never heard of would
not add an unknown frame — it would break that client's socket entirely. A string is ignored
instead, which is what makes the table above extensible. `ProblemCodes` is the same shape for the
same reason.
### Where the shared session will attach
The socket is the seam, and one thing about it is chosen now so that it need not be renegotiated
later: **session data will be binary frames on this same connection, not JSON on the table above.**
Terminal output base64'd into a JSON envelope would cost a third of the bandwidth for nothing, on
the one payload here that is continuous rather than occasional. Control — offer, accept, resize,
end — is JSON like everything else.
That is as far as this ADR goes. Two questions are open and are not being answered by implication:
whether a shared session's bytes go through the API at all or peer-to-peer past it, and what
end-to-end encryption means when the second party is watching a stream rather than holding a key.
Both are ADR 0001 questions and deserve their own decision. What this one buys is that they will not
also be transport questions.
## Consequences
- **Fan-out is in-process, and the deployment is therefore single-node for this feature.** Every
connection is held by the node that accepted it; a write handled by another node produces no
notice on this one. `IVaultEventPublisher` is the seam a backplane implements — PostgreSQL
`LISTEN`/`NOTIFY` needs no infrastructure this stack does not already run — and it is deliberately
**not implemented**, because an untested backplane is worse than a documented gap. Multiple API
replicas do not break: they degrade to polling, which is the state before this ADR. `/api/v1/meta`
advertises `events` so a client knows which it is getting.
- **Per-connection queues are bounded and drop the oldest.** A notice is "pull vault X, which is at
least at sequence N", so the newest is strictly more useful than the one it displaces and the
client's answer is identical either way. A slow reader costs itself latency, never the publisher's
progress — the publish path never blocks and never awaits a socket.
- **The publish happens after the transaction commits**, outside the advisory lock ADR 0003 takes.
A notice sent from inside it would name a sequence a reader cannot yet see, and would hold the
per-vault write lock across a socket write.
- **A client is notified of its own writes.** It pushed, so it already pulled; the extra pass finds
nothing. The client coalesces notices over a short window rather than the server suppressing an
echo, because suppressing it correctly needs a per-*device* identity on the socket and the same
user's other machines must still be told.
- **Connections are capped** per user and per node (`Events:MaxConnectionsPerUser`,
`Events:MaxConnectionsTotal`). A socket is cheap but not free, and an unbounded count of them is a
denial of service that authenticates first.
- The feature can be turned off entirely (`Events:Enabled`). A deployment behind a proxy that will
not upgrade should say so rather than have every client discover it by failing.
### Rejected
- **Shorter polling.** Cheaper to build, converges on a busier server that is still late.
- **Long polling.** No new transport and genuinely immediate, but it holds a request thread and a
connection per client for the same money as a WebSocket while offering none of the bidirectionality
the shared session needs.
- **Pushing the changed items down the socket.** Saves a round trip; forks the apply path in two. See
above.
- **Client-chosen subscriptions.** A `subscribe(vaultId)` frame is an existence oracle for vault ids,
which is the disclosure `SyncPullEndpoint` answers 404 rather than 403 to avoid.
+1 -1
View File
@@ -157,7 +157,7 @@ the chrome, hosts and terminals, file transfer, the vault, teams, and preference
> | **Add Telnet**, and **Serial** in the toolbar | Omitted. `ISshConnection` is the only transport there is. This is also why the card subtitle's `ssh` is a constant today rather than a reading — it is stated in `HostRowViewModel.Summary`, which is the one place in this interface where a constant is printed on purpose. |
> | **+ SSH ID, Certificate, FIDO2** | Omitted. `IDENTITIES` and `CERTIFICATES` have been on this document's list since the first import — neither is even a reserved `SyncEntityType` — and there is no security-key path anywhere in the SSH layer. One control offering three item types that do not exist. |
> | The **Backspace / Default** row | Omitted. It is a terminal setting, and the client has no preferences store and no frame to carry one to the renderer — see the Preferences section. It would be a control whose value could not survive the window closing. |
> | The **chevron beside the vault name** | The name alone, and the move behind the pane's ⋯ menu instead. A host *can* now be moved between vaults, so the gap is no longer that there is nothing to offer — it is that a chevron on a subtitle implies an edit, and this is not one: the two vaults are encrypted under different keys, so it is a re-seal into one and a tombstone in the other, the host takes a new id, and its group and tags stay behind. A control that implied "just change this field" would be describing something else. Where a *new* host goes is still asked in the host editor, as a picker beside the name. A group moves too, from MOVE over the group cards, and takes its nested groups and every host filed under them; keys, passwords and buckets take theirs from the keychain screen's standing picker and cannot be moved yet. |
> | The **chevron beside the vault name** | The name alone, and the move behind the pane's ⋯ menu instead. A host *can* now be moved between vaults, so the gap is no longer that there is nothing to offer — it is that a chevron on a subtitle implies an edit, and this is not one: the two vaults are encrypted under different keys, so it is a re-seal into one and a tombstone in the other, the host takes a new id, and its group and tags stay behind. A control that implied "just change this field" would be describing something else. Where a *new* host goes is still asked in the host editor, as a picker beside the name. A group moves too, from its card's right-click menu, and takes its nested groups and every host filed under them; keys, passwords and buckets take theirs from the keychain screen's standing picker and cannot be moved yet. |
> | **Show more ⌄** | Not drawn as a disclosure. What it would hide — notes, the relay switch, forgetting the host key — is in the editor, one press away, and a second fold inside a pane that already scrolls is a second place for a field to be missing from. |
> | **Port Forwarding** in the sidebar | Nothing, for the third time in this document. |
> | The host grid's toolbar avatar, share and tag-filter controls | Omitted, as in v3 and for the same reasons. |
+115 -11
View File
@@ -324,7 +324,8 @@ phone's, whose list has no room for a row of group cards and draws the whole tre
host filed, the grid says so in a sentence rather than sitting empty.
**Then press a group card once.** It is marked as chosen and **nothing else happens** — the grid is still the
level it was, and EDIT and DELETE now aim at that group. **Then double-press it.** The group opens: its hosts
level it was, and no buttons appear beside the GROUPS heading: editing and deleting a group are on the card's
own right-click menu, which is 7.9. **Then double-press it.** The group opens: its hosts
are the grid, the trail above the cards reads `ALL HOSTS <name> `, each card carrying the group's name as
an accent chip, and the card grid shows what is *inside* that group rather than every group in the keychain.
Pressing ALL HOSTS goes back to the outermost level.
@@ -347,9 +348,9 @@ Make two groups and file one under the other with the parent picker in the group
**Pass:** only the outer group has a card to start with. Double-press it and the inner one is the only card
shown, with the trail reading `ALL HOSTS <outer> `. Double-press that, and the cards disappear entirely —
it has nothing inside it — while the trail, EDIT and DELETE stay: with no card selected the two buttons act
on the group the trail ends with, so a group with nothing in it can still be renamed after being opened.
Pressing the **middle** crumb goes back one level rather than all the way out.
it has nothing inside it — while the trail stays. Pressing the **middle** crumb goes back one level rather
than all the way out, which is also how a group with nothing inside it is renamed: back out to the level
where it has a card, and right-click that.
**Failure means:** cards for groups that are not at this level is `VisibleGroups` having been bound past —
the flat `Groups` is the phone's and the lookups'. A group that cannot be reached at all is worse and is the
@@ -370,13 +371,35 @@ takes only the group that is open.
### 3.3 Deleting a group with hosts in it
Select a group with hosts and press DELETE.
Right-click a group with hosts in it and choose **Delete…**. Do it twice: once leaving the tick alone, and
once — on another group — ticking it.
**Pass:** the question names how many hosts are filed under it and says they stay. Agreeing removes the
group; the hosts lose their chip and are otherwise unchanged.
**Pass:** the question names how many hosts are filed under it, says they stay and move to UNGROUPED, and
offers a tick that would delete them as well. The tick starts clear, and it starts clear again on the next
group even if it was set on the last one. Left clear, agreeing removes the group and the hosts stay, without
a chip and otherwise unchanged. Ticked, the hosts go with it — and only the hosts that were filed under that
group. A group with nothing under it is asked no second question and shows no tick.
**Failure means:** if the hosts vanish, the delete is rewriting host payloads, which it must not — see
`HostGroupRepository`.
**Failure means:** a tick that carries from one question to the next is the reset in `OnPendingDeletionChanged`
having gone, and it deletes machines on the strength of a decision about a different group. Hosts that keep
the chip after an unticked delete are the unfiling not happening: they still name a group that is gone, which
is what this used to do on purpose and no longer should.
### 3.3a Moving a group to another vault · **needs a second vault**
Build `outer inner` with a host in `inner`, all in your personal vault, then right-click **outer** and
choose **Move to another vault…**. Pick the shared vault and press MOVE.
**Pass:** the panel says what travels and what does not before you press anything. Afterwards all three items
carry the destination's badge, `inner` is still inside `outer` and the host is still inside `inner` — every
one of them under an id it did not have a moment ago. The sentence names the vault, the counts, and the fact
that the group now sits at the top level if it was nested. Nothing is left behind in the vault it came from.
**Failure means:** a host under UNGROUPED in the destination is the group id having been carried across
rather than remapped — the ids are the destination's making, so every reference has to be rewritten as its
target lands. Anything still in the source vault is a partial move, which is survivable by design but should
not happen with the network up: the groups are written top-down and the hosts last, so an interruption leaves
hosts behind and never a shelf with nothing on it.
### 3.4 A group deleted on another machine · **needs two machines**
@@ -802,8 +825,9 @@ Delete.
card, rather than back out to ALL HOSTS. Right-clicking the space around the group cards opens no menu.
**Failure means:** the menu is reading `GroupTarget`'s fallback, which is the group whose contents are on
screen. That fallback is right for the EDIT and DELETE buttons beside the heading and wrong for a menu that
opened on a card.
screen rather than the card the pointer is on. This menu is the only way to edit or delete a group on the
desktop — there are no buttons beside the GROUPS heading any more — so a menu aimed wrongly is the whole of
the mistake.
### 7.10 Clicking a host in the palette connects
@@ -1500,3 +1524,83 @@ delivery survives the failure because it is held against the transfer rather tha
**Failure means:** a retry that succeeds but leaves the destination empty is the delivery having been
dropped on the failure. An error saying the staged file is missing is the copy having been deleted at the
stop, which is what `QueueDeliveredDownload` documents it does not do.
## Phase 15 — Changes that arrive without a timer
The socket is covered by tests on both sides: the endpoint suite opens a real one against a real
`TestServer` and proves a push produces a notice, that another account's push does not, and that a frame
carries no ciphertext; the shell suite proves a notice wakes the synchronisation loop long before the
minute. What none of that can reach is **the network in between**, and that is where this feature is most
likely to fail: a reverse proxy that will not upgrade, one that drops an idle socket without telling either
end, a corporate middlebox, a phone moving between Wi-Fi and mobile data. Every one of those looks the same
from inside a test host, which has no proxy and no radio.
The pass condition throughout is *two* things, and the second matters as much as the first: it arrives
quickly, **and** it still arrives when the socket is gone. A build where the timer had stopped working would
pass every "it was fast" check here and fail nobody until somebody's proxy changed.
### 15.1 A colleague's edit appears while you are looking at it
Two accounts sharing a vault, both unlocked, both on the Hosts screen. On the first machine, rename a host
in the shared vault and save.
**Pass:** the second machine's list shows the new name within a second or two, with nothing pressed and no
screen flicker — the row updates, the selection does not move, and the status line is not repainted with a
sync report.
**Failure means:** nothing within a minute, then the new name, is the socket not being established at all —
that is the timer doing its job, which is the correct fallback and not the feature. Check `/api/v1/meta`
lists `events`, then whether the proxy in front of the API forwards `Upgrade` and `Connection`. A list that
never updates at all is a synchronisation failure and has nothing to do with this phase.
### 15.2 A vault shared with you turns up as it is shared
The second account signed in and unlocked, sitting on the VAULTS screen. From the first, add them to a team
and press SHARE KEY.
**Pass:** the vault appears in their list within a second or two of the key being wrapped, and reads as
waiting for a key until the share, then as readable.
**Failure means:** the vault appearing only on the minute is the `vaults.changed` notice not being published
or not being followed. Both the membership add and the grant publish one; if the membership arrives promptly
and the key does not, the grant path is the one to look at.
### 15.3 It still works with the socket taken away
On the second machine, block the WebSocket — the simplest way is a proxy rule rejecting the upgrade, or
setting `Events:Enabled` to `false` on the server and restarting it.
**Pass:** everything above still happens, within the minute rather than within seconds. Nothing on the
screen says anything is wrong, because nothing is: no error, no OFFLINE badge, no repeated status message.
The Sync button still works and still reports.
**Failure means:** an error message, a titlebar claiming to be offline, or a status line that repaints with
a socket failure is the client treating an absent push channel as a fault. It is not one — the timer is the
guarantee and the socket is the optimisation, and a user with a strict proxy must never be told their
keychain is broken.
### 15.4 A laptop that slept comes back on its own
With the second machine idle and connected, close the lid for a few minutes — or disable Wi-Fi for two
minutes and re-enable it. Then make a change on the first machine.
**Pass:** the change arrives quickly again, without the vault having been locked or the application
restarted. The reconnection is invisible.
**Failure means:** changes that arrive only on the timer from then on are the stream having given up after
its socket died — the reconnection loop is what should make that impossible, and a client that reconnects
once and not twice is the specific defect its tests exist to catch. Changes that never arrive again, timer
included, are a different and worse bug in the synchronisation loop rather than in the socket.
### 15.5 An expiring token does not end the push
This one needs a short access-token lifetime in the identity provider — the dev realm's Keycloak client can
be set to a couple of minutes. Leave a machine unlocked and idle for longer than that, then make a change
elsewhere.
**Pass:** the change still arrives quickly. The socket is closed by the server at the token's expiry and the
client reconnects with a fresh one, which should be invisible.
**Failure means:** notices stopping at roughly the token's lifetime is the reconnection not asking for a new
token — it would be dialling with the spent one and being closed again immediately. A burst of reconnection
attempts in the server log is the same defect seen from the other end.