Merge branch 'claude/vault-realtime-push-d64c61'
ci / build and test (push) Successful in 1m33s
ci / android head (push) Failing after 5s
ci / api image (push) Canceled after 36s

This commit is contained in:
2026-08-04 16:38:42 +02:00
31 changed files with 3285 additions and 13 deletions
+5 -1
View File
@@ -58,7 +58,11 @@ the Npgsql connection string** — the default; do not enable multiplexing.
- One place enforces revision, change-log and ACL invariants. That halves both the endpoint
count and the authorization surface, which is the main reason for the single write path.
- Delta pull makes frequent polling cheap, so multi-device feels live; push notification over
SSE or the existing WebSocket can layer on with polling as the fallback.
SSE or the existing WebSocket can layer on with polling as the fallback. **That has since been
built — see [ADR 0012](0012-realtime-push.md)** — and nothing in this ADR changed to accommodate
it. The socket carries a notice naming a vault and a sequence, whose answer is the delta pull
above, so there is still exactly one path that applies a change; and polling is still what
guarantees a pass rather than a legacy route kept for old clients.
- Conflict resolution is entirely client-side. The client retains a `BaseCiphertext` common
ancestor and performs a field-level three-way merge for structured items, or creates a
visible conflicted copy for opaque ones. **It must never silently drop a key or a host.**
+167
View File
@@ -0,0 +1,167 @@
# ADR 0012 — A WebSocket that carries notices, not data
- Status: accepted
- Date: 2026-08-04
## Context
[ADR 0003](0003-sync-protocol.md) built a delta pull that is cheap enough to run on a timer, and
the client does: one pass a minute. That is the difference between a colleague's change appearing
"soon" and appearing *now*, and it shows up in three places that are not equally forgivable.
- **A vault shared with you** arrives on the next pass. `AdmitNewVaultsAsync` says so in its own
remarks — "the recipient is handed nothing — there is no push channel" — and the README repeats
it. Sharing works and looks broken.
- **Two people editing one keychain** see each other up to a minute late, which is long enough to
make the same edit twice and produce a conflict that nobody needed to have.
- **A revoked grant** keeps serving a client that has not noticed yet, for up to a pass.
Shortening the interval is the obvious answer and the wrong one: it costs a request per client per
interval whether or not anything happened, and it does not converge on *immediate* — it converges
on a busier server that is still late.
There is also a second thing coming that this decision has to not preclude. The intended feature is
a **shared terminal session** — one person's shell, watched or driven by another, TeamViewer-shaped.
That is bidirectional, continuous, and latency-sensitive in a way a keychain notice is not.
## Decision
### One WebSocket per signed-in client, at `GET /api/v1/events`
Subprotocol `dodossh.events.v1`. The client opens it after unlock and keeps it open; the server
sends a notice whenever something the client can read has changed.
**Not SSE.** Server-sent events would carry today's notices perfectly well and would be less code.
It is one-directional, so the shared-session feature would need a second mechanism next to it, and
then two transports would need reconnection, authorization and lifetime rules that agree. The cost
of a WebSocket over SSE is small; the cost of two transports is not.
**Not SignalR.** It brings hub protocol negotiation, its own serialisation and transport fallbacks,
none of which are wanted here: `DodoSSH.Contracts` and its source-generated serialiser are "the
actual contract between the two sides", and a second wire format alongside it is exactly the silent
drift `Setup/Json.cs` records having already cost this project once.
### The notice carries no ciphertext
A `vault.changed` frame is `{ kind, vaultId, sequence }` and nothing else. The client's answer to it
is the pull it would have done on the timer anyway.
This is the load-bearing decision, and it is worth being explicit about why the tempting alternative
is refused. Pushing the changed items themselves would save a round trip and would fork the code
path that applies a change into two — one that arrives by pull and one that arrives by socket — with
the cursor, the merge and the tombstone rules duplicated across both. ADR 0003 put every mutation
through one write path for exactly that reason; this keeps every *read* on one path for the same
one. The socket decides *when* to sync. It never decides *what* a vault contains.
It also means a dropped notice is harmless, which is what lets everything below be simple.
### Polling stays, and is the fallback rather than a legacy path
The one-minute pass is unchanged. The socket makes it *early*; it does not make it *necessary*. A
client on a network that eats WebSockets, an older client, a server that has the feature off, a
notice dropped under backpressure, a second API replica that did not see the write — every one of
those degrades to what the product does today, which is correct and up to a minute late.
Nothing may be reachable only by socket. That is a rule about future features, not an observation
about this one.
### Authorization: the bearer token on the upgrade, not a ticket
[ADR 0004](0004-relay-authorization.md) gives the relay a two-step ticket so its WebSocket carries
no API authority. This one goes the other way and takes the ordinary bearer JWT on the upgrade
request, which is an ordinary authenticated HTTP request. The difference is not inconsistency:
- The relay's socket is a **byte pipe to a third party**, and its whole authorization decision —
which host, which IPs, which port — is made *before* the socket opens and never revisited. It is
also the extraction seam for a standalone relay process that must not hold ACL code.
- This socket is a **view of the caller's own vault list**, and it has to keep answering "what may
this account read" for as long as it is open. It needs the full ACL context, in-process, for the
life of the connection. A ticket would carry that context in a token instead, and it would be
wrong the moment the account's access changed.
A long-lived connection authorised by a short-lived token is the problem this creates, and it is met
head-on rather than ignored:
1. **The socket does not outlive the token.** The `exp` claim is read at accept, and the connection
is closed with `4401` when it passes. The client reconnects with a fresh token; that is a
sub-second gap in a channel whose failure mode is already "poll instead".
2. **The vault set is re-resolved periodically** (`Events:AccessRefreshInterval`, default five
minutes) as well as on the changes that are known to affect it. A withdrawn grant therefore stops
producing notices within that window at the latest, and immediately in the ordinary case.
Both are bounds on **metadata** — the fact that a vault changed and roughly when — because that is
all a notice contains. Nobody's ciphertext is behind this socket, and a client that stayed subscribed
one interval too long could still not read a byte of it: reading requires a vault key grant, which
this server has never held.
### The frames
Text frames, JSON, `DodoSshJsonContext`. Server to client:
| kind | meaning |
| --- | --- |
| `hello` | accepted; carries the heartbeat interval and the vault count subscribed |
| `vault.changed` | `vaultId` moved to `sequence`; pull it |
| `vaults.changed` | the set of vaults this account can reach is different; re-read it |
| `ping` | heartbeat; the client answers `pong` |
Client to server: `ping`, answered with `pong`. Nothing else — subscription is decided by the server
from the caller's access, not asked for by the client, because a client that could ask to subscribe
to a vault id is a client that can probe for vault ids.
`kind` is a **string**, not an enum, and that is deliberate. `UseStringEnumConverter` throws on a
value it does not know, so a newer server sending a kind an older client has never heard of would
not add an unknown frame — it would break that client's socket entirely. A string is ignored
instead, which is what makes the table above extensible. `ProblemCodes` is the same shape for the
same reason.
### Where the shared session will attach
The socket is the seam, and one thing about it is chosen now so that it need not be renegotiated
later: **session data will be binary frames on this same connection, not JSON on the table above.**
Terminal output base64'd into a JSON envelope would cost a third of the bandwidth for nothing, on
the one payload here that is continuous rather than occasional. Control — offer, accept, resize,
end — is JSON like everything else.
That is as far as this ADR goes. Two questions are open and are not being answered by implication:
whether a shared session's bytes go through the API at all or peer-to-peer past it, and what
end-to-end encryption means when the second party is watching a stream rather than holding a key.
Both are ADR 0001 questions and deserve their own decision. What this one buys is that they will not
also be transport questions.
## Consequences
- **Fan-out is in-process, and the deployment is therefore single-node for this feature.** Every
connection is held by the node that accepted it; a write handled by another node produces no
notice on this one. `IVaultEventPublisher` is the seam a backplane implements — PostgreSQL
`LISTEN`/`NOTIFY` needs no infrastructure this stack does not already run — and it is deliberately
**not implemented**, because an untested backplane is worse than a documented gap. Multiple API
replicas do not break: they degrade to polling, which is the state before this ADR. `/api/v1/meta`
advertises `events` so a client knows which it is getting.
- **Per-connection queues are bounded and drop the oldest.** A notice is "pull vault X, which is at
least at sequence N", so the newest is strictly more useful than the one it displaces and the
client's answer is identical either way. A slow reader costs itself latency, never the publisher's
progress — the publish path never blocks and never awaits a socket.
- **The publish happens after the transaction commits**, outside the advisory lock ADR 0003 takes.
A notice sent from inside it would name a sequence a reader cannot yet see, and would hold the
per-vault write lock across a socket write.
- **A client is notified of its own writes.** It pushed, so it already pulled; the extra pass finds
nothing. The client coalesces notices over a short window rather than the server suppressing an
echo, because suppressing it correctly needs a per-*device* identity on the socket and the same
user's other machines must still be told.
- **Connections are capped** per user and per node (`Events:MaxConnectionsPerUser`,
`Events:MaxConnectionsTotal`). A socket is cheap but not free, and an unbounded count of them is a
denial of service that authenticates first.
- The feature can be turned off entirely (`Events:Enabled`). A deployment behind a proxy that will
not upgrade should say so rather than have every client discover it by failing.
### Rejected
- **Shorter polling.** Cheaper to build, converges on a busier server that is still late.
- **Long polling.** No new transport and genuinely immediate, but it holds a request thread and a
connection per client for the same money as a WebSocket while offering none of the bidirectionality
the shared session needs.
- **Pushing the changed items down the socket.** Saves a round trip; forks the apply path in two. See
above.
- **Client-chosen subscriptions.** A `subscribe(vaultId)` frame is an existence oracle for vault ids,
which is the disclosure `SyncPullEndpoint` answers 404 rather than 403 to avoid.
+80
View File
@@ -1500,3 +1500,83 @@ delivery survives the failure because it is held against the transfer rather tha
**Failure means:** a retry that succeeds but leaves the destination empty is the delivery having been
dropped on the failure. An error saying the staged file is missing is the copy having been deleted at the
stop, which is what `QueueDeliveredDownload` documents it does not do.
## Phase 15 — Changes that arrive without a timer
The socket is covered by tests on both sides: the endpoint suite opens a real one against a real
`TestServer` and proves a push produces a notice, that another account's push does not, and that a frame
carries no ciphertext; the shell suite proves a notice wakes the synchronisation loop long before the
minute. What none of that can reach is **the network in between**, and that is where this feature is most
likely to fail: a reverse proxy that will not upgrade, one that drops an idle socket without telling either
end, a corporate middlebox, a phone moving between Wi-Fi and mobile data. Every one of those looks the same
from inside a test host, which has no proxy and no radio.
The pass condition throughout is *two* things, and the second matters as much as the first: it arrives
quickly, **and** it still arrives when the socket is gone. A build where the timer had stopped working would
pass every "it was fast" check here and fail nobody until somebody's proxy changed.
### 15.1 A colleague's edit appears while you are looking at it
Two accounts sharing a vault, both unlocked, both on the Hosts screen. On the first machine, rename a host
in the shared vault and save.
**Pass:** the second machine's list shows the new name within a second or two, with nothing pressed and no
screen flicker — the row updates, the selection does not move, and the status line is not repainted with a
sync report.
**Failure means:** nothing within a minute, then the new name, is the socket not being established at all —
that is the timer doing its job, which is the correct fallback and not the feature. Check `/api/v1/meta`
lists `events`, then whether the proxy in front of the API forwards `Upgrade` and `Connection`. A list that
never updates at all is a synchronisation failure and has nothing to do with this phase.
### 15.2 A vault shared with you turns up as it is shared
The second account signed in and unlocked, sitting on the VAULTS screen. From the first, add them to a team
and press SHARE KEY.
**Pass:** the vault appears in their list within a second or two of the key being wrapped, and reads as
waiting for a key until the share, then as readable.
**Failure means:** the vault appearing only on the minute is the `vaults.changed` notice not being published
or not being followed. Both the membership add and the grant publish one; if the membership arrives promptly
and the key does not, the grant path is the one to look at.
### 15.3 It still works with the socket taken away
On the second machine, block the WebSocket — the simplest way is a proxy rule rejecting the upgrade, or
setting `Events:Enabled` to `false` on the server and restarting it.
**Pass:** everything above still happens, within the minute rather than within seconds. Nothing on the
screen says anything is wrong, because nothing is: no error, no OFFLINE badge, no repeated status message.
The Sync button still works and still reports.
**Failure means:** an error message, a titlebar claiming to be offline, or a status line that repaints with
a socket failure is the client treating an absent push channel as a fault. It is not one — the timer is the
guarantee and the socket is the optimisation, and a user with a strict proxy must never be told their
keychain is broken.
### 15.4 A laptop that slept comes back on its own
With the second machine idle and connected, close the lid for a few minutes — or disable Wi-Fi for two
minutes and re-enable it. Then make a change on the first machine.
**Pass:** the change arrives quickly again, without the vault having been locked or the application
restarted. The reconnection is invisible.
**Failure means:** changes that arrive only on the timer from then on are the stream having given up after
its socket died — the reconnection loop is what should make that impossible, and a client that reconnects
once and not twice is the specific defect its tests exist to catch. Changes that never arrive again, timer
included, are a different and worse bug in the synchronisation loop rather than in the socket.
### 15.5 An expiring token does not end the push
This one needs a short access-token lifetime in the identity provider — the dev realm's Keycloak client can
be set to a couple of minutes. Leave a machine unlocked and idle for longer than that, then make a change
elsewhere.
**Pass:** the change still arrives quickly. The socket is closed by the server at the token's expiry and the
client reconnects with a fresh one, which should be invisible.
**Failure means:** notices stopping at roughly the token's lifetime is the reconnection not asking for a new
token — it would be dialling with the spent one and being closed again immediately. A burst of reconnection
attempts in the server log is the same defect seen from the other end.