Public Access
Merge branch 'claude/vault-realtime-push-d64c61'
This commit is contained in:
@@ -58,7 +58,11 @@ the Npgsql connection string** — the default; do not enable multiplexing.
|
||||
- One place enforces revision, change-log and ACL invariants. That halves both the endpoint
|
||||
count and the authorization surface, which is the main reason for the single write path.
|
||||
- Delta pull makes frequent polling cheap, so multi-device feels live; push notification over
|
||||
SSE or the existing WebSocket can layer on with polling as the fallback.
|
||||
SSE or the existing WebSocket can layer on with polling as the fallback. **That has since been
|
||||
built — see [ADR 0012](0012-realtime-push.md)** — and nothing in this ADR changed to accommodate
|
||||
it. The socket carries a notice naming a vault and a sequence, whose answer is the delta pull
|
||||
above, so there is still exactly one path that applies a change; and polling is still what
|
||||
guarantees a pass rather than a legacy route kept for old clients.
|
||||
- Conflict resolution is entirely client-side. The client retains a `BaseCiphertext` common
|
||||
ancestor and performs a field-level three-way merge for structured items, or creates a
|
||||
visible conflicted copy for opaque ones. **It must never silently drop a key or a host.**
|
||||
|
||||
@@ -0,0 +1,167 @@
|
||||
# ADR 0012 — A WebSocket that carries notices, not data
|
||||
|
||||
- Status: accepted
|
||||
- Date: 2026-08-04
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0003](0003-sync-protocol.md) built a delta pull that is cheap enough to run on a timer, and
|
||||
the client does: one pass a minute. That is the difference between a colleague's change appearing
|
||||
"soon" and appearing *now*, and it shows up in three places that are not equally forgivable.
|
||||
|
||||
- **A vault shared with you** arrives on the next pass. `AdmitNewVaultsAsync` says so in its own
|
||||
remarks — "the recipient is handed nothing — there is no push channel" — and the README repeats
|
||||
it. Sharing works and looks broken.
|
||||
- **Two people editing one keychain** see each other up to a minute late, which is long enough to
|
||||
make the same edit twice and produce a conflict that nobody needed to have.
|
||||
- **A revoked grant** keeps serving a client that has not noticed yet, for up to a pass.
|
||||
|
||||
Shortening the interval is the obvious answer and the wrong one: it costs a request per client per
|
||||
interval whether or not anything happened, and it does not converge on *immediate* — it converges
|
||||
on a busier server that is still late.
|
||||
|
||||
There is also a second thing coming that this decision has to not preclude. The intended feature is
|
||||
a **shared terminal session** — one person's shell, watched or driven by another, TeamViewer-shaped.
|
||||
That is bidirectional, continuous, and latency-sensitive in a way a keychain notice is not.
|
||||
|
||||
## Decision
|
||||
|
||||
### One WebSocket per signed-in client, at `GET /api/v1/events`
|
||||
|
||||
Subprotocol `dodossh.events.v1`. The client opens it after unlock and keeps it open; the server
|
||||
sends a notice whenever something the client can read has changed.
|
||||
|
||||
**Not SSE.** Server-sent events would carry today's notices perfectly well and would be less code.
|
||||
It is one-directional, so the shared-session feature would need a second mechanism next to it, and
|
||||
then two transports would need reconnection, authorization and lifetime rules that agree. The cost
|
||||
of a WebSocket over SSE is small; the cost of two transports is not.
|
||||
|
||||
**Not SignalR.** It brings hub protocol negotiation, its own serialisation and transport fallbacks,
|
||||
none of which are wanted here: `DodoSSH.Contracts` and its source-generated serialiser are "the
|
||||
actual contract between the two sides", and a second wire format alongside it is exactly the silent
|
||||
drift `Setup/Json.cs` records having already cost this project once.
|
||||
|
||||
### The notice carries no ciphertext
|
||||
|
||||
A `vault.changed` frame is `{ kind, vaultId, sequence }` and nothing else. The client's answer to it
|
||||
is the pull it would have done on the timer anyway.
|
||||
|
||||
This is the load-bearing decision, and it is worth being explicit about why the tempting alternative
|
||||
is refused. Pushing the changed items themselves would save a round trip and would fork the code
|
||||
path that applies a change into two — one that arrives by pull and one that arrives by socket — with
|
||||
the cursor, the merge and the tombstone rules duplicated across both. ADR 0003 put every mutation
|
||||
through one write path for exactly that reason; this keeps every *read* on one path for the same
|
||||
one. The socket decides *when* to sync. It never decides *what* a vault contains.
|
||||
|
||||
It also means a dropped notice is harmless, which is what lets everything below be simple.
|
||||
|
||||
### Polling stays, and is the fallback rather than a legacy path
|
||||
|
||||
The one-minute pass is unchanged. The socket makes it *early*; it does not make it *necessary*. A
|
||||
client on a network that eats WebSockets, an older client, a server that has the feature off, a
|
||||
notice dropped under backpressure, a second API replica that did not see the write — every one of
|
||||
those degrades to what the product does today, which is correct and up to a minute late.
|
||||
|
||||
Nothing may be reachable only by socket. That is a rule about future features, not an observation
|
||||
about this one.
|
||||
|
||||
### Authorization: the bearer token on the upgrade, not a ticket
|
||||
|
||||
[ADR 0004](0004-relay-authorization.md) gives the relay a two-step ticket so its WebSocket carries
|
||||
no API authority. This one goes the other way and takes the ordinary bearer JWT on the upgrade
|
||||
request, which is an ordinary authenticated HTTP request. The difference is not inconsistency:
|
||||
|
||||
- The relay's socket is a **byte pipe to a third party**, and its whole authorization decision —
|
||||
which host, which IPs, which port — is made *before* the socket opens and never revisited. It is
|
||||
also the extraction seam for a standalone relay process that must not hold ACL code.
|
||||
- This socket is a **view of the caller's own vault list**, and it has to keep answering "what may
|
||||
this account read" for as long as it is open. It needs the full ACL context, in-process, for the
|
||||
life of the connection. A ticket would carry that context in a token instead, and it would be
|
||||
wrong the moment the account's access changed.
|
||||
|
||||
A long-lived connection authorised by a short-lived token is the problem this creates, and it is met
|
||||
head-on rather than ignored:
|
||||
|
||||
1. **The socket does not outlive the token.** The `exp` claim is read at accept, and the connection
|
||||
is closed with `4401` when it passes. The client reconnects with a fresh token; that is a
|
||||
sub-second gap in a channel whose failure mode is already "poll instead".
|
||||
2. **The vault set is re-resolved periodically** (`Events:AccessRefreshInterval`, default five
|
||||
minutes) as well as on the changes that are known to affect it. A withdrawn grant therefore stops
|
||||
producing notices within that window at the latest, and immediately in the ordinary case.
|
||||
|
||||
Both are bounds on **metadata** — the fact that a vault changed and roughly when — because that is
|
||||
all a notice contains. Nobody's ciphertext is behind this socket, and a client that stayed subscribed
|
||||
one interval too long could still not read a byte of it: reading requires a vault key grant, which
|
||||
this server has never held.
|
||||
|
||||
### The frames
|
||||
|
||||
Text frames, JSON, `DodoSshJsonContext`. Server to client:
|
||||
|
||||
| kind | meaning |
|
||||
| --- | --- |
|
||||
| `hello` | accepted; carries the heartbeat interval and the vault count subscribed |
|
||||
| `vault.changed` | `vaultId` moved to `sequence`; pull it |
|
||||
| `vaults.changed` | the set of vaults this account can reach is different; re-read it |
|
||||
| `ping` | heartbeat; the client answers `pong` |
|
||||
|
||||
Client to server: `ping`, answered with `pong`. Nothing else — subscription is decided by the server
|
||||
from the caller's access, not asked for by the client, because a client that could ask to subscribe
|
||||
to a vault id is a client that can probe for vault ids.
|
||||
|
||||
`kind` is a **string**, not an enum, and that is deliberate. `UseStringEnumConverter` throws on a
|
||||
value it does not know, so a newer server sending a kind an older client has never heard of would
|
||||
not add an unknown frame — it would break that client's socket entirely. A string is ignored
|
||||
instead, which is what makes the table above extensible. `ProblemCodes` is the same shape for the
|
||||
same reason.
|
||||
|
||||
### Where the shared session will attach
|
||||
|
||||
The socket is the seam, and one thing about it is chosen now so that it need not be renegotiated
|
||||
later: **session data will be binary frames on this same connection, not JSON on the table above.**
|
||||
Terminal output base64'd into a JSON envelope would cost a third of the bandwidth for nothing, on
|
||||
the one payload here that is continuous rather than occasional. Control — offer, accept, resize,
|
||||
end — is JSON like everything else.
|
||||
|
||||
That is as far as this ADR goes. Two questions are open and are not being answered by implication:
|
||||
whether a shared session's bytes go through the API at all or peer-to-peer past it, and what
|
||||
end-to-end encryption means when the second party is watching a stream rather than holding a key.
|
||||
Both are ADR 0001 questions and deserve their own decision. What this one buys is that they will not
|
||||
also be transport questions.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Fan-out is in-process, and the deployment is therefore single-node for this feature.** Every
|
||||
connection is held by the node that accepted it; a write handled by another node produces no
|
||||
notice on this one. `IVaultEventPublisher` is the seam a backplane implements — PostgreSQL
|
||||
`LISTEN`/`NOTIFY` needs no infrastructure this stack does not already run — and it is deliberately
|
||||
**not implemented**, because an untested backplane is worse than a documented gap. Multiple API
|
||||
replicas do not break: they degrade to polling, which is the state before this ADR. `/api/v1/meta`
|
||||
advertises `events` so a client knows which it is getting.
|
||||
- **Per-connection queues are bounded and drop the oldest.** A notice is "pull vault X, which is at
|
||||
least at sequence N", so the newest is strictly more useful than the one it displaces and the
|
||||
client's answer is identical either way. A slow reader costs itself latency, never the publisher's
|
||||
progress — the publish path never blocks and never awaits a socket.
|
||||
- **The publish happens after the transaction commits**, outside the advisory lock ADR 0003 takes.
|
||||
A notice sent from inside it would name a sequence a reader cannot yet see, and would hold the
|
||||
per-vault write lock across a socket write.
|
||||
- **A client is notified of its own writes.** It pushed, so it already pulled; the extra pass finds
|
||||
nothing. The client coalesces notices over a short window rather than the server suppressing an
|
||||
echo, because suppressing it correctly needs a per-*device* identity on the socket and the same
|
||||
user's other machines must still be told.
|
||||
- **Connections are capped** per user and per node (`Events:MaxConnectionsPerUser`,
|
||||
`Events:MaxConnectionsTotal`). A socket is cheap but not free, and an unbounded count of them is a
|
||||
denial of service that authenticates first.
|
||||
- The feature can be turned off entirely (`Events:Enabled`). A deployment behind a proxy that will
|
||||
not upgrade should say so rather than have every client discover it by failing.
|
||||
|
||||
### Rejected
|
||||
|
||||
- **Shorter polling.** Cheaper to build, converges on a busier server that is still late.
|
||||
- **Long polling.** No new transport and genuinely immediate, but it holds a request thread and a
|
||||
connection per client for the same money as a WebSocket while offering none of the bidirectionality
|
||||
the shared session needs.
|
||||
- **Pushing the changed items down the socket.** Saves a round trip; forks the apply path in two. See
|
||||
above.
|
||||
- **Client-chosen subscriptions.** A `subscribe(vaultId)` frame is an existence oracle for vault ids,
|
||||
which is the disclosure `SyncPullEndpoint` answers 404 rather than 403 to avoid.
|
||||
@@ -1500,3 +1500,83 @@ delivery survives the failure because it is held against the transfer rather tha
|
||||
**Failure means:** a retry that succeeds but leaves the destination empty is the delivery having been
|
||||
dropped on the failure. An error saying the staged file is missing is the copy having been deleted at the
|
||||
stop, which is what `QueueDeliveredDownload` documents it does not do.
|
||||
|
||||
## Phase 15 — Changes that arrive without a timer
|
||||
|
||||
The socket is covered by tests on both sides: the endpoint suite opens a real one against a real
|
||||
`TestServer` and proves a push produces a notice, that another account's push does not, and that a frame
|
||||
carries no ciphertext; the shell suite proves a notice wakes the synchronisation loop long before the
|
||||
minute. What none of that can reach is **the network in between**, and that is where this feature is most
|
||||
likely to fail: a reverse proxy that will not upgrade, one that drops an idle socket without telling either
|
||||
end, a corporate middlebox, a phone moving between Wi-Fi and mobile data. Every one of those looks the same
|
||||
from inside a test host, which has no proxy and no radio.
|
||||
|
||||
The pass condition throughout is *two* things, and the second matters as much as the first: it arrives
|
||||
quickly, **and** it still arrives when the socket is gone. A build where the timer had stopped working would
|
||||
pass every "it was fast" check here and fail nobody until somebody's proxy changed.
|
||||
|
||||
### 15.1 A colleague's edit appears while you are looking at it
|
||||
|
||||
Two accounts sharing a vault, both unlocked, both on the Hosts screen. On the first machine, rename a host
|
||||
in the shared vault and save.
|
||||
|
||||
**Pass:** the second machine's list shows the new name within a second or two, with nothing pressed and no
|
||||
screen flicker — the row updates, the selection does not move, and the status line is not repainted with a
|
||||
sync report.
|
||||
|
||||
**Failure means:** nothing within a minute, then the new name, is the socket not being established at all —
|
||||
that is the timer doing its job, which is the correct fallback and not the feature. Check `/api/v1/meta`
|
||||
lists `events`, then whether the proxy in front of the API forwards `Upgrade` and `Connection`. A list that
|
||||
never updates at all is a synchronisation failure and has nothing to do with this phase.
|
||||
|
||||
### 15.2 A vault shared with you turns up as it is shared
|
||||
|
||||
The second account signed in and unlocked, sitting on the VAULTS screen. From the first, add them to a team
|
||||
and press SHARE KEY.
|
||||
|
||||
**Pass:** the vault appears in their list within a second or two of the key being wrapped, and reads as
|
||||
waiting for a key until the share, then as readable.
|
||||
|
||||
**Failure means:** the vault appearing only on the minute is the `vaults.changed` notice not being published
|
||||
or not being followed. Both the membership add and the grant publish one; if the membership arrives promptly
|
||||
and the key does not, the grant path is the one to look at.
|
||||
|
||||
### 15.3 It still works with the socket taken away
|
||||
|
||||
On the second machine, block the WebSocket — the simplest way is a proxy rule rejecting the upgrade, or
|
||||
setting `Events:Enabled` to `false` on the server and restarting it.
|
||||
|
||||
**Pass:** everything above still happens, within the minute rather than within seconds. Nothing on the
|
||||
screen says anything is wrong, because nothing is: no error, no OFFLINE badge, no repeated status message.
|
||||
The Sync button still works and still reports.
|
||||
|
||||
**Failure means:** an error message, a titlebar claiming to be offline, or a status line that repaints with
|
||||
a socket failure is the client treating an absent push channel as a fault. It is not one — the timer is the
|
||||
guarantee and the socket is the optimisation, and a user with a strict proxy must never be told their
|
||||
keychain is broken.
|
||||
|
||||
### 15.4 A laptop that slept comes back on its own
|
||||
|
||||
With the second machine idle and connected, close the lid for a few minutes — or disable Wi-Fi for two
|
||||
minutes and re-enable it. Then make a change on the first machine.
|
||||
|
||||
**Pass:** the change arrives quickly again, without the vault having been locked or the application
|
||||
restarted. The reconnection is invisible.
|
||||
|
||||
**Failure means:** changes that arrive only on the timer from then on are the stream having given up after
|
||||
its socket died — the reconnection loop is what should make that impossible, and a client that reconnects
|
||||
once and not twice is the specific defect its tests exist to catch. Changes that never arrive again, timer
|
||||
included, are a different and worse bug in the synchronisation loop rather than in the socket.
|
||||
|
||||
### 15.5 An expiring token does not end the push
|
||||
|
||||
This one needs a short access-token lifetime in the identity provider — the dev realm's Keycloak client can
|
||||
be set to a couple of minutes. Leave a machine unlocked and idle for longer than that, then make a change
|
||||
elsewhere.
|
||||
|
||||
**Pass:** the change still arrives quickly. The socket is closed by the server at the token's expiry and the
|
||||
client reconnects with a fresh one, which should be invisible.
|
||||
|
||||
**Failure means:** notices stopping at roughly the token's lifetime is the reconnection not asking for a new
|
||||
token — it would be dialling with the spent one and being closed again immediately. A burst of reconnection
|
||||
attempts in the server log is the same defect seen from the other end.
|
||||
|
||||
Reference in New Issue
Block a user