Public Access
Say when a vault has moved, so nobody waits out the minute
The delta pull was cheap enough to run on a timer and the client did, once a minute. That is fine for a machine and wrong for two people: an edit a colleague makes is up to a minute stale, which is long enough for both of them to make it and produce a conflict neither needed to have. Shortening the interval is the obvious answer and the wrong one — it costs a request per client per interval whether or not anything happened, and it converges on a busier server that is still late. So the server now says so. A client holds a WebSocket open at GET /api/v1/events, subprotocol dodossh.events.v1, and gets a line down it when something it can read has changed. ADR 0012 has the reasoning; three parts of it are worth repeating here, because they are what everything else rests on. **What crosses the socket is a notice, never data.** A frame names a vault and how far its change log has got. No item, no ciphertext, not even which item it was. The client's answer is the delta pull it would have run anyway, so there is still exactly one code path that applies a change to a keychain, and it is not this one. Pushing the items themselves would save a round trip and fork that path in two, with the cursor, the merge and the tombstone rules duplicated across both — ADR 0003 put every mutation through one write path for that reason, and this keeps every read on one for the same one. It also makes a dropped notice harmless, which is what lets the fan-out below be as simple as it is. **Polling stays, and is what guarantees a pass.** The minute timer is unchanged. A network that eats WebSockets, a server with Events:Enabled off, an older server, a proxy that will not upgrade, a notice dropped under backpressure — every one of those leaves a client behaving exactly as it did before this commit. Nothing is reachable only over the socket and nothing is meant to become so; VaultViewModel's AutoSyncInterval remark now says that where somebody changing it will read it. **The bearer token authorises the upgrade, unlike the relay's ticket.** Not an inconsistency with ADR 0004: the relay's socket is a byte pipe whose whole authorization decision — which host, which IPs, which port — is made before it opens and never revisited, and it is the extraction seam for a process that must hold no ACL code. This one is a view of the caller's own vault list and has to keep answering "what may this account read" for as long as it is held. A ticket would carry that answer in a token and be wrong the moment the account's access changed. The two bounds that arrangement needs are met rather than waved at: the socket is closed at the token's exp with close code 4401 and the client comes straight back with a fresh one, and the vault set is re-resolved every few minutes as well as on the changes known to affect it. Both bound *metadata*, because a notice contains nothing else and reading a vault still needs a key this server has never held. **The fan-out.** VaultEventHub is a singleton holding the sockets this node accepted; publishing walks them and asks each whether it cares, rather than keeping a vault-to-subscriber index that every re-subscription would have to move entries between under a lock publishing also takes. At a few hundred sockets per node and an event rate bounded by how often people edit keychains, the walk is not measurable and its races are obvious. Per-connection queues are bounded and drop the *oldest*: a notice means "pull vault X, which is at least at sequence N", so the newest subsumes what it displaces and the client's answer is identical either way — which is what lets the publish path be void, never block, and never fail. Announced from the endpoint rather than from SyncService, and that placement is the point: by then the push has committed and released the per-vault advisory lock. From inside it would name a sequence no reader can see yet and would hold the lock that serialises writers across a socket write. Only the highest *applied* sequence, so a batch of pure conflicts announces nothing, and a duplicate — already announced when it first landed — announces nothing either. Grants and membership publish too, and those take the *recipient* rather than the actor. This is what AdmitNewVaultsAsync has been apologising for since sharing shipped — "the recipient is handed nothing, there is no push channel" — and the README with it. A vault shared with somebody now turns up as it is shared. The comment and the README paragraph both say what is true now, and both keep saying that the pass is what *discovers* the vault, because a client with no socket has to arrive at the same place. **On the client**, VaultEventStream is really a reconnection policy wrapped round a ClientWebSocket: a dropped socket is the ordinary case here — laptops sleep, proxies time out, tokens expire, servers are redeployed — so nothing in it treats a failure as exceptional, and every path ends in "wait, then dial again". A connection that lived long enough to say hello resets the backoff, so a laptop that woke, worked, and lost its network an hour later does not inherit a minute-long wait it has already proved it need not take. A 4401 close skips the backoff entirely and asks the token provider again, which is the whole reason that close code is distinct. A server that does not advertise the events feature gets IdleVaultEventStream, which never delivers — so IVaultServer.Events is never null and every caller stays on one shape, because the correct behaviour without a socket is the behaviour with a silent one. The shell's background loop now selects between the timer and a notice, and both waits are held across iterations. That is load-bearing rather than tidy: PeriodicTimer permits one outstanding WaitForNextTickAsync and throws on a second, and an abandoned channel read stays registered and consumes the next notice written. Either defect leaves the first notice working and every one after it silently lost, which is why NoticesKeepWakingTheLoop_NotJustTheFirst pushes three and not one. Notices are coalesced over a quarter of a second, so one person's save — a host and its log entry are two items — and a colleague clearing a folder each cost one pass rather than a dozen. **The kind is a string, not an enum**, and that is a compatibility decision. UseStringEnumConverter throws on a value it does not know, so a newer server sending a kind an older client had never heard of would not add an unreadable frame — it would break that client's socket outright. A string is ignored instead. ProblemCodes is the same shape for the same reason. **Tested on both sides, through the real pipeline.** The endpoint suite opens a genuine socket against TestServer and proves a push produces a notice, that another account's push does not reach it, that a ping is answered, and that a frame this server cannot parse does not end the connection. Two of those assert on *ordering* rather than on absence within a timeout — the stranger's write goes first, so a socket that leaked would have announced it before the one the test waits for — because "nothing arrived in two seconds" is a test that passes on a slow machine for the wrong reason. And ANoticeCarriesNoCiphertext asserts on the bytes that crossed the wire rather than on the record's fields, since the latter would only prove that this type has no payload member, which is a tautology; the former is what catches a field added later without anybody thinking about disclosure. The client suite drives VaultEventStream through an injected connector, because the one thing a test cannot do to a real network is make it fail on cue — and failure is the entire subject. The shell suite proves a notice produces a pull inside ten seconds against a sixty-second timer, so the timer cannot be what caused it. **Two limits, stated rather than left to be discovered.** Fan-out is in-process, so a deployment running more than one API replica only pushes for writes its own replica handled and the rest arrive on the timer. IVaultEventPublisher is the seam a PostgreSQL LISTEN/NOTIFY backplane implements and it is deliberately not implemented: an untested backplane is worse than a documented gap, and multiple replicas degrade to the behaviour before this commit rather than breaking. And a client is notified of its own writes; it pushed, so it already pulled, and the extra pass finds nothing. Suppressing that echo correctly needs a per-device identity on the socket, and the same user's other machines must still be told. Manual checks phase 15 covers what no test here can reach, which is the network in between: a proxy that will not upgrade, one that drops an idle socket without telling either end, a laptop lid, a token expiring. Every one of those is invisible inside a test host, and every check there passes only if the change arrives quickly *and* still arrives with the socket taken away. ADR 0012 also fixes one thing about the shared terminal session this is the transport for, so it need not be renegotiated later: session data will be binary frames on this same socket, because base64 in a JSON envelope is the wrong shape for the one payload here that is continuous rather than occasional. Two questions it explicitly does not answer by implication — whether those bytes go through the API at all, and what end-to-end encryption means when the second party watches a stream rather than holding a key — are ADR 0001 questions and get their own decision. 1512 tests pass. DodoSSH.SystemTests was not run — it needs the whole compose stack — so the end-to-end path is unverified for this change beyond what the manual checks describe.
This commit is contained in:
@@ -10,6 +10,7 @@ using DodoSSH.Client.Session;
|
||||
using DodoSSH.Client.Ssh;
|
||||
using DodoSSH.Client.Sync;
|
||||
using DodoSSH.Client.Terminal;
|
||||
using DodoSSH.Contracts;
|
||||
|
||||
namespace DodoSSH.Client.Shell.ViewModels;
|
||||
|
||||
@@ -1025,13 +1026,29 @@ internal sealed partial class VaultViewModel(
|
||||
VaultVisibility? visibility = null) : ObservableObject, IAsyncDisposable
|
||||
{
|
||||
/// <remarks>
|
||||
/// <para>
|
||||
/// A minute. The pull is a delta keyed on a cursor, so an idle pass is one small request and costs the
|
||||
/// server almost nothing; the number that matters is how stale a teammate's change may look, and a
|
||||
/// minute is short enough not to be noticed. Anything much shorter would be polling for its own sake,
|
||||
/// and a change made on this machine does not wait for the timer anyway — saving pushes immediately.
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// Unchanged by the push channel, and deliberately so. The socket makes a pass <em>early</em>; this is
|
||||
/// what makes one happen at all, for a client whose network eats WebSockets, whose server has the
|
||||
/// feature off, or whose notice was dropped. See <see cref="WaitForWorkAsync"/> and ADR 0012.
|
||||
/// </para>
|
||||
/// </remarks>
|
||||
private static readonly TimeSpan AutoSyncInterval = TimeSpan.FromMinutes(1);
|
||||
|
||||
/// <summary>How long a pushed notice waits, in case more are on their way.</summary>
|
||||
/// <remarks>
|
||||
/// A quarter of a second, which is below what anybody perceives and above the gap between the
|
||||
/// notices one person's save produces — a host and its activity log entry are two items in one
|
||||
/// push, and a colleague clearing a folder is a burst. Without it each notice would run its own
|
||||
/// full pass, and the pass a burst deserves is one.
|
||||
/// </remarks>
|
||||
private static readonly TimeSpan NoticeDebounce = TimeSpan.FromMilliseconds(250);
|
||||
|
||||
/// <summary>How often the logs are pruned, at most.</summary>
|
||||
/// <remarks>
|
||||
/// Hours rather than minutes, because pruning writes tombstones that sync. Retention is measured in days
|
||||
@@ -4030,10 +4047,16 @@ internal sealed partial class VaultViewModel(
|
||||
/// <para>
|
||||
/// <b>This is the whole of how a shared vault arrives.</b> Sharing is two acts on two machines: the
|
||||
/// person sharing wraps the vault key to the recipient, and the recipient's own client has to notice.
|
||||
/// The recipient is handed nothing — there is no push channel — so without this the vault list stayed
|
||||
/// exactly as it was cached at sign-in, and a vault shared with somebody appeared on their machine only
|
||||
/// if they happened to sign in through the browser again. Everything else was already right, which is
|
||||
/// why it looked like sharing was broken rather than like a list that was never re-read.
|
||||
/// Without this the vault list stayed exactly as it was cached at sign-in, and a vault shared with
|
||||
/// somebody appeared on their machine only if they happened to sign in through the browser again.
|
||||
/// Everything else was already right, which is why it looked like sharing was broken rather than like a
|
||||
/// list that was never re-read.
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// The server now says when this is worth doing — a <c>vaults.changed</c> notice wakes the pass, so the
|
||||
/// vault turns up as it is shared rather than within the minute — but that only decides <em>when</em>.
|
||||
/// This call is still what discovers the vault, on the notice and on every timed pass alike, because a
|
||||
/// client with no socket has to arrive at the same place. See ADR 0012.
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// A failure is left to the caller, which treats it as the pass failing: the call is to the same server
|
||||
@@ -4080,6 +4103,7 @@ internal sealed partial class VaultViewModel(
|
||||
private async Task RunAutoSyncLoopAsync(CancellationToken cancellationToken)
|
||||
{
|
||||
using var timer = new PeriodicTimer(AutoSyncInterval);
|
||||
var waits = new AutoSyncWaits();
|
||||
|
||||
try
|
||||
{
|
||||
@@ -4094,7 +4118,7 @@ internal sealed partial class VaultViewModel(
|
||||
// user is doing something.
|
||||
await SyncOnOpenAsync(cancellationToken).ConfigureAwait(true);
|
||||
|
||||
while (await timer.WaitForNextTickAsync(cancellationToken).ConfigureAwait(true))
|
||||
while (await WaitForWorkAsync(timer, waits, cancellationToken).ConfigureAwait(true))
|
||||
{
|
||||
await AutoSyncAsync(cancellationToken).ConfigureAwait(true);
|
||||
}
|
||||
@@ -4105,6 +4129,98 @@ internal sealed partial class VaultViewModel(
|
||||
}
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Waits for the timer to come round, or for the server to say there is something to fetch.
|
||||
/// </summary>
|
||||
/// <returns>Whether to run a pass. False means the loop is over.</returns>
|
||||
/// <remarks>
|
||||
/// <para>
|
||||
/// The timer is unchanged and is still what guarantees a pass. The socket only makes one
|
||||
/// <em>early</em>, which is why nothing here treats its absence as a problem: no connection, a
|
||||
/// server without the feature, a network that eats WebSockets, or a notice dropped under
|
||||
/// backpressure all leave a loop that behaves exactly as it did before this existed. See ADR 0012.
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// <b>Both waits are held across iterations, and that is load-bearing rather than an
|
||||
/// optimisation.</b> <see cref="PeriodicTimer"/> permits only one outstanding
|
||||
/// <c>WaitForNextTickAsync</c> and throws on a second, and an abandoned channel read stays
|
||||
/// registered and consumes the next notice written — which would silently lose exactly the wake-up
|
||||
/// this is for. Whichever wait did not win is kept and awaited again.
|
||||
/// </para>
|
||||
/// </remarks>
|
||||
private async Task<bool> WaitForWorkAsync(
|
||||
PeriodicTimer timer,
|
||||
AutoSyncWaits waits,
|
||||
CancellationToken cancellationToken)
|
||||
{
|
||||
// Re-read every time, because signing out and back in replaces the connection — and with it
|
||||
// the stream. A read still pending against the old one is left to be cancelled with it.
|
||||
var stream = connection()?.Events;
|
||||
|
||||
if (!ReferenceEquals(stream, waits.Watching))
|
||||
{
|
||||
waits.Watching = stream;
|
||||
waits.Notice = null;
|
||||
}
|
||||
|
||||
waits.Tick ??= timer.WaitForNextTickAsync(cancellationToken).AsTask();
|
||||
waits.Notice ??= stream?.ReadAsync(cancellationToken).AsTask();
|
||||
|
||||
if (waits.Notice is null)
|
||||
{
|
||||
var only = waits.Tick;
|
||||
waits.Tick = null;
|
||||
|
||||
return await only.ConfigureAwait(true);
|
||||
}
|
||||
|
||||
var first = await Task.WhenAny(waits.Tick, waits.Notice).ConfigureAwait(true);
|
||||
|
||||
if (ReferenceEquals(first, waits.Tick))
|
||||
{
|
||||
var ticked = waits.Tick;
|
||||
waits.Tick = null;
|
||||
|
||||
return await ticked.ConfigureAwait(true);
|
||||
}
|
||||
|
||||
// Observed so a faulted read does not go unhandled, and so a stream that has been disposed
|
||||
// ends this wait rather than being asked again.
|
||||
await waits.Notice.ConfigureAwait(true);
|
||||
waits.Notice = null;
|
||||
|
||||
// A burst — one person's save is two items, and a colleague tidying a folder is a dozen —
|
||||
// deserves one pass rather than one each.
|
||||
await Task.Delay(NoticeDebounce, cancellationToken).ConfigureAwait(true);
|
||||
|
||||
while (stream!.TryRead(out _))
|
||||
{
|
||||
// Swallowed on purpose. Every notice means the same thing, which is what the pass about to
|
||||
// run already does; what they say about *which* vault is not read, because a pass syncs
|
||||
// every vault this session can reach anyway.
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
/// <summary>The two waits the background loop keeps alive between passes.</summary>
|
||||
/// <remarks>
|
||||
/// A class rather than three locals because <see cref="WaitForWorkAsync"/> has to hand them back
|
||||
/// changed, and a method that took three <c>ref</c> parameters could not be <c>async</c>. See that
|
||||
/// method for why abandoning either of them is a defect rather than a tidiness question.
|
||||
/// </remarks>
|
||||
private sealed class AutoSyncWaits
|
||||
{
|
||||
/// <summary>The pending timer tick, or null when the last one has been consumed.</summary>
|
||||
internal Task<bool>? Tick { get; set; }
|
||||
|
||||
/// <summary>The pending read from the server's push channel.</summary>
|
||||
internal Task<VaultEvent>? Notice { get; set; }
|
||||
|
||||
/// <summary>The stream <see cref="Notice"/> was taken from, to notice a reconnection.</summary>
|
||||
internal IVaultEventStream? Watching { get; set; }
|
||||
}
|
||||
|
||||
/// <summary>Shows one kind of item, if nothing is being edited.</summary>
|
||||
/// <remarks>
|
||||
/// Takes the section rather than there being one command per kind, so a third kind is an enum member and
|
||||
|
||||
Reference in New Issue
Block a user