Come back from a sync position the server will not accept
ci / build and test (push) Failing after 2s

"The server returned 400: The sync cursor is not valid for this vault. Resync
from the beginning." told the user exactly what to do and gave them no way to
do it. The cursor is the only thing a pull sends, so the refusal was permanent:
the next pass read the same stored cursor and was told the same thing, once a
minute, for ever. And because the pull runs first, the exception ended the pass
before it reached the outbox — so the vault stopped receiving other machines'
changes and stopped sending its own. A machine that met this went quietly
read-only until somebody deleted its cache.

The engine now does what the message asks. A pull refused with the
invalid-cursor problem code — the code, never the prose, which is free to
change — drops this vault's position, writes that down, and reads the log again
from the beginning. The restarted request carries no cursor, which is the one
position a server cannot reject, so the retry cannot loop; a refusal of that is
rethrown rather than retried, and a restart is allowed once per pull. The
position is saved before the replay starts, so a process that dies halfway
through begins the next one from the beginning too rather than meeting the same
refusal again.

The mirror is deliberately kept. Replaying rewrites every row the server still
has and applying a change is a blind overwrite, so the re-pull repairs the
mirror on its way past; clearing it first would claim more than the evidence
supports — the position was refused, not the contents — and would leave a
machine that lost its connection mid-replay with less than it started with.
That leaves one gap, named in the remarks rather than left to be discovered:
once tombstone collection exists, a replay stops carrying deletions older than
the retention window.

None of the causes are the user's doing — a rotated cursor signing key, a vault
served from a restored database, a cache copied between machines — so nothing
asks them to decide anything. The report carries ResyncedFromStart and the
status line says the position was not recognised and the vault was read again.
It is kept out of NeedsAttention, because nothing is outstanding, but the
background pass breaks its usual silence for it: a sync that pulled the whole
vault on a day nobody changed anything otherwise reads as a fault.

The fake server grew a switch that refuses cursors the way a rotated signing
key does, including ones it minted itself. Three cases: the vault is re-read
and the change on the far side of the refused position arrives; the edits
waiting in the outbox are still pushed in that same pass, which is the half
that made this worth recovering from rather than merely reporting; and a server
that refuses the beginning itself is surfaced instead of replayed against.

dotnet build is clean at zero warnings, dotnet format is clean, and the sync
and app suites pass — 109 and 101.
This commit is contained in:
2026-07-31 11:32:14 +02:00
parent 240aadb746
commit 9608d73747
5 changed files with 216 additions and 11 deletions
@@ -1,3 +1,5 @@
using DodoSSH.Client.Api;
using DodoSSH.Contracts;
using static DodoSSH.Client.Sync.Tests.SyncHarness;
namespace DodoSSH.Client.Sync.Tests;
@@ -118,6 +120,70 @@ public sealed class SyncEngineTests
after.Cursor.ShouldBe(before.Cursor);
}
[Fact]
public async Task ARefusedCursor_ReadsTheVaultAgainFromTheBeginning()
{
// The server's cursor signing key was rotated, or the vault is being served from a restored
// database. The stored position is refused for good, so a pass that only reported the 400 would
// leave this machine frozen at it — asking once a minute, for ever, and being told the same thing.
using var harness = await CreateAsync();
await harness.First.SyncAsync();
var entityId = await harness.Second.CreateAsync(Host("prod-db"));
await harness.Second.SyncAsync();
harness.Server.RefuseCursors = true;
var report = await harness.First.SyncAsync();
report.ResyncedFromStart.ShouldBeTrue();
report.Pulled.ShouldBe(1);
// And the point of all of it: the change that was on the far side of the refused position is here.
(await harness.First.FindAsync(entityId)).Secret.Label.ShouldBe("prod-db");
}
[Fact]
public async Task ARefusedCursor_DoesNotStrandWhatIsWaitingToBePushed()
{
// The half that made this worth recovering from rather than merely reporting. The pull runs first,
// so a pass that gave up on the refusal never reached the outbox at all: every edit made on this
// machine stayed queued behind a position the server was never going to accept again.
using var harness = await CreateAsync();
await harness.First.SyncAsync();
var entityId = await harness.First.CreateAsync(Host("prod-db"));
harness.Server.RefuseCursors = true;
var report = await harness.First.SyncAsync();
report.ResyncedFromStart.ShouldBeTrue();
report.Pushed.ShouldBe(1);
harness.Server.Find(entityId).ShouldNotBeNull();
}
[Fact]
public async Task ARefusalOfTheBeginningItself_IsReportedRatherThanRetried()
{
// "From the beginning" is the one position a client is allowed to ask for, so a server that
// refuses it is one this code cannot reason about. Saying so beats replaying the log against it
// until a page bound runs out.
using var harness = await CreateAsync();
await harness.First.SyncAsync();
harness.Server.RefuseCursors = true;
harness.Server.RefuseEvenTheBeginning = true;
var failure = await Should.ThrowAsync<DodoSshApiException>(
async () => await harness.First.SyncAsync());
failure.Code.ShouldBe(ProblemCodes.InvalidCursor);
}
[Fact]
public async Task ClockSkew_IsRecordedAndNotActedOn()
{