Commit Graph
13 Commits
Author SHA1 Message Date
jaap-jan 573f5d5668 Keep the device key in the TPM, behind a consent Windows enforces
The last of ADR 0007's three pieces, and it does not implement what that ADR
originally decided — because writing it exposed a flaw in the decision.

The ADR said "a Windows Hello gesture gating a protected blob". That does not
deliver what the rest of the document claims for it: a gate inside the process is
not a gate. A store that showed a prompt and then read a DPAPI blob would be
bypassed by malware that skipped the prompt, read the file and called
CryptUnprotectData itself — which is exactly the attacker the whole decision was
made against, and exactly the reason DPAPI alone was rejected. The presence
requirement has to be a condition of using the key, enforced below the
application, or it is decoration.

So the device key is encrypted to an RSA key created in the Microsoft Platform
Crypto Provider — the TPM — under CngUIProtectionLevels.ProtectKey. Windows
requires consent to use that key, so the prompt is not something this code can be
talked out of showing. Malware can ask for the key; it cannot answer the dialog.
That is strictly stronger than the ADR described, and most of what option D was
being saved for: the wrapping key genuinely never leaves hardware. The X25519
device key still lands in memory to open the wrap, because DSH1 fixes that wrap at
a curve the TPM cannot do — the remaining gap, and now a smaller step than it was.

CngKey is in-box, so this needed no WinRT projection and no Windows target
framework. Which is worth stating plainly because the opposite was planned: the
piece was scoped as "where the Windows TFM lands", and it turned out a platform
guard on one class was enough. Client.App and its two test projects stay on
net10.0.

Two things were measured on real hardware rather than assumed, and the second
changed the shape of the work.

The platform provider works here and holds an RSA key — confirmed by creating and
deleting one before writing anything that depended on it.

And ProtectKey prompts at key *creation*, not only at use. The comment in the
first draft of this file said the opposite, with a confident explanation: sealing
uses only the public half, so it should be silent. It is not. CngKey.Create blocks
on a dialog, because the policy means "protect this key with a PIN" and Windows
asks the user to set that up there and then. Found by writing tests around save
and forget and watching the suite hang for ten minutes waiting for somebody to
type one.

That has two consequences worth knowing before touching this file. SaveAsync is
user-facing code — it belongs on a UI thread, behind a button somebody pressed,
never on a background pass. And almost nothing in the store can be covered
automatically: two tests remain, availability and the empty-blob case, both of
which provably reach no dialog. Disabling the UI policy to make the rest testable
would remove the one property worth having.

The interface offers two things and hides both where they cannot work. "Use
Windows Hello" appears on the unlock screen only when this machine has a cached
wrap and a keystore still willing to release the key; "Use Windows Hello here"
appears in the account bar only when the machine can keep a key and has not
already registered one, so it is spent once used. Absent rather than disabled, in
both cases: a greyed-out button on a machine that never had a TPM reads as
something broken, and the passphrase box beside it is not a fallback — it is the
ordinary way in.

Both unlock paths now share AdoptAsync rather than each opening the known-host
store, building the vault and starting auto-sync. The ordering in there is
load-bearing and a second copy would be a second chance to get it wrong.

The shell's tests drive a fake keystore. Not for speed: the real one prompts on
every save and load, so a suite using it would block forever. What the shell has
to get right is which buttons appear and what happens when one is pressed, and a
fake answers exactly that. It is shared from Client.Session.Tests by source link
rather than reimplemented.

882 tests green, 6 of them new. Zero warnings, dotnet format clean.

Not verified, and not verifiable here: the dialogs. Whether the consent prompt
appears at the right moments, reads sensibly, and returns to a usable window when
declined needs the application run by a person on a machine with a TPM. That is
the remaining half of outstanding item #7, and it is now the only thing between
this feature and being finished.
2026-07-30 15:17:30 +02:00
jaap-jan 211eba0666 Keep host key trust in the vault, and make it withdrawable
ci / build and test (ubuntu) (push) Canceled after 0s
ci / build (windows) (push) Canceled after 0s
A fingerprint approved once is now approved on every machine and survives a
restart, because host key trust is a vault item type rather than a dictionary
that dies with the process. InMemoryKnownHostStore was what shipped, so the user
was asked to verify a fingerprint on every single connection — which is the gap
most likely to train somebody to click through the one warning that actually
matters. A warning that appears when nothing is wrong teaches that nothing is
ever wrong.

The fourth item type, and like the third it cost no sync logic: a row, an EF
configuration, a migration, a server kind; a secret, a codec, a merge, a cipher,
a repository facade and a session property. One row in the client registry. The
reconciler, the mirror, the repository, the outbox and the pull filter were not
touched. SyncEntityType.KnownHostKey and AadResourceType.KnownHostKey were
already reserved, so neither the contract nor docs/crypto.md changed.

One item per (host, port, algorithm), because a server legitimately offers
several host keys and which one gets negotiated is not ours to predict. Pinning
per endpoint would make an algorithm change indistinguishable from an attack.

The label is derived rather than stored, which is the one place this type
departs from the other three. A user never names a pin — there is nothing to
name it after but the three fields it already has — and a stored label is a
second copy of data that can disagree with the first after a merge. Relabel
returns the secret unchanged, and says why.

The store answers the handshake without touching the disk. SshNetConnectionFactory
calls FindAsync from inside SSH.NET's synchronous HostKeyReceived event, over
.GetAwaiter().GetResult(), which cannot be avoided; doing SQLite I/O plus an AEAD
open per lookup there would put the handshake behind the cache. So decryption
happens in OpenAsync and RefreshAsync — on unlock and after each sync pass,
exactly where the host and key lists already reload — and FindAsync is a
dictionary read under a lock with no await inside it.

That snapshot is where the one real bug in this change lived. Install originally
merged the live pins over the freshly loaded snapshot, to protect a TrustAsync
that had landed while the read was in flight. It would also have resurrected
every pin the user had just forgotten, and stopped a withdrawal made on another
machine from ever taking effect — the store would have healed the deletion back
into existence on every refresh. Replacing wholesale and discarding the read
instead is correct because writes are the rare case: every write bumps a
generation counter, and a refresh whose stamp is stale throws itself away rather
than winning. Nothing found this but reading the method again; it is the kind of
mistake that passes every test written before it, because the test that catches
it is the one the bug tells you to write.

Forgetting is new, and persistence is what made it mandatory rather than
convenient. A mismatch is a hard refusal with no way to continue — deliberately,
and that stays — so pinning a key permanently is also a way to make a
legitimately rebuilt server permanently unreachable. Before this change the pin
died at exit and the problem solved itself; now it does not.

ForgetAsync drops every algorithm for an endpoint, and it is reachable from the
host editor rather than from the warning. Putting it on the mismatch banner would
have made it two clicks from "this may be an attack" to "connect anyway", which
is the affordance the hard refusal exists to deny. The banner already promised
the key could be removed in the host's settings; that promise is now true and
points at the button.

Trust recorded on another machine becomes visible at the next sync pass, not
immediately, and that is a decision rather than an oversight. The failure it
produces is a first-contact prompt for a host a colleague approved a minute ago:
answerable, and self-correcting on the next pass. The opposite trade — polling
the vault on the handshake thread to close a one-minute window — buys nothing
and costs the property above. The dangerous direction is not reachable at all: a
pin recorded here enters the snapshot as part of recording it, so a refresh can
never discard a local trust decision.

The server learns nothing, and this is the item type where the temptation was
real. A plaintext host column would let a known-hosts screen sort and page
without decrypting anything, and it would hand the operator the map of every
user's estate — assembled, as these things are, out of facts that are each
individually harmless. A host row concedes an address only when relay is
switched on and the database refuses to store one otherwise (ADR 0004); there is
no equivalent excuse here. The table has no column to put one in, and the EF
configuration says so where somebody adding it would be standing.

Two things about the migration in this commit are worth knowing, because both
came out of getting it wrong.

It was hand-written first, including its .Designer.cs, and that version is not
what is here. Verifying it turned up something that had been quietly assumed:
Migration_AppliedCleanly_WithNoPendingModelChanges does not check the model
snapshot. It asserts that migrations applied and that none are pending, which a
wrong snapshot satisfies perfectly — the snapshot only matters as the diff base
for the *next* migrations add, so an incorrect one passes the whole suite and
corrupts the following migration instead. The real check is to generate a
throwaway migration and confirm its Up and Down come out empty. They did, and
the generated designer was byte-identical to the transcribed one across all 1255
lines, so the hand-written work was in fact correct.

Then dotnet ef migrations remove --no-build deleted the wrong migration. With
--no-build the tool reads the previously compiled assembly rather than the files
on disk, and the probe had just changed which migration was last, so it removed
AddKnownHostKeyItem and reverted the snapshot. That turned out to leave exactly
the right diff base, so the migration here is EF's own output rather than a
transcription — a better outcome than the one that was interrupted, arrived at
by accident. Never pass --no-build to migrations remove.

Mutation tested, all three sabotages detected: dropping the algorithm from
KnownHostIdentity.For, merging instead of replacing in Install, and pointing
KnownHostKeyCipher at PortForward — which is what a cast from the wire enum's 10
would silently produce. Each is caught both by an assertion about the mechanism
and by a behavioural test that never mentions it; the resource-type sabotage is
caught by the table from d10a38d and nothing else, which is what that table is
for.

The end-to-end slice now approves the real sshd's host key through the vault,
pushes it, and reads it back on the second simulated machine — including a check
that the server learned no address, and that the second machine answers null for
an algorithm never offered.

845 tests green. Zero warnings, dotnet format clean.

Three things are deliberately not fixed. A tombstone queued over a create that
was never pushed is refused by the server as Invalid and parked; that is
pre-existing for all four item types, and the fix belongs in
VaultItemRepository.DeleteAsync rather than here. Deleting a host, or changing
its address, orphans its pins — both are correct as trust decisions, since a pin
describes an endpoint and not a bookmark, but nothing surfaces the leftovers.
And there is no interface listing pins at all: trust is created at the connect
prompt and withdrawn in the host editor. A known-hosts list is where the orphans
would become visible, and it wants the vault column rework first, for the same
reason the credential editor does.
2026-07-30 11:00:39 +02:00
jaap-jan 70b3290a77 Bind an SSH key to a host instead of picking one per connection
A host now names the key it authenticates with, or none, as a field in its
encrypted payload — so the choice follows the host to every machine rather than
being made again each time somebody connects. The per-connection "Use key"
switch it replaces was a stopgap for not having this, and keeping both would
have left two mechanisms answering one question.

This is the first payload schema version bump, and it does not work the obvious
way. A host is written at the *lowest* schema version that can represent it: one
that binds a key is written at 2, one that does not is still written at 1, byte
for byte as it was before the field existed. The version is what makes an older
client refuse to edit an item, so stamping 2 unconditionally would mean
upgrading a single machine and renaming a single host made that host uneditable
on every machine that had not upgraded yet. Confining the cost to the hosts that
actually use the field is the difference between a team noticing a bump and a
team being blocked by one. HostSecretCodec states the rule so the next field
added follows it, and a test pins the version-1 bytes against a literal rather
than against the codec, because the claim is about history: every host already in
every vault has to re-encode to what it encoded before, or the first sync after
an upgrade would push the whole vault as changed.

A binding is an item id, not a copy of the key — a second copy of a private key
is one that goes stale — which means the reference can dangle when the key is
deleted on another machine. Both places that meets are handled the same way, by
refusing rather than falling back:

- Connecting to a host whose key is gone is refused outright. A host somebody
  deliberately set up for key-only access must not quietly start offering a
  password.
- Opening such a host in the editor keeps the binding, selected, labelled as
  missing. The quieter version of the same failure is someone editing the port
  and saving, silently converting the host to password authentication with
  nothing ever having said so.

Two things this found by being falsified:

- The merge was untested for the new field, and "just take the server's value"
  passed the entire suite — a local binding change would have been discarded with
  no conflict recorded. HostSecretMergeTests already had a test written for
  exactly this class of omission; it simply had not been extended.

- Adding a nullable field exposed a defect in HostSecretMerge.Field: it
  short-circuited when the discarded value was null, so the formatter never ran
  for the one case where null is a value rather than an absence, and a field
  whose absence has a name could not report it. Now the formatter always runs,
  and "no key" appears in the conflict log where an empty string used to.

Also fixes eight nullable warnings in SyncEndpointTests left by the server-side
SSH key commit, which had omitted the null-forgiving operator the rest of that
file uses. They were invisible until an unrelated change forced the project to
recompile.

The end-to-end slice now binds its host to its key, so a schema-version-2
payload goes through the real API, the real PostgreSQL and back out on a second
machine.

745 tests green. Zero warnings, dotnet format clean.
2026-07-29 20:42:51 +02:00
jaap-jan e3fd3e1728 Sync and authenticate with SSH keys on the client
Completes the client half of SSH keys: they sync alongside hosts, appear in
their own list, and can be selected to authenticate a connection instead of
typing a password.

The reconciler and the repository were Host-typed throughout, so the choice was
to generalise them or to keep a second copy per item type. Generalised, because
ItemReconciler's whole premise is that the pull and the push paths must answer
the same collision the same way — two copies would drift the first time one of
them was fixed. What is genuinely per-type now arrives through
IItemKind<TSecret>: the cipher, the merge, the plaintext columns, and the noun
to use when telling a person what happened to their item. Generic where the
server's IItemKind is not, and for the reason that reverses there — the client
needs the concrete type, because it merges field by field.

The pull filter is derived from the same registry that builds the reconcilers.
That is the specific failure being designed out: an item type that encrypts,
merges and lists perfectly and is never once requested from the server, so it
works on the machine that made it and exists nowhere else.

No client cache migration. The item table's primary key and the outbox's unique
index already carry the entity type, and AadResourceTypes already mapped SshKey
— so a host and a key may share an id and never see each other's rows, which
SshKeySyncTests now arranges deliberately.

A key hands the server nothing in plaintext. There is a public_key_fingerprint
column and it would be accepted; leaving it null is deliberate. A fingerprint is
not secret but it is a stable identifier for a key pair, so filling it would let
an operator tell which of their users hold the same key and correlate one across
vaults, for a column nothing reads. The design allows itself one plaintext
concession — the relay address, which the relay cannot work without — and this
is not that.

A key is chosen per connection rather than bound to a host, which works the way
ssh -i does. Binding one needs a field on HostSecret and therefore a payload
schema bump, which makes every host written afterwards read-only on an older
build; worth doing deliberately rather than as a side effect of adding keys.

Three things this found, all of them by being falsified rather than by review:

- Making the reconciler generic silently turned a record comparison into
  reference equality, because == on a type parameter is not value equality. The
  effect would have been a conflict recorded on every pass for an unacknowledged
  create that had in fact landed. Sabotaging the fix left all 73 tests passing —
  nothing covered that branch — so ConflictMatrixTests now has
  AnUnacknowledgedCreateThatDidLand_IsDroppedQuietly, which fails without it.

- A test asserting that a blank passphrase reaches SSH.NET as null was vacuous:
  it exercised the editor, not the credential path, and passed with the guard
  deleted. Resolved by making SshKeySecret.Passphrase normalise an empty string
  to null, so there is one spelling of one state — which also keeps two clients
  from producing different payload bytes for an identical key. That exposed a
  wider gap: SshKeySecret, its codec and its merge had no direct unit tests at
  all. They have 25 now.

- The reason first given for that normalisation was false. It claimed SSH.NET
  rejects a passphrase supplied for an unprotected key; measured against a real
  sshd it ignores it and authenticates anyway. Corrected everywhere it was
  stated and recorded in docs/platform-flags.md. The same test file also closes
  a real hole: SshPrivateKeyCredential had never been exercised against a
  server, because the existing key test builds SSH.NET's auth method directly
  and bypasses the path a vault-held key actually takes.

Only one editor may be open at a time. Both sit in the same 340-pixel column as
Auto rows and their heights together exceed it at the window's minimum size, so
two open editors put the lower one's Save and Cancel past the bottom edge — the
same failure this window already shipped once with the setup screens. Expressed
as a state rule because that is the only form of it this repository can check:
nothing here loads a .axaml. The refusal keeps what was typed, since in the key
editor that is a pasted private key the user may have nowhere else.

The end-to-end slice now carries a key as well as a host, so both item types go
through the real API, the real PostgreSQL and the real crypto in one pass — the
three hand-kept mappings between enums that do not line up are the reason that
is worth doing rather than trusting the unit suites.

735 tests green, including the container-backed SSH and end-to-end suites. Zero
warnings, dotnet format clean.
2026-07-29 20:27:23 +02:00
jaap-jan 586cb303d5 Merge branch 'claude/gallant-brahmagupta-1f8244'
ci / build and test (ubuntu) (push) Has been cancelled
ci / build (windows) (push) Has been cancelled
Writes down that locking the vault leaves shells running, and shows the count on
the unlock screen rather than leaving it to be inferred.

Conflict resolution:

- ShellFlowTests' fixture keeps main's FakeSshConnectionFactory. The branch added
  an IdleSshConnectionFactory for exactly what main's fake already does — a shell
  that is open, silent and never closes on its own — so FakeSshConnections.cs is
  dropped rather than merged, leaving one fake SSH stack in the suite instead of
  two that would drift apart.
- MainWindowViewModel and TerminalWorkspace: both sides added their own members,
  so both are kept.
- TerminalWorkspaceTests was added by both branches, with the renderer gate on
  one side and session lifetime on the other. Merged into one class over one set
  of helpers; the gate tests now use FakeConnectionFactory rather than an
  NSubstitute stub, since the suite already has the fake. gallant's polling
  Timeout constant is PollTimeout, which no longer reads as the renderer's.
- platform-flags.md keeps main's measured focus section and drops the short
  "nothing hands the terminal keyboard focus" entry the branch still carried,
  which that section supersedes.

One genuine disagreement between the branches, left visible rather than
flattened: this branch measured that a collapsed WebView cannot be typed into and
attributed it to a hidden WS_CHILD window being ineligible for keyboard focus,
while main's focus work measured Win32 focus still held by that hidden window and
added a lock path that moves the keyboard off it. Both results stand; the
mechanism sentence now defers to the focus entry, which makes the input barrier
something the lock path maintains rather than something the platform guarantees.

Full suite green, including the container-backed SSH tests.
2026-07-29 15:46:41 +02:00
jaap-jan 74341d41e0 Merge branch 'claude/distracted-ritchie-53fc70'
Bounds the renderer wait, so a WebView2 that never initialises reports itself
instead of hanging Connect with the busy flag stuck.

Conflict resolution, all of it in the App test suite, which main had changed
under the branch when sleepy-chebyshev landed:

- The workspace fixture keeps main's fake SSH factory and its FakeRenderer-aware
  page, and takes the branch's RendererTimeout on top. One second rather than
  the branch's 250 ms, because the timeout now also bounds FakeRenderer's own
  wait for the attach it just made.
- FakeRenderer arrived on main after the branch was cut and still called the
  no-argument WaitForRendererAsync. Both sides merged cleanly and left the build
  broken; it now passes its own token.
- ConnectingWithNoRenderer's remark claimed the suite never starts the workspace
  and never attaches a renderer. Both are false here, so it now says what is
  true of the test: it is the one connect test that attaches no renderer.
2026-07-29 15:40:25 +02:00
jaap-jan 9270d0cba5 Merge branch 'claude/sleepy-chebyshev-cda68d' 2026-07-29 15:35:02 +02:00
jaap-jan d459dac600 Stop a dead WebView2 hanging Connect with the busy flag stuck
VaultViewModel.ConnectAsync awaited TerminalWorkspace.WaitForRendererAsync
with no timeout and no token, and RunAsync clears IsBusy only after the
work returns. Whether the renderer attaches at all depends on a runtime
this application does not install: with a missing or policy-blocked
Evergreen runtime, or an AppContainer that cannot reach loopback, the
socket never arrives — so Connect never returned, the window stayed
disabled on "Connecting…" for the rest of the session, and nothing on
screen said why. Left out of 0500e43 to keep that change focused, and
recorded in docs/platform-flags.md as worth fixing on its own merits.

The gate itself is unchanged and has to stay: TerminalDataPlane.SendAsync
drops frames when no renderer is attached rather than queueing them, so a
session opened before the renderer arrives loses its SessionOpened frame
and then streams output at a terminal that was never created. Only the
wait changed — RendererAttached.WaitAsync(timeout, cancellationToken),
with the command's own token threaded through.

Fifteen seconds, on TerminalWorkspaceOptions.RendererTimeout. Attaching is
normally near-instant, since WebView2 starts with the window and the page
has usually attached while the passphrase was still being typed, but a
first run on a cold profile creates a user-data directory and starts a
process tree of some thirty-five processes first, which on a loaded
machine is seconds rather than milliseconds. A renderer that will never
attach will not attach however long the wait is, so being generous costs
only how long a broken runtime takes to say so, while being tight costs
telling someone their runtime is broken when it was merely slow.
Injectable because both new tests would otherwise sit out that budget.

The timeout is caught in VaultViewModel rather than left to RunAsync's
generic handler, because TimeoutException.Message is "The operation has
timed out" — which sends someone looking at their network or their host.
The status now names the WebView2 runtime and says to install it.

TerminalWorkspaceTests covers the half that was missing: the wait gives up
(329 ms against a 250 ms budget) and obeys its token (2 ms against a
five-minute one). Before the bound, the first of those would have hung
rather than failed. ShellFlowTests never starts its workspace, which from
the view model's side is indistinguishable from a WebView2 that failed to
initialise, so it asserts that the status names WebView2 and that IsBusy
is cleared; changing the catch to another exception type makes it fail
with "The operation has timed out.", so neither assertion is vacuous. The
success path is untouched and still covered end to end by
TerminalEndToEndTests against a real sshd container, which now passes the
test's cancellation token.

One byproduct: the doc comment on WaitForRendererAsync carried two
double-encoded em dashes, fixed now that the block is rewritten.
2026-07-29 15:26:53 +02:00
jaap-jan c6fc19bbbd Sync the vault automatically instead of only on a button press
Three triggers: once when the vault opens, straight after any local change,
and every minute while it stays open. The Sync button stays, because someone
just handed a credential wants to know now rather than within the minute, but
nothing depends on it being pressed any more.

A background pass is deliberately not the button's code path. Routing it
through RunAsync would raise the busy flag every minute — disabling Connect
and Save for the duration — and repaint the status line over whatever the user
was reading. So it is quiet: the status changes only when a pass actually
moved an item or produced something needing attention, and a pass is skipped
outright while a command is running rather than queueing behind it. Both
guards are covered; removing either fails a test.

A shared semaphore serialises every pass, taken with a zero timeout rather
than awaited — a pass arriving while another runs has nothing to add by
waiting, and queueing them would turn a slow server into a backlog of
identical work.

Failures are swallowed, which is right in exactly this one place: a laptop
closed all afternoon would otherwise replace the status line with a socket
error once a minute. It is quiet rather than hidden — the account bar already
shows when there is no connection, and pressing Sync reports the real reason.
What earns that is the outbox: a test proves a change left queued by a failed
pass is still sent by the next sync, so quiet never means lost.

Two existing tests asserted the opposite behaviour — that a save queued and
pushed nothing until Sync was pressed — and were rewritten rather than
deleted; the local-first guarantee they were really protecting is that the
list updates with no server, which the offline test still covers.

Two things the tests caught in my own work. ReloadAsync had to be split out
of LoadAsync because rebuilding the list repainted the status line
unconditionally, which made "the background pass is quiet" false on the one
path that mattered. And the yields-to-a-command test was vacuous as first
written: saving pushes, so there was no pending change left and the assertion
held with the guard deleted. It now fails the automatic push first to arrange
a real queue.
2026-07-29 14:53:19 +02:00
jaap-jan 5a899afd78 Decide what Lock does to a running shell, and say it
Pressing Lock nulled and disposed the vault view model and touched nothing else.
TerminalWorkspace is injected from App.axaml.cs and outlives every lock, so the SSH
connection, the pty and the pump all kept running while the window said "Unlock your
vault" — and since 0500e43 collapsed the WebView while locked, that live session was
invisible as well as unstopped. CloseSessionAsync was reachable in production only from
DisposeAsync, i.e. shutdown. None of this was written down anywhere, so it was neither a
policy nor a bug, which is the actual problem.

Shells now deliberately outlive the lock, and every layer says so.

The reason to prefer this over making Lock a disconnect: locking is what a person does
when they walk away from the machine, which is exactly when a long upgrade, build or
transfer is most likely to be in flight. Ending every shell would make Lock a button that
destroys work, and the predictable response is to stop pressing it and leave the vault
open instead. The idle auto-lock this will grow decides it outright — an unattended
timeout that killed a running job would be worse than the exposure it removes. Closing
the channel also buys less than it looks: the session was authorised at connect time by a
credential the remote verified itself, and no vault key participates in keeping it alive,
so locking cannot retroactively un-authorise it any more than removing a member can.

Stated honestly rather than implied, because the lock screen is what hides it:

- The unlock screen shows how many shells are still connected, and that locking closes
  the vault and not the connections — so a machine still holding authenticated SSH
  channels does not present itself as merely "locked". Shown only when there is something
  to disclose. Quitting is what ends them, and the text admits that.
- The Lock button carries the same thing in a tooltip, since its name implies the
  opposite of what it does to a shell.
- README lists it as a third architecture consequence beside non-retroactive revocation,
  which is the same shape of honest limit; docs/crypto.md §10 records it as a threat-model
  boundary; TerminalWorkspace and LockAsync carry the argument next to the code.

LiveSessionCount deliberately does not count dictionary entries. Nothing removes a
session when the remote closes the channel by itself — RunSessionAsync only drops the
renderer registration — so sessions.Count would report a shell that exited half an hour
ago as still running, on the one screen where a user is deciding whether it is safe to
walk away. A completed Run task is what "the shell is gone" actually looks like. While
locked the number can only fall, since opening a session needs the vault, so a stale
value over-reports rather than under-reports.

Both new tests fail when the policy is reverted: the count test times out against
sessions.Count, and the shell test reports "workspace.LiveSessionCount should be 1 but was
0" when Lock closes sessions. ShellFlowTests also stops building its workspace with a real
SshNetConnectionFactory that nothing ever called, which had made the suite's independence
from the network a coincidence rather than a property.

Verified by hand with a live shell, which nothing had done: a harness mirroring
MainWindow.axaml's 340,* grid with a real NativeWebView, the shipped WebAssets, a real
sshd in a container, and an ISshShellSession decorator recording every window-change the
remote is actually told about. Across lock and unlock, no window-change reached the remote
at all, stty size answered 50 118 before and after, the renderer's own buffer came back
byte for byte with the wrapped line intact, and the session stayed live throughout. A
control run that never hides the WebView behaves identically, so nothing above is startup
or idle behaviour. Keystrokes injected while locked reach nothing: twelve of twelve
SendInput events accepted with the harness confirmed as the foreground window, no probe
character in the remote's output, and a following Ctrl-U answered BEL, so nothing was
queued in the line editor either. A hidden WS_CHILD window is not eligible for keyboard
focus, which is what makes surviving the lock defensible rather than merely convenient.

Correction to a claim made in f80b3d4: terminal.js's guard comment listed "a host that
hides the WebView while the vault is locked" among the paths that reach a degenerate fit.
It does not. Collapsing the control hides a native child window without resizing it, so
the page still reports paneWidth 840 and paneHeight 760 with unchanged cols and rows, no
ResizeObserver callback fires and the fit never runs. Establishing that rather than
assuming it: the same cycle with MINIMUM_FITTABLE_PIXELS patched to 0 — the guard fully
disabled — is equally clean. The guard is still right for minimising and for a splitter
dragged to the edge; it is simply not what makes locking safe, and must not be cited as
though it were.

Recorded, not fixed:

- Nothing closes one terminal from the interface, so a user reading "1 shell is still
  connected" can only act on it by quitting. CloseSessionAsync is tested and correct;
  VaultViewModel discards the session id it would need.
- A session whose remote exits keeps its ISshConnection, and the thread ShellStream parks,
  until the process ends.
- Suspected and seen once: before the harness waited for the window's scale to settle, a
  DPI settle pushed a 2202x1328 pane for a window 1180 logical units wide and a later
  re-push reflowed the wrapped line. Three later runs at RenderScaling 1.00 never showed
  it, so it is filed as a lead, not a finding.
- WebView2 fails to initialise with CO_E_SERVER_EXEC_FAILURE when the host executable
  sits under a very long path. Cost an hour on the harness; relevant to packaging.
2026-07-29 14:44:02 +02:00
jaap-jan dbddbcd711 Hand the terminal the keyboard on connect, and take it back on lock
After a successful connect the first keystrokes went to the shell's UI rather
than the remote shell. The page's own term.focus() focuses the textarea inside
the document, which does nothing while the window's keyboard focus is still on
the Connect button, so the terminal had to be clicked before it would accept
anything.

The obvious guess about the fix — that reaching a native child window needs
SetFocus through P/Invoke — is backwards, and measuring it first is what kept
this small. NativeWebView overrides Focusable to true and its OnGotFocus calls
the adapter's Focus(), which on Windows is
ICoreWebView2Controller::MoveFocus(PROGRAMMATIC). So a plain Avalonia
Terminal.Focus() really does move Win32 focus into WebView2. Measured in a
standalone harness with no DodoSSH code, on the same 340,* grid as the shell,
reporting GetFocus() and the page's own document.hasFocus() at each step: focus
lands on the Chrome_WidgetWin_1 child and the page reports hasFocus: true.

It is the return trip the package does not implement. OnLostFocus calls the
adapter's ResignFocus(), and on Windows that method body is empty, so Avalonia's
focus and Win32's diverge: after textBox.Focus() the focused element is the text
box while the keyboard is still on WebView2 — a caret that silently receives
nothing. Window.Activate() and Window.Focus() were both measured and neither
recovers it, so the hand-back is a SetFocus on the top-level, in
Views/NativeKeyboardFocus.cs. A real mouse click does recover it, because
Avalonia's window sets focus on pointer input, which is why this is invisible to
anyone who clicks before typing.

That turned up a worse defect than the one being fixed, and it shipped in
0500e43. Collapsing the WebView does not release the keyboard: focus stays on
the hidden holder — measured held by a window reporting visible=False — while
Avalonia's focused element becomes (none). So a user who had clicked the
terminal and then pressed Lock got an unlock screen that swallowed the
passphrase. Locking now hands the keyboard back and focuses that box.

Ctrl+Shift+F6 is the way out for someone using only a keyboard. It has to be
handled in terminal.js and posted to the host as a web message, because once the
child window owns Win32 focus Avalonia receives no key events and no KeyBinding
could fire; the package also subscribes MoveFocusRequested and discards it, so
there is no Tab-out to lean on. Not Escape, which vim alone rules out, and not a
bare F6, which TUIs bind — Ctrl+Shift is the range terminal emulators
conventionally keep for themselves and never forward to the remote. Verified
rather than assumed: the posted string arrives verbatim in Body, and the chord
reaches the page as F6 with both modifiers.

Order matters and is now recorded. Focus() on a collapsed control is a measured
no-op and is not replayed when it is revealed, so focus survives a lock/unlock
cycle only because a session can be opened solely from an unlocked vault, which
is what reveals the control in the first place.

The view models still reference no view. VaultViewModel raises SessionOpened on
the success path only, the shell forwards it as TerminalSessionOpened through the
generated OnVaultChanged hook so unlock, lock and dispose all attach and detach
in one place, and the view holds the whole focus policy. An event rather than a
bound flag because connecting a second host while one is open has to move focus
again, and no state change describes that.

Three tests, and what they do not cover is the point. They cover the plumbing:
focus is asked for once per session, a failed connect does not ask at all — a
host-key prompt needs the keyboard on its own buttons — and locking stops the
forwarding. They cannot cover the focus call, because headless Avalonia has no
native window, so a headless test would focus correctly and confirm the wrong
belief; that is measured in the harness and written down in docs instead.
Dropping the forwarding fails two of them and dropping the detach fails one;
deleting the raise outright does not compile, since the event would be unused.

Reaching the connect path at all needed two new fakes. FakeRenderer attaches the
way the real page does — fetch the served page, read back the token and socket
URL the host substituted into it, then open the socket with both subprotocols —
rather than being handed the token, so the part of the handshake that has been
got wrong before stays under test. FakeSsh replaces a factory that would need a
reachable sshd, which DodoSSH.Client.Ssh.Tests already covers against a
container. The suite also never called workspace.Start(), so nothing served the
page and no renderer could have attached.

DllImport rather than the source-generated LibraryImport, which requires
AllowUnsafeBlocks for the whole project. The signature is blittable so there is
no marshalling stub to improve on, and turning unsafe code on across a client
that handles key material to gain nothing is a poor trade.

Correcting an earlier entry: docs/platform-flags.md described this as a
focus-plumbing gap and offered "click inside the terminal first" as the
workaround. Both true, and both stop short of the half that matters — focus
crosses into the WebView readily and never comes back on its own, which is the
same mechanism as the text boxes that mysteriously stopped accepting keystrokes
in the airspace entry above it, not a separate fault.
2026-07-29 14:31:09 +02:00
jaap-jan 0500e43e02 Stop the terminal's WebView painting over the setup screens
The shell layered its setup and unlock screens over the terminal, which does
not work: NativeWebView attaches a real Win32 child HWND through
NativeControlHost, and a child window composites above everything its parent
paints regardless of visual-tree z-order. The cards rendered sliced at the
terminal column's left edge; at the window's default width every one of their
buttons fell inside the WebView's rectangle, so the flow could only be
completed by keyboard, and a click in that region handed Win32 focus to
WebView2 so the text boxes silently stopped accepting keystrokes.

The WebView is now collapsed while the vault is not unlocked. The comment
that previously forbade this — hiding it means never realising it — was
wrong: NativeControlHost creates the native attachment on attach to the
visual tree, never consulting layout or visibility, and NativeWebView replays
a Source assigned before its adapter exists. A collapsed WebView still starts
WebView2, loads the page and lets the renderer attach. Confirmed: 35
msedgewebview2 processes with the control collapsed. What the first
connection after unlocking actually depends on is the existing await on
WaitForRendererAsync, since the data plane drops frames when no renderer is
attached.

Also fixes the second visible defect: the default server URL was
https://localhost:7217, the API's *second* launch profile, while the README,
its appsettings and a plain `dotnet run` all use http://localhost:5233 — so
nothing was listening, and an HTTPS client against a plaintext port reports
"The SSL connection could not be established", which reads as a certificate
problem. The default now matches, a missing scheme is rejected by name
instead of parsing as scheme "localhost", and that specific TLS failure now
suggests http://. Both new tests fail when the fixes are reverted.

Corrections to claims I made earlier and should not have:

- docs/platform-flags.md asserted the opposite of the mechanism above and
  cited an established msedgewebview2 connection as verification. That
  observation was taken while the overlay was showing but, because of this
  very bug, the WebView was uncovered and in plain view — so it confirmed
  only that a visible WebView is realised. A process-level check cannot
  verify a rendering claim. The entry was also filed under "Local cache".
- ITerminalHost was documented as the live seam the app plugs into, with a
  stub standing in for headless tests. It has no implementation anywhere and
  no test uses it; the view navigates the control directly. It also counted
  Avalonia.Controls.WebView and NativeWebView as two interchangeable
  backends when they are one component, with the Linux backend backwards.
- The README claimed the shell's whole path was covered by tests. Its state
  machine is; its layout is covered by nothing, and a headless test could
  not have caught this — headless has no native window, so it would have
  rendered correctly and confirmed the wrong belief.

Verified by screenshotting the running app: the card renders complete and
centred at the default size, with the button clickable.
2026-07-29 13:26:30 +02:00
jaap-jan 49f617b450 Wire the Avalonia shell to the vault
The host list now comes from the vault instead of from a form. A fresh
machine takes a server URL, signs in through the browser, enrolls, and
from then on opens with the passphrase alone.

DodoSSH.Client.Session is the composition layer: where a profile lives,
how it unlocks, and how a machine gets one. ClientPaths picks a
non-roaming per-OS directory — %LOCALAPPDATA% and never %APPDATA%,
because a SQLite cache that roams between two machines is a corrupt one,
and each machine's outbox is its own. SessionOpener needs no transport at
all and could not reach one if it wanted to; that is the offline unlock,
asserted rather than asserted about. A wrong passphrase, a stale KDF and a
grant revoked by a rekey are three different answers, because the remedies
are three different things and telling someone to retype a passphrase that
was never the problem is worse than saying nothing.

The shell's states are the onboarding story. The recovery code gets its
own state that cannot be clicked past: it exists for one moment, losing it
with the passphrase loses the vault, and there is no server-side reset by
design. It is dropped from memory on confirmation rather than merely
hidden.

Sign-in is a delegate over IVaultServer, so the whole state machine runs
in a test against an in-memory server — no browser, no identity provider,
no toolkit. The view models are plain observable objects, which is what
makes that possible. What it does not cover is whether the XAML binds to
the right names; that needs a rendered tree and Avalonia.Headless, and is
its own piece of work.

Three things found by doing it rather than by reading it:

- Pooled SQLite connections keep the database file open after the last
  context is disposed. On Windows that means locked, so the application
  could never replace its own cache — and a test could not clean up after
  itself, which is how it surfaced. Dispose now clears the pool.
- EF's SQLite provider puts the database in WAL mode, so the cache is
  three files. A comment in ClientCacheFactory claimed the opposite;
  reading PRAGMA journal_mode off a real launch settled it. WAL is the
  right mode here — a sync pass writes while the interface reads — so the
  comment was wrong on the merits as well as on the fact.
- Enrolling a device key with nowhere to keep the private half would put a
  wrap on the server nobody can open and make the device list claim this
  machine can unlock without a passphrase. Device binding is now optional
  and the shell declines it until the OS keystore is wired.

Verified on Windows: the client created %LOCALAPPDATA%\DodoSSH\cache.db
and migrated it on first launch, and msedgewebview2 held an established
connection to the data plane while the unlock overlay covered it — which
is the point of covering the WebView rather than collapsing it, since a
NativeWebView that is never laid out is never realised.

630 tests, up from 593. The recovery-code gate and the offline unlock were
each verified by breaking them and watching the right test fail.

Still to do for M1's actual definition of done: the manual run against the
real API and a real Keycloak. Credentials are not a synced entity type
yet, so a connection still asks for a password, and the interface says so
rather than implying otherwise.
2026-07-29 11:02:19 +02:00