Commit Graph
157 Commits
Author SHA1 Message Date
jaap-janandClaude Opus 5 9629b7d938 Write down how the phone will add hosts, before it adds any
ci / build and test (pull_request) Successful in 1m12s
ci / android head (pull_request) Failing after 5s
ci / api image (pull_request) Successful in 3s
The + button the design has asked for twice needs three things that do not
exist: a tag, a group's parent, and a group's defaults. Two of them are
refused on the record — HostGroupSecret argues groups are flat, and the
theme argues an unused style is a claim the control exists.

So the plan goes in first, with the decisions and the reasons. Tags become a
real item over the reserved slot, a host names them, and membership merges
through ThreeWayMerge.Map rather than as a whole value, so two people tagging
one host both keep theirs — which is what HostTag was going to buy. Group
defaults inherit rather than copy, shown as the field's placeholder, which is
what makes editing a group afterwards mean anything. A parent arrives with a
visited-set walk, because with inheritance a cycle is no longer an undrawable
sidebar — it is a shell that never opens.

Also written down: that Port has to go nullable and everything that touches,
the byte pin it will trip, and the schema-version branch whose absence would
make an inheriting host unreadable rather than read-only on every client that
has not been upgraded.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 22:07:30 +02:00
jaap-janandClaude Opus 5 359087f1ce Give the tabs their hover back, and the plus the shape its comment claims
ci / build and test (push) Successful in 1m12s
ci / android head (push) Failing after 5s
ci / api image (push) Successful in 27s
Two things v2 broke in the last commit, both found by reading resolved brushes rather than markup, and both
the same mistake: Avalonia has no specificity, so the later declaration wins, and a rule that restated the
base instead of excepting from it went in below the rules it was supposed to be underneath.

Making a tab a pill gave it a Background of its own, which the flat tab it replaced never had. That one
detail moved where the hover has to live. `Button.flat:pointerover` is declared far above and had been
supplying it; the moment `Button.tab`'s template rule set a Background, it won, and every tab in the strip
stopped answering the pointer. Silently — a tab that no longer lights is not a crash and not a layout
change, and nothing in a suite that measures heights and reachability can see it.

The `+` lost more than that. It sits below as an exception — no outline, because it is not one of the
things being chosen between — and a `.tab` rule declared after it was overriding the exception itself. It
drew as a filled, outlined pill identical to a tab, contradicting the comment directly above it.

So the base pill and its hover come first now and the exceptions follow, which is the order the rest of
this file already uses and the order the Border.rowmark note further down was written about. The `+` clears
the fill as well as the border, because an exception to a rule that sets both has to say both.

The titlebar's search box had the same shape of error in geometry rather than colour. The design draws it
at exactly 380 and centred, and stating that as a Width on the inner Border is what made it wrong: the
button around it is free to shrink when the account name or the vault chip beside it is long, and a Border
that will not shrink with it arranges outside its own parent — over the name on one side and over the
window buttons on the other. MaxWidth on a stretching button gives the same 380 whenever there is room and
gives way when there is not.

The regression test is the point of this commit rather than an afterthought. It is the only test in that
suite that reads a brush, and the gap it fills is exactly the one these two went through: everything else
measures rectangles. It hovers a tab through the real input path and asserts the fill changes, then asserts
the `+` is neither filled like a tab nor outlined like one. Checked against the broken ordering before
being kept — it fails there and passes here, which is the only thing that makes a regression test worth
committing.

Verified by the whole suite: 1310 tests over nineteen projects, none failing, the layout suite now 70
cases. Both heads build. Still nothing seen on a display.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AZE3u99BNt6LzgTC5jhbz2
2026-08-02 19:54:37 +02:00
jaap-janandClaude Opus 5 3627021420 Give the desktop the second design too, and the window the size it now needs
The desktop v2 design is the other half of the one the phone took last commit, and this is its chrome: a
190-pixel labelled sidebar where the 54-pixel icon rail was, a titlebar with the search box centred in it,
and session tabs drawn as pills. The palette was already here — it is shared, and moved when the phone's
did — so what this changes is shape rather than colour.

**The window's minimum grew, and by exactly what v2 added.** The sidebar is 136 wider and the chrome 14
taller, so 880x560 became 1016x574. That is not a round number somebody liked: it leaves every screen the
same 826x464 it was designed against, which is the arithmetic the layout suite is built on. Four of the
tables stop fitting at 690 wide, so widening the sidebar and leaving the window alone would have broken
them somewhere no test was looking. LayoutHarness carries the new constants and the suite still passes at
the minimum, which is the whole reason it exists.

The rail's five-character abbreviations are gone with the width that caused them — PINS and SNIPS are Pins
and Snippets again — and each row gains a glyph and a count. A count is drawn only where one is real, so
SFTP, Logs and Preferences show nothing rather than a zero: a transfer queue's depth is not how many files
a screen holds, and a log has no total until it is read. The count beside Pins is the vault's own, not the
Pins screen's VisiblePins, which is the filtered list and would have made the sidebar count whatever
somebody had just typed into a filter box on another screen. Teams has no count for a related reason: they
are read from the server when that screen is opened rather than on unlock, so a number there would read 0
until somebody had already been to look.

One colour moved with it, finishing what the repalette started: the live-session summaries on the unlock and
sign-out cards were Info, so the two heads disagreed about a fact the phone paints green. They match again.

**Buckets became a destination rather than a mode**, which is what the design draws and what the phone
already does. The HOST / BUCKET pair inside the files screen is gone; ShellScreen.Buckets draws the same
TransfersScreen with the other picker, and the sidebar entry is what sets it. That also settles an old
disagreement rather than merely moving it: TotalItemCount is keys plus passwords and excludes buckets, so
the number beside the keychain used to disagree with the list under it, and now counts what that screen
shows.

There is one session behind both file destinations, so asking for the other kind while something is open is
refused rather than obeyed — and refusing means staying put. An earlier turn of this had it move anyway and
only decline to switch the picker, which put the S3 entry in the sidebar over a screen still listing an
SFTP host: two pieces of chrome disagreeing about where you are, which is worse than the navigation simply
not happening. The message that says so goes to Transfers.Status, which turned out to be drawn in the same
grid cell as the connected chip — survivable while it was mostly read before connecting, and not once a
refusal reports itself there. It has its own column now.

The design has nine entries' worth of screens and draws five. Pins, Teams, Import and Preferences are
built, working screens, so they keep their entries — the sidebar is labelled now and has the room, and
dropping an entry would have stranded a screen rather than simplified anything. The Team vault card the
design pins to the foot is not drawn: it is a second route to a screen already in the list, carrying a seat
count nothing here produces.

**The status bar survives the design that deletes it**, cut down to one thing. Two of the three facts it
carried moved into the titlebar with v2 — the sync word is beside its dot and the shortcut hint is inside
the box that uses it — so those are gone from it rather than printed twice. The third is Vault.Status, the
only channel this application has for saying a save failed or a merge picked a winner. The design is a
mock-up of an afternoon that goes well and has nowhere to put a sentence like that; dropping the bar would
have meant dropping the sentence or repeating it on nine screens.

What v2 draws and this does not is in docs/design-import-gaps.md, and it is the same list as the phone's
for the same reasons: the forwarding screen and both its chips, the host detail's fingerprint, tags and
last-session cards, the keychain's rotate button, the logs' FOLLOW pill and severity filters, and the
session footer's latency. The terminal is not inset behind a rounded frame either — it is a native child
window that composites above everything Avalonia paints, so the frame would clip nothing, which is the same
answer the phone gave.

**The light theme is not built.** Its accent is #6D5AE6, a different hue rather than a tint of the dark
one, so it needs every colour doubled, a variant to switch on, the renderer's own page switching with it,
and contrast checked twice. That is a piece of work rather than a setting, and it is separable from the
layout — which is why this commit is the layout.

The screens themselves are restyled through the shared vocabulary rather than rebuilt: corner radii,
chips, cards and the accent's ink, all in App.axaml, so every screen moves at once. Their layouts are left
alone deliberately. The design draws read-only detail panes and these screens carry the editors and forms
it has no equivalent of, so replacing a layout with the mock-up's would have lost the half that is
actually used.

Verified by the whole suite: 1309 tests over nineteen projects, none failing, including the 68 layout cases
that stand up real Avalonia and measure every screen at the new minimum. Both heads build. Not run on a
machine with a display — see docs/manual-checks.md for what wants looking at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AZE3u99BNt6LzgTC5jhbz2
2026-08-02 19:29:21 +02:00
jaap-janandClaude Opus 5 5593f337b6 Give the phone the second design, and both heads the palette it arrives with
The Android v2 design is what this head draws now: four destinations in a bottom bar — Hosts,
Terminal, Keychain, More — with snippets, SFTP, S3, logs and preferences one tap deeper behind the
last. The first design's four had nothing behind them, which is what made a hub worth building.

The palette moved from green-black to blue-black, and it moved in the shared project because that is
where it lives and the desktop v2 specifies the same seventeen tokens. One colour changed meaning
rather than value, and it is the only semantic change in the file. Green used to *be* the accent, so
Ellipse.dot.live filled with Accent and "the thing to press" and "a shell is open on this host" were
the same colour by construction. v2 makes the accent blue and keeps a green for status alone, which
finally separates them: Live is that green and nothing merely interactive may use it. The accent is
also two colours now — Accent fills, AccentText writes — because a row of chips in the fill colour is
a row of things that all look like the primary action.

A palette is not one file, which is the part worth knowing before the next one. Nine hex literals
lived outside it: the nav bar's own label colours, the accessory keys and their Ctrl-latched state,
two scrims, the window background Android paints before Avalonia has a frame, and the launcher
vector. The two C# sites now resolve from the dictionary by name rather than restating it. The
renderer's page cannot — it is served to a WebView over a loopback socket — so terminal.css and
terminal.js keep hand-copied values and say so at both sites.

ShellScreen gained More and Buckets, appended rather than slotted in. SFTP and S3 are one screen over
one TransfersViewModel differing only in which picker they offer, and the kind is set by the button
that navigates rather than on arrival — doing it in OnScreenChanged made every arrival at Transfers
force the picker back to hosts, including the desktop's own rail arriving at a screen with a bucket
already open. It refuses to change kind while a session is live, because there is one session behind
both destinations and switching under it would title a screen S3 while it listed an SFTP host.

What the design draws and this does not, on the usual grounds. The FORWARDING screen: nothing here
forwards anything, so every toggle would be a control with no effect — it is a paragraph on the hub
naming the absence, for the reason the desktop keeps TEAMS in its rail. The terminal's `23 ms · fwd
5432`. An ED25519 badge and a SHA256 line on keychain cards, which need an algorithm field and a
fingerprint the item type does not have. An `agent` chip, for an agent that does not exist. Snippet
run history and exit codes. The Logs FOLLOW pill, which claims a live tail over records that are
written once at close and read when the screen opens, and the severity filter, which has nothing to
count — that chip row is spent on the real choice, which of the two logs. S3 bucket totals and
lifecycle. And the + on HOSTS, which would open a host editor this head has not got.

SFTP is browse, open and delete. Both transfer commands work, and what they work against is the local
pane: QueueDownloads writes to Path.Combine(LocalPath, name), and LocalPath starts at
SpecialFolder.UserProfile, which on Android is the application's own private directory. A download
would have reported success and left the file where the person who asked for it cannot open it, which
is worse than not offering it — a refusal is visible and a file in /data/user/0/ is not. The queue is
not drawn either, since nothing here can put anything in it. Both return with the document picker.
The foreground service still counts zero transfers, and the reason moved rather than went away.

Four defects worth naming, because three of them are the kind that compile. A Button as a ListBox
ItemTemplate swallows the pointer press before the list sees it, so the files listing selected
nothing and every command reading the selection did nothing — the row is a Border now and the
phone-only single-tap-to-open is a Tapped handler, which also keeps a desktop single click from
walking into directories. Avalonia type selectors are exact, so TextBlock.fingerprint never matched
SelectableTextBlock and every fingerprint on this head rendered proportional and unwrapped: that was
breaking the never-truncated rule on the host-key sheet already. The new two-level hierarchy had no
handler for the system back gesture, so back left the application from a log screen. And the tab's
close cross had shrunk to a 30x32 target flush against the select target, which is the one control
here that ends a shell with no confirmation and no undo.

Fingerprint unlock is raised on arriving at the lock screen rather than waiting for its button, which
is still there. Only at launch: a lock the user asked for is not answered with an immediate request
to unlock, which makes LOCK look inert and trains the reflex of authenticating at a prompt nobody
asked for. And once, because a declined gesture leaves the passphrase box exactly where it was and a
prompt that came back after being dismissed would be a modal you cannot get out of to type into it.

Two fixes fall on the desktop. Its file listing coloured directories with Info and executables with
Accent, which was blue against green and is now two steps of one blue; an executable is Live now.
And a bucket's folders were drawn with a 0001-01-01 timestamp, because a prefix has no modification
time — blank now, for the reason a directory's size is blank.

Verified by the whole suite: 1309 tests over nineteen projects, none failing, including the layout
suite that stands up real Avalonia and parses every desktop screen. Both heads build. Not verified on
a device — nothing in this head ever has been; see docs/android-port.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AZE3u99BNt6LzgTC5jhbz2
2026-08-02 18:23:53 +02:00
jaap-janandClaude Opus 5 c00e5dbc5c Let the terminal's text be made bigger, and remember how big
ci / build and test (push) Successful in 1m12s
ci / android head (push) Failing after 4s
ci / api image (push) Successful in 24s
Taking pinch-zoom off the phone left nothing in its place, and there was nothing on the desktop
either. This is the replacement, and it is deliberately not the thing that was removed: zoom scales
what has already been drawn, so the remote goes on wrapping to a width that is no longer on screen.
Changing the font size refits the grid and reports the new column count, so the far end is told it
has fewer columns. That round trip is the feature.

The size is one number, owned by the shell. It has to be, for two reasons that pull the same way: it
must survive a relaunch, and it must be reachable from a phone that has no Ctrl key to press. So the
page asks and the host decides — a signed step over a new client opcode, answered with a size over a
new server opcode. The phone's buttons and the desktop's chords arrive at the same place, and a size
set by either is the size both remember.

Stored in settings.json beside the cache rather than in it, and that is not laziness about a
migration. The cache is encrypted and unreadable until a vault is unlocked, and the first terminal of
a locked launch needs the size already. Nothing secret may go in that file; ClientSettings says so
out loud, because the next person to add a preference is the one who needs to read it.

Where it is reachable from differs per head, and only here. The phone gets A− and A+ on the
connection line — not in the accessory row, which scrolls, and a control that fixes unreadable text
must never be the thing that is off-screen. The desktop gets the three chords every terminal
emulator has, answered by the page while a terminal has focus and by the window when it does not,
plus a row in preferences that shows the current value and names the chords rather than replacing
them. Someone whose terminal is too small to read is not in a position to go looking.

Clamped 8 to 32. Below eight a monospace grid stops being legible and becomes a texture, and every
column of it is still a column the remote is being told exists; above thirty-two a phone in portrait
has too few columns to hold a prompt. The buttons disable at the ends rather than accepting presses
that do nothing, which on a terminal reads as the application having stopped responding.

The preferences screen's header comment claimed none of the design's terminal settings could be
saved, and listed the three things that were missing to make one work. All three now exist, so it
says which one is real and why the other five still are not.

Verified with the protocol suite — including that the step byte round-trips signed, since read
unsigned a step down arrives as 255 and clamps to the largest font, making "smaller" do the most
dramatic available version of "larger" — a data-plane test that the chord is heard with no session
registered, and five shell tests: the default matches the renderer's, both clamps hold, reset works,
and a size chosen in one shell is there in a second one over the same profile directory. Layout
suite and both heads build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 13:15:11 +02:00
jaap-janandClaude Opus 5 5cb9ffaf9d Stop the phone's terminal being a page you can pinch
ci / build and test (push) Successful in 1m3s
ci / android head (push) Failing after 4s
ci / api image (push) Successful in 36s
The terminal on Android could be pinch-zoomed and double-tap zoomed, which a terminal must not
be: the grid is sized to the window by the fit addon and the column count is told to the remote,
so zooming makes the visible width disagree with what the far end is wrapping to, and puts the
cell under your thumb somewhere other than where you tapped.

Three things, and the first is the one that made it look worst. The page had no viewport meta tag
at all. An Android WebView with no viewport lays out at a notional 980 CSS pixels and scales the
result down to fit, so a terminal built to fill the window was drawn small and then offered as
something to zoom around in. width=device-width makes one CSS pixel one layout pixel, which is
what the fit addon has been assuming all along and what a desktop WebView gives without being
asked.

Second, the gestures. touch-action: pan-y keeps the only one a terminal wants — dragging the
scrollback — and refuses pinch-zoom and the double-tap zoom that fired on every attempt to place
a cursor. text-size-adjust stops Android's own font inflation, which resizes text it judges too
small without telling the page and leaves the characters no longer matching the grid that was
measured.

Third, the knob that actually disables zoom: BuiltInZoomControls on the WebView. user-scalable=no
is in the viewport tag for completeness and does nothing on its own — Blink has ignored it since
Chrome 48 and WebView follows Blink. That is worth knowing before somebody removes the C# and
trusts the meta tag.

The page and stylesheet are shared with the desktop head, deliberately, and none of it costs
anything there: a desktop WebView already lays out at device width, and a Windows touchscreen
should not be pinch-zooming a terminal either.

This does not settle whether the phone should keep a browser-based terminal at all, which is the
question actually asked. It does remove the reason it was asked, and every complaint about this
renderer so far has turned out to be configuration rather than the approach.

Verified by building the head in Debug and Release and by the desktop layout suite, which draws
the same assets. Not verified on a device — nothing in this head is; see docs/android-port.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 22:45:11 +02:00
jaap-janandClaude Opus 5 52596aac76 Ask the Linux WebView for the one mode it can draw in this window
ci / build and test (push) Successful in 1m12s
ci / android head (push) Failing after 5s
ci / api image (push) Successful in 22s
The terminal renders nothing on Linux, and does it in the most misleading way available: the
page loads, the scripts run, the renderer connects, InvokeScript answers. Everything works
except the pixels, so it reads as a broken terminal rather than as a host with nowhere to
paint.

Measured on Fedora 44 with Avalonia.Controls.WebView 12.0.1, by a spike that hosts a
NativeWebView and reads AdapterInfo. The backend is WebKitGTK 2.52.5 — not the WPE one this
repository's platform notes predicted, and Fedora packages no WPE WebKit at all, so that path
was never going to be the answer here. In its default mode the adapter reports
SupportedScenarios = NativeDialog: a window of its own, and nothing that can be hosted in
place. Identical under X11 and Wayland, so it is the adapter's answer rather than a session
problem.

Setting ExperimentalOffscreen on the GTK environment arguments changes the same adapter's
answer to OffscreenRenderer — the compositor-drawn mode, which is what the
NativeWebViewCompositorHost mentioned in those same notes exists to host. MainWindow now sets
it as the environment is settled. Windows and macOS are untouched by construction rather than
by an OS check: the argument is a GTK type there and the handler does nothing.

The platform notes carried this as "unproven, and still the largest risk in the plan". They
carry the measurement now, including the part that is still unproven and the reason the spike
could not settle it.

What is NOT verified is that it now paints. An XWayland root capture is black under a Wayland
compositor and RenderTargetBitmap does not capture a compositor surface, so both ways of
looking at it from here failed. It needs eyes on a running client, and if the terminal is
still blank the next question is whether it takes input at all — that separates "not drawing"
from "not hosted".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 22:35:19 +02:00
jaap-janandClaude Opus 5 a0d53b0c9d Give the phone its screens back by not stealing a data context
Signing in on the phone reached an unlocked shell showing PREFERENCES over the middle of the
screen, with a bottom bar whose four buttons did nothing. The buttons were fine. Every one
of them changed the shell's screen exactly as asked, and nothing moved, because an opaque
panel was sitting on top of the whole page area and never came down.

PendingScreen set DataContext = this in its constructor, so that Heading and Detail could be
written as plain bindings inside its own XAML. That is not a private arrangement: a binding
the parent writes ON one of these — IsVisible="{Binding IsPreferencesShowing}" in
PhoneShell — resolves against this control's data context, which was no longer the shell.
The binding looked for a shell property on a PendingScreen, found nothing, and left
IsVisible at its default. Its default is true. So the panel that says "this screen is not
built yet" was permanently visible, last in the Panel and therefore on top of the host list
and the keychain both — and the terminal underneath them.

The two properties are read with $parent now and the control inherits its context like every
other screen. There is no x:DataType on it any more either, so a plain binding here is a
compile error rather than a silently missing screen.

The desktop head has the other half of this lesson written down already: MainWindow gives the
terminal's IsVisible a FallbackValue precisely because an unresolved visibility binding does
not hide anything, it shows everything. That note was about the previewer. This is what it
looks like at runtime.

Verified by reproducing the mechanism rather than by reasoning about it: a control that owns
its data context ignores a parent's IsVisible binding and stays visible; one that inherits
obeys it. The head builds in Debug and Release. Nothing has been run on a device, as ever —
see docs/android-port.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 22:35:04 +02:00
jaap-janandClaude Opus 5 ffab2be22a Drop the formatting step, which spent minutes agreeing with the build
ci / build and test (push) Successful in 1m11s
ci / android head (push) Failing after 4s
ci / api image (push) Successful in 43s
`dotnet format --verify-no-changes` re-analysed the whole solution before the build did, to
reach a verdict the build reaches on its own: IDE0055 is an error in .editorconfig,
EnforceCodeStyleInBuild is on and warnings are errors, so a misformatted file fails the
build step. What the separate step bought was hearing about it a few minutes earlier, and
it charged those minutes on every run.

Checked rather than assumed, because the whole justification rests on it: appending a
badly-spaced member to a source file produces three `error IDE0055` lines and a failed
build with no format step in sight.

Three places said the old arrangement out loud and would now be wrong on their own — the
comment on the IDE0055 line, the conventions list in the README, and a note in
platform-flags telling people to run dotnet format before pushing or CI would fail them.
They say the build enforces it now. dotnet format is still how to fix what the build
complains about; it just no longer gates anything.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 21:43:32 +02:00
jaap-janandClaude Opus 5 ddf0dd6a2b Stop the transfer tests depending on which thread ran them first
Every test in TransferQueueingTests failed on the Linux runner with "The calling thread
cannot access this object because a different thread owns it", and none of them had
anything to do with the commits in that run. The class drained its rows through
Dispatcher.UIThread.RunJobs(). That dispatcher is process-wide and belongs to whichever
thread touched it first, and xunit runs each test class as its own parallel collection — so
the moment a runner scheduled another class onto that thread ahead of this one, all nine
died inside DispatcherOperation.Execute having asserted nothing about transfers at all. It
passes locally and fails on a machine that schedules differently, which is the whole of why
this took a CI run to find.

TransfersViewModel now takes the poster it marshals through, defaulting to
Dispatcher.UIThread.Post — the seam VaultViewModel's clipboard already is, for the same
reason: a view model that reaches a process-wide UI object directly makes every test of it
depend on a thread it does not choose. No head passes the parameter, so nothing about the
running application changes.

The test supplies a queue of its own and drains it, which is the same shape the dispatcher
gave it. A poster that ran the action inline was tried first and is wrong: the transfer
queue raises Changed from its pump thread as well as from the call that enqueued, so inline
execution has a background thread adding rows to an ObservableCollection while the test
reads it — it passed once and then failed a different test on the next run. Draining keeps
every mutation on the thread doing the asserting, which is the one thing the dispatcher was
providing that was worth keeping.

One test added for the property that broke: queueing is reachable from any thread and must
not care which. The class as a whole guards the seam — remove it and nothing drains, so
every assertion about a row fails.

Verified by reproducing the failure first: a throwaway probe that touched the dispatcher on
one thread and posted and drained on another produced exactly the CI message. Then six
consecutive Release runs of the app suite, all green, plus the layout suite, which builds a
TransfersViewModel of its own.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 21:43:21 +02:00
jaap-janandClaude Opus 5 093f3904c1 Apply pending migrations at startup instead of asking for a second command
ci / build and test (push) Failing after 1m48s
ci / api image (push) Skipped
ci / android head (push) Failing after 5s
The API deliberately never migrated: it failed readiness while a migration was pending and
named it, and a separate step applied them. That is the right split for a deployment with a
release pipeline and the wrong one for a self-hosted server, where it means an image that
boots, refuses traffic, and waits for somebody to know that dotnet ef exists. The schema and
the code that expects it ship in the same image, so the image is where the two are
reconciled now.

Before RunAsync rather than in the background. A migration racing the first requests would
let them through against a half-applied schema, and the first authenticated request is the
one that provisions accounts. Failing to migrate therefore fails to start, which is the
loudest signal available and the one an orchestrator already acts on.

Concurrent starts take a Postgres advisory lock first. Without it two replicas rolled out
together read the same empty history table, both apply the same migration, and the second
dies on an object that already exists — a crash loop on the day of a schema change, which
is the worst day to have one. The lock is held on a connection of its own because EF opens
and closes one per command, and a session lock belongs to the connection that took it.

The exception is a database that does not exist yet: there is nothing to hold a lock in, so
that path migrates without one and says so. Two instances creating it at once still
converges — one wins, the other restarts into the ordinary locked path — and refusing to
start would leave a fresh deployment stuck on the step this removes.

Database:AutoMigrate turns it off for the deployments that own their schema: a migrator job,
a rollout where new code must run against the old schema first, or a database user denied
DDL. With it off the behaviour is exactly what it was, and the health check now explains
which of the two situations a pending migration means.

Verified against a throwaway PostgreSQL container: an empty database gets all seven
migrations applied before the port opens, the tables land in the dodo schema, and a second
start logs the schema up to date and serves. The API suite passes, which exercises the
startup path once per assembly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 21:12:22 +02:00
jaap-janandClaude Opus 5 73c7e2a1e3 Point a release build at the hosted server, and a debug build at the clone
The shipped default was http://localhost:5233, which is the address the API serves under
dotnet run and a machine an installed application is not running. Somebody who installs a
release and accepts the field unread is signing in to nothing.

Two defaults now, because the two audiences never overlap. A release build offers
https://ssh.dodotech.cloud, so a first launch needs no address typed at all. A debug build
keeps the loopback address, and that half matters as much: shipping the hosted address into
a clone would point every development launch at production, and sign-in is the call that
provisions an account there.

The remark carries the reason the schemes differ, since the pair now looks like an
oversight rather than the deliberate thing it is — the API's first launch profile is
plaintext on 5233, and an HTTPS client meeting a plaintext port reports a TLS failure that
reads like a certificate problem.

The test spells out both branches rather than asserting the constant, which would pass
however it were edited. What it is really guarding is that a release never ships a
developer's loopback address and a debug build never points a clone at production, and it
can only guard those by naming them.

Verified by running the shell suite in Debug and in Release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 21:12:00 +02:00
jaap-janandClaude Opus 5 39b7e5620f Say which half of signing in is happening, and which half failed
A sign-in against a production Keycloak was reported as the shell hanging on "Opening your
browser to sign in…" and then, some time later, saying "The server returned 500". Both
halves of that are the message's fault. The browser flow had already succeeded — the
provider authenticated the user, the code came back, the tokens were exchanged — and what
was actually happening was a round trip to the DodoSSH server for the account. The screen
went on describing a browser nobody was waiting for.

The status now moves when the browser half ends, so the wait that follows is attributed to
the server being asked rather than to the browser that has already answered.

The failure gets the same treatment. "The server returned 500" is the API client's phrase
for any server it talks to, and read underneath a sign-in button it is naturally taken as
the sign-in having failed — which sends somebody to their identity provider's logs to find
out why a thing that worked did not work. It now says signing in succeeded, names the host
that failed afterwards, and says the reason is in that server's logs, because this side
cannot know more than that.

Nothing here fixes the 500. It changes which of the two servers the next person goes and
looks at, which was the actual cost of the old message.

Verified against the shell suite, including the case that asserts a failed command leaves
the window enabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 21:11:38 +02:00
jaap-janandClaude Opus 5 1db8bed872 Let a cancelled sign-in end the sign-in rather than the timeout
Backing out of the login page on Android left the shell showing "Opening your browser to
sign in…" with the button disabled for five minutes. Nothing was wrong except that nobody
told it: the redirect callback only ever completed when an intent arrived, so a user who
pressed back was waiting on OidcClient's browser timeout to expire before the flow failed
and the button came back.

There is no cancel event to subscribe to on this platform. Pressing back, dismissing the
browser and closing a provider's error page are indistinguishable from here — the browser
goes away and this application is foreground again with nothing delivered — so being
resumed while a sign-in is still waiting is the signal, and the only one there is. The
launcher records that a browser took the intent, OnResume fails the wait, and the guard
means the resumes that have nothing to do with signing in (a launch, recents, the
keystore's fingerprint prompt) go through untouched.

An exception rather than a cancellation, because OidcClient reads a cancelled wait as its
own timeout expiring and would report five minutes passing to somebody who waited two
seconds. It cannot steal a successful sign-in either: Android delivers the redirect to
OnNewIntent before resuming the activity, so the completion is already settled and the
attempt does nothing.

The enrollment key-binding trip through the browser is covered by the same change, since it
waits on the same callback.

The OnNewIntent remark had been sitting above OnResume, describing a method two below it.
Moved back, since the new remark wanted the space and the old one was wrong where it was.

Verified by building the head in Debug and Release. The behaviour itself is unverified for
the reason docs/android-port.md gives about this whole head: nothing has been run on a
device.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 21:10:37 +02:00
jaap-janandClaude Opus 5 08a820adcf Stop the login step interpolating a comment I wrote in it
ci / build and test (push) Successful in 1m40s
ci / android head (push) Failing after 5s
ci / api image (push) Successful in 40s
The secrets were never the problem. The log shows both arriving masked, which is what a
runner does with a value it holds — so the repository secrets were configured correctly the
whole time, and the guidance about Actions Variables was wrong.

What broke was the guard added to diagnose them. Its comment contained an expression
delimiter written out literally to explain what an unset secret renders as, and a shell
comment is not a comment yet at that point: the runner substitutes the whole script before
any shell sees it, so it tried to evaluate an empty expression and failed the step with a
parse error carrying no line number. The step never ran, and push then reached the registry
with nothing to authenticate as — "no basic auth credentials", which looks precisely like
the missing-secret problem the guard was added to rule out.

The comment now describes the delimiter instead of containing one, and warns the next
person, since the failure is invisible to review and to every local check: the file is
valid YAML and the script is valid shell.

Verified with a scan for empty expressions across every run block in the file — one before,
none after — and by running the step's script with credentials set, which passes the guard
and gets a 401 from the real registry. That is the right answer for an invented password,
and it means the endpoint is reachable and the path through this step is sound.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 17:01:06 +02:00
jaap-janandClaude Opus 5 2bc0d4d89f Say which credential is missing instead of letting docker guess
ci / build and test (push) Successful in 1m47s
ci / android head (push) Failing after 5s
ci / api image (push) Failing after 4s
The registry secrets are reportedly not arriving, and the job could not have told anybody
which one or why. An unset secret is not an error anywhere upstream: ${{ }} renders a
missing value as an empty string, so docker gets --username "" and replies with something
about credentials — which reads as the registry rejecting a login rather than as a value
that never left the settings page.

Checked before use now, and reported by length rather than by value. Gitea masks known
secret values in logs, but a mask is only as good as the runner's bookkeeping, and a length
answers the only question actually being asked: did anything arrive at all. The message
names the page to look at, and names the neighbouring one too, since Actions Variables and
Actions Secrets sit next to each other and only one of them is readable through the secrets
context.

This does not fix the credentials. It converts a confusing failure into a specific one, so
the next run distinguishes "the secret is empty here" from "the registry refused these" —
two problems with nothing in common that currently look identical.

Verified by running the step's script with both variables set empty, which is the reported
symptom: it names both, points at the settings page and exits 1 before docker is called.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 16:51:12 +02:00
jaap-janandClaude Opus 5 a515a35804 Build the image with BuildKit rather than the builder Docker is retiring
ci / build and test (push) Successful in 1m38s
ci / android head (push) Failing after 4s
ci / api image (push) Failing after 21s
"DEPRECATED: The legacy builder is deprecated and will be removed in a future release."
Not a failure — the image was built and the job carried on — but a countdown, and one the
last commit walked straight into: Alpine's docker-cli package does not carry buildx, so
giving the job a working client left it building the old way.

Two packages instead of one now. With the plugin present `docker build` routes through
BuildKit on its own, which also stops the Dockerfile's independent stages being serialised,
so this is slightly faster as well as not deprecated.

buildx is wanted rather than required, and the difference is deliberate. Missing it costs a
warning and a slower build; the image is still correct. So each install branch ends in
`|| true` and the check afterwards reports instead of exiting — a distribution with no
package for it should not be able to turn a release into a red build over a plugin.

The comment above the build step said this job needed "no buildx plugin", which was true
when the build was the only thing being weighed and is not true now. It says what is
actually wanted, and what is still not: no QEMU, no builder instance to create and tear
down, no third-party action to re-pin.

Verified in Alpine containers with the socket mounted, in all three states this can be in:
nothing installed, the client present and buildx missing — which is exactly what produced
the warning — and buildx unavailable with no package manager to fix it, which warns and
exits 0. The API image builds through BuildKit with no deprecation notice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 15:19:19 +02:00
jaap-janandClaude Opus 5 5ddbca49d3 Give the image job a docker client to go with the daemon it already had
ci / android head (push) Failing after 5s
ci / build and test (push) Successful in 1m50s
ci / api image (push) Failing after 1m6s
Exit 127, `docker: command not found`, from the build step of the image job. The daemon was
never the problem and never missing: Testcontainers speaks to /var/run/docker.sock from a
.NET library, so every integration suite in the build job had been starting PostgreSQL,
Keycloak and an sshd on this runner while `docker` was not a command on it at all. Having a
socket and having a client are two different things to have, and this runner had one.

It failed late for the same reason it was easy to miss. Node, git, the SDK and the tags all
came up fine, so the job looked healthy right until the line that actually needed the
binary.

The client only. There is a daemon answering on that socket already — installing an engine
would start a second one beside the one in use, which is a worse outcome than the error.

Verified by running this step's own script in an Alpine container with the socket mounted:
it installs docker-cli, the client then reports server 29.6.2 across the socket, and the
API image builds to completion from inside that container with the repository as its
context. Which is as close to the runner as this can be checked without being it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 15:14:28 +02:00
jaap-janandClaude Opus 5 338c1a8647 Move the plaintext exemption out of the test body it made too long
ci / build and test (push) Successful in 1m47s
ci / android head (push) Failing after 4s
ci / api image (push) Failing after 2s
MA0051: TheWholeSlice reached 68 lines against a limit of 60, because the last commit put a
nine-line paragraph and a five-line call in the middle of it. dotnet format --verify-no-changes
runs the analysers, so the build stopped there and never reached the two fixes that paragraph
was explaining.

The explanation was worth keeping and the place was wrong. It is a fact about one call, not
about the slice, and this file already keeps its steps in named methods under a "Steps"
heading. SignInToTheStackAsync now holds both, which leaves the test body reading as the
sequence it is meant to be — sign in, enroll, unlock, write, read elsewhere — rather than a
sequence with an essay in it.

Nothing about the behaviour changed: same call, same configureOidc, same exemption claimed
by the same single caller that starts the provider it is talking to.

I should also say how this reached CI, since the answer is not that it was hard to catch. I
ran the format gate locally before the last push and read the exit code of a pipeline it was
piped into rather than the tool's own, so a failing command reported as passing. Run again
against the tool's exit status it is 0, and the end-to-end suite still passes in the Alpine
container that reproduces the runner.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 14:55:32 +02:00
jaap-janandClaude Opus 5 cf1a321d1e Pin the font the layout suite measures, and let the slice say it means plaintext
ci / build and test (push) Failing after 42s
ci / api image (push) Skipped
ci / android head (push) Failing after 5s
Two failures left on the runner, with nothing in common except that both only appear on a
machine unlike the one anybody develops on. The runner is Alpine, musl, inside a container,
with no fonts installed at all — and that combination is now reproducible locally, which is
how these were fixed rather than guessed at. Both are verified by running the suite in it.

The layout suite had two causes stacked, and the first hid the second completely. Missing
libfontconfig stops libSkiaSharp loading, which the last commit fixed and which then
revealed the real one: Avalonia takes its default font family from the platform, and on an
image with no fonts there is no answer, so FontManager throws "Default font family name
can't be null or empty" inside AppBuilder.SetupUnsafe — before a single test body runs, for
all sixty-eight of them, naming none of their subjects. WithInterFont does not prevent it:
it registers a collection without nominating a default. HeadlessApp's own comment already
claimed it measured "the same Inter font the application registers", which was an intention
the code never carried out.

Both heads now name it, through FontManagerOptions.DefaultFamilyName. That is worth more
than getting CI green: a suite whose entire job is measuring text was taking its metrics
from whatever the machine happened to have — Segoe UI here, DejaVu there — and reporting
the two as one number. It also means the application uses the font it has been shipping and
declining to use since it first referenced the package; almost nothing moves visually,
because App.axaml already sets MonoFont on essentially everything that draws.

The end-to-end slice was the product being right and the test leaning on an accident.
ServerConnection permits an http authority only when it is loopback. Testcontainers reports
the host a container can actually be reached at, so running the suite directly gives
localhost and passes, while running it inside a container gives the bridge gateway
172.17.0.1 and is refused — correctly, since a client that accepted plaintext metadata from
a routable address would be a weakness for everyone who is not a test. Loosening that rule
was the wrong repair. The slice now passes configureOidc and says out loud that it accepts
plaintext from the Keycloak it started itself.

Verified by reproducing the runner rather than approximating it: dotnet/sdk:10.0-alpine,
musl-x64, fc-list returning zero, the docker socket mounted so Testcontainers resolves the
gateway exactly as it does in CI. The whole solution passes there — 19 suites, 0 failures,
4 skipped — and the end-to-end failure was confirmed causal by reverting only that one file
and watching it fail again in the same container. The layout suite also still passes on a
Fedora desktop with 595 fonts, so the two agree now.

Not verified on Windows, and it should be said plainly rather than left to be discovered:
pinning the family changed the measured metrics there too, so a tight layout assertion
could have moved. platform-flags.md records that, and corrects the entry the last commit
added — "libfontconfig, and nothing else" was true of the container it was tested in and
false of the runner.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 14:48:30 +02:00
jaap-janandClaude Opus 5 208aca1191 Make a failing test run say what went wrong
ci / build and test (push) Failing after 1m38s
ci / api image (push) Skipped
ci / android head (push) Failing after 5s
Two suites fail on the runner and pass everywhere else, and every attempt to work out why
has been an inference from a filename. The runner prints the path of a log written to a
disk nobody has a shell on, and the log is where the exception type, the message and the
stack all live — so a red build has been a guess, and the last guess was wrong: 69 layout
failures looked like missing fonts and were a missing shared library instead.

This prints the log, and three facts about the machine that no log will ever carry: which
distribution it is and who the job runs as, whether docker answers, and — the one that
matters for the layout suite — ldd against the libSkiaSharp.so the test project carries,
filtered to its unresolved rows. A managed TypeInitializationException on SKImageInfo is a
symptom several missing libraries share; ldd names the library. The fontconfig step ahead
of this exits early when ldconfig already reports one, so if that is present and Skia still
will not load, the answer is a different dependency and this is what says which.

head rather than tail on the log, which is the whole trick. A suite that fails wholesale
writes one stack per test and they are the same stack; the first explains it and the last
two hundred lines are that sentence repeated.

if: failure() and exit 0, so it runs only on a red build and reports without becoming a
second failure on top of the first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 14:32:32 +02:00
jaap-janandClaude Opus 5 43d76d0f2d Let the suite run on Linux, and fix the three things that stopped it
ci / build and test (push) Failing after 1m47s
ci / api image (push) Skipped
ci / android head (push) Failing after 5s
The pipeline finally reached the tests and found four failures. None was the pipeline's,
and only one of the four was a test being fussy about a platform rather than telling the
truth about one.

The local pane's roots bar was the real bug. LocalDirectory.Roots built it from
DriveInfo.GetDrives on every platform, and its own summary — "the drives on Windows, and
the root elsewhere" — had been describing an intention rather than the code for as long as
nobody ran it off Windows. On Unix that call answers with every mount the kernel holds:
/proc, /sys/fs/bpf, one per installed snap, /run/user/1000/doc, some forty on an ordinary
laptop. The transfers screen draws a button per root, so the bar ran to about five thousand
pixels inside an eight-hundred pixel window. Anybody running the Linux build has been
looking at that.

Filtering GetDrives is not the fix and the comment now says why at length, because it is
the obvious thing to try: DriveType answers Fixed for / and /home and equally for every
squashfs snap, for efivarfs and for tracefs, while /boot/efi comes back Removable, and
DriveFormat would need a hand-kept list of every virtual filesystem Linux might grow. So
Unix now names what somebody would want instead of subtracting what they would not — the
root, their home, and whatever is mounted under /run/media/<user>, /media, /mnt or
/Volumes. Anything else is still reachable by navigating from /, which is what the pane is
for. Windows is untouched.

ClientPathsTests looked for "odoSSH" in the profile directory. ClientPaths spells it
DodoSSH on Windows and dodossh on Unix deliberately, one per platform convention, and that
substring was clever enough to survive either spelling of the leading D while still only
ever matching one of them. Now OrdinalIgnoreCase.

WhyTheWindowItselfIsNeverShown asserted a COMException with HResult RPC_E_CHANGED_MODE,
which is WebView2 refusing an MTA thread — a Win32 component raising a COM error. On Linux
the adapter is a different implementation with no apartment to disagree about, so showing
the window works and Should.Throw catches nothing. Skipped there rather than loosened to
accept both outcomes: the assertion is the documentation in that test, and one that passed
everywhere would have stopped recording the constraint it exists to record.

The fourth was CI's alone, and the diagnosis is the useful part. All 69 layout tests failed
on the runner while 6 failed here, which looked like missing fonts and was not: Avalonia's
headless renderer is Skia, libSkiaSharp.so links against libfontconfig, and without it the
suite dies in HeadlessUnitTestSession with a TypeInitializationException on SKImageInfo
naming none of its actual subjects. The job installs the one library now. Verified in a
container where fc-list returns zero and the suite passes regardless, because the
application carries Inter itself — fonts were never the problem, only the thing that would
have looked for them.

The whole solution now passes on Linux: 19 suites, 1295 tests, 0 failures, 4 skipped, the
end-to-end Testcontainers suite included. README and platform-flags.md said testing was
Windows-only, which CI now contradicts on every push, so both say what is true instead and
the two findings are written down where the next person will look for them. macOS is still
untested and now says so on its own rather than hiding inside "not Windows".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 14:28:06 +02:00
jaap-janandClaude Opus 5 71c0bd8882 Ask global.json for an SDK version that exists
ci / build and test (push) Failing after 2m27s
ci / api image (push) Skipped
ci / android head (push) Failing after 52s
setup-dotnet refused the file outright: "Version '10.0.0' is not valid for the 'sdk.version'
value in global.json. When 'rollForward' is specified, a full SDK version is required."

It is right, and the mistake is a category one rather than a typo. 10.0.0 is a runtime
version; SDK versions carry a feature band, so the first SDK of this major is 10.0.100 and
there has never been a 10.0.0 to roll forward from. The local dotnet accepted it because it
resolves a floor loosely, which is exactly why this survived to CI — nothing on a developer
machine ever disagreed with it.

10.0.100 with the same latestMinor keeps what the file meant: any 10.x SDK, newest wins.
Verified against both SDKs in play, 10.0.109 locally and 10.0.302 in the build container.

Left floating rather than pinned, and worth being honest that this is the shakier half.
IsTrimmable on Contracts and Crypto pulls in Microsoft.NET.ILLink.Tasks, whose version
tracks the SDK's patch and is therefore written into packages.lock.json — so the day a
newer 10.x SDK appears on the runner, --locked-mode fails until the lock files are
regenerated against it. Pinning an exact version with rollForward disabled would end that,
at the cost of everyone installing that SDK exactly; it is a real choice and not one to
make silently inside a fix for something else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 14:07:30 +02:00
jaap-janandClaude Opus 5 a6af93148b Give the runner a node before asking it to run an action
ci / build and test (push) Failing after 5s
ci / api image (push) Skipped
ci / android head (push) Failing after 3s
Every job died on its first line: "Cannot find: node in PATH", from actions/checkout.
act_runner executes each `uses:` action with node inside the job container, and the image
this runner is configured with has none — so nothing in the pipeline had run yet, including
the tests the image job gates on.

A `run:` step is shell rather than node, so one placed ahead of the first action can fix the
job it is in. It installs via apt-get, apk or dnf, whichever is there, and says plainly what
to do when none of them is. git goes in alongside, named in the step rather than smuggled
into it: checkout shells out to git the moment node has loaded it, so an image thin enough
to lack one usually lacks the other, and finding that out separately costs another round
trip through CI.

The version is warned about, not enforced. Distributions pin nodejs to whatever shipped
with the release — Ubuntu 24.04 still serves 18, past end of life and older than these
actions declare — but act_runner hands an action whichever node is on PATH regardless of
what it asked for, and it generally works. A warning is the right weight for something that
explains a later inexplicable failure without being one.

Repeated verbatim in all three jobs. It cannot be a local composite action, since that
needs the checkout it exists to unblock, and YAML anchors that would deduplicate it are
rejected by GitHub's parser. Byte-identical across the three so a diff shows drift.

This is still a workaround. The fix is one line of the runner's own config.yaml pointing
container.image at an image that ships node, as Gitea's default
catthehacker/ubuntu:act-latest does; the step then costs a version check and nothing else.
Kept regardless, because a pipeline that silently depends on a runner being configured
correctly elsewhere fails confusingly when it is not.

Verified by running the step's own script in ubuntu:24.04 and alpine:3.20, which have
neither, and node:20-bookworm, which has both: installs where needed, no-ops where not, and
warns only on the node 18 that Ubuntu gives.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 14:04:50 +02:00
jaap-janandClaude Opus 5 57d4b30557 Put the app's own mark on the launcher
ci / build and test (push) Failing after 3s
ci / api image (push) Skipped
ci / android head (push) Failing after 2s
The sign-in and locked screens both draw the same thing — a square outline in the accent
with >_ inside it — and the launcher was still showing the stock Android silhouette, so the
icon somebody taps and the icon the app opens onto had nothing to do with each other.

Redrawn as a vector rather than exported from the screen as a bitmap. There is one geometry
here and no set of density buckets to update four of and forget the fifth, and the accent
stays a number that can be diffed against Palette.axaml rather than a colour baked into a
PNG. The hex is written out because an Android resource cannot reference a XAML dictionary
— the same duplication colors.xml already carries for the window background, with the same
obligation attached.

Adaptive only, no raster fallback. Adaptive icons landed in API 26 and this head requires
28, so there is no device it ships to that would need the bitmaps; density buckets exist to
choose between PNGs and there is nothing to choose. The background layer is the same
@color/dodo_window as the window and the status bar, so the mark sits on the app's own
near-black rather than on a second one almost like it.

Two departures from the screen, both because a launcher is looked at much smaller than a
sign-in header. The box is 42 across rather than the 48 that first suggested itself: 72 of
the 108 survives masking, but that is a width, and a square meets a circular mask at its
corners — at 48 they land 33.9 out against a radius of 36 and read as clipped despite
technically clearing it. And the strokes are 2.2 and 2.8 where proportional fidelity to a
1px border on 44px would be 1.0, which a launcher drawing this at 48dp would render as half
a pixel of nothing.

A monochrome layer too, for the Android 13+ themed-icon setting. Without one a launcher
with themed icons on falls back to the full-colour icon, which would leave this the single
green thing on an otherwise recoloured home screen.

Verified in the packaged APK: the icon resolves at all five densities, the three layers
resolve, and the compiled vector carries the geometry above. The launcher rendering itself
was checked against local renders under circular and squircle masks at 144 and 64 px, not
on the device — the phone was locked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 13:46:40 +02:00
jaap-janandClaude Opus 5 8a568117df Give the API an image, and unbreak the restore that had to run first
registry-docker.dodotech.cloud/dodotech/dodossh-api, built and pushed by a third ci job
that needs the first. Gating on the tests costs a few minutes on every main commit and buys
the only thing worth having here: an image is not an artefact somebody inspects before
using it, so a red commit must not be able to produce one. Pull requests build the image
and stop, which is where a broken Dockerfile should be found.

Tags are :sha-<short> on every build, :main on main, and for a v* tag :1.2.3, :1.2 and
:latest — the last two only when the version has no prerelease suffix, since v1.3.0-rc1
sorts above v1.2.9 and would otherwise walk :latest onto somebody's server. Only sha- is
immutable, and it is the one to pin a deployment to.

No docker/* actions. The build is single-architecture, so it needs the daemon this runner
already has for the Testcontainers suites and nothing else — no buildx, no QEMU, and no
third-party action whose SHA has to be audited and re-pinned. Step outputs and secrets
reach the shell through env rather than ${{ }} interpolation, because a git tag may contain
a semicolon and interpolation is textual substitution performed before the shell parses the
line.

The image is chiseled: no shell, no package manager, uid 1654. Affordable because
Directory.Build.props already sets InvariantGlobalization, so the ICU and tzdata a normal
base carries are exactly what this product decided not to use. The cost is stated in the
Dockerfile rather than hidden — there is no HEALTHCHECK, because there is nothing to run
one with, and /healthz/ready is anonymous precisely so the orchestrator can ask instead.
Nothing migrates the schema from inside the container either; readiness fails while a
migration is pending and names it, which is the design.

And the restore that all of this depends on did not work. 7a3a521 committed lock files
carrying a net10.0/android-arm64 section into fourteen projects — written there by the
Android head's -p:RuntimeIdentifier=android-arm64 packaging build, which restores the
shared projects with a RID and updates their lock files as a side effect. Any restore
without that RID then fails NU1004 in locked mode, which is every other build there is:
`dotnet restore DodoSSH.slnx --locked-mode` has been failing for eleven projects on a clean
checkout of main since that commit. The sections are removed here and nothing else changed
— deletions only, ILLink.Tasks stays at 10.0.10.

Verified: the solution restores in locked mode, the image builds, and it runs. /healthz/live
answers 200 and /healthz/ready answers 503 naming the database it cannot reach, from a
67 MB image as uid 1654, configured entirely through DODOSSH_-prefixed variables.

The Android head's own lock file still carries the RID and is untouched, because that job
restores it separately and is outside DodoSSH.slnx. Whether its packaging step re-dirties
these fourteen on every CI run is worth a look; it is the same mechanism.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 13:40:12 +02:00
jaap-janandClaude Opus 5 215e73b07f Let the phone's theme past the activity it is attached to
The head has now run on a device, and the first thing it did was die on the way up.
DodoTheme parented @android:style/Theme.Material.NoActionBar, but AvaloniaMainActivity
descends from AndroidX's AppCompatActivity, which asserts its own theme attributes while
inflating and throws — "You need to use a Theme.AppCompat theme (or descendant)" — before
a single Avalonia frame exists. The platform's own parents are the ones that look right,
which is why the audit read as correct and the launcher icon still opened onto a splash
screen and then nothing.

Theme.AppCompat.NoActionBar instead, dark rather than .Light because every override below
it repaints the window near-black regardless. The no-action-bar and status-bar decisions
those overrides carry are untouched, so the reason they are there — a header that has to
hold the vault name, and a clock that would otherwise be dark-on-dark — still holds.

Verified on a OnePlus CPH2765: builds, deploys, and reaches the sign-in screen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 13:16:56 +02:00
jaap-jan 4300d917a8 Stop making people wait for a handshake, and give the host list a pointer
ci / build and test (push) Failing after 3s
ci / android head (push) Failing after 2s
Connecting held the vault's busy gate, which meant a window that did nothing visible for
as long as a machine took to answer — and against one that is merely asleep, that is the
whole timeout. The gate is gone from that one command. A tab now appears in the strip in
the same turn as the click, carrying "connecting…" rather than a pane, and the terminal's
rectangle draws a card naming the host and the address being dialled. Every other screen
stays usable, and two connections can be in flight at once.

That splits the vault's one connection event into three, carrying an attempt id, because
"which tab is this about" can no longer be answered by "the most recent one". The id also
buys the two kinds of not-connecting their different endings: a refusal stays in the strip
as a tab holding its reason, since by then the user is quite likely three screens away and
a status line they are not looking at is not where a failure should end; a host key
question takes the tab away and puts the window back on HOSTS, because the prompt is drawn
there and a tab claiming failure would be competing with the thing about to resume it.

ConnectAsync takes no CancellationToken any more, and that is load-bearing rather than
tidying. A [RelayCommand] over a method that takes one generates a command that cancels
the previous execution's token on every invocation — so asking for a second machine
silently abandoned the first, measured as the first tab disappearing with "Cancelled." the
instant the second was asked for. Giving up on a connection is closing its tab, and a
session that lands after that is adopted rather than dropped: a shell running with nothing
naming it cannot be closed at all.

A tab is marked active on IsShowing rather than IsSelected. The selection survives
navigating away — that is what makes the strip a way back to a terminal instead of a way
to lose one — so a tab lit while preferences filled the window was a second "you are here"
mark pointing at something nobody could see. The nav rail's own entries have always made
this distinction.

The host list grows the two gestures it looked like it already had. A right click selects
the row under the pointer before opening a menu of Connect, Edit and Delete — the menu is
on the list rather than in the item template, so its entries are the vault's own commands
and not a row's, and it is cancelled outright over a group heading. Dragging a host onto a
heading files it there, onto a host files it beside that one, and onto UNGROUPED takes it
out of a group; the write is one field of one host through the same repository a save
uses, refused while the editor is open because a drop is a gesture on the list and not on
a half-typed form.

Clicking a result in the palette connects, which is what a list of hosts under a search
box looks like it does. It went through the shell's own command, so the pointer and Enter
take one path.

And the files screen's two pickers followed the vault's lists once, at unlock: a host or a
bucket created afterwards could not be picked until the keychain had been locked and
opened again, with nothing on screen explaining why the machine plainly in the host list
was missing. They follow the collections now, re-finding the selection by id across the
rebuild a sync pass causes every minute.

165 shell tests and 69 layout tests green, including the connecting tab, both failure
endings, two connections at once, a connection in flight across a lock, and the right
click acting on the row under the pointer rather than on the selection. The drag itself is
in docs/manual-checks.md with the rest of phase 7 — headless Avalonia has no platform
drag, and a test that claimed to have dropped something would pass while confirming
nothing.
2026-07-31 22:59:33 +02:00
jaap-jan 7a3a521c59 Give the phone the rest of its screens, and a way in
ci / build and test (push) Failing after 2s
ci / android head (push) Failing after 1s
All seven screens of the design, plus the two it does not draw because it starts at an
enrolled phone: naming a server, and choosing a passphrase.

The five states docs/android-port.md worried about losing at 360dp are all here and none
of them softened. The changed-key refusal is a full-screen panel rather than a bottom
sheet, because a sheet is swipe-to-dismiss by convention and that screen must have no way
forward. The recovery code raises FLAG_SECURE for its own state and lowers it afterwards,
so the sentence about screenshots is true rather than decorative. The delete
confirmations keep their counts and replace the row in place.

Signing in works, and the seam it needed is worth more than the implementation:
IAuthorizationCallback now sits between OidcClient and the loopback listener, so the two
heads differ in where the response arrives and in nothing else. PKCE, the state check,
discovery, the token exchange and the key binding stay one implementation — a second OIDC
client would be a second place for a security bug to live. The phone registers a
private-use scheme with the system rather than binding a loopback port, which on a shared
device any other app can do first.

The accessory key row needed TerminalWorkspace.SendInputAsync: ordinary typing goes from
the renderer straight down the socket, and there was no way in for the keys a software
keyboard does not have. Ctrl latches, because one thumb cannot chord, and the latch is
drawn — a modifier that is on and does not look on is how somebody sends ^L to a database
prompt believing they typed an l.

597 client tests green, including two new ones for the input path and one for the
terminal surface command. Nothing has run on a device.
2026-07-31 21:43:11 +02:00
jaap-jan 81e7e6d939 Write down what the phone found, and stop it rotting
docs/android-port.md was an audit of work not started; it now says what is built. Three of
its statements needed correcting rather than extending, and they are marked where they sit:
the Android version question is settled and was never as open as it looked, because
Avalonia.Controls.WebView ships only a net10.0-android36.0 assembly and nothing lower can
resolve it; cleartext to loopback has to be permitted explicitly, which the audit missed
entirely; and the spike produced a structural change it did not anticipate, in
DodoSSH.Client.Shell.

A CI job of its own, because the head is deliberately not in DodoSSH.slnx and a project
outside the solution is a project nobody notices breaking. It packages as well as builds:
a native library with no Android ABI and an assembly that will not dex are both invisible
to a compile, and both are exactly what this head is exposed to.

The README says plainly that signing in is not built, that a fingerprint re-enrolment
destroys the device key, that a notification appears while a shell is open, and that none
of it has run on a device.
2026-07-31 21:09:42 +02:00
jaap-jan 2caedd93ff Merge branch 'main' into the Android head
Main grew the screens the host-management plan called for — hosts, pins, snippets, logs,
import, teams — plus the ObjectStore and Import projects behind two of them, and moved
WindowsDeviceKeyStore into the desktop head's Platform folder.

Five of those view models landed in a directory this branch had already moved, so they
join the rest in DodoSSH.Client.Shell: git spotted the rename and put them there, and the
namespaces followed. Shell picks up ObjectStore and Import as a result, which the Android
head then gets transitively and will use neither of at first — scoped storage means there
is no ~/.ssh/config to import, and file transfer is out of its first scope.

Desktop suites green at 155 and 64.
2026-07-31 21:03:22 +02:00
jaap-jan fe9d7fc289 Give DodoSSH a phone, and a shared shell for both heads to drive
The Android head from docs/android-port.md, taken as far as its step 6.

Step 3, the spike, is answered and its throwaway screen is gone: libsodium.so and
libe_sqlite3.so are both in the arm64 APK, so NSec resolves its native half on Android
despite shipping no Android build, and the local cache opens. Two findings the audit
could not have had: Avalonia.Controls.WebView only ships net10.0-android36.0, which
settles the open "which Android versions" question at targetSdk 36; and Android has
blocked cleartext HTTP since API 28, so the terminal renderer needs a network security
config scoped to 127.0.0.1 or the WebView loads nothing.

DodoSSH.Client.Shell is new and is why the phone can exist: the view models, the terminal
renderer files and the palette moved there so both heads drive one state machine and draw
from one set of tokens. The desktop head is otherwise untouched and its 144 tests still
pass.

The platform pieces behind interfaces that already existed: the profile directory from
filesDir, a device key wrapped by a StrongBox-backed key that a fingerprint releases, and
a foreground service so a shell outliving a vault lock stays true on a platform that
stops backgrounded processes.

Sign-in is deliberately absent rather than approximated. It needs an app link, because
reusing the desktop loopback listener is the attack RFC 8252 section 8.3 names.
2026-07-31 20:58:48 +02:00
jaap-jan 5cbda59a34 Merge branch 'main' into claude/host-management-ui-plan-7f20ab
Seven files needed a hand. Most were two branches adding something in the same
place, but three were one branch changing what the other had moved or renamed,
and those are the ones worth reading.

The shell keeps both new fields and both constructor lines: the connection
recorder this branch built and the teams view model main did. Where main put a
teams load inside OnScreenChanged, it now sits beside the logs refresh rather
than inside RaiseSurfaceState — this branch extracted that notification block
and it is called from two properties, so a screen-specific side effect in there
would fire on every terminal switch as well.

Main gave four row types a vault id and a vault name, and this branch had moved
one of them — KnownHostRowViewModel — into its own file when the pinned keys
became a screen. Git resolved that as "deleted here, modified there" and took
the delete, which compiles as long as nobody looks: the moved copy still had
the two-argument constructor and the call site had grown to four. Carried over
by hand, along with the ordering the pins list now does on them.

The status line's quiet rule was the subtle one. Main extracted it into
IsWorthReporting; this branch had changed the same condition to read item
counts rather than raw ones, because every user action queues a log entry a
moment later and this machine reads its own entries back on the next pull. Take
main's structure and the merge builds, passes, and silently restores a bug this
branch existed partly to fix — every save's message overwritten a second after
it appears. The method now reads PulledItems and PushedItems, with the reason
in its remarks.

Two conflicts were prose that had gone stale rather than code. The keychain
screen's comment said team vaults are refused by the server's access service,
which was true when it was written and is not now; main's replacement stands,
in this branch's vocabulary. The design-gaps row for groups was claimed by both
— real host groups here, per-vault headings there — and they are different
things, so both rows stay and the difference is stated: a group is a shelf the
user chose, a vault is who can read the item.

One defect the tests found and the compiler could not. Generating a key opens
the same editor as pasting one, but not through NewKey — so it never set the
target vault main added, and a generated key was filed into whatever vault was
edited last, or none. Both key-generation tests failed on it. Fixed where the
editor opens, with the reason recorded there.

One gap is left deliberately and is written down rather than half-built. Hosts,
keys, credentials and pins are read across every vault this session holds a key
for; groups are read from the active vault alone, so a host a teammate filed
shows under UNGROUPED. Nothing is lost or misfiled — it is what the sidebar
already shows for a group that has been deleted — but closing it needs a vault
id on every group row for rename and delete, and a way to tell two vaults'
identically-named groups apart under a layout with one heading per group. Both
are worth doing and neither is a merge's business. It is in the remarks on
ReloadGroupsAsync and in docs/design-import-gaps.md.

dotnet build, dotnet test and dotnet format --verify-no-changes are all clean:
1282 tests, including the end-to-end suite against real containers.
2026-07-31 20:44:39 +02:00
jaap-jan d07b336868 Free the terminal from the Hosts screen, and fill the room it left
The WebView sat inside the Hosts grid, so navigating to Files or the keychain
hid every open terminal and the strip that named them. A connection you had
opened was invisible from four of the five screens. The window now has two
surfaces rather than one: a nav rail that says which page you are on, and a
terminal strip that is always there and switches the whole content area to a
shell. Screen keeps meaning "which page" and never becomes a sixth kind of
page, which is why this is two properties instead of one enum with a terminal
member in it.

Every screen lives inside one wrapper panel that collapses when a terminal is
showing. That is not tidiness — the WebView hosts a Win32 child window that
composites above everything Avalonia draws, so a screen left visible over its
rectangle is a screen sliced in half, and this window has shipped that defect
once already. One decision point, IsTerminalShowing, and a nested panel rather
than five compound bindings nobody would remember to extend.

The focus choreography is the part no test in this repo can see. Every reveal
path now focuses in the same turn the WebView appeared, so all three of them
post at DispatcherPriority.Loaded and let the native control re-push its bounds
first. Going the other way had a real bug: the screen-changed branch called a
bare Focus() where it had to release the keyboard from the native child, so
switching from a terminal to Files silently ate the first keystrokes. Rare
before this commit and the primary gesture after it.

The tab strip grew a cross inside each tab, a plus that opens the quick-connect
palette, and middle-click close. Nested buttons are correct here: Avalonia
handles a left press on the cross and deliberately does not handle other
buttons, which is exactly what lets middle-click bubble up from the cross as
well as the tab. The test is PointerUpdateKind rather than
IsMiddleButtonPressed, because the latter reports button state and is also true
for a left press made while the middle button happens to be held. The handler
is on the tab and not the strip, so the background closes nothing by
construction. Plus opens the palette rather than a flyout, since a menu
dropping into the WebView's rectangle may or may not composite above a child
HWND and this repo does not make rendering claims it has not photographed.

Everything a user reads now says keychain. The wire, the database and the
cryptographic spec still say vault, deliberately: renaming those is a migration
and a protocol change for a word. That split is written down rather than left
to be rediscovered as an inconsistency.

Four things that were squeezed into the keychain's category rail, or into
nothing at all, now have screens. Pinned host keys get one, with fingerprints
never truncated and a filter that matches them, because comparing what you have
against what the operator published is the whole workflow; the approved date is
read out of the item's UUIDv7 rather than added as a column, and says so, since
it means first approval and not last use. Keys can be generated in the client,
which needed the openssh-key-v1 container written by hand — there is no BCL or
NSec helper, and the PKCS#8 route is unverified in the SSH library this uses.
The armour carries no passphrase: encrypting it needs bcrypt_pbkdf, which is
Blowfish with a swizzle, in a project whose crypto is otherwise entirely
libsodium, for a protection the key's own remarks argue is redundant inside a
vault. Generation fills the existing editor and stops, so SAVE stays the one
thing that writes. ~/.ssh/config can be imported behind a preview that is
ticked per row and writes nothing until the button; IdentityFile records the
path and imports the key material only on an explicit opt-in, because reading
somebody's private key into a vault is precisely the act this product exists to
make deliberate. Match blocks and ProxyJump are reported rather than obeyed —
one cannot be evaluated statically and the other has nothing behind it to route
with, and a preview that implied otherwise would be worse than one that admits
it.

Files can be dragged in all four directions that are honestly available. Remote
to Explorer does not ship and is not pretended to: the shell wants the bytes
during the drop, which needs a virtual file and a native COM data object,
outside what Avalonia offers. Note for the next person that Avalonia 12
replaced the drag model outright — DataObject and DataFormats are no-op stubs
and IDataObject is not in the reference assembly, so every tutorial written for
11 does not compile here.

Hosts can be grouped, flat and never nested. A parent id merged as a scalar
lets two offline clients each re-parent A under B and B under A, producing a
cycle inside an encrypted payload that no server can police and every reader
would have to detect for ever. Membership lives in that payload rather than in
the one plaintext concession ADR 0001 allows, whose test is that the relay
cannot function without it — nothing on the server reads a group, so what
plaintext would hand over is a clustering of the estate for nothing. The
plaintext column reserved for it is dropped, provably always null, and the
server now refuses a client that sends one; it was never populated, was copied
on apply, and was not cleared on delete, so a group id would have outlived the
host it described.

Snippets insert through xterm rather than through the pump, because xterm is
the only thing that knows whether the remote has bracketed paste on, and that
is what makes a shell treat embedded newlines as text instead of as execute.
The host process moves opaque bytes and never parses output, so it would have
to guess, and guessing wrong runs every line. Running is off by default and the
copy says the text goes into whatever is there — the terminal has no notion of
being at a prompt, and may be in vi or at a password prompt with echo off, so
the Enter the user presses themselves is the entire safety property.

Connections and keychain changes are recorded as synced encrypted items, which
is what makes them auditable by a team later and costs the server knowledge of
connection rate and timing from row counts alone. ADR 0001 already concedes it
cannot hide that class of metadata; the trade is now written into it rather
than left implicit. A connection entry is written once, at close, which is what
makes a synced log tractable: nothing to merge, one outbox row, no chance of
colliding with itself. Live sessions come from memory, not from the log. The
write is void by contract and posts to a bounded channel, because putting an
encrypt-and-write on the teardown path of every session is how closing the
application comes to take four seconds. A ticket opened before a lock still
closes afterwards, since a shell outlives the vault. The activity log hooks the
one generic repository every kind writes through, so it cannot miss a caller —
which is also why the log kinds themselves declare they are not audited, or the
first entry would write an entry about writing an entry. It records the names
of the fields that changed and never their values; a log with an old password
in it would be a plaintext credential store with no vault around it. Retention
is 90 days or 5,000 entries, whichever bites first, pruned on the sync loop
rather than on a second timer.

That log traffic then broke the status line, which is worth recording because
the fix is a shape and not a patch: background sync counted its own log rows as
pushed items, so the quiet rule stopped being quiet and every action's message
was overwritten a second later by a sync report. The report now separates log
rows from user items and the rule reads the latter.

S3 buckets appear as a remote in the file browser, behind the same interface an
SFTP session implements, so the queue and both panes did not have to learn what
they are talking to. Uploads go through a pipe, because the queue wants to
write and the SDK wants to read; memory is then bounded by the part size
instead of buffering a file to disk twice.

Finally, the Windows device key store moved out of the session project, which
was the one thing keeping it from being portable — everything else in it is
platform-neutral, and a Windows CNG dependency in the middle of the vault code
meant a second head could not reference it without dragging Windows along. The
seam that made the move free was already there. docs/android-port.md is the
audit behind that: what ports, what does not, in order of cost, the four
decisions taken, and an inventory of every screen and state the interface has
to carry, written so a design can be made from it directly.

dotnet build, dotnet test and dotnet format --verify-no-changes are all clean:
1240 tests at zero warnings, including the end-to-end suite against real
containers. The manual checks that headless Avalonia cannot make — the drag
from Explorer, a generated key against a real host, twelve tabs at the minimum
window width — are listed in docs/manual-checks.md and are still outstanding.
2026-07-31 20:30:05 +02:00
jaap-jan 23eca3a21b Merge branch 'main' into claude/m3-implementation-57f9d7
ci / build and test (push) Failing after 2s
Three files conflicted, and two of the resolutions are more than a choice of
side.

QuickConnectTests had both branches fixing the same build break — main's M2
merge left the shell's constructor with an ISftpSessionFactory nobody passed.
Main's version wins because it carries a comment saying why the palette never
needs a session.

VaultSession's conflict is adjacent edits: main added the remembered sign-in
members and this branch changed SyncAsync's summary from "the active vault" to
"one vault". Both kept.

VaultViewModel is the one that matters. Main taught the background pass to
report a sync that had to start over, on the grounds that a machine which
silently re-read a whole vault has had something happen to it; this branch
turned a pass into one report per readable vault. Taking either side alone
would have lost the other, so ResyncedFromStart is now one of the conditions
IsWorthReporting checks, per vault.

Merging also broke something neither branch could have caught alone, and the
build would not have said a word. SyncOnceAsync cleared LastSyncFailed
unconditionally, which was right while a pass was one vault and a failure was
an exception that never reached that line. A failure is now a report — one
unreachable team vault must not stop the others syncing — so the flag was being
cleared over a vault that had just failed, lighting the titlebar SYNCED. It is
computed from the report instead, in the one place both callers go through, so
the manual command gets it as well as the loop. The background pass still
swallows the message and keeps the fact, which is what
AnAutomaticPassThatFails_LeavesTheStatusAlone is there to hold it to.

Two comments the auto-merge left describing a world with one vault in it: the
SCOPES rail's, which said team vaults are refused by the access service, and
the host sidebar's "One heading, for one vault".
2026-07-31 12:26:59 +02:00
jaap-jan 95816de0c5 Share a vault with a team, without the server holding a key
M3's teams, sharing and ACLs. Teams with roles, a public-key directory, the
append-only key log served for clients to check it against, team-owned vaults,
and vault key grants wrapped by a client and stored opaquely by the server.
VaultAccessService resolves team membership to PermissionFlags, so a viewer may
pull and may not push; the desktop client reads and syncs every vault it holds
a key for, and a real TEAMS screen replaces the one that said it did not exist.
No migration: team, team_membership, vault.team_id and vault_key_grant have all
been there since the first one, which is what carrying two unused tables bought.

Membership is authorisation. A grant is access. The obvious model is one
concept — "access", with a role attached, handed out by the server — and this
architecture cannot implement it: a vault key is sealed to each member's X25519
key, and only a client holding the plaintext can seal it for somebody else. So
"give Bob access" decomposes into a database write and a wrap, which happen on
different machines. Adding a member makes the server serve them the vault; it
cannot make it readable. VaultSummary.WrappedVaultKey is null in the meantime
and the vault appears in their list saying it is waiting for a key, because
hiding it until a grant existed would have been tidier and would have implied
the server was the thing granting access. The screen says the same thing after
every add, in the status line. ADR 0009 records the whole decision.

Sharing verifies or refuses. A directory lookup is a claim by the server about
a third party's public key, and wrapping to an unverified claim hands the vault
to whoever made it — no amount of transport security helps, because the server
is inside the threat model. KeyLogAudit reads the whole log, recomputes every
entry's hash from its own contents, checks the chain from genesis, and refuses
unless the offered key appears in it unchanged. There is no override flag: one
that exists gets used on the day the log is briefly unreachable, and the
resulting grant is indistinguishable from a correct one afterwards. What it
still cannot promise is that the key is the right person's, so the fingerprint
comes back for an out-of-band comparison and the success message says so every
time. A test corrupts the fake server's log by one byte and watches the client
refuse rather than warn.

The roles are only the ones that are enforceable. There is no ConnectOnly,
despite the design asking for one and TeamRole having room: SSH terminates on
the client, so a session needs the credential's plaintext on that machine, and
"may connect but may not read the key" cannot be enforced here. Shipping it as
an option in a dropdown would have been a lie. Connect rides along with Read
and is documented as an interface hint. Removal is named for what it does — it
revokes grants and flags the vault for rekey, and claims nothing about what is
already on somebody's laptop.

Three things are deliberately absent, and each is a refusal rather than an
omission. The rekey itself, because re-wrapping every item's data key under a
new vault key needs a client holding the current one; the server records that a
rotation is owed and the interface reports it, which is more honest than a
button that only appears to do it. Ownership transfer, because allowing an
owner to be removed without one leaves a team nobody can administer. And
cross-vault host key trust: a pin in a team vault is listed but not consulted
at connect time, because any member with Write could otherwise pre-approve a
fingerprint another member's client then trusts silently for a host in their
own vault. Scoping trust properly needs a scope on the SSH connect path, which
IKnownHostStore has not got; until then the narrow direction is the safe one
and the cost is in the README rather than hidden.

Reading now spans vaults and writing still does not. Every list on the vault
and hosts screens covers each vault the keyring opened, rows carry the vault
they came from, and an edit goes back to that vault rather than to the active
one — writing it to the active vault would fork the item and only show up when
a colleague wondered why their change never arrived. A new item goes wherever a
picker says, defaulting to the personal vault and never moving on its own,
because an item filed into a team's vault is visible to that team and moving it
back means deleting and retyping. The sidebar heading stops naming one vault
once there are two, and each row names its own.

The server checks what it can and nothing it cannot. It will not record a grant
for a key its recipient no longer holds, for a superseded generation, or for
somebody who is not in the team — each of those would otherwise surface days
later at the far end as a tag failure indistinguishable from corruption. It
does not verify the wrap or the signature, and the grant service says so: that
would be a convenience and never the boundary, and would put an asymmetric
implementation on a machine that is supposed to hold no keys.

Two bugs the tests found. TeamsViewModel's busy gate blocked its own reload, so
a team created a moment earlier was missing from the list it had just been
added to. And syncing every vault turned a failure from an exception into a
report, which made a background pass announce an unreachable vault once a
minute — the exact behaviour AnAutomaticPassThatFails_LeavesTheStatusAlone
exists to prevent. The fact is recorded and the message swallowed, as it was
before; pressing Sync still names the vault and the reason.

Also fixes a build break this branch started with: QuickConnectTests was never
updated when M2 added ISftpSessionFactory to the shell's constructor, so
nothing built at all.
2026-07-31 12:18:28 +02:00
jaap-jan 03e902a2d2 Colour the host's file rows by what their mode says
ci / build and test (push) Failing after 2s
The remote pane's NAME column was blue for a directory and plain for everything
else, and the PERMS column was faint whatever it said. Two colours now come off
the mode, split across those two columns on purpose: NAME says what a row is,
so a file with an execute bit is green there, and PERMS says what is notable
about how it is set, so a file anyone may write to is amber over the characters
that actually say so. Because the two never compete for one TextBlock, a
world-writable executable shows both facts instead of one winning an argument.
No new blue is spent, which is what App.axaml asks for: it reserves blue for a
directory, a distinct scope, and calls it deliberately rare.

Both are files only, and each exclusion is a wrong answer avoided rather than a
case not got to. Every symbolic link is lrwxrwxrwx by convention and its mode
governs nothing — what may be written is the target, whose mode an lstat
listing never fetched — so amber there would fire on every link on the host. A
world-writable directory is /tmp, made safe by a sticky bit PosixMode does not
render, and warning about it would be warning about the half of the mode that
is on screen while the half that answers the warning is not. And the execute
bit on a directory means "may be searched", which is true of very nearly every
directory a host has, so green there would paint the whole pane and mark
nothing.

The two questions read back the string PosixMode wrote rather than carrying its
nine booleans through SftpEntry as well. That is the point rather than a
shortcut: two representations of one fact is how a row ends up coloured for a
bit the column beside it does not show. A mode of the wrong length answers
false rather than throwing, since these decide a colour and a listing is not
worth failing over one.

The amber is Warn rather than WarnText, which is the muted amber a warning card
writes its sentences in. At 9.5px against TextFaint that one is a shade rather
than a signal, and a marker nobody notices is the same as no marker.

The local pane is untouched, on the grounds it already gives for having no
PERMS column at all: a POSIX mode is not a fact about a file on Windows, and
colouring one there would invent exactly what the column declines to print.

Twenty cases in RemotePathTests, which needs no container — the execute bit in
any of the three triples rather than only the owner's, the others-write bit
alone, a mode of the wrong length, and the file-only rule for both questions
from all three kinds. dotnet format is clean and the app and layout suites pass
at 109 and 35.
2026-07-31 12:17:56 +02:00
jaap-jan 1292084af9 Merge branch 'claude/delete-confirmations-becf4a'
ci / build and test (push) Failing after 2s
2026-07-31 11:52:27 +02:00
jaap-jan 91438fb382 Ask before deleting, and connect a host by double-clicking it
DELETE on a host, an SSH key, a stored password or a file on the host now puts
a question where the button was, and only answering it deletes anything. It is
a state rather than a dialog, which is the arrangement signing out already had
and for the same reason: this is the moment that has to be able to say what is
about to go before it goes.

What the question says is counted rather than generic, because a confirmation
that only asks whether you are sure is a click to train people out of. A key
names the hosts that authenticate with it and says they will refuse to connect
afterwards rather than falling back to a typed password, which is what the
connect path actually does. A host discloses a terminal open on it, because
deleting the host does not close the session. Every vault deletion says how far
it travels and whether this machine can push the tombstone yet or is queuing
it. Deleting on the host carries the strongest warning of the four on purpose:
everything else here is a tombstone against a copy the server still holds, and
a file on somebody's machine is bytes with nothing behind them — so that one
names the full path, since a bare name identifies nothing.

The armed request carries the item's entity id, so nothing that moves the
selection between the question and the answer can redirect it, and answering
about something that has since gone says so instead of doing nothing quietly.
Disarming compares ids rather than rows, which is the subtle half: a reload
replaces every row object, so the naive rule would have let the pass that runs
every minute take the card away from somebody halfway through reading it.

Forgetting a pinned host key is deliberately still unguarded. It costs one
fingerprint check on the next connection and it is the safe direction to be
wrong in — the dangerous button there is the one that adds trust, and that one
is already a prompt at connect time. Discarding a stopped transfer is likewise
unguarded: it removes a resumable part file and leaves the source alone.

Double-clicking a host in the sidebar connects to it, wired as a gesture in the
control exactly as the transfers screen opens a directory. CONNECT stays, since
it is the button with the password box beside it.

Ten existing delete call sites now go through arm-and-confirm helpers, and
eight new flow tests cover asking first, cancelling, the counted warning,
disarming on a selection change and on an editor opening, surviving a sync, and
the stale-item guard. Three layout tests measure the new shapes — the sidebar
card is the one card in the application a user cannot scroll — and one of them
also asserts the card renders its text, because a card whose compiled bindings
did not resolve would lay out perfectly as empty rows. The double-click test
performs the real gesture and proves it reached the connect command through a
refusal that never touches a network.

dotnet build, dotnet test and dotnet format --verify-no-changes are all clean:
853 tests, including the end-to-end suite against real containers.
2026-07-31 11:52:13 +02:00
jaap-jan 9608d73747 Come back from a sync position the server will not accept
ci / build and test (push) Failing after 2s
"The server returned 400: The sync cursor is not valid for this vault. Resync
from the beginning." told the user exactly what to do and gave them no way to
do it. The cursor is the only thing a pull sends, so the refusal was permanent:
the next pass read the same stored cursor and was told the same thing, once a
minute, for ever. And because the pull runs first, the exception ended the pass
before it reached the outbox — so the vault stopped receiving other machines'
changes and stopped sending its own. A machine that met this went quietly
read-only until somebody deleted its cache.

The engine now does what the message asks. A pull refused with the
invalid-cursor problem code — the code, never the prose, which is free to
change — drops this vault's position, writes that down, and reads the log again
from the beginning. The restarted request carries no cursor, which is the one
position a server cannot reject, so the retry cannot loop; a refusal of that is
rethrown rather than retried, and a restart is allowed once per pull. The
position is saved before the replay starts, so a process that dies halfway
through begins the next one from the beginning too rather than meeting the same
refusal again.

The mirror is deliberately kept. Replaying rewrites every row the server still
has and applying a change is a blind overwrite, so the re-pull repairs the
mirror on its way past; clearing it first would claim more than the evidence
supports — the position was refused, not the contents — and would leave a
machine that lost its connection mid-replay with less than it started with.
That leaves one gap, named in the remarks rather than left to be discovered:
once tombstone collection exists, a replay stops carrying deletions older than
the retention window.

None of the causes are the user's doing — a rotated cursor signing key, a vault
served from a restored database, a cache copied between machines — so nothing
asks them to decide anything. The report carries ResyncedFromStart and the
status line says the position was not recognised and the vault was read again.
It is kept out of NeedsAttention, because nothing is outstanding, but the
background pass breaks its usual silence for it: a sync that pulled the whole
vault on a day nobody changed anything otherwise reads as a fault.

The fake server grew a switch that refuses cursors the way a rotated signing
key does, including ones it minted itself. Three cases: the vault is re-read
and the change on the far side of the refused position arrives; the edits
waiting in the outbox are still pushed in that same pass, which is the half
that made this worth recovering from rather than merely reporting; and a server
that refuses the beginning itself is surfaced instead of replayed against.

dotnet build is clean at zero warnings, dotnet format is clean, and the sync
and app suites pass — 109 and 101.
2026-07-31 11:32:14 +02:00
jaap-jan 240aadb746 Merge branch 'main' into claude/vault-unlock-logout-autosync-a84c35
ci / build and test (push) Failing after 3s
Four files needed a hand, and all four were two branches adding something in
the same place rather than either changing what the other did.

The shell's constructor now takes both new parameters: main's SFTP session
factory, which it must have because it builds the transfers view model, and
this branch's optional resume handler, which stays last so every existing test
that constructs a shell without one still gets a shell that can only be online
because somebody signed in during this run. App.axaml.cs, ShellFlowTests and
QuickConnectTests pass the pair; the layout suite keeps both of its new fields.

Signing out now detaches the transfers screen exactly as locking does, and the
confirmation says that an open transfer session survives it. That is the same
policy both sides already argue for their own case: signing out destroys this
machine's copy of the vault, not work that authenticated before it.

QuickConnectTests did not compile on main — the SFTP commit added a constructor
parameter and the quick-connect suite, merged from a parallel branch just
before it, was still calling the old one. Fixed here rather than worked around,
since the merged tree has to build.

dotnet build, dotnet test and dotnet format --verify-no-changes are all clean:
980 tests, including the end-to-end suite against real containers.
2026-07-31 11:16:49 +02:00
jaap-jan d1700f5a34 Merge branch 'claude/m2-file-transfer-1b9951'
ci / build and test (push) Failing after 3s
2026-07-31 11:08:05 +02:00
jaap-jan 0b261c4d39 Stay signed in, come back online by itself, and let a machine be given up
Three things a machine that has been set up could not do. Unlock now takes
Enter, which is the gesture everybody makes after typing a password and which
did nothing until they found the button.

Signing in survives a relaunch. The refresh token is kept in the local cache,
sealed under the vault's own cache key, so a later launch resumes the session
through the refresh grant with no browser and nobody present — and because it
is sealed under that key, only an unlocked vault can resume it. A locked
client therefore cannot reach the server at all, which is a consequence worth
stating rather than working around; docs/crypto.md §3.2 records it. Every sync
pass asks the shell for a connection rather than reading one captured at
unlock, so a laptop that unlocked on a train is online within a minute of
finding a network, with nothing pressed. Unlocking itself still never waits on
a socket.

Signing out empties this machine: the profile, the cached items, the outbox
and this machine's device key, with the account's row withdrawn when the
server can be reached. It asks first and says what it costs — the outbox count
when the vault is open, an admission that it cannot be counted when it is not,
and the shells that keep running either way. The vault is on the server and is
untouched, which is what makes the same button the only honest answer to a
forgotten passphrase, so it is on the unlock screen as well as in preferences.
It cannot end the session at the identity provider, and says so.

Two defects surfaced on the way. The synchronisation pass that runs when the
vault opens never ran at all: the loop is started from inside the unlock
command, so the busy flag it yields to was raised by that command — the first
sync was a minute late on every launch. And signing in from preferences while
unlocked threw an unlock screen over an open vault whose keys were still in
memory.

The unlock card and the new confirmation live in their own controls because
MainWindow cannot be laid out headless, so markup left inside it is markup no
test can measure; both are now measured at the window's minimum size in the
shapes that grow. What is still unverified is the composed window itself.
2026-07-31 11:07:36 +02:00
jaap-jan 04faef6597 Move files to and from a host over SFTP
M2's file transfer, built bottom-up: an SFTP session on the SSH layer, a
transfer queue in a project of its own, and the two-pane browser the design
asked for replacing the screen that said it did not exist. Remote listings
carry names, sizes, modification times and a real drwxr-xr-x — nothing in this
repository could render a POSIX mode before — and the queue moves one file at a
time with progress, throughput and resume.

The design import assumed this would be an SFTP subsystem channel on
ISshConnection, beside the shell on a transport that is already up. SSH.NET
does not offer that: SftpClient derives from BaseClient and owns its own
transport, and there is no supported way to hand it an SshClient's session. So
file transfer opens a second authenticated connection, and it is named for
that rather than dressed up as a channel — OpenSftpAsync is on
ISftpSessionFactory, not on a connection. The difference is visible to a user:
the host records a second login, and a host whose password is typed each time
asks for it again on this screen. It goes through the same host key gate, the
same pin and the same two refusals a shell does, so a fingerprint approved for
a terminal is approved here and one approved here reaches the other machines
with the next sync. docs/design-import-gaps.md is corrected, and marked as the
one row where what shipped differs from what it predicted.

Nothing is written at its final name until it is complete. Every transfer goes
to a .dodossh-part file beside its destination and is renamed into place at the
end, so an interrupted transfer can never be mistaken for a finished one —
which matters most for what this screen is actually for, which is copying a
build artefact onto a server and then running it. A destination that already
exists is refused outright rather than overwritten: the queue has no way to
ask, and silently replacing a file somebody's process is serving is the worse
of the two failures. The remote pane has DELETE and MKDIR so that refusal is
not a dead end. A test against the container pins the assumption underneath all
of this — that SFTP's rename does not clobber.

Resume works within a run of the application and not across a restart, and the
limit is deliberate rather than unfinished. Nothing records which source wrote
a part file, and resuming one on the strength of its name matching is how a
corrupt artefact gets delivered with nothing reporting a failure; a part file
found at startup is started over. Making it survive a restart needs the
preferences store this client still has not got. The offset a resume starts at
is the part file's own length rather than the transfer's recorded progress: a
cancellation can land between a write completing and the counter moving, and
only one of those two is a fact about the bytes that are there.

The queue and its connection outlive a lock, as shells do. LockAsync already
argues that locking must not destroy work in flight — it is what somebody does
when they walk away from the machine, which is exactly when a long transfer is
most likely to be running — so TransfersViewModel is created once and the vault
is attached on unlock and detached on lock. What locking takes is the host
list, and it has to: those rows carry decrypted secrets.

DodoSSH.Client.Transfer is a new project rather than more of Client.Ssh. The
two answer different questions — one is about reaching a host, the other about
moving bytes and what to do when moving them stops halfway — and this is the
only client project that deliberately touches the local filesystem.

Three defects the tests found, none of which review would have. SftpPath.Name
answered an empty string for the root. NavigateRemoteAsync wrapped itself in
the busy guard, so navigating from inside another command did nothing at all
and the remote pane simply stayed empty after connecting, with no failure
anywhere to explain it. And opening an SFTP session per test made two
handshakes per test — this client learns a host key by being refused — which
pushed the SSH assembly past sshd's MaxStartups and failed a different few
unrelated tests each run; the session is shared through the fixture now, with
the reason written where the next person will hit it.

1004 tests green across 18 projects, 24 of them new: the SFTP subsystem against
the OpenSSH container, the queue against a real temporary directory and a fake
host, and three more layout measurements because a screen this window has never
laid out is a screen never checked.

Not verified: the screen has not been looked at running. The layout harness
measures it at the window's minimum in three shapes, which is the class of
defect that has shipped here before, but reaching it in the application needs
the compose stack, the migrations, the API and a browser sign-in. What is still
absent — the status bar's transfer count, dragging between the panes,
transferring a directory, and sftp over a bastion — is in
docs/design-import-gaps.md.
2026-07-31 11:07:29 +02:00
jaap-jan f7c5096bc6 Keep the stub servers on loopback
Running the tests raised a Windows Firewall prompt, and raised it again from
every worktree. WireMockServer.Start() with no settings listens on 0.0.0.0
and [::], and the prompt is keyed to the binary that opened the socket — so
each test executable asks once per bin path, which a new worktree or a switch
between Debug and Release makes new again. The three suites that hold a
firewall rule on this machine are exactly the three that use WireMock; every
other listener in the repository already binds 127.0.0.1.

The stubs now say so explicitly. Port 0 is still WireMock's own free-port
search and still comes back on server.Url, which is what each stub builds its
base URL from, so the authority the API validates against and the issuer its
tokens claim follow the binding rather than being pinned to a host name.

Sampling the listening sockets of a full DodoSSH.Api.Tests run afterwards
finds one, 127.0.0.1, where there were previously three.
2026-07-31 11:06:46 +02:00
jaap-jan 4eaa8eae6d Merge branch 'claude/search-modal-closing-cd05c4'
ci / build and test (push) Failing after 2s
2026-07-31 10:45:46 +02:00
jaap-jan 66271faaae Update .github/workflows/ci.yml
ci / build and test (push) Failing after 16s
2026-07-31 08:44:19 +00:00
jaap-jan 9c3edb078e Update .github/workflows/ci.yml
ci / build and test (push) Canceled after 0s
2026-07-31 08:43:39 +00:00
jaap-jan 312d766c30 Let the quick-connect palette answer for itself
Clicking outside the palette did nothing, because nothing was listening: the
wash took no pointer input at all, so the only ways out were a key and the
button that opened it. It now closes on a press whose source is the wash
itself, which is what separates outside from inside — a press on the card
bubbles through the same handler on its way to the window, and closing on
those would make the palette impossible to click into.

The caret never reached the query box either. The window focused it from the
view model's PropertyChanged, and that handler runs before the binding which
reveals the control — measured, with the same wiring, in a replica window. So
it focused a control that was still collapsed, which Avalonia treats as a
no-op and does not replay when the control is revealed, and the keyboard
stayed wherever the click that opened the palette had left it. Becoming
visible is now what triggers it, posted rather than called: a control that has
never been laid out has no visual children, and at the instant IsVisible turns
true the box still reports IsAttachedToVisualTree() == false.

Escape, Enter and the arrows move to the palette as a tunnelled handler.
Answering them only on the window was fragile in the way that matters here:
anything on the route that took a key first would silence them, and with the
focus never landing in the palette the key was being pressed at whatever the
opening click had focused — a focused Button eats Enter. The window keeps
Ctrl+K, which has to work when the palette is not showing, and forwards the
rest as the net for a press that arrives from outside the palette.

Which is also why this moved out of MainWindow rather than being fixed there.
Showing MainWindow initialises WebView2 on a thread it refuses, so nothing on
that window can be tested — the palette shipped with no test of any kind. As a
UserControl it hosts in a bare window and takes real key and pointer input,
and there are now six: press on the wash closes, press on the card does not,
Escape closes, the arrows move the selection without taking the caret out of
the box, Enter takes the highlighted host, and the palette takes the keyboard
when it appears.
2026-07-31 10:43:30 +02:00
jaap-jan f0002b683c Update .github/workflows/ci.yml
ci / build and test (push) Canceled after 0s
2026-07-31 08:39:25 +00:00