Files
DodoSSH/README.md
T
jaap-jan 95816de0c5 Share a vault with a team, without the server holding a key
M3's teams, sharing and ACLs. Teams with roles, a public-key directory, the
append-only key log served for clients to check it against, team-owned vaults,
and vault key grants wrapped by a client and stored opaquely by the server.
VaultAccessService resolves team membership to PermissionFlags, so a viewer may
pull and may not push; the desktop client reads and syncs every vault it holds
a key for, and a real TEAMS screen replaces the one that said it did not exist.
No migration: team, team_membership, vault.team_id and vault_key_grant have all
been there since the first one, which is what carrying two unused tables bought.

Membership is authorisation. A grant is access. The obvious model is one
concept — "access", with a role attached, handed out by the server — and this
architecture cannot implement it: a vault key is sealed to each member's X25519
key, and only a client holding the plaintext can seal it for somebody else. So
"give Bob access" decomposes into a database write and a wrap, which happen on
different machines. Adding a member makes the server serve them the vault; it
cannot make it readable. VaultSummary.WrappedVaultKey is null in the meantime
and the vault appears in their list saying it is waiting for a key, because
hiding it until a grant existed would have been tidier and would have implied
the server was the thing granting access. The screen says the same thing after
every add, in the status line. ADR 0009 records the whole decision.

Sharing verifies or refuses. A directory lookup is a claim by the server about
a third party's public key, and wrapping to an unverified claim hands the vault
to whoever made it — no amount of transport security helps, because the server
is inside the threat model. KeyLogAudit reads the whole log, recomputes every
entry's hash from its own contents, checks the chain from genesis, and refuses
unless the offered key appears in it unchanged. There is no override flag: one
that exists gets used on the day the log is briefly unreachable, and the
resulting grant is indistinguishable from a correct one afterwards. What it
still cannot promise is that the key is the right person's, so the fingerprint
comes back for an out-of-band comparison and the success message says so every
time. A test corrupts the fake server's log by one byte and watches the client
refuse rather than warn.

The roles are only the ones that are enforceable. There is no ConnectOnly,
despite the design asking for one and TeamRole having room: SSH terminates on
the client, so a session needs the credential's plaintext on that machine, and
"may connect but may not read the key" cannot be enforced here. Shipping it as
an option in a dropdown would have been a lie. Connect rides along with Read
and is documented as an interface hint. Removal is named for what it does — it
revokes grants and flags the vault for rekey, and claims nothing about what is
already on somebody's laptop.

Three things are deliberately absent, and each is a refusal rather than an
omission. The rekey itself, because re-wrapping every item's data key under a
new vault key needs a client holding the current one; the server records that a
rotation is owed and the interface reports it, which is more honest than a
button that only appears to do it. Ownership transfer, because allowing an
owner to be removed without one leaves a team nobody can administer. And
cross-vault host key trust: a pin in a team vault is listed but not consulted
at connect time, because any member with Write could otherwise pre-approve a
fingerprint another member's client then trusts silently for a host in their
own vault. Scoping trust properly needs a scope on the SSH connect path, which
IKnownHostStore has not got; until then the narrow direction is the safe one
and the cost is in the README rather than hidden.

Reading now spans vaults and writing still does not. Every list on the vault
and hosts screens covers each vault the keyring opened, rows carry the vault
they came from, and an edit goes back to that vault rather than to the active
one — writing it to the active vault would fork the item and only show up when
a colleague wondered why their change never arrived. A new item goes wherever a
picker says, defaulting to the personal vault and never moving on its own,
because an item filed into a team's vault is visible to that team and moving it
back means deleting and retyping. The sidebar heading stops naming one vault
once there are two, and each row names its own.

The server checks what it can and nothing it cannot. It will not record a grant
for a key its recipient no longer holds, for a superseded generation, or for
somebody who is not in the team — each of those would otherwise surface days
later at the far end as a tag failure indistinguishable from corruption. It
does not verify the wrap or the signature, and the grant service says so: that
would be a convenience and never the boundary, and would put an asymmetric
implementation on a machine that is supposed to hold no keys.

Two bugs the tests found. TeamsViewModel's busy gate blocked its own reload, so
a team created a moment earlier was missing from the list it had just been
added to. And syncing every vault turned a failure from an exception into a
report, which made a background pass announce an unreachable vault once a
minute — the exact behaviour AnAutomaticPassThatFails_LeavesTheStatusAlone
exists to prevent. The fact is recorded and the message swallowed, as it was
before; pressing Sync still names the vault and the reason.

Also fixes a build break this branch started with: QuickConnectTests was never
updated when M2 added ISftpSessionFactory to the shell's constructor, so
nothing built at all.
2026-07-31 12:18:28 +02:00

327 lines
20 KiB
Markdown

# DodoSSH
A self-hosted, team-oriented SSH client with an end-to-end encrypted vault.
Manage hosts, credentials and keys in a desktop app; sync them across your devices and share
them with teammates through a server you run yourself. **The server stores ciphertext and never
holds a key** — the operator cannot read the credentials it stores.
> Status: early development. See [the milestone plan](#milestones) for what exists today.
## Why
Teams either scatter SSH credentials across individual `~/.ssh` directories with no sharing
story, or pay per-seat for a hosted product that holds their infrastructure credentials.
DodoSSH keeps the convenience of a synced, shareable vault while remaining self-hostable and
zero-knowledge.
## Architecture
| Component | Choice |
| --- | --- |
| Backend | ASP.NET Core on .NET 10, PostgreSQL + EF Core |
| Client | Avalonia (C#) for Windows/Linux/macOS; terminal pane is a WebView running xterm.js |
| Auth | OIDC, provider-agnostic (Entra ID, Keycloak, Auth0, Authentik) |
| Vault | End-to-end encrypted; X25519 + Ed25519 + XChaCha20-Poly1305, Argon2id unlock |
| Connections | Client-direct SSH by default, with an optional raw-TCP server relay |
Three consequences worth knowing before you read further:
- **Revocation is not retroactive.** A removed member keeps what they already downloaded. The
real remediation is rotating the SSH credential, so offboarding is built around a rotation
checklist rather than a button that implies more than it delivers.
- **No session recording in relay mode.** The relay forwards SSH ciphertext, so it cannot see
commands. That is the cost of the relay not being able to read your traffic.
- **Locking the vault does not close your shells.** Lock closes the vault and zeroes every key it
held; a session that authenticated before it keeps running, because the remote host never
consulted the vault and the credential was already spent. That is deliberate — you lock when you
walk away from the machine, which is exactly when a long upgrade or transfer is most likely to be
in flight, and an idle auto-lock that killed it would be worse than the exposure it removed. The
honest reading is that *locked* describes the vault and not this machine's access to your hosts.
The unlock screen therefore shows how many shells are still connected, and quitting DodoSSH is
what ends them.
The reasoning behind each major decision is recorded in [`docs/adr/`](docs/adr/), starting with
[the E2EE trust model](docs/adr/0001-e2ee-trust-model.md).
The desktop client's interface was built from a design covering more product than exists yet — file
transfer, teams, saved snippets, port forwarding. Everything that design asked for and this build has not
got is written down in [`docs/design-import-gaps.md`](docs/design-import-gaps.md), with the layer each
piece would land in and what the interface shows in its place. Nothing was rendered with invented data to
fill a screen.
## Repository layout
```
src/
DodoSSH.Contracts DTOs shared with the client — the real API contract
DodoSSH.Crypto DSH1 envelope, AAD derivation, the key hierarchy
DodoSSH.Domain entities and invariants, no EF
DodoSSH.Infrastructure DbContext, configurations, migrations
DodoSSH.Api the server
DodoSSH.Client.Auth OIDC code+PKCE on a loopback redirect, and the key binding
DodoSSH.Client.Api the typed server client, and client-side enrollment
DodoSSH.Client.Domain the decrypted item model and the three-way merge — no I/O at all
DodoSSH.Client.Storage the local cache: ciphertext mirror, outbox, offline unlock material
DodoSSH.Client.Sync the pull/apply/push loop and the conflict policy
DodoSSH.Client.Session where a profile lives, unlocking it, and getting one in the first place
DodoSSH.Client.Ssh connections, PTY shells, SFTP, host key trust
DodoSSH.Client.Terminal the loopback data plane and credit-based flow control
DodoSSH.Client.Transfer the transfer queue, part files and resume, and the local file listing
DodoSSH.Client.App Avalonia; the only project that knows about a UI toolkit
tests/ one test project per source project
docs/adr/ architecture decision records
docs/design-import-gaps.md what the client's design asked for and this build has not got
```
Everything under `src/DodoSSH.Client.*` except `App` is deliberately free of Avalonia. That is the
seam that lets the SSH layer, the terminal's flow control and the OIDC flow be tested without a UI
toolkit or a browser engine — which is most of why they are testable at all.
## Building
Requires the .NET SDK pinned in [`global.json`](global.json) (10.0.x).
```bash
dotnet build DodoSSH.slnx
```
```bash
dotnet test DodoSSH.slnx
```
The tests need a Docker daemon. Everything that touches the database, the identity provider or an SSH
server uses Testcontainers rather than a stub or a shared instance, so there is nothing to start first and
nothing to clean up after — but with no daemon those suites fail rather than skip.
## Running it
Four commands, in order. The first two are once per machine.
**1. The development dependencies** — PostgreSQL and Keycloak, with the `dodossh` realm imported:
```bash
docker compose -f deploy/docker-compose.dev.yml up -d
```
**2. The schema.** The API never migrates anything: it fails readiness while a migration is pending, and
says which one. `dotnet-ef` is pinned in `.config/dotnet-tools.json`, so run `dotnet tool restore` first if
you have not:
```bash
dotnet ef database update --project src/DodoSSH.Infrastructure
```
With nothing else configured this targets the compose stack above. Set `DODOSSH_DESIGN_CONNECTION` to point
it at another database.
**3. The server:**
```bash
dotnet run --project src/DodoSSH.Api
```
It listens on `http://localhost:5233`, serving `/healthz/live`, `/healthz/ready` and — in
Development — `/openapi/v1.json`.
**4. The desktop client:**
```bash
dotnet run --project src/DodoSSH.Client.App
```
In the app, enter `http://localhost:5233` as the server. Your browser opens for sign-in — the realm ships
`alice` / `alice` — then choose a vault passphrase and **write down the recovery code**, which cannot be
skipped and cannot be recovered from the server. You can then add a host and open a shell on it. Keycloak's
admin console is at `http://localhost:18080` (`admin` / `admin`).
You can also add an SSH key, which is stored in the vault like a host and synced the same way: paste the
private key, then edit a host and pick that key from its **key** dropdown. From then on that host
authenticates with it — on every machine, since the choice travels inside the host's encrypted payload —
and its password box disappears.
The first time you connect to a host you are asked to check its key fingerprint. That decision is stored in
the vault, so it is asked once per host rather than once per launch and it reaches your other machines with
the next sync. If a server is legitimately rebuilt and offers a new key, the connection is refused outright
with no way to continue from the warning — edit the host and choose **Forget host key**, which is deliberately
somewhere you have to go on purpose.
Two of M1's known gaps are visible immediately, so they are worth expecting rather than diagnosing: password
authentication asks for the password every time, because nothing in the interface can create a vault
credential yet (they do sync — there is just no editor for one); and unlock asks for the passphrase on every
launch, because no device key is registered.
### Moving files
**FILES** in the nav rail is a two-pane browser: this machine on the left, the host on the right, and a
queue underneath. Choose a host, press **CONNECT**, then select a file in either pane and press the arrow
pointing the way you want it to go.
Two things about it are worth expecting rather than discovering.
**It is a second connection, not a second channel.** SSH itself would allow the SFTP subsystem to open
beside a shell on the transport that is already up; SSH.NET does not offer that — its `SftpClient` owns its
own transport — so pressing CONNECT here authenticates again. The host records a second login, and a host
whose password you type each time will ask for it again on this screen. Host key trust is shared: a
fingerprint approved for a terminal is approved here, and one approved here reaches your other machines with
the next sync.
**Nothing is written at its final name until it is complete.** Every transfer goes to a `.dodossh-part` file
beside its destination and is renamed into place at the end, so an interrupted transfer can never be
mistaken for a finished one — which matters most for what people actually use this for, which is copying a
build artefact onto a server and then running it. A destination that already exists is refused outright
rather than overwritten; the remote pane has **DELETE** and **MKDIR** so that refusal is not a dead end.
**RESUME** on a stopped transfer carries on from what the part file already holds.
Resume works within a run of the application and not across a restart, and that limit is deliberate: nothing
records which source wrote a part file, and resuming one on the strength of its name matching is how a
corrupt artefact gets delivered with nothing reporting a failure. A part file found at startup is started
over.
What is not here: transferring a directory, dragging between the panes, and routing a transfer through a
bastion — the last needs jump hosts the connection layer has not got. All three are in
[`docs/design-import-gaps.md`](docs/design-import-gaps.md).
### Working as a team
**TEAMS** in the nav rail creates a team, adds members and shares vaults. One distinction runs through the
whole screen and is worth having before you use it.
**Adding somebody to a team and giving them a key are two different acts, and only the first is something
the server can do.** Adding a member changes what the server will *serve* them: the team's vaults appear in
their list immediately. It cannot make those vaults readable, because a vault key is sealed to each member's
public key and this server never holds one — so until somebody presses **SHARE KEY** from a machine that has
the key, their vault sits in the list saying it is waiting for one. That is not a rough edge to be smoothed
over later; it is what "the operator cannot read the credentials it stores" costs, and the screen says so
rather than implying the server handed anything out.
Sharing verifies before it wraps. The client reads the server's append-only key log, checks its hash chain
from the first entry, and refuses unless the key the directory just offered appears in that log unchanged.
That converts a key substitution by the server from invisible into visible — a substituted key has to be
published in a log every other client also reads. **It does not prove the key is the right person's.**
Compare the fingerprint with them over something this server does not carry; that is the only step that
closes it, and the success message says so every time.
Three limits, stated rather than discovered:
- **Removing a member is not retroactive.** It revokes their grants and flags the team's vaults for rekey,
and blocks future reads. Everything they already pulled is on their machine. Rotate the SSH credentials
that matter — that is the actual remediation, and it is why there is no button labelled anything stronger.
- **The rekey is flagged, never performed.** See the milestone note above.
- **Host key trust stays in your personal vault.** A pin approved for a team's host is recorded and used
from your own vault, not the team's, so a teammate cannot pre-approve a fingerprint that your client will
then trust silently for a host you defined. The cost is that each member approves a team host's key once
on each of their machines. Team vaults' pins are still *listed* on the Vault screen, so you can see what
has been trusted.
Items are filed into one vault at a time. When more than one vault is writable, the host and vault editors
show a picker; it defaults to your personal vault and never moves on its own, because an item put in a team
vault is visible to everybody in that team and moving it back means deleting and retyping.
### End-to-end verification
One suite runs against a real server rather than a stub. It needs a Docker daemon and nothing else, so it
is part of the ordinary test run:
```bash
dotnet test tests/DodoSSH.SystemTests
```
It brings up PostgreSQL, Keycloak and an OpenSSH server in containers, applies the committed migrations,
starts the API as a child process out of its own build output, and then drives the real client: sign in
through Keycloak, enroll, unlock, create an SSH key and a host bound to it, sync them, open a shell on the
`sshd` and approve its host key at the real first-contact refusal, then read all three back on a second
simulated machine and unlock again with no network. Roughly 25 seconds once the images are pulled.
What makes it worth its weight is that it consumes the artefacts that ship — the realm file from
`deploy/keycloak`, the EF migrations, the API's own `appsettings` — rather than a fixture written to match
them. On its first run it found a loopback redirect URI the realm registered in a form Keycloak rejects,
and a JSON configuration gap that made the whole sync surface unreachable from the real client while every
other test passed. Both are the same class of bug: two sides of a stub agreeing with each other about
something the specification never said.
The one value it cannot take from a committed file is `Oidc:Authority`, since the container's port is
assigned at start. Everything that authority points at is still the real realm.
Development and testing are currently **Windows-only**. Anything known or suspected to differ on
Linux and macOS is tracked in [`docs/platform-flags.md`](docs/platform-flags.md), along with the
deployment gotchas that have already cost time once. Read it before assuming something works
off-Windows.
### Conventions the build enforces
- Warnings are errors. `dotnet format --verify-no-changes` gates CI.
- Package versions are centralised in `Directory.Packages.props`; `packages.lock.json` is
committed and CI restores in locked mode.
- [`BannedSymbols.txt`](BannedSymbols.txt) bans `DateTime.UtcNow` (use `TimeProvider`),
`Guid.NewGuid` (use `CreateVersion7`), sync-over-async, MD5/SHA1 and PBKDF2.
- Public members of `DodoSSH.Contracts` must be declared in `PublicAPI.Unshipped.txt`, so a
contract change is a build error rather than a client-side surprise.
## Milestones
- **M0 — foundation.** Repo structure, build conventions, CI, ADRs. *Done.*
- **M1 — vertical slice.** OIDC login → enroll → create a host → open a shell.
*Server done:* the DSH1 crypto core, the data model, sync push/pull for hosts, `/me`, and
enrollment with the identity-provider key binding.
*Client done:* the key hierarchy, the OIDC flow with the key binding, SSH connections with host key
trust, the terminal data plane, the encrypted local cache with the sync client — offline unlock, an
outbox and a field-level three-way merge, conflict matrix green — and an Avalonia shell that is
vault-backed: server URL → browser sign-in → enroll → unlock → host list → terminal. The shell's *state
machine* is covered by tests against an in-memory server, so the states that matter most (the recovery
code that cannot be skipped, the unlock that needs no network) are checked rather than remembered.
Its *layout* is not covered by anything, and that gap has already cost a shipped defect: the setup and
unlock screens were layered over the terminal's WebView, which on Windows is a native child window that
cannot be covered, so they rendered sliced with their buttons unclickable. No test in this repository
loads a `.axaml` file, and a headless one could not have caught this — there is no native window in
headless, so it would have rendered perfectly and confirmed the wrong belief. Screens get looked at, or
they are unverified.
*Verified end to end:* `tests/DodoSSH.SystemTests` drives the whole slice against a real Keycloak, a
real API, a real PostgreSQL and a real `sshd` — sign-in, the identity-provider key binding, enrollment,
offline unlock, a host and an SSH key through the vault to a second machine, an interactive shell, and the
host key approved at that shell's prompt reaching the second machine as well. See
[End-to-end verification](#end-to-end-verification).
Known gaps in the client, stated rather than implied by the interface: nothing in the interface can create
a vault credential yet, so password authentication still asks for the password each time — SSH keys *are*
editable, and binding one to a host is the way to connect without typing anything; and no device key is
registered, so the passphrase is needed on every launch until the OS keystore is wired.
Host key trust *is* in the vault, which is what makes trust-on-first-use worth having: a fingerprint
approved on one machine is approved on all of them and survives a restart, and the server cannot drop a
pin to force a fresh first-use decision without the item visibly going missing. A changed host key stays a
hard refusal with no way past it; withdrawing a pin is a separate, deliberate act in the host's editor.
Binding a key introduced the first payload schema version bump, and it is worth knowing how it behaves:
a host is written at the *lowest* schema version that can represent it, so only hosts that actually bind
a key are written at version 2 and become read-only on an older build. Hosts that do not are still
written at version 1, byte-identically to before the field existed — which is what keeps upgrading one
machine from making a team's whole vault uneditable everywhere else.
- **M2 — full personal vault**, robust sync, relay. *File transfer done:* an SFTP session, a two-pane file
browser with a real remote listing — names, sizes, modification times and `drwxr-xr-x` permission bits —
and a queue that moves one file at a time with progress, throughput and resume. See
[Moving files](#moving-files) for the two things about it worth knowing before you use it, both of which
are consequences rather than choices.
- **M3 — teams**, sharing, ACLs. *Done, except rekey.* Teams with roles, a public-key directory, the
append-only key log served for clients to verify against, team-owned vaults, and vault key grants
wrapped by a client and stored opaquely by the server. `VaultAccessService` now resolves team
membership to permissions, so a viewer may pull and may not push; the desktop client reads and syncs
every vault it holds a key for, and a real TEAMS screen replaces the placeholder. See
[Working as a team](#working-as-a-team) for the one distinction the whole design rests on, and the limits worth
knowing before you rely on it; the reasoning is in
[ADR 0009](docs/adr/0009-team-access-model.md).
**What is deliberately not here: the rekey itself.** Removing a member revokes their grants and flags
every team vault `RekeyRequired`, and nothing acts on that flag. A rekey re-wraps every item's data key
under a fresh vault key and can only be performed by a client that holds the current one; that is M5's
key rotation. Until it lands the flag is what the interface reads to say a rotation is owed, which is
more honest than a button that only appears to do it. Ownership transfer is absent for the same kind of
reason — the owner cannot be removed or demoted, because nothing can appoint a replacement.
- **M4 — hardening and ops**, packaging, self-hosting guide.
- **M5 — multi-provider OIDC**, key rotation, per-item content keys.
## Licence
[MIT](LICENSE).