9801a744baec628410c8c48082d1e4273c13496f
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
093f3904c1 |
Apply pending migrations at startup instead of asking for a second command
The API deliberately never migrated: it failed readiness while a migration was pending and named it, and a separate step applied them. That is the right split for a deployment with a release pipeline and the wrong one for a self-hosted server, where it means an image that boots, refuses traffic, and waits for somebody to know that dotnet ef exists. The schema and the code that expects it ship in the same image, so the image is where the two are reconciled now. Before RunAsync rather than in the background. A migration racing the first requests would let them through against a half-applied schema, and the first authenticated request is the one that provisions accounts. Failing to migrate therefore fails to start, which is the loudest signal available and the one an orchestrator already acts on. Concurrent starts take a Postgres advisory lock first. Without it two replicas rolled out together read the same empty history table, both apply the same migration, and the second dies on an object that already exists — a crash loop on the day of a schema change, which is the worst day to have one. The lock is held on a connection of its own because EF opens and closes one per command, and a session lock belongs to the connection that took it. The exception is a database that does not exist yet: there is nothing to hold a lock in, so that path migrates without one and says so. Two instances creating it at once still converges — one wins, the other restarts into the ordinary locked path — and refusing to start would leave a fresh deployment stuck on the step this removes. Database:AutoMigrate turns it off for the deployments that own their schema: a migrator job, a rollout where new code must run against the old schema first, or a database user denied DDL. With it off the behaviour is exactly what it was, and the health check now explains which of the two situations a pending migration means. Verified against a throwaway PostgreSQL container: an empty database gets all seven migrations applied before the port opens, the tables land in the dodo schema, and a second start logs the schema up to date and serves. The API suite passes, which exercises the startup path once per assembly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d3b14e6bc0 |
Add configuration, OIDC auth wiring and discovery endpoints (M1)
Options, JWT bearer validation, the /meta and .well-known endpoints, and a dev compose stack with Keycloak. Verified end to end: compose up, migrate, run, both discovery endpoints return correct payloads, and readiness reports the schema current. Configuration: - Strongly-typed options for Server, Oidc, Relay and Sync, all ValidateOnStart. A self-hosted server that boots half-configured and fails later per-request is far harder to diagnose than one that refuses to start and names the bad setting. - Cross-field validation the annotations cannot express: relay needs a WebSocketUrl when enabled, idle timeout must be under max session duration, item payload cap under batch cap. - Startup warnings for combinations that are individually valid but dangerous together: RequireHttpsMetadata false outside Development, and AllowEmailLinking (which turns any token bearing a victim's email into account takeover, hence default false). Auth: - JwtBearer with ClockSkew cut to 30s from the 5-minute default; five minutes of slack on a credential granting vault ciphertext access is more than any clock needs. - IncludeErrorDetails off, and a FallbackPolicy so an endpoint without an explicit policy still requires a caller rather than silently being public. Discovery, per ADR 0002: - /api/v1/meta reports versions, features and push caps. - /.well-known/dodossh-configuration is the onboarding story: the user types one server URL and the client discovers OIDC authority, client id, scopes and relay endpoint. Two environment problems found by actually running the stack: - PostgreSQL 18 changed its data mount point. Mounting /var/lib/postgresql/data — correct through 17 — makes the image refuse to start; 18+ wants a single mount at /var/lib/postgresql with the cluster in a subdirectory. - Keycloak moved to host port 18080. An unrelated Apache Tomcat on this machine holds 127.0.0.1:8080, and a loopback-specific bind beats Docker's 0.0.0.0 publish for "localhost". It presents as Keycloak 404ing every realm while its own log says the import succeeded, which is a genuinely misleading failure. Also: CA1848 is enforced, not advisory — warnings are errors, so the .editorconfig comment claiming otherwise was wrong. Startup and health logging now uses [LoggerMessage]. And a clean rebuild is back to zero warnings; the incremental build had been hiding 40 in test projects (banned Guid.NewGuid, an obsolete Testcontainers constructor, and two analyzer families that are genuinely noise under a test host). Verified: 0 warnings on a clean rebuild, 122 tests pass, format clean. |