Stop the test sshd penalising the suite for its own host-key refusals

The SSH suite has failed intermittently for months with SshConnectionException
"The connection was closed by the remote host", within milliseconds, on
whichever class happened to be running. Two previous attempts guessed at the
cause and said so honestly; this one has a mechanism and a before/after.

◆ THE CAUSE IS PerSourcePenalties, WHICH THIS SUITE PROVOKES BY DESIGN.

OpenSSH 9.8 added per-source penalties and 10.x enables them by default; the
image runs 10.3 and its config never mentions the keyword, so the compiled-in
default was what ran. A source address that repeatedly disconnects without
attempting authentication gets penalised, and while the penalty holds every
connection from it is answered with the clear-text line "Not allowed at this
time" and then closed.

That is exactly the traffic this suite generates. This client's first contact
with an unknown host is a connection deliberately refused at the host key —
a disconnect with no authentication attempt — so every helper that learns a
host key by being turned away first, plus RefusingTheHostKey_AbortsTheConnection
and AnUntrustedHost_IsRefusedExactlyAsAShellWouldBe, feeds the penalty counter.
Enough of them close together and sshd stops talking to the test host for a
while, then starts again.

Measured on a fresh container, probing 200 times with connections of that shape:
with the image default, the first refusal came back at probe 18 and 183 of the
200 were refused. With PerSourcePenalties no, none of 200 were. That is the
before/after the earlier attempts could not produce.

It also explains the shape of the failure, which never fitted a throttle. The
class that failed lost EVERY connection it made rather than a random few —
including the one test that expects a refusal, which passed throughout for the
wrong reason — while the classes around it were untouched. That is a window in
which the server refuses one source, not a probabilistic drop.

Both earlier diagnoses are recorded in the fixture so they are not tried again.
MaxStartups was blamed on the reasoning that xUnit runs test classes in
parallel, so ten unauthenticated connections would be in flight at once; but
every class touching this server shares one collection and xUnit parallelises
collections, not classes, so they run one after another and never have more than
a connection or two open. The reload window was blamed next, and a wait for the
banner was written and removed as unproven — it was unproven because the banner
answers perfectly right up until the penalty lands, so a check that stopped at
the first "SSH-" ran entirely inside the good part.

MaxStartups is kept, on the narrower argument that it is right regardless: a
connection throttle is hardening a test server has no business reproducing.
Removing it would be a second change riding along with this one.

The readiness gate that replaces the reconfigure's silence is a guard rather
than a wait. It requires 25 connections answered back to back, which is the
specific provocation rather than a soak test: 25 is above the measured
threshold of 18 on purpose, and it costs under a second when the setting is off.
Ten was tried first and was worse than useless — it sits below the threshold, so
it passed against a server that was still penalising. With the fix removed the
gate now fails in a minute naming PerSourcePenalties and quoting the server's
own "Not allowed at this time", instead of the suite failing later somewhere
unrelated.

The gate also closes a hole the container's own readiness cannot: a log line and
netstat showing :2222 both pass on a container whose sshd has gone, because
Docker publishes the port with a host-side proxy that accepts before it has
anything to forward to. It is probed from the host rather than with docker exec
for the same reason it matters — that is the path the tests take, and penalties
are counted per source address.

Rejected: patching sshd_config from /custom-cont-init.d to avoid the reload
entirely. It looks like the right hook and is not — the container's log puts
"sshd is listening on port 2222" before "[custom-init] Files found, executing",
so a script there edits a file the running server has already read. It leaves a
config that greps correctly and a server behaving as though it were never
touched, which is the same trap as patching the wrong one of the image's two
config files. Twenty-eight tests failed before that was noticed; the finding is
in the fixture.

Four consecutive full-solution runs clean, and the SSH suite green on every run
since. 1,861 tests, none failing.
This commit is contained in:
2026-08-11 22:58:11 +02:00
parent 9f73893e14
commit 766fe6aebe
@@ -1,4 +1,6 @@
using System.Net.Sockets;
using System.Security.Cryptography; using System.Security.Cryptography;
using System.Text;
using DotNet.Testcontainers.Builders; using DotNet.Testcontainers.Builders;
using DotNet.Testcontainers.Containers; using DotNet.Testcontainers.Containers;
using Xunit; using Xunit;
@@ -40,6 +42,21 @@ public sealed class SshServerFixture : IAsyncLifetime
private const int SshPort = 2222; private const int SshPort = 2222;
/// <summary>
/// How many connections in a row the server has to answer before this fixture calls it ready.
/// </summary>
/// <remarks>
/// Twenty-five, and the number is measured rather than picked. Probing a fresh container 200 times with
/// penalties left at the image's default, the first <c>Not allowed at this time</c> came back at probe
/// 18 and 183 of the 200 were refused; with <c>PerSourcePenalties no</c> applied, none of 200 were. Ten
/// was tried first and is useless — it sits below the threshold, so the guard passed happily against a
/// server that was still penalising. See <see cref="WaitUntilServingAsync"/>.
/// </remarks>
private const int RequiredStreak = 25;
/// <summary>How long to keep trying before giving up on the server entirely.</summary>
private static readonly TimeSpan ReadyTimeout = TimeSpan.FromSeconds(60);
private readonly SemaphoreSlim sftpGate = new(1, 1); private readonly SemaphoreSlim sftpGate = new(1, 1);
private IContainer? container; private IContainer? container;
@@ -80,101 +97,94 @@ public sealed class SshServerFixture : IAsyncLifetime
.Build(); .Build();
await container.StartAsync(); await container.StartAsync();
await AllowTcpForwardingAsync(); await ReconfigureAsync();
await WaitUntilServingAsync();
} }
/// <summary> /// <summary>
/// Lets this server open the direct-tcpip channels a forward is made of. /// Turns off the hardening this suite trips over, and makes the running server re-read its config.
/// </summary> /// </summary>
/// <remarks> /// <remarks>
/// <para> /// <para>
/// ◆ <b>The image ships <c>AllowTcpForwarding no</c>, and nothing says so at the point it bites.</b> A /// ◆ <b><c>PerSourcePenalties no</c> is the fix for the flake this suite had for months, and the other
/// dynamic forward starts perfectly happily — it is a local listener, and opening it asks the server /// two settings here are not.</b> OpenSSH 9.8 added per-source penalties and 10.x has them on by
/// nothing — and then every connection through it is refused when the channel is opened. SSH.NET /// default; this image runs 10.3. A source address that keeps disconnecting without authenticating is
/// reports that as <c>SOCKS5: General failure</c> from the proxy, which names neither the server nor /// penalised, and while the penalty holds every connection from it is answered with the clear-text line
/// the setting, and is what the first run of <c>LoopbackProxyTests</c> collected. /// <c>Not allowed at this time</c> and then closed.
/// </para> /// </para>
/// <para> /// <para>
/// Patched after start rather than baked in, because the image's entrypoint writes its configuration /// <b>This suite generates exactly that traffic, by design.</b> This client's first contact with an
/// itself on every boot — a mounted file would be overwritten before sshd read it. sshd re-reads on /// unknown host is a connection deliberately refused at the host key — which is a disconnect with no
/// <c>SIGHUP</c> and applies the result to connections made after that, and the readiness wait has /// authentication attempt — and several tests do nothing else:
/// already run, so nothing here races the boot. /// <c>RefusingTheHostKey_AbortsTheConnection</c>, <c>AnUntrustedHost_IsRefusedExactlyAsAShellWouldBe</c>,
/// and every helper that learns a host key by being turned away first. Enough of them close together and
/// sshd stops talking to the test host altogether, for a while, and then starts again.
/// </para>
/// <para>
/// From the client that is <c>SshConnectionException: The connection was closed by the remote host</c>
/// within milliseconds — no banner, nothing to say which of the many reasons it was. It hits whichever
/// class is running when the penalty lands and spares the rest, which is why it read as random and why
/// the class it hit lost <em>every</em> connection it made rather than a random few. The one test in that
/// class that expects a refusal passed throughout, for the wrong reason.
/// </para>
/// <para>
/// ◆ <b>Two earlier diagnoses were wrong, and are recorded here so they are not tried again.</b>
/// <c>MaxStartups</c> was blamed on the reasoning that xUnit runs test classes in parallel, so ten
/// unauthenticated connections would be in flight at once — but every class that touches this server
/// shares <see cref="SshCollection"/>, and xUnit's unit of parallelism is the collection, so they run one
/// after another and never have more than a connection or two open. The reload window was blamed next,
/// and a wait for the banner to answer was written and removed as unproven; it was unproven because the
/// banner does answer, right up until the penalty lands.
/// </para>
/// <para>
/// The line is appended rather than replaced in place, unlike the two below it, because the image's
/// config does not mention the keyword at all — there is no line to replace, and sshd takes the first
/// value it finds for a keyword that appears more than once.
/// </para>
/// <para>
/// ◆ <b><c>AllowTcpForwarding</c> is what a dynamic forward needs</b>, and the image ships it off as
/// hardening. Without it a forward opens perfectly happily — a local listener asks the server nothing —
/// and then every connection through it is refused when the channel is opened. SSH.NET reports that as
/// <c>SOCKS5: General failure</c>, which names neither the server nor the setting, and is what the first
/// run of <c>LoopbackProxyTests</c> collected. That suite is also the alarm if this method ever silently
/// stops working.
/// </para>
/// <para>
/// <c>MaxStartups</c> is raised for the reason it should have been in the first place rather than as a
/// fix for anything: the compiled-in default refuses connections at random past ten unauthenticated ones
/// in flight, and a throttle is hardening a test server has no business reproducing. It is kept, not
/// because it was ever shown to matter here, but because removing it would be a second change riding
/// along with this one.
/// </para>
/// <para>
/// Both are replaced in place rather than appended, because sshd_config takes the <em>first</em> value
/// it finds for a keyword: an appended line would be dead the day the image ships an uncommented one of
/// its own.
/// </para> /// </para>
/// <para> /// <para>
/// ◆ <b><c>/config/sshd/sshd_config</c>, and there are two.</b> The image also carries /// ◆ <b><c>/config/sshd/sshd_config</c>, and there are two.</b> The image also carries
/// <c>/etc/ssh/sshd_config</c>, which looks like the file to patch, reads identically, and is not the /// <c>/etc/ssh/sshd_config</c>, which looks like the file to patch, reads identically, and is not the one
/// one the running server was started with — patching it changes the text and nothing else, which is a /// the running server was started with — patching it changes the text and nothing else, which is a fix
/// fix that appears to work and leaves the failure exactly where it was. Measured with <c>find</c> /// that appears to work and leaves the failure exactly where it was.
/// rather than assumed, after the first version of this method did precisely that.
/// </para> /// </para>
/// <para> /// <para>
/// It is on for the whole assembly rather than for the one test that needs it. Forwarding is off in /// ◆ <b>Patched after boot and reloaded, rather than injected before it — which was tried and does not
/// this image as hardening, not as a behaviour worth reproducing: nothing else here opens a channel of /// work.</b> This image family runs <c>/custom-cont-init.d</c> scripts, which look like the right hook
/// any kind, so allowing it changes what exactly one suite can do and what none of the others see. /// and are not: the container's own log puts <c>sshd is listening on port 2222</c> <em>before</em>
/// </para> /// <c>[custom-init] Files found, executing</c>, so a script there edits a file the running server has
/// <para> /// already read. It leaves a config that greps correctly and a server behaving as though it had never
/// ◆ <b><c>MaxStartups</c> is raised here too, against a flake this suite has and that this change is /// been touched — the same trap as the wrong file, one layer up. Measured from the log, after a version
/// a mitigation for rather than a proven cure.</b> The distinction is stated because the evidence /// of this fixture did exactly that and failed twenty-eight tests.
/// stops short of the claim, and a later reader deserves to know which.
/// </para>
/// <para>
/// What is established: sshd's compiled-in default is <c>10:30:100</c> — past ten
/// <em>unauthenticated</em> connections in flight it refuses new ones at random, thirty percent of the
/// time, rising to always at a hundred — and the image ships the line commented out, so that default
/// was what ran. xUnit runs test classes in parallel and most classes here open a connection, so ten
/// in flight is reachable in the opening seconds. A refused connection presents to the client as
/// <c>SshConnectionException: The connection was closed by the remote host</c> within tens of
/// milliseconds, on whichever test connects at the wrong moment — which is exactly the observed
/// failure, seen in CI and reproduced locally.
/// </para>
/// <para>
/// What is <em>not</em> established is that this limit is the only cause, because the flake rate could
/// not be measured reliably. On the development machine the identical unmodified suite ran 85/85 clean
/// and, an hour later, failed 13 runs out of 15 — Docker throughput on that host swings far enough to
/// swamp the effect being measured. Any before/after comparison taken there is noise, and two were,
/// before that was noticed.
/// </para>
/// <para>
/// It is committed anyway, on the narrower argument that it is right regardless: a connection throttle
/// is hardening this suite has no interest in reproducing. It exists to test an SSH client, not to
/// survive a rate limit, and a test server that drops connections at random is a bad test server
/// whether or not it is the cause of this particular flake.
/// </para>
/// <para>
/// <b>Not fixed by serialising the suite</b>, which would have hidden it and cost the parallelism, and
/// not by retrying the connect, which would have made the client's own reconnect behaviour untestable
/// by burying it in the fixture. The limit is a property of a hardened server that this suite has no
/// interest in reproducing — it exists to test an SSH client, not to survive a throttle.
/// </para>
/// <para>
/// Replaced in place rather than appended, because sshd_config takes the <em>first</em> value it finds
/// for a keyword: an appended line would be dead the day the image ships an uncommented one of its own.
/// </para>
/// <para>
/// ◆ <b>The reload window is the other candidate, and it is deliberately not guarded against.</b>
/// <c>SIGHUP</c> makes sshd close its listeners and re-execute itself, and <c>pkill</c> returns when
/// the signal is delivered rather than when that has finished — so in principle a connection made
/// immediately afterwards is refused, producing this same exception. A wait that opened connections
/// until the server answered with its banner three times running was written, and then removed: it
/// could not be shown to change anything either, and a fixture carrying two unproven fixes for one
/// symptom is worse than one, because the next person has to disprove both.
/// </para>
/// <para>
/// If this flake returns, that is the next thing to try. Two things to know before trying it: the two
/// causes are indistinguishable from the client, so a fix can only be judged by a repeat run and never
/// by whether the next run passes — and the repeat run has to happen somewhere with stable Docker
/// throughput, which the development machine is not. Better still, make sshd say why: raise its
/// <c>LogLevel</c> here, disable Ryuk so the container outlives the run, and read
/// <c>docker logs</c>. A <c>MaxStartups</c> refusal names itself there; a reload does not.
/// </para> /// </para>
/// </remarks> /// </remarks>
private async Task AllowTcpForwardingAsync() private async Task ReconfigureAsync()
{ {
var result = await container!.ExecAsync([ var result = await container!.ExecAsync([
"sh", "sh",
"-c", "-c",
"sed -i 's/^AllowTcpForwarding no/AllowTcpForwarding yes/' /config/sshd/sshd_config" "sed -i 's/^AllowTcpForwarding no/AllowTcpForwarding yes/' /config/sshd/sshd_config"
+ " && sed -i 's/^#*MaxStartups .*/MaxStartups 200/' /config/sshd/sshd_config" + " && sed -i 's/^#*MaxStartups .*/MaxStartups 200/' /config/sshd/sshd_config"
+ " && printf '\\nPerSourcePenalties no\\n' >> /config/sshd/sshd_config"
+ " && pkill -HUP sshd", + " && pkill -HUP sshd",
]); ]);
@@ -183,7 +193,107 @@ public sealed class SshServerFixture : IAsyncLifetime
throw new InvalidOperationException( throw new InvalidOperationException(
$"Could not reconfigure the test server: {result.Stderr}"); $"Could not reconfigure the test server: {result.Stderr}");
} }
}
/// <summary>
/// Blocks until the server answers <see cref="RequiredStreak"/> connections in a row with its banner.
/// </summary>
/// <remarks>
/// <para>
/// ◆ <b>This is a guard rather than a wait, and what it guards against is
/// <c>PerSourcePenalties</c> coming back.</b> Reconfiguring above turns it off; this proves it is off,
/// immediately and by name, instead of letting the suite discover it later as an unrelated-looking
/// failure in whichever class happened to be running.
/// </para>
/// <para>
/// <b>Consecutive, and deliberately with no pause between them.</b> Each probe opens a connection, reads
/// the identification string and disconnects without authenticating — which is exactly the shape of
/// connection <c>PerSourcePenalties</c> punishes, and exactly what this suite does all day: a first
/// contact with an unknown host is a connection this client deliberately refuses at the host key.
/// <see cref="RequiredStreak"/> back to back is therefore not a soak test, it is the specific
/// provocation, sized above the measured threshold on purpose, and it costs well under a second when the
/// setting is off.
/// </para>
/// <para>
/// It is also the one check that can tell a listening socket from a running server. The container's own
/// readiness — a log line and <c>netstat</c> showing <c>:2222</c> — passes on a container whose sshd has
/// gone: the socket is published by a host-side proxy that accepts before it has anything to forward to,
/// so a dead server presents as a connection accepted and closed rather than as one refused.
/// </para>
/// <para>
/// Probed from the host rather than with <c>docker exec</c>, deliberately: that is the path the tests
/// take, proxy included, and penalties are counted per source address — from inside the container the
/// source would be the loopback rather than the address every test connects from.
/// </para>
/// </remarks>
private async Task WaitUntilServingAsync()
{
// TimeProvider.System rather than DateTimeOffset.UtcNow, which this repository bans so that time can
// be faked — and rather than a fake, because what is being waited on is a real container starting.
var deadline = TimeProvider.System.GetUtcNow() + ReadyTimeout;
var streak = 0;
var last = "no probe ran";
while (streak < RequiredStreak)
{
if (TimeProvider.System.GetUtcNow() >= deadline)
{
throw new InvalidOperationException(
$"The test server did not answer {RequiredStreak} connections in a row within "
+ $"{ReadyTimeout}. The last probe said: {last}. If it says \"Not allowed at this "
+ "time\", sshd is penalising this source address and PerSourcePenalties is no longer "
+ "being turned off — see ReconfigureAsync.");
}
var (answered, what) = await ProbeAsync();
last = what;
if (answered)
{
streak++;
continue;
}
// Only pause when it is not working. Back-to-back probes are the point while they succeed;
// hammering a server that has not finished starting is just noise.
streak = 0;
await Task.Delay(TimeSpan.FromMilliseconds(200));
}
}
/// <summary>Opens a socket and reads far enough to see OpenSSH's identification string.</summary>
/// <remarks>
/// The description comes back with the answer because the interesting failures are not exceptions. A
/// penalised source is told <c>Not allowed at this time</c> in clear text before the socket closes, and
/// a suite that only knew "no banner" would have to go and find that out again — which is what happened
/// the first time, at some length.
/// </remarks>
private async Task<(bool Answered, string What)> ProbeAsync()
{
try
{
using var probe = new TcpClient();
using var timeout = new CancellationTokenSource(TimeSpan.FromSeconds(5));
await probe.ConnectAsync(Host, Port, timeout.Token);
var buffer = new byte[64];
var read = await probe.GetStream().ReadAtLeastAsync(
buffer, 4, throwOnEndOfStream: false, timeout.Token);
var answered = read >= 4 && "SSH-"u8.SequenceEqual(buffer.AsSpan(0, 4));
return (
answered,
answered
? "SSH-"
: $"{read} bytes: "
+ Encoding.ASCII.GetString(buffer, 0, Math.Max(read, 0)).ReplaceLineEndings(" "));
}
catch (Exception exception) when (exception is SocketException or OperationCanceledException or IOException)
{
return (false, $"{exception.GetType().Name}: {exception.Message}");
}
} }
/// <summary> /// <summary>
@@ -225,12 +335,17 @@ public sealed class SshServerFixture : IAsyncLifetime
/// </summary> /// </summary>
/// <remarks> /// <remarks>
/// <para> /// <para>
/// Shared rather than opened per test, and that is a limit of the server rather than an optimisation. /// Shared rather than opened per test. This was once explained as a way of staying under the server's
/// sshd's <c>MaxStartups</c> drops connections at random once enough are part-way through a handshake, /// <c>MaxStartups</c> throttle, on the belief that the suite ran its classes in parallel and made two
/// and this client's first contact with an unknown host is a connection deliberately <em>refused</em> at /// handshakes per test — this client's first contact with an unknown host is a connection deliberately
/// the host key so a suite that opened its own session per test made two handshakes per test and /// <em>refused</em> at the host key, so every session costs two. The parallelism was not real: every
/// pushed the whole assembly over the threshold. What that looks like is unrelated tests failing with /// class here shares one collection and xUnit runs collections, not classes, in parallel. See
/// "the connection was closed by the remote host", a different few each run. /// <see cref="WaitUntilServingAsync"/>, which is where that mistake was found and what the failure it
/// was blamed for turned out to be.
/// </para>
/// <para>
/// It stays shared regardless, on the plainer argument: one session is enough, and a handshake per test
/// would be seconds of the suite's runtime spent proving nothing this file has not already proved.
/// </para> /// </para>
/// <para> /// <para>
/// Safe to share because an SFTP session holds no per-test state: every test here works in a directory /// Safe to share because an SFTP session holds no per-test state: every test here works in a directory