diff --git a/CHANGES.md b/CHANGES.md
index 6153cefc..f8f9f229 100644
--- a/CHANGES.md
+++ b/CHANGES.md
@@ -1,5 +1,30 @@
# Changes
+## 11.5.0.0 13/09/2026
+
+* **What happened to the connection is readable from the type, instead of the message text** (the follow-up to #179). 11.4.0 made the behaviour correct - one owner per transition, an operation that was overtaken says so - but gave the caller no way to read that answer. `NotConnectedException` carried five different events and `OperationCanceledException` two, so the only way to tell "the consumer disconnected the client" from "this endpoint is not answering" was to classify by message text - which the release notes of 11.3.2.0 told consumers not to do, while the library gave them no type capable of it.
+ * six new exception types, all deriving from the ones thrown today, so no `catch` clause changes meaning and no task changes status: `ClientDisconnectedException` and `ReconnectExhaustedException` (attempts spent, budget configured), `RequestRefusedException` for a request the caller asked not to have wait, `ConnectHandlerFailedException` (how many times the handler failed, and the handler's own exception), `ConnectionClosedPermanentlyException` for a node that closed with a code this client does not retry after, and `NotConnectingException` for a client with no attempt in progress. `ConnectionSupersededException` derives from `OperationCanceledException` and names the transition that took over and where it left the client
+ * **a broken `OnConnected` handler is no longer reported as a disconnect the consumer performed.** The give-up path ends by calling `Disconnect()` itself, so every point reading the permanently-disconnected flag answered "the client has been disconnected" - for a client that is down because its own handler is broken, where the node is answering and failing over would leave a healthy server. The cause now travels with the disconnect, written in the same critical section as the flag it qualifies. The two existing tests on that path made the defect plain: the same broken handler produced one type when it failed immediately and another when it failed a moment later
+ * **a request swept while the connection moved says which transition swept it, and where the client went.** This is the failure consumers meet most often, and it arrived as a bare `OperationCanceledException` reading "Connection was intentionally closed", indistinguishable from a cancellation of their own. Sweeps caused by the connection failing on its own - a network drop, a close being processed - deliberately keep the plain cancellation: that choice is what keeps an ordinary network drop out of consumers' critical logs, and the reason now reaches them on the status stream instead
+ * `ConnectionStatusInfo.StopReason` says why the client stopped, and only when it stopped: the first handshake failure against a server that is not up is announced and then retried, so it names no reason - whether one is named is decided by the same condition that decides whether the retry happens. A consumer reading any reason as terminal would otherwise fail over to another server while this one was still being dialled.
+ * `ConnectionStatusInfo.StopReason` says why the client stopped. `Disconnected` is announced from ten places and they differed only in text, so "still trying" against "gave up" was derivable only from the absence of `ReconnectInfo` - which is also what a client that never had a loop looks like. The reason goes on the notification rather than into `ReconnectInfo`, so `Reconnect != null` keeps its one meaning
+ * `WaitForConnectionOutcomeAsync` answers "did it come back?" with a value rather than an exception, on `Connection`, on `XrplClient` and on `IXrplClient` - where the wait was previously unreachable except through the connection object. `ConnectionWaitOutcome` names the case rather than folding "timed out", "gave up" and "nothing is running" into one `false`. `HasConnectionAsync` is untouched: adding a `CancellationToken` overload beside it would make argument-less calls ambiguous at the call site (CS0121)
+ * **the wait no longer polls.** It slept 100 ms at a time, which is slower than the event it waits for and, on a single-threaded host such as Blazor WebAssembly, more expensive than it looks - browser timers are throttled in a hidden tab, so a wait the event would satisfy at once stretched into seconds. It now sleeps on a signal completed by the one funnel every connection state passes through. Every check it made per pass is unchanged, so the answers are identical; the unit suite runs 43 s to 30 s
+ * `ChangeServer` reads the network id the way `Connect` does. `Connect` has carried that read across a teardown since 11.4.0, because a socket really does open for a moment before a failing handler brings it down; `ChangeServer` read it once, directly, so a connection that needed a second attempt failed the switch
+ * `ConnectionManager` is fixed rather than left alone. The readiness signal above is deliberately not built on it - it releases waiters when a connection is retired, and a retirement has to carry a waiting request over to the new connection rather than fail it - but it is public, reachable as `client.connection.connectionManager`, and notified from nine places on the connection's own threads. Nothing inside the SDK awaits it, so its defects had never shown: a waiter resumed **inside** `ResolveAllAwaiting`, which is called from inside `OnceOpen` before the `OnConnected` handler, and a registration landing during a notification either threw "Collection was modified" or was dropped and never resumed. The list is guarded, waiters are released outside the lock and resume asynchronously, completions are `TrySet*`, and a cancellation is `TrySetCanceled` rather than a faulted task
+ * **`StopAfterMaxAttempts` now actually stops the client.** Found while writing a test for one of the paths above, and present since before this change: a client that spent its reconnect budget announced `Disconnected` and then ran a second full series from attempt #1, announcing it again. The loop's exit clears the two fields that say a sequence is running for this generation, which is exactly what "none is running" looks like, so the close of the attempt that failed last was indistinguishable from the close that began the whole thing - and, with the cancellation source already released, started a fresh sequence with the counter at zero. The generation that gave up is now recorded and refused a new loop. Keyed by generation rather than flagged, so nothing has to reset it: generations only increase, and a `Connect()` or `ChangeServer` begins a new one - which is when asking again is the consumer's decision. A client that stopped on its own still answers the consumer asking, and that is asserted alongside
+ * **the wait and the status stream cannot disagree about how the connection ended.** Three ways they could, all found by the cold review of this change and all in the new code. The wait asked "is anything in progress?" before it asked "did something end?", and after a sequence ends there is no socket, no cancellation source and a `Disconnected` state - which is also exactly what a client nobody has called `Connect()` on looks like: a consumer who heard `ReconnectExhausted` on the status stream and confirmed it on the wait before failing over was told `NotConnecting`, and the outcome this whole change exists to deliver was reachable only by a caller who happened to be parked already. A close code the client does not reconnect after - 1002, 1003, 1007, 1010 - was announced as `ClosedPermanently` and left the wait nothing to recognise, so a parked caller was woken, found nothing, parked again, and was told a whole acquisition timeout later that the connection "was not established in time"; `ConnectionWaitOutcome.ClosedPermanently` and `ConnectionClosedPermanentlyException` are its counterpart. And the exhaustion the loop records, the permanent close, and the generation all used `0` for "none", which is the generation of a brand-new client - so the fields answered their own question with "yes" until the entry check above happened to mask it
+ * **one ownership check instead of three and a missing one.** A status announcement that speaks for a single transition of the connection is only true while that transition still owns it, and the check sat at the call sites: three had it, one did not, and the missing one was invisible for as long as the deduplication happened to swallow what it let through. Making the stop reason a change in its own right stopped it swallowing. The check now lives inside the one funnel every announcement passes through, in the same critical section that publishes the state - which also closes the window where a caller read "the sequence is still running", lost the processor, and published a `RestoringConnection` after the loop had ended the sequence and announced the ending
+ * a wait whose timeout is an exact multiple of the system clock tick - 15.625 ms, so one second and the default five minutes both are - spun on the connection's lock for the rest of the tick instead of timing out, because the deadline comparison was strict while the remaining time was already zero
+ * **every reason the client can stop for now has one outcome, under the same name, and one record behind both.** Each ending used to leave its own residue for a caller to recognise - a flag for the consumer's own disconnect, a generation for a spent reconnect budget - so an ending that left none was invisible to everything except the status stream, and a caller parked in the wait sat out its whole timeout to be told the connection "was not established in time" about a connection it had already been told was over. Two such endings were found, in three places. `ConnectionStopReason` is now recorded where the announcement is published, by the one method every status passes through, and the wait, its outcome value and the check every request makes all read that one record - so the three ways of asking cannot come to know different things. A request issued after the endpoint gave up used to answer "no connection attempt in progress. Call Connect() first", the one distinction the exception family exists to draw. The correspondence is asserted over the enums themselves rather than a list, which is how the last mismatch was found: `ConnectionStopReason.UserDisconnected` had been paired with a `ConnectionWaitOutcome.Disconnected`, and the outcome is renamed to match
+ * **every value of both enums is asserted by a test, and one value was deleted for failing to be.** Coverage was measured rather than assumed, and three holes came out of it: two endings whose exception was pinned but whose outcome value was not, and `ConnectionStopReason.InitialConnectionFailed`, which no test named because nothing can produce it. It is unreachable by construction - the branch that would name it is reached only for a socket that never opened as a session, and such a socket is always retried, which is the same fact `willReconnect` reads - and that was measured, not argued: the branch was instrumented and the whole suite run twice, once on the reason and once on `willReconnect` itself, for zero hits in 1328 tests. A value no consumer can observe invites a branch that never runs, so it is gone before release rather than kept as a defensive one; adding an enum member later is not a breaking change, and removing one is. The ending a consumer gets for a connection that never came up and stopped being retried is `ReconnectExhausted`, from the loop that stopped retrying it
+ * a runnable sample of the whole thing, `Tests/TestsClients/ConnectionLifecycleSample`. It needs no wallet and no funded account, because the subject is the connection: it produces each ending in turn - nothing in progress, an endpoint that is down, a request with no connection, a switch to a node that is down and the recovery from it, a consumer disconnect, a broken `OnConnected` handler, and an operation another one overtook - printing the status stream each one produced and then the type the caller got, with one line per ending saying what a consumer is supposed to do about it. `dotnet run --project Tests/TestsClients/ConnectionLifecycleSample` against the CI stand, or pass any node's URL
+ * **a client that has given up no longer calls itself connected.** The give-up path announces the ending, then rejects the requests in flight - which is what resumes the caller inside `Connect()` - and only then disconnects, so between the second step and the third the socket is still installed and still open. The wait began with a fast `IsConnected()` check and preferred "connected" to any ending, so a consumer was told the connection was up one instant after being told why it was over. The fast path is gone, the endings are read whether or not a socket happens to be open, and a recorded ending outranks a socket on its way out; an established connection clears the record, which is the counterpart. In the same window the cause of the ending was not yet recorded, so the answer named a consumer disconnect - the one thing that had not happened - and it is now recorded before the ending is announced rather than after. Found by CI: the whole suite passed on the development machine, where the socket usually closed in time
+ * **`Connect()` no longer reports success off a socket that is on its way out.** It opened with a fast path on `IsConnected()` alone, and in the window above the socket is still open - so a consumer reacting to a notification whose message ends "Call Connect() to retry" was told they were already connected, and nothing was started. The fast path now asks what the wait asks, under the same lock: an open socket, no permanent disconnect, and no recorded ending. Found by CodeRabbit on the commit that fixed the defect above
+ * **the ownership guard reached the four announcements that still lacked it.** A review found one - the reconnect loop's own `catch`, which checks ownership a line above the announcement rather than in the same critical section, so a takeover landing between the two overwrote the winner's state with a `RestoringConnection` from a loop that had already been superseded. Rather than fix the instance, every announcement in the file was audited: the fast reconnect's two, the handler-failure path that has not given up yet, and the one reported. Announcements that speak for the client as a whole rather than one transition - `Connecting` from the takeover that just took it, the `Disconnect()` paths, `Connected` from a connection that just opened - carry no generation by design
+ * **the ownership guard is now the compiler's job.** After the audit above, every status announcement carried its generation, so the parameter became required rather than optional. This is the only change of the set that closes the class instead of an instance: five reviews found five sites of it on this branch, one at a time. It can no longer be omitted - only passed wrongly, which is a smaller mistake and a visible one at the call site. A `Connected` that really happened is exempt from the spent-sequence seal, because suppressing it would leave the ending standing over a client whose socket is carrying traffic
+ * pinned by 46 tests, each asserting a type or a value and never a message. Four of them exist because they failed first: `Task.WhenAll` does **not** lose the subtype unless a faulted task is alongside it, a readiness signal armed only on takeover leaves a waiter spinning after a close that took over nothing, and a retry filter that cannot tell the client's own teardown from a peer operation reports a different failure depending on timing
+
## 11.4.0.0 07/09/2026
* **A transition of the connection has one owner** (#179, the follow-up to #178). Every operation that moves the connection - `ChangeServer`, `Connect`, `Disconnect`, `DisconnectAndWaitAsync`, the health check's fast reconnect, the reconnect loop and the path taken when an `OnConnected` handler fails - used to decide for itself what happened to the socket, and two of them running at once were reconciled by `ReferenceEquals(ws, ...)` checks placed after whichever await somebody had noticed. #178 added three such checks and its review found the next window each time. The checks were right where they were; the pattern was what did not scale.
diff --git a/Tests/TestsClients/Blazor-WebAssembly/Pages/Index.razor b/Tests/TestsClients/Blazor-WebAssembly/Pages/Index.razor
index 808edad7..651f645b 100644
--- a/Tests/TestsClients/Blazor-WebAssembly/Pages/Index.razor
+++ b/Tests/TestsClients/Blazor-WebAssembly/Pages/Index.razor
@@ -547,8 +547,15 @@
CurrentConnectionState = statusInfo.ConnectionState;
IsConnected = statusInfo.ConnectionState == XrpConnectionState.Connected;
- AddStatusMessage($"[{statusInfo.ConnectionState}] {statusInfo.Message}", type);
- Console.WriteLine($"Connection Status: [{statusInfo.ConnectionState}] {statusInfo.Message}");
+ // The reason is shown only where it means something: on a notification that is not
+ // an ending it is None, and printing that on every line would bury the one place
+ // where "still trying" and "gave up" finally differ by more than their wording.
+ string stopped = statusInfo.StopReason == ConnectionStopReason.None
+ ? string.Empty
+ : $" [stopped: {statusInfo.StopReason}]";
+
+ AddStatusMessage($"[{statusInfo.ConnectionState}]{stopped} {statusInfo.Message}", type);
+ Console.WriteLine($"Connection Status: [{statusInfo.ConnectionState}]{stopped} {statusInfo.Message}");
StateHasChanged();
});
};
diff --git a/Tests/TestsClients/ConnectionLifecycleSample/ConnectionLifecycleSample.csproj b/Tests/TestsClients/ConnectionLifecycleSample/ConnectionLifecycleSample.csproj
new file mode 100644
index 00000000..80e7ddf9
--- /dev/null
+++ b/Tests/TestsClients/ConnectionLifecycleSample/ConnectionLifecycleSample.csproj
@@ -0,0 +1,16 @@
+
+
+
+ Exe
+ net10.0
+ enable
+ enable
+ latest
+ Xrpl.Samples.ConnectionLifecycle
+
+
+
+
+
+
+
diff --git a/Tests/TestsClients/ConnectionLifecycleSample/Program.cs b/Tests/TestsClients/ConnectionLifecycleSample/Program.cs
new file mode 100644
index 00000000..7e5d91f7
--- /dev/null
+++ b/Tests/TestsClients/ConnectionLifecycleSample/Program.cs
@@ -0,0 +1,420 @@
+using System.Net;
+using System.Net.Sockets;
+
+using Xrpl.Client;
+using Xrpl.Client.Exceptions;
+
+namespace Xrpl.Samples.ConnectionLifecycle;
+
+///
+/// Every way an XrplClient connection can end, and what a consumer is supposed to do about each.
+///
+///
+///
+/// Nothing here submits a transaction or needs a funded account: the subject is the connection
+/// itself. Run it against any node - a local standalone rippled is the easiest:
+///
+///
+/// docker compose -f .ci-config/docker-compose.ci.yml up -d
+/// dotnet run --project Tests/TestsClients/ConnectionLifecycleSample
+/// dotnet run --project Tests/TestsClients/ConnectionLifecycleSample -- wss://s.altnet.rippletest.net:51233
+///
+///
+/// The scenarios that need a server which is not answering use a port nothing is listening on, so
+/// only the one node is required.
+///
+///
+internal static class Program
+{
+ ///
+ /// The address the CI stand publishes on, and deliberately not localhost: that name
+ /// resolves to ::1 first on a dual-stack host while the container publishes on IPv4
+ /// only, so every connection pays a failed IPv6 attempt first. Harmless with the default
+ /// attempt timeout and fatal with the short ones below, which is a trap worth not shipping in
+ /// a sample.
+ ///
+ private static string _server = "ws://127.0.0.1:6006";
+
+ private static async Task Main(string[] args)
+ {
+ if (args.Length > 0)
+ {
+ _server = args[0];
+ }
+
+ Console.WriteLine($"Node under test: {_server}");
+ Console.WriteLine("Each scenario prints the status stream it produced, then the answer a caller gets.");
+
+ try
+ {
+ await NothingIsInProgress();
+ await TheEndpointIsNotAnswering();
+ await AConnectionThatWorks();
+ await ARequestWithNoConnection();
+ await ASwitchToAnEndpointThatIsDown();
+ await AConsumerDisconnect();
+ await ABrokenConnectHandler();
+ await AnOperationThatWasOvertaken();
+ }
+ catch (Exception unexpected)
+ {
+ Console.WriteLine();
+ Console.WriteLine($"The sample itself failed: {unexpected.GetType().Name}: {unexpected.Message}");
+ return 1;
+ }
+
+ Console.WriteLine();
+ Console.WriteLine("Done. Every ending above is a distinct type and a distinct ConnectionStopReason,");
+ Console.WriteLine("which is what lets a consumer react without reading message text.");
+ return 0;
+ }
+
+ ///
+ /// A client nobody has called Connect() on. Not an error state - an actionable one.
+ ///
+ private static async Task NothingIsInProgress()
+ {
+ Scenario("A client that was never told to connect");
+
+ XrplClient client = new XrplClient(UnusedEndpoint());
+
+ ConnectionWaitOutcome outcome =
+ await client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(1));
+
+ Console.WriteLine($" outcome: {outcome}");
+ Report(await Caught(() => client.connection.WaitForConnectionAsync(TimeSpan.FromSeconds(1))));
+ }
+
+ ///
+ /// The endpoint is not answering and the client has stopped trying. This is the one ending
+ /// that justifies failing over to another server.
+ ///
+ private static async Task TheEndpointIsNotAnswering()
+ {
+ Scenario("An endpoint that is down, with a reconnect budget that runs out");
+
+ // Without StopAfterMaxAttempts the loop keeps trying for ever and there is no ending to
+ // report - which is the right default for a long-lived client, and the wrong one for a
+ // client that is supposed to move to another node.
+ XrplClient client = new XrplClient(UnusedEndpoint(), new XrplClient.ClientOptions
+ {
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(200),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(400),
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ UseCustomPing = false,
+ });
+
+ Trace(client);
+ Report(await Caught(() => client.Connect()));
+
+ // The same answer from a caller arriving afterwards, and from a request. All three ways of
+ // asking read one record, so a consumer cannot get "this endpoint gave up" from one and
+ // "you never connected" from another.
+ Console.WriteLine($" a later caller: {await client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(1))}");
+ Report(await Caught(() => ServerInfo(client)));
+
+ await client.Disconnect();
+ }
+
+ private static async Task AConnectionThatWorks()
+ {
+ Scenario("A connection that comes up");
+
+ XrplClient client = await Connected();
+
+ Console.WriteLine($" outcome: {await client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(5))}");
+ Console.WriteLine($" IsConnected: {client.connection.IsConnected()}");
+
+ await client.Disconnect();
+ }
+
+ ///
+ /// What a request does while there is no connection is a policy, not an accident.
+ ///
+ private static async Task ARequestWithNoConnection()
+ {
+ Scenario("A request issued while the client is not connected");
+
+ XrplClient refuses = new XrplClient(UnusedEndpoint(), new XrplClient.ClientOptions
+ {
+ RequestPolicy = RequestFailurePolicy.ImmediateFail,
+ MaxReconnectAttempts = 1,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ UseCustomPing = false,
+ });
+
+ // Started and deliberately not awaited: the request has to be issued while an attempt is
+ // in progress, which is the case the policy is about.
+ Task connecting = Swallow(refuses.Connect());
+ await Task.Delay(300);
+
+ Console.WriteLine(" RequestFailurePolicy.ImmediateFail:");
+ Report(await Caught(() => ServerInfo(refuses)), indent: " ");
+
+ await refuses.Disconnect();
+ await connecting;
+
+ Console.WriteLine(" RequestFailurePolicy.WaitForConnection: the request waits for the connection instead,");
+ Console.WriteLine(" and fails with the same typed ending if the connection never arrives.");
+ }
+
+ ///
+ /// A switch to a node that is down leaves the client stopped - and still recoverable.
+ ///
+ private static async Task ASwitchToAnEndpointThatIsDown()
+ {
+ Scenario("Switching to an endpoint that is down, then coming back");
+
+ XrplClient client = await Connected(new XrplClient.ClientOptions
+ {
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(200),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(400),
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ UseCustomPing = false,
+ });
+
+ Trace(client);
+ Report(await Caught(() => client.connection.ChangeServer(UnusedEndpoint())));
+
+ // The client stays stopped until the consumer decides otherwise - that is what a terminal
+ // ending means - and deciding otherwise is one call.
+ Console.WriteLine(" recovering by switching back to a node that is up:");
+ await client.connection.ChangeServer(_server);
+ Console.WriteLine($" IsConnected: {client.connection.IsConnected()}");
+
+ await client.Disconnect();
+ }
+
+ private static async Task AConsumerDisconnect()
+ {
+ Scenario("The consumer disconnects the client");
+
+ XrplClient client = await Connected();
+ Trace(client);
+
+ await client.Disconnect();
+
+ Console.WriteLine($" outcome: {await client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(2))}");
+ Report(await Caught(() => ServerInfo(client)));
+
+ // Nothing is going to bring this connection back on its own, which is exactly why the
+ // ending has a type of its own: failing over here would be leaving a healthy node.
+ Console.WriteLine(" Connect() is the way out, and it works:");
+ await client.Connect();
+ Console.WriteLine($" IsConnected: {client.connection.IsConnected()}");
+ await client.Disconnect();
+ }
+
+ ///
+ /// The node is fine and this side is broken. Failing over would be leaving a healthy server.
+ ///
+ private static async Task ABrokenConnectHandler()
+ {
+ Scenario("An OnConnected handler that keeps throwing");
+
+ XrplClient client = new XrplClient(_server, new XrplClient.ClientOptions
+ {
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(200),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(400),
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ UseCustomPing = false,
+ });
+
+ client.connection.OnConnected += () =>
+ throw new InvalidOperationException("restoring subscriptions failed");
+
+ Trace(client);
+ Report(await Caught(() => client.Connect()));
+
+ await client.Disconnect();
+ }
+
+ ///
+ /// An operation another one overtook reports the winner rather than a bare cancellation.
+ ///
+ ///
+ /// The overtaking operation here is a second switch. A Disconnect() overtaking a switch
+ /// is the other half of the same rule and reports ClientDisconnectedException, because
+ /// the reaction a consumer owes it is the opposite one.
+ ///
+ private static async Task AnOperationThatWasOvertaken()
+ {
+ Scenario("A switch that another switch overtook");
+
+ XrplClient client = await Connected(new XrplClient.ClientOptions
+ {
+ MaxReconnectAttempts = 20,
+ ReconnectBaseDelay = TimeSpan.FromSeconds(2),
+ ReconnectMaxDelay = TimeSpan.FromSeconds(2),
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ Trace(client);
+
+ // Issued from the first switch's own status notification, which lands it inside that
+ // switch every time. A consumer meets this shape whenever a status handler reacts to the
+ // connection - "this node is not answering, go to the other one" is exactly such a
+ // handler.
+ bool overtaking = false;
+ client.connection.OnConnectionStatus += status =>
+ {
+ if (status.ConnectionState == XrpConnectionState.RestoringConnection && !overtaking)
+ {
+ overtaking = true;
+ _ = Swallow(client.connection.ChangeServer(_server));
+ }
+ };
+
+ Report(await Caught(() => client.connection.ChangeServer(UnusedEndpoint())));
+
+ Console.WriteLine($" the winner is where the client ended up - IsConnected: {client.connection.IsConnected()}");
+
+ // A Disconnect() that overtakes a switch reports ClientDisconnectedException instead: the
+ // client is down because it was asked to be, which calls for the opposite reaction, so the
+ // two cases do not share a type.
+ await client.Disconnect();
+ }
+
+ ///
+ /// The point of the whole exercise: one catch per ending, and a different reaction for
+ /// each. None of it reads message text.
+ ///
+ private static void Report(Exception? error, string indent = " ")
+ {
+ if (error is null)
+ {
+ Console.WriteLine($"{indent}completed without an exception");
+ return;
+ }
+
+ switch (error)
+ {
+ case ReconnectExhaustedException exhausted:
+ Console.WriteLine($"{indent}{nameof(ReconnectExhaustedException)} after {exhausted.Attempts} of {exhausted.MaxAttempts} attempts");
+ Console.WriteLine($"{indent} -> this endpoint is not answering. Fail over to another server.");
+ break;
+
+ case ConnectHandlerFailedException handler:
+ Console.WriteLine($"{indent}{nameof(ConnectHandlerFailedException)} after {handler.Failures} failure(s)");
+ Console.WriteLine($"{indent} caused by: {handler.InnerException?.GetType().Name}: {handler.InnerException?.Message}");
+ Console.WriteLine($"{indent} -> the node is fine and this side is broken. Fix the handler; do not fail over.");
+ break;
+
+ case ConnectionClosedPermanentlyException:
+ Console.WriteLine($"{indent}{nameof(ConnectionClosedPermanentlyException)}");
+ Console.WriteLine($"{indent} -> the node closed with a code that says retrying is pointless. Use another server.");
+ break;
+
+ case ClientDisconnectedException:
+ Console.WriteLine($"{indent}{nameof(ClientDisconnectedException)}");
+ Console.WriteLine($"{indent} -> the consumer took the client down. Only Connect() brings it back.");
+ break;
+
+ case NotConnectingException:
+ Console.WriteLine($"{indent}{nameof(NotConnectingException)}");
+ Console.WriteLine($"{indent} -> nothing is being attempted. Call Connect().");
+ break;
+
+ case RequestRefusedException:
+ Console.WriteLine($"{indent}{nameof(RequestRefusedException)}");
+ Console.WriteLine($"{indent} -> the connection was being rebuilt and this caller asked not to wait. Retry once connected.");
+ break;
+
+ case ConnectionSupersededException superseded:
+ Console.WriteLine($"{indent}{nameof(ConnectionSupersededException)}: overtaken by {superseded.Kind}");
+ Console.WriteLine($"{indent} -> a later operation owns the connection. Its result is the one that counts.");
+ break;
+
+ case NotConnectedException other:
+ Console.WriteLine($"{indent}{other.GetType().Name}: {other.Message}");
+ Console.WriteLine($"{indent} -> the base type still catches every one of the above.");
+ break;
+
+ case System.TimeoutException:
+ Console.WriteLine($"{indent}{nameof(System.TimeoutException)}");
+ Console.WriteLine($"{indent} -> it did not come up in the time allowed. Nothing has given up; waiting longer may still work.");
+ break;
+
+ default:
+ Console.WriteLine($"{indent}{error.GetType().Name}: {error.Message}");
+ break;
+ }
+ }
+
+ /// Prints the status stream, which carries the same endings as a value.
+ private static void Trace(XrplClient client)
+ {
+ client.connection.OnConnectionStatus += status =>
+ {
+ string stopped = status.StopReason == ConnectionStopReason.None
+ ? string.Empty
+ : $" [stopped: {status.StopReason}]";
+
+ Console.WriteLine($" status: {status.ConnectionState}{stopped} - {status.Message}");
+ };
+ }
+
+ /// The cheapest real request there is - it needs no account and no funds.
+ private static Task ServerInfo(XrplClient client) =>
+ client.connection.Request(new Dictionary { { "command", "server_info" } });
+
+ private static async Task Connected(XrplClient.ClientOptions? options = null)
+ {
+ XrplClient client = new XrplClient(
+ _server,
+ options ?? new XrplClient.ClientOptions { UseCustomPing = false });
+
+ await client.Connect();
+ return client;
+ }
+
+ private static async Task Caught(Func operation)
+ {
+ try
+ {
+ await operation();
+ return null;
+ }
+ catch (Exception error)
+ {
+ return error;
+ }
+ }
+
+ private static async Task Swallow(Task operation)
+ {
+ try
+ {
+ await operation;
+ }
+ catch
+ {
+ // The scenario reports through its own caller; this one exists only to be started.
+ }
+ }
+
+ /// A loopback port nothing is listening on, so "the server is down" needs no server.
+ private static string UnusedEndpoint()
+ {
+ TcpListener listener = new TcpListener(IPAddress.Loopback, port: 0);
+ listener.Start();
+ int port = ((IPEndPoint)listener.LocalEndpoint).Port;
+ listener.Stop();
+
+ return $"ws://127.0.0.1:{port}";
+ }
+
+ private static void Scenario(string title)
+ {
+ Console.WriteLine();
+ Console.WriteLine($"== {title}");
+ }
+}
diff --git a/Tests/TestsClients/ConnectionLifecycleSample/README.md b/Tests/TestsClients/ConnectionLifecycleSample/README.md
new file mode 100644
index 00000000..2ec33a1d
--- /dev/null
+++ b/Tests/TestsClients/ConnectionLifecycleSample/README.md
@@ -0,0 +1,84 @@
+# Connection lifecycle sample
+
+Every way an `XrplClient` connection can end, and what a consumer is supposed to do about each.
+
+Nothing here submits a transaction or needs a funded account — the subject is the connection
+itself. The sample runs eight scenarios in order, prints the status stream each one produced, and
+then prints the type the caller got and the one sentence that says how to react to it.
+
+## Run it
+
+One node is all it needs. The scenarios that require a server which is **not** answering bind a
+loopback port and release it, so they need nothing running.
+
+Against the CI stand (the default, `ws://127.0.0.1:6006`):
+
+```bash
+docker compose -f .ci-config/docker-compose.ci.yml up -d
+```
+
+```bash
+dotnet run --project Tests/TestsClients/ConnectionLifecycleSample
+```
+
+Against any other node — pass its URL:
+
+```bash
+dotnet run --project Tests/TestsClients/ConnectionLifecycleSample -- wss://s.altnet.rippletest.net:51233
+```
+
+The whole run takes under a minute: most of it is the reconnect backoff of the scenarios that are
+supposed to fail.
+
+Prefer `127.0.0.1` over `localhost` for a local node. On a dual-stack host `localhost` resolves to
+`::1` first while Docker publishes on IPv4 only, so every connection pays a failed IPv6 attempt —
+harmless with the default attempt timeout, fatal with the short ones this sample uses to stay
+quick.
+
+## Reading the output
+
+Two kinds of line. Indented four spaces is the status stream, exactly as `OnConnectionStatus`
+delivers it:
+
+```
+ status: RestoringConnection - Reconnecting in 0.4 seconds... (attempt #2)
+ status: Disconnected [stopped: ReconnectExhausted] - Reconnection stopped after 2 attempts.
+```
+
+`[stopped: ...]` appears only when the client has stopped. A `Disconnected` without it is a
+failure the client is about to retry, which is why the first handshake failure against a server
+that is down names no reason: naming one would tell a consumer to fail over while this node is
+still being dialled.
+
+Indented two spaces is what the caller got:
+
+```
+ ReconnectExhaustedException after 2 of 2 attempts
+ -> this endpoint is not answering. Fail over to another server.
+```
+
+## The scenarios
+
+| Scenario | Ending | What it is for |
+|---|---|---|
+| A client that was never told to connect | `NotConnectingException` | "Nothing is in progress" is an answer, not an error |
+| An endpoint that is down | `ReconnectExhaustedException` | The one ending that justifies failing over. Needs `StopAfterMaxAttempts` |
+| A connection that comes up | `ConnectionWaitOutcome.Connected` | The value-returning wait, with no `catch` |
+| A request with no connection | `RequestRefusedException` | `RequestFailurePolicy` is a decision, not an accident |
+| A switch to a node that is down | `ReconnectExhaustedException`, then recovery | A stopped client stays stopped, and one call brings it back |
+| The consumer disconnects | `ClientDisconnectedException` | Never fail over here — the client was asked to be down |
+| A broken `OnConnected` handler | `ConnectHandlerFailedException` | The node is fine and this side is broken. The handler's own exception is the `InnerException` |
+| A switch another switch overtook | `ConnectionSupersededException` | The overtaken operation names the winner instead of a bare cancellation. A `Disconnect()` overtaking a switch is the other half of the rule and reports `ClientDisconnectedException` — the opposite reaction, so not the same type |
+
+## What to take from it
+
+`Report` in `Program.cs` is the part worth copying: one `catch` per ending, each with a different
+reaction, and not one line of it reads message text. That is the whole point of the typed
+connection outcomes — before them all of these arrived as the same `NotConnectedException`, and
+telling them apart meant matching on strings that were free to change.
+
+The same endings are readable three ways, and all three agree:
+
+- as an exception from `WaitForConnectionAsync`, `Connect`, `ChangeServer` or a request;
+- as a value from `WaitForConnectionOutcomeAsync`;
+- as `ConnectionStatusInfo.StopReason` on the status stream.
diff --git a/Tests/Xrpl.Tests/Client/ClosesWithCodeServer.cs b/Tests/Xrpl.Tests/Client/ClosesWithCodeServer.cs
new file mode 100644
index 00000000..2113538b
--- /dev/null
+++ b/Tests/Xrpl.Tests/Client/ClosesWithCodeServer.cs
@@ -0,0 +1,78 @@
+using System.Net.Sockets;
+using System.Text;
+using System.Text.Json;
+using System.Threading;
+using System.Threading.Tasks;
+
+namespace Xrpl.Tests
+{
+ ///
+ /// WebSocket server that answers requests normally until told to stop, then closes the
+ /// connection with a fixed status code.
+ ///
+ ///
+ /// The close codes the client does not reconnect after - 1002, 1003, 1007, 1010 - are the only
+ /// way a connection ends with neither a consumer Disconnect() nor a reconnect sequence
+ /// behind it, and no other test server can produce one: the shared mock never closes, and
+ /// CloseAfterHandshakeServer sends a close frame carrying no code at all, which the
+ /// client reads as "reconnect".
+ ///
+ internal sealed class ClosesWithCodeServer : WebSocketTestServerBase
+ {
+ private const string ServerInfoEnvelope =
+ "{\"id\":__ID__,\"status\":\"success\",\"type\":\"response\",\"result\":{\"info\":" +
+ "{\"build_version\":\"test-mock\",\"complete_ledgers\":\"1-1\",\"server_state\":\"full\"}}}";
+
+ private readonly int _closeCode;
+ private readonly SemaphoreSlim _closeNow = new SemaphoreSlim(initialCount: 0);
+
+ public ClosesWithCodeServer(int closeCode)
+ {
+ _closeCode = closeCode;
+ StartAccepting();
+ }
+
+ protected override bool ServesManyClients => true;
+
+ /// Releases the serving loop to send the close frame.
+ public void CloseNow() => _closeNow.Release();
+
+ protected override async Task ServeAsync(NetworkStream stream)
+ {
+ Task closeRequested = _closeNow.WaitAsync(Token);
+
+ while (!Token.IsCancellationRequested)
+ {
+ Task nextRequest = ReadTextFrameAsync(stream);
+ Task first = await Task.WhenAny(nextRequest, closeRequested).ConfigureAwait(false);
+
+ if (first == closeRequested)
+ {
+ // FIN + close opcode, two payload bytes, the code big-endian.
+ byte[] close =
+ {
+ 0x88, 0x02, (byte)(_closeCode >> 8), (byte)(_closeCode & 0xFF),
+ };
+
+ await stream.WriteAsync(close, Token).ConfigureAwait(false);
+ await stream.FlushAsync(Token).ConfigureAwait(false);
+ return;
+ }
+
+ string? request = await nextRequest.ConfigureAwait(false);
+ if (request == null)
+ {
+ return;
+ }
+
+ using JsonDocument document = JsonDocument.Parse(request);
+ string id = document.RootElement.TryGetProperty("id", out JsonElement requestId)
+ ? requestId.GetRawText()
+ : "null";
+
+ byte[] response = Encoding.UTF8.GetBytes(ServerInfoEnvelope.Replace("__ID__", id));
+ await WriteFragmentedMessageAsync(stream, response, fragments: 1).ConfigureAwait(false);
+ }
+ }
+ }
+}
diff --git a/Tests/Xrpl.Tests/Client/DropsFirstServerInfoServer.cs b/Tests/Xrpl.Tests/Client/DropsFirstServerInfoServer.cs
new file mode 100644
index 00000000..bba59272
--- /dev/null
+++ b/Tests/Xrpl.Tests/Client/DropsFirstServerInfoServer.cs
@@ -0,0 +1,74 @@
+using System.Net.Sockets;
+using System.Text;
+using System.Text.Json;
+using System.Threading;
+using System.Threading.Tasks;
+
+namespace Xrpl.Tests
+{
+ ///
+ /// WebSocket server that drops the connection the first time it is asked for
+ /// server_info, and behaves normally on every connection after that.
+ ///
+ ///
+ ///
+ /// The handshake succeeds, so a client reaches the point where it believes it is connected and
+ /// sends its first request - and only then does the connection go away. That is the window in
+ /// which reading the network id happens: a caller that reads it once, directly, fails the whole
+ /// operation, while one that carries the read across a teardown recovers and completes.
+ ///
+ ///
+ /// The shared mock cannot produce this: it answers every request on the connection it accepted,
+ /// so the teardown never lands between "connected" and "first answer". Dropping the first
+ /// request only, rather than every one, is what makes the recovery observable instead of just
+ /// the failure.
+ ///
+ ///
+ internal sealed class DropsFirstServerInfoServer : WebSocketTestServerBase
+ {
+ private const string ServerInfoEnvelope =
+ "{\"id\":__ID__,\"status\":\"success\",\"type\":\"response\",\"result\":{\"info\":" +
+ "{\"build_version\":\"test-mock\",\"complete_ledgers\":\"1-1\",\"server_state\":\"full\"}}}";
+
+ private int _dropsLeft = 1;
+
+ public DropsFirstServerInfoServer()
+ {
+ StartAccepting();
+ }
+
+ /// The client reconnects after the drop, so the next connection has to be served.
+ protected override bool ServesManyClients => true;
+
+ protected override async Task ServeAsync(NetworkStream stream)
+ {
+ while (!Token.IsCancellationRequested)
+ {
+ string request = await ReadTextFrameAsync(stream).ConfigureAwait(false);
+ if (request == null)
+ {
+ return;
+ }
+
+ using JsonDocument document = JsonDocument.Parse(request);
+ string command = document.RootElement.TryGetProperty("command", out JsonElement value)
+ ? value.GetString()
+ : null;
+
+ if (command == "server_info" && Interlocked.Decrement(ref _dropsLeft) >= 0)
+ {
+ // Returning closes the socket without answering: the request the client is
+ // waiting for dies with the connection.
+ return;
+ }
+
+ string id = document.RootElement.TryGetProperty("id", out JsonElement requestId)
+ ? requestId.GetRawText()
+ : "null";
+
+ byte[] response = Encoding.UTF8.GetBytes(ServerInfoEnvelope.Replace("__ID__", id));
+ await WriteFragmentedMessageAsync(stream, response, fragments: 1).ConfigureAwait(false);
+ }
+ }
+ }
+}
diff --git a/Tests/Xrpl.Tests/Client/Exceptions/TestUConnectionOutcomeTypes.cs b/Tests/Xrpl.Tests/Client/Exceptions/TestUConnectionOutcomeTypes.cs
new file mode 100644
index 00000000..b64ff2f0
--- /dev/null
+++ b/Tests/Xrpl.Tests/Client/Exceptions/TestUConnectionOutcomeTypes.cs
@@ -0,0 +1,242 @@
+using Microsoft.VisualStudio.TestTools.UnitTesting;
+
+using System;
+using System.Linq;
+using System.Threading;
+using System.Threading.Tasks;
+
+using Xrpl.Client;
+using Xrpl.Client.Exceptions;
+
+namespace Xrpl.Tests.Client.Exceptions
+{
+ ///
+ /// The types that say what happened to the connection, rather than leaving it in the message
+ /// text - see specs/2026-09-09-connection-outcome-api.md.
+ ///
+ ///
+ ///
+ /// carried five different events and
+ /// two, so a consumer that had to react differently to
+ /// each was left classifying by text - which the release notes of 11.3.2.0 told them not to do,
+ /// while the library gave them no type capable of it.
+ ///
+ ///
+ /// These tests are about the types themselves: what they derive from, and what they carry.
+ /// Which path produces which type is a question about the connection, and lives in
+ /// TestUConnectionOutcomes.
+ ///
+ ///
+ [TestClass]
+ public class TestUConnectionOutcomeTypes
+ {
+ ///
+ /// The failure that caused a refusal survives on the exception that reports it.
+ ///
+ ///
+ /// had a single message constructor, so a subtype that
+ /// has a cause - the handler exception behind
+ /// - had nowhere to put it: there is no setter
+ /// for . The whole taxonomy rests on this one
+ /// constructor existing.
+ ///
+ [TestMethod]
+ public void TestUNotConnectedExceptionKeepsTheFailureThatCausedIt()
+ {
+ InvalidOperationException cause = new InvalidOperationException("the handler threw");
+
+ NotConnectedException error = new NotConnectedException("gave up connecting", cause);
+
+ Assert.AreSame(cause, error.InnerException);
+ Assert.AreEqual("gave up connecting", error.Message);
+ }
+
+ ///
+ /// Every new type is still a .
+ ///
+ ///
+ /// This is what makes the change additive rather than breaking: the existing
+ /// catch (NotConnectedException) in consumer code, and the ones inside this library,
+ /// keep catching all five events. Only code that wants to tell them apart has to look.
+ ///
+ [TestMethod]
+ public void TestUEveryOutcomeIsStillANotConnectedException()
+ {
+ Assert.IsInstanceOfType(new ClientDisconnectedException("x"));
+ Assert.IsInstanceOfType(new ReconnectExhaustedException("x", attempts: 1, maxAttempts: 1));
+ Assert.IsInstanceOfType(new RequestRefusedException("x"));
+ Assert.IsInstanceOfType(new ConnectHandlerFailedException("x", failures: 1));
+ Assert.IsInstanceOfType(new NotConnectingException("x"));
+ }
+
+ ///
+ /// Exhaustion says how many attempts were spent and what the budget was.
+ ///
+ ///
+ /// Both numbers are reported, not just the limit, because a consumer deciding whether to
+ /// fail over wants to know the budget it actually configured was really spent. The raw
+ /// counter in the connection stands at MaxAttempts + 1 when the loop stops - it is
+ /// incremented at the head of a pass and the loop breaks on the pass that exceeds the
+ /// budget - so "6 of 5" is what a naive reading would publish. The throw site passes the
+ /// number of attempts made.
+ ///
+ [TestMethod]
+ public void TestUExhaustionCarriesTheAttemptsItSpent()
+ {
+ ReconnectExhaustedException error = new ReconnectExhaustedException(
+ "Connection failed permanently after 5 attempts. Reconnection has been stopped.",
+ attempts: 5,
+ maxAttempts: 5);
+
+ Assert.AreEqual(5, error.Attempts);
+ Assert.AreEqual(5, error.MaxAttempts);
+ }
+
+ ///
+ /// A failing OnConnected handler reports how often it failed, and with what.
+ ///
+ ///
+ /// This is the one outcome where the node is answering and the fault is on this side, so a
+ /// consumer that reacts to it by failing over to another server would be moving away from
+ /// a healthy node. The handler's own exception is what says why, and it used to reach the
+ /// caller only as text inside the message.
+ ///
+ [TestMethod]
+ public void TestUAFailingConnectHandlerCarriesItsCountAndItsCause()
+ {
+ InvalidOperationException cause = new InvalidOperationException("subscription refused");
+
+ ConnectHandlerFailedException error = new ConnectHandlerFailedException(
+ "Gave up connecting: the OnConnected handler failed 3 time(s) in a row.",
+ failures: 3,
+ innerException: cause);
+
+ Assert.AreEqual(3, error.Failures);
+ Assert.AreSame(cause, error.InnerException);
+ }
+
+ ///
+ /// Being overtaken is still a cancellation, and now says by what and where it went.
+ ///
+ ///
+ /// The four supersession sites already threw , so
+ /// deriving from it keeps every catch and every task status exactly as they were.
+ /// What the caller could not do before is tell "another operation took the connection over"
+ /// apart from "my own token was cancelled".
+ ///
+ [TestMethod]
+ public void TestUSupersessionIsACancellationThatNamesItsWinner()
+ {
+ ConnectionSupersededException error = new ConnectionSupersededException(
+ "Superseded by a later ChangeServer to wss://example.test:6006.",
+ ConnectionTransitionKind.ChangeServer,
+ supersededBy: "wss://example.test:6006");
+
+ Assert.IsInstanceOfType(error, "catch (OperationCanceledException) must keep catching this.");
+ Assert.AreEqual(ConnectionTransitionKind.ChangeServer, error.Kind);
+ Assert.AreEqual("wss://example.test:6006", error.SupersededBy);
+ }
+
+ ///
+ /// It carries no cancellation token, because nobody cancelled anything.
+ ///
+ ///
+ /// A caller that filters its own cancellations with
+ /// catch (OperationCanceledException ex) when (ex.CancellationToken == myToken)
+ /// must not have that filter match here: the operation was overtaken by another operation,
+ /// not cancelled by the caller. This is the reason the constructor does not take a token.
+ ///
+ [TestMethod]
+ public void TestUSupersessionCarriesNobodysToken()
+ {
+ using CancellationTokenSource callerToken = new CancellationTokenSource();
+
+ ConnectionSupersededException error = new ConnectionSupersededException(
+ "Superseded by a later Connect().",
+ ConnectionTransitionKind.Connect);
+
+ Assert.AreNotEqual(callerToken.Token, error.CancellationToken);
+ Assert.AreEqual(CancellationToken.None, error.CancellationToken);
+ Assert.IsNull(error.SupersededBy, "A Connect() that won left the client where it already was.");
+ }
+
+ ///
+ /// Awaited directly, the subtype reaches the caller - and the task reports itself cancelled.
+ ///
+ ///
+ /// AsyncTaskMethodBuilder turns an escaping
+ /// into a cancelled task whatever its token says, so this is the behaviour the supersession
+ /// sites already had; the test pins that inheritance did not change it.
+ ///
+ [TestMethod]
+ public async Task TestUSupersessionSurvivesADirectAwait()
+ {
+ Task direct = SupersededAsync();
+
+ ConnectionSupersededException caught =
+ await Assert.ThrowsExactlyAsync(async () => await direct);
+
+ Assert.AreEqual(ConnectionTransitionKind.Reconnect, caught.Kind);
+ Assert.AreEqual(
+ TaskStatus.Canceled,
+ direct.Status,
+ "An OperationCanceledException leaving an async method cancels its task.");
+ }
+
+ ///
+ /// Through Task.WhenAll the subtype survives too, as long as nothing faulted.
+ ///
+ ///
+ /// Measured rather than assumed, and the measurement contradicted the expectation this test
+ /// was written from: WhenAll with cancellations and no faults stores the first
+ /// cancellation exception and rethrows that very instance, so
+ /// is still readable. Documenting the
+ /// pessimistic rule would have sent consumers looking for a workaround they do not need.
+ ///
+ [TestMethod]
+ public async Task TestUSupersessionSurvivesWhenAllWithNoFaults()
+ {
+ Task combined = Task.WhenAll(SupersededAsync(), Task.CompletedTask);
+
+ ConnectionSupersededException caught =
+ await Assert.ThrowsExactlyAsync(async () => await combined);
+
+ Assert.AreEqual(ConnectionTransitionKind.Reconnect, caught.Kind);
+ }
+
+ ///
+ /// A fault alongside it loses the supersession entirely.
+ ///
+ ///
+ /// WhenAll prefers faults to cancellations: with any faulted task the combined task
+ /// faults, and - measured here rather than assumed - the cancellation is not recorded at
+ /// all, so it is absent from Task.Exception.InnerExceptions as well as from the
+ /// await. This is the one case where a caller cannot learn it was overtaken, and it
+ /// is why the XML documentation of the type says to await the operation itself.
+ ///
+ [TestMethod]
+ public async Task TestUAFaultAlongsideItLosesTheSupersession()
+ {
+ static async Task FaultedAsync()
+ {
+ await Task.Yield();
+ throw new InvalidOperationException("something else went wrong");
+ }
+
+ Task combined = Task.WhenAll(SupersededAsync(), FaultedAsync());
+
+ await Assert.ThrowsExactlyAsync(async () => await combined);
+
+ Assert.AreEqual(TaskStatus.Faulted, combined.Status);
+ Assert.IsFalse(
+ combined.Exception!.InnerExceptions.Any(e => e is ConnectionSupersededException),
+ "WhenAll records faults only: a cancellation alongside a fault is dropped, not merely hidden.");
+ }
+
+ private static async Task SupersededAsync()
+ {
+ await Task.Yield();
+ throw new ConnectionSupersededException("overtaken", ConnectionTransitionKind.Reconnect);
+ }
+ }
+}
diff --git a/Tests/Xrpl.Tests/Client/SilentOnPingAndLedgerServer.cs b/Tests/Xrpl.Tests/Client/SilentOnPingAndLedgerServer.cs
new file mode 100644
index 00000000..cad1cdb7
--- /dev/null
+++ b/Tests/Xrpl.Tests/Client/SilentOnPingAndLedgerServer.cs
@@ -0,0 +1,63 @@
+using System.Net.Sockets;
+using System.Text;
+using System.Text.Json;
+using System.Threading.Tasks;
+
+namespace Xrpl.Tests
+{
+ ///
+ /// WebSocket server that answers neither ping nor ledger, and answers everything
+ /// else with the same server_info body.
+ ///
+ ///
+ /// Two silences, for two halves of one scenario. Not answering ping is what drives the
+ /// health check to declare the connection dead and hand it to the fast-reconnect path - the
+ /// shared mock answers pings itself, so its activity clock never runs out. Not answering
+ /// ledger is what keeps a request in flight while that happens, so the sweep the
+ /// reconnect performs has something to sweep. Answering everything else is what keeps the
+ /// connection up long enough for either to matter.
+ ///
+ internal sealed class SilentOnPingAndLedgerServer : WebSocketTestServerBase
+ {
+ private const string ServerInfoEnvelope =
+ "{\"id\":__ID__,\"status\":\"success\",\"type\":\"response\",\"result\":{\"info\":" +
+ "{\"build_version\":\"test-mock\",\"complete_ledgers\":\"1-1\",\"server_state\":\"full\"}}}";
+
+ public SilentOnPingAndLedgerServer()
+ {
+ StartAccepting();
+ }
+
+ /// The client reconnects when the health check gives up on the silence.
+ protected override bool ServesManyClients => true;
+
+ protected override async Task ServeAsync(NetworkStream stream)
+ {
+ while (!Token.IsCancellationRequested)
+ {
+ string request = await ReadTextFrameAsync(stream).ConfigureAwait(false);
+ if (request == null)
+ {
+ return;
+ }
+
+ using JsonDocument document = JsonDocument.Parse(request);
+ string command = document.RootElement.TryGetProperty("command", out JsonElement value)
+ ? value.GetString()
+ : null;
+
+ if (command == "ping" || command == "ledger")
+ {
+ continue;
+ }
+
+ string id = document.RootElement.TryGetProperty("id", out JsonElement requestId)
+ ? requestId.GetRawText()
+ : "null";
+
+ byte[] response = Encoding.UTF8.GetBytes(ServerInfoEnvelope.Replace("__ID__", id));
+ await WriteFragmentedMessageAsync(stream, response, fragments: 1).ConfigureAwait(false);
+ }
+ }
+ }
+}
diff --git a/Tests/Xrpl.Tests/Client/TestUConnectionManagerWaiters.cs b/Tests/Xrpl.Tests/Client/TestUConnectionManagerWaiters.cs
new file mode 100644
index 00000000..4ceb21e8
--- /dev/null
+++ b/Tests/Xrpl.Tests/Client/TestUConnectionManagerWaiters.cs
@@ -0,0 +1,145 @@
+using Microsoft.VisualStudio.TestTools.UnitTesting;
+
+using System;
+using System.Collections.Generic;
+using System.Threading;
+using System.Threading.Tasks;
+
+using Xrpl.Client;
+
+namespace Xrpl.Tests
+{
+ ///
+ /// as something a consumer can actually call.
+ ///
+ ///
+ ///
+ /// It is reachable from outside the library - client.connection.connectionManager is a
+ /// public field on a public type - and the connection notifies it from nine places, on its own
+ /// threads. Nothing inside the SDK awaits it, so its defects never showed; that makes them
+ /// latent, not absent, and the first consumer to call AwaitConnection() would meet all
+ /// of them at once.
+ ///
+ ///
+ /// These are unit tests of the class alone: no socket, no client. What they pin is the
+ /// behaviour any waiter is entitled to - resume off the notifying thread, survive a concurrent
+ /// notification, and be cancelled rather than faulted.
+ ///
+ ///
+ [TestClass]
+ public class TestUConnectionManagerWaiters
+ {
+ ///
+ /// A waiter resumes after the notification returns, not inside it.
+ ///
+ ///
+ /// The completion sources were created without
+ /// , so every parked waiter
+ /// resumed synchronously inside - which
+ /// the connection calls from inside OnceOpen, before the OnConnected handler
+ /// and while it holds the connection together. Consumer code running there is the defect
+ /// issue #177 was about, in a new place.
+ ///
+ [TestMethod]
+ public async Task TestUAWaiterDoesNotResumeInsideTheNotification()
+ {
+ ConnectionManager manager = new ConnectionManager();
+
+ int resumed = 0;
+ Task waiter = ResumeMarker();
+
+ async Task ResumeMarker()
+ {
+ await manager.AwaitConnection();
+ Volatile.Write(ref resumed, 1);
+ }
+
+ manager.ResolveAllAwaiting();
+ int resumedInsideTheCall = Volatile.Read(ref resumed);
+
+ await waiter;
+
+ Assert.AreEqual(
+ 0,
+ resumedInsideTheCall,
+ "The waiter resumed inside ResolveAllAwaiting - which the connection calls from inside OnceOpen.");
+ Assert.AreEqual(1, Volatile.Read(ref resumed), "It still has to resume, just not there.");
+ }
+
+ ///
+ /// A cancelled wait is cancelled, not faulted.
+ ///
+ ///
+ /// SetException(new OperationCanceledException(...)) produces a faulted task: the
+ /// rule that turns a cancellation into a cancelled task belongs to the async method
+ /// builder, not to . The difference is visible
+ /// to anyone reading Task.Status, combining the task with others, or leaving it
+ /// unobserved.
+ ///
+ [TestMethod]
+ public async Task TestUACancelledWaitIsCancelledRatherThanFaulted()
+ {
+ ConnectionManager manager = new ConnectionManager();
+
+ Task waiter = manager.AwaitConnection();
+ manager.RejectAllAwaitingWithCancellation();
+
+ await Assert.ThrowsAsync(async () => await waiter);
+ Assert.AreEqual(TaskStatus.Canceled, waiter.Status);
+ }
+
+ ///
+ /// Registering while a notification is going out does not corrupt the list or lose a waiter.
+ ///
+ ///
+ /// The waiters lived in a plain List<T> mutated without synchronisation, while
+ /// the connection notifies from its socket, reconnect and caller threads. Both failures are
+ /// real: an InvalidOperationException from a list modified while being iterated, and
+ /// a waiter added in that window and dropped by the reassignment that follows - which is the
+ /// worse of the two, because it never comes back and never says anything.
+ ///
+ [TestMethod]
+ public async Task TestURegisteringDuringANotificationLosesNobody()
+ {
+ for (int round = 0; round < 200; round++)
+ {
+ ConnectionManager manager = new ConnectionManager();
+ List waiters = new List();
+
+ Task registering = Task.Run(() =>
+ {
+ for (int i = 0; i < 20; i++)
+ {
+ lock (waiters)
+ {
+ waiters.Add(manager.AwaitConnection());
+ }
+ }
+ });
+
+ Task notifying = Task.Run(() =>
+ {
+ for (int i = 0; i < 20; i++)
+ {
+ manager.ResolveAllAwaiting();
+ }
+ });
+
+ await Task.WhenAll(registering, notifying);
+
+ // Whatever the interleaving was, everyone still parked is released by this one.
+ manager.ResolveAllAwaiting();
+
+ Task all;
+ lock (waiters)
+ {
+ all = Task.WhenAll(waiters);
+ }
+
+ Task finished = await Task.WhenAny(all, Task.Delay(TimeSpan.FromSeconds(5)));
+ Assert.AreSame(all, finished, $"A waiter was dropped and never resumed (round {round}).");
+ await all;
+ }
+ }
+ }
+}
diff --git a/Tests/Xrpl.Tests/Client/TestUConnectionOutcomes.cs b/Tests/Xrpl.Tests/Client/TestUConnectionOutcomes.cs
new file mode 100644
index 00000000..65440a91
--- /dev/null
+++ b/Tests/Xrpl.Tests/Client/TestUConnectionOutcomes.cs
@@ -0,0 +1,2112 @@
+using Microsoft.VisualStudio.TestTools.UnitTesting;
+
+using System;
+using System.Collections.Generic;
+using System.Diagnostics;
+using System.Threading;
+using System.Threading.Tasks;
+
+using Xrpl.Client;
+using Xrpl.Client.Exceptions;
+
+namespace Xrpl.Tests
+{
+ ///
+ /// Which path produces which outcome - see specs/2026-09-09-connection-outcome-api.md.
+ ///
+ ///
+ ///
+ /// The types themselves are covered in TestUConnectionOutcomeTypes. What is pinned here
+ /// is the mapping: an event of the connection, and the type a caller gets for it. Every
+ /// assertion is on a type or on a value, never on message text - which is the entire point of
+ /// the change.
+ ///
+ ///
+ /// Each test also asserts that the base type still catches, because the promise made to
+ /// consumers is that this is additive: a catch (NotConnectedException) written against
+ /// 11.4.0 keeps working.
+ ///
+ ///
+ [TestClass]
+ public class TestUConnectionOutcomes
+ {
+ private XrplClient _client;
+
+ private static Dictionary ServerInfoResponse() => new Dictionary
+ {
+ { "type", "response" },
+ { "status", "success" },
+ { "result", new Dictionary
+ {
+ { "info", new Dictionary
+ {
+ { "build_version", "test-mock" },
+ { "complete_ledgers", "1-1" },
+ { "server_state", "full" },
+ }
+ },
+ }
+ },
+ };
+
+ private static CreateMockRippled StartMock(int port)
+ {
+ CreateMockRippled mock = new CreateMockRippled(port) { suppressOutput = true };
+ mock.AddResponse("server_info", ServerInfoResponse());
+
+ Thread listenerThread = new Thread(() => mock.Start()) { IsBackground = true };
+ listenerThread.Start();
+ return mock;
+ }
+
+ [TestCleanup]
+ public async Task MyTestCleanup()
+ {
+ if (_client != null)
+ {
+ await _client.Disconnect();
+ _client = null;
+ }
+ }
+
+ ///
+ /// A client that was never told to connect says so, rather than saying it is disconnected.
+ ///
+ ///
+ /// "Nothing is in progress" is a distinct, actionable answer - the caller has to call
+ /// Connect() - and it is neither "the consumer took the client down" nor "the
+ /// reconnect loop gave up". It used to arrive as a bare
+ /// alongside both of those.
+ ///
+ [TestMethod]
+ public async Task TestUAClientThatNeverConnectedReportsThatNothingIsInProgress()
+ {
+ int port = TestUtils.GetFreePort(); // nothing is listening, and nothing will be asked to
+ _client = new XrplClient($"ws://127.0.0.1:{port}");
+
+ NotConnectingException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.WaitForConnectionAsync(TimeSpan.FromSeconds(1)));
+
+ Assert.IsInstanceOfType(error, "catch (NotConnectedException) must keep catching this.");
+ }
+
+ ///
+ /// A caller already waiting when the reconnect loop spends its budget is told the budget is
+ /// spent, and how big it was.
+ ///
+ ///
+ ///
+ /// This is the outcome that justifies a failover, and the one it was impossible to
+ /// recognise: it arrived as the same as a consumer's
+ /// own Disconnect(), which calls for the opposite reaction.
+ ///
+ ///
+ /// The wait has to begin before the loop gives up. Once the loop has released its
+ /// cancellation source and reported Disconnected, a fresh call finds no attempt in
+ /// progress at all and is answered by - a different,
+ /// equally correct answer to a different question. That is why the waiter here is
+ /// ChangeServer's own: it starts the attempt and then waits for it, so it is
+ /// already parked when the loop it left behind runs out of budget. Stopping a server and
+ /// racing to park a waiter before the loop finishes would test the same thing by timing.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUAWaiterLearnsTheReconnectBudgetWasSpent()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+ Assert.IsTrue(_client.connection.IsConnected(), "Precondition: connected to the mock.");
+
+ int deadPort = TestUtils.GetFreePort(); // nothing is listening there, and never will be
+
+ ReconnectExhaustedException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.ChangeServer($"ws://127.0.0.1:{deadPort}"));
+
+ Assert.AreEqual(2, error.MaxAttempts, "The budget the client was configured with.");
+ Assert.AreEqual(
+ error.MaxAttempts,
+ error.Attempts,
+ "Attempts made, not the raw counter - which stands one past the budget when the loop stops.");
+ Assert.IsInstanceOfType(error, "catch (NotConnectedException) must keep catching this.");
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// Giving up on a broken OnConnected handler says the handler broke - not that the
+ /// consumer disconnected the client.
+ ///
+ ///
+ ///
+ /// The give-up path ends by calling Disconnect() itself, which sets the same
+ /// permanently-disconnected flag a consumer's own Disconnect() sets. Every point
+ /// that reads that flag then answered "the client has been disconnected" - the one thing
+ /// this event is not. The node is answering; what failed is code on this side, so a
+ /// consumer reacting by failing over would be leaving a healthy server.
+ ///
+ ///
+ /// The distinction was invisible and timing-dependent: the same broken handler produced one
+ /// type when it failed immediately and another when it failed a moment later, because only
+ /// the second left a request in flight to be rejected with the real reason.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUGivingUpOnABrokenConnectHandlerSaysTheHandlerBroke()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ InvalidOperationException thrownByHandler = new InvalidOperationException("handler is permanently broken");
+ _client.connection.OnConnected += () => throw thrownByHandler;
+
+ ConnectHandlerFailedException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.Connect());
+
+ Assert.IsGreaterThanOrEqualTo(1, error.Failures, "The handler failed at least once before the client gave up.");
+ Assert.AreSame(thrownByHandler, error.InnerException, "The handler's own failure is what says why.");
+ Assert.IsInstanceOfType(error, "catch (NotConnectedException) must keep catching this.");
+
+ // The value-returning way of asking has to name the same case. It is a separate
+ // path - the exception is caught and translated - so a mapping that is wrong here
+ // is wrong for every consumer who prefers a value to a catch.
+ ConnectionWaitOutcome outcome =
+ await _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(5));
+
+ Assert.AreEqual(
+ ConnectionWaitOutcome.ConnectHandlerFailed,
+ outcome,
+ "The outcome and the exception are two ways of asking one question.");
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// An operation that was overtaken says what overtook it and where the client ended up.
+ ///
+ ///
+ ///
+ /// The second ChangeServer is issued from the first one's own session-ended handler,
+ /// which lands it inside the first one's yields every time - the shape a consumer actually
+ /// hits, and deterministic where a race would not be.
+ ///
+ ///
+ /// Before this, the overtaken call got a bare ,
+ /// indistinguishable from the caller's own token being cancelled. A consumer that showed
+ /// the user "could not connect, pick another node" for it was reporting a failure on a
+ /// client that was at that moment connecting normally somewhere else.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUAnOvertakenSwitchNamesTheSwitchThatWon()
+ {
+ int firstPort = TestUtils.GetFreePort();
+ int secondPort = TestUtils.GetFreePort();
+ int thirdPort = TestUtils.GetFreePort();
+
+ CreateMockRippled first = StartMock(firstPort);
+ CreateMockRippled second = StartMock(secondPort);
+ CreateMockRippled third = StartMock(thirdPort);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{firstPort}", new XrplClient.ClientOptions
+ {
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ string thirdUrl = $"ws://127.0.0.1:{thirdPort}";
+ int nested = 0;
+ _client.OnSessionEnded += async (reason, _) =>
+ {
+ if (reason == SessionEndReason.ServerChanged && Interlocked.Exchange(ref nested, 1) == 0)
+ {
+ await _client.connection.ChangeServer(thirdUrl);
+ }
+ };
+
+ ConnectionSupersededException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.ChangeServer($"ws://127.0.0.1:{secondPort}"));
+
+ Assert.AreEqual(ConnectionTransitionKind.ChangeServer, error.Kind);
+ Assert.AreEqual(thirdUrl, error.SupersededBy, "The caller is told where the client actually is.");
+ Assert.IsInstanceOfType(
+ error,
+ "catch (OperationCanceledException) must keep catching this.");
+ }
+ finally
+ {
+ first.Stop();
+ second.Stop();
+ third.Stop();
+ }
+ }
+
+ ///
+ /// A request refused because the caller asked not to wait says that, and not that the
+ /// endpoint is dead.
+ ///
+ ///
+ /// ImmediateFail is a policy the consumer chose, and the refusal it produces says
+ /// nothing about the server: the connection was being rebuilt and the caller asked not to
+ /// wait for it. Retrying once connected is the reaction. It used to arrive as the same
+ /// as "the reconnect loop gave up", which calls for a
+ /// failover instead.
+ ///
+ [TestMethod]
+ public async Task TestUARequestRefusedByPolicySaysSoRatherThanBlamingTheServer()
+ {
+ int firstPort = TestUtils.GetFreePort();
+ int secondPort = TestUtils.GetFreePort();
+
+ CreateMockRippled first = StartMock(firstPort);
+ CreateMockRippled second = StartMock(secondPort);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{firstPort}", new XrplClient.ClientOptions
+ {
+ RequestPolicy = RequestFailurePolicy.ImmediateFail,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ // Issued from inside the switch, where the old socket is gone and the new one is not
+ // open yet - the window the policy exists for.
+ Exception refusal = null;
+ _client.OnSessionEnded += async (reason, _) =>
+ {
+ if (reason == SessionEndReason.ServerChanged && refusal == null)
+ {
+ try
+ {
+ await _client.Request(new Dictionary { { "command", "server_info" } });
+ }
+ catch (Exception error)
+ {
+ refusal = error;
+ }
+ }
+ };
+
+ await _client.connection.ChangeServer($"ws://127.0.0.1:{secondPort}");
+
+ Assert.IsInstanceOfType(
+ refusal,
+ $"A request refused by ImmediateFail must say so, got: {refusal?.GetType().Name ?? "no exception"}.");
+ Assert.IsInstanceOfType(refusal, "catch (NotConnectedException) must keep catching this.");
+ }
+ finally
+ {
+ first.Stop();
+ second.Stop();
+ }
+ }
+
+ ///
+ /// Switching servers survives a connection that settles on the second try, the way
+ /// connecting does.
+ ///
+ ///
+ ///
+ /// Connect() is two operations - the connection, and the server_info that
+ /// reads the network id - and it has carried the second across a teardown since 11.4.0,
+ /// because a socket really does open for a moment before a failing handler brings it down.
+ /// ChangeServer is the same two operations against a different server and had no
+ /// such protection: it read the network id once, directly, so a connection that needed a
+ /// second attempt failed the switch.
+ ///
+ ///
+ /// The second server drops the connection the first time it is asked for
+ /// server_info and serves normally afterwards, which puts the teardown exactly
+ /// between "the switch is connected" and "the switch has read the network id" - the only
+ /// window where the two behave differently.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUSwitchingServersSurvivesAConnectionThatSettlesOnRetry()
+ {
+ int firstPort = TestUtils.GetFreePort();
+
+ CreateMockRippled first = StartMock(firstPort);
+ DropsFirstServerInfoServer dropping = new DropsFirstServerInfoServer();
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{firstPort}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 4,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(15),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ await _client.ChangeServer(dropping.Url);
+
+ Assert.IsTrue(_client.connection.IsConnected(), "The switch has to end connected, not thrown.");
+ }
+ finally
+ {
+ first.Stop();
+ dropping.Dispose();
+ }
+ }
+
+ ///
+ /// A request swept by a server switch is told which switch took the connection, and where.
+ ///
+ ///
+ ///
+ /// This is the failure consumers meet most often - not an overtaken transition, but their
+ /// own request dying while the connection moved underneath it. It arrived as a bare
+ /// reading "Connection was intentionally closed",
+ /// identical to a request the caller cancelled itself.
+ ///
+ ///
+ /// The new type still derives from , so a caller
+ /// that treated this as cancellation keeps working unchanged and the task keeps the status
+ /// it had. That is asserted here alongside the new information, because it is the promise
+ /// that makes the change safe to ship in a minor version.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUARequestSweptByASwitchNamesTheSwitch()
+ {
+ int firstPort = TestUtils.GetFreePort();
+ int secondPort = TestUtils.GetFreePort();
+
+ CreateMockRippled first = new CreateMockRippled(firstPort) { suppressOutput = true };
+ first.AddResponse("server_info", ServerInfoResponse());
+ // Sat on long enough that the switch below lands while the request is still in flight.
+ first.AddDelayedResponse("ledger", ServerInfoResponse(), TimeSpan.FromSeconds(10));
+ new Thread(() => first.Start()) { IsBackground = true }.Start();
+
+ CreateMockRippled second = StartMock(secondPort);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{firstPort}", new XrplClient.ClientOptions
+ {
+ RequestTimeout = TimeSpan.FromSeconds(30),
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ Task inFlight = _client.Request(new Dictionary { { "command", "ledger" } });
+ await Task.Delay(TimeSpan.FromMilliseconds(300)); // it is on the wire, and unanswered
+
+ string secondUrl = $"ws://127.0.0.1:{secondPort}";
+ await _client.connection.ChangeServer(secondUrl);
+
+ ConnectionSupersededException error = await Assert.ThrowsExactlyAsync(
+ async () => await inFlight);
+
+ Assert.AreEqual(ConnectionTransitionKind.ChangeServer, error.Kind);
+ Assert.AreEqual(secondUrl, error.SupersededBy);
+ Assert.IsInstanceOfType(
+ error,
+ "catch (OperationCanceledException) must keep catching a swept request.");
+ }
+ finally
+ {
+ first.Stop();
+ second.Stop();
+ }
+ }
+
+ ///
+ /// The notification that says the client stopped also says why it stopped.
+ ///
+ ///
+ ///
+ /// Disconnected is announced from ten places - a consumer's disconnect, a first
+ /// connection that failed, a connection closed for good, a broken handler, and the reconnect
+ /// loop running out of budget - and none of them carried anything but text. "Still trying"
+ /// against "gave up" was derivable only from the absence of ReconnectInfo, which is
+ /// also what a client that never had a loop looks like.
+ ///
+ ///
+ /// The reason goes on the notification rather than into ReconnectInfo: filling that
+ /// in on a terminal notification would take Reconnect != null, which consumers read
+ /// as "a loop is running", and give it a second meaning. It stays null here, and that is
+ /// asserted.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUTheTerminalNotificationSaysWhyTheClientStopped()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ List statuses = new List();
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ _client.connection.OnConnectionStatus += status =>
+ {
+ lock (statuses)
+ {
+ statuses.Add(status);
+ }
+ };
+
+ int deadPort = TestUtils.GetFreePort();
+ try
+ {
+ await _client.connection.ChangeServer($"ws://127.0.0.1:{deadPort}");
+ }
+ catch (ReconnectExhaustedException)
+ {
+ // The subject of this test is the notification, not the exception.
+ }
+
+ ConnectionStatusInfo terminal;
+ lock (statuses)
+ {
+ terminal = statuses.FindLast(s => s.ConnectionState == XrpConnectionState.Disconnected);
+ }
+
+ Assert.IsNotNull(terminal, "The client has to announce that it stopped.");
+ Assert.AreEqual(
+ ConnectionStopReason.ReconnectExhausted,
+ terminal.StopReason,
+ "Giving up on the budget is a different event from the consumer disconnecting.");
+ Assert.IsNull(
+ terminal.Reconnect,
+ "Reconnect stays null on a terminal notification, so 'a loop is running' keeps its one meaning.");
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// A consumer's own disconnect is named as such, and is not confused with giving up.
+ ///
+ [TestMethod]
+ public async Task TestUAConsumerDisconnectIsNamedInTheStatusStream()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ List statuses = new List();
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ _client.connection.OnConnectionStatus += status =>
+ {
+ lock (statuses)
+ {
+ statuses.Add(status);
+ }
+ };
+
+ XrplClient disconnected = _client;
+ await _client.Disconnect();
+ _client = null;
+
+ ConnectionStatusInfo terminal;
+ lock (statuses)
+ {
+ terminal = statuses.FindLast(s => s.ConnectionState == XrpConnectionState.Disconnected);
+ }
+
+ Assert.IsNotNull(terminal);
+ Assert.AreEqual(ConnectionStopReason.UserDisconnected, terminal.StopReason);
+
+ // And the wait says the same, under the same name. The two enums are two views of
+ // one event, and the pair for the most ordinary ending of all was the one the
+ // suite never asked for - it was also the pair whose names did not match.
+ ConnectionWaitOutcome outcome =
+ await disconnected.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(5));
+
+ Assert.AreEqual(ConnectionWaitOutcome.UserDisconnected, outcome);
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// "Did it come back?" is answerable without catching anything - and the answer says which
+ /// of the ways it did not.
+ ///
+ ///
+ ///
+ /// Both known consumers wrapped the throwing wait to get a value back, because "it did not
+ /// come back in time" is an answer a caller has to act on rather than an exceptional event.
+ /// A bool would have been the obvious shape and the wrong one: it folds "timed out",
+ /// "gave up" and "nothing is running" into one false, which is the problem this
+ /// whole change is about, moved into a new method.
+ ///
+ ///
+ /// The caller's own cancellation and an invalid timeout stay exceptions: the first is the
+ /// .NET convention, the second is a mistake by the caller rather than an outcome of the
+ /// connection.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUTheOutcomeOfAWaitIsAValueAndNamesTheCase()
+ {
+ int port = TestUtils.GetFreePort();
+ _client = new XrplClient($"ws://127.0.0.1:{port}");
+
+ ConnectionWaitOutcome outcome =
+ await _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(1));
+
+ Assert.AreEqual(ConnectionWaitOutcome.NotConnecting, outcome);
+ }
+
+ ///
+ /// It is callable through the interface and through the class alike.
+ ///
+ ///
+ /// The interface member is defaulted - forwarding to connection is the only
+ /// implementation that means anything, and an external implementer of
+ /// should not have to add one. A default member is only visible
+ /// through an interface-typed reference, though, so the client carries its own as well;
+ /// both are exercised here because a consumer holding either must be able to ask.
+ ///
+ [TestMethod]
+ public async Task TestUTheOutcomeIsReachableThroughTheInterfaceAndTheClass()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ IXrplClient asInterface = _client;
+
+ Assert.AreEqual(
+ ConnectionWaitOutcome.Connected,
+ await asInterface.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(5)));
+ Assert.AreEqual(
+ ConnectionWaitOutcome.Connected,
+ await _client.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(5)));
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// A request swept by the client rebuilding its own connection names that rebuild.
+ ///
+ ///
+ /// The fourth kind, and the one a consumer must not read as a failure of the node: the
+ /// health check found the connection silent and replaced it, the server is where it always
+ /// was, and the request is worth sending again once the new connection is up. Reached
+ /// through the health check because that is the only path that rebuilds a connection
+ /// nobody asked it to rebuild.
+ ///
+ [TestMethod]
+ public async Task TestUARequestSweptByAReconnectNamesTheReconnect()
+ {
+ using SilentOnPingAndLedgerServer silent = new SilentOnPingAndLedgerServer();
+
+ _client = new XrplClient(silent.Url, new XrplClient.ClientOptions
+ {
+ RequestTimeout = TimeSpan.FromSeconds(30),
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(500),
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(3),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = true,
+ HealthCheckInterval = TimeSpan.FromMilliseconds(200),
+ InactivityTimeout = TimeSpan.FromMilliseconds(500),
+ });
+
+ await _client.Connect();
+
+ Task inFlight = _client.Request(new Dictionary { { "command", "ledger" } });
+
+ ConnectionSupersededException error = await Assert.ThrowsExactlyAsync(
+ async () => await inFlight);
+
+ Assert.AreEqual(ConnectionTransitionKind.Reconnect, error.Kind);
+ Assert.IsInstanceOfType(error);
+ }
+
+ ///
+ /// A waiter parked while the client gives up on a broken handler is answered, not left to
+ /// time out.
+ ///
+ ///
+ ///
+ /// Two things have to line up for this to work, and neither is obvious. The give-up
+ /// announces itself before it performs the disconnect that records why, so a waiter woken
+ /// by that announcement finds nothing terminal yet and parks again. What has to reach it
+ /// then is the disconnect's own notification - and that one says Disconnected on a
+ /// client already reported as Disconnected, so the status stream suppresses it as a
+ /// duplicate.
+ ///
+ ///
+ /// Waking waiters therefore cannot be a side effect of emitting a status event: whether an
+ /// event is worth showing a consumer and whether the state changed are different
+ /// questions. The timeout here is long against a give-up that takes well under a second,
+ /// so "answered" and "gave up waiting" cannot be confused.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUAWaiterIsAnsweredWhenTheClientGivesUpOnItsHandler()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(20),
+ UseCustomPing = false,
+ });
+
+ _client.connection.OnConnected += () => throw new InvalidOperationException("permanently broken");
+
+ // Parked from the status stream rather than after a sleep: the socket is open for
+ // the moment the handler runs in, so a waiter started by the clock can find the
+ // client connected and answer that instead. RestoringConnection is the state where
+ // there is no socket and the client is between attempts - which is where a real
+ // caller waiting for the connection to come back sits.
+ Task waiting = null;
+ _client.connection.OnConnectionStatus += status =>
+ {
+ if (status.ConnectionState == XrpConnectionState.RestoringConnection
+ && Interlocked.CompareExchange(ref waiting, null, null) == null)
+ {
+ Interlocked.CompareExchange(
+ ref waiting,
+ _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(20)),
+ null);
+ }
+ };
+
+ System.Diagnostics.Stopwatch clock = System.Diagnostics.Stopwatch.StartNew();
+
+ await Assert.ThrowsExactlyAsync(async () => await _client.Connect());
+
+ Task parked = Interlocked.CompareExchange(ref waiting, null, null);
+ Assert.IsNotNull(parked, "Precondition: the client has to report RestoringConnection at least once.");
+
+ ConnectionWaitOutcome outcome = await parked;
+ clock.Stop();
+
+ // Not asserted as one particular outcome: the socket really does open for the
+ // moment the handler runs in, so a waiter can legitimately be answered
+ // "Connected" by an attempt that is about to fail, or "ConnectHandlerFailed" by
+ // the give-up. Which one wins is timing. What must never happen is neither - a
+ // waiter left parked because the wake-up that concerned it was swallowed.
+ Assert.AreNotEqual(
+ ConnectionWaitOutcome.TimedOut,
+ outcome,
+ "The waiter sat out its whole timeout: a state change it was waiting for did not reach it.");
+ Assert.IsLessThan(
+ TimeSpan.FromSeconds(10),
+ clock.Elapsed,
+ "The waiter has to be answered by the client, not by its own deadline.");
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// A failure the client is about to retry names no reason for stopping, because it has not
+ /// stopped.
+ ///
+ ///
+ ///
+ /// is defined as why the client stopped, and
+ /// as "this notification is not an ending". A
+ /// reason on a notification that is followed by a reconnect breaks that definition in the
+ /// way that costs something: a consumer reading any reason as terminal fails over to
+ /// another server while this one is still being dialled.
+ ///
+ ///
+ /// The first handshake against a server that is not up is exactly that case - it is
+ /// reported, and then retried. Whether a reason is named now follows the same condition
+ /// that decides whether the retry happens, so the two cannot say different things.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUAFailureTheClientWillRetryNamesNoStopReason()
+ {
+ int port = TestUtils.GetFreePort(); // nothing is listening, and nothing will be
+
+ List statuses = new List();
+
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 1000,
+ StopAfterMaxAttempts = false, // it will keep trying, so it never stops
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(3),
+ UseCustomPing = false,
+ });
+
+ _client.connection.OnConnectionStatus += status =>
+ {
+ lock (statuses)
+ {
+ statuses.Add(status);
+ }
+ };
+
+ try
+ {
+ await _client.Connect();
+ }
+ catch (Exception)
+ {
+ // The connection never comes up; the subject here is what was announced on the way.
+ }
+
+ await Task.Delay(TimeSpan.FromMilliseconds(500));
+
+ List claimingAnEnding;
+ bool stillTrying;
+ string trace;
+ lock (statuses)
+ {
+ claimingAnEnding = statuses.FindAll(s => s.StopReason != ConnectionStopReason.None);
+ stillTrying = statuses.Exists(s => s.ConnectionState == XrpConnectionState.RestoringConnection);
+ trace = string.Join(" | ", statuses.ConvertAll(x => $"{x.ConnectionState}/{x.StopReason}"));
+ }
+
+ Assert.IsTrue(stillTrying, $"Precondition: the client has to be retrying. Sequence was: {trace}");
+ Assert.AreEqual(
+ 0,
+ claimingAnEnding.Count,
+ $"A client that is still dialling must not name a reason for having stopped. Sequence was: {trace}");
+ }
+
+ ///
+ /// A client that spent its reconnect budget stays stopped.
+ ///
+ ///
+ ///
+ /// StopAfterMaxAttempts is a promise that the client will stop asking, and the
+ /// consumer is told it has: Disconnected with
+ /// . A second series behind that
+ /// notification breaks the promise twice over - the budget is spent again without being
+ /// granted again, and a consumer that failed over on the first notification now has a
+ /// client quietly dialling the endpoint it moved away from.
+ ///
+ ///
+ /// What made it possible: the loop's exit clears the two fields that say "a sequence is
+ /// running for this generation", which is exactly what "no sequence is running" looks
+ /// like. The close of the last failed attempt arrives after that and is indistinguishable
+ /// from the first one.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUAClientThatSpentItsReconnectBudgetStaysStopped()
+ {
+ SilentOnPingAndLedgerServer silent = new SilentOnPingAndLedgerServer();
+
+ List statuses = new List();
+
+ try
+ {
+ _client = new XrplClient(silent.Url, new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 1,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(1),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(5),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ _client.connection.OnConnectionStatus += status =>
+ {
+ lock (statuses)
+ {
+ statuses.Add(status);
+ }
+ };
+
+ silent.Dispose();
+
+ // Long enough for a second series to have run and announced itself: the budget
+ // above is spent in a few hundred milliseconds.
+ await Task.Delay(TimeSpan.FromSeconds(4));
+
+ int gaveUp;
+ string trace;
+ lock (statuses)
+ {
+ gaveUp = statuses.FindAll(s =>
+ s.ConnectionState == XrpConnectionState.Disconnected &&
+ s.StopReason == ConnectionStopReason.ReconnectExhausted).Count;
+ trace = string.Join(" | ", statuses.ConvertAll(x => $"{x.ConnectionState}/{x.StopReason}: {x.Message}"));
+ }
+
+ Assert.AreEqual(1, gaveUp, $"Giving up is announced once and meant once. Sequence was: {trace}");
+ Assert.IsFalse(_client.connection.IsConnected());
+
+ // The other half of the promise, and the risk of keeping it: stopping must not mean
+ // wedged. Asking again is the consumer's decision, and a consumer command begins a
+ // new generation - which is what lifts the refusal, with no flag to reset and no
+ // way for it to outlive the sequence it belongs to.
+ int livePort = TestUtils.GetFreePort();
+ CreateMockRippled live = StartMock(livePort);
+ try
+ {
+ await _client.ChangeServer($"ws://127.0.0.1:{livePort}");
+ Assert.IsTrue(
+ _client.connection.IsConnected(),
+ "A client that stopped asking on its own must still answer the consumer asking.");
+ }
+ finally
+ {
+ live.Stop();
+ }
+ }
+ finally
+ {
+ silent.Dispose();
+ }
+ }
+
+ ///
+ /// The same answer whether the handler fails at once or a moment later.
+ ///
+ ///
+ ///
+ /// The timing decides which path reports the failure. A handler that throws at once leaves
+ /// nothing in flight and the caller is answered by the wait; one that works for a moment
+ /// first - a subscribe that gets some way in before falling over - lets the caller reach
+ /// the network-id read, and the teardown then finds a request in flight to reject. Both
+ /// used to be the same bare exception, and attaching the reason to only one of them would
+ /// have handed the caller two different types for one scenario depending on how fast the
+ /// machine was.
+ ///
+ ///
+ /// Holding server_info back is what makes the second path certain rather than lucky:
+ /// answered at once, the request is gone before the handler gives up and the test would
+ /// pass without ever exercising the case.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUABrokenConnectHandlerAnswersTheSameWhicheverWayItFails()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = new CreateMockRippled(port) { suppressOutput = true };
+ mock.AddDelayedResponse("server_info", ServerInfoResponse(), TimeSpan.FromSeconds(5));
+ new Thread(() => mock.Start()) { IsBackground = true }.Start();
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 1,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ InvalidOperationException thrownByHandler = new InvalidOperationException("broken, but not straight away");
+ _client.connection.OnConnected += async () =>
+ {
+ await Task.Delay(TimeSpan.FromMilliseconds(400));
+ throw thrownByHandler;
+ };
+
+ ConnectHandlerFailedException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.Connect());
+
+ Assert.AreSame(thrownByHandler, error.InnerException);
+ Assert.IsGreaterThanOrEqualTo(1, error.Failures);
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// The status stream names a broken handler as such, and says nothing about a notification
+ /// that is not terminal.
+ ///
+ ///
+ /// The two halves belong together: a reason that appeared on every notification would be
+ /// as useless as none at all, and on the
+ /// RestoringConnection that precedes the give-up is what lets a consumer treat the
+ /// field as "this one is terminal, and here is why".
+ ///
+ [TestMethod]
+ public async Task TestUTheStatusStreamNamesABrokenHandlerAndOnlyWhenTerminal()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ List statuses = new List();
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ _client.connection.OnConnectionStatus += status =>
+ {
+ lock (statuses)
+ {
+ statuses.Add(status);
+ }
+ };
+
+ _client.connection.OnConnected += () => throw new InvalidOperationException("permanently broken");
+
+ try
+ {
+ await _client.Connect();
+ }
+ catch (ConnectHandlerFailedException)
+ {
+ // The subject here is the status stream.
+ }
+
+ ConnectionStatusInfo terminal;
+ List nonTerminal;
+ lock (statuses)
+ {
+ terminal = statuses.FindLast(s => s.ConnectionState == XrpConnectionState.Disconnected);
+ nonTerminal = statuses.FindAll(s => s.ConnectionState != XrpConnectionState.Disconnected);
+ }
+
+ string seq;
+ lock (statuses)
+ {
+ seq = string.Join(" | ", statuses.ConvertAll(x => $"{x.ConnectionState}/{x.StopReason}"));
+ }
+
+ Assert.IsNotNull(terminal);
+
+ // The last word, deliberately: the give-up announces the handler failure and then
+ // performs a disconnect of its own, which announces again. Both have to name the
+ // same event, or the status stream ends by contradicting the exception the same
+ // failure produced. Reading the last one is what catches that.
+ Assert.AreEqual(
+ ConnectionStopReason.ConnectHandlerFailed,
+ terminal.StopReason,
+ $"The last thing said about a broken handler must still be the handler. Sequence was: {seq}");
+
+ foreach (ConnectionStatusInfo status in nonTerminal)
+ {
+ Assert.AreEqual(
+ ConnectionStopReason.None,
+ status.StopReason,
+ $"A {status.ConnectionState} notification is not an ending and must not claim a reason.");
+ }
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// A switch overtaken at the client level is reported, not retried away.
+ ///
+ ///
+ /// XrplClient.ChangeServer reads the network id after the switch and carries that
+ /// read across a teardown, which means it has a retry loop around an operation that can
+ /// fail because the client moved. The loop must not swallow a supersession: the caller's
+ /// switch did not happen, and asking again would read the network id of a server it never
+ /// named. The line is drawn by the kind of transition, which is why this is asserted
+ /// through the client rather than through the connection.
+ ///
+ [TestMethod]
+ public async Task TestUAnOvertakenSwitchIsReportedThroughTheClientToo()
+ {
+ int firstPort = TestUtils.GetFreePort();
+ int secondPort = TestUtils.GetFreePort();
+ int thirdPort = TestUtils.GetFreePort();
+
+ CreateMockRippled first = StartMock(firstPort);
+ CreateMockRippled second = StartMock(secondPort);
+ CreateMockRippled third = StartMock(thirdPort);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{firstPort}", new XrplClient.ClientOptions
+ {
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ string thirdUrl = $"ws://127.0.0.1:{thirdPort}";
+ int nested = 0;
+ _client.OnSessionEnded += async (reason, _) =>
+ {
+ if (reason == SessionEndReason.ServerChanged && Interlocked.Exchange(ref nested, 1) == 0)
+ {
+ await _client.connection.ChangeServer(thirdUrl);
+ }
+ };
+
+ ConnectionSupersededException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.ChangeServer($"ws://127.0.0.1:{secondPort}"));
+
+ Assert.AreEqual(ConnectionTransitionKind.ChangeServer, error.Kind);
+ Assert.AreEqual(thirdUrl, error.SupersededBy);
+ }
+ finally
+ {
+ first.Stop();
+ second.Stop();
+ third.Stop();
+ }
+ }
+
+ ///
+ /// Starts a mock that sits on one command, so a request can still be in flight when
+ /// something happens to the connection.
+ ///
+ private static CreateMockRippled StartMockHoldingOnto(int port, string command, TimeSpan delay)
+ {
+ CreateMockRippled mock = new CreateMockRippled(port) { suppressOutput = true };
+ mock.AddResponse("server_info", ServerInfoResponse());
+ mock.AddDelayedResponse(command, ServerInfoResponse(), delay);
+
+ new Thread(() => mock.Start()) { IsBackground = true }.Start();
+ return mock;
+ }
+
+ private XrplClient ClientHoldingRequests(int port) =>
+ new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ RequestTimeout = TimeSpan.FromSeconds(30),
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ ///
+ /// A request swept by the consumer disconnecting is told so, and is still a cancellation.
+ ///
+ ///
+ /// The kind matters more here than anywhere: a request that died because the consumer took
+ /// the client down needs no retry and no failover, and it used to be indistinguishable from
+ /// one that died because the connection moved. It stays an
+ /// rather than becoming a
+ /// - the request was cancelled, and changing that
+ /// would turn a cancellation into a failure for every caller that already handles it.
+ ///
+ [TestMethod]
+ public async Task TestUARequestSweptByADisconnectNamesTheDisconnect()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMockHoldingOnto(port, "ledger", TimeSpan.FromSeconds(10));
+
+ try
+ {
+ _client = ClientHoldingRequests(port);
+ await _client.Connect();
+
+ Task inFlight = _client.Request(new Dictionary { { "command", "ledger" } });
+ await Task.Delay(TimeSpan.FromMilliseconds(300));
+
+ await _client.Disconnect();
+ _client = null;
+
+ ConnectionSupersededException error = await Assert.ThrowsExactlyAsync(
+ async () => await inFlight);
+
+ Assert.AreEqual(ConnectionTransitionKind.Disconnect, error.Kind);
+ Assert.IsNull(error.SupersededBy, "A disconnect took the client nowhere.");
+ Assert.IsInstanceOfType(error);
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// Connect() on a client that is already connected disturbs nothing.
+ ///
+ ///
+ /// Written while trying to reach the sweep that Connect() performs, and kept because
+ /// of what it found instead: Connect() returns at once when the client is already
+ /// connected, so it never gets as far as taking the socket over. The sweep is therefore
+ /// reachable only from a client that is not connected - where there is no request in flight
+ /// to sweep - and the Connect kind is exercised through supersession instead. What
+ /// is worth pinning here is that a redundant Connect() does not quietly kill the
+ /// requests a caller has outstanding.
+ ///
+ [TestMethod]
+ public async Task TestURedundantConnectDoesNotDisturbRequestsInFlight()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMockHoldingOnto(port, "ledger", TimeSpan.FromSeconds(1));
+
+ try
+ {
+ _client = ClientHoldingRequests(port);
+ await _client.Connect();
+
+ Task inFlight = _client.Request(new Dictionary { { "command", "ledger" } });
+ await Task.Delay(TimeSpan.FromMilliseconds(200));
+
+ await _client.connection.Connect(CancellationToken.None);
+
+ await inFlight; // answered by the connection it was written to
+ Assert.IsTrue(_client.connection.IsConnected());
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// A request killed by the connection failing on its own stays a plain cancellation.
+ ///
+ ///
+ ///
+ /// This is the deliberate border of the change. A connection that failed has no transition
+ /// to name and no destination to point at, so inventing one would be worse than saying
+ /// little: a consumer switching on would be told the
+ /// client moved somewhere when nothing moved it.
+ ///
+ ///
+ /// What such a request actually gets is , reported by
+ /// the close path with the code and reason the peer gave - measured here rather than
+ /// assumed, because the sweep this test was written to check turned out not to be the one
+ /// that answers first. The assertion that matters either way is the negative one: no
+ /// transition is named.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUARequestKilledByANetworkDropNamesNoTransition()
+ {
+ DropsFirstServerInfoServer dropping = new DropsFirstServerInfoServer();
+
+ try
+ {
+ _client = new XrplClient(dropping.Url, new XrplClient.ClientOptions
+ {
+ RequestTimeout = TimeSpan.FromSeconds(30),
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ MaxReconnectAttempts = 4,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ // The connection alone, so nothing asks for server_info before the test does.
+ await _client.connection.Connect(CancellationToken.None);
+
+ Exception failure = null;
+ try
+ {
+ await _client.Request(new Dictionary { { "command", "server_info" } });
+ }
+ catch (Exception error)
+ {
+ failure = error;
+ }
+
+ Assert.IsNotNull(failure, "The request died with the connection and has to say so.");
+ Assert.IsNotInstanceOfType(
+ failure,
+ "No transition took this connection anywhere - it failed. Naming a transition would be inventing one.");
+ Assert.IsInstanceOfType(
+ failure,
+ $"A peer that went away is reported by the close path, got: {failure.GetType().Name}.");
+ }
+ finally
+ {
+ dropping.Dispose();
+ }
+ }
+
+ ///
+ /// Starts a client that will keep trying to reach a server that is not there, so a wait
+ /// actually parks instead of being answered on entry.
+ ///
+ ///
+ /// An attempt has to be in progress for the wait to reach its parking spot at all: with
+ /// nothing running it is answered by straight away,
+ /// which is a different test. The connect task is left running on purpose and torn down by
+ /// the cleanup.
+ ///
+ private XrplClient StartClientWaitingForever(int port)
+ {
+ XrplClient client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 1000,
+ StopAfterMaxAttempts = false,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(60),
+ UseCustomPing = false,
+ });
+
+ _ = client.Connect();
+ return client;
+ }
+
+ ///
+ /// Waiting for a connection that never comes: the timeout is an answer, the caller's own
+ /// cancellation is not, and a bad timeout is the caller's mistake.
+ ///
+ ///
+ /// These three are the boundary of the outcome contract and the reason it is not simply
+ /// "every failure becomes a value". A timeout is something the caller has to act on and so
+ /// it is reported as ; cancellation through the
+ /// caller's own token stays an exception because that is the .NET convention and because
+ /// the caller already knows it cancelled; an invalid timeout is a bug at the call site
+ /// rather than an outcome of the connection. This is also the code that changed most when
+ /// the wait stopped polling, so it is pinned rather than assumed.
+ ///
+ [TestMethod]
+ public async Task TestUTimeoutIsAnAnswerAndCancellationIsNot()
+ {
+ int port = TestUtils.GetFreePort(); // nothing is listening, and nothing will be
+ _client = StartClientWaitingForever(port);
+
+ // The wait has to be parked, not answered on entry.
+ await Task.Delay(TimeSpan.FromMilliseconds(300));
+
+ Assert.AreEqual(
+ ConnectionWaitOutcome.TimedOut,
+ await _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromMilliseconds(300)),
+ "Not coming back in time is an answer, not a failure.");
+
+ using CancellationTokenSource alreadyCancelled = new CancellationTokenSource();
+ alreadyCancelled.Cancel();
+ await Assert.ThrowsAsync(async () =>
+ await _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(30), alreadyCancelled.Token));
+
+ using CancellationTokenSource cancelledWhileWaiting = new CancellationTokenSource();
+ Task waiting =
+ _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(30), cancelledWhileWaiting.Token);
+ cancelledWhileWaiting.CancelAfter(TimeSpan.FromMilliseconds(200));
+ await Assert.ThrowsAsync(async () => await waiting);
+
+ await Assert.ThrowsExactlyAsync(async () =>
+ await _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.Zero));
+ }
+
+ ///
+ /// Waiters do not interfere: one timing out and one cancelling leave the third waiting.
+ ///
+ ///
+ /// The wait sleeps on a signal shared by everyone waiting, so a per-waiter deadline must
+ /// not touch it: an implementation that cancelled the shared signal to serve its own
+ /// timeout would take every other waiter down with it, and one that never re-armed the
+ /// signal after completing it would leave the survivors awake and spinning. Both are the
+ /// kind of mistake that shows only with more than one waiter, which is why this test has
+ /// three.
+ ///
+ [TestMethod]
+ public async Task TestUWaitersDoNotTakeEachOtherDown()
+ {
+ int port = TestUtils.GetFreePort();
+ _client = StartClientWaitingForever(port);
+
+ await Task.Delay(TimeSpan.FromMilliseconds(300));
+
+ using CancellationTokenSource cancelling = new CancellationTokenSource();
+
+ Task timesOut =
+ _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromMilliseconds(400));
+ Task cancelled =
+ _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(30), cancelling.Token);
+ Task survives =
+ _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(30));
+
+ cancelling.CancelAfter(TimeSpan.FromMilliseconds(200));
+
+ Assert.AreEqual(ConnectionWaitOutcome.TimedOut, await timesOut);
+ await Assert.ThrowsAsync(async () => await cancelled);
+
+ Assert.IsFalse(survives.IsCompleted, "The third waiter has no reason to be finished yet.");
+
+ // And it is still a live waiter rather than a spinning one: the server comes up, and it
+ // is the connection that ends the wait.
+ CreateMockRippled mock = StartMock(port);
+ try
+ {
+ Assert.AreEqual(ConnectionWaitOutcome.Connected, await survives);
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// A transition a Disconnect() overtook is told the client is down, not that it was
+ /// overtaken.
+ ///
+ ///
+ /// The two answers call for opposite reactions, which is why the disconnect branch keeps
+ /// reporting a while every other winner reports a
+ /// cancellation: after a Disconnect() nothing is coming back on its own, and a
+ /// consumer that treated this as "the switch was overtaken, carry on" would be waiting for
+ /// a connection nobody is building.
+ ///
+ [TestMethod]
+ public async Task TestUASwitchADisconnectOvertookIsToldTheClientIsDown()
+ {
+ int firstPort = TestUtils.GetFreePort();
+ int secondPort = TestUtils.GetFreePort();
+
+ CreateMockRippled first = StartMock(firstPort);
+ CreateMockRippled second = StartMock(secondPort);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{firstPort}", new XrplClient.ClientOptions
+ {
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ int disconnected = 0;
+ _client.OnSessionEnded += async (reason, _) =>
+ {
+ if (reason == SessionEndReason.ServerChanged && Interlocked.Exchange(ref disconnected, 1) == 0)
+ {
+ await _client.Disconnect();
+ }
+ };
+
+ ClientDisconnectedException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.ChangeServer($"ws://127.0.0.1:{secondPort}"));
+
+ Assert.IsInstanceOfType(error, "catch (NotConnectedException) must keep catching this.");
+ Assert.IsFalse(_client.connection.IsConnected(), "The disconnect won, and it has to stay won.");
+ _client = null;
+ }
+ finally
+ {
+ first.Stop();
+ second.Stop();
+ }
+ }
+
+ ///
+ /// A switch overtaken while it waits for its own connection reports it too.
+ ///
+ ///
+ /// This is the check that runs after the wait rather than before it, and it is reached by a
+ /// different route than the others: the switch connected, and only then found the client
+ /// somewhere else. Its condition can prove only that the client is not where this call
+ /// asked for, so the kind is read from the transition that owns the connection instead of
+ /// being assumed from the mismatch - and no existing test in this repository passed through
+ /// it at all.
+ ///
+ [TestMethod]
+ public async Task TestUASwitchOvertakenWhileWaitingReportsTheWinner()
+ {
+ int firstPort = TestUtils.GetFreePort();
+ int secondPort = TestUtils.GetFreePort();
+ int thirdPort = TestUtils.GetFreePort();
+
+ CreateMockRippled first = StartMock(firstPort);
+ CreateMockRippled second = StartMock(secondPort);
+ CreateMockRippled third = StartMock(thirdPort);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{firstPort}", new XrplClient.ClientOptions
+ {
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ string thirdUrl = $"ws://127.0.0.1:{thirdPort}";
+ string secondUrl = $"ws://127.0.0.1:{secondPort}";
+
+ // Issued from the new server's own OnConnected, which runs while the first switch is
+ // still inside its wait - so the takeover lands after the connection came up and
+ // before the switch checks where it ended up.
+ int nested = 0;
+ _client.connection.OnConnected += async () =>
+ {
+ if (string.Equals(_client.connection.GetUrl(), secondUrl, StringComparison.Ordinal)
+ && Interlocked.Exchange(ref nested, 1) == 0)
+ {
+ await _client.connection.ChangeServer(thirdUrl);
+ }
+ };
+
+ ConnectionSupersededException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.ChangeServer(secondUrl));
+
+ Assert.AreEqual(thirdUrl, error.SupersededBy, "The caller is told where the client actually is.");
+ Assert.AreEqual(ConnectionTransitionKind.ChangeServer, error.Kind);
+ }
+ finally
+ {
+ first.Stop();
+ second.Stop();
+ third.Stop();
+ }
+ }
+
+ ///
+ /// A Connect() that overtakes a switch is named as a Connect().
+ ///
+ ///
+ /// The kind is what a consumer reacts to: another ChangeServer put the client on a
+ /// server it did not ask for, while a Connect() rebuilt the connection to the one it
+ /// was already heading for. Both were the same bare cancellation before.
+ ///
+ [TestMethod]
+ public async Task TestUASwitchAConnectOvertookNamesTheConnect()
+ {
+ int firstPort = TestUtils.GetFreePort();
+ int secondPort = TestUtils.GetFreePort();
+
+ CreateMockRippled first = StartMock(firstPort);
+ CreateMockRippled second = StartMock(secondPort);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{firstPort}", new XrplClient.ClientOptions
+ {
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ int reconnected = 0;
+ _client.OnSessionEnded += async (reason, _) =>
+ {
+ if (reason == SessionEndReason.ServerChanged && Interlocked.Exchange(ref reconnected, 1) == 0)
+ {
+ await _client.connection.Connect(CancellationToken.None);
+ }
+ };
+
+ ConnectionSupersededException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.ChangeServer($"ws://127.0.0.1:{secondPort}"));
+
+ Assert.AreEqual(ConnectionTransitionKind.Connect, error.Kind);
+ }
+ finally
+ {
+ first.Stop();
+ second.Stop();
+ }
+ }
+
+ ///
+ /// A waiter learns that the client gave up as soon as it gives up, not when its own timeout
+ /// runs out.
+ ///
+ ///
+ ///
+ /// This guards the wait itself. While it polled, every terminal state was noticed within a
+ /// tick whether or not anything announced it; waiting on a signal instead means a terminal
+ /// condition that forgets to wake its waiters is not slow but invisible - the caller sits
+ /// out the whole acquisition timeout and is then told it timed out, which is the wrong
+ /// answer as well as a late one.
+ ///
+ ///
+ /// The acquisition timeout here is thirty seconds against a reconnect budget that is spent
+ /// in well under one, so the assertion can tell "woken by the client" from "gave up
+ /// waiting" without depending on how fast the machine is.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUAWaiterIsWokenWhenTheClientGivesUpNotWhenItsOwnTimeoutExpires()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ int deadPort = TestUtils.GetFreePort();
+
+ System.Diagnostics.Stopwatch clock = System.Diagnostics.Stopwatch.StartNew();
+ await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.ChangeServer($"ws://127.0.0.1:{deadPort}"));
+ clock.Stop();
+
+ Assert.IsLessThan(
+ TimeSpan.FromSeconds(15),
+ clock.Elapsed,
+ "The waiter has to be woken by the client giving up, not by its own 30s timeout.");
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// A caller who asks after the reconnect loop has given up is told the budget was spent -
+ /// the same thing the status stream told them a moment earlier.
+ ///
+ ///
+ ///
+ /// The natural shape of a failover is to hear
+ /// on the status stream and then
+ /// confirm it on the wait before switching servers. That caller arrives after the sequence
+ /// has ended - no socket, no cancellation source, a Disconnected state - which is also,
+ /// exactly, what a client nobody has called Connect() on looks like. The wait used
+ /// to answer that reading first, so the same client reported ReconnectExhausted or
+ /// NotConnecting depending only on whether the caller had happened to park before
+ /// the loop stopped, and the outcome this API exists to deliver was the one a consumer
+ /// could not get.
+ ///
+ ///
+ /// The assertion is made against the status stream rather than against a literal: what
+ /// matters is that the two ways of asking agree.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUTheWaitAndTheStatusStreamAgreeAfterTheBudgetIsSpent()
+ {
+ int deadPort = TestUtils.GetFreePort(); // nothing is listening there, and never will be
+ _client = new XrplClient($"ws://127.0.0.1:{deadPort}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(50),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(100),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromMilliseconds(500),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ TaskCompletionSource stopped =
+ new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously);
+
+ _client.connection.OnConnectionStatus += info =>
+ {
+ if (info.StopReason == ConnectionStopReason.ReconnectExhausted)
+ {
+ stopped.TrySetResult(true);
+ }
+ };
+
+ try
+ {
+ await _client.Connect();
+ }
+ catch (NotConnectedException)
+ {
+ // What this caller was told is not the subject; what a later one is told is.
+ }
+
+ await stopped.Task.WaitAsync(TimeSpan.FromSeconds(20));
+
+ ConnectionWaitOutcome outcome =
+ await _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(1));
+
+ Assert.AreEqual(
+ ConnectionWaitOutcome.ReconnectExhausted,
+ outcome,
+ "The status stream said the budget was spent; a caller asking straight afterwards must be told the same.");
+
+ ReconnectExhaustedException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.WaitForConnectionAsync(TimeSpan.FromSeconds(1)));
+
+ Assert.AreEqual(2, error.MaxAttempts, "The budget the client was configured with.");
+ }
+
+ ///
+ /// A close code this client does not reconnect after ends the wait, rather than leaving it
+ /// to run out its timeout.
+ ///
+ ///
+ ///
+ /// 1002, 1003, 1007 and 1010 are the codes after which no reconnect is started at all -
+ /// the node is saying that retrying against it is pointless. The status stream has reported
+ /// that ending for as long as has
+ /// existed; the wait could not, because the ending leaves nothing behind to recognise it
+ /// by: no socket, no reconnect loop, and an attempt counter the close path resets to zero.
+ /// A caller already parked was woken by the announcement, found nothing, and parked again -
+ /// to be told, a whole acquisition timeout later, that the connection "was not established
+ /// in time" about a connection the consumer had already been told was over.
+ ///
+ ///
+ /// The waiter here arrives after the close, which is answered through the wait's entry;
+ /// the parked case is the same terminal reached through the wake.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUAPermanentCloseEndsTheWaitInsteadOfRunningItOut()
+ {
+ using ClosesWithCodeServer server = new ClosesWithCodeServer(closeCode: 1003);
+
+ _client = new XrplClient(server.Url, new XrplClient.ClientOptions
+ {
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(5),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+ Assert.IsTrue(_client.connection.IsConnected(), "Precondition: connected to the server.");
+
+ TaskCompletionSource closed =
+ new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously);
+
+ _client.connection.OnConnectionStatus += info =>
+ {
+ if (info.StopReason == ConnectionStopReason.ClosedPermanently)
+ {
+ closed.TrySetResult(true);
+ }
+ };
+
+ server.CloseNow();
+ await closed.Task.WaitAsync(TimeSpan.FromSeconds(20));
+
+ // A generous timeout on purpose: a wait that has to be answered by a terminal must not
+ // be seen to pass because its own deadline was short.
+ ConnectionClosedPermanentlyException error =
+ await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.WaitForConnectionAsync(TimeSpan.FromSeconds(30)));
+
+ Assert.IsInstanceOfType(error, "catch (NotConnectedException) must keep catching this.");
+
+ ConnectionWaitOutcome outcome =
+ await _client.connection.WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(30));
+
+ Assert.AreEqual(
+ ConnectionWaitOutcome.ClosedPermanently,
+ outcome,
+ "The outcome and the stop reason are two views of one event and must name it the same.");
+ }
+
+ ///
+ /// A wait whose timeout is an exact multiple of the system clock tick still ends on its own
+ /// deadline, and ends promptly.
+ ///
+ ///
+ /// Both readings of the clock are quantised to the tick - 15.625 ms on Windows - so the
+ /// elapsed time is always a whole number of ticks, and a timeout that is itself a whole
+ /// number of them is hit exactly rather than passed. A strict deadline comparison then
+ /// declined to fire while the remaining time was already zero, and the loop went round
+ /// again with nothing to await: a spin on the connection's own lock until the clock moved.
+ /// One second is 64 ticks exactly, as is the default acquisition timeout of five minutes.
+ ///
+ [TestMethod]
+ public async Task TestUAWaitWhoseTimeoutLandsOnAClockTickStillTimesOut()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromSeconds(30),
+ ReconnectMaxDelay = TimeSpan.FromSeconds(30),
+ MaxReconnectAttempts = 50,
+ StopAfterMaxAttempts = false,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(4),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+
+ int deadPort = TestUtils.GetFreePort();
+ Task switching = _client.connection.ChangeServer($"ws://127.0.0.1:{deadPort}");
+
+ Stopwatch clock = Stopwatch.StartNew();
+ await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.WaitForConnectionAsync(TimeSpan.FromSeconds(1)));
+ clock.Stop();
+
+ Assert.IsLessThan(
+ TimeSpan.FromSeconds(10),
+ clock.Elapsed,
+ "The wait has to end on its own deadline rather than spin past it.");
+
+ try
+ {
+ await switching;
+ }
+ catch (Exception)
+ {
+ // The switch to a dead port is scenery; it never settles, and how it gives up
+ // is the subject of other tests.
+ }
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// A client that gave up comes back when the consumer asks it to, and the switch that asks
+ /// is not answered with the ending the previous one had.
+ ///
+ ///
+ ///
+ /// The whole point of a terminal ending is that the consumer decides what happens next, and
+ /// the decision they make is usually this one: move to another node. Recovery was covered
+ /// through Connect() and, on the WebAssembly stand, through ChangeServer;
+ /// this pins the second in the suite. What it guards against specifically is the wait
+ /// concluding, from state the previous sequence left behind, that a switch which has made
+ /// no attempts has already run out of them - every field it reads to decide that is keyed
+ /// by generation, and this is the test that says so.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUAClientThatGaveUpSwitchesToALiveServer()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ await _client.Connect();
+ Assert.IsTrue(_client.connection.IsConnected(), "Precondition: connected to the mock.");
+
+ int deadPort = TestUtils.GetFreePort(); // nothing is listening there, and never will be
+
+ await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.ChangeServer($"ws://127.0.0.1:{deadPort}"));
+
+ // The consumer's decision, and the client has to honour it rather than answer with
+ // the ending of the sequence that is over.
+ await _client.connection.ChangeServer($"ws://127.0.0.1:{port}");
+
+ Assert.IsTrue(
+ _client.connection.IsConnected(),
+ "A client that gave up must still switch to a server that answers.");
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ ///
+ /// Every reason the client can stop for has exactly one outcome a waiter can be told, under
+ /// the same name.
+ ///
+ ///
+ ///
+ /// The two enums are two views of one event, and the defect they exist to remove comes back
+ /// the moment they stop corresponding: an ending announced on the status stream that the
+ /// wait has no way to express leaves a parked caller to sit out its whole timeout and be
+ /// told the connection "was not established in time" about a connection the consumer has
+ /// already been told was over. That happened twice - a close the client does not reconnect
+ /// after, and a first attempt with no retry behind it - and both were found by readers
+ /// rather than by this suite.
+ ///
+ ///
+ /// Asserted over the enum itself rather than over a list written out here, so that a reason
+ /// added later fails this test until it is given its counterpart.
+ /// is excluded: it is the absence of an ending.
+ ///
+ ///
+ [TestMethod]
+ public void TestUEveryStopReasonHasAnOutcomeOfTheSameName()
+ {
+ List missing = new List();
+
+ foreach (ConnectionStopReason reason in Enum.GetValues())
+ {
+ if (reason == ConnectionStopReason.None)
+ {
+ continue;
+ }
+
+ if (!Enum.TryParse(typeof(ConnectionWaitOutcome), reason.ToString(), ignoreCase: false, out object _))
+ {
+ missing.Add(reason.ToString());
+ }
+ }
+
+ Assert.AreEqual(
+ 0,
+ missing.Count,
+ $"ConnectionStopReason values with no ConnectionWaitOutcome of the same name: {string.Join(", ", missing)}. " +
+ "A consumer told this on the status stream has no way to be told it by the wait.");
+ }
+
+ ///
+ /// A request issued after the client gave up is told that it gave up, not that it was never
+ /// asked to connect.
+ ///
+ ///
+ /// The check every request passes through asks the same question the wait used to ask
+ /// first - is a socket or a reconnect source installed, and is the state anything other
+ /// than Disconnected - and a client that spent its budget answers no to all three, exactly
+ /// as a client nobody has called Connect() on does. So the status stream said the
+ /// endpoint had given up, the wait agreed, and a request on the same client raised
+ /// "no connection attempt in progress. Call Connect() first" - the one distinction this
+ /// exception family exists to draw, contradicted by the third way of asking.
+ ///
+ [TestMethod]
+ public async Task TestUARequestAfterTheClientGaveUpNamesGivingUp()
+ {
+ int deadPort = TestUtils.GetFreePort(); // nothing is listening there, and never will be
+ _client = new XrplClient($"ws://127.0.0.1:{deadPort}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(50),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(100),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromMilliseconds(500),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(10),
+ UseCustomPing = false,
+ });
+
+ TaskCompletionSource stopped =
+ new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously);
+
+ _client.connection.OnConnectionStatus += info =>
+ {
+ if (info.StopReason == ConnectionStopReason.ReconnectExhausted)
+ {
+ stopped.TrySetResult(true);
+ }
+ };
+
+ try
+ {
+ await _client.Connect();
+ }
+ catch (NotConnectedException)
+ {
+ // What this caller was told is not the subject.
+ }
+
+ await stopped.Task.WaitAsync(TimeSpan.FromSeconds(20));
+
+ ReconnectExhaustedException error = await Assert.ThrowsExactlyAsync(
+ async () => await _client.connection.Request(new Dictionary
+ {
+ { "command", "server_info" },
+ }));
+
+ Assert.AreEqual(2, error.MaxAttempts, "The budget the client was configured with.");
+ Assert.IsInstanceOfType(error, "catch (NotConnectedException) must keep catching this.");
+ }
+
+ ///
+ /// A client that has been announced as stopped does not report itself connected, however
+ /// open the socket it is still holding happens to be.
+ ///
+ ///
+ ///
+ /// The give-up path announces the ending, then rejects the requests in flight - which is
+ /// what resumes the caller inside Connect() - and only then disconnects. Between the
+ /// rejection and the disconnect the socket is still installed and still open, so a caller
+ /// asking in that moment was told the connection was up, one instant after being told why
+ /// it was over. The first failure shape there is: an operation reporting success while the
+ /// state it describes is gone.
+ ///
+ ///
+ /// The status notification is the pause point that makes the window deterministic rather
+ /// than raced. It is raised before the rejection, so inside the handler the socket is
+ /// guaranteed to be open and the ending is guaranteed to be recorded - the exact
+ /// interleaving that CI produced and this machine did not.
+ ///
+ ///
+ [TestMethod]
+ public async Task TestUAStoppedClientDoesNotCallItselfConnected()
+ {
+ int port = TestUtils.GetFreePort();
+ CreateMockRippled mock = StartMock(port);
+
+ try
+ {
+ _client = new XrplClient($"ws://127.0.0.1:{port}", new XrplClient.ClientOptions
+ {
+ ReconnectBaseDelay = TimeSpan.FromMilliseconds(100),
+ ReconnectMaxDelay = TimeSpan.FromMilliseconds(200),
+ MaxReconnectAttempts = 2,
+ StopAfterMaxAttempts = true,
+ ConnectionAttemptTimeout = TimeSpan.FromSeconds(2),
+ ConnectionAcquisitionTimeout = TimeSpan.FromSeconds(30),
+ UseCustomPing = false,
+ });
+
+ _client.connection.OnConnected += () => throw new InvalidOperationException("handler is permanently broken");
+
+ bool? socketWasStillOpen = null;
+ ConnectionWaitOutcome? answeredInsideTheWindow = null;
+
+ _client.connection.OnConnectionStatus += status =>
+ {
+ if (status.StopReason != ConnectionStopReason.ConnectHandlerFailed ||
+ answeredInsideTheWindow != null)
+ {
+ return;
+ }
+
+ socketWasStillOpen = _client.connection.IsConnected();
+ answeredInsideTheWindow = _client.connection
+ .WaitForConnectionOutcomeAsync(TimeSpan.FromSeconds(5))
+ .GetAwaiter()
+ .GetResult();
+ };
+
+ await Assert.ThrowsExactlyAsync(async () => await _client.Connect());
+
+ Assert.IsNotNull(answeredInsideTheWindow, "The terminal notification has to be raised for this to test anything.");
+
+ Assert.AreEqual(
+ ConnectionWaitOutcome.ConnectHandlerFailed,
+ answeredInsideTheWindow,
+ socketWasStillOpen == true
+ ? "Asked while the socket was still open, and the client had already been announced as stopped."
+ : "Asked after the socket had closed, so this run did not exercise the window - but the answer is the same either way.");
+ }
+ finally
+ {
+ mock.Stop();
+ }
+ }
+
+ }
+}
diff --git a/Xrpl/Client/ConnectionManager.cs b/Xrpl/Client/ConnectionManager.cs
index 1b5291f0..dbeb9106 100644
--- a/Xrpl/Client/ConnectionManager.cs
+++ b/Xrpl/Client/ConnectionManager.cs
@@ -1,4 +1,4 @@
-using System;
+using System;
using System.Collections.Generic;
using System.Threading.Tasks;
@@ -6,45 +6,92 @@
namespace Xrpl.Client
{
+ ///
+ /// Holds the callers waiting for the connection to come up, and releases them when it does.
+ ///
+ ///
+ ///
+ /// The connection notifies this from its socket, reconnect and caller threads, while consumers
+ /// register from theirs: client.connection.connectionManager is reachable from outside
+ /// the library. The list is therefore shared state and is guarded as such - it was not, and a
+ /// registration landing inside a notification threw
+ /// ("Collection was modified") or, worse, was dropped
+ /// by the reassignment that followed and never resumed.
+ ///
+ ///
+ /// Waiters are released outside the lock and resume asynchronously. Both matter for the same
+ /// reason: is called from inside the connection's own critical
+ /// path, before the OnConnected handler, and a waiter resuming there would run consumer
+ /// code in the middle of a connection being assembled.
+ ///
+ ///
public class ConnectionManager
{
- private List<(Action resolve, Action reject)> PromisesAwaitingConnection = new List<(Action resolve, Action reject)>();
+ private readonly object _waitersLock = new object();
- public void ResolveAllAwaiting()
+ private List> _promisesAwaitingConnection = new List>();
+
+ ///
+ /// Takes the waiters out, leaving an empty list behind for anyone registering next.
+ ///
+ private List> TakeWaiters()
{
- foreach (var (resolve, _) in PromisesAwaitingConnection)
+ lock (_waitersLock)
{
- resolve();
+ List> waiting = _promisesAwaitingConnection;
+ _promisesAwaitingConnection = new List>();
+ return waiting;
}
+ }
- PromisesAwaitingConnection = new List<(Action resolve, Action reject)>();
+ /// Releases everyone waiting: the connection is up.
+ public void ResolveAllAwaiting()
+ {
+ foreach (TaskCompletionSource