perf: reuse SSH connections per server, with a retry rule that never repeats work
Every ssh.exec opened its own connection: a TCP handshake, a key exchange and an
authentication round trip per command. A key rotation paid for that eight times,
a deployment six, and refreshing M profile states M times.
Connections are now kept per server. The three risks that made this worth doing
carefully are handled explicitly:
- Staleness. A pooled connection can be dead exactly when it matters. Liveness is
tracked through error, close and end, and a lease that finds a dead entry opens
a new one. The remaining race, where the connection dies between the check and
the command, is caught by the retry rule below.
- Retrying. Only a failure that proves the command never reached the server is
retried, and only once, and only on a connection that was already established
before this call. execClient marks exactly that case, when the channel fails to
open. A command that opened a stream is never repeated, because the server may
already be acting on it - repeating a deployment is not this layer's decision.
Two tests hold that line: widening the rule to any failure fails both.
- Lifetime. Idle connections close after a minute, the pool is reference counted
so a shared connection survives until its last user is done, closeAll runs
during quit, and every pooled client keeps a standing error listener so an
error while idle cannot reach the uncaughtException handler.
A trust-on-first-use connection is never pooled: it was established without
verifying the fingerprint, so it must not serve a later verified call. A change
to host, port, user, auth type, key path or trusted fingerprint invalidates the
pooled connection.
ssh-service coverage rises from 61% to 90% of lines and 97% of functions.
Also in this commit, the smaller items from the same review:
- Diagnostics batched records that queue up while a write is in flight into one
append, and chmod runs once per file instead of once per record. At the debug
level every IPC call writes a line, which is exactly when troubleshooting.
- The set that suppresses duplicate deployment notifications is trimmed instead
of growing for the lifetime of the process.
- The updater kept the same once('error') pattern on its spawned helper that
took the app down through the SSH client.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
beeafdcba7
commit
cb9bdcd713
@@ -40,6 +40,7 @@ const {
|
||||
|
||||
let mainWindow;
|
||||
let repositoryMonitor;
|
||||
let sshService;
|
||||
let operationTimer;
|
||||
let diagnostics;
|
||||
let configStore;
|
||||
@@ -252,6 +253,7 @@ app
|
||||
const repositories = new RepositoryService(store, git, gitea, diagnostics);
|
||||
const deployments = new DeploymentService(store, gitea, git, diagnostics);
|
||||
const ssh = new SshService({ store, diagnostics });
|
||||
sshService = ssh;
|
||||
const auditedOperationStates = new Set();
|
||||
const reportOperationChange = (payload) => {
|
||||
broadcast("operations:changed", payload);
|
||||
@@ -262,6 +264,11 @@ app
|
||||
) {
|
||||
const key = `${operation.id}:${operation.status}`;
|
||||
if (!auditedOperationStates.has(key)) {
|
||||
// One entry per completed deployment, so the set is trimmed rather
|
||||
// than kept for the lifetime of the process.
|
||||
if (auditedOperationStates.size >= 500) {
|
||||
auditedOperationStates.delete(auditedOperationStates.values().next().value);
|
||||
}
|
||||
auditedOperationStates.add(key);
|
||||
notify(
|
||||
`Deployment ${operation.status}`,
|
||||
@@ -450,6 +457,7 @@ app.on("before-quit", (event) => {
|
||||
event.preventDefault();
|
||||
quitCleanupStarted = true;
|
||||
repositoryMonitor?.stop();
|
||||
sshService?.closeAll();
|
||||
if (operationTimer) clearTimeout(operationTimer);
|
||||
Promise.resolve()
|
||||
.then(() => diagnostics?.info("app.quitting", {}))
|
||||
|
||||
Reference in New Issue
Block a user