fix(nodes): stop a restarting panel from reporting itself as down

Adding a node fails right after that node's panel restarts. nodes/add
probes the node's /panel/api/server/status first, and that endpoint
returns whatever the @2s ticker last sampled - nil until the first tick
lands, so the master reads a healthy panel as unreachable and rejects it
with "Add node (remote returned success=false: )", an error whose
message is empty because the node answered success with a null obj.

The window is far wider than one tick: GetStatus resolved the public
IPv4/IPv6 addresses inline and held s.mu across every lookup, so a box
with no IPv6 route spent 3s per service - about 15s of nil status after
each restart, and the same stall on a fresh panel's first sample.

- status now answers from CurrentStatus, which samples on demand when
  the ticker has not run yet instead of returning a null obj
- the public-IP lookups run in the background and outside s.mu, so a
  status sample never waits on them
- probe tells "no status yet" apart from a genuine success=false, so the
  master's error says something when it meets an older node
This commit is contained in:
Farhan Zare
2026-09-18 06:25:24 -04:00
committed by GitHub
parent 1c0ce80e8e
commit 95f19b192f
6 changed files with 237 additions and 29 deletions
+7 -1
View File
@@ -1323,10 +1323,16 @@ func (s *NodeService) probe(ctx context.Context, n *model.Node, proxyURL string)
patch.LastError = "decode response: " + err.Error()
return patch, err
}
if !envelope.Success || envelope.Obj == nil {
if !envelope.Success {
patch.LastError = "remote returned success=false: " + envelope.Msg
return patch, errors.New(patch.LastError)
}
// A panel that has not sampled its status yet answers success with a null
// obj; saying so beats "success=false: " with nothing after the colon.
if envelope.Obj == nil {
patch.LastError = "remote panel reported no status yet; it may still be starting up"
return patch, errors.New(patch.LastError)
}
o := envelope.Obj
patch.CpuPct = o.CpuPct
if o.Mem.Total > 0 {