feat: multi-backend gRPC addresses via static pick_first resolver - #256
Conversation
Support comma-separated reporter.grpc.backend_service (>=2 after normalize)
with a static endpoint list and pick_first, aligned with Node sw-static
failover. Single-address paths keep historical behavior.
Multi-backend only: dial/RPC timeouts, proxy bypass, auth-failure throttled
logs, Collect Send/CloseAndRecv bounds, in-place RecreateConnection on
half-open peers, and UNAVAILABLE retries limited to reportInstanceProperties
so streaming Collect cannot replay segments.
GetConnection shares the connection-manager mutex with Recreate/Release/Peek.
Multi-backend status watcher skips Connect on Shutdown ClientConns.
MultiBackendSend cancels the stream context on deadline and uses a buffered
done channel so a timed-out Send does not need a second drain goroutine;
Recreate closes the prior conn to unblock it.
gRPC service stubs and PprofTaskClient are published as immutable bundles via
atomic.Pointer so bind/recreate cannot race with send, heartbeat, profile, or
pprof poll/upload goroutines; multi-backend pprof upload retries once after
refreshing the stub.
Tests: mid-stream kill->standby, concurrent Get/Recreate/Release, TestRace*
(including client-bundle and pprof-client swap), service-config retry scope,
hung-send unblock, reporter auto-failover for trace/metrics/log without
manual RecreateConnection. CI runs go test -race on reporter packages. E2E
resolves active via unique /sw-failover-probe/{token} and requires standby
POST:/info + toolkit log growth, with bounded waits/retries so the case
stays under the GHA job timeout.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved critical and moderate findings remain in connection lifecycle, stream retry/recovery, pprof handling, and E2E configuration.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 4
Open (5)
Goroutine panics bypass sendWithRecover and can crash the process · New Concurrent acquisition can receive a connection closed during recreation · New Retrying streaming Collect replays segments and duplicates data · New Expected fixture path does not match the e2e case layout · New Recreate connection after shutdown and leak the new channel · New
What changed in this PR
Adds multi-address gRPC backend support with static resolution and pick_first failover while preserving single-address behavior.
Changes:
- Adds backend normalization, static resolution, reconnection, and bounded operations.
- Updates reporter, pprof, CDS, tooling, CI, and configuration support.
- Adds unit, race, and end-to-end failover coverage and documentation.
| File | Summary |
|---|---|
tools/go-agent/tools/dst.go |
Extends package-reference rewriting. |
tools/go-agent/tools/dst_test.go |
Adds AST regression tests. |
tools/go-agent/config/agent.default.yaml |
Documents multi-address configuration. |
test/e2e/case/grpc-multi-backend/verify-failover.sh |
Verifies backend failover. |
test/e2e/case/grpc-multi-backend/expected/failover-ok.yml |
Defines expected E2E output. |
test/e2e/case/grpc-multi-backend/e2e.yaml |
Registers the E2E scenario. |
test/e2e/case/grpc-multi-backend/docker-compose.yml |
Defines the multi-backend topology. |
test/e2e/base/consumer/main.go |
Adds failover probe endpoints. |
plugins/core/reporter/static_backend_resolver.go |
Implements static gRPC resolution. |
plugins/core/reporter/pprof_manager.go |
Adds swappable pprof clients and retries. |
plugins/core/reporter/multi_backend_stream_test.go |
Tests stream failover and concurrency. |
plugins/core/reporter/grpc/multi_backend_failover_test.go |
Tests reporter failover. |
plugins/core/reporter/grpc/grpc.go |
Integrates multi-backend reporter logic. |
plugins/core/reporter/grpc/client_bundle_race_test.go |
Tests atomic client swaps. |
plugins/core/reporter/conn_manager.go |
Manages multi-backend connections and recreation. |
plugins/core/reporter/cds_manager.go |
Adds bounded multi-backend CDS calls. |
plugins/core/reporter/backend_addresses.go |
Normalizes backend address lists. |
plugins/core/reporter/backend_addresses_test.go |
Tests backend address parsing. |
plugins/core/reporter/auth_failure_log.go |
Throttles authentication failure logs. |
plugins/core/reporter/auth_failure_log_test.go |
Tests authentication logging. |
docs/menu.yml |
Adds troubleshooting navigation. |
docs/en/advanced-features/grpc-backend-troubleshooting.md |
Documents backend failover and diagnostics. |
agent/reporter/imports.go |
Adds generated reporter imports. |
.github/workflows/skywalking-go.yaml |
Adds reporter race testing. |
.github/workflows/e2e.yaml |
Registers multi-backend E2E testing. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| cases: | ||
| - name: fail over to standby after active collector stops | ||
| query: bash test/e2e/case/grpc-multi-backend/verify-failover.sh | ||
| expected: expected/failover-ok.yml |
|
Cross-agent alignment follow-up, updated for
The invalid-entry policy is now implemented. The remaining decision is whether to accept the Go/Node.js DNS and authority behavior or require Python's startup-DNS and first-usable-authority semantics. |
|
Thanks for the professional follow‑up and related handling |
mrproliu
left a comment
There was a problem hiding this comment.
Thanks for the PR.
My main concern is that this adds a second implementation (dialing, connection status, send loops, heartbeat) that only runs when there are two or more addresses. I don't think the fork is needed.
gRPC's ClientConn already handles failover across addresses: pick_first is the default LB policy, and when the active backend dies the reporter just sees a failed Send, then re-opens the stream on the same ClientConn. That is exactly what the existing single-address loops already do.
Proposed shape
- Always parse
backend_serviceinto a list. When it has 2+ entries, dial through the static resolver; otherwise dial as today. That should be the only branch. - Keep the existing connection status checker and the existing long-lived stream loops. Remove
multiBackendTraceSendLoop, theIsMultiBackend()branches in the metrics/log/profile loops, and the multi-only heartbeat logic. - The
Sendtimeout and "cancel stream when the channel leaves READY" protections are useful, but they fix a real problem for single-address too (a half-open connection blocksSendforever). Enable them for both modes. serviceClients,pprofClientBundle, CDS stub re-binding andshutdownCtxonly exist to support connection re-creation, which this PR says does not happen. They can go.
This should make the change much smaller and keep one code path to maintain.
A few things to fix regardless
- In multi-backend mode the first
ReportInstancePropertieswaits a fullcheck_interval(20s) becauseConnectingis treated as disconnected. Single-address registers immediately. - The per-batch trace stream (
CloseAndRecvafter every batch) makes throughput RTT-bound. The existing long-lived stream already has the same "no replay" behavior. - TLS verifies every backend against the first address's host. Setting
resolver.Address.ServerNameper address avoids this instead of documenting it. - An all-invalid list also disables the Kafka reporter, since
initManagerruns on that path too. - Please remove the "Codex P2" / "Codex: ..." comments and the Node.js / Python references from code and CI.
| @@ -0,0 +1,81 @@ | |||
| # gRPC Backend Connectivity Troubleshooting | |||
|
|
|||
| ## Multi-address failover | |||
There was a problem hiding this comment.
Can we merge this selection into https://github.com/apache/skywalking-go/blob/main/docs/en/advanced-features/grpc-tls.md?
Address review feedback: dial is the only multi-address branch; reuse long-lived Collect loops with BoundSend (including CloseAndRecv) and cancel-on-unready for both modes; unify connection status under lock (Connecting is Connected; closed conn reports Shutdown immediately); refresh instance properties after disconnect and periodically without blocking heartbeats; drop recreate support structures; set per-address TLS ServerName; keep Kafka usable when gRPC backend_service is all-invalid; merge troubleshooting docs into grpc-tls.md and strip Codex/cross-agent noise from code/CI.
|
Thanks for the review. We converged the multi-backend work onto a single reporter path, as requested. Structural changes
Other fixes
StatusThese changes are on a single commit and have passed fork CI (unit/race/lint, plugin tests, Windows plugin tests, and multi-backend E2E). |
mrproliu
left a comment
There was a problem hiding this comment.
BoundSend: goroutine + timer per message on the hot path
After the convergence, BoundSend runs for every segment, metric batch and log line for all users. Each call spawns a goroutine, allocates a result channel and creates a timer.
grpc-go's SendMsg already returns as soon as the stream context is cancelled (the write-quota wait is interrupted by stream done), so the only thing we actually need is "cancel the context if Send takes too long". That can be done without a goroutine:
watchdog := time.AfterFunc(timeout, cancel)
err := stream.Send(s)
watchdog.Stop()with one timer reused per stream via Reset. The "send ignores cancel, wait for grace" branch guards against a situation grpc-go does not produce on a real stream; TestMultiBackendSendReturnsWhenSendIgnoresCancel only reproduces it with a time.Sleep.
Remaining multi-address-only branches
Two places still behave differently depending on the address count. Since everything else is now one path, I'd suggest removing both:
-
checkConnectionStatuspolls everycheckInterval(20s by default) for multi-address but every 5s for single-address. That makes Disconnect → Connected detection four times slower in multi mode for no reason. Theconn.Connect()nudge can also go: on grpc-go 1.55Connect()is a no-op while inTransientFailure. With those two gone, themultivariable in this function is unused. -
directTCPContextDialeris only installed for multi-address. Its effects are bypassingHTTP_PROXY/HTTPS_PROXYand changing TCP keepalive from 15s to 10s. Dial time is already bounded byMinConnectTimeout: 5s. Dropping the custom dialer removes the last behavioral difference between the two modes, and the proxy paragraph in the docs can be removed with it.
| @@ -47,6 +47,8 @@ jobs: | |||
| case: | |||
| - name: gRPC | |||
There was a problem hiding this comment.
| - name: gRPC | |
| - name: gRPC single-backend |
| } | ||
| } | ||
| r.closeTracingStream(stream) | ||
| cancel() |
There was a problem hiding this comment.
This runs on the normal shutdown path (channel closed by Close()), and it cancels the stream context before CloseAndRecv. Two consequences:
CloseAndRecvreturnsCanceledimmediately, so we no longer wait for the OAP ack. Segments still sitting in the transport write buffer at shutdown can be lost. The previous code waited for the ack here.- Every graceful shutdown now logs
send closing error context canceled.
Suggest: on this path call closeTracingStream first (BoundSend already cancels on timeout, so it can't hang), then cancel(). The error path at L266 is fine to keep as is, since the stream is already broken there.
Same applies to the metrics (L322), log (L368) and profile (L439) loops.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Clean stream shutdown can discard drained telemetry, and key regression tests do not reliably exercise their claimed failure paths.
Review effort: Balanced
Findings: 2
Open (5)
Canceling before bounded stream close can drop final trace messages · New Expected fixture path does not match the e2e case layout Per-address ServerName conflicts with fixed channel authority contract · New Test may stop the wrong backend and skip failover coverage · New Test does not observe RPC errors before assertions run · New
| cancel() | ||
| r.closeTracingStream(cancel, stream) |
| // configuredAddressesAsResolverState builds resolver addresses with ServerName | ||
| // set from each endpoint's host so TLS SNI / certificate verification follows | ||
| // the dialed backend rather than a fixed channel authority. | ||
| func configuredAddressesAsResolverState(backends []string) []resolver.Address { | ||
| addresses := make([]resolver.Address, 0, len(backends)) | ||
| for _, cfg := range backends { | ||
| host, _, err := net.SplitHostPort(cfg) | ||
| if err != nil { | ||
| addresses = append(addresses, resolver.Address{Addr: cfg}) | ||
| continue | ||
| } | ||
| addresses = append(addresses, resolver.Address{Addr: cfg, ServerName: host}) |
| aGS.Stop() | ||
| _ = aLis.Close() |
| _, err := stream.Recv() | ||
| if err == io.EOF { | ||
| return status.Error(s.code, "collector failed after receiving the segment") | ||
| } |
Close Collect streams before canceling on graceful drain; simplify BoundSend to a context watchdog without a per-message goroutine; use one status poll interval with no Connect nudge, and drop the multi-only custom TCP dialer; tighten failover/RPC-error tests and rename the single-backend E2E case.
Shutdown and BoundSend
Single-path cleanup
Docs, CI naming, and tests
StatusThese changes are on a single commit and have passed fork CI (unit/race/lint, plugin tests, Windows plugin tests, and multi-backend E2E). |
| run: make test | ||
| - name: Test Race | ||
| run: make test-race | ||
| - name: Test Race (reporter packages) |
There was a problem hiding this comment.
What’s the difference with the existing make test-race?
Add the multi-address gRPC backend feature to CHANGES.md, and remove the extra reporter race step that duplicates make test-race (TestRace*).
CHANGES.md and CI
StatusThese changes are on a single commit and have passed fork CI (unit/race/lint, plugin tests, Windows plugin tests, and multi-backend E2E). |
mrproliu
left a comment
There was a problem hiding this comment.
Generally LGTM. Once the CI all passed, it can be merge.
|
Thanks for the review |
wu-sheng
left a comment
There was a problem hiding this comment.
Thanks for converging onto a single path, it is much easier to follow now. I re-reviewed at aab0e69. Three issues should be fixed before merge (1–3); the rest can be follow-ups.
Blocking
1. TLS enabled + no valid backend_service panics at startup (instrument.go#L229)
In the generated initManager, tc, err := generateTLSCredential(...) declares a new err inside the TLS branch, so connManager, err = NewConnectionManager(...) assigns to that inner variable. The outer err stays nil when NewConnectionManager returns (nil, errNoValidBackendService), NewCDSManager then calls GetConnection on a nil *ConnectionManager, and the application crashes. That contradicts the "reporting is disabled, the application continues" behavior from dc4aa2c. The shadowing was harmless before because NewConnectionManager never returned an error.
Fix:
- tc, err := generateTLSCredential({{.Config.Reporter.GRPC.TLS.CAPath.ToGoStringValue}},
+ tc, tlsErr := generateTLSCredential({{.Config.Reporter.GRPC.TLS.CAPath.ToGoStringValue}},
{{.Config.Reporter.GRPC.TLS.ClientKeyPath.ToGoStringValue}},
{{.Config.Reporter.GRPC.TLS.ClientCertChainPath.ToGoStringValue}},
{{.Config.Reporter.GRPC.TLS.InsecureSkipVerify.ToGoBoolValue}})
- if err != nil {
- panic(fmt.Sprintf("generate go agent tls credential error: %v", err))
+ if tlsErr != nil {
+ panic(fmt.Sprintf("generate go agent tls credential error: %v", tlsErr))
}Regression test: tools/go-agent/instrument/reporter/instrument_test.go (fails at aab0e69, passes with the fix)
Lint cannot see variable scopes inside a template string, so this renders initManagerFunc, stubs the constructors, and runs the emitted Go for TLS off and on.
package reporter
import (
"html"
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
"github.com/apache/skywalking-go/tools/go-agent/config"
"github.com/apache/skywalking-go/tools/go-agent/tools"
)
func TestGeneratedInitManagerPropagatesConnectionError(t *testing.T) {
if err := config.LoadConfig(""); err != nil {
t.Fatal(err)
}
generated := html.UnescapeString(tools.ExecuteTemplate(initManagerFunc, struct {
Config *config.Config
}{Config: config.GetConfig()}))
// Exercise the emitted Go code: lint cannot inspect variable scopes inside
// a template string. Stub external constructors to isolate error propagation.
generated = strings.ReplaceAll(generated, "operator.LogOperator", "interface{}")
source := "package main\nimport (\"fmt\"; \"os\"; \"strconv\"; \"strings\"; \"time\")\n" +
generated + initManagerErrorHarness
sourcePath := filepath.Join(t.TempDir(), "main.go")
if err := os.WriteFile(sourcePath, []byte(source), 0o600); err != nil {
t.Fatal(err)
}
cmd := exec.Command("go", "run", sourcePath)
if output, err := cmd.CombinedOutput(); err != nil {
t.Fatalf("generated initManager failed to propagate connection error: %v\n%s", err, output)
}
}
const initManagerErrorHarness = `
type ConnectionManager struct{}
type CDSManager struct{}
type PprofTaskManager struct{}
var connectionError = fmt.Errorf("no valid backend service addresses")
func generateTLSCredential(string, string, string, bool) (interface{}, error) {
return struct{}{}, nil
}
func NewConnectionManager(interface{}, time.Duration, string, string, interface{}) (*ConnectionManager, error) {
return nil, connectionError
}
func NewCDSManager(interface{}, string, time.Duration, *ConnectionManager) (*CDSManager, error) {
panic("CDS must not start after connection initialization fails")
}
func NewPprofTaskManager(interface{}, string, time.Duration, *ConnectionManager, string) (*PprofTaskManager, error) {
panic("pprof must not start after connection initialization fails")
}
func main() {
for _, tlsEnabled := range []string{"false", "true"} {
if err := os.Setenv("SW_AGENT_REPORTER_GRPC_TLS_ENABLE", tlsEnabled); err != nil {
panic(err)
}
conn, cds, pprof, err := initManager(nil, time.Second)
if err != connectionError || conn != nil || cds != nil || pprof != nil {
panic(fmt.Sprintf("TLS=%s: connection error was not propagated: %v", tlsEnabled, err))
}
}
}
`2. A first address that silently drops connections prevents failover (conn_manager.go#L206)
In grpc-go v1.55.0, pick_first keeps all addresses in one addrConn, and resetTransport computes a single connectDeadline that tryAllAddrs shares across every address (clientconn.go#L1150-L1220). With MinConnectTimeout: 5s, if the first address in the shuffled list silently drops the connection attempt (host powered off or network-partitioned, no RST), each attempt spends the whole deadline on it. The next address starts with the deadline already passed and fails immediately. The channel goes TransientFailure, backs off, and repeats. The shuffle is fixed per channel, so the healthy standby is never used for the life of the process. The unit and e2e tests stop servers gracefully, so the connection is refused immediately and this case is never exercised.
Suggestion: give each address its own dial timeout, e.g. a grpc.WithContextDialer that caps each attempt at a slice of the remaining deadline, plus a test with a non-responding listener as the first address.
3. Regression: gRPC target URIs are now rejected (backend_addresses.go#L70)
On main, backend_service is passed as-is to grpc.Dial, so dns:///oap-headless:11800, unix:///path.sock, etc. work. dc4aa2c preserved that by parsing only comma-separated values. Since fd76a1a, every entry goes through net.SplitHostPort, which rejects these targets ("too many colons"). parseBackendServiceList then returns errNoValidBackendService and the reporter becomes a discard reporter with only a warning. Please restore passthrough for a single entry that is not host:port.
Should fix
4. No keepalive, so a backend that stops responding without closing the connection never triggers failover (conn_manager.go#L300)
No grpc.WithKeepaliveParams is set. BoundSend and WatchConnCancelOnUnready cancel the stream, but the transport stays READY, so pick_first keeps choosing the dead backend. Each reopened stream buffers about 64KB of Sends that return nil and are lost, then blocks for 8s, and this repeats until the OS gives up on the TCP connection (about 15 minutes). Client keepalive would close the transport and let pick_first move on. For the same reason, WatchConnCancelOnUnready (L374) rarely fires: when a READY transport is lost, pick_first goes to IDLE, not TransientFailure, and a backend that stops responding without closing the connection stays READY.
5. The 8s BoundSend also applies to single-backend setups (grpc.go#L221)
I understand this was requested for both modes, but note the tradeoff: an OAP that stops reading for more than 8s (load, long GC) now gets the Collect stream reset, and everything buffered in it is dropped, where it previously just applied backpressure. With keepalive from (4) handling dead peers, this bound could be much longer or configurable.
6. One timer is created and stopped for every message sent (conn_manager.go#L307)
time.AfterFunc + Stop per Send. One timer per stream reused via Reset gives the same bound without per-message timers.
Pre-existing, touched by this PR (fine as follow-ups)
- CDS / profile-task / pprof calls attach no auth metadata (cds_manager.go#L82, also
GetProfileTaskCommands,GetPprofTaskCommands, pprofCollect). Same onmain, but in multi-backend mode the newauthFailureLoggernow logs "check reporter.grpc.authentication" every 30s even when the token is correct. case ConnectionStatusShutdown: breakinInitCDS(cds_manager.go#L74) only exits theswitch, so the loop keeps calling a closed connection forever. This PR fixed the same pattern ingrpc.gowithreturn.
Cleanup
- Test-only / unused exports ship in every instrumented binary (conn_manager.go#L289): the no-op "deprecated"
*CancelGrace*ForTestshims, theMultiBackend*ForTestaliases,MultiBackendSend,PeekConnection,ResolvedBackendAddresses,isIPLiteralHost. The comment says they keep existing tests compiling, but those tests are new in this PR; please callBoundSend/SetBoundSendTimeoutForTestdirectly and drop the rest. BackendRPCContext/BackendStreamContexttakeserverAddrand discard it with_ = serverAddr(L319, L334); please remove the parameter.- The four
close*Streamfunctions (grpc.go#L450-L497) and the four open/watch blocks are near-identical copies; one helper each would keep future fixes in one place. - E2E: verify-failover.sh#L95 counts
POST:/infoon the standby, but the provider also produces aPOST:/infoentry span and reports to its own, independently shuffled backend. If the provider picked the standby from the start, the trace half passes without the consumer failing over. Only the log check is specific to the consumer. - The PR description is out of date: it still describes batched trace sends, waiting for READY, and the first endpoint as the fixed TLS identity, and says single-address behavior is unchanged. Please update it to match the current code.
Propagate connection errors when TLS is enabled; give each multi-backend address its own dial timeout; restore single-target gRPC URI passthrough; add client keepalive and a longer reusable BoundSend watchdog; attach auth metadata to CDS/profile/pprof RPCs; fix CDS shutdown; harden failover E2E; and drop unused test-only exports.
Move resolver-address inspection into a test-only helper so instrumented binaries no longer ship the diagnostics accessor called out in review.
|
Thanks for the review. Blocking
Should fix
Cleanup
StatusThese changes are on this PR and have passed fork CI (unit/race/lint, plugin tests, Windows plugin tests, and multi-backend E2E). |
|
Thanks for the quick turnaround. I re-reviewed at Blocking items: verified
Keepalive vs. the server ping policyThe new client keepalive (conn_manager.go#L76-L80) pings every 30s. OAP's I reproduced this with this PR's With the default Suggestion: derive the keepalive time from the heartbeat, e.g. Three or more addresses share the connect budgetEach address gets 2s (conn_manager.go#L74), but grpc-go still shares one Leftover duplication
func closeStream[R any](r *gRPCReporter, cancel context.CancelFunc,
stream interface{ CloseAndRecv() (R, error) }, errLog string) {
if err := reporter.BoundSend(cancel, func() error {
if _, err := stream.CloseAndRecv(); err != io.EOF {
return err
}
return nil
}, 0); err != nil {
r.logger.Errorf("%s %v", errLog, err)
}
}Also, |
Derive client keepalive Time from check_interval to avoid too_many_pings against default server policy; scale multi-backend MinConnectTimeout by address count; collapse CloseAndRecv into a generic helper; and run the generated initManager under stubs for TLS on/off.
Satisfy gocritic typeDefFirst so lint passes in CI.
|
Thanks for the review. Keepalive and connect budget
Cleanup and tests
StatusThese changes are on this PR and have passed fork CI (unit/race/lint, plugin tests, Windows plugin tests, and multi-backend E2E). |
|
Two additional issues remain at
|
Clear the per-address dial deadline only after the first post-handshake read (HTTP/2 SETTINGS) so a TLS-only silent peer cannot starve failover. Stop treating channel Idle as stream failure so graceful GOAWAY can drain.
Rely on embedded TransportCredentials for OverrideServerName so staticcheck SA1019 no longer fails CI lint.
|
Thanks for the review. 1. TLS handshake clearing the per-address dial deadline — Fixed. The context dialer still sets a per-address deadline for the full TCP/TLS/HTTP/2 window, but no longer clears it on the first raw socket read. A credentials wrapper clears the deadline only after 2. Canceling streams on channel Idle (graceful GOAWAY) — Fixed. Please take another look when convenient. |
…cher Keep the per-address dial deadline until the first complete HTTP/2 frame (the server SETTINGS preface) has been read. Clearing it on the first post-handshake read let a peer that sends only the 9-byte frame header and withholds the payload consume the shared connect deadline, so pick_first never reached the healthy standby. Remove WatchConnCancelOnUnready. Channel connectivity state describes the channel, not the transport a stream runs on: after a graceful GOAWAY, a refused reconnect moved the channel to TRANSIENT_FAILURE and the watcher canceled a stream that was still draining. A stream on a dead transport already fails on its own, and keepalive plus BoundSend cover dead peers and stuck sends. Add regression tests for a partial SETTINGS first peer and for a draining stream after GracefulStop with a refused reconnect. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q5YiM91L3aNmFHTZWTWTMV
The connection-state watcher was removed in the previous commit. Describe the remaining mechanisms instead: keepalive closes half-open transports and the bounded send timeout unblocks a stuck Send. Also update the dial-deadline comments to say the deadline clears after the server's first complete HTTP/2 frame, not the first post-handshake read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q5YiM91L3aNmFHTZWTWTMV



Summary
Support comma-separated
reporter.grpc.backend_serviceaddresses on one shared gRPC channel with nativepick_firstfailover. Dial is the only multi-address branch; reporters reuse the existing long-lived Collect loops for both single- and multi-address modes.Behavior
host:porttarget (for exampledns:///...orunix:///...) is passed through togrpc.Dial.ServerNamefrom that host (explicit credential overrides still win).ClientConn: a failed Send reopens Collect on the same channel without replaying telemetry. Only idempotentreportInstancePropertiesretriesUNAVAILABLE.check_interval. Client keepalive and per-address dial timeouts help leave dead or silent peers sopick_firstcan move on.Validation