Skip to content

fix: cancel stream attempts abandoned by the establishment timeout - #388

Merged
merlimat merged 2 commits into
oxia-db:mainfrom
merlimat:fix/cancel-abandoned-attempts
Sep 25, 2026
Merged

merlimat merged 2 commits into
oxia-db:mainfrom
merlimat:fix/cancel-abandoned-attempts

Conversation

@merlimat

Copy link
Copy Markdown
Collaborator

Motivation

Since #370, read, list, rangeScan and getShardAssignments bound the establishment of each attempt with requestTimeout. When it fires, TIMEOUT is reported to the observer, but the gRPC call is not cancelled: its later events are dropped by GuardedStreamObserver or by the terminated CancelableStreamObserver, yet it stays open on the server. getShardAssignments is bounded by the subscription max age, while read, list and rangeScan have no deadline. So a server that hangs but still passes the connection health checks accumulates one open call per timed-out attempt, until it answers.

Modifications

Apply to these four RPCs the approach that #386 uses for getNotifications: run each attempt in its own Context.current().withCancellation(), and cancel it in .exceptionally, once the error has been reported to the observer.

  • The error is reported first, so that the observer gets TIMEOUT rather than the CANCELLED status of the cancelled call.
  • The context is cancelled with a null cause: with a TimeoutException cause, gRPC would close the call with DEADLINE_EXCEEDED.
  • For list and rangeScan, CancelableStreamObserver.cancel() can't be used, as it does nothing once the observer is terminated.

Only the attempt that timed out can be left open: TIMEOUT is not retryable, so it is always the last attempt, and earlier attempts end when their call is closed or fails to start. The contexts of successful attempts are not cancelled, which doesn't leak: a call registers its cancellation listener on the context when it starts, and removes it when it closes.

Verification

  • The *TimesOutSilentAttempt tests of these four RPCs now check that the server sees the cancellation. Their silent handlers no longer block on a latch, which also held back the server's onCancel callback, and register ServerCallStreamObserver.setOnCancelHandler instead. Without the fix, all four fail on the new check; with it, they passed 10 repeated runs.
  • GrpcRpcProviderTest and the rest of the unit tests pass locally. The Docker-based ITs are left to CI.

Since oxia-db#370, read, list, rangeScan and getShardAssignments bound the
establishment of each attempt with requestTimeout. When it fires, TIMEOUT
is reported to the observer, but the gRPC call is not cancelled: its
later events are dropped, yet it stays open on the server.
getShardAssignments is bounded by the subscription max age, while read,
list and rangeScan have no deadline. So a server that hangs but still
passes the connection health checks accumulates one open call per
timed-out attempt, until it answers.

Run each attempt in its own cancellable gRPC context, and cancel it once
the error has been reported. The error is reported first, so that the
observer gets TIMEOUT rather than the CANCELLED status of the cancelled
call. The context is cancelled without a cause, since a TimeoutException
cause would close the call with DEADLINE_EXCEEDED. For list and
rangeScan, CancelableStreamObserver.cancel() can't be used, as it does
nothing once the observer is terminated.

The *TimesOutSilentAttempt tests now check that the server sees the
cancellation. Their silent handlers no longer block, since a blocked
handler also holds back the server's onCancel callback.

Signed-off-by: Matteo Merli <mmerli@apache.org>
…bd2ec2

Signed-off-by: Matteo Merli <mmerli@apache.org>
@merlimat
merlimat merged commit 12b7541 into oxia-db:main Sep 25, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant