Repository navigation
Conversation
# Conflicts: # hadoop-hdds/container-service/src/main/java/org/apache/hadoop/ozone/container/common/transport/server/ratis/ContainerStateMachine.java
| * | ||
| * @return the number of bytes written to the stream | ||
| */ | ||
| static long streamReadBlock(ContainerDispatcher dispatcher, ContainerCommandRequestProto request, |
There was a problem hiding this comment.
Curious why we make this as static? Is it for easy testing?
There was a problem hiding this comment.
just because it doesnt depend on ContainerStateMachine.
| CertificateClient caClient, StateContext context) throws IOException { | ||
| Parameters parameters = createTlsParameters( | ||
| new SecurityConfig(ozoneConf), caClient); | ||
| if (parameters == null) { |
There was a problem hiding this comment.
In this case, createTlsParameters could just return new Parameters() than null?
| } catch (IOException e) { | ||
| error.set(e); | ||
| // Throw to stop reading the block | ||
| throw new UncheckedIOException(e); |
There was a problem hiding this comment.
Could we let UncheckedIOException propagate through KeyValueHandler.readBlock? It currently becomes CONTAINER_INTERNAL_ERROR and triggers a container scan on client disconnect. Please add a test asserting no scan is triggered.
| } | ||
|
|
||
| final Container<?> container = containerController.getContainer(requestProto.getContainerID()); | ||
| if (container == null || container.getContainerState() != CLOSED) { |
There was a problem hiding this comment.
Should this also accept QUASI_CLOSED? SCM returns a read pipeline without an existing Raft group for that state too.
|
@peterxcli thanks for the patch! |
Adapt the tests to HDDS-16258: ReadBlockResponseProto no longer has checksumData, and buildReadBlockCommandProto takes includeChecksums.
…elled When a reply cannot be written, the Ratis read stream now cancels the read with a CANCELLED status. KeyValueHandler.readBlock passes a CANCELLED status on instead of turning it into CONTAINER_INTERNAL_ERROR, which made the dispatcher scan the container. This also covers the gRPC streaming read, whose onNext throws CANCELLED once the client has gone away.
… is not open or closing SCM gives every container that is not open or closing a read pipeline with a random ID, so the resolver now serves those containers instead of only CLOSED ones.
# Conflicts: # hadoop-hdds/container-service/src/test/java/org/apache/hadoop/ozone/container/common/transport/server/ratis/ContainerStateMachineTests.java
…ir own Also say in the resolver why it checks isClosing, and in the Ratis read stream why it throws CANCELLED.
What changes were proposed in this pull request?
HDDS-9904 adds block reads over Ratis data streams. Ratis 3.3.1 supports read-only data streams (RATIS-1240): the client sends one request, and the server sends the data back in a sequence of replies. This PR adds the datanode side for ReadBlock. The client side comes in later sub-tasks of HDDS-9904.
Changes:
ContainerStateMachine#transferToserves ReadBlock and rejects other commands. It passes the request to the dispatcher's streaming read (streamDataReadOnly, the same one the gRPC streaming read uses) and writes each response to the stream as one reply, with layout:[int: metadata length][the response without its data][data].CONTAINER_NOT_FOUND, is sent as a reply with no data, so the client still gets the result code.CANCELLEDstatus.KeyValueHandler.readBlockpasses aCANCELLEDstatus on instead of returningCONTAINER_INTERNAL_ERROR, so the datanode does not scan the container. This also covers the gRPC streaming read, whoseonNextthrowsCANCELLEDonce the client is gone.ClosedContainerReadResolveris a RatisDataStreamApi.Resolver(RATIS-2603), registered inXceiverServerRatis. Ratis asks it before it looks up the Raft group. It serves the read by container ID when the request is a ReadBlock, the container is not open or closing, and the request names this datanode. SCM gives every such container a read pipeline with a random ID. For any other request it returns null, and Ratis uses the Raft group as before. Such a container no longer changes through Raft, so its read needs neither the group nor the leader.Block tokens are checked as in the gRPC read path: the dispatcher validates the token when the
DispatcherContextis null.Each reply is built in one heap buffer, which copies the data once more. Removing that copy is left to a later sub-task, because it requires:
streamDataReadOnlyandHandler.readBlocktake a new observer type that keeps data outside the protobuf and supplies the buffer to read into.readBlockImplso that it reads into that buffer. Now that HDDS-16258. Refactor streaming block reads and fix checksum verification for variable-sized chunks #11302 is merged, this is small:BlockReadCursorknows the length and the chunks of each read before the read.What is the link to the Apache JIRA
https://issues.apache.org/jira/browse/HDDS-16722
How was this patch tested?
ContainerStateMachineTests(run byTestContainerStateMachineLeaderandTestContainerStateMachineFollower):TestClosedContainerReadResolver: the resolver serves ReadBlock only for a container on this datanode that is not open or closing.TestHddsDispatcher#testReadBlockScansContainerOnlyIfTheStreamIsNotCancelled: a cancelled ReadBlock stream does not make the datanode scan the container; any other failure still does.TestStreamRead#testRatisStreamReadBlock: on a mini cluster, a plain Ratis client reads a block on a read-only stream, first from the open container through its Raft group, then, after the container is closed, through the read pipeline from SCM, which has a random ID and no Raft group. Both reads match the block file.