Fix DataNode status handling for Pipe receiver disk failures - #18484
Open
Caideyipi wants to merge 2 commits into
Open
Fix DataNode status handling for Pipe receiver disk failures#18484Caideyipi wants to merge 2 commits into
Caideyipi wants to merge 2 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Avoid duplicate node status transitions
CommonConfig#setNodeStatusnow returns immediately when the requested status is already active. This preserves the existing status reason and prevents repeatedReadOnly -> ReadOnlytransition logs from causing a log storm.Preserve the existing storage-disk recovery semantics
DataNode heartbeat entry and recovery decisions continue to use the original aggregate available/total disk ratio. Recovery does not require every
FileStoreto exceed the warning threshold, so the conditions for restoringRunningare not made stricter.Pipe receiver-only directories are not added to the system/storage disk aggregate. If Pipe and the storage engine use separate disks, a full Pipe disk therefore does not affect the DataNode status. If they share a disk, the storage-engine directory still makes that disk part of the normal aggregate.
Report Pipe disk failures without setting global ReadOnly
FolderManagerkeeps its existing behavior by default, while allowing Pipe receiver managers to disable the globalReadOnlyside effect. Both the regular DataNode Pipe receiver and the IoTConsensusV2 receiver use this mode. Disk-space exceptions still propagate through the existing receiver response path, so the sender receives an error and can report/retry it.Tests
CommonConfigTest: verifies that a repeatedReadOnlyupdate preservesDISK_FULL.DataNodeInternalRPCServiceImplDiskTest: verifies aggregate-ratioRunningrecovery remains unchanged and a below-threshold storage aggregate still entersReadOnly.FolderManagerTest: verifies all directory strategies throwDiskSpaceInsufficientExceptionfor a full Pipe receiver without changing the global node status.Executed:
This PR has:
Key changed/added classes (or packages if there are too many classes) in this PR
CommonConfigFolderManagerDirectoryStrategyDataNodeInternalRPCServiceImplIoTDBDataNodeReceiverIoTConsensusV2Receiver