HDDS-16014. [Ozone Versioning] [T3] Enabled Write Path - #10895
Open
symious wants to merge 13 commits into
Open
Conversation
symious
force-pushed
the
HDDS-16014
branch
2 times, most recently
from
July 31, 2026 07:20
0207cb8 to
0737502
Compare
symious
force-pushed
the
HDDS-16014
branch
4 times, most recently
from
August 24, 2026 07:51
dfbc533 to
90cd7d0
Compare
S3 gives a bucket three versioning states, while Ozone has a single isVersionEnabled boolean. This adds the three-state status alongside the flag rather than in place of it, so existing buckets and older clients go on working unchanged. BucketVersioningStatus holds the three states and the state machine that governs them: UNVERSIONED may move anywhere, but once versioning has been enabled or suspended a bucket can never return to UNVERSIONED. The proto gains a matching enum and an optional versioningStatus on both BucketInfo and BucketArgs. OmBucketInfo keeps the two representations in sync in both directions: a status derives the flag (ENABLED -> true), and a record carrying only the flag derives a status, so a bucket written before this change still answers getVersioningStatus(). Disabling the flag is the asymmetric case - it leaves an explicitly SUSPENDED status alone, since the state machine has no way back to UNVERSIONED. Nothing enforces the state machine yet; this commit only defines it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A bucket's versioning status may only change along the state machine the previous commit defined. OMBucketSetPropertyRequest now checks the requested status against the one the bucket already holds and rejects with INVALID_REQUEST what the state machine forbids - the return to UNVERSIONED once versioning has been enabled or suspended. A request carrying only the legacy flag is mapped onto the same machine before that check: enabling always means ENABLED, while disabling means SUSPENDED, except on a bucket that is still UNVERSIONED, where it stays UNVERSIONED. The status is refused outright at bucket creation. S3 has no way to create a bucket already in a versioning state: CreateBucket carries no such parameter, and the state is set afterwards through PutBucketVersioning. versioningStatus sits on BucketInfo because that message is the bucket's on-disk record and the shape InfoBucket and ListBuckets return, not because CreateBucket needs it; honouring it there would let a caller land on any status in one step, with none of the above applied. Nothing populates the field on a create today, so the request is rejected rather than quietly ignored - ignoring it would leave a future caller believing it had created a versioned bucket. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Each section below was a separate commit; they are folded here so the task's review fixes land as one change. * Address comments * Use a 0x00 separator in versionedKeyTable dbKeys Key names in OBJECT_STORE buckets contain '/' verbatim, so a '/' separator interleaves a key's versions with those of keys nested under it, breaking the single-seek promotion and the merged ListObjectVersions order. * Do not derive a versioning status from the legacy flag Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds VersionIdGenerator, the pluggable source of the id an object version is numbered with, and UniqueIdVersionIdGenerator as the cluster default. The id is proposed on the OM that received the request, in preExecute, so it travels in the replicated request and every OM applies a version that is already numbered. Nothing about it depends on the transaction carrying the write, or on OM being replicated by Ratis. The default numbers a version with the time it was written, through the scheme Ozone already uses for block local IDs: currentTimeMillis << 16 with a 16-bit counter separating ids proposed inside one millisecond. It needs no allocator state and no coordination, which is what makes it safe to read on any OM. The interface has one abstract method, generateVersionId(), plus a default versionIdFor(proposed, hasCurrentVersion) that lets a generator number some versions specially at apply time. The implementation is selected cluster-wide by ozone.om.versioning.version-id-generator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds VersionIdAllocator, which turns the id proposed for a version into the id it is applied with. versionedKeyTable orders a key's versions by Long.MAX_VALUE - versionId, so the ids of one key have to increase in the order the versions were written. A proposal is a clock reading and cannot promise that: ids proposed inside one millisecond can exhaust the counter separating them, and a leader change onto a lagging clock proposes a lower value. So a proposal is a floor. The applied id is the later of it and the id after the key's current version, which the write path already holds - no read of its own, no global state, and identical on every OM. Under a clock regression an affected key's ids climb by one until proposals overtake them again: the versions stay ordered and only the id's reading as a time degrades. propose() runs in preExecute on the OM that received the request; allocate() runs under the write's lock on every OM. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds PinnedFirstVersionIdGenerator, which numbers versions like the default except that a key's first version takes FIRST_VERSION_ID, so it can be referenced without listing the key's versions first. Whether the key already has a version is not known when the id is proposed, so the generator decides it in versionIdFor, under the write's lock, from the current version the allocator was handed. The sentinel is 1: below every proposed id, so a pinned version sorts at the old end of the key in versionedKeyTable, and above the unset value a pre-versioning record carries. It says nothing about the null version, which carries a proposed id like any other and is marked by isNullVersion. Known trade-off: once every version of a key has been permanently deleted, a recreated key takes the sentinel again, so an external reference to the first version resolves to the new content. The generator is off by default and selected by ozone.om.versioning.version-id-generator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
On a versioning-enabled bucket a commit no longer reclaims the version it overwrites: the previous current version moves to the versionedKeyTable and the new record becomes current, in one WriteBatch. A record written before versioning was enabled carries no versionId and becomes the key's null version. The new version is numbered from KeyArgs.proposedVersionId, which preExecute stamps into the request on the OM that received it, so every OM applies the same id and nothing is generated during apply. The allocator raises the proposal to come after the key's current version when it does not already. An hsync re-commit keeps updating the version it opened, so it keeps its versionId and moves nothing. Both reclaim branches now depend on the versioning status rather than on the legacy isVersionEnabled flag being kept in sync with it, so dropping that sync cannot strand a version record by reclaiming the blocks it still refers to. S3MultipartUploadCompleteRequest is guarded the same way; it does not yet record a version for the key it supersedes, so an MPU overwrite on a versioned bucket leaks those blocks until T5 lands. OMKeyCommitRequestWithFSO is left alone: isS3VersioningEnabled() requires the OBJECT_STORE layout, so the check is structurally false there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A delete without a versionId on a versioning-enabled bucket removes no data: a delete marker becomes the key's current version and the version it supersedes moves to the versionedKeyTable. As in S3, the marker is inserted even when the key does not exist. The marker is a version, so its id comes from KeyArgs.proposedVersionId the same way a commit's does, proposed in preExecute and raised to come after the key's current version at apply time. The failure response declares the same tables as the successful one: the double buffer cleans the table cache from the response's CleanupTableInfo, so a marker request that failed after touching the versionedKeyTable cache would otherwise leave an entry behind that is in no DB and never cleaned. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A batch delete is the same operation over more keys, but it goes through a request of its own and that one still hard-deleted on a versioned bucket: the current version's blocks were reclaimed and the versions underneath were left in the versionedKeyTable with nothing in the keyTable above them. Those versions then read as absent, hold quota, and cannot be promoted, because promotion only runs when a version is deleted by id. The gap opens with T3.1, where the two tables first diverge, so it is closed here rather than later. It is reachable today: the S3 gateway wires DeleteObjects straight to this request, and so does the client's deleteKeys API. Rather than write a second marker implementation, the one T3.2 added moves to OMKeyRequest and returns what it changed, so the single-key and batch requests build their own responses from the same insertion. One proposed versionId covers a batch: an id only has to increase within one key, and the applying OM raises it per key against that key's own current version. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
This PR only includes the commits with prefix of [T3.x], please ignore the commits prefixed with "T1" and "T2".
This ticket is related to the write part of versioning feature.
What is the link to the Apache JIRA
https://issues.apache.org/jira/browse/HDDS-16014
How was this patch tested?
unit test.