Repository navigation
fix(kafka): preserve stable log.dirs ordering across storageConfigs changes - #350
Open
dobrerazvan wants to merge 1 commit into
Open
dobrerazvan wants to merge 1 commit into
dobrerazvan wants to merge 1 commit into
Conversation
…hanges Previously, getEffectiveLogDirsMountPaths() seeded the effective list directly from mountPathsNew (the freshly generated order from the current storageConfigs spec), so simply reordering entries in storageConfigs reshuffled log.dirs to match. In KRaft mode an unset metadata.log.dir defaults to log.dirs[0], so this could silently move the metadata log to a different disk whenever storageConfigs was reordered (e.g. while removing/re-adding a disk), with no change intended to the actual broker storage layout. Now the effective list is built by walking mountPathsOld first, keeping each existing path in its original position (including paths kept temporarily during a pending disk removal/rebalance), and only appending genuinely new paths from mountPathsNew at the end, in their declared relative order. Example: old log.dirs = [a, b, c]; storageConfigs reordered to declare [c, a, b] -> before: log.dirs became [c, a, b]; after: log.dirs stays [a, b, c]. Example: disk 'b' removal confirmed, then re-added before 'a' in storageConfigs (old=[b], new=[a, b]) -> before: log.dirs became [a, b]; after: log.dirs becomes [b, a] (b keeps its slot, a is appended as new). Adds 5 table-driven test cases to TestGetEffectiveLogDirsMountPaths covering reordering, new-path append order, pending-removal position retention under reordering, confirmed-removal-then-readd append, and a mixed scenario.
Author
|
@azun @eduardagarici Please review. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes
getEffectiveLogDirsMountPathsso that reordering entries inspec.brokers[].brokerConfig.storageConfigs(or the sharedbrokerConfigGroupsequivalent) no longer reshuffles the broker's effectivelog.dirs.Why this matters
In KRaft mode, when
metadata.log.diris not explicitly set, Kafka defaults it tolog.dirs[0]. Koperator never setsmetadata.log.direxplicitly, so whichever path ends up first inlog.dirssilently becomes the metadata log directory.Previously, the merge function started from
mountPathsNew(the list freshly generated from the currentstorageConfigs), so any reordering ofstorageConfigs— even with no actual disks added or removed — propagated straight intolog.dirs. This could silently move the KRaft metadata log to a different disk as a side effect of an unrelated spec edit (e.g., someone removing and then re-adding a disk entry, or just tidying up the YAML).Behavior change
storageConfigsdeclaration orderlog.dirsstorageConfigsorderlog.dirs)storageConfigsstorageConfigsrelative orderExample 1 — pure reordering, no disk changes
log.dirs:[/a/kafka, /b/kafka, /c/kafka]storageConfigsreordered to declare:[/c, /a, /b]Before:
log.dirsbecomes[/c/kafka, /a/kafka, /b/kafka]→/cis nowlog.dirs[0], silently becoming the KRaft metadata dir.After:
log.dirsstays[/a/kafka, /b/kafka, /c/kafka]→/aremainslog.dirs[0].Example 2 — disk removal confirmed, then re-added before an existing disk
log.dirs:[/b/kafka](disk/awas already fully removed)storageConfigsnow declares:[/a, /b](re-adding/aahead of/b)Before:
log.dirsbecomes[/a/kafka, /b/kafka]→/abecomeslog.dirs[0].After:
log.dirsbecomes[/b/kafka, /a/kafka]→/bkeeps its existing slot,/ais appended as the newly (re-)added disk.Example 3 — mixed: pending removal + confirmed-removal re-add
log.dirs:[/b/kafka, /c/kafka], with/cmid graceful-disk-removal (GracefulDiskRemovalRequired)storageConfigsnow declares:[/a, /b](adds/a, drops/cfrom the spec,/c's removal is still pending)Before:
log.dirsbecomes[/a/kafka, /b/kafka](and/chandling depended on where it ended up relative to the new order).After:
log.dirsbecomes[/b/kafka, /c/kafka, /a/kafka]→/band/ckeep their original relative order (with/cretained because its removal hasn't completed yet), and/ais appended at the end as the new disk.Changes
pkg/resources/kafka/configmap.go:getEffectiveLogDirsMountPathsnow builds the effective list by walkingmountPathsOldfirst (preserving order, including paths kept during a pending removal/rebalance), then appends only genuinely new paths frommountPathsNewat the end. Added a doc comment explaining the KRaftmetadata.log.dirrationale.pkg/resources/kafka/configmap_test.go: added 5 new table-driven cases toTestGetEffectiveLogDirsMountPaths:storageConfigsdoes not reorderlog.dirsAll 14 sub-tests in
TestGetEffectiveLogDirsMountPaths(9 existing + 5 new) pass.This PR alone does not guarantee
__cluster_metadata-0stays onlog.dirs[0]This fix only stabilizes
log.dirsordering when a previous broker ConfigMap exists to merge against (mountPathsOldnon-empty). It does not verify or correct where the physical__cluster_metadata-0partition actually lives on disk, and there are still code paths wherelog.dirs[0]can end up pointing at a mount path that was never the metadata directory:getEffectiveLogDirsMountPathsshort-circuits tomountPathsNewas-is:NotFound/API error while fetching it — regardless of whether the attached PVCs already contain real__cluster_metadata-0data from a prior broker life (e.g. disaster recovery / PV restore).log.dirs[0]then becomes whatever is first instorageConfigs, with no check against existing on-disk metadata state.log.dirs[0]is removed andshouldKeepRemovedLogDirInConfigreports the removal as no longer pending (noGracefulActionStateentry, or no broker status at all), that path is dropped and whatever is next becomeslog.dirs[0], with no verification that__cluster_metadata-0was actually moved off it beforehand.Recommendation: keep (or add) an init script/container that scans
log.dirsfor the directory actually holding__cluster_metadata-0andmvs it underlog.dirs[0]before Kafka starts, as a defense-in-depth safety net. This PR reduces how often that script needs to do anything — the common "user reorderedstorageConfigs" case becomes a no-op for it — but it does not replace the need for it in the scenarios above.Testing