Skip to content

kvm: fill the template's holes before taking its LUKS clone snapshot - #14

Open
calvix wants to merge 1 commit into
integration/all-fixes-4.23.0.0from
fix/rbd-encryption-template-holes-min
Open

calvix wants to merge 1 commit into
integration/all-fixes-4.23.0.0from
fix/rbd-encryption-template-holes-min

Conversation

@calvix

@calvix calvix commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Description

An encrypted root disk on RBD is a librbd CoW clone of the (plaintext) template with a LUKS2 header applied to the clone. librbd serves the ranges the clone's own objects don't hold from that parent, as plaintext, which is what makes the inherited filesystem readable at all.

That stops as soon as an object exists in the clone. One 4 KiB guest write copies up the whole 4 MiB object; librbd re-encrypts the parent data it copies, but the ranges the parent never materialised (a sparse template) aren't written at all. The object now exists, so reads of those ranges no longer go to the parent - they go through the crypto layer, and zeros decrypted with AES-XTS are not zeros. The guest gets garbage where it wrote nothing and the template held nothing.

So the corruption is latent and spreads with ordinary writes: the VM boots and runs, and breaks later in whatever file happens to sit in a copied-up object. A mostly empty separate /boot shows it first, and ext4lazyinit walks the whole disk.

The fix

The template is already prepared once for encrypted clones: it gets grown by the LUKS2 header reserve and a protected snapshot is taken to clone from. The only thing wrong is that this snapshot has the template's holes. So right before taking it, the template image is rewritten from its own cloudstack-base-snap with qemu-img convert -n -S 0, which writes out every zero range. The snapshot taken after that has nothing left to materialise, and every range of an encrypted clone stays readable. Clones stay thin.

  • The snapshot is now called cloudstack-base-snap-luks2 instead of -luks, so templates that were prepared before (and carry a sparse -luks snapshot) get a dense one too. Roots already cloned from the old snapshot keep it.
  • Plaintext clones aren't affected: they're cloned from cloudstack-base-snap, which is only read.
  • If another host prepares the same template at the same moment, the snapshot it created is used instead of failing the deploy.
  • QemuImg gets setWriteZeroRanges() for the -S 0, since qemu-img convert skips the source's zero ranges by default.

The cost is that a template used for encrypted roots becomes fully allocated in the pool, once per template per pool.

Types of changes

  • Breaking change (fix or feature that would cause existing functionality to change)
  • New feature (non-breaking change which adds functionality)
  • Bug fix (non-breaking change which fixes an issue)
  • Enhancement (improves an existing feature and functionality)
  • Cleanup (Code refactoring and cleanup, that may add test cases)
  • Build/CI
  • Test (unit or integration test code)

Feature/Enhancement Scale or Bug Severity

Bug Severity

  • BLOCKER
  • Critical
  • Major
  • Minor
  • Trivial

How Has This Been Tested?

3-node KVM cluster (4.23.0.0), RBD primary storage.

Unit tests: RbdEncryptionTest (10) and LibvirtStorageAdaptorTest (23) pass. RbdEncryptionTest checks that the copy asks for -S 0, reads from the template's snapshot and writes into the template image itself, unencrypted. QemuImgTest covers what -S 0 does to a local file, but only runs where qemu-img and libvirt are available.

On the cluster, with an Ubuntu 24.04 template that already had the old sparse -luks snapshot (2.2 of 4 GiB used):

  • Deploying an encrypted root created cloudstack-base-snap-luks2 (~31 s for the 4 GiB template), fully allocated (4.0 of 4.0 GiB). The old -luks snapshot was left alone.
  • The root is a clone of @cloudstack-base-snap-luks2 and the VM booted from it.
  • A plaintext VM from the same template is still a clone of @cloudstack-base-snap and runs.

To check the actual bug I made two clones of that template, one from the old -luks snapshot and one from the new -luks2, both with a LUKS2 header. In 10 objects that hold data but also have holes, I read a hole sector, wrote 4 KiB of zeros elsewhere in the same object (forcing the copy-up), and read the sector again:

                    hole reads zeros   hole reads zeros
                    before the write   after the copy-up
                    ----------------   -----------------
old -luks (sparse)        10/10              0/10
new -luks2 (dense)        10/10             10/10

Not tested: two hosts preparing the same template at the same time, and templates bigger than 4 GiB.

An encrypted root is a librbd CoW clone of a plaintext template carrying a
LUKS2 header. librbd serves the ranges the clone's own objects do not hold
from that parent, as plaintext, which is what makes the inherited filesystem
readable. That stops the moment an object exists in the clone: one 4 KiB guest
write copies up the whole 4 MiB object, and the ranges inside it that the
parent never materialised are from then on read through the crypto layer -
where zeros decrypted with AES-XTS are not zeros.

Before taking the snapshot encrypted roots are cloned from, rewrite the
template image from its own cloudstack-base-snap with qemu-img convert -S 0,
so the snapshot has no holes left. The snapshot gets a new name, -luks2, so
templates prepared before get a dense one too; roots already cloned from the
old -luks snapshot keep it. If another host prepares the same template at the
same time, its protected snapshot is used instead of failing the deploy.

QemuImg grows setWriteZeroRanges() for the -S 0 this needs, since qemu-img
skips the source's zero ranges by default.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant