Please confirm
Distribution
Ubuntu
Distribution version
24.04
Output of snap list --all lxd core20 core22 core24 core26 snapd
Name Version Rev Tracking Publisher Notes
core26 20260629 462 latest/stable canonical✓ base
lxd git-9ec42b5 40719 6/edge canonical✓ -
snapd 2.76.3 27738 latest/stable canonical✓ snapd
System info
driver_version: 6.0.6 | 10.2.1
kernel_version: 6.8.0-138-generic
server_version: "6.9"
storage: zfs
Instance log
n/a
Expected behavior
When deleting an instance, LXD should either return an error without deleting anything, or clean up everything.
I realise this is probably quite difficult to implement. The bare minimum, I think, is to retry removing ZFS volumes if the initial remove fails with dataset is busy. This error is expected, but transient, when lxc delete is run concurrently with virtiofsd, due to the issue discussed here: #18918 (comment).
Actual behavior
For a VM with a ZFS root disk, whose metadata volume is busy, the initial delete removes the ZFS block device but not the metadata disk or the database entry. On retry, LXD deletes the database entry but still doesn't delete the metadata volume, even if it's not busy anymore.
Steps to reproduce
$ lxc launch --vm ubuntu:24.04 v1
Launching v1
$ until lxc exec v1 -- systemctl is-system-running &>/dev/null; do sleep 1; done
$ sudo mount -t zfs default/virtual-machines/v1 /mnt
$ lxc delete -f v1
Error: Failed deleting instance "v1" in project "default": Error deleting storage volume: Failed running: zfs destroy -r default/virtual-machines/v1: exit status 1 (cannot destroy 'default/virtual-machines/v1': dataset is busy)
$ sudo umount /mnt
$ lxc delete v1
$ zfs list -r default/virtual-machines
NAME USED AVAIL REFER MOUNTPOINT
default/virtual-machines 99.5K 3.54G 24K legacy
default/virtual-machines/v1 75.5K 99.9M 78K legacy
Information to attach
Please confirm
Distribution
Ubuntu
Distribution version
24.04
Output of
snap list --all lxd core20 core22 core24 core26 snapdSystem info
Instance log
n/a
Expected behavior
When deleting an instance, LXD should either return an error without deleting anything, or clean up everything.
I realise this is probably quite difficult to implement. The bare minimum, I think, is to retry removing ZFS volumes if the initial remove fails with
dataset is busy. This error is expected, but transient, whenlxc deleteis run concurrently withvirtiofsd, due to the issue discussed here: #18918 (comment).Actual behavior
For a VM with a ZFS root disk, whose metadata volume is busy, the initial delete removes the ZFS block device but not the metadata disk or the database entry. On retry, LXD deletes the database entry but still doesn't delete the metadata volume, even if it's not busy anymore.
Steps to reproduce
Information to attach
dmesg)lxc config show <instance> --expanded)/var/log/lxd/lxd.logor/var/snap/lxd/common/lxd/logs/lxd.log)--debug--debug(or uselxc monitorwhile reproducing the issue)