Files
Centos-kernel-stream-10/include/uapi/linux/raid
Nigel Croxon 15e8399d92 md: allow configuring logical block size
JIRA: https://issues.redhat.com/browse/RHEL-129337

commit 62ed1b58224636185fa689db81224b8c8af46473
Author: Li Nan <linan122@huawei.com>
Date:   Mon Nov 3 20:57:57 2025 +0800

    md: allow configuring logical block size

    Previously, raid array used the maximum logical block size (LBS)
    of all member disks. Adding a larger LBS disk at runtime could
    unexpectedly increase RAID's LBS, risking corruption of existing
    partitions. This can be reproduced by:

    ```
      # LBS of sd[de] is 512 bytes, sdf is 4096 bytes.
      mdadm -CRq /dev/md0 -l1 -n3 /dev/sd[de] missing --assume-clean

      # LBS is 512
      cat /sys/block/md0/queue/logical_block_size

      # create partition md0p1
      parted -s /dev/md0 mklabel gpt mkpart primary 1MiB 100%
      lsblk | grep md0p1

      # LBS becomes 4096 after adding sdf
      mdadm --add -q /dev/md0 /dev/sdf
      cat /sys/block/md0/queue/logical_block_size

      # partition lost
      partprobe /dev/md0
      lsblk | grep md0p1
    ```

    Simply restricting larger-LBS disks is inflexible. In some scenarios,
    only disks with 512 bytes LBS are available currently, but later, disks
    with 4KB LBS may be added to the array.

    Making LBS configurable is the best way to solve this scenario.
    After this patch, the raid will:
      - store LBS in disk metadata
      - add a read-write sysfs 'mdX/logical_block_size'

    Future mdadm should support setting LBS via metadata field during RAID
    creation and the new sysfs. Though the kernel allows runtime LBS changes,
    users should avoid modifying it after creating partitions or filesystems
    to prevent compatibility issues.

    Only 1.x metadata supports configurable LBS. 0.90 metadata inits all
    fields to default values at auto-detect. Supporting 0.90 would require
    more extensive changes and no such use case has been observed.

    Note that many RAID paths rely on PAGE_SIZE alignment, including for
    metadata I/O. A larger LBS than PAGE_SIZE will result in metadata
    read/write failures. So this config should be prevented.

    Link: https://lore.kernel.org/linux-raid/20251103125757.1405796-6-linan666@huaweicloud.com
    Signed-off-by: Li Nan <linan122@huawei.com>
    Reviewed-by: Xiao Ni <xni@redhat.com>
    Signed-off-by: Yu Kuai <yukuai@fnnas.com>

Signed-off-by: Nigel Croxon <ncroxon@redhat.com>
2026-05-15 19:35:26 -04:00
..
2025-01-15 15:59:53 -05:00