ext4: fix dirtyclusters double decrement on fs shutdown

JIRA: https://redhat.atlassian.net/browse/RHEL-145583
Conflicts: Drop hunk from non-existent test code.

commit 94a8cea54cd935c54fa2fba70354757c0fc245e3
Author: Brian Foster <bfoster@redhat.com>
Date:   Tue Jan 13 12:19:05 2026 -0500

    ext4: fix dirtyclusters double decrement on fs shutdown

    fstests test generic/388 occasionally reproduces a warning in
    ext4_put_super() associated with the dirty clusters count:

      WARNING: CPU: 7 PID: 76064 at fs/ext4/super.c:1324 ext4_put_super+0x48c/0x590 [ext4]

    Tracing the failure shows that the warning fires due to an
    s_dirtyclusters_counter value of -1. IOW, this appears to be a
    spurious decrement as opposed to some sort of leak. Further tracing
    of the dirty cluster count deltas and an LLM scan of the resulting
    output identified the cause as a double decrement in the error path
    between ext4_mb_mark_diskspace_used() and the caller
    ext4_mb_new_blocks().

    First, note that generic/388 is a shutdown vs. fsstress test and so
    produces a random set of operations and shutdown injections. In the
    problematic case, the shutdown triggers an error return from the
    ext4_handle_dirty_metadata() call(s) made from
    ext4_mb_mark_context(). The changed value is non-zero at this point,
    so ext4_mb_mark_diskspace_used() does not exit after the error
    bubbles up from ext4_mb_mark_context(). Instead, the former
    decrements both cluster counters and returns the error up to
    ext4_mb_new_blocks(). The latter falls into the !ar->len out path
    which decrements the dirty clusters counter a second time, creating
    the inconsistency.

    To avoid this problem and simplify ownership of the cluster
    reservation in this codepath, lift the counter reduction to a single
    place in the caller. This makes it more clear that
    ext4_mb_new_blocks() is responsible for acquiring cluster
    reservation (via ext4_claim_free_clusters()) in the !delalloc case
    as well as releasing it, regardless of whether it ends up consumed
    or returned due to failure.

    Fixes: 0087d9fb3f ("ext4: Fix s_dirty_blocks_counter if block allocation failed with nodelalloc")
    Signed-off-by: Brian Foster <bfoster@redhat.com>
    Reviewed-by: Baokun Li <libaokun1@huawei.com>
    Link: https://patch.msgid.link/20260113171905.118284-1-bfoster@redhat.com
    Signed-off-by: Theodore Ts'o <tytso@mit.edu>
    Cc: stable@kernel.org

Signed-off-by: Brian Foster <bfoster@redhat.com>
This commit is contained in:
Brian Foster
2026-03-24 08:54:40 -04:00
parent 382af06e11
commit 7898d6f279
+5 -16
View File
@@ -3798,8 +3798,7 @@ void ext4_exit_mballoc(void)
* Returns 0 if success or error code * Returns 0 if success or error code
*/ */
static noinline_for_stack int static noinline_for_stack int
ext4_mb_mark_diskspace_used(struct ext4_allocation_context *ac, ext4_mb_mark_diskspace_used(struct ext4_allocation_context *ac, handle_t *handle)
handle_t *handle, unsigned int reserv_clstrs)
{ {
struct buffer_head *bitmap_bh = NULL; struct buffer_head *bitmap_bh = NULL;
struct ext4_group_desc *gdp; struct ext4_group_desc *gdp;
@@ -3885,13 +3884,6 @@ ext4_mb_mark_diskspace_used(struct ext4_allocation_context *ac,
ext4_unlock_group(sb, ac->ac_b_ex.fe_group); ext4_unlock_group(sb, ac->ac_b_ex.fe_group);
percpu_counter_sub(&sbi->s_freeclusters_counter, ac->ac_b_ex.fe_len); percpu_counter_sub(&sbi->s_freeclusters_counter, ac->ac_b_ex.fe_len);
/*
* Now reduce the dirty block count also. Should not go negative
*/
if (!(ac->ac_flags & EXT4_MB_DELALLOC_RESERVED))
/* release all the reserved blocks if non delalloc */
percpu_counter_sub(&sbi->s_dirtyclusters_counter,
reserv_clstrs);
if (sbi->s_log_groups_per_flex) { if (sbi->s_log_groups_per_flex) {
ext4_group_t flex_group = ext4_flex_group(sbi, ext4_group_t flex_group = ext4_flex_group(sbi,
@@ -6085,7 +6077,7 @@ repeat:
ext4_mb_pa_put_free(ac); ext4_mb_pa_put_free(ac);
} }
if (likely(ac->ac_status == AC_STATUS_FOUND)) { if (likely(ac->ac_status == AC_STATUS_FOUND)) {
*errp = ext4_mb_mark_diskspace_used(ac, handle, reserv_clstrs); *errp = ext4_mb_mark_diskspace_used(ac, handle);
if (*errp) { if (*errp) {
ext4_discard_allocated_blocks(ac); ext4_discard_allocated_blocks(ac);
goto errout; goto errout;
@@ -6116,12 +6108,9 @@ errout:
out: out:
if (inquota && ar->len < inquota) if (inquota && ar->len < inquota)
dquot_free_block(ar->inode, EXT4_C2B(sbi, inquota - ar->len)); dquot_free_block(ar->inode, EXT4_C2B(sbi, inquota - ar->len));
if (!ar->len) { /* release any reserved blocks */
if ((ar->flags & EXT4_MB_DELALLOC_RESERVED) == 0) if (reserv_clstrs)
/* release all the reserved blocks if non delalloc */ percpu_counter_sub(&sbi->s_dirtyclusters_counter, reserv_clstrs);
percpu_counter_sub(&sbi->s_dirtyclusters_counter,
reserv_clstrs);
}
trace_ext4_allocate_blocks(ar, (unsigned long long)block); trace_ext4_allocate_blocks(ar, (unsigned long long)block);