100 Commits
Author SHA1 Message Date
Rafael Aquini 30eb170c14 mm/list_lru: drain before clearing xarray entry on reparent
JIRA: https://redhat.atlassian.net/browse/RHEL-227151
CVE: CVE-2026-53153

commit 98733f3f0becb1ae0701d021c1748e974e5fa55c
Author: Shakeel Butt <shakeel.butt@linux.dev>
Date:   Mon Jun 1 09:15:01 2026 -0700

    mm/list_lru: drain before clearing xarray entry on reparent

    memcg_reparent_list_lrus() clears the dying memcg's xarray entry with
    xas_store(&xas, NULL) before reparenting its per-node lists into the
    parent.  This opens a window where a concurrent list_lru_del() arriving
    for the dying memcg sees xa_load() == NULL, walks to the parent in
    lock_list_lru_of_memcg(), takes the parent's per-node lock, and calls
    list_del_init() on an item still physically linked on the dying memcg's
    list.

    If another in-flight thread holds the dying memcg's per-node lock at the
    same moment (another list_lru_del, or a list_lru_walk_one running an
    isolate callback), both threads modify ->next/->prev pointers on the same
    physical list under different locks.  Adjacent items can corrupt each
    other's links.

    Fix it by reversing the order: reparent each per-node list and mark the
    child's list lru dead and then clear the xarray entry.  Any concurrent
    list_lru op that finds the still-set xarray entry either takes the dying
    memcg's per-node lock (synchronizing with the drain) or sees LONG_MIN and
    walks to the parent, where the items now live.

    Link: https://lore.kernel.org/20260601161501.1444829-1-shakeel.butt@linux.dev
    Fixes: fb56fdf8b9a2 ("mm/list_lru: split the lock to per-cgroup scope")
    Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
    Reported-by: Chris Mason <clm@fb.com>
    Reviewed-by: Kairui Song <kasong@tencent.com>
    Acked-by: Muchun Song <muchun.song@linux.dev>
    Cc: Dave Chinner <david@fromorbit.com>
    Cc: Johannes Weiner <hannes@cmpxchg.org>
    Cc: Roman Gushchin <roman.gushchin@linux.dev>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-08-26 08:50:59 -04:00
Rafael Aquini a35e50db9a mm/gup: fix GUP-fast fallback for NULL-mapping order-0 folios
JIRA: https://redhat.atlassian.net/browse/RHEL-231964
Upstream status: git://git.kernel.org/pub/scm/linux/kernel/git/next/linux-next.git

commit c494788faffe67216c56623d240541fde50139c3
Author: John Hubbard <jhubbard@nvidia.com>
Date:   Tue Jul 7 17:57:45 2026 -0700

    mm/gup: fix GUP-fast fallback for NULL-mapping order-0 folios

    Since commit f002882ca3 ("mm: merge folio_is_secretmem() and
    folio_fast_pin_allowed() into gup_fast_folio_allowed()"),
    gup_fast_folio_allowed() falls back to the slow path for any order-0 folio
    with a NULL mapping when CONFIG_SECRETMEM=y.  This causes a performance
    regression for drivers that allocate pages with alloc_page() and insert
    them into VMAs via vm_insert_page().  These pages legitimately have a NULL
    folio->mapping, but they cannot be secretmem pages.

    Secretmem pages are always added to the secretmem inode's page cache via
    filemap_add_folio(), which sets folio->mapping to the inode's i_mapping.
    A folio with a NULL mapping can never be a secretmem folio.  The
    NULL-mapping check was intended to handle truncated file-backed pages (a
    reject_file_backed concern), not secretmem detection.

    When only check_secretmem is true (and reject_file_backed is false), a
    NULL mapping is sufficient to prove the folio is not secretmem, so the
    fast path can proceed.

    Link: https://lore.kernel.org/20260708005745.164928-1-jhubbard@nvidia.com
    Fixes: f002882ca3 ("mm: merge folio_is_secretmem() and folio_fast_pin_allowed() into gup_fast_folio_allowed()")
    Signed-off-by: John Hubbard <jhubbard@nvidia.com>
    Tested-by: Sourab Gupta <sougupta@nvidia.com>
    Acked-by: David Hildenbrand (Arm) <david@kernel.org>
    Cc: Alistair Popple <apopple@nvidia.com>
    Cc: Balbir Singh <balbirs@nvidia.com>
    Cc: Zi Yan <ziy@nvidia.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-08-11 11:07:57 -04:00
Rafael Aquini 9d6e7c89f6 mm/hugetlb: fix folio is still mapped when deleted
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-40006

commit 7b7387650dcf2881fd8bb55bcf3c8bd6c9542dd7
Author: Jinjiang Tu <tujinjiang@huawei.com>
Date:   Fri Sep 12 15:41:39 2025 +0800

    mm/hugetlb: fix folio is still mapped when deleted

    Migration may be raced with fallocating hole.  remove_inode_single_folio
    will unmap the folio if the folio is still mapped.  However, it's called
    without folio lock.  If the folio is migrated and the mapped pte has been
    converted to migration entry, folio_mapped() returns false, and won't
    unmap it.  Due to extra refcount held by remove_inode_single_folio,
    migration fails, restores migration entry to normal pte, and the folio is
    mapped again.  As a result, we triggered BUG in filemap_unaccount_folio.

    The log is as follows:
     BUG: Bad page cache in process hugetlb  pfn:156c00
     page: refcount:515 mapcount:0 mapping:0000000099fef6e1 index:0x0 pfn:0x156c00
     head: order:9 mapcount:1 entire_mapcount:1 nr_pages_mapped:0 pincount:0
     aops:hugetlbfs_aops ino:dcc dentry name(?):"my_hugepage_file"
     flags: 0x17ffffc00000c1(locked|waiters|head|node=0|zone=2|lastcpupid=0x1fffff)
     page_type: f4(hugetlb)
     page dumped because: still mapped when deleted
     CPU: 1 UID: 0 PID: 395 Comm: hugetlb Not tainted 6.17.0-rc5-00044-g7aac71907bde-dirty #484 NONE
     Hardware name: QEMU Ubuntu 24.04 PC (i440FX + PIIX, 1996), BIOS 0.0.0 02/06/2015
     Call Trace:
      <TASK>
      dump_stack_lvl+0x4f/0x70
      filemap_unaccount_folio+0xc4/0x1c0
      __filemap_remove_folio+0x38/0x1c0
      filemap_remove_folio+0x41/0xd0
      remove_inode_hugepages+0x142/0x250
      hugetlbfs_fallocate+0x471/0x5a0
      vfs_fallocate+0x149/0x380

    Hold folio lock before checking if the folio is mapped to avold race with
    migration.

    Link: https://lkml.kernel.org/r/20250912074139.3575005-1-tujinjiang@huawei.com
    Fixes: 4aae8d1c05 ("mm/hugetlbfs: unmap pages if page fault raced with hole punch")
    Signed-off-by: Jinjiang Tu <tujinjiang@huawei.com>
    Cc: David Hildenbrand <david@redhat.com>
    Cc: Kefeng Wang <wangkefeng.wang@huawei.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Muchun Song <muchun.song@linux.dev>
    Cc: Oscar Salvador <osalvador@suse.de>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:09 -04:00
Rafael Aquini d863ac6ebc mm/filemap: fix page_cache_prev_miss() when no hole is found
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 88e09fffeef5825931e6374b9e88d4b1a1d5f6f8
Author: Tal Zussman <tz2294@columbia.edu>
Date:   Tue May 12 16:45:59 2026 -0400

    mm/filemap: fix page_cache_prev_miss() when no hole is found

    page_cache_prev_miss() is documented to return a value outside the
    searched range when no gap is found.  However, the no-gap-found path
    returns xas.xa_index, which after a successful loop is the first index in
    the range.  As such, that index is misreported as a gap.

    The sole caller, page_cache_sync_ra(), uses the return value to estimate
    the cached run preceding a sequential read.  In some cases, the buggy
    return value can undercount the contiguous range by one, shrinking the
    readahead window or pushing borderline requests into the small-random-read
    branch.

    Fix this by returning the start of the range - 1 when no hole is found.
    Update page_cache_next_miss() for clarity as well.

    Both helpers were previously fixed together in commit 9425c591e0 ("page
    cache: fix page_cache_next/prev_miss off by one"), but the fix was
    reverted because it caused a hugetlb performance regression.  hugetlb no
    longer uses these functions and next_miss was subsequently refixed in
    commit 901a269ff3 ("filemap: fix page_cache_next_miss() when no hole
    found") and commit bbcaee20e03e ("readahead: fix return value of
    page_cache_next_miss() when no hole is found"), but prev_miss was not
    addressed.

    This was found by pointing Claude Opus 4.7 at mm/filemap.c.

    Link: https://lore.kernel.org/20260512-prev_miss_fix-v2-1-4af8e5c1ae62@columbia.edu
    Fixes: 0d3f929666 ("page cache: Convert hole search to XArray")
    Assisted-by: Claude:claude-opus-4-7
    Signed-off-by: Tal Zussman <tz2294@columbia.edu>
    Reviewed-by: Jan Kara <jack@suse.cz>
    Reviewed-by: Vishal Moola <vishal.moola@gmail.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:08 -04:00
Rafael Aquini e17229043a mm/kconfig: make BALLOON_COMPACTION depend on MIGRATION
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 7cf3318a25877c0908e450919f7e1517908e24f1
Author: David Hildenbrand (Red Hat) <david@kernel.org>
Date:   Tue Jan 20 00:01:29 2026 +0100

    mm/kconfig: make BALLOON_COMPACTION depend on MIGRATION

    Migration support for balloon memory depends on MIGRATION not COMPACTION.
    Compaction is simply another user of page migration.

    The last dependency on compaction.c was effectively removed with commit
    3d388584d599 ("mm: convert "movable" flag in page->mapping to a page
    flag").  Ever since, everything for handling movable_ops page migration
    resides in core migration code.

    So let's change the dependency and adjust the description + help text.

    We'll rename BALLOON_COMPACTION separately next.

    Link: https://lkml.kernel.org/r/20260119230133.3551867-22-david@kernel.org
    Signed-off-by: David Hildenbrand (Red Hat) <david@kernel.org>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Acked-by: Michael S. Tsirkin <mst@redhat.com>
    Cc: Arnd Bergmann <arnd@arndb.de>
    Cc: Christophe Leroy <christophe.leroy@csgroup.eu>
    Cc: Eugenio Pérez <eperezma@redhat.com>
    Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
    Cc: Jason Wang <jasowang@redhat.com>
    Cc: Jerrin Shaji George <jerrin.shaji-george@broadcom.com>
    Cc: Jonathan Corbet <corbet@lwn.net>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
    Cc: Michael Ellerman <mpe@ellerman.id.au>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Mike Rapoport <rppt@kernel.org>
    Cc: Nicholas Piggin <npiggin@gmail.com>
    Cc: Oscar Salvador <osalvador@suse.de>
    Cc: SeongJae Park <sj@kernel.org>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Xuan Zhuo <xuanzhuo@linux.alibaba.com>
    Cc: Zi Yan <ziy@nvidia.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:08 -04:00
Rafael Aquini 687d0c45e9 Docs/mm/damon/maintainer-profile: fix a typo on mm-untable link
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 7ad58e009dd159d7004592d826749003197f4083
Author: SeongJae Park <sj@kernel.org>
Date:   Wed Nov 12 07:41:11 2025 -0800

    Docs/mm/damon/maintainer-profile: fix a typo on mm-untable link

    Commit 0b473f9e6eac ("Docs/mm/damon/maintainer-profile: update for mm-new
    tree") mistakenly forgot putting a space between a link and the next word.
    Fix it.

    Link: https://lkml.kernel.org/r/20251112154114.66053-9-sj@kernel.org
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Cc: Bill Wendling <morbo@google.com>
    Cc: Brendan Higgins <brendan.higgins@linux.dev>
    Cc: David Gow <davidgow@google.com>
    Cc: David Hildenbrand <david@kernel.org>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Jonathan Corbet <corbet@lwn.net>
    Cc: Justin Stitt <justinstitt@google.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Miguel Ojeda <ojeda@kernel.org>
    Cc: Mike Rapoport <rppt@kernel.org>
    Cc: Nathan Chancellor <nathan@kernel.org>
    Cc: Shuah Khan <shuah@kernel.org>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:08 -04:00
Rafael Aquini 920fae0575 mm/damon/core: fix wrong comment of damon_call() return timing
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 50ca6423643cdb26871970b4f8870b4940b4c498
Author: SeongJae Park <sj@kernel.org>
Date:   Sun Oct 26 11:22:06 2025 -0700

    mm/damon/core: fix wrong comment of damon_call() return timing

    Patch series "mm/damon: misc documentation fixups".

    First three patches fix up issues in the documents, including wrong
    explanation of a behavior, wrong link, and a contextual typo.  Following
    five patches update documents for not yet documented features and
    behaviors.

    This patch (of 8):

    damon_call() works asynchronously and synchronously for repeat and
    non-repeat mode requests, respectively.  The comment about the behavior is
    wrong, though.  Fix it.

    The wrong comment was introduced together with the repeat mode, by commit
    43df7676e550 ("mm/damon/core: introduce repeat mode damon_call()").

    Link: https://lkml.kernel.org/r/20251026182216.118200-1-sj@kernel.org
    Link: https://lkml.kernel.org/r/20251026182216.118200-2-sj@kernel.org
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Cc: David Hildenbrand <david@redhat.com>
    Cc: Jonathan Corbet <corbet@lwn.net>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Mike Rapoport <rppt@kernel.org>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:08 -04:00
Rafael Aquini fe1cb410a1 mm/shmem: remove unused entry_order after large swapin rework
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 5919f1282141f29345432a4f1dadf34716f3dbec
Author: Jackie Liu <liuyun01@kylinos.cn>
Date:   Mon Sep 8 14:26:14 2025 +0800

    mm/shmem: remove unused entry_order after large swapin rework

    After commit 93c0476e7057 ("mm/shmem, swap: rework swap entry and index
    calculation for large swapin"), xas_get_order() will never return a
    non-zero value for `entry_order` in shmem_split_large_entry().  As a
    result, the local variable `entry_order` is effectively unused.

    Clean up the code by removing `entry_order` and directly using
    `cur_order`.  This change is purely a refactor and has no functional
    impact.

    No functional change intended.

    Link: https://lkml.kernel.org/r/20250908062614.89880-1-liu.yun@linux.dev
    Signed-off-by: Jackie Liu <liuyun01@kylinos.cn>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Kairui Song <kasong@tencent.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:08 -04:00
Rafael Aquini 36a9fb9af6 mm/damon/core: clear walk_control on inactive context in damos_walk()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2026-43388

commit d210fdcac9c0d1380eab448aebc93f602c1cd4e6
Author: Raul Pazemecxas De Andrade <raul_pazemecxas@hotmail.com>
Date:   Mon Feb 23 17:10:59 2026 -0800

    mm/damon/core: clear walk_control on inactive context in damos_walk()

    damos_walk() sets ctx->walk_control to the caller-provided control
    structure before checking whether the context is running.  If the context
    is inactive (damon_is_running() returns false), the function returns
    -EINVAL without clearing ctx->walk_control.  This leaves a dangling
    pointer to a stack-allocated structure that will be freed when the caller
    returns.

    This is structurally identical to the bug fixed in commit f9132fbc2e83
    ("mm/damon/core: remove call_control in inactive contexts") for
    damon_call(), which had the same pattern of linking a control object and
    returning an error without unlinking it.

    The dangling walk_control pointer can cause:
    1. Use-after-free if the context is later started and kdamond
       dereferences ctx->walk_control (e.g., in damos_walk_cancel()
       which writes to control->canceled and calls complete())
    2. Permanent -EBUSY from subsequent damos_walk() calls, since the
       stale pointer is non-NULL

    Nonetheless, the real user impact is quite restrictive.  The
    use-after-free is impossible because there is no damos_walk() callers who
    starts the context later.  The permanent -EBUSY can actually confuse
    users, as DAMON is not running.  But the symptom is kept only while the
    context is turned off.  Turning it on again will make DAMON internally
    uses a newly generated damon_ctx object that doesn't have the invalid
    damos_walk_control pointer, so everything will work fine again.

    Fix this by clearing ctx->walk_control under walk_control_lock before
    returning -EINVAL, mirroring the fix pattern from f9132fbc2e83.

    Link: https://lkml.kernel.org/r/20260224011102.56033-1-sj@kernel.org
    Fixes: bf0eaba0ff9c ("mm/damon/core: implement damos_walk()")
    Reported-by: Raul Pazemecxas De Andrade <raul_pazemecxas@hotmail.com>
    Closes: https://lore.kernel.org/CPUPR80MB8171025468965E583EF2490F956CA@CPUPR80MB8171.lamprd80.prod.outlook.com
    Signed-off-by: Raul Pazemecxas De Andrade <raul_pazemecxas@hotmail.com>
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Cc: <stable@vger.kernel.org>    [6.14+]
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:08 -04:00
Rafael Aquini 4afee2fc9e mm/damon/Kconfig: make DAMON_STAT_ENABLED_DEFAULT depend on DAMON_STAT
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 4e915656a38afe8aeebb283493f49c22d675a9fc
Author: Enze Li <lienze@kylinos.cn>
Date:   Fri Aug 15 17:21:10 2025 +0800

    mm/damon/Kconfig: make DAMON_STAT_ENABLED_DEFAULT depend on DAMON_STAT

    The DAMON_STAT_ENABLED_DEFAULT option is strongly tied to DAMON_STAT
    option -- enabling it alone is meaningless.  This patch makes
    DAMON_STAT_ENABLED_DEFAULT depend on DAMON_STAT, ensuring functional
    consistency.

    Link: https://lkml.kernel.org/r/20250815092110.811757-1-lienze@kylinos.cn
    Fixes: 369c415e6073 ("mm/damon: introduce DAMON_STAT module")
    Signed-off-by: Enze Li <lienze@kylinos.cn>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:08 -04:00
Rafael Aquini 2b3324bc12 mm/mremap: honour writable bit in mremap pte batching
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 04d1c9d60c6ec4c0003d433572eaa45f8b217788
Author: Dev Jain <dev.jain@arm.com>
Date:   Tue Oct 28 12:09:52 2025 +0530

    mm/mremap: honour writable bit in mremap pte batching

    Currently mremap folio pte batch ignores the writable bit during figuring
    out a set of similar ptes mapping the same folio.  Suppose that the first
    pte of the batch is writable while the others are not - set_ptes will end
    up setting the writable bit on the other ptes, which is a violation of
    mremap semantics.  Therefore, use FPB_RESPECT_WRITE to check the writable
    bit while determining the pte batch.

    Link: https://lkml.kernel.org/r/20251028063952.90313-1-dev.jain@arm.com
    Signed-off-by: Dev Jain <dev.jain@arm.com>
    Fixes: f822a9a81a31 ("mm: optimize mremap() by PTE batching")
    Reported-by: David Hildenbrand <david@redhat.com>
    Debugged-by: David Hildenbrand <david@redhat.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Acked-by: Pedro Falcato <pfalcato@suse.de>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Jann Horn <jannh@google.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: <stable@vger.kernel.org>    [6.17+]
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:07 -04:00
Rafael Aquini f70d7e4675 mm/hugetlb: fix incorrect error return from hugetlb_reserve_pages()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 9ee5d1766c8bfa4924bd47e31c4dd193493f5a45
Author: Shameer Kolothum <skolothumtho@nvidia.com>
Date:   Tue Nov 25 17:13:50 2025 +0000

    mm/hugetlb: fix incorrect error return from hugetlb_reserve_pages()

    The function hugetlb_reserve_pages() returns the number of pages added
    to the reservation map on success and a negative error code on failure
    (e.g. -EINVAL, -ENOMEM). However, in some error paths, it may return -1
    directly.

    For example, a failure at:

        if (hugetlb_acct_memory(h, gbl_reserve) < 0)
            goto out_put_pages;

    results in returning -1 (since add = -1), which may be misinterpreted
    in userspace as -EPERM.

    Fix this by explicitly capturing and propagating the return values from
    helper functions, and using -EINVAL for all other failure cases.

    Link: https://lkml.kernel.org/r/20251125171350.86441-1-skolothumtho@nvidia.com
    Fixes: 986f5f2b4be3 ("mm/hugetlb: make hugetlb_reserve_pages() return nr of entries updated")
    Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
    Reviewed-by: Joshua Hahn <joshua.hahnjy@gmail.com>
    Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
    Acked-by: Oscar Salvador <osalvador@suse.de>
    Cc: Matthew R. Ochs <mochs@nvidia.com>
    Cc: Muchun Song <muchun.song@linux.dev>
    Cc: Nicolin Chen <nicolinc@nvidia.com>
    Cc: Vivek Kasireddy <vivek.kasireddy@intel.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:07 -04:00
Rafael Aquini c260816d5c mm/memory_hotplug: maintain N_NORMAL_MEMORY during hotplug
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 2ecbe06abf9bfb2261cd6464a6bc3a3615625402
Author: Hao Li <hao.li@linux.dev>
Date:   Mon Mar 30 11:57:49 2026 +0800

    mm/memory_hotplug: maintain N_NORMAL_MEMORY during hotplug

    N_NORMAL_MEMORY is initialized from zone population at boot, but memory
    hotplug currently only updates N_MEMORY.  As a result, a node that gains
    normal memory via hotplug can remain invisible to users iterating over
    N_NORMAL_MEMORY, while a node that loses its last normal memory can stay
    incorrectly marked as such.

    The most visible effect is that
    /sys/devices/system/node/has_normal_memory does not report a node even
    after that node has gained normal memory via hotplug.

    Also, list_lru-based shrinkers can undercount objects on such a node
    and may skip reclaim on that node entirely, which can lead to a higher
    memory footprint than expected.

    Restore N_NORMAL_MEMORY maintenance directly in online_pages() and
    offline_pages().  Set the bit when a node that currently lacks normal
    memory onlines pages into a zone <= ZONE_NORMAL, and clear it when
    offlining removes the last present pages from zones <= ZONE_NORMAL.

    This restores the intended semantics without bringing back the old
    status_change_nid_normal notifier plumbing which was removed in
    8d2882a8edb8.

    Current users that benefit include list_lru, zswap, nfsd filecache,
    hugetlb_cgroup, and has_normal_memory sysfs reporting.

    Link: https://lkml.kernel.org/r/20260330035941.518186-1-hao.li@linux.dev
    Fixes: 8d2882a8edb8 ("mm,memory_hotplug: remove status_change_nid_normal and update documentation")
    Signed-off-by: Hao Li <hao.li@linux.dev>
    Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
    Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
    Reviewed-by: Joshua Hahn <joshua.hahnjy@gmail.com>
    Acked-by: David Hildenbrand (Arm) <david@kernel.org>
    Cc: Oscar Salvador <osalvador@suse.de>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:07 -04:00
Rafael Aquini 475bb8a20e mm: restore per-memcg proactive reclaim with !CONFIG_NUMA
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 16aca2c98a6fdf071e5a1a765a295995d7c7e346
Author: Yosry Ahmed <yosry.ahmed@linux.dev>
Date:   Fri Jan 16 20:52:47 2026 +0000

    mm: restore per-memcg proactive reclaim with !CONFIG_NUMA

    Commit 2b7226af730c ("mm/memcg: make memory.reclaim interface generic")
    moved proactive reclaim logic from memory.reclaim handler to a generic
    user_proactive_reclaim() helper to be used for per-node proactive reclaim.

    However, user_proactive_reclaim() was only defined under CONFIG_NUMA, with
    a stub always returning 0 otherwise.  This broke memory.reclaim on
    !CONFIG_NUMA configs, causing it to report success without actually
    attempting reclaim.

    Move the definition of user_proactive_reclaim() outside CONFIG_NUMA, and
    instead define a stub for __node_reclaim() in the !CONFIG_NUMA case.
    __node_reclaim() is only called from user_proactive_reclaim() when a write
    is made to sys/devices/system/node/nodeX/reclaim, which is only defined
    with CONFIG_NUMA.

    Link: https://lkml.kernel.org/r/20260116205247.928004-1-yosry.ahmed@linux.dev
    Fixes: 2b7226af730c ("mm/memcg: make memory.reclaim interface generic")
    Signed-off-by: Yosry Ahmed <yosry.ahmed@linux.dev>
    Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
    Acked-by: Michal Hocko <mhocko@suse.com>
    Cc: Axel Rasmussen <axelrasmussen@google.com>
    Cc: David Hildenbrand <david@kernel.org>
    Cc: Davidlohr Bueso <dave@stgolabs.net>
    Cc: Johannes Weiner <hannes@cmpxchg.org>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Mike Rapoport <rppt@kernel.org>
    Cc: Qi Zheng <zhengqi.arch@bytedance.com>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Wei Xu <weixugc@google.com>
    Cc: Yuanchu Xie <yuanchu@google.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:07 -04:00
Rafael Aquini 17e53b8c73 mm/damon/core: remove call_control in inactive contexts
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2026-23012

commit f9132fbc2e83baf2c45a77043672a63a675c9394
Author: SeongJae Park <sj@kernel.org>
Date:   Tue Dec 30 17:23:13 2025 -0800

    mm/damon/core: remove call_control in inactive contexts

    If damon_call() is executed against a DAMON context that is not running,
    the function returns error while keeping the damon_call_control object
    linked to the context's call_controls list.  Let's suppose the object is
    deallocated after the damon_call(), and yet another damon_call() is
    executed against the same context.  The function tries to add the new
    damon_call_control object to the call_controls list, which still has the
    pointer to the previous damon_call_control object, which is deallocated.
    As a result, use-after-free happens.

    This can actually be triggered using the DAMON sysfs interface.  It is not
    easily exploitable since it requires the sysfs write permission and making
    a definitely weird file writes, though.  Please refer to the report for
    more details about the issue reproduction steps.

    Fix the issue by making two changes.  Firstly, move the final
    kdamond_call() for cancelling all existing damon_call() requests from
    terminating DAMON context to be done before the ctx->kdamond reset.  This
    makes any code that sees NULL ctx->kdamond can safely assume the context
    may not access damon_call() requests anymore.  Secondly, let damon_call()
    to cleanup the damon_call_control objects that were added to the
    already-terminated DAMON context, before returning the error.

    Link: https://lkml.kernel.org/r/20251231012315.75835-1-sj@kernel.org
    Fixes: 004ded6bee11 ("mm/damon: accept parallel damon_call() requests")
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Reported-by: JaeJoon Jung <rgbi3307@gmail.com>
    Closes: https://lore.kernel.org/20251224094401.20384-1-rgbi3307@gmail.com
    Cc: <stable@vger.kernel.org> # 6.17.x
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:07 -04:00
Rafael Aquini 422756ac18 mm/damon/core: fix list_add_tail() call on damon_call()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit c3fa5b1bfd8380d935fa961f2ac166bdf000f418
Author: SeongJae Park <sj@kernel.org>
Date:   Tue Oct 14 13:59:36 2025 -0700

    mm/damon/core: fix list_add_tail() call on damon_call()

    Each damon_ctx maintains callback requests using a linked list
    (damon_ctx->call_controls).  When a new callback request is received via
    damon_call(), the new request should be added to the list.  However, the
    function is making a mistake at list_add_tail() invocation: putting the
    new item to add and the list head to add it before, in the opposite order.
    Because of the linked list manipulation implementation, the new request
    can still be reached from the context's list head.  But the list items
    that were added before the new request are dropped from the list.

    As a result, the callbacks are unexpectedly not invocated.  Worse yet, if
    the dropped callback requests were dynamically allocated, the memory is
    leaked.  Actually DAMON sysfs interface is using a dynamically allocated
    repeat-mode callback request for automatic essential stats update.  And
    because the online DAMON parameters commit is using a non-repeat-mode
    callback request, the issue can easily be reproduced, like below.

        # damo start --damos_action stat --refresh_stat 1s
        # damo tune --damos_action stat --refresh_stat 1s

    The first command dynamically allocates the repeat-mode callback request
    for automatic essential stat update.  Users can see the essential stats
    are automatically updated for every second, using the sysfs interface.

    The second command calls damon_commit() with a new callback request that
    was made for the commit.  As a result, the previously added repeat-mode
    callback request is dropped from the list.  The automatic stats refresh
    stops working, and the memory for the repeat-mode callback request is
    leaked.  It can be confirmed using kmemleak.

    Fix the mistake on the list_add_tail() call.

    Link: https://lkml.kernel.org/r/20251014205939.1206-1-sj@kernel.org
    Fixes: 004ded6bee11 ("mm/damon: accept parallel damon_call() requests")
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Cc: <stable@vger.kernel.org>    [6.17+]
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:07 -04:00
Rafael Aquini c73e0c3263 mm/damon/core: fix memory leak of repeat mode damon_call_control objects
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 817383b34db1e7d2a74d2d2b51cb0eed1586253b
Author: Enze Li <lienze@kylinos.cn>
Date:   Tue Dec 2 16:23:40 2025 +0800

    mm/damon/core: fix memory leak of repeat mode damon_call_control objects

    A memory leak exists in the handling of repeat mode damon_call_control
    objects by kdamond_call().  While damon_call() correctly allows multiple
    repeat mode objects (with ->repeat set to true) to be added to the
    per-context list, kdamond_call() incorrectly processes them.

    The function moves all repeat mode objects from the context's list to a
    temporary list (repeat_controls).  However, it only moves the first object
    back to the context's list for future calls, leaving the remaining objects
    on the temporary list where they are abandoned and leaked.

    This patch fixes the leak by ensuring all repeat mode objects are properly
    re-added to the context's list.

    Note that the leak is not in the real world, and therefore no user is
    impacted.  It is only potential for imaginaray damon_call() use cases that
    do not exist in the tree for now.  In more detail, the leak happens only
    when the multiple repeat mode objects are assumed to be deallocated by
    kdamond_call() (damon_call_control->dealloc_on_cancel is set).  There is
    no such damon_call() use cases at the moment.

    Link: https://lkml.kernel.org/r/20251202082340.34178-1-lienze@kylinos.cn
    Fixes: 43df7676e550 ("mm/damon/core: introduce repeat mode damon_call()")
    Signed-off-by: Enze Li <lienze@kylinos.cn>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:06 -04:00
Rafael Aquini cdf413ab65 mm/damon/sysfs: check contexts->nr in repeat_call_fn
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit d0bde8e2f3d2fb9aaec15d9c01da0a01526c7a56
Author: Josh Law <objecting@objecting.org>
Date:   Sat Mar 21 10:54:26 2026 -0700

    mm/damon/sysfs: check contexts->nr in repeat_call_fn

    damon_sysfs_repeat_call_fn() calls damon_sysfs_upd_tuned_intervals(),
    damon_sysfs_upd_schemes_stats(), and
    damon_sysfs_upd_schemes_effective_quotas() without checking contexts->nr.
    If nr_contexts is set to 0 via sysfs while DAMON is running, these
    functions dereference contexts_arr[0] and cause a NULL pointer
    dereference.  Add the missing check.

    For example, the issue can be reproduced using DAMON sysfs interface and
    DAMON user-space tool (damo) [1] like below.

        $ sudo damo start --refresh_interval 1s
        $ echo 0 | sudo tee \
                /sys/kernel/mm/damon/admin/kdamonds/0/contexts/nr_contexts

    Link: https://patch.msgid.link/20260320163559.178101-3-objecting@objecting.org
    Link: https://lkml.kernel.org/r/20260321175427.86000-4-sj@kernel.org
    Link: https://github.com/damonitor/damo [1]
    Fixes: d809a7c64ba8 ("mm/damon/sysfs: implement refresh_ms file internal work")
    Signed-off-by: Josh Law <objecting@objecting.org>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Cc: <stable@vger.kernel.org>    [6.17+]
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:06 -04:00
Rafael Aquini 40e3caab87 mm/damon/sysfs: change next_update_jiffies to a global variable
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 9fd7bb5083d1e1027b8ac1e365c29921ab88b177
Author: Quanmin Yan <yanquanmin1@huawei.com>
Date:   Thu Oct 30 10:07:46 2025 +0800

    mm/damon/sysfs: change next_update_jiffies to a global variable

    In DAMON's damon_sysfs_repeat_call_fn(), time_before() is used to compare
    the current jiffies with next_update_jiffies to determine whether to
    update the sysfs files at this moment.

    On 32-bit systems, the kernel initializes jiffies to "-5 minutes" to make
    jiffies wrap bugs appear earlier. However, this causes time_before() in
    damon_sysfs_repeat_call_fn() to unexpectedly return true during the first
    5 minutes after boot on 32-bit systems (see [1] for more explanation,
    which fixes another jiffies-related issue before). As a result, DAMON
    does not update sysfs files during that period.

    There is also an issue unrelated to the system's word size[2]: if the
    user stops DAMON just after next_update_jiffies is updated and restarts
    it after 'refresh_ms' or a longer delay, next_update_jiffies will retain
    an older value, causing time_before() to return false and the update to
    happen earlier than expected.

    Fix these issues by making next_update_jiffies a global variable and
    initializing it each time DAMON is started.

    Link: https://lkml.kernel.org/r/20251030020746.967174-3-yanquanmin1@huawei.com
    Link: https://lkml.kernel.org/r/20250822025057.1740854-1-ekffu200098@gmail.com [1]
    Link: https://lore.kernel.org/all/20251029013038.66625-1-sj@kernel.org/ [2]
    Fixes: d809a7c64ba8 ("mm/damon/sysfs: implement refresh_ms file internal work")
    Suggested-by: SeongJae Park <sj@kernel.org>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Signed-off-by: Quanmin Yan <yanquanmin1@huawei.com>
    Cc: Kefeng Wang <wangkefeng.wang@huawei.com>
    Cc: ze zuo <zuoze1@huawei.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:06 -04:00
Rafael Aquini 04da67f20c selftests/proc: fix string literal warning in proc-maps-race.c
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit ab5ac789efa985e42bdb5b5e8a7a4ad84935d44e
Author: Sukrut Heroorkar <hsukrut3@gmail.com>
Date:   Tue Aug 5 00:56:14 2025 +0200

    selftests/proc: fix string literal warning in proc-maps-race.c

    This change resolves non literal string format warning invoked for
    proc-maps-race.c while compiling.

    proc-maps-race.c:205:17: warning: format not a string literal and no format arguments [-Wformat-security]
     205 |                 printf(text);
         |                 ^~~~~~
    proc-maps-race.c:209:17: warning: format not a string literal and no format arguments [-Wformat-security]
     209 |                 printf(text);
         |                 ^~~~~~
    proc-maps-race.c: In function `print_last_lines':
    proc-maps-race.c:224:9: warning: format not a string literal and no format arguments [-Wformat-security]
     224 |         printf(start);
         |         ^~~~~~

    Add string format specifier %s for the printf calls in both
    print_first_lines() and print_last_lines() thus resolving the warnings.

    The test executes fine after this change thus causing no effect to the
    functional behavior of the test.

    Link: https://lkml.kernel.org/r/20250804225633.841777-1-hsukrut3@gmail.com
    Fixes: aadc099c480f ("selftests/proc: add verbose mode for /proc/pid/maps tearing tests")
    Signed-off-by: Sukrut Heroorkar <hsukrut3@gmail.com>
    Acked-by: Suren Baghdasaryan <surenb@google.com>
    Cc: David Hunter <david.hunter.linux@gmail.com>
    Cc: Shuah Khan <shuah@kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:06 -04:00
Rafael Aquini 69cbdcc960 mm/mseal: update VMA end correctly on merge
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
Conflicts:
  * minor context diff in the 2nd hunk due to RHEL-10 missing upstream's
    v6.19 commit 9119d6c2095b ("mm: update vma_modify_flags() to handle residual
    flags, document") which ends up const'ifying 'curr_end'. We do not take the
    type qualifier in this backport, keeping consistency with the current codebase.

commit 5ac9c7c2efd0d3c0c2d3bc6e9cd900d3ab6af27a
Author: Lorenzo Stoakes (Oracle) <ljs@kernel.org>
Date:   Fri Mar 27 17:31:04 2026 +0000

    mm/mseal: update VMA end correctly on merge

    Previously we stored the end of the current VMA in curr_end, and then upon
    iterating to the next VMA updated curr_start to curr_end to advance to the
    next VMA.

    However, this doesn't take into account the fact that a VMA might be
    updated due to a merge by vma_modify_flags(), which can result in curr_end
    being stale and thus, upon setting curr_start to curr_end, ending up with
    an incorrect curr_start on the next iteration.

    Resolve the issue by setting curr_end to vma->vm_end unconditionally to
    ensure this value remains updated should this occur.

    While we're here, eliminate this entire class of bug by simply setting
    const curr_[start/end] to be clamped to the input range and VMAs, which
    also happens to simplify the logic.

    Link: https://lkml.kernel.org/r/20260327173104.322405-1-ljs@kernel.org
    Fixes: 6c2da14ae1e0 ("mm/mseal: rework mseal apply logic")
    Signed-off-by: Lorenzo Stoakes (Oracle) <ljs@kernel.org>
    Reported-by: Antonius <antonius@bluedragonsec.com>
    Closes: https://lore.kernel.org/linux-mm/CAK8a0jwWGj9-SgFk0yKFh7i8jMkwKm5b0ao9=kmXWjO54veX2g@mail.gmail.com/
    Suggested-by: David Hildenbrand (ARM) <david@kernel.org>
    Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
    Reviewed-by: Pedro Falcato <pfalcato@suse.de>
    Acked-by: David Hildenbrand (Arm) <david@kernel.org>
    Cc: Jann Horn <jannh@google.com>
    Cc: Jeff Xu <jeffxu@chromium.org>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:06 -04:00
Rafael Aquini f7c212a1a9 init/main.c: fix boot time tracing crash
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 669602b5b7386e4fa00fc67b045ca3fd816e685d
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Sun Aug 24 16:07:59 2025 +0300

    init/main.c: fix boot time tracing crash

    Steven Rostedt reported a crash with "ftrace=function" kernel command
    line:

    [    0.159269] BUG: kernel NULL pointer dereference, address: 000000000000001c
    [    0.160254] #PF: supervisor read access in kernel mode
    [    0.160975] #PF: error_code(0x0000) - not-present page
    [    0.161697] PGD 0 P4D 0
    [    0.162055] Oops: Oops: 0000 [#1] SMP PTI
    [    0.162619] CPU: 0 UID: 0 PID: 0 Comm: swapper Not tainted 6.17.0-rc2-test-00006-g48d06e78b7cb-dirty #9 PREEMPT(undef)
    [    0.164141] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
    [    0.165439] RIP: 0010:kmem_cache_alloc_noprof (mm/slub.c:4237)
    [ 0.166186] Code: 90 90 90 f3 0f 1e fa 0f 1f 44 00 00 55 48 89 e5 41 57 41 56 41 55 41 54 49 89 fc 53 48 83 e4 f0 48 83 ec 20 8b 05 c9 b6 7e 01 <44> 8b 77 1c 65 4c 8b 2d b5 ea 20 02 4c 89 6c 24 18 41 89 f5 21 f0
    [    0.168811] RSP: 0000:ffffffffb2e03b30 EFLAGS: 00010086
    [    0.169545] RAX: 0000000001fff33f RBX: 0000000000000000 RCX: 0000000000000000
    [    0.170544] RDX: 0000000000002800 RSI: 0000000000002800 RDI: 0000000000000000
    [    0.171554] RBP: ffffffffb2e03b80 R08: 0000000000000004 R09: ffffffffb2e03c90
    [    0.172549] R10: ffffffffb2e03c90 R11: 0000000000000000 R12: 0000000000000000
    [    0.173544] R13: ffffffffb2e03c90 R14: ffffffffb2e03c90 R15: 0000000000000001
    [    0.174542] FS:  0000000000000000(0000) GS:ffff9d2808114000(0000) knlGS:0000000000000000
    [    0.175684] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
    [    0.176486] CR2: 000000000000001c CR3: 000000007264c001 CR4: 00000000000200b0
    [    0.177483] Call Trace:
    [    0.177828]  <TASK>
    [    0.178123] mas_alloc_nodes (lib/maple_tree.c:176 (discriminator 2) lib/maple_tree.c:1255 (discriminator 2))
    [    0.178692] mas_store_gfp (lib/maple_tree.c:5468)
    [    0.179223] execmem_cache_add_locked (mm/execmem.c:207)
    [    0.179870] execmem_alloc (mm/execmem.c:213 mm/execmem.c:313 mm/execmem.c:335 mm/execmem.c:475)
    [    0.180397] ? ftrace_caller (arch/x86/kernel/ftrace_64.S:169)
    [    0.180922] ? __pfx_ftrace_caller (arch/x86/kernel/ftrace_64.S:158)
    [    0.181517] execmem_alloc_rw (mm/execmem.c:487)
    [    0.182052] arch_ftrace_update_trampoline (arch/x86/kernel/ftrace.c:266 arch/x86/kernel/ftrace.c:344 arch/x86/kernel/ftrace.c:474)
    [    0.182778] ? ftrace_caller_op_ptr (arch/x86/kernel/ftrace_64.S:182)
    [    0.183388] ftrace_update_trampoline (kernel/trace/ftrace.c:7947)
    [    0.184024] __register_ftrace_function (kernel/trace/ftrace.c:368)
    [    0.184682] ftrace_startup (kernel/trace/ftrace.c:3048)
    [    0.185205] ? __pfx_function_trace_call (kernel/trace/trace_functions.c:210)
    [    0.185877] register_ftrace_function_nolock (kernel/trace/ftrace.c:8717)
    [    0.186595] register_ftrace_function (kernel/trace/ftrace.c:8745)
    [    0.187254] ? __pfx_function_trace_call (kernel/trace/trace_functions.c:210)
    [    0.187924] function_trace_init (kernel/trace/trace_functions.c:170)
    [    0.188499] tracing_set_tracer (kernel/trace/trace.c:5916 kernel/trace/trace.c:6349)
    [    0.189088] register_tracer (kernel/trace/trace.c:2391)
    [    0.189642] early_trace_init (kernel/trace/trace.c:11075 kernel/trace/trace.c:11149)
    [    0.190204] start_kernel (init/main.c:970)
    [    0.190732] x86_64_start_reservations (arch/x86/kernel/head64.c:307)
    [    0.191381] x86_64_start_kernel (??:?)
    [    0.191955] common_startup_64 (arch/x86/kernel/head_64.S:419)
    [    0.192534]  </TASK>
    [    0.192839] Modules linked in:
    [    0.193267] CR2: 000000000000001c
    [    0.193730] ---[ end trace 0000000000000000 ]---

    The crash happens because on x86 ftrace allocations from execmem require
    maple tree to be initialized.

    Move maple tree initialization that depends only on slab availability
    earlier in boot so that it will happen right after mm_core_init().

    Link: https://lkml.kernel.org/r/20250824130759.1732736-1-rppt@kernel.org
    Fixes: 5d79c2be5081 ("x86/ftrace: enable EXECMEM_ROX_CACHE for ftrace allocations")
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Reported-by: Steven Rostedt (Google) <rostedt@goodmis.org>
    Tested-by: Steven Rostedt (Google) <rostedt@goodmis.org>
    Closes: https://lore.kernel.org/all/20250820184743.0302a8b5@gandalf.local.home/
    Reviewed-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
    Reviewed-by: Liam R. Howlett <Liam.Howlett@oracle.com>
    Cc: Borislav Betkov <bp@alien8.de>
    Cc: Ingo Molnar <mingo@redhat.com>
    Cc: Peter Zijlstra <peterz@infradead.org>
    Cc: Thomas Gleinxer <tglx@linutronix.de>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:06 -04:00
Rafael Aquini c849ac9c03 selftests/mm: fix usage of FORCE_READ() in cow tests
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit bce1dabd310e87fefe0645fec9ba98b84d37e418
Author: Kevin Brodsky <kevin.brodsky@arm.com>
Date:   Thu Jan 22 17:02:19 2026 +0000

    selftests/mm: fix usage of FORCE_READ() in cow tests

    Commit 5bbc2b785e63 ("selftests/mm: fix FORCE_READ to read input value
    correctly") modified FORCE_READ() to take a value instead of a pointer.
    It also changed most of the call sites accordingly, but missed many of
    them in cow.c.  In those cases, we ended up with the pointer itself being
    read, not the memory it points to.

    No failure occurred as a result, so it looks like the tests work just fine
    without faulting in.  However, the huge_zeropage tests explicitly check
    that pages are populated, so those became skipped.

    Convert all the remaining FORCE_READ() to fault in the mapped page, as was
    originally intended.  This allows the huge_zeropage tests to run again (3
    tests in total).

    Link: https://lkml.kernel.org/r/20260122170224.4056513-5-kevin.brodsky@arm.com
    Fixes: 5bbc2b785e63 ("selftests/mm: fix FORCE_READ to read input value correctly")
    Signed-off-by: Kevin Brodsky <kevin.brodsky@arm.com>
    Acked-by: SeongJae Park <sj@kernel.org>
    Reviewed-by: wang lian <lianux.mm@gmail.com>
    Acked-by: David Hildenbrand (Red Hat) <david@kernel.org>
    Reviewed-by: Dev Jain <dev.jain@arm.com>
    Cc: Jason Gunthorpe <jgg@nvidia.com>
    Cc: John Hubbard <jhubbard@nvidia.com>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Mark Brown <broonie@kernel.org>
    Cc: Paolo Abeni <pabeni@redhat.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Shuah Khan <shuah@kernel.org>
    Cc: Usama Anjum <Usama.Anjum@arm.com>
    Cc: Yunsheng Lin <linyunsheng@huawei.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:05 -04:00
Rafael Aquini 74e8ec105d mm/damon/sysfs: dealloc repeat_call_control if damon_call() fails
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2026-31653

commit 0199390a6b92fc21860e1b858abf525c7e73b956
Author: SeongJae Park <sj@kernel.org>
Date:   Thu Mar 26 17:32:22 2026 -0700

    mm/damon/sysfs: dealloc repeat_call_control if damon_call() fails

    damon_call() for repeat_call_control of DAMON_SYSFS could fail if somehow
    the kdamond is stopped before the damon_call().  It could happen, for
    example, when te damon context was made for monitroing of a virtual
    address processes, and the process is terminated immediately, before the
    damon_call() invocation.  In the case, the dyanmically allocated
    repeat_call_control is not deallocated and leaked.

    Fix the leak by deallocating the repeat_call_control under the
    damon_call() failure.

    This issue is discovered by sashiko [1].

    Link: https://lkml.kernel.org/r/20260327003224.55752-1-sj@kernel.org
    Link: https://lore.kernel.org/20260320020630.962-1-sj@kernel.org [1]
    Fixes: 04a06b139ec0 ("mm/damon/sysfs: use dynamically allocated repeat mode damon_call_control")
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Cc: <stable@vger.kernel.org>    [6.17+]
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:05 -04:00
Rafael Aquini 128f0c05b2 kmsan: fix out-of-bounds access to shadow memory
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-40008

commit 85e1ff61060a765d91ee62dc5606d4d547d9d105
Author: Eric Biggers <ebiggers@kernel.org>
Date:   Thu Sep 11 12:58:58 2025 -0700

    kmsan: fix out-of-bounds access to shadow memory

    Running sha224_kunit on a KMSAN-enabled kernel results in a crash in
    kmsan_internal_set_shadow_origin():

        BUG: unable to handle page fault for address: ffffbc3840291000
        #PF: supervisor read access in kernel mode
        #PF: error_code(0x0000) - not-present page
        PGD 1810067 P4D 1810067 PUD 192d067 PMD 3c17067 PTE 0
        Oops: 0000 [#1] SMP NOPTI
        CPU: 0 UID: 0 PID: 81 Comm: kunit_try_catch Tainted: G                 N  6.17.0-rc3 #10 PREEMPT(voluntary)
        Tainted: [N]=TEST
        Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS rel-1.17.0-0-gb52ca86e094d-prebuilt.qemu.org 04/01/2014
        RIP: 0010:kmsan_internal_set_shadow_origin+0x91/0x100
        [...]
        Call Trace:
        <TASK>
        __msan_memset+0xee/0x1a0
        sha224_final+0x9e/0x350
        test_hash_buffer_overruns+0x46f/0x5f0
        ? kmsan_get_shadow_origin_ptr+0x46/0xa0
        ? __pfx_test_hash_buffer_overruns+0x10/0x10
        kunit_try_run_case+0x198/0xa00

    This occurs when memset() is called on a buffer that is not 4-byte aligned
    and extends to the end of a guard page, i.e.  the next page is unmapped.

    The bug is that the loop at the end of kmsan_internal_set_shadow_origin()
    accesses the wrong shadow memory bytes when the address is not 4-byte
    aligned.  Since each 4 bytes are associated with an origin, it rounds the
    address and size so that it can access all the origins that contain the
    buffer.  However, when it checks the corresponding shadow bytes for a
    particular origin, it incorrectly uses the original unrounded shadow
    address.  This results in reads from shadow memory beyond the end of the
    buffer's shadow memory, which crashes when that memory is not mapped.

    To fix this, correctly align the shadow address before accessing the 4
    shadow bytes corresponding to each origin.

    Link: https://lkml.kernel.org/r/20250911195858.394235-1-ebiggers@kernel.org
    Fixes: 2ef3cec44c ("kmsan: do not wipe out origin when doing partial unpoisoning")
    Signed-off-by: Eric Biggers <ebiggers@kernel.org>
    Tested-by: Alexander Potapenko <glider@google.com>
    Reviewed-by: Alexander Potapenko <glider@google.com>
    Cc: Dmitriy Vyukov <dvyukov@google.com>
    Cc: Marco Elver <elver@google.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:05 -04:00
Rafael Aquini 66c1958997 mm/hugetlb: fix copy_hugetlb_page_range() to use ->pt_share_count
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 14967a9c7d247841b0312c48dcf8cd29e55a4cc8
Author: Jane Chu <jane.chu@oracle.com>
Date:   Mon Sep 15 18:45:20 2025 -0600

    mm/hugetlb: fix copy_hugetlb_page_range() to use ->pt_share_count

    commit 59d9094df3d79 ("mm: hugetlb: independent PMD page table shared
    count") introduced ->pt_share_count dedicated to hugetlb PMD share count
    tracking, but omitted fixing copy_hugetlb_page_range(), leaving the
    function relying on page_count() for tracking that no longer works.

    When lazy page table copy for hugetlb is disabled, that is, revert commit
    bcd51a3c67 ("hugetlb: lazy page table copies in fork()") fork()'ing with
    hugetlb PMD sharing quickly lockup -

    [  239.446559] watchdog: BUG: soft lockup - CPU#75 stuck for 27s!
    [  239.446611] RIP: 0010:native_queued_spin_lock_slowpath+0x7e/0x2e0
    [  239.446631] Call Trace:
    [  239.446633]  <TASK>
    [  239.446636]  _raw_spin_lock+0x3f/0x60
    [  239.446639]  copy_hugetlb_page_range+0x258/0xb50
    [  239.446645]  copy_page_range+0x22b/0x2c0
    [  239.446651]  dup_mmap+0x3e2/0x770
    [  239.446654]  dup_mm.constprop.0+0x5e/0x230
    [  239.446657]  copy_process+0xd17/0x1760
    [  239.446660]  kernel_clone+0xc0/0x3e0
    [  239.446661]  __do_sys_clone+0x65/0xa0
    [  239.446664]  do_syscall_64+0x82/0x930
    [  239.446668]  ? count_memcg_events+0xd2/0x190
    [  239.446671]  ? syscall_trace_enter+0x14e/0x1f0
    [  239.446676]  ? syscall_exit_work+0x118/0x150
    [  239.446677]  ? arch_exit_to_user_mode_prepare.constprop.0+0x9/0xb0
    [  239.446681]  ? clear_bhb_loop+0x30/0x80
    [  239.446684]  ? clear_bhb_loop+0x30/0x80
    [  239.446686]  entry_SYSCALL_64_after_hwframe+0x76/0x7e

    There are two options to resolve the potential latent issue:
      1. warn against PMD sharing in copy_hugetlb_page_range(),
      2. fix it.
    This patch opts for the second option.
    While at it, simplify the comment, the details are not actually relevant
    anymore.

    Link: https://lkml.kernel.org/r/20250916004520.1604530-1-jane.chu@oracle.com
    Fixes: 59d9094df3d7 ("mm: hugetlb: independent PMD page table shared count")
    Signed-off-by: Jane Chu <jane.chu@oracle.com>
    Reviewed-by: Harry Yoo <harry.yoo@oracle.com>
    Acked-by: Oscar Salvador <osalvador@suse.de>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Jann Horn <jannh@google.com>
    Cc: Liu Shixin <liushixin2@huawei.com>
    Cc: Muchun Song <muchun.song@linux.dev>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:05 -04:00
Rafael Aquini 2d68a654e0 mm/damon/sysfs: use dynamically allocated repeat mode damon_call_control
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 04a06b139ec08aa63d7377f6d3e5218f8ddb1c5d
Author: SeongJae Park <sj@kernel.org>
Date:   Mon Sep 8 13:15:13 2025 -0700

    mm/damon/sysfs: use dynamically allocated repeat mode damon_call_control

    DAMON sysfs interface is using a single global repeat mode
    damon_call_control variable for refresh_ms handling, for all DAMON
    contexts.  As a result, when there are more than one context, the single
    global damon_call_control is unexpectedly over-written (corrupted).
    Particularly the ->link field is overwritten by the multiple contexts and
    this can cause a user hangup, and/or a kernel crash.  Fix it by using
    dynamically allocated damon_call_control object per DAMON context.

    Link: https://lkml.kernel.org/r/20250908201513.60802-3-sj@kernel.org
    Link: https://lore.kernel.org/20250904011738.930-1-yunjeong.mun@sk.com [1]
    Link: https://lore.kernel.org/20250905035411.39501-1-sj@kernel.org [2]
    Fixes: d809a7c64ba8 ("mm/damon/sysfs: implement refresh_ms file internal work")
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Reported-by: Yunjeong Mun <yunjeong.mun@sk.com>
    Closes: https://lore.kernel.org/20250904011738.930-1-yunjeong.mun@sk.com
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:05 -04:00
Rafael Aquini 23566d8021 mm/damon/core: introduce damon_call_control->dealloc_on_cancel
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit e6a0deb6fa5b0fc134ee2aa127d1cfc9456d8445
Author: SeongJae Park <sj@kernel.org>
Date:   Mon Sep 8 13:15:12 2025 -0700

    mm/damon/core: introduce damon_call_control->dealloc_on_cancel

    Patch series "mm/damon/sysfs: fix refresh_ms control overwriting on
    multi-kdamonds usages".

    Automatic esssential DAMON/DAMOS status update feature of DAMON sysfs
    interface (refresh_ms) is broken [1] for multiple DAMON contexts
    (kdamonds) use case, since it uses a global single damon_call_control
    object for all created DAMON contexts.  The fields of the object,
    particularly the list field is over-written for the contexts and it makes
    unexpected results including user-space hangup and kernel crashes [2].
    Fix it by extending damon_call_control for the use case and updating the
    usage on DAMON sysfs interface to use per-context dynamically allocated
    damon_call_control object.

    This patch (of 2):

    When damon_call_control->repeat is set, damon_call() is executed
    asynchronously, and is eventually canceled when kdamond finishes.  If the
    damon_call_control object is dynamically allocated, finding the place to
    deallocate the object is difficult.  Introduce a new damon_call_control
    field, namely dealloc_on_cancel, to ask the kdamond deallocates those
    dynamically allocated objects when those are canceled.

    Link: https://lkml.kernel.org/r/20250908201513.60802-3-sj@kernel.org
    Link: https://lkml.kernel.org/r/20250908201513.60802-2-sj@kernel.org
    Fixes: d809a7c64ba8 ("mm/damon/sysfs: implement refresh_ms file internal work")
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Cc: Yunjeong Mun <yunjeong.mun@sk.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:05 -04:00
Rafael Aquini fa3e26fdc1 mm: folio_may_be_lru_cached() unless folio_test_large()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 2da6de30e60dd9bb14600eff1cc99df2fa2ddae3
Author: Hugh Dickins <hughd@google.com>
Date:   Mon Sep 8 15:23:15 2025 -0700

    mm: folio_may_be_lru_cached() unless folio_test_large()

    mm/swap.c and mm/mlock.c agree to drain any per-CPU batch as soon as a
    large folio is added: so collect_longterm_unpinnable_folios() just wastes
    effort when calling lru_add_drain[_all]() on a large folio.

    But although there is good reason not to batch up PMD-sized folios, we
    might well benefit from batching a small number of low-order mTHPs (though
    unclear how that "small number" limitation will be implemented).

    So ask if folio_may_be_lru_cached() rather than !folio_test_large(), to
    insulate those particular checks from future change.  Name preferred to
    "folio_is_batchable" because large folios can well be put on a batch: it's
    just the per-CPU LRU caches, drained much later, which need care.

    Marked for stable, to counter the increase in lru_add_drain_all()s from
    "mm/gup: check ref_count instead of lru before migration".

    Link: https://lkml.kernel.org/r/57d2eaf8-3607-f318-e0c5-be02dce61ad0@google.com
    Fixes: 9a4e9f3b2d ("mm: update get_user_pages_longterm to migrate pages allocated from CMA region")
    Signed-off-by: Hugh Dickins <hughd@google.com>
    Suggested-by: David Hildenbrand <david@redhat.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: "Aneesh Kumar K.V" <aneesh.kumar@kernel.org>
    Cc: Axel Rasmussen <axelrasmussen@google.com>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Christoph Hellwig <hch@infradead.org>
    Cc: Jason Gunthorpe <jgg@ziepe.ca>
    Cc: Johannes Weiner <hannes@cmpxchg.org>
    Cc: John Hubbard <jhubbard@nvidia.com>
    Cc: Keir Fraser <keirf@google.com>
    Cc: Konstantin Khlebnikov <koct9i@gmail.com>
    Cc: Li Zhe <lizhe.67@bytedance.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Peter Xu <peterx@redhat.com>
    Cc: Rik van Riel <riel@surriel.com>
    Cc: Shivank Garg <shivankg@amd.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Wei Xu <weixugc@google.com>
    Cc: Will Deacon <will@kernel.org>
    Cc: yangge <yangge1116@126.com>
    Cc: Yuanchu Xie <yuanchu@google.com>
    Cc: Yu Zhao <yuzhao@google.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:05 -04:00
Rafael Aquini b984531407 mm: revert "mm: vmscan.c: fix OOM on swap stress test"
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 8d79ed36bfc83d0583ab72216b7980340478cdfb
Author: Hugh Dickins <hughd@google.com>
Date:   Mon Sep 8 15:21:12 2025 -0700

    mm: revert "mm: vmscan.c: fix OOM on swap stress test"

    This reverts commit 0885ef4705: that was a fix to the reverted
    33dfe9204f.

    Link: https://lkml.kernel.org/r/aa0e9d67-fbcd-9d79-88a1-641dfbe1d9d1@google.com
    Signed-off-by: Hugh Dickins <hughd@google.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: "Aneesh Kumar K.V" <aneesh.kumar@kernel.org>
    Cc: Axel Rasmussen <axelrasmussen@google.com>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Christoph Hellwig <hch@infradead.org>
    Cc: Jason Gunthorpe <jgg@ziepe.ca>
    Cc: Johannes Weiner <hannes@cmpxchg.org>
    Cc: John Hubbard <jhubbard@nvidia.com>
    Cc: Keir Fraser <keirf@google.com>
    Cc: Konstantin Khlebnikov <koct9i@gmail.com>
    Cc: Li Zhe <lizhe.67@bytedance.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Peter Xu <peterx@redhat.com>
    Cc: Rik van Riel <riel@surriel.com>
    Cc: Shivank Garg <shivankg@amd.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Wei Xu <weixugc@google.com>
    Cc: Will Deacon <will@kernel.org>
    Cc: yangge <yangge1116@126.com>
    Cc: Yuanchu Xie <yuanchu@google.com>
    Cc: Yu Zhao <yuzhao@google.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:04 -04:00
Rafael Aquini a3ed5b4d5d mm: revert "mm/gup: clear the LRU flag of a page before adding to LRU batch"
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit afb99e9f500485160f34b8cad6d3763ada3e80e8
Author: Hugh Dickins <hughd@google.com>
Date:   Mon Sep 8 15:19:17 2025 -0700

    mm: revert "mm/gup: clear the LRU flag of a page before adding to LRU batch"

    This reverts commit 33dfe9204f: now that
    collect_longterm_unpinnable_folios() is checking ref_count instead of lru,
    and mlock/munlock do not participate in the revised LRU flag clearing,
    those changes are misleading, and enlarge the window during which
    mlock/munlock may miss an mlock_count update.

    It is possible (I'd hesitate to claim probable) that the greater
    likelihood of missed mlock_count updates would explain the "Realtime
    threads delayed due to kcompactd0" observed on 6.12 in the Link below.  If
    that is the case, this reversion will help; but a complete solution needs
    also a further patch, beyond the scope of this series.

    Included some 80-column cleanup around folio_batch_add_and_move().

    The role of folio_test_clear_lru() (before taking per-memcg lru_lock) is
    questionable since 6.13 removed mem_cgroup_move_account() etc; but perhaps
    there are still some races which need it - not examined here.

    Link: https://lore.kernel.org/linux-mm/DU0PR01MB10385345F7153F334100981888259A@DU0PR01MB10385.eurprd01.prod.exchangelabs.com/
    Link: https://lkml.kernel.org/r/05905d7b-ed14-68b1-79d8-bdec30367eba@google.com
    Signed-off-by: Hugh Dickins <hughd@google.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: "Aneesh Kumar K.V" <aneesh.kumar@kernel.org>
    Cc: Axel Rasmussen <axelrasmussen@google.com>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Christoph Hellwig <hch@infradead.org>
    Cc: Jason Gunthorpe <jgg@ziepe.ca>
    Cc: Johannes Weiner <hannes@cmpxchg.org>
    Cc: John Hubbard <jhubbard@nvidia.com>
    Cc: Keir Fraser <keirf@google.com>
    Cc: Konstantin Khlebnikov <koct9i@gmail.com>
    Cc: Li Zhe <lizhe.67@bytedance.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Peter Xu <peterx@redhat.com>
    Cc: Rik van Riel <riel@surriel.com>
    Cc: Shivank Garg <shivankg@amd.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Wei Xu <weixugc@google.com>
    Cc: Will Deacon <will@kernel.org>
    Cc: yangge <yangge1116@126.com>
    Cc: Yuanchu Xie <yuanchu@google.com>
    Cc: Yu Zhao <yuzhao@google.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:04 -04:00
Rafael Aquini 5a291e8fa0 mm/gup: local lru_add_drain() to avoid lru_add_drain_all()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit a09a8a1fbb374e0053b97306da9dbc05bd384685
Author: Hugh Dickins <hughd@google.com>
Date:   Mon Sep 8 15:16:53 2025 -0700

    mm/gup: local lru_add_drain() to avoid lru_add_drain_all()

    In many cases, if collect_longterm_unpinnable_folios() does need to drain
    the LRU cache to release a reference, the cache in question is on this
    same CPU, and much more efficiently drained by a preliminary local
    lru_add_drain(), than the later cross-CPU lru_add_drain_all().

    Marked for stable, to counter the increase in lru_add_drain_all()s from
    "mm/gup: check ref_count instead of lru before migration".  Note for clean
    backports: can take 6.16 commit a03db236aebf ("gup: optimize longterm
    pin_user_pages() for large folio") first.

    Link: https://lkml.kernel.org/r/66f2751f-283e-816d-9530-765db7edc465@google.com
    Signed-off-by: Hugh Dickins <hughd@google.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: "Aneesh Kumar K.V" <aneesh.kumar@kernel.org>
    Cc: Axel Rasmussen <axelrasmussen@google.com>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Christoph Hellwig <hch@infradead.org>
    Cc: Jason Gunthorpe <jgg@ziepe.ca>
    Cc: Johannes Weiner <hannes@cmpxchg.org>
    Cc: John Hubbard <jhubbard@nvidia.com>
    Cc: Keir Fraser <keirf@google.com>
    Cc: Konstantin Khlebnikov <koct9i@gmail.com>
    Cc: Li Zhe <lizhe.67@bytedance.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Peter Xu <peterx@redhat.com>
    Cc: Rik van Riel <riel@surriel.com>
    Cc: Shivank Garg <shivankg@amd.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Wei Xu <weixugc@google.com>
    Cc: Will Deacon <will@kernel.org>
    Cc: yangge <yangge1116@126.com>
    Cc: Yuanchu Xie <yuanchu@google.com>
    Cc: Yu Zhao <yuzhao@google.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:04 -04:00
Rafael Aquini 7f3d97e83e mm/damon/sysfs: fix use-after-free in state_show()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-39877

commit 3260a3f0828e06f5f13fac69fb1999a6d60d9cff
Author: Stanislav Fort <stanislav.fort@aisle.com>
Date:   Fri Sep 5 13:10:46 2025 +0300

    mm/damon/sysfs: fix use-after-free in state_show()

    state_show() reads kdamond->damon_ctx without holding damon_sysfs_lock.
    This allows a use-after-free race:

    CPU 0                         CPU 1
    -----                         -----
    state_show()                  damon_sysfs_turn_damon_on()
    ctx = kdamond->damon_ctx;     mutex_lock(&damon_sysfs_lock);
                                  damon_destroy_ctx(kdamond->damon_ctx);
                                  kdamond->damon_ctx = NULL;
                                  mutex_unlock(&damon_sysfs_lock);
    damon_is_running(ctx);        /* ctx is freed */
    mutex_lock(&ctx->kdamond_lock); /* UAF */

    (The race can also occur with damon_sysfs_kdamonds_rm_dirs() and
    damon_sysfs_kdamond_release(), which free or replace the context under
    damon_sysfs_lock.)

    Fix by taking damon_sysfs_lock before dereferencing the context, mirroring
    the locking used in pid_show().

    The bug has existed since state_show() first accessed kdamond->damon_ctx.

    Link: https://lkml.kernel.org/r/20250905101046.2288-1-disclosure@aisle.com
    Fixes: a61ea561c8 ("mm/damon/sysfs: link DAMON for virtual address spaces monitoring")
    Signed-off-by: Stanislav Fort <disclosure@aisle.com>
    Reported-by: Stanislav Fort <disclosure@aisle.com>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:04 -04:00
Rafael Aquini 19cf64cd76 mm/vmalloc, mm/kasan: respect gfp mask in kasan_populate_vmalloc()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-39910

commit 79357cd06d41d0f5a11b17d7c86176e395d10ef2
Author: Uladzislau Rezki (Sony) <urezki@gmail.com>
Date:   Sun Aug 31 14:10:58 2025 +0200

    mm/vmalloc, mm/kasan: respect gfp mask in kasan_populate_vmalloc()

    kasan_populate_vmalloc() and its helpers ignore the caller's gfp_mask and
    always allocate memory using the hardcoded GFP_KERNEL flag.  This makes
    them inconsistent with vmalloc(), which was recently extended to support
    GFP_NOFS and GFP_NOIO allocations.

    Page table allocations performed during shadow population also ignore the
    external gfp_mask.  To preserve the intended semantics of GFP_NOFS and
    GFP_NOIO, wrap the apply_to_page_range() calls into the appropriate
    memalloc scope.

    xfs calls vmalloc with GFP_NOFS, so this bug could lead to deadlock.

    There was a report here
    https://lkml.kernel.org/r/686ea951.050a0220.385921.0016.GAE@google.com

    This patch:
     - Extends kasan_populate_vmalloc() and helpers to take gfp_mask;
     - Passes gfp_mask down to alloc_pages_bulk() and __get_free_page();
     - Enforces GFP_NOFS/NOIO semantics with memalloc_*_save()/restore()
       around apply_to_page_range();
     - Updates vmalloc.c and percpu allocator call sites accordingly.

    Link: https://lkml.kernel.org/r/20250831121058.92971-1-urezki@gmail.com
    Fixes: 451769ebb7 ("mm/vmalloc: alloc GFP_NO{FS,IO} for vmalloc")
    Signed-off-by: Uladzislau Rezki (Sony) <urezki@gmail.com>
    Reported-by: syzbot+3470c9ffee63e4abafeb@syzkaller.appspotmail.com
    Reviewed-by: Andrey Ryabinin <ryabinin.a.a@gmail.com>
    Cc: Baoquan He <bhe@redhat.com>
    Cc: Michal Hocko <mhocko@kernel.org>
    Cc: Alexander Potapenko <glider@google.com>
    Cc: Andrey Konovalov <andreyknvl@gmail.com>
    Cc: Dmitry Vyukov <dvyukov@google.com>
    Cc: Vincenzo Frascino <vincenzo.frascino@arm.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:04 -04:00
Rafael Aquini 2539fd1cbd mm/mremap: fix regression in vrm->new_addr check
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 78d2d32f0b789d67cbe5cfea0c0714cb2446c37e
Author: Carlos Llamas <cmllamas@google.com>
Date:   Thu Aug 28 14:26:56 2025 +0000

    mm/mremap: fix regression in vrm->new_addr check

    Commit 3215eaceca87 ("mm/mremap: refactor initial parameter sanity
    checks") moved the sanity check for vrm->new_addr from mremap_to() to
    check_mremap_params().

    However, this caused a regression as vrm->new_addr is now checked even
    when MREMAP_FIXED and MREMAP_DONTUNMAP flags are not specified.  In this
    case, vrm->new_addr can be garbage and create unexpected failures.

    Fix this by moving the new_addr check after the vrm_implies_new_addr()
    guard.  This ensures that the new_addr is only checked when the user has
    specified one explicitly.

    Link: https://lkml.kernel.org/r/20250828142657.770502-1-cmllamas@google.com
    Fixes: 3215eaceca87 ("mm/mremap: refactor initial parameter sanity checks")
    Signed-off-by: Carlos Llamas <cmllamas@google.com>
    Reviewed-by: Liam R. Howlett <Liam.Howlett@oracle.com>
    Reviewed-by: Vlastimil Babka <vbabka@suse.cz>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Carlos Llamas <cmllamas@google.com>
    Cc: Jann Horn <jannh@google.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:04 -04:00
Rafael Aquini 7a9a0be1be percpu: fix race on alloc failed warning limit
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 7989fdce69ec0a928e136477da2aa208a191fba2
Author: Vlad Dumitrescu <vdumitrescu@nvidia.com>
Date:   Fri Aug 22 15:55:16 2025 -0700

    percpu: fix race on alloc failed warning limit

    The 'allocation failed, ...' warning messages can cause unlimited log
    spam, contrary to the implementation's intent.

    The warn_limit variable is accessed without synchronization.  If more than
    <warn_limit> threads enter the warning path at the same time, the variable
    will get decremented past 0.  Once it becomes negative, the non-zero check
    will always return true leading to unlimited log spam.

    Use atomic operation to access warn_limit and change condition to test for
    non-negative (>= 0) - atomic_dec_if_positive will return -1 once
    warn_limit becomes 0.  Continue to print disable message alongside the
    last warning.

    While the change cited in Fixes is only adjacent, the warning limit
    implementation was correct before it.  Only non-atomic allocations were
    considered for warnings, and those happened to hold pcpu_alloc_mutex while
    accessing warn_limit.

    [vdumitrescu@nvidia.com: prevent warn_limit from going negative, per Christoph Lameter]
      Link: https://lkml.kernel.org/r/ee87cc59-2717-4dbb-8052-1d2692c5aaaa@nvidia.com
    Link: https://lkml.kernel.org/r/ab22061a-a62f-4429-945b-744e5cc4ba35@nvidia.com
    Fixes: f7d77dfc91 ("mm/percpu.c: print error message too if atomic alloc failed")
    Signed-off-by: Vlad Dumitrescu <vdumitrescu@nvidia.com>
    Reviewed-by: Baoquan He <bhe@redhat.com>
    Cc: Christoph Lameter (Ampere) <cl@gentwo.org>
    Cc: Dennis Zhou <dennis@kernel.org>
    Cc: Tejun Heo <tj@kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:03 -04:00
Rafael Aquini 8366df508b mm/damon/reclaim: avoid divide-by-zero in damon_reclaim_apply_parameters()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-39916

commit e6b543ca9806d7bced863f43020e016ee996c057
Author: Quanmin Yan <yanquanmin1@huawei.com>
Date:   Wed Aug 27 19:58:58 2025 +0800

    mm/damon/reclaim: avoid divide-by-zero in damon_reclaim_apply_parameters()

    When creating a new scheme of DAMON_RECLAIM, the calculation of
    'min_age_region' uses 'aggr_interval' as the divisor, which may lead to
    division-by-zero errors.  Fix it by directly returning -EINVAL when such a
    case occurs.

    Link: https://lkml.kernel.org/r/20250827115858.1186261-3-yanquanmin1@huawei.com
    Fixes: f5a79d7c0c ("mm/damon: introduce struct damos_access_pattern")
    Signed-off-by: Quanmin Yan <yanquanmin1@huawei.com>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Cc: Kefeng Wang <wangkefeng.wang@huawei.com>
    Cc: ze zuo <zuoze1@huawei.com>
    Cc: <stable@vger.kernel.org>    [6.1+]
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:03 -04:00
Rafael Aquini 5a8b23d351 mm/damon/lru_sort: avoid divide-by-zero in damon_lru_sort_apply_parameters()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-39909

commit 711f19dfd783ffb37ca4324388b9c4cb87e71363
Author: Quanmin Yan <yanquanmin1@huawei.com>
Date:   Wed Aug 27 19:58:57 2025 +0800

    mm/damon/lru_sort: avoid divide-by-zero in damon_lru_sort_apply_parameters()

    Patch series "mm/damon: avoid divide-by-zero in DAMON module's parameters
    application".

    DAMON's RECLAIM and LRU_SORT modules perform no validation on
    user-configured parameters during application, which may lead to
    division-by-zero errors.

    Avoid the divide-by-zero by adding validation checks when DAMON modules
    attempt to apply the parameters.

    This patch (of 2):

    During the calculation of 'hot_thres' and 'cold_thres', either
    'sample_interval' or 'aggr_interval' is used as the divisor, which may
    lead to division-by-zero errors.  Fix it by directly returning -EINVAL
    when such a case occurs.  Additionally, since 'aggr_interval' is already
    required to be set no smaller than 'sample_interval' in damon_set_attrs(),
    only the case where 'sample_interval' is zero needs to be checked.

    Link: https://lkml.kernel.org/r/20250827115858.1186261-2-yanquanmin1@huawei.com
    Fixes: 40e983cca9 ("mm/damon: introduce DAMON-based LRU-lists Sorting")
    Signed-off-by: Quanmin Yan <yanquanmin1@huawei.com>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Cc: Kefeng Wang <wangkefeng.wang@huawei.com>
    Cc: ze zuo <zuoze1@huawei.com>
    Cc: <stable@vger.kernel.org>    [6.0+]
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:03 -04:00
Rafael Aquini b98990025d mm/damon/core: set quota->charged_from to jiffies at first charge window
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit ce652aac9c90a96c6536681d17518efb1f660fb8
Author: Sang-Heon Jeon <ekffu200098@gmail.com>
Date:   Fri Aug 22 11:50:57 2025 +0900

    mm/damon/core: set quota->charged_from to jiffies at first charge window

    Kernel initializes the "jiffies" timer as 5 minutes below zero, as shown
    in include/linux/jiffies.h

     /*
     * Have the 32 bit jiffies value wrap 5 minutes after boot
     * so jiffies wrap bugs show up earlier.
     */
     #define INITIAL_JIFFIES ((unsigned long)(unsigned int) (-300*HZ))

    And jiffies comparison help functions cast unsigned value to signed to
    cover wraparound

     #define time_after_eq(a,b) \
      (typecheck(unsigned long, a) && \
      typecheck(unsigned long, b) && \
      ((long)((a) - (b)) >= 0))

    When quota->charged_from is initialized to 0, time_after_eq() can
    incorrectly return FALSE even after reset_interval has elapsed.  This
    occurs when (jiffies - reset_interval) produces a value with MSB=1, which
    is interpreted as negative in signed arithmetic.

    This issue primarily affects 32-bit systems because: On 64-bit systems:
    MSB=1 values occur after ~292 million years from boot (assuming HZ=1000),
    almost impossible.

    On 32-bit systems: MSB=1 values occur during the first 5 minutes after
    boot, and the second half of every jiffies wraparound cycle, starting from
    day 25 (assuming HZ=1000)

    When above unexpected FALSE return from time_after_eq() occurs, the
    charging window will not reset.  The user impact depends on esz value at
    that time.

    If esz is 0, scheme ignores configured quotas and runs without any limits.

    If esz is not 0, scheme stops working once the quota is exhausted.  It
    remains until the charging window finally resets.

    So, change quota->charged_from to jiffies at damos_adjust_quota() when it
    is considered as the first charge window.  By this change, we can avoid
    unexpected FALSE return from time_after_eq()

    Link: https://lkml.kernel.org/r/20250822025057.1740854-1-ekffu200098@gmail.com
    Fixes: 2b8a248d58 ("mm/damon/schemes: implement size quota for schemes application speed control") # 5.16
    Signed-off-by: Sang-Heon Jeon <ekffu200098@gmail.com>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:03 -04:00
Rafael Aquini db9aa121e1 mm/hugetlb: add missing hugetlb_lock in __unmap_hugepage_range()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 21cc2b5c5062a256ae9064442d37ebbc23f5aef7
Author: Jeongjun Park <aha310510@gmail.com>
Date:   Sun Aug 24 03:21:15 2025 +0900

    mm/hugetlb: add missing hugetlb_lock in __unmap_hugepage_range()

    When restoring a reservation for an anonymous page, we need to check to
    freeing a surplus.  However, __unmap_hugepage_range() causes data race
    because it reads h->surplus_huge_pages without the protection of
    hugetlb_lock.

    And adjust_reservation is a boolean variable that indicates whether
    reservations for anonymous pages in each folio should be restored.
    Therefore, it should be initialized to false for each round of the loop.
    However, this variable is not initialized to false except when defining
    the current adjust_reservation variable.

    This means that once adjust_reservation is set to true even once within
    the loop, reservations for anonymous pages will be restored
    unconditionally in all subsequent rounds, regardless of the folio's state.

    To fix this, we need to add the missing hugetlb_lock, unlock the
    page_table_lock earlier so that we don't lock the hugetlb_lock inside the
    page_table_lock lock, and initialize adjust_reservation to false on each
    round within the loop.

    Link: https://lkml.kernel.org/r/20250823182115.1193563-1-aha310510@gmail.com
    Fixes: df7a6d1f64 ("mm/hugetlb: restore the reservation if needed")
    Signed-off-by: Jeongjun Park <aha310510@gmail.com>
    Reported-by: syzbot+417aeb05fd190f3a6da9@syzkaller.appspotmail.com
    Closes: https://syzkaller.appspot.com/bug?extid=417aeb05fd190f3a6da9
    Reviewed-by: Sidhartha Kumar <sidhartha.kumar@oracle.com>
    Cc: Breno Leitao <leitao@debian.org>
    Cc: David Hildenbrand <david@redhat.com>
    Cc: Muchun Song <muchun.song@linux.dev>
    Cc: Oscar Salvador <osalvador@suse.de>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:03 -04:00
Rafael Aquini e8f95c418d mm/memory_hotplug: fix hwpoisoned large folio handling in do_migrate_range()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 397f6d14f9c370e4910e6885294c340f39dedbf5
Author: Jinjiang Tu <tujinjiang@huawei.com>
Date:   Fri Jun 27 20:57:47 2025 +0800

    mm/memory_hotplug: fix hwpoisoned large folio handling in do_migrate_range()

    In do_migrate_range(), the hwpoisoned folio may be large folio, which
    can't be handled by unmap_poisoned_folio().

    I can reproduce this issue in qemu after adding delay in memory_failure()

    BUG: kernel NULL pointer dereference, address: 0000000000000000
    Workqueue: kacpi_hotplug acpi_hotplug_work_fn
    RIP: 0010:try_to_unmap_one+0x16a/0xfc0
      <TASK>
      rmap_walk_anon+0xda/0x1f0
      try_to_unmap+0x78/0x80
      ? __pfx_try_to_unmap_one+0x10/0x10
      ? __pfx_folio_not_mapped+0x10/0x10
      ? __pfx_folio_lock_anon_vma_read+0x10/0x10
      unmap_poisoned_folio+0x60/0x140
      do_migrate_range+0x4d1/0x600
      ? slab_memory_callback+0x6a/0x190
      ? notifier_call_chain+0x56/0xb0
      offline_pages+0x3e6/0x460
      memory_subsys_offline+0x130/0x1f0
      device_offline+0xba/0x110
      acpi_bus_offline+0xb7/0x130
      acpi_scan_hot_remove+0x77/0x290
      acpi_device_hotplug+0x1e0/0x240
      acpi_hotplug_work_fn+0x1a/0x30
      process_one_work+0x186/0x340

    Besides, do_migrate_range() may be called between memory_failure set
    hwpoison flag and isolate the folio from lru, so remove WARN_ON(). In other
    places, unmap_poisoned_folio() is called when the folio is isolated, obey
    it in do_migrate_range() too.

    [david@redhat.com: don't abort offlining, fixed typo, add comment]
    Link: https://lkml.kernel.org/r/3c214dff-9649-4015-840f-10de0e03ebe4@redhat.com
    Fixes: b15c87263a ("hwpoison, memory_hotplug: allow hwpoisoned pages to be offlined")
    Signed-off-by: Jinjiang Tu <tujinjiang@huawei.com>
    Signed-off-by: David Hildenbrand <david@redhat.com>
    Acked-by: Zi Yan <ziy@nvidia.com>
    Reviewed-by: Miaohe Lin <linmiaohe@huawei.com>
    Cc: Kefeng Wang <wangkefeng.wang@huawei.com>
    Cc: Luis Chamberalin <mcgrof@kernel.org>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Michal Hocko <mhocko@kernel.org>
    Cc: Oscar Salvador <osalvador@suse.de>
    Cc: Pankaj Raghav <kernel@pankajraghav.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:03 -04:00
Rafael Aquini 68074b3fec mm/khugepaged: fix the address passed to notifier on testing young
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 394bfac1c7f7b701c2c93834c5761b9c9ceeebcf
Author: Wei Yang <richard.weiyang@gmail.com>
Date:   Fri Aug 22 06:33:18 2025 +0000

    mm/khugepaged: fix the address passed to notifier on testing young

    Commit 8ee53820ed ("thp: mmu_notifier_test_young") introduced
    mmu_notifier_test_young(), but we are passing the wrong address.
    In xxx_scan_pmd(), the actual iteration address is "_address" not
    "address".  We seem to misuse the variable on the very beginning.

    Change it to the right one.

    [akpm@linux-foundation.org fix whitespace, per everyone]
    Link: https://lkml.kernel.org/r/20250822063318.11644-1-richard.weiyang@gmail.com
    Fixes: 8ee53820ed ("thp: mmu_notifier_test_young")
    Signed-off-by: Wei Yang <richard.weiyang@gmail.com>
    Reviewed-by: Dev Jain <dev.jain@arm.com>
    Reviewed-by: Zi Yan <ziy@nvidia.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Liam R. Howlett <Liam.Howlett@oracle.com>
    Cc: Nico Pache <npache@redhat.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:03 -04:00
Rafael Aquini 25d4e1e701 mm: fix possible deadlock in kmemleak
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit c873ccbb2f8db46ad9b4a989ea924b6d8f19abf1
Author: Gu Bowen <gubowen5@huawei.com>
Date:   Fri Aug 22 15:35:41 2025 +0800

    mm: fix possible deadlock in kmemleak

    There are some AA deadlock issues in kmemleak, similar to the situation
    reported by Breno [1].  The deadlock path is as follows:

    mem_pool_alloc()
      -> raw_spin_lock_irqsave(&kmemleak_lock, flags);
          -> pr_warn()
              -> netconsole subsystem
                 -> netpoll
                     -> __alloc_skb
                       -> __create_object
                         -> raw_spin_lock_irqsave(&kmemleak_lock, flags);

    To solve this problem, switch to printk_safe mode before printing warning
    message, this will redirect all printk()-s to a special per-CPU buffer,
    which will be flushed later from a safe context (irq work), and this
    deadlock problem can be avoided.  The proper API to use should be
    printk_deferred_enter()/printk_deferred_exit() [2].  Another way is to
    place the warn print after kmemleak is released.

    Link: https://lkml.kernel.org/r/20250822073541.1886469-1-gubowen5@huawei.com
    Link: https://lore.kernel.org/all/20250731-kmemleak_lock-v1-1-728fd470198f@debian.org/#t [1]
    Link: https://lore.kernel.org/all/5ca375cd-4a20-4807-b897-68b289626550@redhat.com/ [2]
    Signed-off-by: Gu Bowen <gubowen5@huawei.com>
    Reviewed-by: Waiman Long <longman@redhat.com>
    Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
    Reviewed-by: Breno Leitao <leitao@debian.org>
    Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
    Cc: John Ogness <john.ogness@linutronix.de>
    Cc: Lu Jialin <lujialin4@huawei.com>
    Cc: Petr Mladek <pmladek@suse.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:02 -04:00
Rafael Aquini 21dd97697b x86/mm/64: define ARCH_PAGE_TABLE_SYNC_MASK and arch_sync_kernel_mappings()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-39845

commit 6659d027998083fbb6d42a165b0c90dc2e8ba989
Author: Harry Yoo <harry.yoo@oracle.com>
Date:   Mon Aug 18 11:02:06 2025 +0900

    x86/mm/64: define ARCH_PAGE_TABLE_SYNC_MASK and arch_sync_kernel_mappings()

    Define ARCH_PAGE_TABLE_SYNC_MASK and arch_sync_kernel_mappings() to ensure
    page tables are properly synchronized when calling p*d_populate_kernel().

    For 5-level paging, synchronization is performed via
    pgd_populate_kernel().  In 4-level paging, pgd_populate() is a no-op, so
    synchronization is instead performed at the P4D level via
    p4d_populate_kernel().

    This fixes intermittent boot failures on systems using 4-level paging and
    a large amount of persistent memory:

      BUG: unable to handle page fault for address: ffffe70000000034
      #PF: supervisor write access in kernel mode
      #PF: error_code(0x0002) - not-present page
      PGD 0 P4D 0
      Oops: 0002 [#1] SMP NOPTI
      RIP: 0010:__init_single_page+0x9/0x6d
      Call Trace:
       <TASK>
       __init_zone_device_page+0x17/0x5d
       memmap_init_zone_device+0x154/0x1bb
       pagemap_range+0x2e0/0x40f
       memremap_pages+0x10b/0x2f0
       devm_memremap_pages+0x1e/0x60
       dev_dax_probe+0xce/0x2ec [device_dax]
       dax_bus_probe+0x6d/0xc9
       [... snip ...]
       </TASK>

    It also fixes a crash in vmemmap_set_pmd() caused by accessing vmemmap
    before sync_global_pgds() [1]:

      BUG: unable to handle page fault for address: ffffeb3ff1200000
      #PF: supervisor write access in kernel mode
      #PF: error_code(0x0002) - not-present page
      PGD 0 P4D 0
      Oops: Oops: 0002 [#1] PREEMPT SMP NOPTI
      Tainted: [W]=WARN
      RIP: 0010:vmemmap_set_pmd+0xff/0x230
       <TASK>
       vmemmap_populate_hugepages+0x176/0x180
       vmemmap_populate+0x34/0x80
       __populate_section_memmap+0x41/0x90
       sparse_add_section+0x121/0x3e0
       __add_pages+0xba/0x150
       add_pages+0x1d/0x70
       memremap_pages+0x3dc/0x810
       devm_memremap_pages+0x1c/0x60
       xe_devm_add+0x8b/0x100 [xe]
       xe_tile_init_noalloc+0x6a/0x70 [xe]
       xe_device_probe+0x48c/0x740 [xe]
       [... snip ...]

    Link: https://lkml.kernel.org/r/20250818020206.4517-4-harry.yoo@oracle.com
    Fixes: 8d400913c2 ("x86/vmemmap: handle unpopulated sub-pmd ranges")
    Signed-off-by: Harry Yoo <harry.yoo@oracle.com>
    Closes: https://lore.kernel.org/linux-mm/20250311114420.240341-1-gwan-gyeong.mun@intel.com [1]
    Suggested-by: Dave Hansen <dave.hansen@linux.intel.com>
    Acked-by: Kiryl Shutsemau <kas@kernel.org>
    Reviewed-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Alexander Potapenko <glider@google.com>
    Cc: Alistair Popple <apopple@nvidia.com>
    Cc: Andrey Konovalov <andreyknvl@gmail.com>
    Cc: Andrey Ryabinin <ryabinin.a.a@gmail.com>
    Cc: Andy Lutomirski <luto@kernel.org>
    Cc: "Aneesh Kumar K.V" <aneesh.kumar@linux.ibm.com>
    Cc: Anshuman Khandual <anshuman.khandual@arm.com>
    Cc: Ard Biesheuvel <ardb@kernel.org>
    Cc: Arnd Bergmann <arnd@arndb.de>
    Cc: bibo mao <maobibo@loongson.cn>
    Cc: Borislav Betkov <bp@alien8.de>
    Cc: Christoph Lameter (Ampere) <cl@gentwo.org>
    Cc: Dennis Zhou <dennis@kernel.org>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Dmitriy Vyukov <dvyukov@google.com>
    Cc: Ingo Molnar <mingo@redhat.com>
    Cc: Jane Chu <jane.chu@oracle.com>
    Cc: Joao Martins <joao.m.martins@oracle.com>
    Cc: Joerg Roedel <joro@8bytes.org>
    Cc: John Hubbard <jhubbard@nvidia.com>
    Cc: Kevin Brodsky <kevin.brodsky@arm.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Oscar Salvador <osalvador@suse.de>
    Cc: Peter Xu <peterx@redhat.com>
    Cc: Peter Zijlstra <peterz@infradead.org>
    Cc: Qi Zheng <zhengqi.arch@bytedance.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: Tejun Heo <tj@kernel.org>
    Cc: Thomas Gleinxer <tglx@linutronix.de>
    Cc: Thomas Huth <thuth@redhat.com>
    Cc: "Uladzislau Rezki (Sony)" <urezki@gmail.com>
    Cc: Vincenzo Frascino <vincenzo.frascino@arm.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:02 -04:00
Rafael Aquini 801d81d0d5 mm: introduce and use {pgd,p4d}_populate_kernel()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit f2d2f9598ebb0158a3fe17cda0106d7752e654a2
Author: Harry Yoo <harry.yoo@oracle.com>
Date:   Mon Aug 18 11:02:05 2025 +0900

    mm: introduce and use {pgd,p4d}_populate_kernel()

    Introduce and use {pgd,p4d}_populate_kernel() in core MM code when
    populating PGD and P4D entries for the kernel address space.  These
    helpers ensure proper synchronization of page tables when updating the
    kernel portion of top-level page tables.

    Until now, the kernel has relied on each architecture to handle
    synchronization of top-level page tables in an ad-hoc manner.  For
    example, see commit 9b861528a8 ("x86-64, mem: Update all PGDs for direct
    mapping and vmemmap mapping changes").

    However, this approach has proven fragile for following reasons:

      1) It is easy to forget to perform the necessary page table
         synchronization when introducing new changes.
         For instance, commit 4917f55b4e ("mm/sparse-vmemmap: improve memory
         savings for compound devmaps") overlooked the need to synchronize
         page tables for the vmemmap area.

      2) It is also easy to overlook that the vmemmap and direct mapping areas
         must not be accessed before explicit page table synchronization.
         For example, commit 8d400913c2 ("x86/vmemmap: handle unpopulated
         sub-pmd ranges")) caused crashes by accessing the vmemmap area
         before calling sync_global_pgds().

    To address this, as suggested by Dave Hansen, introduce _kernel() variants
    of the page table population helpers, which invoke architecture-specific
    hooks to properly synchronize page tables.  These are introduced in a new
    header file, include/linux/pgalloc.h, so they can be called from common
    code.

    They reuse existing infrastructure for vmalloc and ioremap.
    Synchronization requirements are determined by ARCH_PAGE_TABLE_SYNC_MASK,
    and the actual synchronization is performed by
    arch_sync_kernel_mappings().

    This change currently targets only x86_64, so only PGD and P4D level
    helpers are introduced.  Currently, these helpers are no-ops since no
    architecture sets PGTBL_{PGD,P4D}_MODIFIED in ARCH_PAGE_TABLE_SYNC_MASK.

    In theory, PUD and PMD level helpers can be added later if needed by other
    architectures.  For now, 32-bit architectures (x86-32 and arm) only handle
    PGTBL_PMD_MODIFIED, so p*d_populate_kernel() will never affect them unless
    we introduce a PMD level helper.

    [harry.yoo@oracle.com: fix KASAN build error due to p*d_populate_kernel()]
      Link: https://lkml.kernel.org/r/20250822020727.202749-1-harry.yoo@oracle.com
    Link: https://lkml.kernel.org/r/20250818020206.4517-3-harry.yoo@oracle.com
    Fixes: 8d400913c2 ("x86/vmemmap: handle unpopulated sub-pmd ranges")
    Signed-off-by: Harry Yoo <harry.yoo@oracle.com>
    Suggested-by: Dave Hansen <dave.hansen@linux.intel.com>
    Acked-by: Kiryl Shutsemau <kas@kernel.org>
    Reviewed-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Alexander Potapenko <glider@google.com>
    Cc: Alistair Popple <apopple@nvidia.com>
    Cc: Andrey Konovalov <andreyknvl@gmail.com>
    Cc: Andrey Ryabinin <ryabinin.a.a@gmail.com>
    Cc: Andy Lutomirski <luto@kernel.org>
    Cc: "Aneesh Kumar K.V" <aneesh.kumar@linux.ibm.com>
    Cc: Anshuman Khandual <anshuman.khandual@arm.com>
    Cc: Ard Biesheuvel <ardb@kernel.org>
    Cc: Arnd Bergmann <arnd@arndb.de>
    Cc: bibo mao <maobibo@loongson.cn>
    Cc: Borislav Betkov <bp@alien8.de>
    Cc: Christoph Lameter (Ampere) <cl@gentwo.org>
    Cc: Dennis Zhou <dennis@kernel.org>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Dmitriy Vyukov <dvyukov@google.com>
    Cc: Gwan-gyeong Mun <gwan-gyeong.mun@intel.com>
    Cc: Ingo Molnar <mingo@redhat.com>
    Cc: Jane Chu <jane.chu@oracle.com>
    Cc: Joao Martins <joao.m.martins@oracle.com>
    Cc: Joerg Roedel <joro@8bytes.org>
    Cc: John Hubbard <jhubbard@nvidia.com>
    Cc: Kevin Brodsky <kevin.brodsky@arm.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Oscar Salvador <osalvador@suse.de>
    Cc: Peter Xu <peterx@redhat.com>
    Cc: Peter Zijlstra <peterz@infradead.org>
    Cc: Qi Zheng <zhengqi.arch@bytedance.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: Tejun Heo <tj@kernel.org>
    Cc: Thomas Gleinxer <tglx@linutronix.de>
    Cc: Thomas Huth <thuth@redhat.com>
    Cc: "Uladzislau Rezki (Sony)" <urezki@gmail.com>
    Cc: Vincenzo Frascino <vincenzo.frascino@arm.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:02 -04:00
Rafael Aquini e0c13df4cd mm: move page table sync declarations to linux/pgtable.h
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-39844

commit 7cc183f2e67d19b03ee5c13a6664b8c6cc37ff9d
Author: Harry Yoo <harry.yoo@oracle.com>
Date:   Mon Aug 18 11:02:04 2025 +0900

    mm: move page table sync declarations to linux/pgtable.h

    During our internal testing, we started observing intermittent boot
    failures when the machine uses 4-level paging and has a large amount of
    persistent memory:

      BUG: unable to handle page fault for address: ffffe70000000034
      #PF: supervisor write access in kernel mode
      #PF: error_code(0x0002) - not-present page
      PGD 0 P4D 0
      Oops: 0002 [#1] SMP NOPTI
      RIP: 0010:__init_single_page+0x9/0x6d
      Call Trace:
       <TASK>
       __init_zone_device_page+0x17/0x5d
       memmap_init_zone_device+0x154/0x1bb
       pagemap_range+0x2e0/0x40f
       memremap_pages+0x10b/0x2f0
       devm_memremap_pages+0x1e/0x60
       dev_dax_probe+0xce/0x2ec [device_dax]
       dax_bus_probe+0x6d/0xc9
       [... snip ...]
       </TASK>

    It turns out that the kernel panics while initializing vmemmap (struct
    page array) when the vmemmap region spans two PGD entries, because the new
    PGD entry is only installed in init_mm.pgd, but not in the page tables of
    other tasks.

    And looking at __populate_section_memmap():
      if (vmemmap_can_optimize(altmap, pgmap))
              // does not sync top level page tables
              r = vmemmap_populate_compound_pages(pfn, start, end, nid, pgmap);
      else
              // sync top level page tables in x86
              r = vmemmap_populate(start, end, nid, altmap);

    In the normal path, vmemmap_populate() in arch/x86/mm/init_64.c
    synchronizes the top level page table (See commit 9b861528a8 ("x86-64,
    mem: Update all PGDs for direct mapping and vmemmap mapping changes")) so
    that all tasks in the system can see the new vmemmap area.

    However, when vmemmap_can_optimize() returns true, the optimized path
    skips synchronization of top-level page tables.  This is because
    vmemmap_populate_compound_pages() is implemented in core MM code, which
    does not handle synchronization of the top-level page tables.  Instead,
    the core MM has historically relied on each architecture to perform this
    synchronization manually.

    We're not the first party to encounter a crash caused by not-sync'd top
    level page tables: earlier this year, Gwan-gyeong Mun attempted to address
    the issue [1] [2] after hitting a kernel panic when x86 code accessed the
    vmemmap area before the corresponding top-level entries were synced.  At
    that time, the issue was believed to be triggered only when struct page
    was enlarged for debugging purposes, and the patch did not get further
    updates.

    It turns out that current approach of relying on each arch to handle the
    page table sync manually is fragile because 1) it's easy to forget to sync
    the top level page table, and 2) it's also easy to overlook that the
    kernel should not access the vmemmap and direct mapping areas before the
    sync.

    # The solution: Make page table sync more code robust and harder to miss

    To address this, Dave Hansen suggested [3] [4] introducing
    {pgd,p4d}_populate_kernel() for updating kernel portion of the page tables
    and allow each architecture to explicitly perform synchronization when
    installing top-level entries.  With this approach, we no longer need to
    worry about missing the sync step, reducing the risk of future
    regressions.

    The new interface reuses existing ARCH_PAGE_TABLE_SYNC_MASK,
    PGTBL_P*D_MODIFIED and arch_sync_kernel_mappings() facility used by
    vmalloc and ioremap to synchronize page tables.

    pgd_populate_kernel() looks like this:
    static inline void pgd_populate_kernel(unsigned long addr, pgd_t *pgd,
                                           p4d_t *p4d)
    {
            pgd_populate(&init_mm, pgd, p4d);
            if (ARCH_PAGE_TABLE_SYNC_MASK & PGTBL_PGD_MODIFIED)
                    arch_sync_kernel_mappings(addr, addr);
    }

    It is worth noting that vmalloc() and apply_to_range() carefully
    synchronizes page tables by calling p*d_alloc_track() and
    arch_sync_kernel_mappings(), and thus they are not affected by this patch
    series.

    This series was hugely inspired by Dave Hansen's suggestion and hence
    added Suggested-by: Dave Hansen.

    Cc stable because lack of this series opens the door to intermittent
    boot failures.

    This patch (of 3):

    Move ARCH_PAGE_TABLE_SYNC_MASK and arch_sync_kernel_mappings() to
    linux/pgtable.h so that they can be used outside of vmalloc and ioremap.

    Link: https://lkml.kernel.org/r/20250818020206.4517-1-harry.yoo@oracle.com
    Link: https://lkml.kernel.org/r/20250818020206.4517-2-harry.yoo@oracle.com
    Link: https://lore.kernel.org/linux-mm/20250220064105.808339-1-gwan-gyeong.mun@intel.com [1]
    Link: https://lore.kernel.org/linux-mm/20250311114420.240341-1-gwan-gyeong.mun@intel.com [2]
    Link: https://lore.kernel.org/linux-mm/d1da214c-53d3-45ac-a8b6-51821c5416e4@intel.com [3]
    Link: https://lore.kernel.org/linux-mm/4d800744-7b88-41aa-9979-b245e8bf794b@intel.com  [4]
    Fixes: 8d400913c2 ("x86/vmemmap: handle unpopulated sub-pmd ranges")
    Signed-off-by: Harry Yoo <harry.yoo@oracle.com>
    Acked-by: Kiryl Shutsemau <kas@kernel.org>
    Reviewed-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Reviewed-by: "Uladzislau Rezki (Sony)" <urezki@gmail.com>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Alexander Potapenko <glider@google.com>
    Cc: Alistair Popple <apopple@nvidia.com>
    Cc: Andrey Konovalov <andreyknvl@gmail.com>
    Cc: Andrey Ryabinin <ryabinin.a.a@gmail.com>
    Cc: Andy Lutomirski <luto@kernel.org>
    Cc: "Aneesh Kumar K.V" <aneesh.kumar@linux.ibm.com>
    Cc: Anshuman Khandual <anshuman.khandual@arm.com>
    Cc: Ard Biesheuvel <ardb@kernel.org>
    Cc: Arnd Bergmann <arnd@arndb.de>
    Cc: bibo mao <maobibo@loongson.cn>
    Cc: Borislav Betkov <bp@alien8.de>
    Cc: Christoph Lameter (Ampere) <cl@gentwo.org>
    Cc: Dennis Zhou <dennis@kernel.org>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Dmitriy Vyukov <dvyukov@google.com>
    Cc: Gwan-gyeong Mun <gwan-gyeong.mun@intel.com>
    Cc: Ingo Molnar <mingo@redhat.com>
    Cc: Jane Chu <jane.chu@oracle.com>
    Cc: Joao Martins <joao.m.martins@oracle.com>
    Cc: Joerg Roedel <joro@8bytes.org>
    Cc: John Hubbard <jhubbard@nvidia.com>
    Cc: Kevin Brodsky <kevin.brodsky@arm.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Oscar Salvador <osalvador@suse.de>
    Cc: Peter Xu <peterx@redhat.com>
    Cc: Peter Zijlstra <peterz@infradead.org>
    Cc: Qi Zheng <zhengqi.arch@bytedance.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: Tejun Heo <tj@kernel.org>
    Cc: Thomas Gleinxer <tglx@linutronix.de>
    Cc: Thomas Huth <thuth@redhat.com>
    Cc: Vincenzo Frascino <vincenzo.frascino@arm.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Dave Hansen <dave.hansen@linux.intel.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:02 -04:00
Rafael Aquini 862db7b222 mm: fix accounting of memmap pages
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit c3576889d87b603cb66b417e08844a53c1077a37
Author: Sumanth Korikkar <sumanthk@linux.ibm.com>
Date:   Thu Aug 7 20:35:45 2025 +0200

    mm: fix accounting of memmap pages

    For !CONFIG_SPARSEMEM_VMEMMAP, memmap page accounting is currently done
    upfront in sparse_buffer_init().  However, sparse_buffer_alloc() may
    return NULL in failure scenario.

    Also, memmap pages may be allocated either from the memblock allocator
    during early boot or from the buddy allocator.  When removed via
    arch_remove_memory(), accounting of memmap pages must reflect the original
    allocation source.

    To ensure correctness:
    * Account memmap pages after successful allocation in sparse_init_nid()
      and section_activate().
    * Account memmap pages in section_deactivate() based on allocation
      source.

    Link: https://lkml.kernel.org/r/20250807183545.1424509-1-sumanthk@linux.ibm.com
    Fixes: 15995a3524 ("mm: report per-page metadata information")
    Signed-off-by: Sumanth Korikkar <sumanthk@linux.ibm.com>
    Suggested-by: David Hildenbrand <david@redhat.com>
    Reviewed-by: Wei Yang <richard.weiyang@gmail.com>
    Cc: Alexander Gordeev <agordeev@linux.ibm.com>
    Cc: Gerald Schaefer <gerald.schaefer@linux.ibm.com>
    Cc: Heiko Carstens <hca@linux.ibm.com>
    Cc: Vasily Gorbik <gor@linux.ibm.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:02 -04:00
Rafael Aquini 30be7f6a03 mm/damon/core: prevent unnecessary overflow in damos_set_effective_quota()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 9f68eabab9d9aaa764a8d234c4170119e6518102
Author: Quanmin Yan <yanquanmin1@huawei.com>
Date:   Thu Aug 21 20:55:55 2025 +0800

    mm/damon/core: prevent unnecessary overflow in damos_set_effective_quota()

    On 32-bit systems, the throughput calculation in
    damos_set_effective_quota() is prone to unnecessary multiplication
    overflow.  Using mult_frac() to fix it.

    Andrew Paniakin also recently found and privately reported this issue, on
    64 bit systems.  This can also happen on 64-bit systems, once the charged
    size exceeds ~17 TiB.  On systems running for long time in production,
    this issue can actually happen.

    More specifically, when a DAMOS scheme having the time quota run for
    longtime, throughput calculation can overflow and set esz too small.  As a
    result, speed of the scheme get unexpectedly slow.

    Link: https://lkml.kernel.org/r/20250821125555.3020951-1-yanquanmin1@huawei.com
    Fixes: 1cd2430300 ("mm/damon/schemes: implement time quota")
    Signed-off-by: Quanmin Yan <yanquanmin1@huawei.com>
    Reported-by: Andrew Paniakin <apanyaki@amazon.com>
    Reviewed-by: SeongJae Park <sj@kernel.org>
    Cc: Kefeng Wang <wangkefeng.wang@huawei.com>
    Cc: ze zuo <zuoze1@huawei.com>
    Cc: <stable@vger.kernel.org>    [5.16+]
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:02 -04:00
Rafael Aquini 4da8d23b0d mm/kasan: avoid lazy MMU mode hazards
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit c519c3c0a1133c408e83a383aa4dd30010aa5d71
Author: Alexander Gordeev <agordeev@linux.ibm.com>
Date:   Mon Aug 18 18:39:13 2025 +0200

    mm/kasan: avoid lazy MMU mode hazards

    Functions __kasan_populate_vmalloc() and __kasan_depopulate_vmalloc() use
    apply_to_pte_range(), which enters lazy MMU mode.  In that mode updating
    PTEs may not be observed until the mode is left.

    That may lead to a situation in which otherwise correct reads and writes
    to a PTE using ptep_get(), set_pte(), pte_clear() and other access
    primitives bring wrong results when the vmalloc shadow memory is being
    (de-)populated.

    To avoid these hazards leave the lazy MMU mode before and re-enter it
    after each PTE manipulation.

    Link: https://lkml.kernel.org/r/0d2efb7ddddbff6b288fbffeeb10166e90771718.1755528662.git.agordeev@linux.ibm.com
    Fixes: 3c5c3cfb9e ("kasan: support backing vmalloc space with real shadow memory")
    Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
    Cc: Andrey Ryabinin <ryabinin.a.a@gmail.com>
    Cc: Daniel Axtens <dja@axtens.net>
    Cc: Marc Rutland <mark.rutland@arm.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:01 -04:00
Rafael Aquini fefae2534b mm/kasan: fix vmalloc shadow memory (de-)population races
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 08c7c253e032863199da4f089bd0ccab5d1a4876
Author: Alexander Gordeev <agordeev@linux.ibm.com>
Date:   Mon Aug 18 18:39:12 2025 +0200

    mm/kasan: fix vmalloc shadow memory (de-)population races

    While working on the lazy MMU mode enablement for s390 I hit pretty
    curious issues in the kasan code.

    The first is related to a custom kasan-based sanitizer aimed at catching
    invalid accesses to PTEs and is inspired by [1] conversation.  The kasan
    complains on valid PTE accesses, while the shadow memory is reported as
    unpoisoned:

    [  102.783993] ==================================================================
    [  102.784008] BUG: KASAN: out-of-bounds in set_pte_range+0x36c/0x390
    [  102.784016] Read of size 8 at addr 0000780084cf9608 by task vmalloc_test/0/5542
    [  102.784019]
    [  102.784040] CPU: 1 UID: 0 PID: 5542 Comm: vmalloc_test/0 Kdump: loaded Tainted: G           OE       6.16.0-gcc-ipte-kasan-11657-gb2d930c4950e #340 PREEMPT
    [  102.784047] Tainted: [O]=OOT_MODULE, [E]=UNSIGNED_MODULE
    [  102.784049] Hardware name: IBM 8561 T01 703 (LPAR)
    [  102.784052] Call Trace:
    [  102.784054]  [<00007fffe0147ac0>] dump_stack_lvl+0xe8/0x140
    [  102.784059]  [<00007fffe0112484>] print_address_description.constprop.0+0x34/0x2d0
    [  102.784066]  [<00007fffe011282c>] print_report+0x10c/0x1f8
    [  102.784071]  [<00007fffe090785a>] kasan_report+0xfa/0x220
    [  102.784078]  [<00007fffe01d3dec>] set_pte_range+0x36c/0x390
    [  102.784083]  [<00007fffe01d41c2>] leave_ipte_batch+0x3b2/0xb10
    [  102.784088]  [<00007fffe07d3650>] apply_to_pte_range+0x2f0/0x4e0
    [  102.784094]  [<00007fffe07e62e4>] apply_to_pmd_range+0x194/0x3e0
    [  102.784099]  [<00007fffe07e820e>] __apply_to_page_range+0x2fe/0x7a0
    [  102.784104]  [<00007fffe07e86d8>] apply_to_page_range+0x28/0x40
    [  102.784109]  [<00007fffe090a3ec>] __kasan_populate_vmalloc+0xec/0x310
    [  102.784114]  [<00007fffe090aa36>] kasan_populate_vmalloc+0x96/0x130
    [  102.784118]  [<00007fffe0833a04>] alloc_vmap_area+0x3d4/0xf30
    [  102.784123]  [<00007fffe083a8ba>] __get_vm_area_node+0x1aa/0x4c0
    [  102.784127]  [<00007fffe083c4f6>] __vmalloc_node_range_noprof+0x126/0x4e0
    [  102.784131]  [<00007fffe083c980>] __vmalloc_node_noprof+0xd0/0x110
    [  102.784135]  [<00007fffe083ca32>] vmalloc_noprof+0x32/0x40
    [  102.784139]  [<00007fff608aa336>] fix_size_alloc_test+0x66/0x150 [test_vmalloc]
    [  102.784147]  [<00007fff608aa710>] test_func+0x2f0/0x430 [test_vmalloc]
    [  102.784153]  [<00007fffe02841f8>] kthread+0x3f8/0x7a0
    [  102.784159]  [<00007fffe014d8b4>] __ret_from_fork+0xd4/0x7d0
    [  102.784164]  [<00007fffe299c00a>] ret_from_fork+0xa/0x30
    [  102.784173] no locks held by vmalloc_test/0/5542.
    [  102.784176]
    [  102.784178] The buggy address belongs to the physical page:
    [  102.784186] page: refcount:1 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x84cf9
    [  102.784198] flags: 0x3ffff00000000000(node=0|zone=1|lastcpupid=0x1ffff)
    [  102.784212] page_type: f2(table)
    [  102.784225] raw: 3ffff00000000000 0000000000000000 0000000000000122 0000000000000000
    [  102.784234] raw: 0000000000000000 0000000000000000 f200000000000001 0000000000000000
    [  102.784248] page dumped because: kasan: bad access detected
    [  102.784250]
    [  102.784252] Memory state around the buggy address:
    [  102.784260]  0000780084cf9500: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
    [  102.784274]  0000780084cf9580: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
    [  102.784277] >0000780084cf9600: fd 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
    [  102.784290]                          ^
    [  102.784293]  0000780084cf9680: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
    [  102.784303]  0000780084cf9700: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
    [  102.784306] ==================================================================

    The second issue hits when the custom sanitizer above is not implemented,
    but the kasan itself is still active:

    [ 1554.438028] Unable to handle kernel pointer dereference in virtual kernel address space
    [ 1554.438065] Failing address: 001c0ff0066f0000 TEID: 001c0ff0066f0403
    [ 1554.438076] Fault in home space mode while using kernel ASCE.
    [ 1554.438103] AS:00000000059d400b R2:0000000ffec5c00b R3:00000000c6c9c007 S:0000000314470001 P:00000000d0ab413d
    [ 1554.438158] Oops: 0011 ilc:2 [#1]SMP
    [ 1554.438175] Modules linked in: test_vmalloc(E+) nft_fib_inet(E) nft_fib_ipv4(E) nft_fib_ipv6(E) nft_fib(E) nft_reject_inet(E) nf_reject_ipv4(E) nf_reject_ipv6(E) nft_reject(E) nft_ct(E) nft_chain_nat(E) nf_nat(E) nf_conntrack(E) nf_defrag_ipv6(E) nf_defrag_ipv4(E) nf_tables(E) sunrpc(E) pkey_pckmo(E) uvdevice(E) s390_trng(E) rng_core(E) eadm_sch(E) vfio_ccw(E) mdev(E) vfio_iommu_type1(E) vfio(E) sch_fq_codel(E) drm(E) loop(E) i2c_core(E) drm_panel_orientation_quirks(E) nfnetlink(E) ctcm(E) fsm(E) zfcp(E) scsi_transport_fc(E) diag288_wdt(E) watchdog(E) ghash_s390(E) prng(E) aes_s390(E) des_s390(E) libdes(E) sha3_512_s390(E) sha3_256_s390(E) sha512_s390(E) sha1_s390(E) sha_common(E) pkey(E) autofs4(E)
    [ 1554.438319] Unloaded tainted modules: pkey_uv(E):1 hmac_s390(E):2
    [ 1554.438354] CPU: 1 UID: 0 PID: 1715 Comm: vmalloc_test/0 Kdump: loaded Tainted: G            E       6.16.0-gcc-ipte-kasan-11657-gb2d930c4950e #350 PREEMPT
    [ 1554.438368] Tainted: [E]=UNSIGNED_MODULE
    [ 1554.438374] Hardware name: IBM 8561 T01 703 (LPAR)
    [ 1554.438381] Krnl PSW : 0704e00180000000 00007fffe1d3d6ae (memset+0x5e/0x98)
    [ 1554.438396]            R:0 T:1 IO:1 EX:1 Key:0 M:1 W:0 P:0 AS:3 CC:2 PM:0 RI:0 EA:3
    [ 1554.438409] Krnl GPRS: 0000000000000001 001c0ff0066f0000 001c0ff0066f0000 00000000000000f8
    [ 1554.438418]            00000000000009fe 0000000000000009 0000000000000000 0000000000000002
    [ 1554.438426]            0000000000005000 000078031ae655c8 00000feffdcf9f59 0000780258672a20
    [ 1554.438433]            0000780243153500 00007f8033780000 00007fffe083a510 00007f7fee7cfa00
    [ 1554.438452] Krnl Code: 00007fffe1d3d6a0: eb540008000c        srlg    %r5,%r4,8
               00007fffe1d3d6a6: b9020055           ltgr    %r5,%r5
              #00007fffe1d3d6aa: a784000b           brc     8,00007fffe1d3d6c0
              >00007fffe1d3d6ae: 42301000           stc     %r3,0(%r1)
               00007fffe1d3d6b2: d2fe10011000       mvc     1(255,%r1),0(%r1)
               00007fffe1d3d6b8: 41101100           la      %r1,256(%r1)
               00007fffe1d3d6bc: a757fff9           brctg   %r5,00007fffe1d3d6ae
               00007fffe1d3d6c0: 42301000           stc     %r3,0(%r1)
    [ 1554.438539] Call Trace:
    [ 1554.438545]  [<00007fffe1d3d6ae>] memset+0x5e/0x98
    [ 1554.438552] ([<00007fffe083a510>] remove_vm_area+0x220/0x400)
    [ 1554.438562]  [<00007fffe083a9d6>] vfree.part.0+0x26/0x810
    [ 1554.438569]  [<00007fff6073bd50>] fix_align_alloc_test+0x50/0x90 [test_vmalloc]
    [ 1554.438583]  [<00007fff6073c73a>] test_func+0x46a/0x6c0 [test_vmalloc]
    [ 1554.438593]  [<00007fffe0283ac8>] kthread+0x3f8/0x7a0
    [ 1554.438603]  [<00007fffe014d8b4>] __ret_from_fork+0xd4/0x7d0
    [ 1554.438613]  [<00007fffe299ac0a>] ret_from_fork+0xa/0x30
    [ 1554.438622] INFO: lockdep is turned off.
    [ 1554.438627] Last Breaking-Event-Address:
    [ 1554.438632]  [<00007fffe1d3d65c>] memset+0xc/0x98
    [ 1554.438644] Kernel panic - not syncing: Fatal exception: panic_on_oops

    This series fixes the above issues and is a pre-requisite for the s390
    lazy MMU mode implementation.

    test_vmalloc was used to stress-test the fixes.

    This patch (of 2):

    When vmalloc shadow memory is established the modification of the
    corresponding page tables is not protected by any locks.  Instead, the
    locking is done per-PTE.  This scheme however has defects.

    kasan_populate_vmalloc_pte() - while ptep_get() read is atomic the
    sequence pte_none(ptep_get()) is not.  Doing that outside of the lock
    might lead to a concurrent PTE update and what could be seen as a shadow
    memory corruption as result.

    kasan_depopulate_vmalloc_pte() - by the time a page whose address was
    extracted from ptep_get() read and cached in a local variable outside of
    the lock is attempted to get free, could actually be freed already.

    To avoid these put ptep_get() itself and the code that manipulates the
    result of the read under lock.  In addition, move freeing of the page out
    of the atomic context.

    Link: https://lkml.kernel.org/r/cover.1755528662.git.agordeev@linux.ibm.com
    Link: https://lkml.kernel.org/r/adb258634194593db294c0d1fb35646e894d6ead.1755528662.git.agordeev@linux.ibm.com
    Link: https://lore.kernel.org/linux-mm/5b0609c9-95ee-4e48-bb6d-98f57c5d2c31@arm.com/ [1]
    Fixes: 3c5c3cfb9e ("kasan: support backing vmalloc space with real shadow memory")
    Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
    Cc: Andrey Ryabinin <ryabinin.a.a@gmail.com>
    Cc: Daniel Axtens <dja@axtens.net>
    Cc: Marc Rutland <mark.rutland@arm.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:01 -04:00
Rafael Aquini 3b97ed6a83 kunit: kasan_test: disable fortify string checker on kasan_strings() test
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 7a19afee6fb39df63ddea7ce78976d8c521178c6
Author: Yeoreum Yun <yeoreum.yun@arm.com>
Date:   Fri Aug 1 13:02:36 2025 +0100

    kunit: kasan_test: disable fortify string checker on kasan_strings() test

    Similar to commit 09c6304e38 ("kasan: test: fix compatibility with
    FORTIFY_SOURCE") the kernel is panicing in kasan_string().

    This is due to the `src` and `ptr` not being hidden from the optimizer
    which would disable the runtime fortify string checker.

    Call trace:
      __fortify_panic+0x10/0x20 (P)
      kasan_strings+0x980/0x9b0
      kunit_try_run_case+0x68/0x190
      kunit_generic_run_threadfn_adapter+0x34/0x68
      kthread+0x1c4/0x228
      ret_from_fork+0x10/0x20
     Code: d503233f a9bf7bfd 910003fd 9424b243 (d4210000)
     ---[ end trace 0000000000000000 ]---
     note: kunit_try_catch[128] exited with irqs disabled
     note: kunit_try_catch[128] exited with preempt_count 1
         # kasan_strings: try faulted: last
    ** replaying previous printk message **
         # kasan_strings: try faulted: last line seen mm/kasan/kasan_test_c.c:1600
         # kasan_strings: internal error occurred preventing test case from running: -4

    Link: https://lkml.kernel.org/r/20250801120236.2962642-1-yeoreum.yun@arm.com
    Fixes: 73228c7ecc ("KASAN: port KASAN Tests to KUnit")
    Signed-off-by: Yeoreum Yun <yeoreum.yun@arm.com>
    Cc: Alexander Potapenko <glider@google.com>
    Cc: Andrey Konovalov <andreyknvl@gmail.com>
    Cc: Andrey Ryabinin <ryabinin.a.a@gmail.com>
    Cc: Dmitriy Vyukov <dvyukov@google.com>
    Cc: Vincenzo Frascino <vincenzo.frascino@arm.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:01 -04:00
Rafael Aquini 360ef4cccf selftests/mm: fix FORCE_READ to read input value correctly
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 5bbc2b785e63699cfcaa7adbf739f6e9b771028a
Author: Zi Yan <ziy@nvidia.com>
Date:   Tue Aug 5 13:51:40 2025 -0400

    selftests/mm: fix FORCE_READ to read input value correctly

    FORCE_READ() converts input value x to its pointer type then reads from
    address x.  This is wrong.  If x is a non-pointer, it would be caught it
    easily.  But all FORCE_READ() callers are trying to read from a pointer
    and FORCE_READ() basically reads a pointer to a pointer instead of the
    original typed pointer.  Almost no access violation was found, except the
    one from split_huge_page_test.

    Fix it by implementing a simplified READ_ONCE() instead.

    Link: https://lkml.kernel.org/r/20250805175140.241656-1-ziy@nvidia.com
    Fixes: 3f6bfd4789a0 ("selftests/mm: reuse FORCE_READ to replace "asm volatile("" : "+r" (XXX));"")
    Signed-off-by: Zi Yan <ziy@nvidia.com>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Reviewed-by: wang lian <lianux.mm@gmail.com>
    Reviewed-by: Wei Yang <richard.weiyang@gmail.com>
    Cc: Christian Brauner <brauner@kernel.org>
    Cc: Jann Horn <jannh@google.com>
    Cc: Kairui Song <ryncsn@gmail.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Mark Brown <broonie@kernel.org>
    Cc: SeongJae Park <sj@kernel.org>
    Cc: Shuah Khan <shuah@kernel.org>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Zi Yan <ziy@nvidia.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:01 -04:00
Rafael Aquini 41bd528e82 mm/userfaultfd: fix kmap_local LIFO ordering for CONFIG_HIGHPTE
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-39899

commit 9614d8bee66387501f48718fa306e17f2aa3f2f3
Author: Sasha Levin <sashal@kernel.org>
Date:   Thu Jul 31 10:44:31 2025 -0400

    mm/userfaultfd: fix kmap_local LIFO ordering for CONFIG_HIGHPTE

    With CONFIG_HIGHPTE on 32-bit ARM, move_pages_pte() maps PTE pages using
    kmap_local_page(), which requires unmapping in Last-In-First-Out order.

    The current code maps dst_pte first, then src_pte, but unmaps them in the
    same order (dst_pte, src_pte), violating the LIFO requirement.  This
    causes the warning in kunmap_local_indexed():

      WARNING: CPU: 0 PID: 604 at mm/highmem.c:622 kunmap_local_indexed+0x178/0x17c
      addr \!= __fix_to_virt(FIX_KMAP_BEGIN + idx)

    Fix this by reversing the unmap order to respect LIFO ordering.

    This issue follows the same pattern as similar fixes:
    - commit eca6828403b8 ("crypto: skcipher - fix mismatch between mapping and unmapping order")
    - commit 8cf57c6df8 ("nilfs2: eliminate staggered calls to kunmap in nilfs_rename")

    Both of which addressed the same fundamental requirement that kmap_local
    operations must follow LIFO ordering.

    Link: https://lkml.kernel.org/r/20250731144431.773923-1-sashal@kernel.org
    Fixes: adef440691 ("userfaultfd: UFFDIO_MOVE uABI")
    Signed-off-by: Sasha Levin <sashal@kernel.org>
    Acked-by: David Hildenbrand <david@redhat.com>
    Reviewed-by: Suren Baghdasaryan <surenb@google.com>
    Cc: Andrea Arcangeli <aarcange@redhat.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:01 -04:00
Rafael Aquini dc23cf8bf3 s390/mm: Prevent possible preempt_count overflow
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 57834ce5a6a47df282c8419019ba5495eac58fb9
Author: Gerald Schaefer <gerald.schaefer@linux.ibm.com>
Date:   Thu Aug 21 19:00:03 2025 +0200

    s390/mm: Prevent possible preempt_count overflow

    The s390 implementation of ptep_modify_prot_start() currently does
    preempt_disable(), and the preempt_enable() is done later in
    ptep_modify_prot_commit(). This logic is not really required, because the
    PTE lock must be held over the complete prot_start/commit transaction,
    as described in the comment of the generic implementation of
    ptep_modify_prot_start().

    That comment also mentions that this interface should be batchable,
    and modify_prot_start_ptes() might start a transaction over a batch of
    PTEs, implemented as a simple loop over ptep_modify_prot_start().
    In this case, the preempt_disable() in ptep_modify_prot_start() would
    be called multiple times, before the corresponding preempt_enable()
    calls happen, and this can lead to a preempt_count overflow.

    To fix this, simply remove the preempt_disable/enable() calls in
    ptep_modify_prot_start/commit(), and rely on the PTE lock being held.

    Commit cac1db8c3aad ("mm: optimize mprotect() by PTE batching") made use
    of this PTE batching for the first time, and triggers warnings like this:

     DEBUG_LOCKS_WARN_ON((preempt_count() & PREEMPT_MASK) >= PREEMPT_MASK - 10)
     BUG: sleeping function called from invalid context at mm/mprotect.c:576

    Hence, add a Fixes tag on that commit. Not because it is broken, but to
    make sure that it won't get backported w/o also this fix for s390.

    Fixes: cac1db8c3aad ("mm: optimize mprotect() by PTE batching")
    Reviewed-by: Alexander Gordeev <agordeev@linux.ibm.com>
    Signed-off-by: Gerald Schaefer <gerald.schaefer@linux.ibm.com>
    Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:01 -04:00
Rafael Aquini 0368942c90 memblock: fix kernel-doc for MEMBLOCK_RSRV_NOINIT
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit b3dcc9d1d806fb1e175f85978713eef868531da4
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Tue Aug 26 10:19:46 2025 +0300

    memblock: fix kernel-doc for MEMBLOCK_RSRV_NOINIT

    The kernel-doc description of MEMBLOCK_RSRV_NOINIT and
    memblock_reserved_mark_noinit() do not accurately describe their
    functionality.

    Expand their kernel doc to make it clear that the user of
    MEMBLOCK_RSRV_NOINIT is responsible to properly initialize the struct pages
    for such regions and add more details about effects of using this flag.

    Reviewed-by: David Hildenbrand <david@redhat.com>
    Link: https://lore.kernel.org/r/f8140a17-c4ec-489b-b314-d45abe48bf36@redhat.com
    Link: https://lore.kernel.org/r/20250826071947.1949725-1-rppt@kernel.org
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:00 -04:00
Rafael Aquini 038813df4b mm/slub: avoid accessing metadata when pointer is invalid in object_err()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-39902

commit b4efccec8d06ceb10a7d34d7b1c449c569d53770
Author: Li Qiong <liqiong@nfschina.com>
Date:   Mon Aug 4 10:57:59 2025 +0800

    mm/slub: avoid accessing metadata when pointer is invalid in object_err()

    object_err() reports details of an object for further debugging, such as
    the freelist pointer, redzone, etc. However, if the pointer is invalid,
    attempting to access object metadata can lead to a crash since it does
    not point to a valid object.

    One known path to the crash is when alloc_consistency_checks()
    determines the pointer to the allocated object is invalid because of a
    freelist corruption, and calls object_err() to report it. The debug code
    should report and handle the corruption gracefully and not crash in the
    process.

    In case the pointer is NULL or check_valid_pointer() returns false for
    the pointer, only print the pointer value and skip accessing metadata.

    Fixes: 81819f0fc8 ("SLUB core")
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Li Qiong <liqiong@nfschina.com>
    Reviewed-by: Harry Yoo <harry.yoo@oracle.com>
    Reviewed-by: Matthew Wilcox (Oracle) <willy@infradead.org>
    Signed-off-by: Vlastimil Babka <vbabka@suse.cz>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:00 -04:00
Rafael Aquini a6bafe67d2 mm: numa,memblock: Use SZ_1M macro to denote bytes to MB conversion
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 4647c4deadcc17f40858be06bcf416369a8f1d57
Author: Pratyush Brahma <pratyush.brahma@oss.qualcomm.com>
Date:   Wed Aug 20 06:29:34 2025 +0530

    mm: numa,memblock: Use SZ_1M macro to denote bytes to MB conversion

    Replace the manual bitwise conversion of bytes to MB with
    SZ_1M macro, a standard macro used within the mm subsystem,
    to improve readability.

    Signed-off-by: Pratyush Brahma <pratyush.brahma@oss.qualcomm.com>
    Link: https://lore.kernel.org/r/20250820-numa-memblks-refac-v2-1-43bf1af02acd@oss.qualcomm.com
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:00 -04:00
Rafael Aquini db93b7c589 mm/mremap: fix WARN with uffd that has remap events disabled
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-39775

commit 772e5b4a5e8360743645b9a466842d16092c4f94
Author: David Hildenbrand <david@redhat.com>
Date:   Mon Aug 18 19:53:58 2025 +0200

    mm/mremap: fix WARN with uffd that has remap events disabled

    Registering userfaultd on a VMA that spans at least one PMD and then
    mremap()'ing that VMA can trigger a WARN when recovering from a failed
    page table move due to a page table allocation error.

    The code ends up doing the right thing (recurse, avoiding moving actual
    page tables), but triggering that WARN is unpleasant:

    WARNING: CPU: 2 PID: 6133 at mm/mremap.c:357 move_normal_pmd mm/mremap.c:357 [inline]
    WARNING: CPU: 2 PID: 6133 at mm/mremap.c:357 move_pgt_entry mm/mremap.c:595 [inline]
    WARNING: CPU: 2 PID: 6133 at mm/mremap.c:357 move_page_tables+0x3832/0x44a0 mm/mremap.c:852
    Modules linked in:
    CPU: 2 UID: 0 PID: 6133 Comm: syz.0.19 Not tainted 6.17.0-rc1-syzkaller-00004-g53e760d89498 #0 PREEMPT(full)
    Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2~bpo12+1 04/01/2014
    RIP: 0010:move_normal_pmd mm/mremap.c:357 [inline]
    RIP: 0010:move_pgt_entry mm/mremap.c:595 [inline]
    RIP: 0010:move_page_tables+0x3832/0x44a0 mm/mremap.c:852
    Code: ...
    RSP: 0018:ffffc900037a76d8 EFLAGS: 00010293
    RAX: 0000000000000000 RBX: 0000000032930007 RCX: ffffffff820c6645
    RDX: ffff88802e56a440 RSI: ffffffff820c7201 RDI: 0000000000000007
    RBP: ffff888037728fc0 R08: 0000000000000007 R09: 0000000000000000
    R10: 0000000032930007 R11: 0000000000000000 R12: 0000000000000000
    R13: ffffc900037a79a8 R14: 0000000000000001 R15: dffffc0000000000
    FS:  000055556316a500(0000) GS:ffff8880d68bc000(0000) knlGS:0000000000000000
    CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
    CR2: 0000001b30863fff CR3: 0000000050171000 CR4: 0000000000352ef0
    Call Trace:
     <TASK>
     copy_vma_and_data+0x468/0x790 mm/mremap.c:1215
     move_vma+0x548/0x1780 mm/mremap.c:1282
     mremap_to+0x1b7/0x450 mm/mremap.c:1406
     do_mremap+0xfad/0x1f80 mm/mremap.c:1921
     __do_sys_mremap+0x119/0x170 mm/mremap.c:1977
     do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
     do_syscall_64+0xcd/0x4c0 arch/x86/entry/syscall_64.c:94
     entry_SYSCALL_64_after_hwframe+0x77/0x7f
    RIP: 0033:0x7f00d0b8ebe9
    Code: ...
    RSP: 002b:00007ffe5ea5ee98 EFLAGS: 00000246 ORIG_RAX: 0000000000000019
    RAX: ffffffffffffffda RBX: 00007f00d0db5fa0 RCX: 00007f00d0b8ebe9
    RDX: 0000000000400000 RSI: 0000000000c00000 RDI: 0000200000000000
    RBP: 00007ffe5ea5eef0 R08: 0000200000c00000 R09: 0000000000000000
    R10: 0000000000000003 R11: 0000000000000246 R12: 0000000000000002
    R13: 00007f00d0db5fa0 R14: 00007f00d0db5fa0 R15: 0000000000000005
     </TASK>

    The underlying issue is that we recurse during the original page table
    move, but not during the recovery move.

    Fix it by checking for both VMAs and performing the check before the
    pmd_none() sanity check.

    Add a new helper where we perform+document that check for the PMD and PUD
    level.

    Thanks to Harry for bisecting.

    Link: https://lkml.kernel.org/r/20250818175358.1184757-1-david@redhat.com
    Fixes: 0cef0bb836e3 ("mm: clear uffd-wp PTE/PMD state on mremap()")
    Signed-off-by: David Hildenbrand <david@redhat.com>
    Reported-by: syzbot+4d9a13f0797c46a29e42@syzkaller.appspotmail.com
    Closes: https://lkml.kernel.org/r/689bb893.050a0220.7f033.013a.GAE@google.com
    Tested-by: Harry Yoo <harry.yoo@oracle.com>
    Cc: "Liam R. Howlett" <Liam.Howlett@oracle.com>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Jann Horn <jannh@google.com>
    Cc: Pedro Falcato <pfalcato@suse.de>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:00 -04:00
Rafael Aquini a412d45277 mm/damon/sysfs-schemes: put damos dests dir after removing its files
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit ba1dd7ac735d604249f1e614d997dc66b30844ab
Author: SeongJae Park <sj@kernel.org>
Date:   Sat Aug 16 09:55:59 2025 -0700

    mm/damon/sysfs-schemes: put damos dests dir after removing its files

    damon_sysfs_scheme_rm_dirs() puts dests directory kobject before removing
    its internal files.  Sincee putting the kobject frees its container
    struct, and the internal files removal accesses the container,
    use-after-free happens.  Fix it by putting the reference _after_ removing
    the files.

    Link: https://lkml.kernel.org/r/20250816165559.2601-1-sj@kernel.org
    Fixes: 2cd0bf85a203 ("mm/damon/sysfs-schemes: implement DAMOS action destinations directory")
    Signed-off-by: SeongJae Park <sj@kernel.org>
    Reported-by: Alexandre Ghiti <alex@ghiti.fr>
    Closes: https://lore.kernel.org/2d39a734-320d-4341-8f8a-4019eec2dbf2@ghiti.fr
    Tested-by: Alexandre Ghiti <alexghiti@rivosinc.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:00 -04:00
Rafael Aquini 7bbceb62ec mm/migrate: fix NULL movable_ops if CONFIG_ZSMALLOC=m
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 053c8ebe74f7e1f4c072e59428da80b9d78bc4b7
Author: Huacai Chen <chenhuacai@kernel.org>
Date:   Sun Aug 17 23:17:59 2025 +0800

    mm/migrate: fix NULL movable_ops if CONFIG_ZSMALLOC=m

    After commit 84caf98838a3e5f4bdb34 ("mm: stop storing migration_ops in
    page->mapping") we get such an error message if CONFIG_ZSMALLOC=m:

     WARNING: CPU: 3 PID: 42 at mm/migrate.c:142 isolate_movable_ops_page+0xa8/0x1c0
     CPU: 3 UID: 0 PID: 42 Comm: kcompactd0 Not tainted 6.16.0-rc5+ #2133 PREEMPT
     pc 9000000000540bd8 ra 9000000000540b84 tp 9000000100420000 sp 9000000100423a60
     a0 9000000100193a80 a1 000000000000000c a2 000000000000001b a3 ffffffffffffffff
     a4 ffffffffffffffff a5 0000000000000267 a6 0000000000000000 a7 9000000100423ae0
     t0 00000000000000f1 t1 00000000000000f6 t2 0000000000000000 t3 0000000000000001
     t4 ffffff00010eb834 t5 0000000000000040 t6 900000010c89d380 t7 90000000023fcc70
     t8 0000000000000018 u0 0000000000000000 s9 ffffff00010eb800 s0 ffffff00010eb800
     s1 000000000000000c s2 0000000000043ae0 s3 0000800000000000 s4 900000000219cc40
     s5 0000000000000000 s6 ffffff00010eb800 s7 0000000000000001 s8 90000000025b4000
        ra: 9000000000540b84 isolate_movable_ops_page+0x54/0x1c0
       ERA: 9000000000540bd8 isolate_movable_ops_page+0xa8/0x1c0
      CRMD: 000000b0 (PLV0 -IE -DA +PG DACF=CC DACM=CC -WE)
      PRMD: 00000004 (PPLV0 +PIE -PWE)
      EUEN: 00000000 (-FPE -SXE -ASXE -BTE)
      ECFG: 00071c1d (LIE=0,2-4,10-12 VS=7)
     ESTAT: 000c0000 [BRK] (IS= ECode=12 EsubCode=0)
      PRID: 0014c010 (Loongson-64bit, Loongson-3A5000)
     CPU: 3 UID: 0 PID: 42 Comm: kcompactd0 Not tainted 6.16.0-rc5+ #2133 PREEMPT
     Stack : 90000000021fd000 0000000000000000 9000000000247720 9000000100420000
             90000001004236a0 90000001004236a8 0000000000000000 90000001004237e8
             90000001004237e0 90000001004237e0 9000000100423550 0000000000000001
             0000000000000001 90000001004236a8 725a84864a19e2d9 90000000023fcc58
             9000000100420000 90000000024c6848 9000000002416848 0000000000000001
             0000000000000000 000000000000000a 0000000007fe0000 ffffff00010eb800
             0000000000000000 90000000021fd000 0000000000000000 900000000205cf30
             000000000000008e 0000000000000009 ffffff00010eb800 0000000000000001
             90000000025b4000 0000000000000000 900000000024773c 00007ffff103d748
             00000000000000b0 0000000000000004 0000000000000000 0000000000071c1d
             ...
     Call Trace:
     [<900000000024773c>] show_stack+0x5c/0x190
     [<90000000002415e0>] dump_stack_lvl+0x70/0x9c
     [<90000000004abe6c>] isolate_migratepages_block+0x3bc/0x16e0
     [<90000000004af408>] compact_zone+0x558/0x1000
     [<90000000004b0068>] compact_node+0xa8/0x1e0
     [<90000000004b0aa4>] kcompactd+0x394/0x410
     [<90000000002b3c98>] kthread+0x128/0x140
     [<9000000001779148>] ret_from_kernel_thread+0x28/0xc0
     [<9000000000245528>] ret_from_kernel_thread_asm+0x10/0x88

    The reason is that defined(CONFIG_ZSMALLOC) evaluates to 1 only when
    CONFIG_ZSMALLOC=y, we should use IS_ENABLED(CONFIG_ZSMALLOC) instead.  But
    when I use IS_ENABLED(CONFIG_ZSMALLOC), page_movable_ops() cannot access
    zsmalloc_mops because zsmalloc_mops is in a module.

    To solve this problem, we define a set_movable_ops() interface to register
    and unregister offline_movable_ops / zsmalloc_movable_ops in mm/migrate.c,
    and call them at mm/balloon_compaction.c & mm/zsmalloc.c.  Since
    offline_movable_ops / zsmalloc_movable_ops are always accessible, all
    #ifdef / #endif are removed in page_movable_ops().

    Link: https://lkml.kernel.org/r/20250817151759.2525174-1-chenhuacai@loongson.cn
    Fixes: 84caf98838a3 ("mm: stop storing migration_ops in page->mapping")
    Signed-off-by: Huacai Chen <chenhuacai@loongson.cn>
    Acked-by: Zi Yan <ziy@nvidia.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Huacai Chen <chenhuacai@kernel.org>
    Cc: Huacai Chen <chenhuacai@loongson.cn>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: "Michael S. Tsirkin" <mst@redhat.com>
    Cc: Minchan Kim <minchan@kernel.org>
    Cc: Sergey Senozhatsky <senozhatsky@chromium.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:00 -04:00
Rafael Aquini d843f275de selftests/mm: add test for invalid multi VMA operations
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 742d3663a5775cb7b957f4ca2ddb4ccd26badb94
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Sun Aug 3 12:11:23 2025 +0100

    selftests/mm: add test for invalid multi VMA operations

    We can use UFFD to easily assert invalid multi VMA moves, so do so,
    asserting expected behaviour when VMAs invalid for a multi VMA operation
    are encountered.

    We assert both that such operations are not permitted, and that we do not
    even attempt to move the first VMA under these circumstances.

    We also assert that we can still move a single VMA regardless.

    We then assert that a partial failure can occur if the invalid VMA appears
    later in the range of multiple VMAs, both at the very next VMA, and also at
    the end of the range.

    As part of this change, we are using the is_range_valid() helper more
    aggressively. Therefore, fix a bug where stale buffered data would hang
    around on success, causing subsequent calls to is_range_valid() to
    potentially give invalid results.

    We simply have to fflush() the stream on success to resolve this issue.

    Link: https://lkml.kernel.org/r/c4fb86dd5ba37610583ad5fc0e0c2306ddf318b9.1754218667.git.lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: David Hildenbrand <david@redhat.com>
    Cc: Jann Horn <jannh@google.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Mike Rapoport <rppt@kernel.org>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:51:00 -04:00
Rafael Aquini a2e2e65eee mm/mremap: catch invalid multi VMA moves earlier
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit d5f416c7c36456676c2cf5ab98776db2e7601f27
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Sun Aug 3 12:11:22 2025 +0100

    mm/mremap: catch invalid multi VMA moves earlier

    Previously, any attempt to solely move a VMA would require that the
    span specified reside within the span of that single VMA, with no gaps
    before or afterwards.

    After commit d23cb648e365 ("mm/mremap: permit mremap() move of multiple
    VMAs"), the multi VMA move permitted a gap to exist only after VMAs.
    This was done to provide maximum flexibility.

    However, We have consequently permitted this behaviour for the move of
    a single VMA including those not eligible for multi VMA move.

    The change introduced here means that we no longer permit non-eligible
    VMAs from being moved in this way.

    This is consistent, as it means all eligible VMA moves are treated the
    same, and all non-eligible moves are treated as they were before.

    This change does not break previous behaviour, which equally would have
    disallowed such a move (only in all cases).

    [lorenzo.stoakes@oracle.com: do not incorrectly reference invalid VMA in VM_WARN_ON_ONCE()]
      Link: https://lkml.kernel.org/r/b6dbda20-667e-4053-abae-8ed4fa84bb6c@lucifer.local
    Link: https://lkml.kernel.org/r/2b5aad5681573be85b5b8fac61399af6fb6b68b6.1754218667.git.lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reviewed-by: Vlastimil Babka <vbabka@suse.cz>
    Cc: David Hildenbrand <david@redhat.com>
    Cc: Jann Horn <jannh@google.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Mike Rapoport <rppt@kernel.org>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:59 -04:00
Rafael Aquini 434728cdb3 mm/mremap: allow multi-VMA move when filesystem uses thp_get_unmapped_area
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 7c91e0b91aaa3fa1f897efb06565af0ceb75195c
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Sun Aug 3 12:11:21 2025 +0100

    mm/mremap: allow multi-VMA move when filesystem uses thp_get_unmapped_area

    The multi-VMA move functionality introduced in commit d23cb648e365
    ("mm/mremap: permit mremap() move of multiple VMA") doesn't allow moves of
    file-backed mappings which specify a custom f_op->get_unmapped_area
    handler excepting hugetlb and shmem.

    We expand this to include thp_get_unmapped_area to support file-backed
    mappings for filesystems which use large folios.

    Additionally, when the first VMA in a range is not compatible with a
    multi-VMA move, instead of moving the first VMA and returning an error,
    this series results in us not moving anything and returning an error
    immediately.

    Examining this second change in detail:

    The semantics of multi-VMA moves in mremap() very clearly indicate that a
    failure can result in a partial move of VMAs.

    This is in line with other aggregate operations within the kernel, which
    share these semantics.

    There are two classes of failures we're concerned with - eligiblity for
    mutli-VMA move, and transient failures that would occur even if the user
    individually moved each VMA.

    The latter is due to out-of-memory conditions (which, given the
    allocations involved are small, would likely be fatal in any case), or
    hitting the mapping limit.

    Regardless of the cause, transient issues would be fatal anyway, so it
    isn't really material which VMAs succeeded at being moved or not.

    However with when it comes to multi-VMA move eligiblity, we face another
    issue - we must allow a single VMA to succeed regardless of this
    eligiblity (as, of course, it is not a multi-VMA move) - but we must then
    fail multi-VMA operations.

    The two means by which VMAs may fail the eligbility test are - the VMAs
    being UFFD-armed, or the VMA being file-backed and providing its own
    f_op->get_unmapped_area() helper (because this may result in MREMAP_FIXED
    being disregarded), excepting those known to correctly handle
    MREMAP_FIXED.

    It is therefore conceivable that a user could erroneously try to use this
    functionality in these instances, and would prefer to not perform any move
    at all should that occur.

    This series therefore avoids any move of subsequent VMAs should the first
    be multi-VMA move ineligble and the input span exceeds that of the first
    VMA.

    We also add detailed test logic to assert that multi VMA move with
    ineligible VMAs functions as expected.

    This patch (of 3):

    We currently restrict multi-VMA move to avoid filesystems or drivers which
    provide a custom f_op->get_unmapped_area handler unless it is known to
    correctly handle MREMAP_FIXED.

    We do this so we do not get unexpected result when moving from one area to
    another (for instance, if the handler would align things resulting in the
    moved VMAs having different gaps than the original mapping).

    More and more filesystems are moving to using large folios, and typically
    do so (in part) by setting f_op->get_unmapped_area to
    thp_get_unmapped_area.

    When mremap() invokes the file system's get_unmapped MREMAP_FIXED, it does
    so via get_unmapped_area(), called in vrm_set_new_addr().  In order to do
    so, it converts the MREMAP_FIXED flag to a MAP_FIXED flag and passes this
    to the unmapped area handler.

    The __get_unmapped_area() function (called by get_unmapped_area()) in turn
    invokes the filesystem or driver's f_op->get_unmapped_area() handler.

    Therefore this is a point at which thp_get_unmapped_area() may be called
    (also, this is the case for anonymous mappings where the size is huge page
    aligned).

    thp_get_unmapped_area() calls thp_get_unmapped_area_vmflags() and
    __thp_get_unmapped_area() in turn (falling back to
    mm_get_unmapped_area_vm_flags() which is known to handle MAP_FIXED
    correctly).

    The __thp_get_unmapped_area() function in turn does nothing to change the
    address hint, nor the MAP_FIXED flag, only adjusting alignment parameters.
    It hten calls mm_get_unmapped_area_vmflags(), and in turn arch-specific
    unmapped area functions, all of which honour MAP_FIXED correctly.

    Therefore, we can safely add thp_get_unmapped_area to the known-good
    handlers.

    Link: https://lkml.kernel.org/r/cover.1754218667.git.lorenzo.stoakes@oracle.com
    Link: https://lkml.kernel.org/r/4f2542340c29c84d3d470b0c605e916b192f6c81.1754218667.git.lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reviewed-by: Vlastimil Babka <vbabka@suse.cz>
    Cc: David Hildenbrand <david@redhat.com>
    Cc: Jann Horn <jannh@google.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Mike Rapoport <rppt@kernel.org>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:59 -04:00
Rafael Aquini 5631108580 mm/numa_memblks: Use pr_debug instead of printk(KERN_DEBUG)
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit d045c3154080a04beb07726fa311b89d21608981
Author: Pratyush Brahma <pratyush.brahma@oss.qualcomm.com>
Date:   Wed Aug 13 12:51:02 2025 +0530

    mm/numa_memblks: Use pr_debug instead of printk(KERN_DEBUG)

    Replace the direct usage of printk(KERN_DEBUG ...) with pr_debug(...) to
    align with the consistent `pr_*` API usage within the file.

    Reviewed-by: Joshua Hahn <joshua.hahnjy@gmail.com>
    Signed-off-by: Pratyush Brahma <pratyush.brahma@oss.qualcomm.com>
    Link: https://lore.kernel.org/r/20250813-numa-dbg-v3-1-1dcd1234fcc5@oss.qualcomm.com
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:59 -04:00
Rafael Aquini 343c303b20 mm/mremap: avoid expensive folio lookup on mremap folio pte batch
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 0b5be138ce00f421bd7cc5a226061bd62c4ab850
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Thu Aug 7 19:58:19 2025 +0100

    mm/mremap: avoid expensive folio lookup on mremap folio pte batch

    It was discovered in the attached report that commit f822a9a81a31 ("mm:
    optimize mremap() by PTE batching") introduced a significant performance
    regression on a number of metrics on x86-64, most notably
    stress-ng.bigheap.realloc_calls_per_sec - indicating a 37.3% regression in
    number of mremap() calls per second.

    I was able to reproduce this locally on an intel x86-64 raptor lake
    system, noting an average of 143,857 realloc calls/sec (with a stddev of
    4,531 or 3.1%) prior to this patch being applied, and 81,503 afterwards
    (stddev of 2,131 or 2.6%) - a 43.3% regression.

    During testing I was able to determine that there was no meaningful
    difference in efforts to optimise the folio_pte_batch() operation, nor
    checking folio_test_large().

    This is within expectation, as a regression this large is likely to
    indicate we are accessing memory that is not yet in a cache line (and
    perhaps may even cause a main memory fetch).

    The expectation by those discussing this from the start was that
    vm_normal_folio() (invoked by mremap_folio_pte_batch()) would likely be
    the culprit due to having to retrieve memory from the vmemmap (which
    mremap() page table moves does not otherwise do, meaning this is
    inevitably cold memory).

    I was able to definitively determine that this theory is indeed correct
    and the cause of the issue.

    The solution is to restore part of an approach previously discarded on
    review, that is to invoke pte_batch_hint() which explicitly determines,
    through reference to the PTE alone (thus no vmemmap lookup), what the PTE
    batch size may be.

    On platforms other than arm64 this is currently hardcoded to return 1, so
    this naturally resolves the issue for x86-64, and for arm64 introduces
    little to no overhead as the pte cache line will be hot.

    With this patch applied, we move from 81,503 realloc calls/sec to 138,701
    (stddev of 496.1 or 0.4%), which is a -3.6% regression, however accounting
    for the variance in the original result, this is broadly restoring
    performance to its prior state.

    Link: https://lkml.kernel.org/r/20250807185819.199865-1-lorenzo.stoakes@oracle.com
    Fixes: f822a9a81a31 ("mm: optimize mremap() by PTE batching")
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reported-by: kernel test robot <oliver.sang@intel.com>
    Closes: https://lore.kernel.org/oe-lkp/202508071609.4e743d7c-lkp@intel.com
    Acked-by: David Hildenbrand <david@redhat.com>
    Acked-by: Pedro Falcato <pfalcato@suse.de>
    Reviewed-by: Barry Song <baohua@kernel.org>
    Acked-by: Vlastimil Babka <vbabka@suse.cz>
    Reviewed-by: Dev Jain <dev.jain@arm.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Jann Horn <jannh@google.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:59 -04:00
Rafael Aquini 33d3a596c1 userfaultfd: fix a crash in UFFDIO_MOVE when PMD is a migration entry
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-38686

commit aba6faec0103ed8f169be8dce2ead41fcb689446
Author: Suren Baghdasaryan <surenb@google.com>
Date:   Wed Aug 6 15:00:22 2025 -0700

    userfaultfd: fix a crash in UFFDIO_MOVE when PMD is a migration entry

    When UFFDIO_MOVE encounters a migration PMD entry, it proceeds with
    obtaining a folio and accessing it even though the entry is swp_entry_t.
    Add the missing check and let split_huge_pmd() handle migration entries.
    While at it also remove unnecessary folio check.

    [surenb@google.com: remove extra folio check, per David]
      Link: https://lkml.kernel.org/r/20250807200418.1963585-1-surenb@google.com
    Link: https://lkml.kernel.org/r/20250806220022.926763-1-surenb@google.com
    Fixes: adef440691 ("userfaultfd: UFFDIO_MOVE uABI")
    Signed-off-by: Suren Baghdasaryan <surenb@google.com>
    Reported-by: syzbot+b446dbe27035ef6bd6c2@syzkaller.appspotmail.com
    Closes: https://lore.kernel.org/all/68794b5c.a70a0220.693ce.0050.GAE@google.com/
    Reviewed-by: Peter Xu <peterx@redhat.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Andrea Arcangeli <aarcange@redhat.com>
    Cc: Lokesh Gidra <lokeshgidra@google.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:59 -04:00
Rafael Aquini e5f4d19019 mm: pass page directly instead of using folio_page
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit cf1b80dc31a1137b8b4568c138b453bf7453204a
Author: Dev Jain <dev.jain@arm.com>
Date:   Wed Aug 6 20:26:11 2025 +0530

    mm: pass page directly instead of using folio_page

    In commit_anon_folio_batch(), we iterate over all pages pointed to by the
    PTE batch.  Therefore we need to know the first page of the batch;
    currently we derive that via folio_page(folio, 0), but, that takes us to
    the first (head) page of the folio instead - our PTE batch may lie in the
    middle of the folio, leading to incorrectness.

    Bite the bullet and throw away the micro-optimization of reusing the folio
    in favour of code simplicity.  Derive the page and the folio in
    change_pte_range, and pass the page too to commit_anon_folio_batch to fix
    the aforementioned issue.

    Link: https://lkml.kernel.org/r/20250806145611.3962-1-dev.jain@arm.com
    Fixes: cac1db8c3aad ("mm: optimize mprotect() by PTE batching")
    Reported-by: syzbot+57bcc752f0df8bb1365c@syzkaller.appspotmail.com
    Signed-off-by: Dev Jain <dev.jain@arm.com>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Debugged-by: David Hildenbrand <david@redhat.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Anshuman Khandual <anshuman.khandual@arm.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Catalin Marinas <catalin.marinas@arm.com>
    Cc: Christophe Leroy <christophe.leroy@csgroup.eu>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Jann Horn <jannh@google.com>
    Cc: Joey Gouly <joey.gouly@arm.com>
    Cc: Kevin Brodsky <kevin.brodsky@arm.com>
    Cc: Lance Yang <ioworker0@gmail.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Peter Xu <peterx@redhat.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Will Deacon <will@kernel.org>
    Cc: Yang Shi <yang@os.amperecomputing.com>
    Cc: Yicong Yang <yangyicong@hisilicon.com>
    Cc: Zhenhua Huang <quic_zhenhuah@quicinc.com>
    Cc: Zi Yan <ziy@nvidia.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:59 -04:00
Rafael Aquini f6e2efbd92 mm/vmscan: fix inverted polarity in lru_gen_seq_show()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit eb5ca9094a18fb98777bf4814ea84c93bf7c271d
Author: Danilo Krummrich <dakr@kernel.org>
Date:   Sun Jul 27 12:59:06 2025 +0200

    mm/vmscan: fix inverted polarity in lru_gen_seq_show()

    Commit a7694ff11aa9 ("vmscan: don't bother with debugfs_real_fops()")
    started using debugfs_get_aux_num() to distinguish between the RW
    "lru_gen" and the RO "lru_gen_full" file [1].

    Willy reported the inverted polarity [2] and Al fixed it up in [3].

    However, the patch in [1] was applied. Hence, fix this up accordingly.

    Cc: Alexander Viro <viro@zeniv.linux.org.uk>
    Cc: Matthew Wilcox <willy@infradead.org>
    Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
    Link: https://lore.kernel.org/all/20250704040720.GP1880847@ZenIV/ [1]
    Link: https://lore.kernel.org/all/aGZu3Z730FQtqxsE@casper.infradead.org/ [2]
    Link: https://lore.kernel.org/all/20250704040720.GP1880847@ZenIV/ [3]
    Fixes: a7694ff11aa9 ("vmscan: don't bother with debugfs_real_fops()")
    Acked-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
    Link: https://lore.kernel.org/r/20250727105937.7480-1-dakr@kernel.org
    Signed-off-by: Danilo Krummrich <dakr@kernel.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:58 -04:00
Rafael Aquini c87f76eb6c mm: fix the race between collapse and PT_RECLAIM under per-vma lock
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 366a4532d96fc357998465133db34d34edb79e4c
Author: Barry Song <baohua@kernel.org>
Date:   Tue Aug 5 11:54:47 2025 +0800

    mm: fix the race between collapse and PT_RECLAIM under per-vma lock

    The check_pmd_still_valid() call during collapse is currently only
    protected by the mmap_lock in write mode, which was sufficient when
    pt_reclaim always ran under mmap_lock in read mode.  However, since
    madvise_dontneed can now execute under a per-VMA lock, this assumption is
    no longer valid.  As a result, a race condition can occur between collapse
    and PT_RECLAIM, potentially leading to a kernel panic.

     [   38.151897] Oops: general protection fault, probably for non-canonical address 0xdffffc0000000003: 0000 [#1] SMP KASI
     [   38.153519] KASAN: null-ptr-deref in range [0x0000000000000018-0x000000000000001f]
     [   38.154605] CPU: 0 UID: 0 PID: 721 Comm: repro Not tainted 6.16.0-next-20250801-next-2025080 #1 PREEMPT(voluntary)
     [   38.155929] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS rel-1.16.3-0-ga6ed6b701f0a-prebuilt.qemu.org4
     [   38.157418] RIP: 0010:kasan_byte_accessible+0x15/0x30
     [   38.158125] Code: 03 0f 1f 40 00 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 66 0f 1f 00 48 b8 00 00 00 00 00 fc0
     [   38.160461] RSP: 0018:ffff88800feef678 EFLAGS: 00010286
     [   38.161220] RAX: dffffc0000000000 RBX: 0000000000000001 RCX: 1ffffffff0dde60c
     [   38.162232] RDX: 0000000000000000 RSI: ffffffff85da1e18 RDI: dffffc0000000003
     [   38.163176] RBP: ffff88800feef698 R08: 0000000000000001 R09: 0000000000000000
     [   38.164195] R10: 0000000000000000 R11: ffff888016a8ba58 R12: 0000000000000018
     [   38.165189] R13: 0000000000000018 R14: ffffffff85da1e18 R15: 0000000000000000
     [   38.166100] FS:  0000000000000000(0000) GS:ffff8880e3b40000(0000) knlGS:0000000000000000
     [   38.167137] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
     [   38.167891] CR2: 00007f97fadfe504 CR3: 0000000007088005 CR4: 0000000000770ef0
     [   38.168812] PKRU: 55555554
     [   38.169275] Call Trace:
     [   38.169647]  <TASK>
     [   38.169975]  ? __kasan_check_byte+0x19/0x50
     [   38.170581]  lock_acquire+0xea/0x310
     [   38.171083]  ? rcu_is_watching+0x19/0xc0
     [   38.171615]  ? __sanitizer_cov_trace_const_cmp4+0x1a/0x20
     [   38.172343]  ? __sanitizer_cov_trace_const_cmp8+0x1c/0x30
     [   38.173130]  _raw_spin_lock+0x38/0x50
     [   38.173707]  ? __pte_offset_map_lock+0x1a2/0x3c0
     [   38.174390]  __pte_offset_map_lock+0x1a2/0x3c0
     [   38.174987]  ? __pfx___pte_offset_map_lock+0x10/0x10
     [   38.175724]  ? __pfx_pud_val+0x10/0x10
     [   38.176308]  ? __sanitizer_cov_trace_const_cmp1+0x1e/0x30
     [   38.177183]  unmap_page_range+0xb60/0x43e0
     [   38.177824]  ? __pfx_unmap_page_range+0x10/0x10
     [   38.178485]  ? mas_next_slot+0x133a/0x1a50
     [   38.179079]  unmap_single_vma.constprop.0+0x15b/0x250
     [   38.179830]  unmap_vmas+0x1fa/0x460
     [   38.180373]  ? __pfx_unmap_vmas+0x10/0x10
     [   38.180994]  ? __sanitizer_cov_trace_const_cmp4+0x1a/0x20
     [   38.181877]  exit_mmap+0x1a2/0xb40
     [   38.182396]  ? lock_release+0x14f/0x2c0
     [   38.182929]  ? __pfx_exit_mmap+0x10/0x10
     [   38.183474]  ? __pfx___mutex_unlock_slowpath+0x10/0x10
     [   38.184188]  ? mutex_unlock+0x16/0x20
     [   38.184704]  mmput+0x132/0x370
     [   38.185208]  do_exit+0x7e7/0x28c0
     [   38.185682]  ? __this_cpu_preempt_check+0x21/0x30
     [   38.186328]  ? do_group_exit+0x1d8/0x2c0
     [   38.186873]  ? __pfx_do_exit+0x10/0x10
     [   38.187401]  ? __this_cpu_preempt_check+0x21/0x30
     [   38.188036]  ? _raw_spin_unlock_irq+0x2c/0x60
     [   38.188634]  ? lockdep_hardirqs_on+0x89/0x110
     [   38.189313]  do_group_exit+0xe4/0x2c0
     [   38.189831]  __x64_sys_exit_group+0x4d/0x60
     [   38.190413]  x64_sys_call+0x2174/0x2180
     [   38.190935]  do_syscall_64+0x6d/0x2e0
     [   38.191449]  entry_SYSCALL_64_after_hwframe+0x76/0x7e

    This patch moves the vma_start_write() call to precede
    check_pmd_still_valid(), ensuring that the check is also properly
    protected by the per-VMA lock.

    Link: https://lkml.kernel.org/r/20250805035447.7958-1-21cnbao@gmail.com
    Fixes: a6fde7add78d ("mm: use per_vma lock for MADV_DONTNEED")
    Signed-off-by: Barry Song <v-songbaohua@oppo.com>
    Tested-by: "Lai, Yi" <yi1.lai@linux.intel.com>
    Reported-by: "Lai, Yi" <yi1.lai@linux.intel.com>
    Closes: https://lore.kernel.org/all/aJAFrYfyzGpbm+0m@ly-workstation/
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: David Hildenbrand <david@redhat.com>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Qi Zheng <zhengqi.arch@bytedance.com>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Cc: Jann Horn <jannh@google.com>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: Lokesh Gidra <lokeshgidra@google.com>
    Cc: Tangquan Zheng <zhengtangquan@oppo.com>
    Cc: Lance Yang <ioworker0@gmail.com>
    Cc: Zi Yan <ziy@nvidia.com>
    Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Liam R. Howlett <Liam.Howlett@oracle.com>
    Cc: Nico Pache <npache@redhat.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Dev Jain <dev.jain@arm.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:58 -04:00
Rafael Aquini f4c7f8bb5e s390/mm: Allocate page table with PAGE_SIZE granularity
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit daa8af80d283ee9a7d42dd6f164a65036665b9d4
Author: Sumanth Korikkar <sumanthk@linux.ibm.com>
Date:   Mon Aug 4 11:57:03 2025 +0200

    s390/mm: Allocate page table with PAGE_SIZE granularity

    Make vmem_pte_alloc() consistent by always allocating page table of
    PAGE_SIZE granularity, regardless of whether page_table_alloc() (with
    slab) or memblock_alloc() is used. This ensures page table can be fully
    freed when the corresponding page table entries are removed.

    Fixes: d08d4e7cd6 ("s390/mm: use full 4KB page for 2KB PTE")
    Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
    Reviewed-by: Alexander Gordeev <agordeev@linux.ibm.com>
    Signed-off-by: Sumanth Korikkar <sumanthk@linux.ibm.com>
    Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:58 -04:00
Rafael Aquini 18dbc4ab12 mm: mempool: fix crash in mempool_free() for zero-minimum pools
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit a2152fef29020e740ba0276930f3a24440012505
Author: Yadan Fan <ydfan@suse.com>
Date:   Fri Aug 1 02:14:45 2025 +0800

    mm: mempool: fix crash in mempool_free() for zero-minimum pools

    The mempool wake-up fix introduced in commit a5867a218d7c ("mm: mempool:
    fix wake-up edge case bug for zero-minimum pools") inlined the
    add_element() logic in mempool_free() to return the element to the
    zero-minimum pool:

    pool->elements[pool->curr_nr++] = element;

    This causes crash, because mempool_init_node() does not initialize with
    real allocation for zero-minimum pool, it only returns ZERO_SIZE_PTR to
    the elements array which is unable to be dereferenced, and the
    pre-allocation of this array never happened since the while test:

    while (pool->curr_nr < pool->min_nr)

    can never be satisfied as min_nr is zero, so the pool does not actually
    reserve any buffer, the only way so far is to call alloc_fn() to get
    buffer from SLUB, but if the memory is under high pressure the alloc_fn()
    could never get any buffer, the waiting thread would be in an indefinite
    loop of wake-sleep in a period until there is free memory to get.

    This patch changes mempool_init_node() to allocate 1 element for the
    elements array of zero-minimum pool, so that the pool will have reserved
    buffer to use.  This will fix the crash issue and let the waiting thread
    can get the reserved element when alloc_fn() failed to get buffer under
    high memory pressure.

    Also modify add_element() to support zero-minimum pool with simplifying
    codes of zero-minimum handling in mempool_free().

    Link: https://lkml.kernel.org/r/e01f00f3-58d9-4ca7-af54-bfa42fec9527@suse.com
    Fixes: a5867a218d7c ("mm: mempool: fix wake-up edge case bug for zero-minimum pools")
    Signed-off-by: Yadan Fan <ydfan@suse.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:58 -04:00
Rafael Aquini 7a4f5d7012 mm: correct type for vmalloc vm_flags fields
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit f04fd85f15945f3ff189701050e3ce303c1a4d98
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Tue Jul 29 12:49:06 2025 +0100

    mm: correct type for vmalloc vm_flags fields

    Several functions refer to the unfortunately named 'vm_flags' field when
    referencing vmalloc flags, which happens to be the precise same name used
    for VMA flags.

    As a result these were erroneously changed to use the vm_flags_t type
    (which currently is a typedef equivalent to unsigned long).

    Currently this has no impact, but in future when vm_flags_t changes this
    will result in issues, so change the type to unsigned long to account for
    this.

    [lorenzo.stoakes@oracle.com: fixup very disguised vmalloc flags parameter]
      Link: https://lkml.kernel.org/r/e74dd8de-7e60-47ab-8a45-2c851f3c5d26@lucifer.local
    Link: https://lkml.kernel.org/r/20250729114906.55347-1-lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reported-by: Harry Yoo <harry.yoo@oracle.com>
    Closes: https://lore.kernel.org/all/aIgSpAnU8EaIcqd9@hyeyoo/
    Reviewed-by: Pedro Falcato <pfalcato@suse.de>
    Acked-by: David Hildenbrand <david@redhat.com>
    Reviewed-by: Harry Yoo <harry.yoo@oracle.com>
    Acked-by: Vlastimil Babka <vbabka@suse.cz>
    Cc: Jann Horn <jannh@google.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Mike Rapoport <rppt@kernel.org>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Cc: "Uladzislau Rezki (Sony)" <urezki@gmail.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:58 -04:00
Rafael Aquini 1558fc43df mm/shmem, swap: fix major fault counting
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit de55be42379cc0561aadfd9e1459239dea70be32
Author: Kairui Song <kasong@tencent.com>
Date:   Mon Jul 28 15:53:06 2025 +0800

    mm/shmem, swap: fix major fault counting

    If the swapin failed, don't update the major fault count.  There is a long
    existing comment for doing it this way, now with previous cleanups, we can
    finally fix it.

    Link: https://lkml.kernel.org/r/20250728075306.12704-9-ryncsn@gmail.com
    Signed-off-by: Kairui Song <kasong@tencent.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Baoquan He <bhe@redhat.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Kemeng Shi <shikemeng@huaweicloud.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Nhat Pham <nphamcs@gmail.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:58 -04:00
Rafael Aquini 9bb19d9e8d mm/shmem, swap: rework swap entry and index calculation for large swapin
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 93c0476e705768c7ca902cffea4efb500b9678b4
Author: Kairui Song <kasong@tencent.com>
Date:   Mon Jul 28 15:53:05 2025 +0800

    mm/shmem, swap: rework swap entry and index calculation for large swapin

    Instead of calculating the swap entry differently in different swapin
    paths, calculate it early before the swap cache lookup and use that for
    the lookup and later swapin.  And after swapin have brought a folio,
    simply round it down against the size of the folio.

    This is simple and effective enough to verify the swap value.  A folio's
    swap entry is always aligned by its size.  Any kind of parallel split or
    race is acceptable because the final shmem_add_to_page_cache ensures that
    all entries covered by the folio are correct, and thus there will be no
    data corruption.

    This also prevents false positive cache lookup.  If a shmem read request's
    index points to the middle of a large swap entry, previously, shmem will
    try the swap cache lookup using the large swap entry's starting value
    (which is the first sub swap entry of this large entry).  This will lead
    to false positive lookup results if only the first few swap entries are
    cached but the actual requested swap entry pointed by the index is
    uncached.  This is not a rare event, as swap readahead always tries to
    cache order 0 folios when possible.

    And this shouldn't cause any increased repeated faults.  Instead, no
    matter how the shmem mapping is split in parallel, as long as the mapping
    still contains the right entries, the swapin will succeed.

    The final object size and stack usage are also reduced due to simplified
    code:

    ./scripts/bloat-o-meter mm/shmem.o.old mm/shmem.o
    add/remove: 0/0 grow/shrink: 0/1 up/down: 0/-145 (-145)
    Function                                     old     new   delta
    shmem_swapin_folio                          4056    3911    -145
    Total: Before=33242, After=33097, chg -0.44%

    Stack usage (Before vs After):
    mm/shmem.c:2314:12:shmem_swapin_folio   264     static
    mm/shmem.c:2314:12:shmem_swapin_folio   256     static

    And while at it, round down the index too if swap entry is round down.
    The index is used either for folio reallocation or confirming the mapping
    content.  In either case, it should be aligned with the swap folio.

    Link: https://lkml.kernel.org/r/20250728075306.12704-8-ryncsn@gmail.com
    Signed-off-by: Kairui Song <kasong@tencent.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Tested-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Baoquan He <bhe@redhat.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Kemeng Shi <shikemeng@huaweicloud.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Nhat Pham <nphamcs@gmail.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:58 -04:00
Rafael Aquini 98bfdbf91a mm/shmem, swap: simplify swapin path and result handling
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 1326359f22805b2b0e9567ec0099980b8956fc29
Author: Kairui Song <kasong@tencent.com>
Date:   Mon Jul 28 15:53:04 2025 +0800

    mm/shmem, swap: simplify swapin path and result handling

    Slightly tidy up the different handling of swap in and error handling for
    SWP_SYNCHRONOUS_IO and non-SWP_SYNCHRONOUS_IO devices.  Now swapin will
    always use either shmem_swap_alloc_folio or shmem_swapin_cluster, then
    check the result.

    Simplify the control flow and avoid a redundant goto label.

    Link: https://lkml.kernel.org/r/20250728075306.12704-7-ryncsn@gmail.com
    Signed-off-by: Kairui Song <kasong@tencent.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Baoquan He <bhe@redhat.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Kemeng Shi <shikemeng@huaweicloud.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Nhat Pham <nphamcs@gmail.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:57 -04:00
Rafael Aquini 11ec0a0f26 mm/shmem, swap: never use swap cache and readahead for SWP_SYNCHRONOUS_IO
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 69805ea79db6634d4e7d596f3f36667924dc6cbf
Author: Kairui Song <kasong@tencent.com>
Date:   Mon Jul 28 15:53:03 2025 +0800

    mm/shmem, swap: never use swap cache and readahead for SWP_SYNCHRONOUS_IO

    For SWP_SYNCHRONOUS_IO devices, if a cache bypassing THP swapin failed due
    to reasons like memory pressure, partially conflicting swap cache or ZSWAP
    enabled, shmem will fallback to cached order 0 swapin.

    Right now the swap cache still has a non-trivial overhead, and readahead
    is not helpful for SWP_SYNCHRONOUS_IO devices, so we should always skip
    the readahead and swap cache even if the swapin falls back to order 0.

    So handle the fallback logic without falling back to the cached read.

    Link: https://lkml.kernel.org/r/20250728075306.12704-6-ryncsn@gmail.com
    Signed-off-by: Kairui Song <kasong@tencent.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Baoquan He <bhe@redhat.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Kemeng Shi <shikemeng@huaweicloud.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Nhat Pham <nphamcs@gmail.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:57 -04:00
Rafael Aquini b9273c152f mm/shmem, swap: tidy up swap entry splitting
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 91ab656ece137c368a3189dfd42f8c9203a6285c
Author: Kairui Song <kasong@tencent.com>
Date:   Mon Jul 28 15:53:02 2025 +0800

    mm/shmem, swap: tidy up swap entry splitting

    Instead of keeping different paths of splitting the entry before the swap
    in start, move the entry splitting after the swapin has put the folio in
    swap cache (or set the SWAP_HAS_CACHE bit).  This way we only need one
    place and one unified way to split the large entry.  Whenever swapin
    brought in a folio smaller than the shmem swap entry, split the entry and
    recalculate the entry and index for verification.

    This removes duplicated codes and function calls, reduces LOC, and the
    split is less racy as it's guarded by swap cache now.  So it will have a
    lower chance of repeated faults due to raced split.  The compiler is also
    able to optimize the coder further:

    bloat-o-meter results with GCC 14:

    With DEBUG_SECTION_MISMATCH (-fno-inline-functions-called-once):
    ./scripts/bloat-o-meter mm/shmem.o.old mm/shmem.o
    add/remove: 0/0 grow/shrink: 0/1 up/down: 0/-143 (-143)
    Function                                     old     new   delta
    shmem_swapin_folio                          2358    2215    -143
    Total: Before=32933, After=32790, chg -0.43%

    With !DEBUG_SECTION_MISMATCH:
    add/remove: 0/1 grow/shrink: 1/0 up/down: 1069/-749 (320)
    Function                                     old     new   delta
    shmem_swapin_folio                          2871    3940   +1069
    shmem_split_large_entry.isra                 749       -    -749
    Total: Before=32806, After=33126, chg +0.98%

    Since shmem_split_large_entry is only called in one place now. The
    compiler will either generate more compact code, or inlined it for
    better performance.

    Link: https://lkml.kernel.org/r/20250728075306.12704-5-ryncsn@gmail.com
    Signed-off-by: Kairui Song <kasong@tencent.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Tested-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Baoquan He <bhe@redhat.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Kemeng Shi <shikemeng@huaweicloud.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Nhat Pham <nphamcs@gmail.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:57 -04:00
Rafael Aquini aff1b4ed2c mm/shmem, swap: tidy up THP swapin checks
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit c262ffd72c8539d16ada8641a6348c5a88f0c542
Author: Kairui Song <kasong@tencent.com>
Date:   Mon Jul 28 15:53:01 2025 +0800

    mm/shmem, swap: tidy up THP swapin checks

    Move all THP swapin related checks under CONFIG_TRANSPARENT_HUGEPAGE, so
    they will be trimmed off by the compiler if not needed.

    And add a WARN if shmem sees a order > 0 entry when
    CONFIG_TRANSPARENT_HUGEPAGE is disabled, that should never happen unless
    things went very wrong.

    There should be no observable feature change except the new added WARN.

    Link: https://lkml.kernel.org/r/20250728075306.12704-4-ryncsn@gmail.com
    Signed-off-by: Kairui Song <kasong@tencent.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Baoquan He <bhe@redhat.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Kemeng Shi <shikemeng@huaweicloud.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Nhat Pham <nphamcs@gmail.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:57 -04:00
Rafael Aquini a7e68c57d7 mm/shmem, swap: avoid redundant Xarray lookup during swapin
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 0cfc0e7e3d062b93e9eec6828de000981cdfb152
Author: Kairui Song <kasong@tencent.com>
Date:   Mon Jul 28 15:53:00 2025 +0800

    mm/shmem, swap: avoid redundant Xarray lookup during swapin

    Patch series "mm/shmem, swap: bugfix and improvement of mTHP swap in", v6.

    The current THP swapin path have several problems.  It may potentially
    hang, may cause redundant faults due to false positive swap cache lookup,
    and it issues redundant Xarray walks.  !CONFIG_TRANSPARENT_HUGEPAGE builds
    may also contain unnecessary THP checks.

    This series fixes all of the mentioned issues, the code should be more
    robust and prepared for the swap table series.  Now 4 walks is reduced to
    3 (get order & confirm, confirm, insert folio),
    !CONFIG_TRANSPARENT_HUGEPAGE build overhead is also minimized, and comes
    with a sanity check now.

    The performance is slightly better after this series, sequential swap in
    of 24G data from ZRAM, using transparent_hugepage_tmpfs=always (24 samples
    each):

    Before:         avg: 10.66s, stddev: 0.04
    After patch 1:  avg: 10.58s, stddev: 0.04
    After patch 2:  avg: 10.65s, stddev: 0.05
    After patch 3:  avg: 10.65s, stddev: 0.04
    After patch 4:  avg: 10.67s, stddev: 0.04
    After patch 5:  avg: 9.79s,  stddev: 0.04
    After patch 6:  avg: 9.79s,  stddev: 0.05
    After patch 7:  avg: 9.78s,  stddev: 0.05
    After patch 8:  avg: 9.79s,  stddev: 0.04

    Several patches improve the performance by a little, which is about ~8%
    faster in total.

    Build kernel test showed very slightly improvement, testing with make -j48
    with defconfig in a 768M memcg also using ZRAM as swap, and
    transparent_hugepage_tmpfs=always (6 test runs):

    Before:         avg: 3334.66s, stddev: 43.76
    After patch 1:  avg: 3349.77s, stddev: 18.55
    After patch 2:  avg: 3325.01s, stddev: 42.96
    After patch 3:  avg: 3354.58s, stddev: 14.62
    After patch 4:  avg: 3336.24s, stddev: 32.15
    After patch 5:  avg: 3325.13s, stddev: 22.14
    After patch 6:  avg: 3285.03s, stddev: 38.95
    After patch 7:  avg: 3287.32s, stddev: 26.37
    After patch 8:  avg: 3295.87s, stddev: 46.24

    This patch (of 7):

    Currently shmem calls xa_get_order to get the swap radix entry order,
    requiring a full tree walk.  This can be easily combined with the swap
    entry value checking (shmem_confirm_swap) to avoid the duplicated lookup
    and abort early if the entry is gone already.  Which should improve the
    performance.

    Link: https://lkml.kernel.org/r/20250728075306.12704-1-ryncsn@gmail.com
    Link: https://lkml.kernel.org/r/20250728075306.12704-3-ryncsn@gmail.com
    Signed-off-by: Kairui Song <kasong@tencent.com>
    Reviewed-by: Kemeng Shi <shikemeng@huaweicloud.com>
    Reviewed-by: Dev Jain <dev.jain@arm.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Baoquan He <bhe@redhat.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Chris Li <chrisl@kernel.org>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
    Cc: Nhat Pham <nphamcs@gmail.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:57 -04:00
Rafael Aquini 090953f63a x86/ftrace: enable EXECMEM_ROX_CACHE for ftrace allocations
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 5d79c2be508143559c65ace445e7a951ef92881b
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Sun Jul 13 10:17:30 2025 +0300

    x86/ftrace: enable EXECMEM_ROX_CACHE for ftrace allocations

    For the most part ftrace uses text poking and can handle ROX memory.  The
    only place that requires writable memory is create_trampoline() that
    updates the allocated memory and in the end makes it ROX.

    Use execmem_alloc_rw() in x86::ftrace::alloc_tramp() and enable ROX cache
    for EXECMEM_FTRACE when configuration and CPU features allow that.

    Link: https://lkml.kernel.org/r/20250713071730.4117334-9-rppt@kernel.org
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Acked-by: Steven Rostedt (Google) <rostedt@goodmis.org>
    Cc: Daniel Gomez <da.gomez@samsung.com>
    Cc: Masami Hiramatsu (Google) <mhiramat@kernel.org>
    Cc: Petr Pavlu <petr.pavlu@suse.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:57 -04:00
Rafael Aquini 3712080d20 x86/kprobes: enable EXECMEM_ROX_CACHE for kprobes allocations
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 36de1e4238c1243866eaec515ef59972c490367f
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Sun Jul 13 10:17:29 2025 +0300

    x86/kprobes: enable EXECMEM_ROX_CACHE for kprobes allocations

    x86::alloc_insn_page() always allocates ROX memory.

    Instead of overriding this method, add EXECMEM_KPROBES entry in
    execmem_info with pgprot set to PAGE_KERNEL_ROX and use ROX cache when
    configuration and CPU features allow it.

    Link: https://lkml.kernel.org/r/20250713071730.4117334-8-rppt@kernel.org
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Acked-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
    Cc: Daniel Gomez <da.gomez@samsung.com>
    Cc: Petr Pavlu <petr.pavlu@suse.com>
    Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:56 -04:00
Rafael Aquini 5e4aec77f8 execmem: drop writable parameter from execmem_fill_trapping_insns()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit ab674b6871b049aab2e86d1d7375526368ed175a
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Sun Jul 13 10:17:28 2025 +0300

    execmem: drop writable parameter from execmem_fill_trapping_insns()

    After update of execmem_cache_free() that made memory writable before
    updating it, there is no need to update read only memory, so the writable
    parameter to execmem_fill_trapping_insns() is not needed.  Drop it.

    Link: https://lkml.kernel.org/r/20250713071730.4117334-7-rppt@kernel.org
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Cc: Daniel Gomez <da.gomez@samsung.com>
    Cc: Masami Hiramatsu (Google) <mhiramat@kernel.org>
    Cc: Petr Pavlu <petr.pavlu@suse.com>
    Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:56 -04:00
Rafael Aquini e591ba43a3 execmem: add fallback for failures in vmalloc(VM_ALLOW_HUGE_VMAP)
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 3bd4e0ac61b2fd87d64572e866f58940d1d5fbdf
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Sun Jul 13 10:17:27 2025 +0300

    execmem: add fallback for failures in vmalloc(VM_ALLOW_HUGE_VMAP)

    When execmem populates ROX cache it uses vmalloc(VM_ALLOW_HUGE_VMAP).
    Although vmalloc falls back to allocating base pages if high order
    allocation fails, it may happen that it still cannot allocate enough
    memory.

    Right now ROX cache is only used by modules and in majority of cases the
    allocations happen at boot time when there's plenty of free memory, but
    upcoming enabling ROX cache for ftrace and kprobes would mean that execmem
    allocations can happen when the system is under memory pressure and a
    failure to allocate large page worth of memory becomes more likely.

    Fallback to regular vmalloc() if vmalloc(VM_ALLOW_HUGE_VMAP) fails.

    Link: https://lkml.kernel.org/r/20250713071730.4117334-6-rppt@kernel.org
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Cc: Daniel Gomez <da.gomez@samsung.com>
    Cc: Masami Hiramatsu (Google) <mhiramat@kernel.org>
    Cc: Petr Pavlu <petr.pavlu@suse.com>
    Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:56 -04:00
Rafael Aquini 83f5304c66 execmem: move execmem_force_rw() and execmem_restore_rox() before use
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 888b5a847ba9650f454cd0842ccf8497268da959
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Sun Jul 13 10:17:26 2025 +0300

    execmem: move execmem_force_rw() and execmem_restore_rox() before use

    to avoid static declarations.

    Link: https://lkml.kernel.org/r/20250713071730.4117334-5-rppt@kernel.org
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Cc: Daniel Gomez <da.gomez@samsung.com>
    Cc: Masami Hiramatsu (Google) <mhiramat@kernel.org>
    Cc: Petr Pavlu <petr.pavlu@suse.com>
    Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:56 -04:00
Rafael Aquini 8d4177d54d execmem: rework execmem_cache_free()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 187fd8521dd8b202cbacd7af57f4301da4d5b52d
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Sun Jul 13 10:17:25 2025 +0300

    execmem: rework execmem_cache_free()

    Currently execmem_cache_free() ignores potential allocation failures that
    may happen in execmem_cache_add().  Besides, it uses text poking to fill
    the memory with trapping instructions before returning it to cache
    although it would be more efficient to make that memory writable, update
    it using memcpy and then restore ROX protection.

    Rework execmem_cache_free() so that in case of an error it will defer
    freeing of the memory to a delayed work.

    With this the happy fast path will now change permissions to RW, fill the
    memory with trapping instructions using memcpy, restore ROX permissions,
    add the memory back to the free cache and clear the relevant entry in
    busy_areas.

    If any step in the fast path fails, the entry in busy_areas will be marked
    as pending_free.  These entries will be handled by a delayed work and
    freed asynchronously.

    To make the fast path faster, use __GFP_NORETRY for memory allocations and
    let asynchronous handler try harder with GFP_KERNEL.

    Link: https://lkml.kernel.org/r/20250713071730.4117334-4-rppt@kernel.org
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Cc: Daniel Gomez <da.gomez@samsung.com>
    Cc: Masami Hiramatsu (Google) <mhiramat@kernel.org>
    Cc: Petr Pavlu <petr.pavlu@suse.com>
    Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:56 -04:00
Rafael Aquini f200bc43a9 execmem: introduce execmem_alloc_rw()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 838955f64ae7582f009a3538889bb9244f37ab26
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Sun Jul 13 10:17:24 2025 +0300

    execmem: introduce execmem_alloc_rw()

    Some callers of execmem_alloc() require the memory to be temporarily
    writable even when it is allocated from ROX cache.  These callers use
    execemem_make_temp_rw() right after the call to execmem_alloc().

    Wrap this sequence in execmem_alloc_rw() API.

    Link: https://lkml.kernel.org/r/20250713071730.4117334-3-rppt@kernel.org
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Reviewed-by: Daniel Gomez <da.gomez@samsung.com>
    Reviewed-by: Petr Pavlu <petr.pavlu@suse.com>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Cc: Masami Hiramatsu (Google) <mhiramat@kernel.org>
    Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:56 -04:00
Rafael Aquini 62259ab9b7 execmem: drop unused execmem_update_copy()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit fcd90ad31e29d0b403f3a074a64cd7f0876175dd
Author: Mike Rapoport (Microsoft) <rppt@kernel.org>
Date:   Sun Jul 13 10:17:23 2025 +0300

    execmem: drop unused execmem_update_copy()

    Patch series "x86: enable EXECMEM_ROX_CACHE for ftrace and kprobes", v3.

    These patches enable use of EXECMEM_ROX_CACHE for ftrace and kprobes
    allocations on x86.

    They also include some ground work in execmem.

    Since the execmem model for caching large ROX pages changed from the
    initial assumption that the memory that is allocated from ROX cache is
    always ROX to the current state where memory can be temporarily made RW
    and then restored to ROX, we can stop using text poking to update it.
    This also saves the hassle of trying lock text_mutex in
    execmem_cache_free() when kprobes already hold that mutex.

    This patch (of 8):

    The execmem_update_copy() that used text poking was required when memory
    allocated from ROX cache was always read-only.  Since now its permissions
    can be switched to read-write there is no need in a function that updates
    memory with text poking.

    Remove it.

    Link: https://lkml.kernel.org/r/20250713071730.4117334-1-rppt@kernel.org
    Link: https://lkml.kernel.org/r/20250713071730.4117334-2-rppt@kernel.org
    Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Cc: Daniel Gomez <da.gomez@samsung.com>
    Cc: Masami Hiramatsu (Google) <mhiramat@kernel.org>
    Cc: Petr Pavlu <petr.pavlu@suse.com>
    Cc: Steven Rostedt (Google) <rostedt@goodmis.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:55 -04:00
Rafael Aquini be6adf38e9 mm: fix a UAF when vma->mm is freed after vma->vm_refcnt got dropped
JIRA: https://redhat.atlassian.net/browse/RHEL-145695
CVE: CVE-2025-38554

commit 9bbffee67ffd16360179327b57f3b1245579ef08
Author: Suren Baghdasaryan <surenb@google.com>
Date:   Mon Jul 28 10:53:55 2025 -0700

    mm: fix a UAF when vma->mm is freed after vma->vm_refcnt got dropped

    By inducing delays in the right places, Jann Horn created a reproducer for
    a hard to hit UAF issue that became possible after VMAs were allowed to be
    recycled by adding SLAB_TYPESAFE_BY_RCU to their cache.

    Race description is borrowed from Jann's discovery report:
    lock_vma_under_rcu() looks up a VMA locklessly with mas_walk() under
    rcu_read_lock().  At that point, the VMA may be concurrently freed, and it
    can be recycled by another process.  vma_start_read() then increments the
    vma->vm_refcnt (if it is in an acceptable range), and if this succeeds,
    vma_start_read() can return a recycled VMA.

    In this scenario where the VMA has been recycled, lock_vma_under_rcu()
    will then detect the mismatching ->vm_mm pointer and drop the VMA through
    vma_end_read(), which calls vma_refcount_put().  vma_refcount_put() drops
    the refcount and then calls rcuwait_wake_up() using a copy of vma->vm_mm.
    This is wrong: It implicitly assumes that the caller is keeping the VMA's
    mm alive, but in this scenario the caller has no relation to the VMA's mm,
    so the rcuwait_wake_up() can cause UAF.

    The diagram depicting the race:
    T1         T2         T3
    ==         ==         ==
    lock_vma_under_rcu
      mas_walk
              <VMA gets removed from mm>
                          mmap
                            <the same VMA is reallocated>
      vma_start_read
        __refcount_inc_not_zero_limited_acquire
                          munmap
                            __vma_enter_locked
                              refcount_add_not_zero
      vma_end_read
        vma_refcount_put
          __refcount_dec_and_test
                              rcuwait_wait_event
                                <finish operation>
          rcuwait_wake_up [UAF]

    Note that rcuwait_wait_event() in T3 does not block because refcount was
    already dropped by T1.  At this point T3 can exit and free the mm causing
    UAF in T1.

    To avoid this we move vma->vm_mm verification into vma_start_read() and
    grab vma->vm_mm to stabilize it before vma_refcount_put() operation.

    [surenb@google.com: v3]
      Link: https://lkml.kernel.org/r/20250729145709.2731370-1-surenb@google.com
    Link: https://lkml.kernel.org/r/20250728175355.2282375-1-surenb@google.com
    Fixes: 3104138517fc ("mm: make vma cache SLAB_TYPESAFE_BY_RCU")
    Signed-off-by: Suren Baghdasaryan <surenb@google.com>
    Reported-by: Jann Horn <jannh@google.com>
    Closes: https://lore.kernel.org/all/CAG48ez0-deFbVH=E3jbkWx=X3uVbd8nWeo6kbJPQ0KoUD+m2tA@mail.gmail.com/
    Reviewed-by: Vlastimil Babka <vbabka@suse.cz>
    Acked-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Jann Horn <jannh@google.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: <stable@vger.kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:55 -04:00
Rafael Aquini ca015bd615 mm/rmap: add anon_vma lifetime debug check
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit a222439e1e273fa0f4e37ce17aeb109f3e91824f
Author: Jann Horn <jannh@google.com>
Date:   Fri Jul 25 14:16:24 2025 +0200

    mm/rmap: add anon_vma lifetime debug check

    If an anon folio is mapped into userspace, its anon_vma must be alive,
    otherwise rmap walks can hit UAF.

    There have been syzkaller reports a few months ago[1][2] of UAF in rmap
    walks that seems to indicate that there can be pages with elevated
    mapcount whose anon_vma has already been freed, but I think we never
    figured out what the cause is; and syzkaller only hit these UAFs when
    memory pressure randomly caused reclaim to rmap-walk the affected pages,
    so it of course didn't manage to create a reproducer.

    Add a VM_WARN_ON_FOLIO() when we add/remove mappings of anonymous folios
    to hopefully catch such issues more reliably.

    [1] https://lore.kernel.org/r/67abaeaf.050a0220.110943.0041.GAE@google.com
    [2] https://lore.kernel.org/r/67a76f33.050a0220.3d72c.0028.GAE@google.com

    Link: https://lkml.kernel.org/r/20250725-anonvma-uaf-debug-v2-1-bc3c7e5ba5b1@google.com
    Signed-off-by: Jann Horn <jannh@google.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Acked-by: Vlastimil Babka <vbabka@suse.cz>
    Acked-by: Harry Yoo <harry.yoo@oracle.com>
    Cc: David Hildenbrand <david@redhat.com>
    Cc: Jann Horn <jannh@google.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Rik van Riel <riel@surriel.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:55 -04:00
Rafael Aquini 1ea9960714 mm: remove mm/io-mapping.c
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 9a4f90e246615d1f42a9b907deb9b4c0a418d996
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Fri Jul 25 15:29:01 2025 +0100

    mm: remove mm/io-mapping.c

    This is dead code, which was used from commit b739f125e4 ("i915: use
    io_mapping_map_user") but reverted a month later by commit 0e4fe0c9f2
    ("Revert "i915: use io_mapping_map_user"") back in 2021.

    Since then nobody has used it, so remove it.

    [akpm@linux-foundation.org: update Documentation/core-api/mm-api.rst, per Vlastimil]
    Link: https://lkml.kernel.org/r/20250725142901.81502-1-lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Acked-by: Vlastimil Babka <vbabka@suse.cz>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Mike Rapoport <rppt@kernel.org>
    Cc: Suren Baghdasaryan <surenb@google.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:55 -04:00
Rafael Aquini dcf21bbe46 khugepaged: optimize collapse_pte_mapped_thp() by PTE batching
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 22d0229093b92db2fe6ca6ba946bad1f246024e8
Author: Dev Jain <dev.jain@arm.com>
Date:   Thu Jul 24 10:53:01 2025 +0530

    khugepaged: optimize collapse_pte_mapped_thp() by PTE batching

    Use PTE batching to batch process PTEs mapping the same large folio. An
    improvement is expected due to batching mapcount manipulation on the
    folios, and for arm64 which supports contig mappings, the number of
    TLB flushes is also reduced.

    Note that we do not need to make a change to the check
    "if (folio_page(folio, i) != page)"; if i'th page of the folio is equal
    to the first page of our batch, then i + 1, .... i + nr_batch_ptes - 1
    pages of the folio will be equal to the corresponding pages of our
    batch mapping consecutive pages.

    Link: https://lkml.kernel.org/r/20250724052301.23844-4-dev.jain@arm.com
    Signed-off-by: Dev Jain <dev.jain@arm.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reviewed-by: Zi Yan <ziy@nvidia.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Mariano Pache <npache@redhat.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:55 -04:00
Rafael Aquini 935d2319fe khugepaged: optimize __collapse_huge_page_copy_succeeded() by PTE batching
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 4ea3594a47412f9dd20fbda0dc70b0cbec9cba43
Author: Dev Jain <dev.jain@arm.com>
Date:   Thu Jul 24 10:53:00 2025 +0530

    khugepaged: optimize __collapse_huge_page_copy_succeeded() by PTE batching

    Use PTE batching to batch process PTEs mapping the same large folio. An
    improvement is expected due to batching refcount-mapcount manipulation on
    the folios, and for arm64 which supports contig mappings, the number of
    TLB flushes is also reduced.

    Link: https://lkml.kernel.org/r/20250724052301.23844-3-dev.jain@arm.com
    Signed-off-by: Dev Jain <dev.jain@arm.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Barry Song <baohua@kernel.org>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Mariano Pache <npache@redhat.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Zi Yan <ziy@nvidia.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:55 -04:00
Rafael Aquini fc63214189 mm: add get_and_clear_ptes() and clear_ptes()
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 3dfde97800e06882960cc926d2c428f2128b7c70
Author: David Hildenbrand <david@redhat.com>
Date:   Thu Jul 24 10:52:59 2025 +0530

    mm: add get_and_clear_ptes() and clear_ptes()

    Patch series "Optimizations for khugepaged", v4.

    If the underlying folio mapped by the ptes is large, we can process those
    ptes in a batch using folio_pte_batch().

    For arm64 specifically, this results in a 16x reduction in the number of
    ptep_get() calls, since on a contig block, ptep_get() on arm64 will
    iterate through all 16 entries to collect a/d bits.  Next, ptep_clear()
    will cause a TLBI for every contig block in the range via
    contpte_try_unfold().  Instead, use clear_ptes() to only do the TLBI at
    the first and last contig block of the range.

    For split folios, there will be no pte batching; the batch size returned
    by folio_pte_batch() will be 1.  For pagetable split folios, the ptes will
    still point to the same large folio; for arm64, this results in the
    optimization described above, and for other arches, a minor improvement is
    expected due to a reduction in the number of function calls and batching
    atomic operations.

    This patch (of 3):

    Let's add variants to be used where "full" does not apply -- which will
    be the majority of cases in the future. "full" really only applies if
    we are about to tear down a full MM.

    Use get_and_clear_ptes() in existing code, clear_ptes() users will
    be added next.

    Link: https://lkml.kernel.org/r/20250724052301.23844-2-dev.jain@arm.com
    Signed-off-by: David Hildenbrand <david@redhat.com>
    Signed-off-by: Dev Jain <dev.jain@arm.com>
    Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
    Reviewed-by: Barry Song <baohua@kernel.org>
    Reviewed-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reviewed-by: Zi Yan <ziy@nvidia.com>
    Cc: Liam Howlett <liam.howlett@oracle.com>
    Cc: Mariano Pache <npache@redhat.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:55 -04:00
Rafael Aquini e088062036 mm/mincore: hold PTL in mincore_hugetlb
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 1623717b057f904d558eb0489fbd592a18750c1e
Author: Jinjiang Tu <tujinjiang@huawei.com>
Date:   Thu Jul 24 17:09:58 2025 +0800

    mm/mincore: hold PTL in mincore_hugetlb

    Hold PTL in mincore_hugetlb() to avoid operating on stale page, as
    mincore_pte_range() have done.

    Link: https://lkml.kernel.org/r/20250724090958.455887-4-tujinjiang@huawei.com
    Signed-off-by: Jinjiang Tu <tujinjiang@huawei.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Andrei Vagin <avagin@gmail.com>
    Cc: Andrii Nakryiko <andrii@kernel.org>
    Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Brahmajit Das <brahmajit.xyz@gmail.com>
    Cc: Catalin Marinas <catalin.marinas@arm.com>
    Cc: Christophe Leroy <christophe.leroy@csgroup.eu>
    Cc: David Rientjes <rientjes@google.com>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Joern Engel <joern@logfs.org>
    Cc: Kefeng Wang <wangkefeng.wang@huawei.com>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Thiago Jung Bauermann <thiago.bauermann@linaro.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:54 -04:00
Rafael Aquini 45c0c4091a mm/memory-failure: hold PTL in hwpoison_hugetlb_range
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 9109bd52559b44a66e4dbde69d0dd36f3e4dcae8
Author: Jinjiang Tu <tujinjiang@huawei.com>
Date:   Fri Jul 25 11:31:12 2025 +0800

    mm/memory-failure: hold PTL in hwpoison_hugetlb_range

    Hold PTL in hwpoison_hugetlb_range() to avoid operating on stale page, as
    hwpoison_pte_range() have done.

    This change is not known to address any issues which users have
    experienced.

    Link: https://lkml.kernel.org/r/20250725033112.2690158-1-tujinjiang@huawei.com
    Signed-off-by: Jinjiang Tu <tujinjiang@huawei.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Andrei Vagin <avagin@gmail.com>
    Cc: Andrii Nakryiko <andrii@kernel.org>
    Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
    Cc: Brahmajit Das <brahmajit.xyz@gmail.com>
    Cc: Catalin Marinas <catalin.marinas@arm.com>
    Cc: Christophe Leroy <christophe.leroy@csgroup.eu>
    Cc: David Rientjes <rientjes@google.com>
    Cc: Dev Jain <dev.jain@arm.com>
    Cc: Hugh Dickins <hughd@google.com>
    Cc: Joern Engel <joern@logfs.org>
    Cc: Kefeng Wang <wangkefeng.wang@huawei.com>
    Cc: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Cc: Michal Hocko <mhocko@suse.com>
    Cc: Ryan Roberts <ryan.roberts@arm.com>
    Cc: Thiago Jung Bauermann <thiago.bauermann@linaro.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:54 -04:00
Rafael Aquini 36fd0515f6 mm/mseal: rework mseal apply logic
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 6c2da14ae1e0a0146587381594559027bd46c059
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Fri Jul 25 09:29:45 2025 +0100

    mm/mseal: rework mseal apply logic

    The logic can be simplified - firstly by renaming the inconsistently named
    apply_mm_seal() to mseal_apply().

    We then wrap mseal_fixup() into the main loop as the logic is simple
    enough to not require it, equally it isn't a hugely pleasant pattern in
    mprotect() etc.  so it's not something we want to perpetuate.

    We eliminate the need for invoking vma_iter_end() on each loop by directly
    determining if the VMA was merged - the only thing we need concern
    ourselves with is whether the start/end of the (gapless) range are offset
    into VMAs.

    This refactoring also avoids the rather horrid 'pass pointer to prev
    around' pattern used in mprotect() et al.

    No functional change intended.

    Link: https://lkml.kernel.org/r/ddfa4376ce29f19a589d7dc8c92cb7d4f7605a4c.1753431105.git.lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reviewed-by: Pedro Falcato <pfalcato@suse.de>
    Reviewed-by: Liam R. Howlett <Liam.Howlett@oracle.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Acked-by: Jeff Xu <jeffxu@chromium.org>
    Cc: Jann Horn <jannh@google.com>
    Cc: Kees Cook <kees@kernel.org>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:54 -04:00
Rafael Aquini 891cf145a7 mm/mseal: simplify and rename VMA gap check
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 530e090964130d538dfa74874012ca461ef692fa
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Fri Jul 25 09:29:44 2025 +0100

    mm/mseal: simplify and rename VMA gap check

    The check_mm_seal() function is doing something general - checking whether
    a range contains only VMAs (or rather that it does NOT contain any
    unmapped regions).

    So rename this function to range_contains_unmapped().

    Additionally simplify the logic, we are simply checking whether the last
    vma->vm_end has either a VMA starting after it or ends before the end
    parameter.

    This check is rather dubious, so it is sensible to keep it local to
    mm/mseal.c as at a later stage it may be removed, and we don't want any
    other mm code to perform such a check.

    No functional change intended.

    [lorenzo.stoakes@oracle.com: add comment explaining why we disallow gaps on mseal()]
      Link: https://lkml.kernel.org/r/d85b3d55-09dc-43ba-8204-b48267a96751@lucifer.local
    Link: https://lkml.kernel.org/r/dd50984eff1e242b5f7f0f070a3360ef760e06b8.1753431105.git.lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reviewed-by: Liam R. Howlett <Liam.Howlett@oracle.com>
    Acked-by: David Hildenbrand <david@redhat.com>
    Acked-by: Jeff Xu <jeffxu@chromium.org>
    Reviewed-by: Pedro Falcato <pfalcato@suse.de>
    Cc: Jann Horn <jannh@google.com>
    Cc: Kees Cook <kees@kernel.org>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:54 -04:00
Rafael Aquini ddc7217a11 mm/mseal: small cleanups
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit 8b2914162aa3a56062d4b7c716149946672d48a6
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Fri Jul 25 09:29:43 2025 +0100

    mm/mseal: small cleanups

    Drop the wholly unnecessary set_vma_sealed() helper(), which is used only
    once, and place VMA_ITERATOR() declarations in the correct place.

    Retain vma_is_sealed(), and use it instead of the confusingly named
    can_modify_vma(), so it's abundantly clear what's being tested, rather
    then a nebulous sense of 'can the VMA be modified'.

    No functional change intended.

    Link: https://lkml.kernel.org/r/98cf28d04583d632a6eb698e9ad23733bb6af26b.1753431105.git.lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reviewed-by: Liam R. Howlett <Liam.Howlett@oracle.com>
    Reviewed-by: Pedro Falcato <pfalcato@suse.de>
    Acked-by: David Hildenbrand <david@redhat.com>
    Acked-by: Jeff Xu <jeffxu@chromium.org>
    Cc: Jann Horn <jannh@google.com>
    Cc: Kees Cook <kees@kernel.org>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:54 -04:00
Rafael Aquini e0e1ed6ae5 mm/mseal: update madvise() logic
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit d0b47a6866f1047247061f3a38f12a981825b265
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Fri Jul 25 09:29:42 2025 +0100

    mm/mseal: update madvise() logic

    The madvise() logic is inexplicably performed in mm/mseal.c - this ought
    to be located in mm/madvise.c.

    Additionally can_modify_vma_madv() is inconsistently named and, in
    combination with is_ro_anon(), is very confusing logic.

    Put a static function in mm/madvise.c instead - can_madvise_modify() -
    that spells out exactly what's happening.  Also explicitly check for an
    anon VMA.

    Also add commentary to explain what's going on.

    Essentially - we disallow discarding of data in mseal()'d mappings in
    instances where the user couldn't otherwise write to that data.

    We retain the existing behaviour here regarding MAP_PRIVATE mappings of
    file-backed mappings, which entails some complexity - while this, strictly
    speaking - appears to violate mseal() semantics, it may interact badly
    with users which expect to be able to madvise(MADV_DONTNEED) .text
    mappings for instance.

    We may revisit this at a later date.

    No functional change intended.

    Link: https://lkml.kernel.org/r/492a98d9189646e92c8f23f4cce41ed323fe01df.1753431105.git.lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reviewed-by: Liam R. Howlett <Liam.Howlett@oracle.com>
    Reviewed-by: Pedro Falcato <pfalcato@suse.de>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Jann Horn <jannh@google.com>
    Cc: Jeff Xu <jeffxu@chromium.org>
    Cc: Kees Cook <kees@kernel.org>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:54 -04:00
Rafael Aquini 7c2ddaa6ee mm/mseal: always define VM_SEALED
JIRA: https://redhat.atlassian.net/browse/RHEL-145695

commit f225b34f1e6c81c50e48f6207ddb6d290be1b932
Author: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
Date:   Fri Jul 25 09:29:41 2025 +0100

    mm/mseal: always define VM_SEALED

    Patch series "mseal cleanups", v4.

    Perform a number of cleanups to the mseal logic.  Firstly, VM_SEALED is
    treated differently from every other VMA flag, it really doesn't make
    sense to do this, so we start by making this consistent with everything
    else.

    Next we place the madvise logic where it belongs - in mm/madvise.c.  It
    really makes no sense to abstract this elsewhere.  In doing so, we go to
    great lengths to explain very clearly the previously very confusing logic
    as to what sealed mappings are impacted here.

    In doing so, we retain existing logic regarding treatment of madvise()
    discard operations for a sealed, read-only MAP_PRIVATE file-backed
    mapping.  This is something we likely need to revisit.

    We then abstract out and explain the 'are there are any gaps in this range
    in the mm?' check being performed as a prerequisite to mseal being
    performed.

    Finally, we simplify the actual mseal logic which is really quite
    straightforward.

    No functional change is intended.

    This patch (of 4):

    There is no reason to treat VM_SEALED in a special way, in each other case
    in which a VMA flag is unavailable due to configuration, we simply assign
    that flag to VM_NONE, so make VM_SEALED consistent with all other VMA
    flags in this respect.

    Additionally, use the next available bit for VM_SEALED, 42, rather than
    arbitrarily putting it at 63 and update the declaration to match all other
    VMA flags.

    No functional change intended.

    Link: https://lkml.kernel.org/r/cover.1753431105.git.lorenzo.stoakes@oracle.com
    Link: https://lkml.kernel.org/r/aeb398a77029b6e7377cd944328bc9bbc3c90537.1753431105.git.lorenzo.stoakes@oracle.com
    Signed-off-by: Lorenzo Stoakes <lorenzo.stoakes@oracle.com>
    Reviewed-by: Liam R. Howlett <Liam.Howlett@oracle.com>
    Reviewed-by: Pedro Falcato <pfalcato@suse.de>
    Acked-by: David Hildenbrand <david@redhat.com>
    Cc: Jann Horn <jannh@google.com>
    Cc: Jeff Xu <jeffxu@chromium.org>
    Cc: Kees Cook <kees@kernel.org>
    Cc: Vlastimil Babka <vbabka@suse.cz>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Signed-off-by: Rafael Aquini <raquini@redhat.com>
2026-07-31 14:50:53 -04:00