100 Commits
Author SHA1 Message Date
Anusha Srivatsa f1906fd17d drm/i915/display: enable ccs modifiers on dg2
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit aab3d205a086233c612fee86009265451793e0c2
Author: Juha-Pekka Heikkila <juhapekka.heikkila@gmail.com>
Date:   Mon Apr 27 19:57:15 2026 +0300

    drm/i915/display: enable ccs modifiers on dg2

    Since Xe driver aux ccs enablement dg2 ccs modifiers have been
    disabled on i915 driver. Here allow dg2 to use ccs again for framebuffers.

    Fixes: 6a99e91a6ca8 ("drm/i915/display: Detect AuxCCS support via display parent interface")
    Signed-off-by: Juha-Pekka Heikkila <juhapekka.heikkila@gmail.com>
    Reviewed-by: Ville Syrjälä <ville.syrjala@linux.intel.com>
    Signed-off-by: Mika Kahola <mika.kahola@intel.com>
    Link: https://patch.msgid.link/20260427165715.864721-1-juhapekka.heikkila@gmail.com
    (cherry picked from commit aee13ba1448213975f36942ba5d1ce693eb5c002)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:01 -05:00
Anusha Srivatsa 8ce0fb3b0b drm/amd/display: Write REFCLK to 48MHz on DCN21
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit ec78a85d95e9c37b6ca16d6ed1639fa64d5dd6dc
Author: Ivan Lipski <ivan.lipski@amd.com>
Date:   Thu May 14 11:53:50 2026 -0400

    drm/amd/display: Write REFCLK to 48MHz on DCN21

    [Why&How]
    dccg21_init() calls dccg2_init() which hardcodes 100MHz refclk values
    for MICROSECOND_TIME_BASE_DIV and MILLISECOND_TIME_BASE_DIV. DCN21
    uses 48MHz refclk, so the wrong values corrupt DCCG timing and cause eDP
    link training failure on cold boot.

    Write the correct 48MHz values directly instead of calling dccg2_init().

    v2:
    Fixed typo

    Fixes: e6e2b956fc81 ("drm/amd/display: Add missing DCCG register entries for DCN20-DCN316")
    Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5272
    Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5311
    Reported-by: Max Chernoff <git@maxchernoff.ca>
    Tested-by: Max Chernoff <git@maxchernoff.ca>
    Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
    Acked-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 08236c3ef284cd2d110e5e3d51fc9615e551f9dc)
    Cc: stable@vger.kernel.org

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:01 -05:00
Anusha Srivatsa 76e42f4a13 drm/panthor: Extend VM locked region for remap case to be a superset
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 8867262d993d249309a8d3ec28ce095378cf1720
Author: Adrián Larumbe <adrian.larumbe@collabora.com>
Date:   Wed Apr 8 20:12:23 2026 +0100

    drm/panthor: Extend VM locked region for remap case to be a superset

    In the event of an sm_step_remap() that leads to a partial unmap of a
    transparent huge page, the new locked region required by an extended unmap
    might not be a superset of the original one. Then, if it leaves a portion
    of the initially requested one out, the ensuing map will trigger a warning.

    Fixes: 8e7460eac786 ("drm/panthor: Support partial unmaps of huge pages")
    Reviewed-by: Boris Brezillon <boris.brezillon@collabora.com>
    Reviewed-by: Steven Price <steven.price@arm.com>
    Reviewed-by: Liviu Dudau <liviu.dudau@arm.com>
    Link: https://patch.msgid.link/20260408191228.537625-1-adrian.larumbe@collabora.com
    Signed-off-by: Adrián Larumbe <adrian.larumbe@collabora.com>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:01 -05:00
Anusha Srivatsa fee2737774 drm/amdkfd: Check if there are kfd porcesses using adev by kfd_processes_count
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 81665e35f143d93adef654f3be1360def9196e72
Author: Xiaogang Chen <xiaogang.chen@amd.com>
Date:   Fri Apr 24 13:47:01 2026 -0500

    drm/amdkfd: Check if there are kfd porcesses using adev by kfd_processes_count

    During gpu hot-unplug need check if there are kfd porcesses still using the
    being removed gpu before clean resources of the device. Current driver checks
    if kfd_processes_table is empty. kfd processes are not terminated after
    removed from kfd_processes_table immediately. They are still alive and may
    access the device until kfd_process_wq work queue got ran.

    Check kfd->kfd_processes_count value that is updated after kfd process got
    uninitialized when its ref becomes zero.

    Fixes: 6cca686dfce7 ("drm/amdkfd: kfd driver supports hot unplug/replug amdgpu devices")
    Signed-off-by: Xiaogang Chen <xiaogang.chen@amd.com>
    Reviewed-by: Felix Kuehling <felix.kuehling@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit d12d05c4bc4c15585130af43e897923ff292df7b)

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:01 -05:00
Anusha Srivatsa 5706fd9ec7 dma-buf: fix UAF in dma_buf_put() tracepoint
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 2d76319c4cbb19eccfca71fa05d40a6b4ce7fc3d
Author: Andi Shyti <andi.shyti@kernel.org>
Date:   Wed Apr 8 14:39:15 2026 +0200

    dma-buf: fix UAF in dma_buf_put() tracepoint

    dma_buf_put() may drop the final file reference via fput(), which
    can free the dma-buf. The new tracepoint invocation was added
    after fput(), and DMA_BUF_TRACE() dereferences dmabuf and takes
    dmabuf->name_lock.

    This leads to a use-after-free on the final put, visible for
    example as a spinlock bad magic fault on a poisoned 0x6b6b6b...
    lock.

    Move the dma_buf_put tracepoint before fput().

    Reported-by: Janusz Krzysztofik <janusz.krzysztofik@linux.intel.com>
    Fixes: 281a22631423 ("dma-buf: add some tracepoints to debug.")
    Signed-off-by: Andi Shyti <andi.shyti@linux.intel.com>
    Reviewed-by: Christian König <christian.koenig@amd.com>
    Signed-off-by: Christian König <christian.koenig@amd.com>
    Link: https://lore.kernel.org/r/20260408123916.2604101-1-andi.shyti@kernel.org

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:01 -05:00
Anusha Srivatsa 267916ac99 drm/amd/pm: fix runtime PM imbalance issue in amdgpu_pm.c
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 25fd8095a868cfbeb9ef3118131d2ba1f7057846
Author: Yang Wang <kevinyang.wang@amd.com>
Date:   Thu Apr 16 18:17:30 2026 +0800

    drm/amd/pm: fix runtime PM imbalance issue in amdgpu_pm.c

    Fix runtime PM counter imbalance to prevent device from failing to enter low power state

    Fixes: a50d32c41fb2 ("drm/amd/pm: Deprecate print_clock_levels interface")
    Signed-off-by: Yang Wang <kevinyang.wang@amd.com>
    Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa e1356f0b72 drm/amdgpu: drop retry loop in amdgpu_hmm_range_get_pages
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit dd03ceed9fc71fa8f6c3ae44d395ac53ea2a0dd0
Author: Honglei Huang <honghuan@amd.com>
Date:   Fri May 29 10:23:17 2026 +0800

    drm/amdgpu: drop retry loop in amdgpu_hmm_range_get_pages

    commit 342981fff32802a819d6fc7cf3c9fedf9f3d9d60 upstream.

    Since commit c08972f55594 ("drm/amdgpu: fix amdgpu_hmm_range_get_pages")
    moved mmu_interval_read_begin() out of the per-chunk loop, the
    captured notifier_seq is no longer refreshed across retries. As a
    result, the existing -EBUSY retry path can never make progress:

      hmm_range_fault() returns -EBUSY only when
      mmu_interval_check_retry(notifier, notifier_seq) reports that the
      sequence is stale. Once the sequence has advanced, the stored seq
      will never match again, so every subsequent call within the same
      invocation returns -EBUSY immediately.

    The "goto retry" therefore degenerates into a busy spin that simply
    burns CPU for the full HMM_RANGE_DEFAULT_TIMEOUT (~1s) window before
    finally bailing out with -EAGAIN. This is pure latency with no chance
    of recovery, and it actively hurts the KFD userptr stack: the caller
    ends up blocked for a second while holding mmap_lock, only to return
    -EAGAIN to the restore worker (or to userspace) which would have
    re-driven the operation immediately anyway.

    Drop the retry/timeout entirely and let -EBUSY propagate straight to
    out_free_pfns, where it is already translated to -EAGAIN. Recovery is
    handled at a higher level: the KFD restore_userptr_worker reschedules
    itself, and the userptr ioctl path returns -EAGAIN to userspace.

    No functional regression: the previous behaviour on -EBUSY was already
    to fail with -EAGAIN after a 1s stall; we just skip the stall.

    Reviewed-by: Christian König <christian.koenig@amd.com>
    Signed-off-by: Honglei Huang <honghuan@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa 9f70900688 drm/amd/display: Use krealloc_array() in dal_vector_reserve()
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit de988c7a31f0774f07894cfe4802996f318e2870
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Tue May 5 11:52:15 2026 -0400

    drm/amd/display: Use krealloc_array() in dal_vector_reserve()

    commit da48bc4461b8a5ebfb9264c9b191a701d8e99009 upstream.

    [Why & How]
    dal_vector_reserve() computes the allocation size as
    "capacity * vector->struct_size" using uint32_t arithmetic, which can
    silently wrap to a small value on overflow. This would cause krealloc to
    return a smaller buffer than expected, leading to heap overflows on
    subsequent vector appends.

    Replace krealloc() with krealloc_array() which performs an internal
    overflow check and returns NULL on wrap, preventing the issue.

    Fixes: 2004f45ef8 ("drm/amd/display: Use kernel alloc/free")
    Assisted-by: Copilot:claude-opus-4.6
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ray Wu <ray.wu@amd.com>
    Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 37668568641ccc4cc1dbca4923d0a16609dd5707)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa 65b012a240 drm/amd/display: Fix out-of-bounds read in dp_get_eq_aux_rd_interval()
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit dc1490927d79fe9621e29f4a4f5d7b5ccb6aea3e
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Tue May 5 11:44:15 2026 -0400

    drm/amd/display: Fix out-of-bounds read in dp_get_eq_aux_rd_interval()

    commit e8b4d37eba05141ee01794fc6b7f2da808cee83b upstream.

    [Why & How]
    The aux_rd_interval array in struct dc_lttpr_caps is declared with
    MAX_REPEATER_CNT - 1 (7) elements, indexed 0..6. However, the offset
    parameter passed to dp_get_eq_aux_rd_interval() can be as large as
    MAX_REPEATER_CNT (8) when a sink reports 8 LTTPR repeaters via DPCD.
    This leads to an out-of-bounds read of aux_rd_interval[7] when offset
    is 8.

    Fix this by growing aux_rd_interval to MAX_REPEATER_CNT elements to
    accommodate the full range of valid repeater counts defined by the DP
    spec.

    Assisted-by: GitHub Copilot:Claude claude-4-opus
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ray Wu <ray.wu@amd.com>
    Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit a55a458a8df37a65ffda5cf721d554a8f74f6b04)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa b3fdcfa826 drm/amd/display: Fix NULL deref and buffer over-read in SDP debugfs
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit c90954cdea4d6998ec345de0d840d030c145b89e
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Mon May 11 16:46:25 2026 -0400

    drm/amd/display: Fix NULL deref and buffer over-read in SDP debugfs

    commit adf67034b1f61f7119295208085bfd43f85f56af upstream.

    [Why & How]
    dp_sdp_message_debugfs_write() dereferences connector->base.state->crtc
    without checking for NULL. A connector can be connected but not bound to
    any CRTC (e.g. after hot-plug before the next atomic commit), causing a
    kernel crash when writing to the sdp_message debugfs node.

    The function also ignores the user-provided size argument and always
    passes 36 bytes to copy_from_user(), reading past the user buffer when
    size < 36.

    Fix both issues by:
    - Returning -ENODEV when connector->base.state or state->crtc is NULL
    - Clamping write_size to min(size, sizeof(data))

    Fixes: c7ba3653e9 ("drm/amd/display: Generic SDP message access in amdgpu")
    Assisted-by: Copilot:claude-opus-4.6
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ray Wu <ray.wu@amd.com>
    Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 6ab4c36a522842ff70474a1c0af2e40e50fc8300)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa 8f8b12715d drm/amd/display: add missing CSC entries for BT.2020 for DCE IPs
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 98abb179dd513c1c56bf97532e5816d6e927876b
Author: Leorize <leorize+oss@disroot.org>
Date:   Wed May 27 23:58:54 2026 -0700

    drm/amd/display: add missing CSC entries for BT.2020 for DCE IPs

    commit 6590fe323ce2807f5d9454e7fccf3fab875d4352 upstream.

    DCE-based hardware does not have the CSC matrices for BT.2020, which
    causes the driver to fallback to the GPU built-in matrices. This does
    not appear to cause any issues for RGB sinks, but causes major color
    artifacts for YCbCr ones (e.g. black becomes green).

    This commit adds the missing CSC matrices (taken from DC common) to DCE
    CSC tables, resolving the issue.

    Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/3358
    Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5333
    Assisted-by: oh-my-pi:GPT-5.5
    Signed-off-by: Leorize <leorize+oss@disroot.org>
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 51e6668ab4baf55b082c376318d51ef965757196)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa d65e3a8016 drm/amd/display: Clamp VBIOS HDMI retimer register count to array size
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 8aaa7e317fbd4beb9c6a9f77aa4cf52fae78b117
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Mon May 4 15:51:13 2026 -0400

    drm/amd/display: Clamp VBIOS HDMI retimer register count to array size

    commit fb0707ce00eef4e2d60c3020e1c0432739703e4a upstream.

    [Why & How]
    The VBIOS integrated info tables (v1_11 and v2_1) contain HdmiRegNum and
    Hdmi6GRegNum fields that are used as loop bounds when copying retimer I2C
    register settings into fixed-size arrays (dp*_ext_hdmi_reg_settings[9]
    and dp*_ext_hdmi_6g_reg_settings[3]). These u8 fields are not validated
    before use, so a malformed VBIOS can specify values up to 255, causing an
    out-of-bounds heap write during driver probe.

    Clamp each register count to the destination array size using min_t()
    before the copy loops, in both get_integrated_info_v11() and
    get_integrated_info_v2_1().

    Assisted-by: GitHub Copilot:claude-opus-4.6
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ray Wu <ray.wu@amd.com>
    Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 5a7f0ef90195940c54b0f5bb85b87da55f038c69)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa c03a8d0157 drm/amd/display: Clamp HDMI HDCP2 rx_id_list read to buffer size
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 98cfb7530ea91d8e5e928285cdce58e1131f6e83
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Thu May 7 15:38:37 2026 -0400

    drm/amd/display: Clamp HDMI HDCP2 rx_id_list read to buffer size

    commit f0f3981c43b32cadfe373d636d9e9ca522bb3702 upstream.

    [Why & How]
    During HDCP 2.x repeater authentication over HDMI, the driver reads the
    sink's RxStatus register and extracts a 10-bit message size field (max
    value 1023). This value is used as the read length for the ReceiverID
    list without being clamped to the size of the destination buffer
    rx_id_list[177]. A malicious HDMI repeater could advertise a message
    size larger than the buffer, causing an out-of-bounds write during the
    I2C read.

    Clamp the read length in mod_hdcp_read_rx_id_list() to the size of the
    rx_id_list buffer, matching the approach already used in the DP branch.

    Fixes: eff682f83c ("drm/amd/display: Add DDC handles for HDCP2.2")
    Assisted-by: Copilot:claude-opus-4.6
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ray Wu <ray.wu@amd.com>
    Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 229212219e4247d9486f8ba41ef087358490be09)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa 3a3969fce3 drm/amd/display: Bound VBIOS record-chain walk loops
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 2645e3caf7e013189da9c6ff621d006cca5a538b
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Tue May 12 15:24:22 2026 -0400

    drm/amd/display: Bound VBIOS record-chain walk loops

    commit ff287df16a1a58aca78b08d1f3ee09fc44da0351 upstream.

    [Why & How]
    All record-chain walk loops in bios_parser.c and bios_parser2.c use
    for(;;) and only terminate on a 0xFF record_type sentinel or zero
    record_size. A malformed VBIOS image missing the terminator record
    causes unbounded iteration at probe time, potentially hundreds of
    thousands of iterations with record_size=1. In the final iterations
    near the BIOS image boundary, struct casts beyond the 2-byte header
    validated by GET_IMAGE can also read out of bounds.

    Cap all 14 record-chain walk loops to BIOS_MAX_NUM_RECORD (256)
    iterations. The atombios.h defines up to 22 distinct record types
    and atomfirmware.h has 13. Assuming an average of less than 10
    records per type (which is reasonable since most are connector-
    based) 256 is a generous upper bound.

    Fixes: 4562236b3b ("drm/amd/dc: Add dc display driver (v2)")
    Assisted-by: Copilot:claude-opus-4.6 Mythos
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ray Wu <ray.wu@amd.com>
    Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 95700a3d660287ed657d6892f7be9ffc0e294a93)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa 4c73676be3 drm/amd/pm: smu_v14_0_0: use SoftMin for gfxclk in set_soft_freq_limited_range
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 65c7bff058b359711cb7626969331ff2895d2362
Author: Priya Hosur <Priya.Hosur@amd.com>
Date:   Thu May 7 13:31:37 2026 +0530

    drm/amd/pm: smu_v14_0_0: use SoftMin for gfxclk in set_soft_freq_limited_range

    commit 03b70e0d8aa26bab89a0f1394c1c80a871925e42 upstream.

    In smu_v14_0_0_set_soft_freq_limited_range(), the gfxclk floor is
    programmed via SetHardMinGfxClk together with SetSoftMaxGfxClk. Under
    power_dpm_force_performance_level=high this pins HardMin to peak gfxclk.

    In PMFW arbitration HardMin has higher priority than SoftMax, so the
    firmware thermal/PPT throttler cannot clamp gfxclk via SoftMax once
    HardMin is set to peak. Replace SetHardMinGfxClk with SetSoftMinGfxclk
    so the driver still requests peak performance but the firmware
    throttler retains the ability to clamp gfxclk under thermal/PPT
    pressure. SoftMax handling is unchanged and no other clock domains
    are affected.

    Signed-off-by: Priya Hosur <Priya.Hosur@amd.com>
    Acked-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 3ea273267fd29cbf6d83ee72329f59eb5042605b)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:54:00 -05:00
Anusha Srivatsa 78918408e7 drm/amd/pm: mark metrics.energy_accumulator is invalid for smu 14.0.2
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 8f634181b538c4e2cfb9a559e9914b32d38ef540
Author: Yang Wang <kevinyang.wang@amd.com>
Date:   Fri May 29 11:47:31 2026 +0800

    drm/amd/pm: mark metrics.energy_accumulator is invalid for smu 14.0.2

    commit ee193c5bbd5e2b56bbeb54ef554414b43a6fc896 upstream.

    EnergyAccumulator is unsupported on SMU 14.0.2, mark it invalid.

    Signed-off-by: Yang Wang <kevinyang.wang@amd.com>
    Reviewed-by: Asad Kamal <asad.kamal@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 646b05043eeed04b51c14aad22a400a8250af4b7)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:59 -05:00
Anusha Srivatsa 00edd1e329 drm/amd/pm: fix smu13 power limit default/cap calculation
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit f02262d654f9ffdc83a75eb35f1d94766c52f6e2
Author: Yang Wang <kevinyang.wang@amd.com>
Date:   Tue May 19 11:18:12 2026 +0800

    drm/amd/pm: fix smu13 power limit default/cap calculation

    commit bb204f19e4a115f094a6a3c4d82fcf48862d0766 upstream.

    smu_v13_0_0_get_power_limit() and smu_v13_0_7_get_power_limit() mix
    runtime power_limit with PP table limits when reporting default/min/max.

    When current power limit query succeeds, default_power_limit was set to the
    runtime value instead of the PP table default, and min/max could be derived
    from inconsistent bases (MsgLimits/runtime), leading to incorrect cap info.

    Use SocketPowerLimitAc/Dc as the PP default base (pp_limit), keep
    current_power_limit as runtime value, and derive min/max from pp_limit with
    OD percentages.

    Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5227
    Signed-off-by: Yang Wang <kevinyang.wang@amd.com>
    Reviewed-by: Kenneth Feng <kenneth.feng@amd.com>
    Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 1eaf26db95901ca70737503a89b831dd763c8453)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:59 -05:00
Anusha Srivatsa 4947ddefa3 drm/amd/pm: apply SMU 13.0.10 workaround during MP1 unload
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 0fc1c46182fcaaf6eb396273782ca9b5d4f6e5bf
Author: Yang Wang <kevinyang.wang@amd.com>
Date:   Thu May 21 22:36:37 2026 +0800

    drm/amd/pm: apply SMU 13.0.10 workaround during MP1 unload

    commit 2493d87bb4c31ec9ca7f0ef7257e33b8b175f913 upstream.

    On SMU v13.0.10, sending PrepareMp1ForUnload with the default
    parameter may leave the device in an inaccessible state. This can
    affect runtime power management and partial PnP flows.
    e.g: kexec, driver unload, boco/d3cold.

    Pass the required workaround parameter 0x55, when preparing MP1 for
    unload on SMU v13.0.10, keep the existing behavior for other SMU
    versions.

    Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5133
    Signed-off-by: Yang Wang <kevinyang.wang@amd.com>
    Reviewed-by: Kenneth Feng <kenneth.feng@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 4e8ee1afeedb8d24dd22cdd5ae9f98a6d76ebe4b)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:59 -05:00
Anusha Srivatsa b3adaf3a41 drm/amdgpu: Fix incorrect VRAM GART mappings on non-4K page size systems
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 7e8fb27dcdb284c1b0c7a0613ea5bd5c340c12fe
Author: Donet Tom <donettom@linux.ibm.com>
Date:   Wed May 27 18:49:31 2026 +0530

    drm/amdgpu: Fix incorrect VRAM GART mappings on non-4K page size systems

    commit ec4c462e2d8161b32038e21e7187f4a15fe1661d upstream.

    When mapping VRAM pages into the GART page table,
    amdgpu_gart_map_vram_range() assumes that the system page size is the
    same as the GPU page size.

    On systems with non-4K page sizes, multiple GPU pages can exist within
    a single CPU page. As a result, the mappings are created incorrectly
    because fewer page table entries are programmed than required.

    Fix this by programming the mappings correctly for non-4K page size
    systems.

    Fixes: 237d623ae659 ("drm/amdgpu/gart: Add helper to bind VRAM pages (v2)")
    Reviewed-by: Timur Kristóf <timur.kristof@gmail.com>
    Reviewed-by: Christian König <christian.koenig@amd.com>
    Signed-off-by: Donet Tom <donettom@linux.ibm.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit a8f0bc22388f74e0cf4ed8b7d1846c580eaf44cc)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:59 -05:00
Anusha Srivatsa 41bc0bbf2a drm/amdgpu: set noretry=1 as default for GFX 10.1.x (Navi10/12/14)
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit fafd80c0ab641e1eabc93cafa78af5b04f78c753
Author: Vitaly Prosyak <vitaly.prosyak@amd.com>
Date:   Fri May 29 13:50:38 2026 -0400

    drm/amdgpu: set noretry=1 as default for GFX 10.1.x (Navi10/12/14)

    commit e47b0056a08dc70430ffc44bbf62197e7d1ff8ea upstream.

    Problem:
    While developing the amd_close_race IGT test (which intentionally triggers
    execute permission faults by removing VM_PAGE_EXECUTABLE from GPU page table
    entries), we discovered that on Navi10 (GFX 10.1.x) these faults produce
    zero diagnostic output. The GPU simply hangs silently for ~10s until the
    scheduler timeout fires. There is no way to distinguish an execute
    permission fault from any other type of GPU hang.

    Root cause:
    GFX 10.1.x defaults to noretry=0, which sets
    RETRY_PERMISSION_OR_INVALID_PAGE_FAULT=1 in the GFXHUB UTCL2 registers
    (gfxhub_v2_0.c line 313). With this bit set, permission faults (valid PTE,
    wrong R/W/X bits) are handled entirely within the UTCL1/UTCL2 hardware
    loop: UTCL2 returns an XNACK to UTCL1, and UTCL1 re-requests the
    translation indefinitely, expecting software to eventually fix the
    permission bits (as happens in SVM/HMM recovery). No interrupt of any kind
    reaches the IH ring.

    This is different from invalid-page faults (V=0) which DO generate a retry
    fault interrupt that the driver can escalate to a no-retry fault. Permission
    faults with valid PTEs loop silently forever in hardware.

    GFX 10.3+ already defaults to noretry=1, which makes permission faults
    generate immediate L2 protection fault interrupts. GFX 10.1.x was
    inadvertently left out of this default.

    Fix:
    Change the noretry=1 threshold from IP_VERSION(10, 3, 0) to
    IP_VERSION(10, 1, 0) in amdgpu_gmc_noretry_set(). This is a one-line
    change that aligns GFX 10.1.x behavior with GFX 10.3+ and all newer
    generations.

    With noretry=1, the existing non-retry fault handler
    (gmc_v10_0_process_interrupt) already decodes and prints the full
    GCVM_L2_PROTECTION_FAULT_STATUS register including PERMISSION_FAULTS,
    faulting address, VMID, PASID, and process name. No additional logging
    code is needed — the fix is purely routing permission faults to the
    existing, fully-capable non-retry interrupt handler.

    v2: Dropped GFX10-specific logging from gmc_v10_0.c and
    kfd_int_process_v10.c (Felix Kuehling). v1 added logging in the retry
    fault handler, but with noretry=1 permission faults take the non-retry
    path — the v1 retry handler code was dead and would never execute.

    Tested on Navi10 (GFX 10.1.10):
    - Execute permission faults now produce immediate, clear output:
        [gfxhub] page fault (src_id:0 ring:64 vmid:4 pasid:592)
         Process amd_close_race pid 13380 thread amd_close_race pid 13384
          in page at address 0x40001000 from client 0x1b (UTCL2)
        GCVM_L2_PROTECTION_FAULT_STATUS:0x00700881
             PERMISSION_FAULTS: 0x8
    - No regressions with properly-mapped GPU workloads

    Cc: Christian Koenig <christian.koenig@amd.com>
    Cc: Alex Deucher <alexander.deucher@amd.com>
    Cc: Felix Kuehling <felix.kuehling@amd.com>
    Signed-off-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
    Acked-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit eb21edd24c40d81066753f8ac6f23bce15745395)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:59 -05:00
Anusha Srivatsa 6ad09946d7 drm/amdgpu: restart the CS if some parts of the VM are still invalidated
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 1cab6864538592d29524a10e5d5945a2d0b66575
Author: Christian König <christian.koenig@amd.com>
Date:   Wed Feb 25 15:12:02 2026 +0100

    drm/amdgpu: restart the CS if some parts of the VM are still invalidated

    commit 40396ffdf6120e2380706c59e1a84d7e765a37b6 upstream.

    Make sure that we only submit work with full up to date VM page tables.

    Backport to 7.1 and older.

    Signed-off-by: Christian König <christian.koenig@amd.com>
    Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
    Tested-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
    Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 59720bfd8c6dbebeb8d5a7ab64241b007efd9213)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:59 -05:00
Anusha Srivatsa ce1a819a68 drm/amdgpu: fix waiting for all submissions for userptrs
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 6f25f07271a85956f43a7b971bc5130833db3311
Author: Christian König <christian.koenig@amd.com>
Date:   Wed Feb 18 13:05:46 2026 +0100

    drm/amdgpu: fix waiting for all submissions for userptrs

    commit 58bafc666c484b21839a2d27e923ae1b2727a1df upstream.

    Wait for all submissions when userptrs need to be invalidated by the MMU
    notifier, not just the one the userptr was involved into.

    Signed-off-by: Christian König <christian.koenig@amd.com>
    Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
    Tested-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
    Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 91250893cbaa25c86872deca95a540d08de1f91e)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:59 -05:00
Anusha Srivatsa c16f940b8c drm/xe: Clear pending_disable before signaling suspend fence
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 8bff5f60f0bd57322762c2c3ac7e4eb839035338
Author: Tangudu Tilak Tirumalesh <tilak.tirumalesh.tangudu@intel.com>
Date:   Wed Jun 3 12:22:16 2026 +0530

    drm/xe: Clear pending_disable before signaling suspend fence

    commit 54f2a0442a30fe7a0f6bc8345e81f8b2db8effbd upstream.

    In the schedule-disable done path for suspend, we
    signal the suspend fence before clearing pending_disable.

    That wakeup can let suspend_wait complete and resume be queued
    immediately. The resume path may then reach enable_scheduling()
    while pending_disable is still set and hit the
    !exec_queue_pending_disable(q) assertion.

    Fix this by clearing pending_disable before signaling
    the suspend fence, so any resumed transition observes a
    consistent state.

    Fixes: 87651f31ae4e ("drm/xe/guc_submit: fix race around suspend_pending")
    Cc: stable@vger.kernel.org # v7.0+
    Signed-off-by: Tangudu Tilak Tirumalesh <tilak.tirumalesh.tangudu@intel.com>
    Reviewed-by: Thomas Hellstrom <thomas.hellstrom@linux.intel.com>
    Signed-off-by: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
    Link: https://patch.msgid.link/20260603065217.3131066-3-tilak.tirumalesh.tangudu@intel.com
    (cherry picked from commit 4b1ae138b0e103d753773956a84eebc2edbf62c4)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:59 -05:00
Anusha Srivatsa 22ca32df32 drm/xe/multi_queue: skip submit when primary queue is suspended
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 91cffc74c61cf5cf5733d68f8058f55522e7e4f4
Author: Niranjana Vishwanathapura <niranjana.vishwanathapura@intel.com>
Date:   Wed Jun 3 16:39:47 2026 -0700

    drm/xe/multi_queue: skip submit when primary queue is suspended

    commit ec4cbdd163f9bb2a2bd44eb93ecf4a2fa0e912a9 upstream.

    Return early in submit path when the multi-queue primary exec
    queue is suspended to avoid submitting while suspended.

    v2: Remove idle_skip_suspend fix as that feature is being
    reverted here https://patchwork.freedesktop.org/series/167262/

    Fixes: bc5775c59258 ("drm/xe/multi_queue: Add GuC interface for multi queue support")
    Cc: stable@vger.kernel.org # v7.0+
    Assisted-by: GitHub-Copilot:claude-sonnet-4.6
    Reviewed-by: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
    Signed-off-by: Niranjana Vishwanathapura <niranjana.vishwanathapura@intel.com>
    Link: https://patch.msgid.link/20260603233946.863663-2-niranjana.vishwanathapura@intel.com
    (cherry picked from commit b7fb55cc3364ca128cfff9d50649ffd4327cd01e)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:58 -05:00
Anusha Srivatsa 7a86ffb50e drm/xe/display: fix oops in suspend/shutdown without display
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 238bcdaae8f2abc65e182de7d1f69cf8f611a610
Author: Jani Nikula <jani.nikula@intel.com>
Date:   Fri May 15 19:09:20 2026 +0300

    drm/xe/display: fix oops in suspend/shutdown without display

    commit 68938cc08e23a94fd881e845837ff918de005ce7 upstream.

    The xe driver keeps track of whether to probe display, and whether
    display hardware is there, using xe->info.probe_display. It gets set to
    false if there's no display after intel_display_device_probe(). However,
    the display may also be disabled via fuses, detected at a later time in
    intel_display_device_info_runtime_init().

    In this case, the xe driver does for_each_intel_crtc() on uninitialized
    mode config in xe_display_flush_cleanup_work(), leading to a NULL
    pointer dereference, and generally calls display code with display info
    cleared.

    Check for intel_display_device_present() after
    intel_display_device_info_runtime_init(), and reset
    xe->info.probe_display as necessary. Also do unset_display_features()
    for completeness, although display runtime init has already done
    that. This will need to be unified across all cases later.

    Move intel_display_device_info_runtime_init() call slightly earlier,
    similar to i915, to avoid a bunch of unnecessary setup for no display
    cases.

    Note #1: The xe driver has no business doing low level display plumbing
    like for_each_intel_crtc() to begin with. It all needs to happen in
    display code.

    Note #2: The actual bug is present already in commit 44e694958b95
    ("drm/xe/display: Implement display support"), but the oops was likely
    introduced later at commit ddf6492e0e50 ("drm/xe/display: Make display
    suspend/resume work on discrete").

    Fixes: 44e694958b95 ("drm/xe/display: Implement display support")
    Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/7904
    Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/6150
    Cc: stable@vger.kernel.org # v6.8+
    Reviewed-by: Suraj Kandpal <suraj.kandpal@intel.com>
    Link: https://patch.msgid.link/20260515160920.1082842-1-jani.nikula@intel.com
    Signed-off-by: Jani Nikula <jani.nikula@intel.com>
    (cherry picked from commit 7c3eb9f47533220888a67266448185fd0775d4da)
    Signed-off-by: Matthew Brost <matthew.brost@intel.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:58 -05:00
Anusha Srivatsa efc876e84b drm/amdkfd: Fix buffer overflow in SDMA queue checkpoint/restore on GFX11
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit d02f05d30f35b036f7cbaf72de634affb5b38ec6
Author: Andrew Martin <andrew.martin@amd.com>
Date:   Thu May 28 12:54:39 2026 -0400

    drm/amdkfd: Fix buffer overflow in SDMA queue checkpoint/restore on GFX11

    commit 352ea59028ea48a6fff77f19ae28f98f71946a80 upstream.

    The v11 MQD manager incorrectly assigned the CP-compute variants of
    checkpoint_mqd/restore_mqd for KFD_MQD_TYPE_SDMA queues. These functions
    use sizeof(struct v11_compute_mqd) (2048 bytes) instead of sizeof(struct
    v11_sdma_mqd) (512 bytes), causing a 1536-byte overflow.

    During CRIU checkpoint of an SDMA queue on Navi3x:
    - checkpoint_mqd() reads 2048 bytes from a 512-byte SDMA MQD buffer,
      leaking 1536 bytes of adjacent GTT memory to userspace

    During CRIU restore:
    - restore_mqd() writes 2048 bytes into a 512-byte SDMA MQD buffer,
      corrupting 1536 bytes of adjacent GTT memory (often the ring buffer
      or neighboring MQDs)

    This is a copy-paste regression unique to v11. All other ASIC backends
    (cik, vi, v9, v10, v12) correctly use the SDMA-specific variants.

    Add checkpoint_mqd_sdma() and restore_mqd_sdma() functions that properly
    handle the smaller v11_sdma_mqd structure, matching the pattern used in
    other MQD managers.

    Fixes: cc009e613de6 ("drm/amdkfd: Add KFD support for soc21 v3")
    Assisted-by: Claude:Sonnet 4-5
    Signed-off-by: Andrew Martin <andrew.martin@amd.com>
    Acked-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 6fa41db7ffdec97d62433adf03b7b9b759af8c2c)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:58 -05:00
Anusha Srivatsa d4f8bb8883 drm/amdkfd: fix NULL dereference in get_queue_ids()
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit e1965e8913cfbf17622ca12638e7a07f68ba0848
Author: Muhammad Bilal <meatuni001@gmail.com>
Date:   Sat May 23 16:56:46 2026 +0000

    drm/amdkfd: fix NULL dereference in get_queue_ids()

    commit 2bd550b547deabef98bd3b017ff743b7c34d3a6d upstream.

    When usr_queue_id_array is NULL and num_queues is non-zero,
    get_queue_ids() returns NULL. The callers check only IS_ERR() on the
    return value; since IS_ERR(NULL) == false the check passes, and
    suspend_queues() calls q_array_invalidate() which immediately
    dereferences NULL while iterating num_queues times.

    Userspace can trigger this via kfd_ioctl_set_debug_trap() by supplying
    num_queues > 0 with a zero queue_array_ptr, causing a kernel panic.

    A NULL usr_queue_id_array with num_queues == 0 is a legitimate no-op
    (q_array_invalidate never executes, and resume_queues already guards
    all queue_ids dereferences behind a NULL check). Return ERR_PTR(-EINVAL)
    only when num_queues is non-zero and the pointer is absent; both callers
    already propagate IS_ERR() returns correctly to userspace.

    Fixes: a70a93fa568b ("drm/amdkfd: add debug suspend and resume process queues operation")
    Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit f165a82cdf503884bb1797771c61b2fcc72113d4)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:58 -05:00
Anusha Srivatsa 7b6856d598 drm/i915: Fix color blob reference handling in intel_plane_state
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit da8f42b770a36018560e1f6edf6a83e2c7949b8d
Author: Chaitanya Kumar Borah <chaitanya.kumar.borah@intel.com>
Date:   Mon Jun 1 13:59:53 2026 +0530

    drm/i915: Fix color blob reference handling in intel_plane_state

    commit 26eb7c0a7ab09d83eec833db6a5a2bc60b9d4d9a upstream.

    Take proper references for hw color blobs (degamma_lut, gamma_lut,
    ctm, lut_3d) in intel_plane_duplicate_state() and drop them in
    intel_plane_destroy_state().

    v2:
    - handle blobs in hw state clear

    Cc: <stable@vger.kernel.org> #v6.19+
    Fixes: 3b7476e786c2 ("drm/i915/color: Add framework to program PRE/POST CSC LUT")
    Fixes: a78f1b6baf4d ("drm/i915/color: Add framework to program CSC")
    Fixes: 65db7a1f9cf7 ("drm/i915/color: Add 3D LUT to color pipeline")
    Reviewed-by: Pranay Samala <pranay.samala@intel.com> #v1
    Reviewed-by: Uma Shankar <uma.shankar@intel.com>
    Signed-off-by: Chaitanya Kumar Borah <chaitanya.kumar.borah@intel.com>
    Signed-off-by: Uma Shankar <uma.shankar@intel.com>
    Link: https://patch.msgid.link/20260601082953.128539-4-chaitanya.kumar.borah@intel.com
    (cherry picked from commit c6eea1925154b6697fe22b217faab9bb30635e6b)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:58 -05:00
Anusha Srivatsa c6dd792122 drm/gem: Try to fix change_handle ioctl, attempt 4
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 1d9b93df7fc768228906e24220591ec1cddad391
Author: Simona Vetter <simona.vetter@ffwll.ch>
Date:   Thu Jun 4 21:44:37 2026 +0200

    drm/gem: Try to fix change_handle ioctl, attempt 4

    commit 1a4f03d22fb655e5f192244fb2c87d8066fcfca2 upstream.

    [airlied: just added some comments on how to reenable]
    On-list because the cat is out of the bag and we're clearly not good
    enough to figure this out in private. The story thus far:

    5e28b7b94408 ("drm: Set old handle to NULL before prime swap in
    change_handle") tried to fix a race condition between the gem_close and
    gem_change_handle ioctls, but got a few things wrong:

    - There's a confusion with the local variable handle, which is actually
      the new handle, and so the two-stage trick was actually applied to the
      wrong idr slot. 7164d78559b0 ("drm/gem: fix race between
      change_handle and handle_delete") tried to fix that by adding yet
      another code block, but forgot to add the error handling. Which meant
      we now have two paths, both kinda wrong.

    - dc366607c41c ("drm: Replace old pointer to new idr") tried to apply
      another fix, but inconsistently, again because of the handle confusion
      - this would be the right fix (kinda, somewhat, it's a mess) if we'd
      do the two-stage approach for the new handle. Except that wasn't the
      intent of the original fix.

    We also didn't have an igt merged for the original ioctl, which is a big
    no-go. This was attempted to address off-list in the original bugfix,
    and amd QA people claimed the bug was fixed now. Very clearly that's not
    the case. Here's my attempt to sort this out:

    - Rename the local variable to new_handle, the old aliasing with
      args->handle is just too dangerously confusing.

    - Merge the gem obj lookup with the two-stage idr_replace so that we
      avoid getting ourselves confused there.

    - This means we don't have a surplus temporary reference anymore, only
      an inherited from the idr. A concurrent gem_close on the new_handle
      could steal that. Fix that with the same two-stage approach
      create_tail uses. This is a bit overkill as documented in the comment,
      but I also don't trust my ability to understand this all correctly, so
      go with the established pattern we have from other ioctls instead for
      maximum paranoia.

    - Adjust error paths. I've tried to make the error and success paths
      common, because they are identical except for which handle is removed
      and on which we call idr_replace to (re)install the object again. But
      that made things messier to read, so I've left it at the more verbose
      version, which unfortunately hides the symmetry in the entire code
      flow a bit.

    - While at it, also replace the 7 space indent with 1 tab.

    And finally, because I flat out don't trust my abilities here at all
    anymore:

    - Disable the ioctl until we have the igt situation and everything else
      sorted out on-list and with full consensus.

    v2:

    Sashiko noticed that I didn't handle the error path for idr_replace
    correctly, it must be checked with IS_ERR_OR_NULL like in
    gem_handle_delete. So yeah, definitely should just the existing paths
    1:1 because this is endless amounts of tricky.

    Also add the Fixes: line for the original ioctl, I forgot that too.

    Reported-by: DARKNAVY (@DarkNavyOrg) <vr@darknavy.com>
    Signed-off-by: Simona Vetter <simona.vetter@ffwll.ch>
    Fixes: dc366607c41c ("drm: Replace old pointer to new idr")
    Cc: syzbot+d7c9eed171647e421013@syzkaller.appspotmail.com
    Cc: stable@vger.kernel.org
    Cc: Edward Adam Davis <eadavis@qq.com>
    Cc: Dave Airlie <airlied@redhat.com>
    Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com>
    Cc: Maxime Ripard <mripard@kernel.org>
    Cc: Thomas Zimmermann <tzimmermann@suse.de>
    Fixes: 5e28b7b94408 ("drm: Set old handle to NULL before prime swap in change_handle")
    Cc: David Francis <David.Francis@amd.com>
    Cc: Puttimet Thammasaeng <pwn8official@gmail.com>
    Cc: Christian Koenig <Christian.Koenig@amd.com>
    Fixes: 7164d78559b0 ("drm/gem: fix race between change_handle and handle_delete")
    Cc: Zhenghang Xiao <kipreyyy@gmail.com>
    Fixes: 5e28b7b94408 ("drm: Set old handle to NULL before prime swap in change_handle")
    Reviewed-by: David Francis <David.Francis@amd.com>
    Signed-off-by: Dave Airlie <airlied@redhat.com>
    Link: https://patch.msgid.link/20260604194437.1725314-1-simona.vetter@ffwll.ch
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:58 -05:00
Anusha Srivatsa c004703be0 accel/ethosu: reject NPU_OP_RESIZE commands from userspace
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 70090a32f56a4589e7e860e0f9a8fbe4417df0a1
Author: Muhammad Bilal <meatuni001@gmail.com>
Date:   Sat May 23 21:07:52 2026 +0000

    accel/ethosu: reject NPU_OP_RESIZE commands from userspace

    commit ef911805d86a05363d3ec2fa9835a41def83bb7e upstream.

    NPU_OP_RESIZE is a U85-only command that the driver does not yet
    implement. The existing WARN_ON(1) placeholder fires unconditionally
    whenever userspace submits this command via DRM_IOCTL_ETHOSU_GEM_CREATE,
    causing unbounded kernel log spam.

    If panic_on_warn is set the kernel panics, giving any unprivileged user
    with access to the DRM device a trivial denial-of-service primitive.

    Replace the WARN_ON(1) with an explicit -EINVAL return so the ioctl
    rejects the command before it reaches hardware.

    Fixes: 5a5e9c0228e6 ("accel: Add Arm Ethos-U NPU driver")
    Cc: stable@vger.kernel.org
    Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
    Link: https://patch.msgid.link/20260523210840.92039-2-meatuni001@gmail.com
    Signed-off-by: Rob Herring (Arm) <robh@kernel.org>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:58 -05:00
Anusha Srivatsa 4de57cdf2c accel/ethosu: reject DMA commands with uninitialized length
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit fb25c76a820ca8a547aa478bfb503da0a11494ab
Author: Muhammad Bilal <meatuni001@gmail.com>
Date:   Sun May 24 13:03:19 2026 +0000

    accel/ethosu: reject DMA commands with uninitialized length

    commit d9d021218162b6c4fe0bdf42b2b340f1aae23a12 upstream.

    cmd_state_init() initializes the command state with memset(0xff),
    leaving dma->len at U64_MAX to signal missing setup. The only setter
    is NPU_SET_DMA0_LEN; if userspace omits this command and issues
    NPU_OP_DMA_START, dma->len remains U64_MAX.

    In dma_length(), a positive stride added to U64_MAX wraps to a small
    value. With size0 == 1, check_mul_overflow() does not trigger and
    dma_length() returns 0 instead of U64_MAX. The caller's U64_MAX check
    then passes, region_size[] stays 0, and the bounds check in
    ethosu_job.c is bypassed, allowing hardware to execute DMA with stale
    physical addresses.

    Fix by checking for U64_MAX at the start of dma_length() before any
    arithmetic, consistent with the sentinel value used throughout the
    driver to detect uninitialized fields.

    Fixes: 5a5e9c0228e6 ("accel: Add Arm Ethos-U NPU driver")
    Cc: stable@vger.kernel.org
    Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
    Link: https://patch.msgid.link/20260524130319.12747-1-meatuni001@gmail.com
    Signed-off-by: Rob Herring (Arm) <robh@kernel.org>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:58 -05:00
Anusha Srivatsa 2c3e9e39ea accel/ethosu: fix arithmetic issues in dma_length()
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 6bb73845d1855ceaf50e397175e5979a7bdf69bc
Author: Muhammad Bilal <meatuni001@gmail.com>
Date:   Sun May 24 10:37:10 2026 +0000

    accel/ethosu: fix arithmetic issues in dma_length()

    commit ee6d9b6e51626f259c6f0e38d94f91be4fd14754 upstream.

    dma_length() derives DMA region usage from command stream values and
    updates region_size[]:

        len = ((len + stride[0]) * size0 + stride[1]) * size1
        region_size[region] = max(..., len + dma->offset)

    Several arithmetic issues can corrupt the derived region size:

    - signed stride values may underflow when added to len
    - intermediate multiplications may overflow
    - len + dma->offset may overflow during region_size updates
    - dma_length() error returns were not validated by the caller

    region_size[] is later used by ethosu_job.c to validate command stream
    accesses against GEM buffer sizes. Arithmetic wraparound can therefore
    under-report region usage and bypass the bounds validation.

    Fix by validating signed additions, using overflow helpers for
    multiplications and offset updates, and propagating dma_length()
    failures to the caller.

    Fixes: 5a5e9c0228e6 ("accel: Add Arm Ethos-U NPU driver")
    Cc: stable@vger.kernel.org
    Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
    Link: https://patch.msgid.link/20260524103710.47397-1-meatuni001@gmail.com
    Signed-off-by: Rob Herring (Arm) <robh@kernel.org>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:58 -05:00
Anusha Srivatsa d4f975bd7b accel/ethosu: fix wrong weight index in NPU_SET_SCALE1_LENGTH on U85
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 0c50186a1554b3fe512f387d2c1f4e63def9a9e6
Author: Muhammad Bilal <meatuni001@gmail.com>
Date:   Sat May 23 21:07:53 2026 +0000

    accel/ethosu: fix wrong weight index in NPU_SET_SCALE1_LENGTH on U85

    commit e703843f242b28e35ac79408de571ae110c740b5 upstream.

    On non-U65 hardware (e.g. U85), opcode 0x4093 is NPU_SET_WEIGHT2_LENGTH.
    The BASE handler for the same opcode correctly assigns to
    st.weight[2].base, but the LENGTH handler mistakenly assigns cmds[1]
    to st.weight[1].length instead of st.weight[2].length.

    This leaves weight[2].length at its initialised sentinel value of
    0xffffffff and corrupts weight[1].length with the user-supplied value,
    breaking the software bounds-check state for both weight buffers on U85.

    Fix the index to match the BASE handler.

    Fixes: 5a5e9c0228e6 ("accel: Add Arm Ethos-U NPU driver")
    Cc: stable@vger.kernel.org
    Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
    Link: https://patch.msgid.link/20260523210840.92039-3-meatuni001@gmail.com
    Signed-off-by: Rob Herring (Arm) <robh@kernel.org>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:57 -05:00
Anusha Srivatsa 3f927c72fd accel/ethosu: fix IFM region index out-of-bounds in command stream parser
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit ee7bed779def61ebff1b92b0e851f412176fa416
Author: Muhammad Bilal <meatuni001@gmail.com>
Date:   Sat May 23 19:51:59 2026 +0000

    accel/ethosu: fix IFM region index out-of-bounds in command stream parser

    commit 00f547e0dfecf83014fb32bcba587c6b684c1362 upstream.

    NPU_SET_IFM_REGION extracts the region index with param & 0x7f, giving
    a maximum value of 127. However region_size[] and output_region[] in
    struct ethosu_validated_cmdstream_info are both sized to
    NPU_BASEP_REGION_MAX (8), giving valid indices [0..7].

    Every other region assignment in the same switch uses param & 0x7:
      NPU_SET_OFM_REGION:  st.ofm.region  = param & 0x7;
      NPU_SET_IFM2_REGION: st.ifm2.region = param & 0x7;
      NPU_SET_WEIGHT_REGION: st.weight[0].region = param & 0x7;
      NPU_SET_SCALE_REGION:  st.scale[0].region  = param & 0x7;

    The 0x7f mask on IFM is inconsistent and appears to be a typo.

    feat_matrix_length() and calc_sizes() use the region index directly
    as an array subscript into the kzalloc'd info struct:
      info->region_size[fm->region] = max(...);

    A userspace caller supplying NPU_SET_IFM_REGION with param > 7 causes
    a write up to 127*8 = 1016 bytes past the start of region_size[],
    corrupting adjacent kernel heap data.

    Fix by applying the same & 0x7 mask used by all other region
    assignments.

    Fixes: 5a5e9c0228e6 ("accel: Add Arm Ethos-U NPU driver")
    Cc: stable@vger.kernel.org
    Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
    Link: https://patch.msgid.link/20260523195159.55801-1-meatuni001@gmail.com
    Signed-off-by: Rob Herring (Arm) <robh@kernel.org>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:57 -05:00
Anusha Srivatsa 96514e9bac accel/ethosu: fix OOB write in ethosu_gem_cmdstream_copy_and_validate()
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit db6cb3e35cebf487f9a78ebd4cfa4b83708ff40d
Author: Muhammad Bilal <meatuni001@gmail.com>
Date:   Sat May 23 19:08:43 2026 +0000

    accel/ethosu: fix OOB write in ethosu_gem_cmdstream_copy_and_validate()

    commit c0837b9cf6eabbad8b8cbddaff1a46a6d0a2e29d upstream.

    The command stream parsing loop increments the index variable a second
    time when a 64-bit command word is encountered (bit 14 set), but does
    not re-check the loop bound before writing the second word:

        for (i = 0; i < size / 4; i++) {
            bocmds[i] = cmds[0];
            if (cmd & 0x4000) {
                i++;
                bocmds[i] = cmds[1];   /* unchecked */
            }
        }

    The buffer bocmds is backed by a DMA allocation of exactly size bytes
    from drm_gem_dma_create(ddev, size), giving valid indices [0, size/4-1].

    When i == size/4 - 1 on entry to an iteration and bit 14 of cmds[0] is
    set, bocmds[size/4-1] is written in bounds, i is then incremented to
    size/4, and bocmds[size/4] writes four bytes past the end of the
    allocation.

    Userspace controls both the buffer contents and the size argument via
    the ioctl, making this a userspace-triggerable heap out-of-bounds write.

    Fix by checking the incremented index against the buffer bound before
    the second write and returning -EINVAL if the buffer is too small to
    contain the extended command.

    Fixes: 5a5e9c0228e6 ("accel: Add Arm Ethos-U NPU driver")
    Cc: stable@vger.kernel.org
    Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
    Link: https://patch.msgid.link/20260523190843.33977-1-meatuni001@gmail.com
    Signed-off-by: Rob Herring (Arm) <robh@kernel.org>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:57 -05:00
Anusha Srivatsa 3253c878f4 drm/amd/display: Reject gpio_bitshift >= 32 in bios_parser_get_gpio_pin_info()
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 9b0104625451d3931acceabd7c74155764bccf38
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Tue May 5 11:50:07 2026 -0400

    drm/amd/display: Reject gpio_bitshift >= 32 in bios_parser_get_gpio_pin_info()

    commit 49c3da65961fe9857c831d47fa1989084e87514a upstream.

    [Why & How]
    gpio_bitshift is a uint8_t read directly from the VBIOS GPIO pin table.
    If the value is >= 32, the expression "1 << gpio_bitshift" triggers
    undefined behaviour in C (shift count exceeds type width). On x86 the
    shift is silently masked to 5 bits, producing an incorrect GPIO mask
    that may cause wrong MMIO register bits to be toggled.

    Validate gpio_bitshift before use and return BP_RESULT_BADBIOSTABLE for
    out-of-range values.

    Fixes: ae79c310b1 ("drm/amd/display: Add DCE12 bios parser support")
    Assisted-by: Copilot:claude-opus-4.6
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ray Wu <ray.wu@amd.com>
    Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit eadf438ab8d370b9d19acee9359918c85afeb80d)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:57 -05:00
Anusha Srivatsa 3fb231ff89 drm/virtio: fix dma_fence refcount leak on error in virtio_gpu_dma_fence_wait()
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit c0fffc874c264292e769f26194a2a5e66ce31810
Author: Wentao Liang <vulab@iscas.ac.cn>
Date:   Sun Jun 7 09:03:03 2026 +0000

    drm/virtio: fix dma_fence refcount leak on error in virtio_gpu_dma_fence_wait()

    commit 3f26bb732cc136ab20176697c92f32c9c84cb125 upstream.

    dma_fence_unwrap_for_each() internally calls dma_fence_unwrap_first()
    which does cursor->chain = dma_fence_get(head), taking an extra
    reference. On normal loop completion, dma_fence_unwrap_next()
    releases this via dma_fence_chain_walk() -> dma_fence_put().

    When virtio_gpu_do_fence_wait() fails and the function returns early
    from inside the loop, the cursor->chain reference is never released.
    This is the only caller in the entire kernel that does an early return
    inside dma_fence_unwrap_for_each.

    Add dma_fence_put(itr.chain) before the early return.

    Cc: stable@vger.kernel.org
    Fixes: eba57fb5498f ("drm/virtio: Wait for each dma-fence of in-fence array individually")
    Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
    Reviewed-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
    Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
    Link: https://patch.msgid.link/20260607090303.92423-1-vulab@iscas.ac.cn
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:57 -05:00
Anusha Srivatsa bfefbc32c7 drm/i915/gem: Fix phys BO pread/pwrite with offset
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit dd51a2eeb93bc6faa892ff9083911dd23f82c187
Author: Joonas Lahtinen <joonas.lahtinen@linux.intel.com>
Date:   Wed Jun 10 09:03:14 2026 +0300

    drm/i915/gem: Fix phys BO pread/pwrite with offset

    commit d21ad938398bca695a511307de38a65889e3b354 upstream.

    sg_page() returns struct page pointer not (void *) so the scaling
    of pread/pwrite is wrong for phys BO and wrong parts of BO would be
    accessed if non-zero offset is used.

    Last impacted platform with overlay or cursor planes using phys
    mapping was Gen3/945G/Lakeport.

    Reported-by: Matthew Wilcox (Oracle) <willy@infradead.org>
    Fixes: c6790dc223 ("drm/i915: Wean off drm_pci_alloc/drm_pci_free")
    Cc: <stable@vger.kernel.org> # v4.5+
    Cc: Tvrtko Ursulin <tursulin@ursulin.net>
    Cc: Simona Vetter <simona@ffwll.ch>
    Cc: Jani Nikula <jani.nikula@linux.intel.com>
    Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Joonas Lahtinen <joonas.lahtinen@linux.intel.com>
    Reviewed-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
    Link: https://patch.msgid.link/20260610060314.26111-1-joonas.lahtinen@linux.intel.com
    (cherry picked from commit 3e49a2f85070b2fb672c1e0fdba281a4ea3aebe6)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:57 -05:00
Anusha Srivatsa dc4df36c41 Revert "drm/xe: Skip exec queue schedule toggle if queue is idle during suspend"
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit b69b715f48ac7e802c89ed5924795c5b055da91e
Author: Tangudu Tilak Tirumalesh <tilak.tirumalesh.tangudu@intel.com>
Date:   Wed Jun 3 12:22:15 2026 +0530

    Revert "drm/xe: Skip exec queue schedule toggle if queue is idle during suspend"

    commit fa7c84726dc217ce0c183926ef9411636c7a2213 upstream.

    This reverts commit 8533051ce92015e9cc6f75e0d52119b9d91610b6.

    The idle-skip optimization bypasses GuC suspend, so the GPU may not
    perform the context switch that flushes TLB entries for invalidated
    userptr VMAs. In LR/preempt-fence VM mode, this can lead to missed TLB
    invalidation and page faults during userptr invalidation tests.

    Restore unconditional schedule toggling on suspend so the context-switch
    TLB flush is always performed.

    This optimization will be reintroduced with a fix that does not skip
    suspend in LR/preempt-fence VM mode.

    Fixes: 8533051ce920 ("drm/xe: Skip exec queue schedule toggle if queue is idle during suspend")
    Cc: stable@vger.kernel.org # v7.0+
    Suggested-by: Thomas Hellstrom <thomas.hellstrom@linux.intel.com>
    Signed-off-by: Tangudu Tilak Tirumalesh <tilak.tirumalesh.tangudu@intel.com>
    Reviewed-by: Thomas Hellstrom <thomas.hellstrom@linux.intel.com>
    Signed-off-by: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
    Link: https://patch.msgid.link/20260603065217.3131066-2-tilak.tirumalesh.tangudu@intel.com
    (cherry picked from commit 6a1e7934d9a6cf46aecae00a99c2603d1295e170)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:57 -05:00
Anusha Srivatsa 7ba859dd71 accel/ivpu: Fix signed integer truncation in IPC receive
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 2821bf2b79e47f87e1dbdd9d25c78240965a97d6
Author: Andrzej Kacprowski <andrzej.kacprowski@linux.intel.com>
Date:   Mon Jun 1 18:16:43 2026 +0200

    accel/ivpu: Fix signed integer truncation in IPC receive

    commit d9faef564438d1e4579c692c046603e7ada7bdf4 upstream.

    Fix potential buffer overflow where firmware-supplied data_size is cast
    to signed int before being used in min_t(). Large unsigned values
    (>= 0x80000000) become negative, causing unsigned wraparound and
    oversized memcpy operations that can overflow the stack buffer.

    Change min_t(int, ...) to min() as both values are unsigned and can be
    handled by min() without explicit cast.

    Fixes: 3b434a3445ff ("accel/ivpu: Use threaded IRQ to handle JOB done messages")
    Cc: stable@vger.kernel.org # v6.12+
    Signed-off-by: Andrzej Kacprowski <andrzej.kacprowski@linux.intel.com>
    Reviewed-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Signed-off-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Link: https://patch.msgid.link/20260601161643.229342-1-andrzej.kacprowski@linux.intel.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:57 -05:00
Anusha Srivatsa 6ecc80f6ad accel/ivpu: Add buffer overflow check in MS get_info_ioctl
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 4e5047cc94bea1cc7b670b7f503358e9af0542df
Author: Andrzej Kacprowski <andrzej.kacprowski@linux.intel.com>
Date:   Fri May 29 14:08:41 2026 +0200

    accel/ivpu: Add buffer overflow check in MS get_info_ioctl

    commit fb176425837693f50c5c9fc8db6fbb04af22bd0a upstream.

    Add validation that the info size returned from the metric stream info
    query is not exceeded when checked against the allocated buffer size.
    If the firmware returns a size larger than the buffer, reject the
    operation with -EOVERFLOW instead of proceeding with an incorrect
    buffer copy.

    Fixes: cdfad4db7756 ("accel/ivpu: Add NPU profiling support")
    Cc: stable@vger.kernel.org # v6.18+
    Signed-off-by: Andrzej Kacprowski <andrzej.kacprowski@linux.intel.com>
    Reviewed-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Signed-off-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Link: https://patch.msgid.link/20260529120841.135852-1-andrzej.kacprowski@linux.intel.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:57 -05:00
Anusha Srivatsa 0d420116fc accel/ivpu: Add bounds checks for firmware log indices
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 535da9ad8420c3b686a642403d4147ff220255fd
Author: Andrzej Kacprowski <andrzej.kacprowski@linux.intel.com>
Date:   Fri May 29 13:58:42 2026 +0200

    accel/ivpu: Add bounds checks for firmware log indices

    commit dd1311bcf0e62f0c515115f46a3813370f4a4bb1 upstream.

    Add validation that read and write indices in the firmware log buffer
    are within valid bounds (< data_size) before using them. If
    out-of-bounds indices are encountered (from firmware), clamp them to
    safe values instead of proceeding with invalid offsets.

    This prevents potential out-of-bounds buffer access when firmware
    supplies invalid log indices.

    Fixes: 1fc1251149a7 ("accel/ivpu: Refactor functions in ivpu_fw_log.c")
    Cc: stable@vger.kernel.org # v6.18+
    Signed-off-by: Andrzej Kacprowski <andrzej.kacprowski@linux.intel.com>
    Reviewed-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Signed-off-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Link: https://patch.msgid.link/20260529115842.135378-1-andrzej.kacprowski@linux.intel.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:56 -05:00
Anusha Srivatsa 047a99b1d6 accel/ivpu: Add bounds check for firmware runtime memory
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit f8ab60ae9309e76d9a09c601c10cc222e25b3d5b
Author: Andrzej Kacprowski <andrzej.kacprowski@linux.intel.com>
Date:   Fri May 29 14:08:53 2026 +0200

    accel/ivpu: Add bounds check for firmware runtime memory

    commit 1d0b597facdd3c0239c88e8797c1014e1ea0ef15 upstream.

    Validate that the firmware runtime memory specified in the image
    header is properly aligned and sized to hold the firmware image.
    This prevents errors during memory allocation and image transfer.

    Fixes: 2007e210b6a1 ("accel/ivpu: Split FW runtime and global memory buffers")
    Cc: stable@vger.kernel.org # v7.0+
    Signed-off-by: Andrzej Kacprowski <andrzej.kacprowski@linux.intel.com>
    Reviewed-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Signed-off-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Link: https://patch.msgid.link/20260529120853.135876-1-andrzej.kacprowski@linux.intel.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:56 -05:00
Anusha Srivatsa 65ccf84c95 Revert "drm/xe/nvls: Define GuC firmware for NVL-S"
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 9c2a9c6f930f8220cedb16578604ca2a543c616d
Author: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
Date:   Fri May 29 12:36:02 2026 -0700

    Revert "drm/xe/nvls: Define GuC firmware for NVL-S"

    commit 42445de1765547f56f48d107c0b8f3482c98458e upstream.

    This reverts commit 4e88de313ff4d1c67b644b1f39f9fb4089711b71.

    The early GuC FW definition meant for our CI branch was accidentally
    merged to the drm-xe-next branch instead. This GuC FW will never be
    released to linux-firmware, so we do not want the definition to be
    available in the mainline Linux codebase.

    Fixes: 4e88de313ff4 ("drm/xe/nvls: Define GuC firmware for NVL-S")
    Signed-off-by: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
    Cc: Julia Filipchuk <julia.filipchuk@intel.com>
    Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Cc: Matt Roper <matthew.d.roper@intel.com>
    Cc: stable@vger.kernel.org # v7.0+
    Reviewed-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Link: https://patch.msgid.link/20260529193558.185436-11-daniele.ceraolospurio@intel.com
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    (cherry picked from commit 65b8e0ac86e48cfc9128c04dfc53ea3395d030dd)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:56 -05:00
Anusha Srivatsa 5d22e66936 drm/xe: fix job timeout recovery for unstarted jobs and kernel queues
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 35b4d137d25c21e62bd8d57d60715cfb66f31a3e
Author: Rodrigo Vivi <rodrigo.vivi@intel.com>
Date:   Wed Jun 10 11:25:49 2026 -0400

    drm/xe: fix job timeout recovery for unstarted jobs and kernel queues

    [ Upstream commit 347ccc0453fca2c669e8dc8a72000e76ca4adf10 ]

    A job that GuC never scheduled (never started) indicates a GuC
    scheduling failure; previously such jobs were silently errored out
    instead of triggering a GT reset to recover. Trigger a GT reset and
    resubmit them, but only when the queue was not already killed or banned:
    an unstarted job on an already banned queue is the ban working as
    intended and must neither clear the ban nor kick off a reset, otherwise
    a banned userspace queue could be resurrected and spam GT resets.

    Kernel queues are always recovered this way and wedge the device once
    recovery attempts are exhausted, since kernel work must not silently
    fail. A started job that times out on a userspace VM bind queue stays
    banned rather than being reset and retried.

    The queue is banned early in the timeout handler to signal the G2H
    scheduling-done handler so it wakes the disable-scheduling waiter;
    without it the waiter sleeps the full 5s timeout. When a reset is
    warranted the ban is cleared before rearming so that
    guc_exec_queue_start() can resubmit jobs after the GT reset - a
    still-banned queue would block resubmission and cause an infinite TDR
    loop. The already-banned case is gated out before this point via
    skip_timeout_check, so it is unaffected.

    v2: (Himal) Do it for any queue type, not just kernel/migration
    v3: - (Sashiko and Sanjay): don't clear the ban / GT reset for already
          killed/banned queues on unstarted-job timeout
        - Update commit message
        - (Matt) Add Fixes tag

    Fixes: fe05cee4d953 ("drm/xe: Don't short circuit TDR on jobs not started")
    Cc: Matthew Auld <matthew.auld@intel.com>
    Cc: Matthew Brost <matthew.brost@intel.com>
    Cc: Sanjay Yadav <sanjay.kumar.yadav@intel.com>
    Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
    Assisted-by: GitHub-Copilot:claude-sonnet-4.6
    Assisted-by: GitHub-Copilot:claude-opus-4.8
    Tested-by: Sanjay Yadav <sanjay.kumar.yadav@intel.com>
    Reviewed-by: Sanjay Yadav <sanjay.kumar.yadav@intel.com>
    Reviewed-by: Matthew Brost <matthew.brost@intel.com>
    Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
    Link: https://patch.msgid.link/20260610152548.404575-3-rodrigo.vivi@intel.com
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    (cherry picked from commit b1107d085e7e8ed15ba6f80c102528a9c8a6cb0e)
    Signed-off-by: Matthew Brost <matthew.brost@intel.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:56 -05:00
Anusha Srivatsa 3a565809ed drm/xe: fix refcount leak in xe_range_fence_insert()
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 667f5fc1cd0b44c2ee8baee92c0c94834329ccfa
Author: Wentao Liang <vulab@iscas.ac.cn>
Date:   Wed Jun 10 10:27:05 2026 -0700

    drm/xe: fix refcount leak in xe_range_fence_insert()

    [ Upstream commit ba36786b21d19082e696eda85bfcd49e7071944a ]

    xe_range_fence_insert() acquires a reference on fence via
    dma_fence_get() and stores it in rfence->fence.  It then calls
    dma_fence_add_callback() and handles two cases: when the callback
    is successfully registered (err == 0) the fence is transferred to
    the tree for later cleanup; when the fence is already signaled
    (err == -ENOENT) it manually drops the extra reference with
    dma_fence_put(fence).

    However, dma_fence_add_callback() can fail with other errors
    (e.g. -EINVAL) and in that case the code falls through to the free:
    label without releasing the acquired reference, leaking it.

    Fix the leak by adding an else branch that calls dma_fence_put()
    before jumping to free: for any error other than -ENOENT.

    Fixes: 845f64bdbfc9 ("drm/xe: Introduce a range-fence utility")
    Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
    Reviewed-by: Matthew Brost <matthew.brost@intel.com>
    Signed-off-by: Matthew Brost <matthew.brost@intel.com>
    Link: https://patch.msgid.link/20260610172705.3450560-1-matthew.brost@intel.com
    (cherry picked from commit 98c4a4201290823c2c5c7ba21692bd9a64b61021)
    Signed-off-by: Matthew Brost <matthew.brost@intel.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:56 -05:00
Anusha Srivatsa 988dbb829e drm/amd/display: use plane color_mgmt_changed to track colorop changes
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 297cc20ee2dba847393dfcdd364dde2f9d9b248b
Author: Melissa Wen <mwen@igalia.com>
Date:   Tue Jun 9 12:20:21 2026 +0200

    drm/amd/display: use plane color_mgmt_changed to track colorop changes

    [ Upstream commit d79716401a954677a93c4dd51fec65beccb38296 ]

    Ensure the driver tracks changes in any colorop property of a plane
    color pipeline by using the same mechanism of CRTC color management and
    update plane color blocks when any colorop property changes. It fixes an
    issue observed on gamescope settings for night mode which is done via
    shaper/3D-LUT updates.

    Fixes: 9ba25915efba ("drm/amd/display: Add support for sRGB EOTF in DEGAM block")
    Reviewed-by: Harry Wentland <harry.wentland@amd.com>
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Melissa Wen <mwen@igalia.com>
    Signed-off-by: Melissa Wen <melissa.srw@gmail.com>
    Link: https://patch.msgid.link/20260609110420.1298352-5-mwen@igalia.com
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:56 -05:00
Anusha Srivatsa e6c70c02ef drm/atomic: track individual colorop updates
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit dc5b1582adc0268ec55bd12069b3cc24b0a5b842
Author: Melissa Wen <mwen@igalia.com>
Date:   Tue Jun 9 12:20:20 2026 +0200

    drm/atomic: track individual colorop updates

    [ Upstream commit 2e235e2a2784b12b735321e5b42240ca51c49b0f ]

    As we do for CRTC color mgmt properties, use color_mgmt_changed flag to
    track any value changes in the color pipeline of a given plane, so that
    drivers can update color blocks as soon as plane color pipeline or
    individual colorop values change. Since we're here, only announce and
    track changes to plane COLOR_PIPELINE prop if its value is actually
    changing.

    Fixes: 8c5ea1745f4c ("drm/colorop: Add BYPASS property")
    Fixes: 7fa3ee8c0a79 ("drm/colorop: Define LUT_1D interpolation")
    Fixes: 41651f9d42eb ("drm/colorop: Add 1D Curve subtype")
    Fixes: 3410108037d5 ("drm/colorop: Add multiplier type")
    Fixes: db971856bbe0 ("drm/colorop: Add 3D LUT support to color pipeline")
    Fixes: e5719e7f1900 ("drm/colorop: Add 3x4 CTM type")
    Fixes: 99a4e4f08abe ("drm/colorop: Add 1D Curve Custom LUT type")
    Fixes: 2afc3184f3b3 ("drm/plane: Add COLOR PIPELINE property")
    Reviewed-by: Harry Wentland <harry.wentland@amd.com> #v1
    Reviewed-by: Chaitanya Kumar Borah <chaitanya.kumar.borah@intel.com>
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Fixes: 9ba25915efba ("drm/amd/display: Add support for sRGB EOTF in DEGAM block")
    Signed-off-by: Melissa Wen <mwen@igalia.com>
    Signed-off-by: Melissa Wen <melissa.srw@gmail.com>
    Link: https://patch.msgid.link/20260609110420.1298352-4-mwen@igalia.com
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:56 -05:00
Anusha Srivatsa 80327b9c17 drm/colorop: make lut(1/3)d_interpolation props correctly behave as mutable
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 546c2badf51266360146206fe36384e9ee97bc0a
Author: Melissa Wen <mwen@igalia.com>
Date:   Tue Jun 9 12:20:19 2026 +0200

    drm/colorop: make lut(1/3)d_interpolation props correctly behave as mutable

    [ Upstream commit 94ff735296d371045fce163451a3d65e44ac4729 ]

    As interpolation props are actually mutable props, any changes should be
    handled by drm_colorop_state. Move their enum and make it correctly
    behaves as mutable.

    Fixes: 7fa3ee8c0a79 ("drm/colorop: Define LUT_1D interpolation")
    Fixes: db971856bbe0 ("drm/colorop: Add 3D LUT support to color pipeline")
    Reviewed-by: Chaitanya Kumar Borah <chaitanya.kumar.borah@intel.com>
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Fixes: 9ba25915efba ("drm/amd/display: Add support for sRGB EOTF in DEGAM block")
    Signed-off-by: Melissa Wen <mwen@igalia.com>
    Signed-off-by: Melissa Wen <melissa.srw@gmail.com>
    Link: https://patch.msgid.link/20260609110420.1298352-3-mwen@igalia.com
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:56 -05:00
Anusha Srivatsa dd056ab6f6 drm/colorop: Remove read-only comments from interpolation fields
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 6c513c4d50515e38b3ea86a3c4b1125fe0fd79af
Author: Alex Hung <alex.hung@amd.com>
Date:   Tue Jun 9 12:20:18 2026 +0200

    drm/colorop: Remove read-only comments from interpolation fields

    [ Upstream commit e480228cf65583040c894bb9cc02e1d5b328cee0 ]

    The lut1d_interpolation and lut3d_interpolation fields and their
    associated properties were marked as read-only, but userspace
    can set them via drm_atomic_colorop_set_property().

    Fixes: 7fa3ee8c0a79 ("drm/colorop: Define LUT_1D interpolation")
    Fixes: db971856bbe0 ("drm/colorop: Add 3D LUT support to color pipeline")
    Reviewed-by: Chaitanya Kumar Borah <chaitanya.kumar.borah@intel.com>
    Signed-off-by: Alex Hung <alex.hung@amd.com>
    Fixes: 9ba25915efba ("drm/amd/display: Add support for sRGB EOTF in DEGAM block")
    Signed-off-by: Melissa Wen <mwen@igalia.com>
    Signed-off-by: Melissa Wen <melissa.srw@gmail.com>
    Link: https://patch.msgid.link/20260609110420.1298352-2-mwen@igalia.com
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:55 -05:00
Anusha Srivatsa 5c46073901 drm/virtio: Fix driver removal with disabled KMS
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 15e561869a8b4e4db69733be1d6f33770664f989
Author: Dmitry Osipenko <dmitry.osipenko@collabora.com>
Date:   Thu Jun 4 15:27:43 2026 +0300

    drm/virtio: Fix driver removal with disabled KMS

    [ Upstream commit f329e8325e054bd6d84d10904f8dd51137281b92 ]

    DRM atomic and modesetting aren't initialized if virtio-gpu driver built
    with disabled KMS, leading to access of uninitialized data on driver
    removal/unbinding and crashing kernel. Fix it by skipping shutting down
    atomic core with unavailable KMS.

    Fixes: 72122c69d717 ("drm/virtio: Add option to disable KMS support")
    Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
    Tested-by: Ryosuke Yasuoka <ryasuoka@redhat.com>
    Reviewed-by: Ryosuke Yasuoka <ryasuoka@redhat.com>
    Link: https://patch.msgid.link/20260604122743.13383-1-dmitry.osipenko@collabora.com
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:55 -05:00
Anusha Srivatsa 254924bd55 drm/i915/edp: Check supported link rates DPCD read
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 914b780ecf31aa85daa1aabfbfb9bf5a9e308d5f
Author: Nikita Zhandarovich <n.zhandarovich@fintech.ru>
Date:   Fri May 29 17:57:58 2026 +0300

    drm/i915/edp: Check supported link rates DPCD read

    [ Upstream commit 2673cefa99ca918e7ac5b0388ff578a83656c896 ]

    intel_edp_set_sink_rates() reads DP_SUPPORTED_LINK_RATES into a local
    stack array and then parses the array unconditionally. If the read
    fails, the array contents are not valid and may result in bogus sink
    link rates being used.

    Use drm_dp_dpcd_read_data() and clear the sink rate array on failure,
    so the existing parser falls back to the default sink rate handling.

    Found by Linux Verification Center (linuxtesting.org) with static
    analysis tool SVACE.

    Fixes: 68f357cb73 ("drm/i915/dp: generate and cache sink rate array for all DP, not just eDP 1.4")
    Signed-off-by: Nikita Zhandarovich <n.zhandarovich@fintech.ru>
    Reviewed-by: Jani Nikula <jani.nikula@intel.com>
    Link: https://patch.msgid.link/20260529145759.1640646-1-n.zhandarovich@fintech.ru
    Signed-off-by: Jani Nikula <jani.nikula@intel.com>
    (cherry picked from commit bd61c7756b34157e093028225a69383b4b1203cc)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:55 -05:00
Anusha Srivatsa ab2afcb969 drm/imx: Fix three kernel-doc warnings in dcss-scaler.c
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 99bdde74338559ed49e5d06f44b7a0a50dfd4379
Author: Yicong Hui <yiconghui@gmail.com>
Date:   Mon Apr 6 19:00:13 2026 +0100

    drm/imx: Fix three kernel-doc warnings in dcss-scaler.c

    [ Upstream commit ae0383e5a9a4b12d68c76c4769857def4665deff ]

    Fix the following W=1 kerneldoc warnings by adding the missing parameter
    descriptions for @phase0_identity and @nn_interpolation in
    dcss_scaler_filter_design() and @phase0_identity in
    dcss_scaler_gaussian_filter()

    Warning: drivers/gpu/drm/imx/dcss/dcss-scaler.c:173 function parameter 'phase0_identity' not described in 'dcss_scaler_gaussian_filter'
    Warning: drivers/gpu/drm/imx/dcss/dcss-scaler.c:270 function parameter 'phase0_identity' not described in 'dcss_scaler_filter_design'
    Warning: drivers/gpu/drm/imx/dcss/dcss-scaler.c:270 function parameter 'nn_interpolation' not described in 'dcss_scaler_filter_design'

    Fixes: 9021c317b7 ("drm/imx: Add initial support for DCSS on iMX8MQ")
    Signed-off-by: Yicong Hui <yiconghui@gmail.com>
    Reviewed-by: Laurentiu Palcu <laurentiu.palcu@oss.nxp.com>
    Link: https://patch.msgid.link/20260406180013.2442096-1-yiconghui@gmail.com
    Signed-off-by: Liu Ying <victor.liu@nxp.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:55 -05:00
Anusha Srivatsa 5a77ebe0a4 drm/amdgpu: check num_entries in GEM_OP GET_MAPPING_INFO
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 967a00b8e06a7734aded23861faba9ee2462be87
Author: Ziyi Guo <n7l8m4@u.northwestern.edu>
Date:   Sun Feb 8 00:02:55 2026 +0000

    drm/amdgpu: check num_entries in GEM_OP GET_MAPPING_INFO

    commit a1ba4594232c87c3b8defd6f89a2e40f8b08395d upstream.

    kvcalloc(args->num_entries, sizeof(*vm_entries), GFP_KERNEL) at
    amdgpu_gem.c:1050 uses the user-supplied num_entries directly without
    any upper bounds check. Since num_entries is a __u32 and
    sizeof(drm_amdgpu_gem_vm_entry) is 32 bytes, a large num_entries
    produces an allocation exceeding INT_MAX, triggering
    WARNING in __kvmalloc_node_noprof(), causing a kernel WARNING,
    TAINT_WARN, and panic on CONFIG_PANIC_ON_WARN=y systems.

    Add a size bounds check before we invoke the kvzalloc() to
    reject oversized num_entries early with -EINVAL.

    Fixes: 4d82724f7f2b ("drm/amdgpu: Add mapping info option for GEM_OP ioctl")
    Signed-off-by: Ziyi Guo <n7l8m4@u.northwestern.edu>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 1fe7bf5457f6efd7be60b17e23163ba54341d73d)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:55 -05:00
Anusha Srivatsa 3ce85d4fda drm/amdgpu: fix amdgpu_hmm_range_get_pages
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 2fd24407457a6b181ba827705678da70e528dcd0
Author: Christian König <christian.koenig@amd.com>
Date:   Wed Feb 18 12:53:27 2026 +0100

    drm/amdgpu: fix amdgpu_hmm_range_get_pages

    commit 962d684b5dc0741dcd93485d41b450de402d5592 upstream.

    The notifier sequence must only be read once or otherwise we could work
    with invalid pages.

    While at it also fix the coding style, e.g. drop the pre-initialized
    return value and use the common define for 2G range.

    Signed-off-by: Christian König <christian.koenig@amd.com>
    Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
    Tested-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
    Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit c08972f555945cda57b0adb72272a37910153390)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:55 -05:00
Anusha Srivatsa abfe919071 drm/amdgpu: fix calling VM invalidation in amdgpu_hmm_invalidate_gfx
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit f0d97cd476103747b1f0c0d441869dbe99d6a833
Author: Christian König <christian.koenig@amd.com>
Date:   Wed Feb 18 12:31:29 2026 +0100

    drm/amdgpu: fix calling VM invalidation in amdgpu_hmm_invalidate_gfx

    commit 1c824497d8acd3187d585d6187cedc1897dcc871 upstream.

    Otherwise we don't invalidate page tables on next CS.

    Signed-off-by: Christian König <christian.koenig@amd.com>
    Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
    Tested-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
    Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit b6444d1bcbc34f6f2a31a3aab3059be082f3683e)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:55 -05:00
Anusha Srivatsa ae4acde731 drm/amdgpu: fix lock leak on ENOMEM in AMDGPU_GEM_OP_GET_MAPPING_INFO
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 8f643d534ffc6f1b6182e4f3acff8f04890504b9
Author: Michael Bommarito <michael.bommarito@gmail.com>
Date:   Sun May 17 09:17:42 2026 -0400

    drm/amdgpu: fix lock leak on ENOMEM in AMDGPU_GEM_OP_GET_MAPPING_INFO

    commit 2e7f55eb408c3f72ee1957a0d0ad11d8648a6379 upstream.

    The AMDGPU_GEM_OP_GET_MAPPING_INFO branch of amdgpu_gem_op_ioctl()
    holds three cleanup-tracked resources before calling kvcalloc():
    the drm_gem_object reference from drm_gem_object_lookup(), the
    drm_exec lock on the looked-up GEM via drm_exec_lock_obj(), and
    the drm_exec lock on the per-process VM root page directory via
    amdgpu_vm_lock_pd().  All three are released by the out_exec
    label that every other error path in this function jumps to.
    The kvcalloc() failure path returns -ENOMEM directly, skipping
    out_exec and leaking all three.

    The leaked per-process VM root PD dma_resv lock is the
    load-bearing leak: any subsequent operation on the same VM
    (further GEM ops, command-submission, eviction, TTM shrinker
    callbacks) blocks on the held lock.  DRM_IOCTL_AMDGPU_GEM_OP is
    DRM_AUTH | DRM_RENDER_ALLOW, so this is an unprivileged-local
    denial of service against the caller's GPU context, reachable
    by any process with /dev/dri/renderD* access.

    Route the failure through out_exec so drm_exec_fini() and
    drm_gem_object_put() run.

    Reproduced on stock 7.0.0-10, Ryzen 7 5700U / Radeon Vega
    (Lucienne): the failing ioctl returns -ENOMEM and a second
    GET_MAPPING_INFO on the same fd then blocks in
    drm_exec_lock_obj() on the leaked dma_resv.  SIGKILL on the
    caller does not reap the task; the fd-release path during
    process exit goes through amdgpu_gem_object_close() ->
    drm_exec_prepare_obj() on the same lock, leaving the task in D
    state until the box is rebooted.  The patched kernel was not
    rebuilt and re-tested on this hardware; the fix is mechanical.
    Tested on a single Lucienne / Vega box only.

    Ziyi Guo posted an independent INT_MAX-bound check for
    args->num_entries in the same branch [1]; the two patches are
    complementary and can land in either order.

    Fixes: 4d82724f7f2b ("drm/amdgpu: Add mapping info option for GEM_OP ioctl")
    Link: https://lore.kernel.org/all/20260208000255.4073363-1-n7l8m4@u.northwestern.edu/ # [1]
    Signed-off-by: Michael Bommarito <michael.bommarito@gmail.com>
    Assisted-by: Claude:claude-opus-4-7
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit b69d3256d79de15f54c322986ff4da68f1d65b0a)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:55 -05:00
Anusha Srivatsa 88f32cdda7 drm/amdkfd: Check for pdd drm file first in CRIU restore path
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit bbb5dfdb729fa1c9563746c10da2672176ee3168
Author: David Francis <David.Francis@amd.com>
Date:   Thu May 14 10:31:20 2026 -0400

    drm/amdkfd: Check for pdd drm file first in CRIU restore path

    commit 6842b6a4b72da9b2906ffc5ca9d846ace2c54c14 upstream.

    CRIU restore ioctls are meant to be called by CRIU with no
    existing drm file. There's an error path
    for if the drm file unexpectedly exists. It was positioned so
    it was missing a fput(drm_file).

    Do that check earlier, as soon as we have the pdd.

    Signed-off-by: David Francis <David.Francis@amd.com>
    Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 2bab781dac78916c5cc8de76345a4102449267d7)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:54 -05:00
Anusha Srivatsa eed403b2be drm/amdkfd: fix a vulnerability of integer overflow in kfd debugger
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 4f9eeedc3d3151f8a226fd676c314a813edda5a1
Author: Eric Huang <jinhuieric.huang@amd.com>
Date:   Tue May 12 10:19:52 2026 -0400

    drm/amdkfd: fix a vulnerability of integer overflow in kfd debugger

    commit 93f5534b35a05ef8a0109c1eefa800062fee810a upstream.

    get_queue_ids() computes array_size = num_queues * sizeof(uint32_t),
    which could overflow on 32-bit size_t build. using array_size()
    instead, it saturates to SIZE_MAX on overflow.

    Signed-off-by: Eric Huang <jinhuieric.huang@amd.com>
    Acked-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 2d57a0475f085c08b49312dfd8edcb461845f285)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:54 -05:00
Anusha Srivatsa 03206966e5 drm/amdkfd: fix NULL pointer bug in svm_range_set_attr
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit c24eee21f9a943374fd64260a6e17dc3984e3d0e
Author: Eric Huang <jinhuieric.huang@amd.com>
Date:   Thu May 7 15:51:49 2026 -0400

    drm/amdkfd: fix NULL pointer bug in svm_range_set_attr

    commit e984d61d92e702096058f0f828f4b2b8563b88ce upstream.

    The process_info could be NULL if user doesn't call kfd_ioctl_acquire_vm
    before calling kfd_ioctl_svm.

    Signed-off-by: Eric Huang <jinhuieric.huang@amd.com>
    Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 83a26c812e0529eb040d31a76f73e33e637243d4)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:54 -05:00
Anusha Srivatsa d5c928daee drm/amd/pm/si: Disregard vblank time when no displays are connected
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit bde1b7af764bfb16baa2ac67b4f28f5ec682593b
Author: Timur Kristóf <timur.kristof@gmail.com>
Date:   Tue May 19 10:41:54 2026 +0200

    drm/amd/pm/si: Disregard vblank time when no displays are connected

    commit dd4f3ee535b3b0ac027f75dbf9dc5fc88733c765 upstream.

    When no displays are connected, there is no vblank
    happening so the power management code shouldn't
    worry about it.

    This fixes a regression that caused the memory clock
    to be stuck at maximum when there were no displays
    connected to a SI GPU.

    Fixes: 9003a0746864 ("drm/amd/pm: Treat zero vblank time as too short in si_dpm (v3)")
    Fixes: 9d73b107a61b ("drm/amd/pm: Use pm_display_cfg in legacy DPM (v2)")
    Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
    Tested-by: Jeremy Klarenbeek <jeremy.klarenbeek99@gmail.com>
    Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 6d87e0199f7b83735b56e422d59f170a201897a8)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:54 -05:00
Anusha Srivatsa bde9cd6618 drm/i915: Fix potential UAF in TTM object purge
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit a29654d451bbffe63d584a4cf64ad0efce6bcf1c
Author: Janusz Krzysztofik <janusz.krzysztofik@linux.intel.com>
Date:   Fri May 8 14:23:51 2026 +0200

    drm/i915: Fix potential UAF in TTM object purge

    commit 5c4063c87a619e4df954c179d24628636f5db15f upstream.

    TLDR: The bo->ttm object might be changed by calling ttm_bo_validate(),
          move casting it to an i915_tt object later to actually get the right
          pointer.

    A user reported hitting the following bug under heavy use on DG2:

    [26620.095550] Oops: general protection fault, probably for non-canonical address 0xa56b6b6b6b6b6b8b: 0000 1 SMP NOPTI
    [26620.095556] CPU: 2 UID: 0 PID: 631 Comm: Xorg Not tainted 6.18.8 #1 PREEMPT(lazy)
    [26620.095558] Hardware name: ASRock B850M Steel Legend WiFi/B850M Steel Legend WiFi, BIOS 3.50 09/18/2025
    [26620.095559] RIP: 0010:i915_ttm_purge+0x84/0x100 [i915]
    [26620.095604] Code: 00 00 00 48 8d 54 24 10 48 89 e6 48 89 fb e8 83 aa ae ff 85 c0 75 6f 48 83 bb a8 01 00 00 00 74 2c 48 8b 45 78 48 85 c0 74 23 <48> 8b 78 20 48 c7 c2 ff ff ff ff 31 f6 e8 7a 73 e3 e0 48 8b 7d 78
    [26620.095605] RSP: 0018:ffffc90005fd7430 EFLAGS: 00010282
    [26620.095607] RAX: a56b6b6b6b6b6b6b RBX: ffff8881f46c3dc0 RCX: 0000000000000000
    [26620.095608] RDX: 0000000000000000 RSI: 0000000000000246 RDI: 00000000ffffffff
    [26620.095609] RBP: ffff888289610f00 R08: 0000000000000001 R09: ffff88823b022000
    [26620.095609] R10: ffff888103029b28 R11: ffff8881fc7f3800 R12: ffff88810b6150d0
    [26620.095609] R13: ffff888289610f00 R14: 0000000000000000 R15: ffff8881f46c3dc0
    [26620.095610] FS: 00007f1004d86900(0000) GS:ffff88901c858000(0000) knlGS:0000000000000000
    [26620.095611] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
    [26620.095611] CR2: 00007f0fdf489000 CR3: 000000035b0c1000 CR4: 0000000000750ef0
    [26620.095612] PKRU: 55555554
    [26620.095612] Call Trace:
    [26620.095615] <TASK>
    [26620.095615] i915_ttm_move+0x2b9/0x420 [i915]
    [26620.095642] ? ttm_tt_init+0x65/0x80 [ttm]
    [26620.095644] ? i915_ttm_tt_create+0xc6/0x150 [i915]
    [26620.095667] ttm_bo_handle_move_mem+0xb6/0x160 [ttm]
    [26620.095669] ttm_bo_evict+0x100/0x150 [ttm]
    [26620.095671] ? preempt_count_add+0x64/0xa0
    [26620.095673] ? _raw_spin_lock+0xe/0x30
    [26620.095675] ? _raw_spin_unlock+0xd/0x30
    [26620.095675] ? i915_gem_object_evictable+0xb7/0xd0 [i915]
    [26620.095704] ttm_bo_evict_cb+0x6e/0xd0 [ttm]
    [26620.095705] ttm_lru_walk_for_evict+0xa6/0x200 [ttm]
    [26620.095708] ttm_bo_alloc_resource+0x185/0x4f0 [ttm]
    [26620.095709] ? init_object+0x62/0xd0
    [26620.095712] ttm_bo_validate+0x7a/0x180 [ttm]
    [26620.095713] ? _raw_spin_unlock_irqrestore+0x16/0x30
    [26620.095714] __i915_ttm_get_pages+0xb0/0x170 [i915]
    [26620.095737] i915_ttm_get_pages+0x9f/0x150 [i915]
    [26620.095759] ? i915_gem_do_execbuffer+0xedc/0x2b40 [i915]
    [26620.095786] ? alloc_debug_processing+0xd0/0x100
    [26620.095787] ? _raw_spin_unlock_irqrestore+0x16/0x30
    [26620.095788] ? i915_vma_instance+0xa0/0x4e0 [i915]
    [26620.095822] __i915_gem_object_get_pages+0x2f/0x40 [i915]
    [26620.095848] i915_vma_pin_ww+0x706/0x980 [i915]
    [26620.095875] ? i915_gem_do_execbuffer+0xedc/0x2b40 [i915]
    [26620.095904] eb_validate_vmas+0x170/0xa00 [i915]
    [26620.095930] i915_gem_do_execbuffer+0x1201/0x2b40 [i915]
    [26620.095953] ? alloc_debug_processing+0xd0/0x100
    [26620.095954] ? _raw_spin_unlock_irqrestore+0x16/0x30
    [26620.095955] ? i915_gem_execbuffer2_ioctl+0xc9/0x240 [i915]
    [26620.095977] ? __wake_up_sync_key+0x32/0x50
    [26620.095979] ? i915_gem_execbuffer2_ioctl+0xc9/0x240 [i915]
    [26620.096001] ? __slab_alloc.isra.0+0x67/0xc0
    [26620.096003] i915_gem_execbuffer2_ioctl+0x11a/0x240 [i915]

    Results from decode_stacktrace.sh pointed to dereference of a file pointer
    field of a i915 TTM page vector container associated with an object being
    purged on eviction.  That path is taken when the object is marked as no
    longer needed.

    Code analysis revealed a possibility of the i915 TTM page vector container
    being replaced with a new instance inside a function that purges content
    of the object, should it be still busy.  That function is called,
    indirectly via a more general function that changes the object's placement
    and caching policy, before the problematic dereference, but still after
    a pointer to the container is captured, rendering the pointer no longer
    valid.

    Fix the issue by capturing the pointer to the container only after its
    potential replacement.

    v2: Move the container_of() inside the if block (Sebastian),
      - a simplified version of the commit description that explains briefly
        why the change is necessary (Christian).

    Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/14882
    Fixes: 7ae034590ceae ("drm/i915/ttm: add tt shmem backend")
    Signed-off-by: Janusz Krzysztofik <janusz.krzysztofik@linux.intel.com>
    Cc: stable@vger.kernel.org # v5.17+
    Cc: Matthew Auld <matthew.auld@intel.com>
    Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
    Cc: Sebastian Brzezinka <sebastian.brzezinka@intel.com>
    Cc: Christian König <christian.koenig@amd.com>
    Reviewed-by: Andi Shyti <andi.shyti@linux.intel.com>
    Reviewed-by: Christian König <christian.koenig@amd.com>
    Signed-off-by: Andi Shyti <andi.shyti@linux.intel.com>
    Link: https://lore.kernel.org/r/20260508122612.469227-2-janusz.krzysztofik@linux.intel.com
    (cherry picked from commit 4462966a93eb185849b7f174f0d0de53476d00a4)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:54 -05:00
Anusha Srivatsa ca16d695ee drm/i915/psr: Use DC_OFF wake reference to block DC6 on vblank enable
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 9e7e6a073f1d0c842c9ca14813f6888454917a6d
Author: Jouni Högander <jouni.hogander@intel.com>
Date:   Wed May 20 13:49:44 2026 +0300

    drm/i915/psr: Use DC_OFF wake reference to block DC6 on vblank enable

    commit 3549a9649dc7c5fc586ab12f675279283cdcb2a7 upstream.

    We are observing following warnings:

    *ERROR* power well DC_off state mismatch (refcount 0/enabled 1)

    gen9_dc_off_power_well_enabled is considering target state DC_STATE_DISABLE
    as DC_OFF power well being enabled. Fix this by using wakeref for the
    purpose.

    To achieve this we need to modify notification code as well. Currently it
    is possible that PSR gets notified vblank enable/disable twice on same
    status. This is currently not a problem as it is just triggering call to
    intel_display_power_set_target_dc_state with same target state as a
    parameter. When using wakeref this becomes a problem due to reference
    counting. Fix this storing vbank status on last notification and use that
    to ensure there are no more than one notification with same vblank status.

    v2: ensure there is no subsequent notifications with same status

    Fixes: aa451abcffb5 ("drm/i915/display: Prevent DC6 while vblank is enabled for Panel Replay")
    Cc: <stable@vger.kernel.org> # v6.13+
    Signed-off-by: Jouni Högander <jouni.hogander@intel.com>
    Reviewed-by: Michał Grzelak <michal.grzelak@intel.com>
    Link: https://patch.msgid.link/20260520104944.239797-2-jouni.hogander@intel.com
    (cherry picked from commit 35485ac56d878192a3829a58cb26503125ec7104)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:54 -05:00
Anusha Srivatsa bddbfe8304 drm/i915/psr: Block DC states on vblank enable when Panel Replay supported
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit eec1212197051553babdc436370131a1cdd835c3
Author: Jouni Högander <jouni.hogander@intel.com>
Date:   Wed May 20 13:49:43 2026 +0300

    drm/i915/psr: Block DC states on vblank enable when Panel Replay supported

    commit 8bb9093df555f9e89fdbe1405118b11384c03e04 upstream.

    Currently we are blocking DC states only when Panel Replay is enabled on
    vblank enable. It may happen that Panel Replay is getting enabled when
    vblank is already enabled. Fix this by blocking DC states always if Panel
    Replay is supported.

    While at it take care of possible dual eDP case by looping all encoders
    supporting PSR.

    Fixes: 0c427ac78a1d ("drm/i915/psr: Add interface to notify PSR of vblank enable/disable")
    Cc: <stable@vger.kernel.org> # v6.16+
    Signed-off-by: Jouni Högander <jouni.hogander@intel.com>
    Reviewed-by: Michał Grzelak <michal.grzelak@intel.com>
    Link: https://patch.msgid.link/20260520104944.239797-1-jouni.hogander@intel.com
    (cherry picked from commit eb5911f990554f7ce947dd53df00c114362e4465)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:54 -05:00
Anusha Srivatsa 4d35cad88a drm/i915/color: Fix HDR pre-CSC LUT programming loop
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit b09f70cec428423c51c408e87b0e7ccd73817401
Author: Pranay Samala <pranay.samala@intel.com>
Date:   Tue May 19 13:23:08 2026 +0530

    drm/i915/color: Fix HDR pre-CSC LUT programming loop

    commit d196136a988051173f68f91de0b5a1bd32122dd7 upstream.

    The integer lut programming loop never executes completely due to
    incorrect condition (i++ > 130).

    Fix to properly program 129th+ entries for values > 1.0.

    Cc: <stable@vger.kernel.org> #v6.19
    Fixes: 82caa1c8813f ("drm/i915/color: Program Pre-CSC registers")
    Signed-off-by: Pranay Samala <pranay.samala@intel.com>
    Signed-off-by: Chaitanya Kumar Borah <chaitanya.kumar.borah@intel.com>
    Reviewed-by: Uma Shankar <uma.shankar@intel.com>
    Signed-off-by: Suraj Kandpal <suraj.kandpal@intel.com>
    Link: https://patch.msgid.link/20260519075308.383877-1-pranay.samala@intel.com
    (cherry picked from commit f33862ec3e8849ad7c0a3dd46719083b13ade248)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:53 -05:00
Anusha Srivatsa fc2532812c drm/gem: fix race between change_handle and handle_delete
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit cde2c9257cbe8463b9dcf7b1075177b72b5fd938
Author: Zhenghang Xiao <kipreyyy@gmail.com>
Date:   Tue May 26 16:53:13 2026 +0800

    drm/gem: fix race between change_handle and handle_delete

    commit 7164d78559b0ff29931a366a840a9e5dd53d4b7c upstream.

    drm_gem_change_handle_ioctl leaves the old handle live in the IDR
    during the window between spin_unlock(table_lock) and the final
    spin_lock(table_lock). A concurrent drm_gem_handle_delete on the old
    handle succeeds in this window, decrements handle_count to 0, and frees
    the GEM object while the new handle's IDR entry still references it.

    NULL the old handle's IDR entry before dropping table_lock so that any
    concurrent GEM_CLOSE on the old handle sees NULL and returns -EINVAL.
    Restore the old entry on the prime-bookkeeping error path.

    Fixes: 5e28b7b94408 ("drm: Set old handle to NULL before prime swap in change_handle")
    Signed-off-by: Zhenghang Xiao <kipreyyy@gmail.com>
    Cc: stable@vger.kernel.org
    Signed-off-by: Dave Airlie <airlied@redhat.com>
    Link: https://patch.msgid.link/20260526085313.26791-1-kipreyyy@gmail.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:53 -05:00
Anusha Srivatsa a78519a667 drm/hyperv: validate VMBus packet size in receive callback
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit c8974d96b6a5496f33dc69a3ce28a7bf5078def4
Author: Berkant Koc <me@berkoc.com>
Date:   Sat May 23 15:27:47 2026 +0200

    drm/hyperv: validate VMBus packet size in receive callback

    commit 7f87763f47a3c22fb50265a00619ef10f2394b18 upstream.

    hyperv_receive_sub() reads msg->vid_hdr.type and dispatches into one
    of four message-type branches without knowing how many bytes the host
    wrote into hv->recv_buf. The completion path then runs
    memcpy(hv->init_buf, msg, VMBUS_MAX_PACKET_SIZE), so the consumer that
    wakes on wait_for_completion_timeout() can read up to 16 KiB of
    residue from a prior message as if it were the response payload.

    Pass bytes_recvd into hyperv_receive_sub() and reject any packet that
    does not cover the pipe + synthvid header. A single switch on
    msg->vid_hdr.type then computes the type-specific payload size: the
    three completion-driving types (SYNTHVID_VERSION_RESPONSE,
    SYNTHVID_RESOLUTION_RESPONSE, SYNTHVID_VRAM_LOCATION_ACK) fall through
    to a shared exit that requires that size before memcpy/complete, while
    SYNTHVID_FEATURE_CHANGE validates its own payload and returns before
    reading is_dirt_needed. Unknown types are dropped.

    SYNTHVID_RESOLUTION_RESPONSE is variable length: the host fills
    resolution_count entries, not the full SYNTHVID_MAX_RESOLUTION_COUNT
    array. Validate the fixed prefix first so resolution_count can be
    read, bound it against the array, then require only the count-sized
    array, so the shorter responses the host actually sends are accepted.

    Only run the sub-handler when vmbus_recvpacket() returned success. The
    memcpy length is bytes_recvd, which is bounded by VMBUS_MAX_PACKET_SIZE
    only on a successful receive; on -ENOBUFS vmbus_recvpacket() instead
    reports the required length, which can exceed hv->recv_buf, so copying
    bytes_recvd would read and write past the 16 KiB buffers. Gating on the
    success return keeps the copy bounded. The nonzero-return path is itself
    a malformed-message case and is now logged rather than silently skipped;
    channel recovery is not attempted.

    Rejected packets are reported via drm_err_ratelimited() rather than
    silently dropped, matching the CoCo-hardened pattern in
    hv_kvp_onchannelcallback().

    Fixes: 76c56a5aff ("drm/hyperv: Add DRM driver for hyperv synthetic video device")
    Cc: stable@vger.kernel.org # 5.14+
    Signed-off-by: Berkant Koc <me@berkoc.com>
    Assisted-by: Claude:claude-opus-4-7 berkoc-pipeline
    Reviewed-by: Michael Kelley <mhklinux@outlook.com>
    Tested-by: Michael Kelley <mhklinux@outlook.com>
    Signed-off-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
    Link: https://patch.msgid.link/8200dbc199c7a9b75ac7e8af6c748d2189b5ebd5.1779542874.git.me@berkoc.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:53 -05:00
Anusha Srivatsa 885a12aaa7 drm/hyperv: validate resolution_count and fix WIN8 fallback
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 8a114b25b5521eae451b13bce98ae978624962e5
Author: Berkant Koc <me@berkoc.com>
Date:   Tue May 19 22:08:17 2026 +0200

    drm/hyperv: validate resolution_count and fix WIN8 fallback

    commit 13d33b9ef67066c77c84273fac5a1d3fde3533d1 upstream.

    A SYNTHVID_RESOLUTION_RESPONSE with resolution_count > 64 walks past
    the supported_resolution[SYNTHVID_MAX_RESOLUTION_COUNT] array in the
    parse loop. Bound resolution_count against the array size, folded
    into the existing zero-check.

    When the WIN10 resolution probe fails, the caller in
    hyperv_connect_vsp() left hv->screen_*_max / preferred_* unpopulated,
    which sets mode_config.max_width / max_height to 0 and makes
    drm_internal_framebuffer_create() reject every userspace framebuffer
    with -EINVAL. The pre-WIN10 branch had the same gap for
    preferred_width / preferred_height. Use a single post-probe fallback
    guarded by screen_width_max == 0 so both paths converge on the WIN8
    defaults.

    Signed-off-by: Berkant Koc <me@berkoc.com>
    Assisted-by: Claude:claude-opus-4-7 berkoc-pipeline
    Fixes: 76c56a5aff ("drm/hyperv: Add DRM driver for hyperv synthetic video device")
    Cc: stable@vger.kernel.org # 5.14+
    Reviewed-by: Michael Kelley <mhklinux@outlook.com>
    Tested-by: Michael Kelley <mhklinux@outlook.com>
    Signed-off-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
    Link: https://patch.msgid.link/6945b22419c7d404b4954a113de2ac9c900dba93.1779542874.git.me@berkoc.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:53 -05:00
Anusha Srivatsa c96e72ae14 dma-buf: fix UAF in dma_buf_fd() tracepoint
JIRA: https://issues.redhat.com/browse/RHEL-180329

Conflicts:
drivers/dma-buf/dma-buf.c

commit b569f86e2f8dbf6f11d31d3de794d22e18098b23
Author: David Carlier <devnexen@gmail.com>
Date:   Sat May 23 19:14:46 2026 +0100

    dma-buf: fix UAF in dma_buf_fd() tracepoint

    commit ead6680f354f83966c796fc7f9463a3171789616 upstream.

    Once FD_ADD() returns, the fd is live in the file descriptor table
    and a thread sharing that table can close() it before DMA_BUF_TRACE()
    runs. The close drops the last reference, __fput() frees the dma_buf,
    and the tracepoint then dereferences dmabuf to take dmabuf->name_lock
    -- slab-use-after-free.

    Split FD_ADD() back into get_unused_fd_flags() + fd_install() and
    emit the tracepoint between them. While the fdtable slot is reserved
    with a NULL file pointer, a racing close() returns -EBADF without
    entering __fput(), so the dma_buf stays alive across the trace. Same
    approach as commit 2d76319c4cbb ("dma-buf: fix UAF in dma_buf_put()
    tracepoint").

    This undoes the FD_ADD() conversion done in commit 34dfce523c90
    ("dma: convert dma_buf_fd() to FD_ADD()"); FD_ADD() has no place to
    hook the tracepoint safely.

    Reported-by: syzbot+7f4987d0afb97dd090cb@syzkaller.appspotmail.com
    Closes: https://syzkaller.appspot.com/bug?extid=7f4987d0afb97dd090cb
    Fixes: 281a22631423 ("dma-buf: add some tracepoints to debug.")
    Cc: stable@vger.kernel.org # 7.0.x
    Signed-off-by: David Carlier <devnexen@gmail.com>
    Reviewed-by: Christian König <christian.koenig@amd.com>
    Signed-off-by: Sumit Semwal <sumit.semwal@linaro.org>
    Link: https://patch.msgid.link/20260523181446.69525-1-devnexen@gmail.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:53 -05:00
Anusha Srivatsa 8f3ec6ba2b drm/i915/psr: Apply Intel DPCD workaround when SDP on prior line used
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit c03affe0700a1240b19923f353965ae394cc996b
Author: Jouni Högander <jouni.hogander@intel.com>
Date:   Fri May 29 12:38:37 2026 +0300

    drm/i915/psr: Apply Intel DPCD workaround when SDP on prior line used

    commit 4703049f768fc1c1caac754134118bee1a3af189 upstream.

    There is Intel specific workaround DPCD address containing workaround for
    case where SDP is on prior line. Apply this workaround according to values
    in the offset.

    Fixes: 61e887329e33 ("drm/i915/xelpd: Handle PSR2 SDP indication in the prior scanline")
    Cc: <stable@vger.kernel.org> # v5.15+
    Signed-off-by: Jouni Högander <jouni.hogander@intel.com>
    Reviewed-by: Suraj Kandpal <suraj.kandpal@intel.com>
    Link: https://patch.msgid.link/20260515095756.2799483-4-jouni.hogander@intel.com
    (cherry picked from commit c3fe899fbeac86ea4a5ca9dd845b2cbc0da46249)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:52 -05:00
Anusha Srivatsa 8072b6db0f drm/i915/psr: Read Intel DPCD workaround register
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit bb7ede8b396e509af6838b6cfe35af621fae80d2
Author: Jouni Högander <jouni.hogander@intel.com>
Date:   Fri May 29 12:38:36 2026 +0300

    drm/i915/psr: Read Intel DPCD workaround register

    commit f30bece421a4ae34359254e1dc2a187a42b6af9b upstream.

    Read Intel DPCD workaround register and store it into
    intel_connector->dp.psr_caps. psr_caps was chosen as currently it contains
    only PSR workaround for PSR2 SDP on prior scanline implementation.

    Signed-off-by: Jouni Högander <jouni.hogander@intel.com>
    Reviewed-by: Suraj Kandpal <suraj.kandpal@intel.com>
    Link: https://patch.msgid.link/20260515095756.2799483-3-jouni.hogander@intel.com
    (cherry picked from commit c48ff24d0f4ab7ad696b2d35ad64ce7e049c668c)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:52 -05:00
Anusha Srivatsa c051e8d1b5 drm/i915/psr: Add defininitions for INTEL_WA_REGISTER_CAPS DPCD register
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit a21a0a1c269c3be1ff2bbd676b3296c4a541ee42
Author: Jouni Högander <jouni.hogander@intel.com>
Date:   Fri May 29 12:38:35 2026 +0300

    drm/i915/psr: Add defininitions for INTEL_WA_REGISTER_CAPS DPCD register

    commit fbceb39b536e40c2f7cc47ab42037bb7c2b7ced9 upstream.

    EDP specification says:

    "If either VSC SDP is unable to be transmitted 100 ns before the SU region,
    the Source device may optionally transmit the VSC SDP during the prior
    video scan line’s HBlank period There is a Intel specific drm dp register
    currently containing bits related how TCON can support PSR2 with SDP on
    prior line."

    Unfortunately many panels are having problems in implementing this. So
    there is a custom Intel specific DPCD register (INTEL_WA_REGISTER_CAPS) to
    figure out if this is properly implemented on a panel or if panel doesn't
    require that 100 ns delay before the SU region. Here are the definitions in
    this custom DPCD address:

    0 = Panel doesn't support SDP on prior line
    1 = Panel supports SDP on prior line
    2 = Panel doesn't have 100ns requirement
    3 = Reserved

    Add definitions for this new register and it's values into new header
    intel_dpcd.h.

    v2: add INTEL_DPCD_ prefix to definitions

    Bspec: 74741
    Signed-off-by: Jouni Högander <jouni.hogander@intel.com>
    Reviewed-by: Suraj Kandpal <suraj.kandpal@intel.com>
    Link: https://patch.msgid.link/20260515095756.2799483-2-jouni.hogander@intel.com
    (cherry picked from commit 1da1c9294825f08f622c473480d185680c2a3b75)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:52 -05:00
Anusha Srivatsa 046e99c089 drm/xe: Restore IDLEDLY regiter on engine reset
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit f1a32b2056c6d4a3e99cd599effb3f02158c892f
Author: Balasubramani Vivekanandan <balasubramani.vivekanandan@intel.com>
Date:   Fri May 22 22:05:32 2026 +0530

    drm/xe: Restore IDLEDLY regiter on engine reset

    [ Upstream commit f657a6a3ba4c20bc01f5be3752d53498ee1bfe35 ]

    Wa_16023105232 programs the register IDLEDLY. The register is reset
    whenever the engine is reset. Therefore it should be added to the GuC
    save-restore register list for it to be restored after reset.

    Fixes: 7c53ff050ba8 ("drm/xe: Apply Wa_16023105232")
    Reviewed-by: Matt Roper <matthew.d.roper@intel.com>
    Link: https://patch.msgid.link/20260522163531.1365540-2-balasubramani.vivekanandan@intel.com
    Signed-off-by: Balasubramani Vivekanandan <balasubramani.vivekanandan@intel.com>
    (cherry picked from commit df1cfe24743a93b71eab27687e148ab8ae9b69e3)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:52 -05:00
Anusha Srivatsa c6ba01030d drm/i915/aux: use polling when irqs are unavailable
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit ef1add67462324cc1a3d7d57a2efcc27b6ed33d3
Author: Michał Grzelak <michal.grzelak@intel.com>
Date:   Thu Apr 16 18:37:44 2026 +0200

    drm/i915/aux: use polling when irqs are unavailable

    [ Upstream commit 202e77cf2e839e1adc804433322dc5c9ee511c9f ]

    PTL with physically disconnected display was observed to have 40s longer
    execution time when testing xe_fault_injection@xe_guc_mmio_send_recv.
    The issue has not been seen when reverting commit 40a9f77a28fa ("Revert
    "drm/i915/dp: change aux_ctl reg read to polling read"").

    Apparently the configuration suffers from not having AUX enabled when
    using interrupts. One probable cause can be xe enabling interrupts too
    late: interrupts need memory allocations which currently can't be done
    before the display FB takeover is done.

    As for now, use polling for AUX in case interrupts are unavailable.

    Fixes: 40a9f77a28fa ("Revert "drm/i915/dp: change aux_ctl reg read to polling read"")
    Suggested-by: Ville Syrjälä <ville.syrjala@linux.intel.com>
    Signed-off-by: Michał Grzelak <michal.grzelak@intel.com>
    Signed-off-by: Ville Syrjälä <ville.syrjala@linux.intel.com>
    Link: https://patch.msgid.link/20260416163744.288107-1-michal.grzelak@intel.com
    (cherry picked from commit 05e0550b65cd1604bd515fbc65f522bce4c10a87)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:52 -05:00
Anusha Srivatsa b139453f32 accel/ivpu: prevent uninitialized data bug in debugfs
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 4193221708b44d33f5c55c5748bbdcaa2f534a60
Author: Dan Carpenter <error27@gmail.com>
Date:   Mon May 25 10:14:42 2026 +0300

    accel/ivpu: prevent uninitialized data bug in debugfs

    [ Upstream commit 44e151be23deb788d9f6124de93823faf6e04e99 ]

    The simple_write_to_buffer() will only initialize data starting from
    the *pos offset so if it's non-zero then the first part of the buffer
    uninitialized.  Really, if *pos is non-zero then this code won't work
    so just check for that at the start of the function.

    Fixes: 320323d2e545 ("accel/ivpu: Add debugfs interface for setting HWS priority bands")
    Signed-off-by: Dan Carpenter <error27@gmail.com>
    Reviewed-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Signed-off-by: Karol Wachowski <karol.wachowski@linux.intel.com>
    Link: https://patch.msgid.link/ahP24m6Mii9EDL7Q@stanley.mountain
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:51 -05:00
Anusha Srivatsa 436fbeeecf drm/xe/oa: Fix exec_queue leak on width check in stream open
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 4b0c4f0c1b133d4bfa31c167200bcda646873328
Author: Shuicheng Lin <shuicheng.lin@intel.com>
Date:   Thu May 14 20:32:10 2026 +0000

    drm/xe/oa: Fix exec_queue leak on width check in stream open

    [ Upstream commit 4d25342543c01310fc4e0cba7cb17c775e2421e2 ]

    In xe_oa_stream_open_ioctl(), when param.exec_q->width > 1 the
    function returns -EOPNOTSUPP directly, skipping the existing
    err_exec_q cleanup path. The exec_queue reference obtained by
    xe_exec_queue_lookup() is leaked.

    The exec queue holds a reference on the xe_file, which is only
    dropped during queue teardown. The leaked lookup ref is not on
    the file's exec_queue xarray, so file close cannot release it.
    This keeps both the exec queue and the file private state pinned
    indefinitely.

    Jump to err_exec_q instead of returning directly so the reference
    is released.

    Fixes: f0ed39830e60 ("xe/oa: Fix query mode of operation for OAR/OAC")
    Assisted-by: Claude:claude-opus-4.6
    Reviewed-by: Ashutosh Dixit <ashutosh.dixit@intel.com>
    Link: https://patch.msgid.link/20260514203210.593488-1-shuicheng.lin@intel.com
    Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com>
    (cherry picked from commit 339fa0be9e4a5d69fa47e91f4a36574224fb478f)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:51 -05:00
Anusha Srivatsa a3dcf17e4b drm/amdgpu/vce1: Fix VCE 1 firmware size and offsets
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit ce0de178ef08408f6ba8f2e9a13bf52fbe5852f4
Author: Timur Kristóf <timur.kristof@gmail.com>
Date:   Wed May 13 22:04:13 2026 +0200

    drm/amdgpu/vce1: Fix VCE 1 firmware size and offsets

    [ Upstream commit 3e5a1d5bb2ff061e64c7992f8e5404dfd4c2d0f3 ]

    The VCPU BO contains the actual FW at an offset, but
    it was not calculated into the VCPU BO size.
    Subtract this from the FW size to make sure there is
    no out of bounds access.

    Make sure the stack and data offsets are aligned to
    the 32K TLB size.

    Check that the FW microcode actually fits in the
    space that is reserved for it.

    Fixes: d4a640d4b9f3 ("drm/amdgpu/vce1: Implement VCE1 IP block (v2)")
    Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
    Reviewed-by: Christian König <christian.koenig@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit c16fe59f622a080fc457a57b3e8f14c780699449)
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:51 -05:00
Anusha Srivatsa fb73d3c253 drm/amdgpu/vce1: Check that the GPU address is < 128 MiB
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 89ea749336cae9e3a72e6d8f2f09342fbb617c17
Author: Timur Kristóf <timur.kristof@gmail.com>
Date:   Wed May 13 22:04:09 2026 +0200

    drm/amdgpu/vce1: Check that the GPU address is < 128 MiB

    [ Upstream commit 9f907adb66d8369dd45412794a04845011503fa8 ]

    When ensuring the low 32-bit address, make sure it is
    less than 128 MiB, otherwise the VCE seems to fail to initialize.
    This seems to be an undocumented limitation of the firmware
    validation mechanism. Note that in case of VCE1 the BAR
    address is zero and we can't change it also due to the
    firmware validator.

    When programming the mmVCE_VCPU_CACHE_OFFSETn registers,
    don't AND them with a mask. This is incorrect because
    the register mask is actually 0x0fffffff and useless because
    we already ensure the addresses are below the limit.

    Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
    Reviewed-by: Christian König <christian.koenig@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit e729ae5f3ac73c861c062080ac8c3d666c972404)
    Stable-dep-of: 3e5a1d5bb2ff ("drm/amdgpu/vce1: Fix VCE 1 firmware size and offsets")
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:51 -05:00
Anusha Srivatsa 74930a1a2d drm/amdgpu: Align amdgpu_gtt_mgr entries to TLB size on Tahiti (v2)
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 35b07a6762a7485adc4d368329602ebe970612f2
Author: Timur Kristóf <timur.kristof@gmail.com>
Date:   Wed May 13 22:04:08 2026 +0200

    drm/amdgpu: Align amdgpu_gtt_mgr entries to TLB size on Tahiti (v2)

    [ Upstream commit 4d798ea0712fddbd35b439cef32b8ac735eb76f9 ]

    The TLB is organized in groups of 8 entries, each one is 4K.
    On Tahiti, the HW requires these GART entries to be 32K-aligned.

    This fixes a VCE 1 firmware validation failure that can happen
    after suspend/resume since we use amdgpu_gtt_mgr for VCE 1.

    v2:
    - Change variable declaration order
    - Add comment about "V bit HW bug"

    Fixes: 698fa62f56aa ("drm/amdgpu: Add helper to alloc GART entries")
    Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
    Reviewed-by: Christian König <christian.koenig@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit 530411b465ef0b2c0cc18c2e3d7e38422b1117d1)
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:51 -05:00
Anusha Srivatsa d7b0c03196 drm/i915/dp: Fix readback for target_rr in Adaptive Sync SDP
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 3cc3c70770eef8d1f7263923b62c995d5c7bca6f
Author: Ankit Nautiyal <ankit.k.nautiyal@intel.com>
Date:   Mon May 11 18:02:15 2026 +0530

    drm/i915/dp: Fix readback for target_rr in Adaptive Sync SDP

    [ Upstream commit f87abd0c6604fb6cc31cc86fc7ccc6a576924352 ]

    Correct the bit-shift logic to properly readback the 10 bit target_rr from
    DB3 and DB4.

    v2: Align the style with readback for vtotal. (Ville)

    Fixes: 12ea89291603 ("drm/i915/dp: Add Read/Write support for Adaptive Sync SDP")
    Cc: Mitul Golani <mitulkumar.ajitkumar.golani@intel.com>
    Cc: Ankit Nautiyal <ankit.k.nautiyal@intel.com>
    Signed-off-by: Ankit Nautiyal <ankit.k.nautiyal@intel.com>
    Reviewed-by: Ville Syrjälä <ville.syrjala@linux.intel.com>
    Link: https://patch.msgid.link/20260511123218.1589830-2-ankit.k.nautiyal@intel.com
    (cherry picked from commit f7abc4af2b19240a145a221461dfe756cc01d74a)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:50 -05:00
Anusha Srivatsa fcd1ff9064 drm/xe: Define and use MCR version of COMMON_SLICE_CHICKEN4
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 72c4b6cd22a2e1b5934885ffb2fbd8043833a9c1
Author: Gustavo Sousa <gustavo.sousa@intel.com>
Date:   Thu May 14 18:44:46 2026 -0300

    drm/xe: Define and use MCR version of COMMON_SLICE_CHICKEN4

    [ Upstream commit 6df5678b6a94ac80e31e847074c4b30c21025b1f ]

    The register COMMON_SLICE_CHICKEN4 is a MCR register on both Xe2 and
    Xe3. Let's make sure to define a MCR version of it and use it for the
    relevant IP versions.

    Use XEHP_ as prefix for the register name, since it is MCR as of Xe_HP.

    v2:
      - Also change for one entry in lrc_tunnings, which was caught by
        manual testing and add corresponging Fixes tag in commit message.
        (Gustavo)

    Fixes: 8d6f16f1f082 ("drm/xe: Extend Wa_22021007897 to Xe3 platforms")
    Fixes: e5c13e2c505b ("drm/xe/xe2hpg: Add Wa_22021007897")
    Fixes: 8ccf5f6b2295 ("drm/xe/tuning: Apply windower hardware filtering setting on Xe3 and Xe3p")
    Bspec: 66534, 71185, 74417
    Reviewed-by: Matt Roper <matthew.d.roper@intel.com>
    Link: https://patch.msgid.link/20260514-rtp-mcr-check-v3-3-30dd47855fee@intel.com
    Signed-off-by: Gustavo Sousa <gustavo.sousa@intel.com>
    (cherry picked from commit 75f65f1a4c06da1d87f28570a9d4cdad28f13360)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:50 -05:00
Anusha Srivatsa 7b76fb94a5 drm/xe/tuning: Apply windower hardware filtering setting on Xe3 and Xe3p
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit aad838731edf97aebc7604eb645a275fba822f8e
Author: Matt Roper <matthew.d.roper@intel.com>
Date:   Tue Feb 24 15:50:56 2026 -0800

    drm/xe/tuning: Apply windower hardware filtering setting on Xe3 and Xe3p

    [ Upstream commit 8ccf5f6b2295164962bbee5b0770f4366fd9bee2 ]

    A recent bspec tuning guide update asks us to program
    COMMON_SLICE_CHICKEN4[5] on Xe3 and Xe3p platforms.  Add this setting to
    our LRC tuning RTP table so that the setting will become part of each
    context's LRC.

    Bspec: 72161, 55902
    Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
    Link: https://patch.msgid.link/20260224235055.3038710-2-matthew.d.roper@intel.com
    Signed-off-by: Matt Roper <matthew.d.roper@intel.com>
    Stable-dep-of: 6df5678b6a94 ("drm/xe: Define and use MCR version of COMMON_SLICE_CHICKEN4")
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:50 -05:00
Anusha Srivatsa 363a4306f8 drm/xe: Define and use MCR version of COMMON_SLICE_CHICKEN1
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 64bbac33d8ad8c6724978806518e37642d064713
Author: Gustavo Sousa <gustavo.sousa@intel.com>
Date:   Thu May 14 18:44:45 2026 -0300

    drm/xe: Define and use MCR version of COMMON_SLICE_CHICKEN1

    [ Upstream commit a4660bd949733fd6ea621fdb50fabac2608155e9 ]

    The register COMMON_SLICE_CHICKEN1 is a MCR register on Xe2.
    Let's make sure to define a MCR version of it and use it for the
    relevant IP versions.

    Use XEHP_ as prefix for the register name, since it is MCR as of Xe_HP.

    Fixes: a5d221924e13 ("drm/xe/xe2_hpg: Add set of workarounds")
    Fixes: 9f18b55b6d3f ("drm/xe/xe2: Add workaround 18033852989")
    Bspec: 66534, 71185
    Reviewed-by: Matt Roper <matthew.d.roper@intel.com>
    Link: https://patch.msgid.link/20260514-rtp-mcr-check-v3-2-30dd47855fee@intel.com
    Signed-off-by: Gustavo Sousa <gustavo.sousa@intel.com>
    (cherry picked from commit a672725fdbfc3ea430130039d677c7dc98d59df8)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:50 -05:00
Anusha Srivatsa c59658e26a drm/xe: Consolidate workaround entries for Wa_18033852989
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 986a69e2a2abbff2d7ab099ac84cf06d2f46de9c
Author: Matt Roper <matthew.d.roper@intel.com>
Date:   Fri Feb 20 09:27:41 2026 -0800

    drm/xe: Consolidate workaround entries for Wa_18033852989

    [ Upstream commit fe681e7b44d78fd77d79de21eca58c3b6bdcda0e ]

    Wa_18033852989 applies to all graphics versions from 20.01 through 20.04
    (inclusive).  Consolidate the RTP entries into a single range-based entry.

    Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
    Link: https://patch.msgid.link/20260220-forupstream-wa_cleanup-v2-19-b12005a05af6@intel.com
    Signed-off-by: Matt Roper <matthew.d.roper@intel.com>
    Stable-dep-of: a4660bd94973 ("drm/xe: Define and use MCR version of COMMON_SLICE_CHICKEN1")
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:49 -05:00
Anusha Srivatsa a51acd01bb drm/xe: Consolidate workaround entries for Wa_14019988906
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 5758fa346aa7a98e2bf89509cabfa8f72a9e2692
Author: Matt Roper <matthew.d.roper@intel.com>
Date:   Fri Feb 20 09:27:40 2026 -0800

    drm/xe: Consolidate workaround entries for Wa_14019988906

    [ Upstream commit c2142a1a841525d897ef69b3e6a5ab48183e1fcf ]

    Wa_14019988906 applies to all graphics versions from 20.01 through 20.04
    (inclusive).  Consolidate the RTP entries into a single range-based entry.

    Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
    Link: https://patch.msgid.link/20260220-forupstream-wa_cleanup-v2-18-b12005a05af6@intel.com
    Signed-off-by: Matt Roper <matthew.d.roper@intel.com>
    Stable-dep-of: a4660bd94973 ("drm/xe: Define and use MCR version of COMMON_SLICE_CHICKEN1")
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:49 -05:00
Anusha Srivatsa 6b4e2e7b1c drm/xe/pf: Fix CFI failure in debugfs access
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 9fba83e404e6142428308344d61adb7cdc107a2b
Author: Mohanram Meenakshisundaram <mohanram.meenakshisundaram@intel.com>
Date:   Thu May 14 23:19:18 2026 +0530

    drm/xe/pf: Fix CFI failure in debugfs access

    [ Upstream commit 96bf49b526e2d03a2b7f6e861925a08f46ed0d28 ]

    Reading debugfs file (/sys/kernel/debug/dri/0/gt*/pf/adverse_events)
    with CFI (Control Flow Integrity) enabled, the kernel panics at
    xe_gt_debugfs_simple_show+0x82/0xc0.

    xe_gt_debugfs_simple_show() declare a function pointer expecting int
    return type, but xe_gt_sriov_pf_monitor_print_events() is void return
    type, leading to CFI failure and kernel panic.

    [507620.973657] CFI failure at xe_gt_debugfs_simple_show+0x82/0xc0 [xe]
    (target: xe_gt_sriov_pf_monitor_print_events+0x0/0x130 [xe]; expected
    type: 0xd72c7139)

    Fix xe_gt_sriov_pf_monitor_print_events() function by updating to return
    an int type.

    Fixes: 1c99d3d3edab ("drm/xe/pf: Expose PF monitor details via debugfs")
    Signed-off-by: Mohanram Meenakshisundaram <mohanram.meenakshisundaram@intel.com>
    Reviewed-by: Michal Wajdeczko <michal.wajdeczko@intel.com>
    Signed-off-by: Michal Wajdeczko <michal.wajdeczko@intel.com>
    Link: https://patch.msgid.link/20260514174918.1556357-2-mohanram.meenakshisundaram@intel.com
    (cherry picked from commit ff1d386a8359746d9699ac30336e3b0684c68958)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:49 -05:00
Anusha Srivatsa 6085ce10f1 drm/xe/vf: Fix signature of print functions
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit a0b154d60043d151b5d9ee4f798eea502b650292
Author: Michal Wajdeczko <michal.wajdeczko@intel.com>
Date:   Thu May 14 17:57:26 2026 +0200

    drm/xe/vf: Fix signature of print functions

    [ Upstream commit 9bb2f1d7e6e58b8e434ddc2048c661bf87ccdf2a ]

    We have plugged-in existing VF print functions into our GT debugfs
    show helper as-is, but we missed that the helper expects functions
    to return int, while they were defined as void. This can lead to
    errors being reported when CFI is enabled.

    Fixes: 63d8cb8fe3dd ("drm/xe/vf: Expose SR-IOV VF attributes to GT debugfs")
    Signed-off-by: Michal Wajdeczko <michal.wajdeczko@intel.com>
    Cc: Mohanram Meenakshisundaram <mohanram.meenakshisundaram@intel.com>
    Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
    Link: https://patch.msgid.link/20260514155726.7165-1-michal.wajdeczko@intel.com
    (cherry picked from commit 314e31c9a8a1c421ee4f7f755b9348aefbbca090)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:49 -05:00
Anusha Srivatsa c080079d52 drm/xe/gsc: Fix double-free of managed BO in error path
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 889f70de2b51a877339e1979aab95111b41bed75
Author: Shuicheng Lin <shuicheng.lin@intel.com>
Date:   Mon May 11 15:41:34 2026 +0000

    drm/xe/gsc: Fix double-free of managed BO in error path

    [ Upstream commit d3ded53fab90996e7d94a39049e11962dd066725 ]

    The error path in xe_gsc_init_post_hwconfig() explicitly frees a BO
    allocated with xe_managed_bo_create_pin_map() via
    xe_bo_unpin_map_no_vm(). Since the managed BO already has a devm
    cleanup action registered, this causes a double-free when devm
    unwinds during probe failure.

    Remove the explicit free and let devm handle it, consistent with
    all other xe_managed_bo_create_pin_map() callers.

    Fixes: 2e5d47fe7839 ("drm/xe/uc: Use managed bo for HuC and GSC objects")
    Reviewed-by: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
    Assisted-by: Claude:claude-opus-4.6
    Link: https://patch.msgid.link/20260511154134.223696-1-shuicheng.lin@intel.com
    Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com>
    (cherry picked from commit 71d61e3e299a17139e47f980a4d6f425b2c59bf7)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:48 -05:00
Anusha Srivatsa 4e4f5814ba drm/gem: Make the GEM LRU lock part of drm_device
JIRA: https://issues.redhat.com/browse/RHEL-180329

Conflicts:
drivers/gpu/drm/drm_drv.c
drivers/gpu/drm/drm_gem.c
include/drm/drm_device.h
include/drm/drm_gem.h

Bringing in MSM bloates the MR by adding ~1400 lines if code.
Hence, MSM driver changes dropped - CONFIG_DRM_MSM is not set for cs9.
Conflicting files skipped:
drivers/gpu/drm/msm/msm_drv.c
drivers/gpu/drm/msm/msm_drv.h
drivers/gpu/drm/msm/msm_gem.c
drivers/gpu/drm/msm/msm_gem_shrinker.c
drivers/gpu/drm/msm/msm_gem_submit.c
drivers/gpu/drm/msm/msm_gem_vma.c
drivers/gpu/drm/msm/msm_ringbuffer.c

commit 6e13c85cac3bc92f13535fd18a96e206f9f9df19
Author: Boris Brezillon <boris.brezillon@collabora.com>
Date:   Mon May 18 13:41:45 2026 +0200

    drm/gem: Make the GEM LRU lock part of drm_device

    [ Upstream commit 379e8f1ca5e919b130b40d8115d92a536e5f8d7a ]

    Recently, a few races have been discovered in the GEM LRU logic, all
    of them caused by the fact the LRU lock is accessed through
    gem->lru->lock, and that very same lock also protects changes to
    gem->lru, leading to situations where gem->lru needs to first be
    accessed without the lock held, to then get the lru to access the lock
    through and finally take the lock and do the expected operation.

    Currently, the only driver making use of this API (MSM) declares a
    device-wide lock, and the user we're about to add (panthor) will
    do the same. There's no evidence that we will ever have a driver
    that wants different pools of LRUs protected by different locks under
    the same drm_device. So we're better off moving this lock to drm_device
    and always locking it through obj->dev->gem_lru_mutex, or directly
    through dev->gem_lru_mutex.

    If anyone ever needs more fine-grained locking, this can be revisited
    to pass some drm_gem_lru_pool object representing the pool of LRUs
    under a specific lock, but for now, the per-device lock seems to be
    enough.

    Fixes: e7c2af13f811 ("drm/gem: Add LRU/shrinker helper")
    Reported-by: Chia-I Wu <olvaffe@gmail.com>
    Closes: https://gitlab.freedesktop.org/panfrost/linux/-/work_items/86
    Reviewed-by: Rob Clark <rob.clark@oss.qualcomm.com>
    Reviewed-by: Liviu Dudau <liviu.dudau@arm.com>
    Reviewed-by: Steven Price <steven.price@arm.com>
    Reviewed-by: Chia-I Wu <olvaffe@gmail.com>
    Link: https://patch.msgid.link/20260518-panthor-shrinker-fixes-v4-1-1920234470d5@collabora.com
    Signed-off-by: Boris Brezillon <boris.brezillon@collabora.com>
    Signed-off-by: Sasha Levin <sashal@kernel.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:48 -05:00
Anusha Srivatsa 4c8a829c3f drm/amd/display: Validate payload length and link_index in dc_process_dmub_aux_transfer_async
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 1c8c6e912f2945b2a3e669afca6b52174b88e86e
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Thu May 7 16:26:31 2026 -0400

    drm/amd/display: Validate payload length and link_index in dc_process_dmub_aux_transfer_async

    commit 6c92f6d9600efa3ef0d9e560a2b52776d9803c29 upstream.

    [Why&How]
    dc_process_dmub_aux_transfer_async() copies payload->length bytes into a
    16-byte stack buffer (dpaux.data[16]) guarded only by an ASSERT(), which
    is a no-op in release builds. If a caller ever passes length > 16 this
    results in a stack buffer overflow via memcpy.

    Additionally, link_index is used to dereference dc->links[] without
    bounds checking against dc->link_count, risking an out-of-bounds access.

    Replace the ASSERT with a hard runtime check that returns false when
    payload->length exceeds the destination buffer size, and add a bounds
    check for link_index before it is used.

    Assisted-by: GitHub Copilot:Claude claude-4-opus
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
    Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit ba4caa9fecdf7a38f98c878ad05a8a64148b6881)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:47 -05:00
Anusha Srivatsa 1bd5c6f3af drm/amd/display: Validate GPIO pin LUT table size before iterating
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit f2a4827e980ba07de4391fa84d9c39a12726bdd7
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Mon May 4 16:14:11 2026 -0400

    drm/amd/display: Validate GPIO pin LUT table size before iterating

    commit 86d2b20644b11d21fe52c596e6e922b4590a3e3f upstream.

    [Why&How]
    The GPIO pin table parsers in get_gpio_i2c_info() and
    bios_parser_get_gpio_pin_info() derive an element count from the VBIOS
    table_header.structuresize field, then iterate over gpio_pin[] entries.
    However, GET_IMAGE() only validates that the table header itself fits
    within the BIOS image. If the VBIOS reports a structuresize larger than
    the actual mapped data, the loop reads past the end of the BIOS image,
    causing an out-of-bounds read.

    Fix this by calling bios_get_image() to validate that the full claimed
    structuresize is accessible within the BIOS image before entering the
    loop in both functions.

    Assisted-by: GitHub Copilot:claude-opus-4-6
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
    Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit ba5e95b43b773ae1bf1f66ee6b31eb774e65afe3)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:47 -05:00
Anusha Srivatsa ce33d938df drm/amd/display: Fix integer overflow in bios_get_image()
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 0d1aacbbe265fc5d166484b098b9e8a7335cf98b
Author: Harry Wentland <harry.wentland@amd.com>
Date:   Mon May 4 11:14:45 2026 -0400

    drm/amd/display: Fix integer overflow in bios_get_image()

    commit cd86529ec61474a38c3837fb7823790a7c3f8cce upstream.

    [Why&How]
    The bounds check in bios_get_image() computes 'offset + size' using
    unsigned 32-bit arithmetic before comparing against bios_size. If a
    VBIOS image contains a near-UINT32_MAX offset the addition wraps to a
    small value, the comparison passes, and the function returns a wild
    pointer past the VBIOS mapping.

    Additionally, the comparison uses '<' (strict), which incorrectly
    rejects the valid exact-fit case where offset + size == bios_size.

    Fix both issues by restructuring the check to avoid the addition
    entirely: first reject if offset alone exceeds bios_size, then check
    size against the remaining space (bios_size - offset). This eliminates
    the overflow and correctly permits exact-fit accesses.

    Assisted-by: GitHub Copilot:claude-opus-4.6
    Reviewed-by: Alex Hung <alex.hung@amd.com>
    Signed-off-by: Harry Wentland <harry.wentland@amd.com>
    Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
    Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit d40fb392af659c4a02b560319f226842f6ec1a95)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:47 -05:00
Anusha Srivatsa 7bb58b3d76 drm/bridge: megachips: remove bridge when irq request fails
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 3526e6297dd3606e77eeac1fca80ecdbb10d2b02
Author: Osama Abdelkader <osama.abdelkader@gmail.com>
Date:   Thu Apr 30 21:56:59 2026 +0200

    drm/bridge: megachips: remove bridge when irq request fails

    commit d45d5c819f2cd0b6b5d76a194a537a5f4aeefecb upstream.

    If devm_request_threaded_irq() fails after drm_bridge_add(), remove the
    bridge before returning.

    Keep drm_bridge_add() rather than devm_drm_bridge_add(): registration is
    tied to the STDP4028 device while ge_b850v3_register() may complete from
    either I2C probe; devm would not unwind the bridge if the other client's
    probe fails.

    Signed-off-by: Osama Abdelkader <osama.abdelkader@gmail.com>
    Fixes: fcfa0ddc18 ("drm/bridge: Drivers for megachips-stdpxxxx-ge-b850v3-fw (LVDS-DP++)")
    Cc: stable@vger.kernel.org
    Reviewed-by: Luca Ceresoli <luca.ceresoli@bootlin.com>
    Tested-by: Ian Ray <ian.ray@gehealthcare.com>
    Link: https://patch.msgid.link/20260430195700.80317-1-osama.abdelkader@gmail.com
    Signed-off-by: Luca Ceresoli <luca.ceresoli@bootlin.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:46 -05:00
Anusha Srivatsa e335536147 drm/bridge: it66121: acquire reset GPIO in probe
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 95a6d5d3e2a9fdba8e50e61c57bb96611378f7b7
Author: Julien Chauveau <chauveau.julien@gmail.com>
Date:   Tue Mar 24 20:30:11 2026 +0100

    drm/bridge: it66121: acquire reset GPIO in probe

    commit e02b5262fd288cc235f14e12233ea54e78c04611 upstream.

    The it66121_ctx structure has a gpio_reset field, and it66121_hw_reset()
    calls gpiod_set_value() on it. However, the GPIO descriptor is never
    acquired via devm_gpiod_get(), leaving gpio_reset as NULL throughout
    the driver lifetime.

    gpiod_set_value() silently returns when passed a NULL descriptor, so
    the hardware reset sequence in it66121_hw_reset() is a no-op. This
    leaves the chip in an undefined state at probe time, which can prevent
    it from responding on the I2C bus.

    The DT binding marks reset-gpios as a required property, so all
    compliant device trees provide this GPIO. Add the missing
    devm_gpiod_get() call after enabling power supplies and before the
    hardware reset, so the chip is properly reset with power applied.

    Fixes: 988156dc2f ("drm: bridge: add it66121 driver")
    Cc: stable@vger.kernel.org
    Signed-off-by: Julien Chauveau <chauveau.julien@gmail.com>
    Reviewed-by: Javier Martinez Canillas <javierm@redhat.com>
    Tested-by: Javier Martinez Canillas <javierm@redhat.com>
    Link: https://patch.msgid.link/20260324193011.16583-1-chauveau.julien@gmail.com
    Signed-off-by: Javier Martinez Canillas <javierm@redhat.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:46 -05:00
Anusha Srivatsa c4dcbad73e drm/amdgpu/vpe: Force collaborate sync after TRAP
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 5bb70a29c5510fee581085b78bf15c76685340c7
Author: Alan Liu <haoping.liu@amd.com>
Date:   Fri May 1 12:35:48 2026 +0800

    drm/amdgpu/vpe: Force collaborate sync after TRAP

    commit b6074630a461b1322a814988779005cbc43612ea upstream.

    VPE1 could possibly hang and fail to power off at the end of commands in
    collaboration mode. This workaround adds a COLLAB_SYNC after TRAP to
    force instances synchronized to avoid VPE1 fail to power off.

    Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>
    Signed-off-by: Alan liu <haoping.liu@amd.com>
    Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5171
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
    (cherry picked from commit a8b749c5c5afb7e5daa2bfb95d958fb3c6b8f055)
    Cc: stable@vger.kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:46 -05:00
Anusha Srivatsa 4513defa74 drm/xe/multi_queue: Fix secondary queue error case
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 9831e534853ff169660ae7a4fcea699fc8f3367d
Author: Niranjana Vishwanathapura <niranjana.vishwanathapura@intel.com>
Date:   Mon May 18 12:16:40 2026 -0700

    drm/xe/multi_queue: Fix secondary queue error case

    commit 00907da2126ed785451b2a2f0fef282246dad104 upstream.

    If xe_lrc_create() fails, the secondary queue added to the
    multi-queue group list is not removed before freeing the
    queue. Fix error path handling for secondary queues by
    removing it from the multi-queue group list at the right
    place.

    Reported-by: Sebastian Österlund <sebastian.osterlund@intel.com>
    Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/7979
    Fixes: d716a5088c88 ("drm/xe/multi_queue: Handle tearing down of a multi queue")
    Cc: stable@vger.kernel.org # v7.0+
    Signed-off-by: Niranjana Vishwanathapura <niranjana.vishwanathapura@intel.com>
    Reviewed-by: Matthew Auld <matthew.auld@intel.com>
    Link: https://patch.msgid.link/20260518191639.320890-2-niranjana.vishwanathapura@intel.com
    (cherry picked from commit d2d23c12789cf69eddc35b8d38cd8eaabd0168f1)
    Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:45 -05:00
Anusha Srivatsa dc7ac4c780 drm/virtio: use uninterruptible resv lock for plane updates
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit a2359a411b15f495d12cfda6a7db6855ebb7f90f
Author: Deepanshu Kartikey <kartikey406@gmail.com>
Date:   Tue May 19 13:52:47 2026 +0530

    drm/virtio: use uninterruptible resv lock for plane updates

    commit 9af1b6e175c82daf4b423da339a722d8e67a735a upstream.

    virtio_gpu_cursor_plane_update() and virtio_gpu_resource_flush() lock
    the framebuffer BO's dma_resv via virtio_gpu_array_lock_resv() and
    ignore its return value. The function can fail with -EINTR from
    dma_resv_lock_interruptible() (signal during lock wait) or with
    -ENOMEM from dma_resv_reserve_fences() (fence slot allocation),
    leaving the resv lock not held. The queue path then walks the object
    array and calls dma_resv_add_fence(), which requires the lock held;
    with lockdep enabled this trips dma_resv_assert_held():

      WARNING: drivers/dma-buf/dma-resv.c:296 at dma_resv_add_fence+0x71e/0x840
      Call Trace:
       virtio_gpu_array_add_fence
       virtio_gpu_queue_ctrl_sgs
       virtio_gpu_queue_fenced_ctrl_buffer
       virtio_gpu_cursor_plane_update
       drm_atomic_helper_commit_planes
       drm_atomic_helper_commit_tail
       commit_tail
       drm_atomic_helper_commit
       drm_atomic_commit
       drm_atomic_helper_update_plane
       __setplane_atomic
       drm_mode_cursor_universal
       drm_mode_cursor_common
       drm_mode_cursor_ioctl
       drm_ioctl
       __x64_sys_ioctl

    Beyond the WARN, mutating the dma_resv fence list without the lock
    races with concurrent readers/writers and can corrupt the list.

    Both call sites run inside the .atomic_update plane callback, which
    DRM atomic helpers do not allow to fail (by the time it runs, the
    commit has been signed off to userspace and there is no clean
    rollback path). Moving the lock acquisition to .prepare_fb was
    rejected because the broader lock scope deadlocks against other BO
    locking paths in the same atomic commit.

    Introduce virtio_gpu_lock_one_resv_uninterruptible() that uses
    dma_resv_lock() instead of dma_resv_lock_interruptible(). This
    eliminates the -EINTR failure mode -- the realistic syzbot trigger
    -- without extending the lock hold across the commit. The helper
    locks a single BO and rejects nents > 1 with -EINVAL; both fix
    sites lock exactly one BO.

    Use it from virtio_gpu_cursor_plane_update() and
    virtio_gpu_resource_flush(); check the return value to handle the
    remaining -ENOMEM case from dma_resv_reserve_fences() by freeing
    the objs and skipping the plane update for that frame. The
    framebuffer BOs touched here are not shared with other contexts
    and lock contention is expected to be brief, so the loss of
    signal-interruptibility is acceptable.

    Other callers of virtio_gpu_array_lock_resv() (the ioctl paths)
    continue to use the interruptible variant.

    The bug was reported by syzbot, triggered via fault injection
    (fail_nth) on the DRM_IOCTL_MODE_CURSOR path, which forces the
    -ENOMEM branch in dma_resv_reserve_fences().

    Reported-by: syzbot+72bd3dd3a5d5f39a0271@syzkaller.appspotmail.com
    Closes: https://syzkaller.appspot.com/bug?extid=72bd3dd3a5d5f39a0271
    Fixes: 5cfd31c5b3 ("drm/virtio: fix virtio_gpu_cursor_plane_update().")
    Cc: stable@vger.kernel.org
    Signed-off-by: Deepanshu Kartikey <kartikey406@gmail.com>
    Signed-off-by: Dmitry Osipenko <dmitry.osipenko@collabora.com>
    Link: https://patch.msgid.link/20260519082247.34470-1-kartikey406@gmail.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:45 -05:00
Anusha Srivatsa 29e114452e drm/i915/display: Copy color pipeline from plane in the primary joiner pipe
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit 1effd54b8d6b03afcf1e27d5e4393577d699c26d
Author: Chaitanya Kumar Borah <chaitanya.kumar.borah@intel.com>
Date:   Mon May 11 11:02:10 2026 +0530

    drm/i915/display: Copy color pipeline from plane in the primary joiner pipe

    commit 86ed2d96db1965e9008e919b1936145ae66540e3 upstream.

    When copying plane color state in a joiner configuration, use the plane in
    the primary joiner pipe since it carries the pipeline number selected by
    the user-space.

    This assumes that all pipes in the joiner are symmetric in their plane
    color capabilities.

    Cc: stable@vger.kernel.org # v6.19+
    Fixes: a78f1b6baf4d ("drm/i915/color: Add framework to program CSC")
    Tested-by: Vidya Srinivas <vidya.srinivas@intel.com>
    Signed-off-by: Chaitanya Kumar Borah <chaitanya.kumar.borah@intel.com>
    Reviewed-by: Uma Shankar <uma.shankar@intel.com>
    Signed-off-by: Ankit Nautiyal <ankit.k.nautiyal@intel.com>
    Link: https://patch.msgid.link/20260511053213.3122314-2-chaitanya.kumar.borah@intel.com
    (cherry picked from commit e8308fb5e05ca08ddfb8b46f6d947a6e3fd80cd7)
    Signed-off-by: Tvrtko Ursulin <tursulin@ursulin.net>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:45 -05:00
Anusha Srivatsa 25309950d7 drm/bridge: chipone-icn6211: use devm_drm_bridge_add in i2c probe
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit e27ae0c00329ee577bec1c417c2ec02ab4950e5a
Author: Osama Abdelkader <osama.abdelkader@gmail.com>
Date:   Thu Apr 30 21:49:42 2026 +0200

    drm/bridge: chipone-icn6211: use devm_drm_bridge_add in i2c probe

    commit 73d01051e8040c0b1de7fd26b3b8d0c2ffa6895c upstream.

    Use devm_drm_bridge_add() so the bridge is released if probe
    fails after registration, and drop drm_bridge_remove() in chipone_i2c_probe.

    Signed-off-by: Osama Abdelkader <osama.abdelkader@gmail.com>
    Fixes: 8dde6f7452a1 ("drm: bridge: icn6211: Add I2C configuration support")
    Cc: stable@vger.kernel.org
    Reviewed-by: Luca Ceresoli <luca.ceresoli@bootlin.com>
    Link: https://patch.msgid.link/20260430194944.78119-1-osama.abdelkader@gmail.com
    Signed-off-by: Luca Ceresoli <luca.ceresoli@bootlin.com>
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:45 -05:00
Anusha Srivatsa 64418b3185 drm/gma500/oaktrail_lvds: fix i2c adapter leaks on init
JIRA: https://issues.redhat.com/browse/RHEL-180329

commit c96c3d62ac177653a97b6dd0b6210f041447c147
Author: Johan Hovold <johan@kernel.org>
Date:   Fri May 8 16:44:46 2026 +0200

    drm/gma500/oaktrail_lvds: fix i2c adapter leaks on init

    commit 84d1c9b416d54afe760ca4c378bd95c89261254c upstream.

    The LVDS init code looks up an I2C adapter using i2c_get_adapter() and
    tries to read the EDID before falling back to allocating and registering
    its own adapter.

    Make sure to drop the references taken by i2c_get_adapter() when falling
    back to allocating an adapter as well as on late errors to allow the
    looked up adapter to be deregistered.

    Fixes: 1b082ccf59 ("gma500: Add Oaktrail support")
    Cc: stable@vger.kernel.org      # 3.3
    Signed-off-by: Johan Hovold <johan@kernel.org>
    Signed-off-by: Patrik Jakobsson <patrik.r.jakobsson@gmail.com>
    Link: https://patch.msgid.link/20260508144446.59722-4-johan@kernel.org
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Anusha Srivatsa <asrivats@redhat.com>
2026-08-03 14:53:44 -05:00