Commit Graph
100 Commits
Author SHA1 Message Date
Vitaly Kuznetsov d3734a66e9 Buffer overflow in drivers/xen/sys-hypervisor.c
JIRA: https://redhat.atlassian.net/browse/RHEL-172511
CVE: CVE-2026-31786

commit 27fdbab4221b375de54bf91919798d88520c6e28
Author: Juergen Gross <jgross@suse.com>
Date:   Fri Mar 27 14:13:38 2026 +0100

    Buffer overflow in drivers/xen/sys-hypervisor.c

    The build id returned by HYPERVISOR_xen_version(XENVER_build_id) is
    neither NUL terminated nor a string.

    The first causes a buffer overflow as sprintf in buildid_show will
    read and copy till it finds a NUL.

    00000000  f4 91 51 f4 dd 38 9e 9d  65 47 52 eb 10 71 db 50  |..Q..8..eGR..q.P|
    00000010  b9 a8 01 42 6f 2e 32                              |...Bo.2|
    00000017

    So use a memcpy instead of sprintf to have the correct value:

    00000000  f4 91 51 f4 dd 00 9e 9d  65 47 52 eb 10 71 db 50  |..Q.....eGR..q.P|
    00000010  b9 a8 01 42                                       |...B|
    00000014

    (the above have a hack to embed a zero inside and check it's
    returned correctly).

    This is XSA-485 / CVE-2026-31786

    Fixes: 84b7625728 ("xen: add sysfs node for hypervisor build id")
    Signed-off-by: Frediano Ziglio <frediano.ziglio@citrix.com>
    Reviewed-by: Juergen Gross <jgross@suse.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2026-08-26 15:35:19 +02:00
Vitaly Kuznetsov 5d66b806d2 redhat: sign UKI's inner vmlinuz with modules signing key
JIRA: https://redhat.atlassian.net/browse/RHEL-223629
Upstream Status: RHEL only

The existing workflow which puts SB-signed vmlinuz in the UKI and then
SB-signs the UKI itself may be problematic for the situation when build
time signing is unavailable. Switch to using the transient module signing
key for signing vmlinuz which gets included into the UKI. Compared to
putting unsigned vmlinuz in the UKI, this keeps kexec/kdump happy.

Note: 'systemd-sbsign' is not available in RHEL9, use pesign/NSS instead.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2026-08-03 17:02:52 +02:00
Vitaly Kuznetsov 1603e6758a scsi: storvsc: Handle PERSISTENT_RESERVE_IN truncation for Hyper-V vFC
JIRA: https://redhat.atlassian.net/browse/RHEL-159283

commit 9cf351b289fb2be22491fa3964f99126db67aa08
Author: Li Tian <litian@redhat.com>
Date:   Mon Apr 6 09:53:44 2026 +0800

    scsi: storvsc: Handle PERSISTENT_RESERVE_IN truncation for Hyper-V vFC

    The storvsc driver has become stricter in handling SRB status codes
    returned by the Hyper-V host. When using Virtual Fibre Channel (vFC)
    passthrough, the host may return SRB_STATUS_DATA_OVERRUN for
    PERSISTENT_RESERVE_IN commands if the allocation length in the CDB does
    not match the host's expected response size.

    Currently, this status is treated as a fatal error, propagating
    Host_status=0x07 [DID_ERROR] to the SCSI mid-layer. This causes
    userspace storage utilities (such as sg_persist) to fail with transport
    errors, even when the host has actually returned the requested
    reservation data in the buffer.

    Refactor the existing command-specific workarounds into a new helper
    function, storvsc_host_mishandles_cmd(), and add PERSISTENT_RESERVE_IN
    to the list of commands where SRB status errors should be suppressed for
    vFC devices. This ensures that the SCSI mid-layer processes the returned
    data buffer instead of terminating the command.

    Signed-off-by: Li Tian <litian@redhat.com>
    Reviewed-by: Long Li <longli@microsoft.com>
    Reviewed-by: Laurence Oberman <loberman@redhat.com>
    Link: https://patch.msgid.link/20260406015344.12566-1-litian@redhat.com
    Signed-off-by: Martin K. Petersen <martin.petersen@oracle.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2026-04-27 14:44:20 +02:00
Vitaly Kuznetsov bd893fe805 redhat: conflict with unsupported shim on x86/aarch64
JIRA: https://issues.redhat.com/browse/RHEL-126425
Upstream Status: RHEL only

The kernel has recently switched to using 800-series keys for SecureBoot
and this requires shim to have the corresponding CA certificate. The first
version which had it was 15.8-1 so in case the new kernel is installed with
an older shim, 'Security violation' error is going to prevent booting when
SecureBoot=on. Prevent such broken combos by adding an explicit conflict.

The problem can easily be observed on x86 by upgrading the kernel to a
recent version on an old (RHEL9.2 and below) system. Aarch64 systems are
only theoretically affected as SecureBoot was not supported by these old
releases.

Note: UKI is not affected by the issue as it still uses 504 key.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-12-10 10:41:35 +01:00
Vitaly Kuznetsov fe87b7e11d redhat: revert to using redhatsecureboot504 for RHEL UKI
JIRA: https://issues.redhat.com/browse/RHEL-122230
Upstream Status: RHEL only

Azure CVM instances use Full Disk Encryption with the volume key
sealed to PCR7, updating the certificate requires additional action.
Restore the status quo for now.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-10-22 08:43:34 -04:00
Vitaly Kuznetsov 708da51bbd x86/hyperv: Fix kdump on Azure CVMs
Upstream Status: git://git.kernel.org/pub/scm/linux/kernel/git/hyperv/linux.git
JIRA: https://issues.redhat.com/browse/RHEL-70228

commit a883d6d0d8066d691680ee6425aadb83ca36130b
Author: Vitaly Kuznetsov <vkuznets@redhat.com>
Date:   Thu Aug 28 12:16:18 2025 +0300

    x86/hyperv: Fix kdump on Azure CVMs

    Azure CVM instance types featuring a paravisor hang upon kdump. The
    investigation shows that makedumpfile causes a hang when it steps on a page
    which was previously share with the host
    (HVCALL_MODIFY_SPARSE_GPA_PAGE_HOST_VISIBILITY). The new kernel has no
    knowledge of these 'special' regions (which are Vmbus connection pages,
    GPADL buffers, ...). There are several ways to approach the issue:
    - Convey the knowledge about these regions to the new kernel somehow.
    - Unshare these regions before accessing in the new kernel (it is unclear
    if there's a way to query the status for a given GPA range).
    - Unshare these regions before jumping to the new kernel (which this patch
    implements).

    To make the procedure as robust as possible, store PFN ranges of shared
    regions in a linked list instead of storing GVAs and re-using
    hv_vtom_set_host_visibility(). This also allows to avoid memory allocation
    on the kdump/kexec path.

    Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
    Reviewed-by: Michael Kelley <mhklinux@outlook.com>
    Reviewed-by: Tianyu Lan <tiala@microsoft.com>
    Signed-off-by: Wei Liu <wei.liu@kernel.org>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-09-16 12:52:15 +02:00
Vitaly Kuznetsov 9056622a67 x86/tdx: Fix "in-kernel MMIO" check
JIRA: https://issues.redhat.com/browse/RHEL-63318
CVE: CVE-2024-47727

commit d4fc4d01471528da8a9797a065982e05090e1d81
Author: Alexey Gladkov (Intel) <legion@kernel.org>
Date:   Fri Sep 13 19:05:56 2024 +0200

    x86/tdx: Fix "in-kernel MMIO" check

    TDX only supports kernel-initiated MMIO operations. The handle_mmio()
    function checks if the #VE exception occurred in the kernel and rejects
    the operation if it did not.

    However, userspace can deceive the kernel into performing MMIO on its
    behalf. For example, if userspace can point a syscall to an MMIO address,
    syscall does get_user() or put_user() on it, triggering MMIO #VE. The
    kernel will treat the #VE as in-kernel MMIO.

    Ensure that the target MMIO address is within the kernel before decoding
    instruction.

    Fixes: 31d58c4e557d ("x86/tdx: Handle in-kernel MMIO")
    Signed-off-by: Alexey Gladkov (Intel) <legion@kernel.org>
    Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
    Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
    Acked-by: Dave Hansen <dave.hansen@linux.intel.com>
    Cc:stable@vger.kernel.org
    Link: https://lore.kernel.org/all/565a804b80387970460a4ebc67c88d1380f61ad1.1726237595.git.legion%40kernel.org

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-03-20 17:47:55 +01:00
Vitaly Kuznetsov 871087ba5b uki: get rid of RHEL-only walinuxagentcvm module in the initramfs
JIRA: https://issues.redhat.com/browse/RHEL-83003
Upstream Status: RHEL-only

The module was used for early Confidential OS disk encryption
implementation in Azure before it gained support for systemd-style root
volume key sealing. All Azure regions were migrated so the rules are now
useless.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-03-20 15:41:00 +01:00
Vitaly Kuznetsov 9c4fdb3687 Revert "x86/kvm: Override default caching mode for SEV-SNP and TDX"
JIRA: https://issues.redhat.com/browse/RHEL-75512
Upstream Status: RHEL only

This reverts commit b7cf9ba976.

The commit breaks TPM device on Google Clouds's SNP and TDX instances,
the problem was acknowledged upstream but the fix is still being worked
on, see
https://lore.kernel.org/lkml/Z5P_Rj4Uc82lJBDx@google.com/
https://lore.kernel.org/lkml/20250201005048.657470-1-seanjc@google.com/

For now, restore the status quo by reverting the commit.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-20 11:38:51 +01:00
Vitaly Kuznetsov 94093965b8 x86/xen: add FRAME_END to xen_hypercall_hvm()
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit 0bd797b801bd8ee06c822844e20d73aaea0878dd
Author: Juergen Gross <jgross@suse.com>
Date:   Wed Feb 5 10:07:56 2025 +0100

    x86/xen: add FRAME_END to xen_hypercall_hvm()

    xen_hypercall_hvm() is missing a FRAME_END at the end, add it.

    Reported-by: kernel test robot <lkp@intel.com>
    Closes: https://lore.kernel.org/oe-kbuild-all/202502030848.HTNTTuo9-lkp@intel.com/
    Fixes: b4845bb63838 ("x86/xen: add central hypercall functions")
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Reviewed-by: Jan Beulich <jbeulich@suse.com>
    Reviewed-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:06 +01:00
Vitaly Kuznetsov ea7c25b535 x86/xen: fix xen_hypercall_hvm() to not clobber %rbx
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit 98a5cfd2320966f40fe049a9855f8787f0126825
Author: Juergen Gross <jgross@suse.com>
Date:   Wed Feb 5 09:43:31 2025 +0100

    x86/xen: fix xen_hypercall_hvm() to not clobber %rbx

    xen_hypercall_hvm(), which is used when running as a Xen PVH guest at
    most only once during early boot, is clobbering %rbx. Depending on
    whether the caller relies on %rbx to be preserved across the call or
    not, this clobbering might result in an early crash of the system.

    This can be avoided by using an already saved register instead of %rbx.

    Fixes: b4845bb63838 ("x86/xen: add central hypercall functions")
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Reviewed-by: Jan Beulich <jbeulich@suse.com>
    Reviewed-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:06 +01:00
Vitaly Kuznetsov b92405e563 x86/static-call: Remove early_boot_irqs_disabled check to fix Xen PVH dom0
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit 5cc2db37124bb33914996d6fdbb2ddb3811f2945
Author: Andrew Cooper <andrew.cooper3@citrix.com>
Date:   Sat Dec 21 21:10:46 2024 +0000

    x86/static-call: Remove early_boot_irqs_disabled check to fix Xen PVH dom0

    __static_call_update_early() has a check for early_boot_irqs_disabled, but
    is used before early_boot_irqs_disabled is set up in start_kernel().

    Xen PV has always special cased early_boot_irqs_disabled, but Xen PVH does
    not and falls over the BUG when booting as dom0.

    It is very suspect that early_boot_irqs_disabled starts as 0, becomes 1 for
    a time, then becomes 0 again, but as this needs backporting to fix a
    breakage in a security fix, dropping the BUG_ON() is the far safer option.

    Fixes: 0ef8047b737d ("x86/static-call: provide a way to do very early static-call updates")
    Closes: https://bugzilla.kernel.org/show_bug.cgi?id=219620
    Reported-by: Alex Zenla <alex@edera.dev>
    Suggested-by: Peter Zijlstra <peterz@infradead.org>
    Signed-off-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Reviewed-by: Juergen Gross <jgross@suse.com>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Tested-by: Alex Zenla <alex@edera.dev>
    Link: https://lore.kernel.org/r/20241221211046.6475-1-andrew.cooper3@citrix.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:06 +01:00
Vitaly Kuznetsov a05539d312 x86/asm: Make serialize() always_inline
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit ae02ae16b76160f0aeeae2c5fb9b15226d00a4ef
Author: Juergen Gross <jgross@suse.com>
Date:   Wed Dec 18 11:09:18 2024 +0100

    x86/asm: Make serialize() always_inline

    In order to allow serialize() to be used from noinstr code, make it
    __always_inline.

    Fixes: 0ef8047b737d ("x86/static-call: provide a way to do very early static-call updates")
    Closes: https://lore.kernel.org/oe-kbuild-all/202412181756.aJvzih2K-lkp@intel.com/
    Reported-by: kernel test robot <lkp@intel.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Link: https://lore.kernel.org/r/20241218100918.22167-1-jgross@suse.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:06 +01:00
Vitaly Kuznetsov 2512e94725 x86/static-call: fix 32-bit build
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit 349f0086ba8b2a169877d21ff15a4d9da3a60054
Author: Juergen Gross <jgross@suse.com>
Date:   Wed Dec 18 09:02:28 2024 +0100

    x86/static-call: fix 32-bit build

    In 32-bit x86 builds CONFIG_STATIC_CALL_INLINE isn't set, leading to
    static_call_initialized not being available.

    Define it as "0" in that case.

    Reported-by: Stephen Rothwell <sfr@canb.auug.org.au>
    Fixes: 0ef8047b737d ("x86/static-call: provide a way to do very early static-call updates")
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:06 +01:00
Vitaly Kuznetsov 7c2d4a329e x86/xen: remove hypercall page
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit 7fa0da5373685e7ed249af3fa317ab1e1ba8b0a6
Author: Juergen Gross <jgross@suse.com>
Date:   Thu Oct 17 15:27:31 2024 +0200

    x86/xen: remove hypercall page

    The hypercall page is no longer needed. It can be removed, as from the
    Xen perspective it is optional.

    But, from Linux's perspective, it removes naked RET instructions that
    escape the speculative protections that Call Depth Tracking and/or
    Untrain Ret are trying to achieve.

    This is part of XSA-466 / CVE-2024-53241.

    Reported-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Reviewed-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Reviewed-by: Jan Beulich <jbeulich@suse.com>

Conflicts:
	arch/x86/xen/enlighten_pvh.c (skipping 4c006734898a1)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:06 +01:00
Vitaly Kuznetsov dcbec11ae7 x86/xen: use new hypercall functions instead of hypercall page
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit b1c2cb86f4a7861480ad54bb9a58df3cbebf8e92
Author: Juergen Gross <jgross@suse.com>
Date:   Thu Oct 17 14:47:13 2024 +0200

    x86/xen: use new hypercall functions instead of hypercall page

    Call the Xen hypervisor via the new xen_hypercall_func static-call
    instead of the hypercall page.

    This is part of XSA-466 / CVE-2024-53241.

    Reported-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Co-developed-by: Peter Zijlstra <peterz@infradead.org>
    Co-developed-by: Josh Poimboeuf <jpoimboe@redhat.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:06 +01:00
Vitaly Kuznetsov f1f1253847 x86/xen: add central hypercall functions
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit b4845bb6383821a9516ce30af3a27dc873e37fd4
Author: Juergen Gross <jgross@suse.com>
Date:   Thu Oct 17 11:00:52 2024 +0200

    x86/xen: add central hypercall functions

    Add generic hypercall functions usable for all normal (i.e. not iret)
    hypercalls. Depending on the guest type and the processor vendor
    different functions need to be used due to the to be used instruction
    for entering the hypervisor:

    - PV guests need to use syscall
    - HVM/PVH guests on Intel need to use vmcall
    - HVM/PVH guests on AMD and Hygon need to use vmmcall

    As PVH guests need to issue hypercalls very early during boot, there
    is a 4th hypercall function needed for HVM/PVH which can be used on
    Intel and AMD processors. It will check the vendor type and then set
    the Intel or AMD specific function to use via static_call().

    This is part of XSA-466 / CVE-2024-53241.

    Reported-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Co-developed-by: Peter Zijlstra <peterz@infradead.org>

Conflicts:
	arch/x86/xen/xen-ops.h (skipping bcea31e2d1c7a)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:06 +01:00
Vitaly Kuznetsov 1f48e04258 x86/xen: don't do PV iret hypercall through hypercall page
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit a2796dff62d6c6bfc5fbebdf2bee0d5ac0438906
Author: Juergen Gross <jgross@suse.com>
Date:   Wed Oct 16 10:40:26 2024 +0200

    x86/xen: don't do PV iret hypercall through hypercall page

    Instead of jumping to the Xen hypercall page for doing the iret
    hypercall, directly code the required sequence in xen-asm.S.

    This is done in preparation of no longer using hypercall page at all,
    as it has shown to cause problems with speculation mitigations.

    This is part of XSA-466 / CVE-2024-53241.

    Reported-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Reviewed-by: Jan Beulich <jbeulich@suse.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:05 +01:00
Vitaly Kuznetsov 6e825739ea x86/static-call: provide a way to do very early static-call updates
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit 0ef8047b737d7480a5d4c46d956e97c190f13050
Author: Juergen Gross <jgross@suse.com>
Date:   Fri Nov 29 16:15:54 2024 +0100

    x86/static-call: provide a way to do very early static-call updates

    Add static_call_update_early() for updating static-call targets in
    very early boot.

    This will be needed for support of Xen guest type specific hypercall
    functions.

    This is part of XSA-466 / CVE-2024-53241.

    Reported-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Co-developed-by: Peter Zijlstra <peterz@infradead.org>
    Co-developed-by: Josh Poimboeuf <jpoimboe@redhat.com>

Conflicts:
	include/linux/compiler.h (skipping ed2f752e0e0a2, context)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:05 +01:00
Vitaly Kuznetsov 937599d763 objtool/x86: allow syscall instruction
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit dda014ba59331dee4f3b773a020e109932f4bd24
Author: Juergen Gross <jgross@suse.com>
Date:   Fri Nov 29 15:47:49 2024 +0100

    objtool/x86: allow syscall instruction

    The syscall instruction is used in Xen PV mode for doing hypercalls.
    Allow syscall to be used in the kernel in case it is tagged with an
    unwind hint for objtool.

    This is part of XSA-466 / CVE-2024-53241.

    Reported-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Co-developed-by: Peter Zijlstra <peterz@infradead.org>

Conflicts:
	tools/objtool/check.c (skipping 246b2c85487a7)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:05 +01:00
Vitaly Kuznetsov e65e4086c5 x86: make get_cpu_vendor() accessible from Xen code
JIRA: https://issues.redhat.com/browse/RHEL-70669
CVE: CVE-2024-53241

commit efbcd61d9bebb771c836a3b8bfced8165633db7c
Author: Juergen Gross <jgross@suse.com>
Date:   Thu Oct 17 08:29:48 2024 +0200

    x86: make get_cpu_vendor() accessible from Xen code

    In order to be able to differentiate between AMD and Intel based
    systems for very early hypercalls without having to rely on the Xen
    hypercall page, make get_cpu_vendor() non-static.

    Refactor early_cpu_init() for the same reason by splitting out the
    loop initializing cpu_devs() into an externally callable function.

    This is part of XSA-466 / CVE-2024-53241.

    Reported-by: Andrew Cooper <andrew.cooper3@citrix.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:05 +01:00
Vitaly Kuznetsov 102534db12 x86/xen: Avoid relocatable quantities in Xen ELF notes
JIRA: https://issues.redhat.com/browse/RHEL-70669

commit 223abe96ac0d227b22d48ab447dd9384b7a6c9fa
Author: Ard Biesheuvel <ardb@kernel.org>
Date:   Wed Oct 9 18:04:43 2024 +0200

    x86/xen: Avoid relocatable quantities in Xen ELF notes

    Xen puts virtual and physical addresses into ELF notes that are treated
    by the linker as relocatable by default. Doing so is not only pointless,
    given that the ELF notes are only intended for consumption by Xen before
    the kernel boots. It is also a KASLR leak, given that the kernel's ELF
    notes are exposed via the world readable /sys/kernel/notes.

    So emit these constants in a way that prevents the linker from marking
    them as relocatable. This involves place-relative relocations (which
    subtract their own virtual address from the symbol value) and linker
    provided absolute symbols that add the address of the place to the
    desired value.

    Tested-by: Jason Andryuk <jason.andryuk@amd.com>
    Signed-off-by: Ard Biesheuvel <ardb@kernel.org>
    Reviewed-by: Jason Andryuk <jason.andryuk@amd.com>
    Message-ID: <20241009160438.3884381-11-ardb+git@google.com>
    Signed-off-by: Juergen Gross <jgross@suse.com>

Conflicts:
	arch/x86/platform/pvh/head.S (skipping 47ffe0578aee4)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-02-18 17:58:05 +01:00
Vitaly Kuznetsov 8484b839ff redhat: make kernel-debug-uki-virt installable without kernel-debug-core
JIRA: https://issues.redhat.com/browse/RHEL-70874
Upstream Status: RHEL-only

It was noticed that 'kernel-debug-uki-virt' package drags in
'kernel-debug-core' through 'kernel-debug-modules-core' dependency but it
shouldn't as 'kernel-debug-uki-virt' is suffifient. Turns out:

$ rpm -qp --requires kernel-debug-modules-core-6.13.0-0.rc5.42.fc42.x86_64.rpm
...
kernel-uname-r = 6.13.0-0.rc5.42.fc42.x86_64++debug
...

$ rpm -qp --provides kernel-debug-uki-virt-6.13.0-0.rc5.42.fc42.x86_64.rpm
...
kernel-debug-uname-r = 6.13.0-0.rc5.42.fc42.x86_64++debug
...

which doesn't match. On the contrary, 'kernel-debug-core' provides what's
needed:

$ rpm -qp --provides kernel-debug-core-6.13.0-0.rc5.42.fc42.x86_64.rpm
...
kernel-uname-r = 6.13.0-0.rc5.42.fc42.x86_64++debug
...

Make 'kernel-debug-uki-virt' provide 'kernel-uname-r' instead of
'kernel-debug-uname-r', the fact that it is a debug kernel is determined by
the suffix.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2025-01-07 12:29:09 +01:00
Vitaly Kuznetsov 5dc9672eb6 x86/tdx: Enable CPU topology enumeration
JIRA: https://issues.redhat.com/browse/RHEL-29351

commit 7ae15e2f69bad06527668b478dff7c099ad2e6ae
Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Date:   Mon Nov 4 12:38:03 2024 +0200

    x86/tdx: Enable CPU topology enumeration

    TDX 1.0 defines baseline behaviour of TDX guest platform. TDX 1.0
    generates a #VE when accessing topology-related CPUID leafs (0xB and
    0x1F) and the X2APIC_APICID MSR. The kernel returns all zeros on CPUID
    topology. In practice, this means that the kernel can only boot with a
    plain topology. Any complications will cause problems.

    The ENUM_TOPOLOGY feature allows the VMM to provide topology
    information to the guest. Enabling the feature eliminates
    topology-related #VEs: the TDX module virtualizes accesses to
    the CPUID leafs and the MSR.

    Enable ENUM_TOPOLOGY if it is available.

    Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
    Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
    Acked-by: Kai Huang <kai.huang@intel.com>
    Link: https://lore.kernel.org/all/20241104103803.195705-5-kirill.shutemov%40linux.intel.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-31 13:29:11 +01:00
Vitaly Kuznetsov 15936fb7ae x86/tdx: Dynamically disable SEPT violations from causing #VEs
JIRA: https://issues.redhat.com/browse/RHEL-29351

commit f65aa0ad79fca4ace921da0701644f020129043d
Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Date:   Mon Nov 4 12:38:02 2024 +0200

    x86/tdx: Dynamically disable SEPT violations from causing #VEs

    Memory access #VEs are hard for Linux to handle in contexts like the
    entry code or NMIs.  But other OSes need them for functionality.
    There's a static (pre-guest-boot) way for a VMM to choose one or the
    other.  But VMMs don't always know which OS they are booting, so they
    choose to deliver those #VEs so the "other" OSes will work.  That,
    unfortunately has left us in the lurch and exposed to these
    hard-to-handle #VEs.

    The TDX module has introduced a new feature. Even if the static
    configuration is set to "send nasty #VEs", the kernel can dynamically
    request that they be disabled. Once they are disabled, access to private
    memory that is not in the Mapped state in the Secure-EPT (SEPT) will
    result in an exit to the VMM rather than injecting a #VE.

    Check if the feature is available and disable SEPT #VE if possible.

    If the TD is allowed to disable/enable SEPT #VEs, the ATTR_SEPT_VE_DISABLE
    attribute is no longer reliable. It reflects the initial state of the
    control for the TD, but it will not be updated if someone (e.g. bootloader)
    changes it before the kernel starts. Kernel must check TDCS_TD_CTLS bit to
    determine if SEPT #VEs are enabled or disabled.

    [ dhansen: remove 'return' at end of function ]

    Fixes: 373e715e31bf ("x86/tdx: Panic on bad configs that #VE on "private" memory access")
    Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
    Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
    Acked-by: Kai Huang <kai.huang@intel.com>
    Link: https://lore.kernel.org/all/20241104103803.195705-4-kirill.shutemov%40linux.intel.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-31 13:29:08 +01:00
Vitaly Kuznetsov b4df8f805e x86/tdx: Rename tdx_parse_tdinfo() to tdx_setup()
JIRA: https://issues.redhat.com/browse/RHEL-29351

commit b064043d9565786b385f85e6436ca5716bbd5552
Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Date:   Mon Nov 4 12:38:01 2024 +0200

    x86/tdx: Rename tdx_parse_tdinfo() to tdx_setup()

    Rename tdx_parse_tdinfo() to tdx_setup() and move setting NOTIFY_ENABLES
    there.

    The function will be extended to adjust TD configuration.

    Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
    Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
    Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>
    Reviewed-by: Kai Huang <kai.huang@intel.com>
    Link: https://lore.kernel.org/all/20241104103803.195705-3-kirill.shutemov%40linux.intel.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-31 13:29:05 +01:00
Vitaly Kuznetsov 898b8b0a6a x86/tdx: Introduce wrappers to read and write TD metadata
JIRA: https://issues.redhat.com/browse/RHEL-29351

commit 5081e8fadb809253c911b349b01d87c5b4e3fec5
Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Date:   Mon Nov 4 12:38:00 2024 +0200

    x86/tdx: Introduce wrappers to read and write TD metadata

    The TDG_VM_WR TDCALL is used to ask the TDX module to change some
    TD-specific VM configuration. There is currently only one user in the
    kernel of this TDCALL leaf.  More will be added shortly.

    Refactor to make way for more users of TDG_VM_WR who will need to modify
    other TD configuration values.

    Add a wrapper for the TDG_VM_RD TDCALL that requests TD-specific
    metadata from the TDX module. There are currently no users for
    TDG_VM_RD. Mark it as __maybe_unused until the first user appears.

    This is preparation for enumeration and enabling optional TD features.

    Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
    Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
    Reviewed-by: Kai Huang <kai.huang@intel.com>
    Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>
    Link: https://lore.kernel.org/all/20241104103803.195705-2-kirill.shutemov%40linux.intel.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-31 13:28:51 +01:00
Vitaly Kuznetsov 4a2d63e284 efi/x86: Free EFI memory map only when installing a new one.
JIRA: https://issues.redhat.com/browse/RHEL-33045

commit 75dde792d6f6c2d0af50278bd374bf0c512fe196
Author: Ard Biesheuvel <ardb@kernel.org>
Date:   Mon Jun 10 16:02:13 2024 +0200

    efi/x86: Free EFI memory map only when installing a new one.

    The logic in __efi_memmap_init() is shared between two different
    execution flows:
    - mapping the EFI memory map early or late into the kernel VA space, so
      that its entries can be accessed;
    - the x86 specific cloning of the EFI memory map in order to insert new
      entries that are created as a result of making a memory reservation
      via a call to efi_mem_reserve().

    In the former case, the underlying memory containing the kernel's view
    of the EFI memory map (which may be heavily modified by the kernel
    itself on x86) is not modified at all, and the only thing that changes
    is the virtual mapping of this memory, which is different between early
    and late boot.

    In the latter case, an entirely new allocation is created that carries a
    new, updated version of the kernel's view of the EFI memory map. When
    installing this new version, the old version will no longer be
    referenced, and if the memory was allocated by the kernel, it will leak
    unless it gets freed.

    The logic that implements this freeing currently lives on the code path
    that is shared between these two use cases, but it should only apply to
    the latter. So move it to the correct spot.

    While at it, drop the dummy definition for non-x86 architectures, as
    that is no longer needed.

    Cc: <stable@vger.kernel.org>
    Fixes: f0ef652347 ("efi: Fix efi_memmap_alloc() leaks")
    Tested-by: Ashish Kalra <Ashish.Kalra@amd.com>
    Link: https://lore.kernel.org/all/36ad5079-4326-45ed-85f6-928ff76483d3@amd.com
    Signed-off-by: Ard Biesheuvel <ardb@kernel.org>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-11 17:57:26 +01:00
Vitaly Kuznetsov e8c16bcf03 x86/sev: Convert shared memory back to private on kexec
JIRA: https://issues.redhat.com/browse/RHEL-33045

commit 3074152e56c9b0f9b9c67edfbc08b371db050b6d
Author: Ashish Kalra <ashish.kalra@amd.com>
Date:   Thu Aug 1 19:14:50 2024 +0000

    x86/sev: Convert shared memory back to private on kexec

    SNP guests allocate shared buffers to perform I/O. It is done by
    allocating pages normally from the buddy allocator and converting them
    to shared with set_memory_decrypted().

    The second, kexec-ed, kernel has no idea what memory is converted this
    way. It only sees E820_TYPE_RAM.

    Accessing shared memory via private mapping will cause unrecoverable RMP
    page-faults.

    On kexec, walk direct mapping and convert all shared memory back to
    private. It makes all RAM private again and second kernel may use it
    normally. Additionally, for SNP guests, convert all bss decrypted
    section pages back to private.

    The conversion occurs in two steps: stopping new conversions and
    unsharing all memory. In the case of normal kexec, the stopping of
    conversions takes place while scheduling is still functioning. This
    allows for waiting until any ongoing conversions are finished. The
    second step is carried out when all CPUs except one are inactive and
    interrupts are disabled. This prevents any conflicts with code that may
    access shared memory.

    Co-developed-by: Borislav Petkov (AMD) <bp@alien8.de>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Signed-off-by: Ashish Kalra <ashish.kalra@amd.com>
    Reviewed-by: Tom Lendacky <thomas.lendacky@amd.com>
    Link: https://lore.kernel.org/r/05a8c15fb665dbb062b04a8cb3d592a63f235937.1722520012.git.ashish.kalra@amd.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-11 17:57:26 +01:00
Vitaly Kuznetsov 00af4de72c x86/mm: Refactor __set_clr_pte_enc()
JIRA: https://issues.redhat.com/browse/RHEL-33045

commit 2a783066b6f5f5250b838d2acfc716561d2a66e0
Author: Ashish Kalra <ashish.kalra@amd.com>
Date:   Thu Aug 1 19:14:34 2024 +0000

    x86/mm: Refactor __set_clr_pte_enc()

    Refactor __set_clr_pte_enc() and add two new helper functions to
    set/clear PTE C-bit from early SEV/SNP initialization code and later
    during shutdown/kexec especially when all CPUs are stopped and
    interrupts are disabled and set_memory_xx() interfaces can't be used.

    Co-developed-by: Borislav Petkov (AMD) <bp@alien8.de>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Signed-off-by: Ashish Kalra <ashish.kalra@amd.com>
    Reviewed-by: Tom Lendacky <thomas.lendacky@amd.com>
    Link: https://lore.kernel.org/r/5df4aa450447f28294d1c5a890e27b63ed4ded36.1722520012.git.ashish.kalra@amd.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-11 17:57:25 +01:00
Vitaly Kuznetsov 2347a2f004 x86/boot: Skip video memory access in the decompressor for SEV-ES/SNP
JIRA: https://issues.redhat.com/browse/RHEL-33045

commit f30470c190c2f4776e0baeba1f53fd8dd3820394
Author: Ashish Kalra <ashish.kalra@amd.com>
Date:   Thu Aug 1 19:14:17 2024 +0000

    x86/boot: Skip video memory access in the decompressor for SEV-ES/SNP

    Accessing guest video memory/RAM in the decompressor causes guest
    termination as the boot stage2 #VC handler for SEV-ES/SNP systems does
    not support MMIO handling.

    This issue is observed during a SEV-ES/SNP guest kexec as kexec -c adds
    screen_info to the boot parameters passed to the second kernel, which
    causes console output to be dumped to both video and serial.

    As the decompressor output gets cleared really fast, it is preferable to
    get the console output only on serial, hence, skip accessing the video
    RAM during decompressor stage to prevent guest termination.

    Serial console output during decompressor stage works as boot stage2 #VC
    handler already supports handling port I/O.

      [ bp: Massage. ]

    Suggested-by: Borislav Petkov (AMD) <bp@alien8.de>
    Suggested-by: Thomas Lendacky <thomas.lendacky@amd.com>
    Signed-off-by: Ashish Kalra <ashish.kalra@amd.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>
    Reviewed-by: Tom Lendacky <thomas.lendacky@amd.com>
    Link: https://lore.kernel.org/r/8a55ea86524c686e575d273311acbe57ce8cee23.1722520012.git.ashish.kalra@amd.com

Conflicts:
	arch/x86/boot/compressed/misc.c (skipping cd0d9d92c8bb4)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-11 17:57:25 +01:00
Vitaly Kuznetsov 619bc16ca7 uki: enable FIPS mode
JIRA: https://issues.redhat.com/browse/RHEL-37109
Upstream Status: RHEL-only

dracut-057-79.git20241127.el9 adds support for UKIs in the FIPS module,
enable it. Note: RHEL9 already ships 'fips=1' cmdline extension in
kernel-uki-virt-addons, this can now be used.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-11 15:55:52 +01:00
Vitaly Kuznetsov 7574c7f156 x86/tdx: mark TDX guest as fully supported
JIRA: https://issues.redhat.com/browse/RHEL-70465
Upstream Status: RHEL-only

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-12-09 16:44:23 +01:00
Vitaly Kuznetsov 1675a13de8 HID: hyperv: streamline driver probe to avoid devres issues
JIRA: https://issues.redhat.com/browse/RHEL-29299

commit 66ef47faa90d838cda131fe1f7776456cc3b59f2
Author: Vitaly Kuznetsov <vkuznets@redhat.com>
Date:   Mon Nov 11 14:12:40 2024 +0100

    HID: hyperv: streamline driver probe to avoid devres issues

    It was found that unloading 'hid_hyperv' module results in a devres
    complaint:

     ...
     hv_vmbus: unregistering driver hid_hyperv
     ------------[ cut here ]------------
     WARNING: CPU: 2 PID: 3983 at drivers/base/devres.c:691 devres_release_group+0x1f2/0x2c0
     ...
     Call Trace:
      <TASK>
      ? devres_release_group+0x1f2/0x2c0
      ? __warn+0xd1/0x1c0
      ? devres_release_group+0x1f2/0x2c0
      ? report_bug+0x32a/0x3c0
      ? handle_bug+0x53/0xa0
      ? exc_invalid_op+0x18/0x50
      ? asm_exc_invalid_op+0x1a/0x20
      ? devres_release_group+0x1f2/0x2c0
      ? devres_release_group+0x90/0x2c0
      ? rcu_is_watching+0x15/0xb0
      ? __pfx_devres_release_group+0x10/0x10
      hid_device_remove+0xf5/0x220
      device_release_driver_internal+0x371/0x540
      ? klist_put+0xf3/0x170
      bus_remove_device+0x1f1/0x3f0
      device_del+0x33f/0x8c0
      ? __pfx_device_del+0x10/0x10
      ? cleanup_srcu_struct+0x337/0x500
      hid_destroy_device+0xc8/0x130
      mousevsc_remove+0xd2/0x1d0 [hid_hyperv]
      device_release_driver_internal+0x371/0x540
      driver_detach+0xc5/0x180
      bus_remove_driver+0x11e/0x2a0
      ? __mutex_unlock_slowpath+0x160/0x5e0
      vmbus_driver_unregister+0x62/0x2b0 [hv_vmbus]
      ...

    And the issue seems to be that the corresponding devres group is not
    allocated. Normally, devres_open_group() is called from
    __hid_device_probe() but Hyper-V HID driver overrides 'hid_dev->driver'
    with 'mousevsc_hid_driver' stub and basically re-implements
    __hid_device_probe() by calling hid_parse() and hid_hw_start() but not
    devres_open_group(). hid_device_probe() does not call __hid_device_probe()
    for it. Later, when the driver is removed, hid_device_remove() calls
    devres_release_group() as it doesn't check whether hdev->driver was
    initially overridden or not.

    The issue seems to be related to the commit 62c68e7cee33 ("HID: ensure
    timely release of driver-allocated resources") but the commit itself seems
    to be correct.

    Fix the issue by dropping the 'hid_dev->driver' override and using
    hid_register_driver()/hid_unregister_driver() instead. Alternatively, it
    would have been possible to rely on the default handling but
    HID_CONNECT_DEFAULT implies HID_CONNECT_HIDRAW and it doesn't seem to work
    for mousevsc as-is.

    Fixes: 62c68e7cee33 ("HID: ensure timely release of driver-allocated resources")
    Suggested-by: Michael Kelley <mhklinux@outlook.com>
    Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
    Reviewed-by: Michael Kelley <mhklinux@outlook.com>
    Tested-by: Saurabh Sengar <ssengar@linux.microsoft.com>
    Signed-off-by: Jiri Kosina <jkosina@suse.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-11-25 17:34:35 +01:00
Vitaly Kuznetsov 9bcfc2b7b3 filter-modules.sh.rhel: Move squashfs to kernel-modules-core
JIRA: https://issues.redhat.com/browse/RHEL-68817
Upstream Status: RHEL-only

squashfs module is required for kdump but it currently lives in
kernel-modules subpackage which is optional. In some environments,
it is preferable to avoid installing additional modules to e.g. reduce the
attack surface. As kdump functionality is pretty basic, move squashfs to
kernel-modules-core. Note: kernel-modules package depend on
kernel-modules-core so there should be no difference for existing
environments where both packages are installed.

The change is RHEL9 only. RHEL10/ARK are switching to erofs instead.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-11-25 17:14:36 +01:00
Vitaly Kuznetsov 1d9adaffed KVM: selftests: Allow skipping the KVM_RUN sanity check in rseq_test
JIRA: https://issues.redhat.com/browse/RHEL-28186

commit 20ecf595b513b4ee69220794c3317380c4f051b1
Author: Zide Chen <zide.chen@intel.com>
Date:   Thu May 2 14:39:36 2024 -0700

    KVM: selftests: Allow skipping the KVM_RUN sanity check in rseq_test

    The rseq test's migration worker delays 1-10 us, assuming that one KVM_RUN
    iteration only takes a few microseconds.  But if the CPU low power wakeup
    latency is large enough, for example, hundreds or even thousands of
    microseconds for deep C-state exit latencies on x86 server CPUs, it may
    happen that the target CPU is unable to wakeup and run the vCPU before the
    migration worker starts to migrate the vCPU thread to the _next_ CPU.

    If the system workload is light, most CPUs could be at a certain low
    power state, which may result in less successful migrations and fail the
    migration/KVM_RUN ratio sanity check.  But this is not supposed to be
    deemed a test failure.

    Add a command line option to skip the sanity check, along with a comment
    and a verbose assert message to try to help the user resolve the potential
    source of failures without having to resort to disabling the check.

    Co-developed-by: Dongsheng Zhang <dongsheng.x.zhang@intel.com>
    Signed-off-by: Dongsheng Zhang <dongsheng.x.zhang@intel.com>
    Signed-off-by: Zide Chen <zide.chen@intel.com>
    Link: https://lore.kernel.org/r/20240502213936.27619-1-zide.chen@intel.com
    [sean: massage changelog]
    Signed-off-by: Sean Christopherson <seanjc@google.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-11-22 14:39:06 +01:00
Vitaly Kuznetsov 2b845ad56f redhat: create 'crashkernel=' addons for UKI
Upstream Status: RHEL-only
ARK commit 81d765e4091780d01f507a258e2cd0bd96175df8
JIRA: https://issues.redhat.com/browse/RHEL-33051

Create 'crashkernel-default' and specific 'crashkernel=' UKI addons. Note,
x86_64 and aarch64 have slightly different default options already, UKI
addons just copy the default kdump tools behavior.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-11-11 12:38:44 +01:00
Vitaly Kuznetsov 5a502e66f9 redhat: avoid superfluous quotes in UKI cmdline addones
Upstream Status: RHEL-only
ARK commit 36e02cd15671146950e53efab5115d849f0e48ea
JIRA: https://issues.redhat.com/browse/RHEL-33051

The way uki_create_addon.py runs 'ukify' results in superfluous quoting,
e.g.

 # cat /proc/cmdline
 console=tty0 console=ttyS0  "fips=0"

this seems to work in certain cases but is fragile as cmdline arguments may
be parsed both by kernel and userspace. Avoid the extra quoting by simply
switching from '--cmdline="arg"' to ['--cmdline', arg]. Spaces should be
handled correctly this way.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-11-11 12:38:11 +01:00
Vitaly Kuznetsov a5df78e9f9 xen-netfront: Fix NULL sring after live migration
JIRA: https://issues.redhat.com/browse/RHEL-63751
CVE: CVE-2022-48969

commit d50b7914fae04d840ce36491d22133070b18cca9
Author: Lin Liu <lin.liu@citrix.com>
Date:   Fri Dec 2 08:52:48 2022 +0000

    xen-netfront: Fix NULL sring after live migration

    A NAPI is setup for each network sring to poll data to kernel
    The sring with source host is destroyed before live migration and
    new sring with target host is setup after live migration.
    The NAPI for the old sring is not deleted until setup new sring
    with target host after migration. With busy_poll/busy_read enabled,
    the NAPI can be polled before got deleted when resume VM.

    BUG: unable to handle kernel NULL pointer dereference at
    0000000000000008
    IP: xennet_poll+0xae/0xd20
    PGD 0 P4D 0
    Oops: 0000 [#1] SMP PTI
    Call Trace:
     finish_task_switch+0x71/0x230
     timerqueue_del+0x1d/0x40
     hrtimer_try_to_cancel+0xb5/0x110
     xennet_alloc_rx_buffers+0x2a0/0x2a0
     napi_busy_loop+0xdb/0x270
     sock_poll+0x87/0x90
     do_sys_poll+0x26f/0x580
     tracing_map_insert+0x1d4/0x2f0
     event_hist_trigger+0x14a/0x260

     finish_task_switch+0x71/0x230
     __schedule+0x256/0x890
     recalc_sigpending+0x1b/0x50
     xen_sched_clock+0x15/0x20
     __rb_reserve_next+0x12d/0x140
     ring_buffer_lock_reserve+0x123/0x3d0
     event_triggers_call+0x87/0xb0
     trace_event_buffer_commit+0x1c4/0x210
     xen_clocksource_get_cycles+0x15/0x20
     ktime_get_ts64+0x51/0xf0
     SyS_ppoll+0x160/0x1a0
     SyS_ppoll+0x160/0x1a0
     do_syscall_64+0x73/0x130
     entry_SYSCALL_64_after_hwframe+0x41/0xa6
    ...
    RIP: xennet_poll+0xae/0xd20 RSP: ffffb4f041933900
    CR2: 0000000000000008
    ---[ end trace f8601785b354351c ]---

    xen frontend should remove the NAPIs for the old srings before live
    migration as the bond srings are destroyed

    There is a tiny window between the srings are set to NULL and
    the NAPIs are disabled, It is safe as the NAPI threads are still
    frozen at that time

    Signed-off-by: Lin Liu <lin.liu@citrix.com>
    Fixes: 4ec2411980 ([NET]: Do not check netif_running() and carrier state in ->poll())
    Signed-off-by: David S. Miller <davem@davemloft.net>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-10-24 13:45:16 +02:00
Vitaly Kuznetsov 40e5ec97d9 xen/netfront: destroy queues before real_num_tx_queues is zeroed
JIRA: https://issues.redhat.com/browse/RHEL-63751

commit dcf4ff7a48e7598e6b10126cc02177abb8ae4f3f
Author: Marek Marczykowski-Górecki <marmarek@invisiblethingslab.com>
Date:   Wed Feb 23 22:19:54 2022 +0100

    xen/netfront: destroy queues before real_num_tx_queues is zeroed

    xennet_destroy_queues() relies on info->netdev->real_num_tx_queues to
    delete queues. Since d7dac083414eb5bb99a6d2ed53dc2c1b405224e5
    ("net-sysfs: update the queue counts in the unregistration path"),
    unregister_netdev() indirectly sets real_num_tx_queues to 0. Those two
    facts together means, that xennet_destroy_queues() called from
    xennet_remove() cannot do its job, because it's called after
    unregister_netdev(). This results in kfree-ing queues that are still
    linked in napi, which ultimately crashes:

        BUG: kernel NULL pointer dereference, address: 0000000000000000
        #PF: supervisor read access in kernel mode
        #PF: error_code(0x0000) - not-present page
        PGD 0 P4D 0
        Oops: 0000 [#1] PREEMPT SMP PTI
        CPU: 1 PID: 52 Comm: xenwatch Tainted: G        W         5.16.10-1.32.fc32.qubes.x86_64+ #226
        RIP: 0010:free_netdev+0xa3/0x1a0
        Code: ff 48 89 df e8 2e e9 00 00 48 8b 43 50 48 8b 08 48 8d b8 a0 fe ff ff 48 8d a9 a0 fe ff ff 49 39 c4 75 26 eb 47 e8 ed c1 66 ff <48> 8b 85 60 01 00 00 48 8d 95 60 01 00 00 48 89 ef 48 2d 60 01 00
        RSP: 0000:ffffc90000bcfd00 EFLAGS: 00010286
        RAX: 0000000000000000 RBX: ffff88800edad000 RCX: 0000000000000000
        RDX: 0000000000000001 RSI: ffffc90000bcfc30 RDI: 00000000ffffffff
        RBP: fffffffffffffea0 R08: 0000000000000000 R09: 0000000000000000
        R10: 0000000000000000 R11: 0000000000000001 R12: ffff88800edad050
        R13: ffff8880065f8f88 R14: 0000000000000000 R15: ffff8880066c6680
        FS:  0000000000000000(0000) GS:ffff8880f3300000(0000) knlGS:0000000000000000
        CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
        CR2: 0000000000000000 CR3: 00000000e998c006 CR4: 00000000003706e0
        Call Trace:
         <TASK>
         xennet_remove+0x13d/0x300 [xen_netfront]
         xenbus_dev_remove+0x6d/0xf0
         __device_release_driver+0x17a/0x240
         device_release_driver+0x24/0x30
         bus_remove_device+0xd8/0x140
         device_del+0x18b/0x410
         ? _raw_spin_unlock+0x16/0x30
         ? klist_iter_exit+0x14/0x20
         ? xenbus_dev_request_and_reply+0x80/0x80
         device_unregister+0x13/0x60
         xenbus_dev_changed+0x18e/0x1f0
         xenwatch_thread+0xc0/0x1a0
         ? do_wait_intr_irq+0xa0/0xa0
         kthread+0x16b/0x190
         ? set_kthread_struct+0x40/0x40
         ret_from_fork+0x22/0x30
         </TASK>

    Fix this by calling xennet_destroy_queues() from xennet_uninit(),
    when real_num_tx_queues is still available. This ensures that queues are
    destroyed when real_num_tx_queues is set to 0, regardless of how
    unregister_netdev() was called.

    Originally reported at
    https://github.com/QubesOS/qubes-issues/issues/7257

    Fixes: d7dac083414eb5bb9 ("net-sysfs: update the queue counts in the unregistration path")
    Cc: stable@vger.kernel.org
    Signed-off-by: Marek Marczykowski-Górecki <marmarek@invisiblethingslab.com>
    Signed-off-by: David S. Miller <davem@davemloft.net>

Conflicts:
	drivers/net/xen-netfront.c (context - skipping b27d47950e481)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-10-24 13:44:11 +02:00
Vitaly Kuznetsov 65432dd956 x86/vmware: Add TDX hypercall support
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit 57b7b6acb41b51087ceb40c562efe392ec8c9677
Author: Alexey Makhalov <alexey.makhalov@broadcom.com>
Date:   Thu Jun 13 12:16:50 2024 -0700

    x86/vmware: Add TDX hypercall support

    VMware hypercalls use I/O port, VMCALL or VMMCALL instructions.  Add a call to
    __tdx_hypercall() in order to support TDX guests.

    No change in high bandwidth hypercalls, as only low bandwidth ones are supported
    for TDX guests.

      [ bp: Massage, clear on-stack struct tdx_module_args variable. ]

    Co-developed-by: Tim Merrifield <tim.merrifield@broadcom.com>
    Signed-off-by: Tim Merrifield <tim.merrifield@broadcom.com>
    Signed-off-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Link: https://lore.kernel.org/r/20240613191650.9913-9-alexey.makhalov@broadcom.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:04 +02:00
Vitaly Kuznetsov b99ac61bc6 x86/vmware: Remove legacy VMWARE_HYPERCALL* macros
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit 9dfb18031f0df2378b3d33a13fc485ef89caa285
Author: Alexey Makhalov <alexey.makhalov@broadcom.com>
Date:   Thu Jun 13 12:16:49 2024 -0700

    x86/vmware: Remove legacy VMWARE_HYPERCALL* macros

    No more direct use of these macros should be allowed. The vmware_hypercallX API
    still uses the new implementation of VMWARE_HYPERCALL macro internally, but it
    is not exposed outside of the vmware.h.

    Signed-off-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Link: https://lore.kernel.org/r/20240613191650.9913-8-alexey.makhalov@broadcom.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:04 +02:00
Vitaly Kuznetsov 9496e62656 x86/vmware: Correct macro names
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit 86cb65448d07fe516e18d9512ae5786cd90db9bf
Author: Alexey Makhalov <alexey.makhalov@broadcom.com>
Date:   Thu Jun 13 12:16:48 2024 -0700

    x86/vmware: Correct macro names

    VCPU_RESERVED and LEGACY_X2APIC are not VMware hypercall commands.  These are
    bits in the return value of the VMWARE_CMD_GETVCPU_INFO command.  Change
    VMWARE_CMD_ prefix to GETVCPU_INFO_ one. And move the bit-shift
    operation into the macro body.

    Fixes: 4cca6ea04d ("x86/apic: Allow x2apic without IR on VMware platform")
    Signed-off-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Link: https://lore.kernel.org/r/20240613191650.9913-7-alexey.makhalov@broadcom.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:04 +02:00
Vitaly Kuznetsov c9f3fe801d x86/vmware: Use VMware hypercall API
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit b2c13c23ea9c1f748315b8c2c028bb3ae18f1e12
Author: Alexey Makhalov <alexey.makhalov@broadcom.com>
Date:   Thu Jun 13 12:16:47 2024 -0700

    x86/vmware: Use VMware hypercall API

    Remove VMWARE_CMD macro and move to vmware_hypercall API.
    No functional changes intended.

    Use u32/u64 instead of uint32_t/uint64_t across the file.

    Signed-off-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Link: https://lore.kernel.org/r/20240613191650.9913-6-alexey.makhalov@broadcom.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:04 +02:00
Vitaly Kuznetsov 37c52a2e79 drm/vmwgfx: Use VMware hypercall API
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit 90328eaaff34f5617b3ec9603681b08d4a8e72df
Author: Alexey Makhalov <alexey.makhalov@broadcom.com>
Date:   Thu Jun 13 12:16:46 2024 -0700

    drm/vmwgfx: Use VMware hypercall API

    Switch from VMWARE_HYPERCALL macro to vmware_hypercall API. Eliminate arch
    specific code.

    drivers/gpu/drm/vmwgfx/vmwgfx_msg_arm64.h: implement arm64 variant
    of vmware_hypercall. And keep it here until introduction of ARM64
    VMWare hypervisor interface.

    Signed-off-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Link: https://lore.kernel.org/r/20240613191650.9913-5-alexey.makhalov@broadcom.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:04 +02:00
Vitaly Kuznetsov 7f7f4ef9c2 input/vmmouse: Use VMware hypercall API
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit f0db90b4127c0e454cd9a19ec4256221b974b819
Author: Alexey Makhalov <alexey.makhalov@broadcom.com>
Date:   Thu Jun 13 12:16:45 2024 -0700

    input/vmmouse: Use VMware hypercall API

    Switch from VMWARE_HYPERCALL macro to vmware_hypercall API.
    Eliminate arch specific code. No functional changes intended.

    Signed-off-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Link: https://lore.kernel.org/r/20240613191650.9913-4-alexey.makhalov@broadcom.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:04 +02:00
Vitaly Kuznetsov 6d1a51e32b ptp/vmware: Use VMware hypercall API
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit 54651bb4dcfea0949afe72775212511ec4193b85
Author: Alexey Makhalov <alexey.makhalov@broadcom.com>
Date:   Thu Jun 13 12:16:44 2024 -0700

    ptp/vmware: Use VMware hypercall API

    Switch from VMWARE_HYPERCALL macro to vmware_hypercall API.
    Eliminate arch specific code. No functional changes intended.

    Signed-off-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Link: https://lore.kernel.org/r/20240613191650.9913-3-alexey.makhalov@broadcom.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:04 +02:00
Vitaly Kuznetsov 15a93113fe x86/vmware: Introduce VMware hypercall API
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit 34bf25e820ae1ab38f9cd88834843ba76678a2fd
Author: Alexey Makhalov <alexey.makhalov@broadcom.com>
Date:   Thu Jun 13 12:16:43 2024 -0700

    x86/vmware: Introduce VMware hypercall API

    Introduce a vmware_hypercall family of functions. It is a common implementation
    to be used by the VMware guest code and virtual device drivers in architecture
    independent manner.

    The API consists of vmware_hypercallX and vmware_hypercall_hb_{out,in}
    set of functions analogous to KVM's hypercall API. Architecture-specific
    implementation is hidden inside.

    It will simplify future enhancements in VMware hypercalls such as SEV-ES and
    TDX related changes without needs to modify a caller in device drivers code.

    Current implementation extends an idea from

      bac7b4e843 ("x86/vmware: Update platform detection code for VMCALL/VMMCALL hypercalls")

    to have a slow, but safe path vmware_hypercall_slow() earlier during the boot
    when alternatives are not yet applied.  The code inherits VMWARE_CMD logic from
    the commit mentioned above.

    Move common macros from vmware.c to vmware.h.

      [ bp: Fold in a fix:
        https://lore.kernel.org/r/20240625083348.2299-1-alexey.makhalov@broadcom.com ]

    Signed-off-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Link: https://lore.kernel.org/r/20240613191650.9913-2-alexey.makhalov@broadcom.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:04 +02:00
Vitaly Kuznetsov f939aa6b9e drm/vmwgfx: Add unwind hints around RBP clobber
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit a9da8247627eefc73f909bf945031a5431a53993
Author: Josh Poimboeuf <jpoimboe@kernel.org>
Date:   Mon Jun 5 09:12:22 2023 -0700

    drm/vmwgfx: Add unwind hints around RBP clobber

    VMware high-bandwidth hypercalls take the RBP register as input.  This
    breaks basic frame pointer convention, as RBP should never be clobbered.

    So frame pointer unwinding is broken for the instructions surrounding
    the hypercalls.  Fortunately this doesn't break live patching with
    CONFIG_FRAME_POINTER, as it only unwinds from blocking tasks, and stack
    traces from preempted tasks are already marked unreliable anyway.

    However, for live patching with ORC, this could actually be a
    theoretical problem if vmw_port_hb_{in,out}() were still compiled with a
    frame pointer due to having an aligned stack.  In practice that hasn't
    seemed to be an issue since the objtool warnings have only been seen
    with CONFIG_FRAME_POINTER.

    Add unwind hint annotations to tell the ORC unwinder to mark stack
    traces as unreliable.

    Fixes the following warnings:

      vmlinux.o: warning: objtool: vmw_port_hb_in+0x1df: return with modified stack frame
      vmlinux.o: warning: objtool: vmw_port_hb_out+0x1dd: return with modified stack frame

    Fixes: 89da76fde6 ("drm/vmwgfx: Add VMWare host messaging capability")
    Reported-by: kernel test robot <lkp@intel.com>
    Link: https://lore.kernel.org/oe-kbuild-all/202305160135.97q0Elax-lkp@intel.com/
    Link: https://lore.kernel.org/r/4c795f2d87bc0391cf6543bcb224fa540b55ce4b.1685981486.git.jpoimboe@kernel.org
    Signed-off-by: Josh Poimboeuf <jpoimboe@kernel.org>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:03 +02:00
Vitaly Kuznetsov 1ab1227645 objtool: Allow stack operations in UNWIND_HINT_UNDEFINED regions
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit 1e4b619185e83e54aca617cf5070c64a88fe936b
Author: Josh Poimboeuf <jpoimboe@kernel.org>
Date:   Mon Jun 5 09:12:21 2023 -0700

    objtool: Allow stack operations in UNWIND_HINT_UNDEFINED regions

    If the code specified UNWIND_HINT_UNDEFINED, skip the "undefined stack
    state" warning due to a stack operation.  Just ignore the stack op and
    continue to propagate the undefined state.

    Link: https://lore.kernel.org/r/820c5b433f17c84e8761fb7465a8d319d706b1cf.1685981486.git.jpoimboe@kernel.org
    Signed-off-by: Josh Poimboeuf <jpoimboe@kernel.org>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:15:03 +02:00
Vitaly Kuznetsov a2cf11c056 x86,objtool: Split UNWIND_HINT_EMPTY in two
JIRA: https://issues.redhat.com/browse/RHEL-52683

commit fb799447ae2974a07907906dff5bd4b9e47b7123
Author: Josh Poimboeuf <jpoimboe@kernel.org>
Date:   Wed Mar 1 07:13:12 2023 -0800

    x86,objtool: Split UNWIND_HINT_EMPTY in two

    Mark reported that the ORC unwinder incorrectly marks an unwind as
    reliable when the unwind terminates prematurely in the dark corners of
    return_to_handler() due to lack of information about the next frame.

    The problem is UNWIND_HINT_EMPTY is used in two different situations:

      1) The end of the kernel stack unwind before hitting user entry, boot
         code, or fork entry

      2) A blind spot in ORC coverage where the unwinder has to bail due to
         lack of information about the next frame

    The ORC unwinder has no way to tell the difference between the two.
    When it encounters an undefined stack state with 'end=1', it blindly
    marks the stack reliable, which can break the livepatch consistency
    model.

    Fix it by splitting UNWIND_HINT_EMPTY into UNWIND_HINT_UNDEFINED and
    UNWIND_HINT_END_OF_STACK.

    Reported-by: Mark Rutland <mark.rutland@arm.com>
    Signed-off-by: Josh Poimboeuf <jpoimboe@kernel.org>
    Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Acked-by: Steven Rostedt (Google) <rostedt@goodmis.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/r/fd6212c8b450d3564b855e1cb48404d6277b4d9f.1677683419.git.jpoimboe@kernel.org

Conflicts:
	arch/x86/entry/entry_64.S (context, skipping f71e1d2ff8e6a)
	arch/x86/kernel/head_64.S (context, skipping 666e1156b2c51)

RHEL-only:
	arch/x86/entry/entry.S: UNWIND_HINT_EMPTY->UNWIND_HINT_UNDEFINED to
        match upstream.

Omitted-fix: b9f174c811e3 ("x86/unwind/orc: Add ELF section with ORC version identifier")
 see https://issues.redhat.com/browse/RHEL-27234 for discussion

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-18 15:14:40 +02:00
Vitaly Kuznetsov c0b52dafd8 uki: use systemd-pcrphase dracut module
Upstream Status: RHEL-only
JIRA: INTERNAL
ARK commit: 827cdc1f92cb4202c94a25cef5aa79177e9b65b9

dracut in Fedora 39 ships a systemd-pcrphase module.
Use that instead of copying over files manually via
install_items+="..."

Signed-off-by: Gerd Hoffmann <kraxel@redhat.com>

[RHEL9 note]: systemd-pcrphase is shipped with dracut >= 057-67.git20240812.el9

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-03 11:25:37 +02:00
Vitaly Kuznetsov d0d5a2c124 redhat: hmac sign the UKI for FIPS
JIRA: INTERNAL
Upstream Status: RHEL-only
ARK commit: c21494f10acecd92677eb363b38956a4994b2e29

Dracut's FIPS module contains kernel integrity check for traditional
kernels: /boot/vmlinuz-`uname-r`'s HMAC is compared to
/boot/.vmlinuz-`uname-r`.hmac which is created duing kernel
build. In preparation to enabling FIPS mode support for UKI, create
HMAC for the it too.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-02 18:02:54 +02:00
Vitaly Kuznetsov 03685f4585 Also make vmlinuz-virt.efi world readable
JIRA: INTERNAL
Upstream Status: RHEL-only
ARK commit: 231fb53f9cd6b3a678a4b3f9ae6952d99909b8c2

The file is created by dracut, which by default uses 0600, which is reasonable
for initrd images that contain machine-specific info, but is not needed for
packaged initrds which obviously cannot contain any secrets. Adjust the file
mode in the package list.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-09-02 18:02:46 +02:00
Vitaly Kuznetsov ce0a40c60a redhat: Fix the ownership of /lib/modules/<kversion> directory
Upstream Status: RHEL-only
ARK commit 77474b00e974f51edf1be49208795b9d1ecd4351
JIRA: https://issues.redhat.com/browse/RHEL-21034

/lib/modules/<kversion> is currently owned by 'kernel-core' so when this
package is not installed (e.g. UKI use-case), the empty directory is not
removed when kernel RPMs are removed. To make the removal work reliably
regardless of the installed RPM set and the removal order, make
/lib/modules/<kversion> (and /lib/modules for completeness) owned by:
- kernel-core
- kernel-modules-core
- kernel-uki-virt

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-08-12 14:14:25 +02:00
Vitaly Kuznetsov b8b23c1e40 Revert "xen/x2apic: enable x2apic mode when supported for HVM"
JIRA: https://issues.redhat.com/browse/RHEL-34602
Upstream Status: RHEL-only

This reverts commit d667c10dae.

Switching Xen HVM guests to x2apic mode led to CPU onlining problems on
AWS instance types. The issue is also reproducible with upstream kernel.
For the time being, revert RHEL9 change to avoid the regression.

'#include <asm/apic.h>' must be kept in enlighten_hvm.c to not break the
build.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-06-06 15:19:47 +02:00
Vitaly Kuznetsov 94c9dcdd0a x86/coco: Use CC_VENDOR_INTEL for Hyper-V/TDX
JIRA: https://issues.redhat.com/browse/RHEL-38910
Upstream Status: RHEL-only

RHEL backport 30467d1358 ("x86/coco: Get rid of accessor functions")
mistakenly used CC_VENDOR_AMD for Hyper-V/TDX, fix that.

Fixes: 30467d1358 ("x86/coco: Get rid of accessor functions")
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-06-05 18:14:24 +02:00
Vitaly Kuznetsov d454fc39ef xen-netfront: Add missing skb_mark_for_recycle
JIRA: https://issues.redhat.com/browse/RHEL-36573
CVE: CVE-2024-27393

commit 037965402a010898d34f4e35327d22c0a95cd51f
Author: Jesper Dangaard Brouer <hawk@kernel.org>
Date:   Wed Mar 27 13:14:56 2024 +0100

    xen-netfront: Add missing skb_mark_for_recycle

    Notice that skb_mark_for_recycle() is introduced later than fixes tag in
    commit 6a5bcd84e8 ("page_pool: Allow drivers to hint on SKB recycling").

    It is believed that fixes tag were missing a call to page_pool_release_page()
    between v5.9 to v5.14, after which is should have used skb_mark_for_recycle().
    Since v6.6 the call page_pool_release_page() were removed (in
    commit 535b9c61bdef ("net: page_pool: hide page_pool_release_page()")
    and remaining callers converted (in commit 6bfef2ec0172 ("Merge branch
    'net-page_pool-remove-page_pool_release_page'")).

    This leak became visible in v6.8 via commit dba1b8a7ab68 ("mm/page_pool: catch
    page_pool memory leaks").

    Cc: stable@vger.kernel.org
    Fixes: 6c5aa6fc4d ("xen networking: add basic XDP support for xen-netfront")
    Reported-by: Leonidas Spyropoulos <artafinde@archlinux.com>
    Link: https://bugzilla.kernel.org/show_bug.cgi?id=218654
    Reported-by: Arthur Borsboom <arthurborsboom@gmail.com>
    Signed-off-by: Jesper Dangaard Brouer <hawk@kernel.org>
    Link: https://lore.kernel.org/r/171154167446.2671062.9127105384591237363.stgit@firesoul
    Signed-off-by: Jakub Kicinski <kuba@kernel.org>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-05-21 17:45:07 +02:00
Vitaly Kuznetsov 4a68f53ca9 x86/xen: Add some null pointer checking to smp.c
JIRA: https://issues.redhat.com/browse/RHEL-33260
CVE: CVE-2024-26908

commit 3693bb4465e6e32a204a5b86d3ec7e6b9f7e67c2
Author: Kunwu Chan <chentao@kylinos.cn>
Date:   Fri Jan 19 17:49:48 2024 +0800

    x86/xen: Add some null pointer checking to smp.c

    kasprintf() returns a pointer to dynamically allocated memory
    which can be NULL upon failure. Ensure the allocation was successful
    by checking the pointer validity.

    Signed-off-by: Kunwu Chan <chentao@kylinos.cn>
    Reported-by: kernel test robot <lkp@intel.com>
    Closes: https://lore.kernel.org/oe-kbuild-all/202401161119.iof6BQsf-lkp@intel.com/
    Suggested-by: Markus Elfring <Markus.Elfring@web.de>
    Reviewed-by: Juergen Gross <jgross@suse.com>
    Link: https://lore.kernel.org/r/20240119094948.275390-1-chentao@kylinos.cn
    Signed-off-by: Juergen Gross <jgross@suse.com>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-05-21 16:25:16 +02:00
Vitaly Kuznetsov 6b3b23379f x86/sev: Harden #VC instruction emulation somewhat
JIRA: https://issues.redhat.com/browse/RHEL-30031
CVE: CVE-2024-25742
CVE: CVE-2024-25743

commit e3ef461af35a8c74f2f4ce6616491ddb355a208f
Author: Borislav Petkov (AMD) <bp@alien8.de>
Date:   Fri Jan 5 11:14:07 2024 +0100

    x86/sev: Harden #VC instruction emulation somewhat

    Compare the opcode bytes at rIP for each #VC exit reason to verify the
    instruction which raised the #VC exception is actually the right one.

    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Acked-by: Tom Lendacky <thomas.lendacky@amd.com>
    Link: https://lore.kernel.org/r/20240105101407.11694-1-bp@alien8.de

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-25 11:56:28 +01:00
Vitaly Kuznetsov 8f8ca9ff07 x86/fpu/xstate: Fix PKRU covert channel
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 18032b47adf1db7b7f5fb2d1344e65aafe6417df
Author: Jim Mattson <jmattson@google.com>
Date:   Wed Aug 30 21:32:21 2023 -0700

    x86/fpu/xstate: Fix PKRU covert channel

    When XCR0[9] is set, PKRU can be read and written from userspace with
    XSAVE and XRSTOR, even when CR4.PKE is clear.

    Clear XCR0[9] when protection keys are disabled.

    Reported-by: Tavis Ormandy <taviso@google.com>
    Signed-off-by: Jim Mattson <jmattson@google.com>
    Signed-off-by: Ingo Molnar <mingo@kernel.org>
    Acked-by: Dave Hansen <dave.hansen@linux.intel.com>
    Link: https://lore.kernel.org/r/20230831043228.1194256-1-jmattson@google.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:28 -04:00
Vitaly Kuznetsov 483d40c2af x86/mm: fix poking_init() for Xen PV guests
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 26ce6ec364f18d2915923bc05784084e54a5c4cc
Author: Juergen Gross <jgross@suse.com>
Date:   Mon Jan 9 16:09:22 2023 +0100

    x86/mm: fix poking_init() for Xen PV guests

    Commit 3f4c8211d982 ("x86/mm: Use mm_alloc() in poking_init()") broke
    the kernel for running as Xen PV guest.

    It seems as if the new address space is never activated before being
    used, resulting in Xen rejecting to accept the new CR3 value (the PGD
    isn't pinned).

    Fix that by adding the now missing call of paravirt_arch_dup_mmap() to
    poking_init(). That call was previously done by dup_mm()->dup_mmap() and
    it is a NOP for all cases but for Xen PV, where it is just doing the
    pinning of the PGD.

    Fixes: 3f4c8211d982 ("x86/mm: Use mm_alloc() in poking_init()")
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lkml.kernel.org/r/20230109150922.10578-1-jgross@suse.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:28 -04:00
Vitaly Kuznetsov 2b51fbf27b x86/sev: Move sev_setup_arch() to mem_encrypt.c
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 6e74b125155dc8c747d76fb45d8e6d20e9e4fb4d
Author: Alexander Shishkin <alexander.shishkin@linux.intel.com>
Date:   Tue Oct 10 17:52:19 2023 +0300

    x86/sev: Move sev_setup_arch() to mem_encrypt.c

    Since commit:

      4d96f9109109b ("x86/sev: Replace occurrences of sev_active() with cc_platform_has()")

    ... the SWIOTLB bounce buffer size adjustment and restricted virtio memory
    setting also inadvertently apply to TDX: the code is using
    cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT) as a gatekeeping condition,
    which is also true for TDX, and this is also what we want.

    To reflect this, move the corresponding code to generic mem_encrypt.c.

    No functional changes intended.

    Signed-off-by: Alexander Shishkin <alexander.shishkin@linux.intel.com>
    Signed-off-by: Ingo Molnar <mingo@kernel.org>
    Reviewed-by: Tom Lendacky <thomas.lendacky@amd.com>
    Link: https://lore.kernel.org/r/20231010145220.3960055-2-alexander.shishkin@linux.intel.com

[pbonzini: Conflict: RHEL uses arch_has_restricted_virtio_memory_access
 instead of virtio_set_mem_acc_cb]

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov b6c3ec8d96 x86/mem_encrypt: Remove stale mem_encrypt_init() declaration
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 1b2c92a1cb2469d8c0079dbf496ab86e22e1cb7c
Author: Linus Torvalds <torvalds@linux-foundation.org>
Date:   Wed Jun 28 12:47:30 2023 -0700

    x86/mem_encrypt: Remove stale mem_encrypt_init() declaration

    The memory encryption initialization logic was moved from init/main.c
    into arch_cpu_finalize_init() in commit 439e17576eb4 ("init, x86: Move
    mem_encrypt_init() into arch_cpu_finalize_init()"), but a stale
    declaration for the init function was left in <linux/init.h>.

    And didn't cause any problems if you had X86_MEM_ENCRYPT enabled, which
    apparently everybody involved did have.  See also commit 0a9567ac5e6a
    ("x86/mem_encrypt: Unbreak the AMD_MEM_ENCRYPT=n build") in this whole
    sad saga of conflicting declarations for different situations.

    Reported-by: Matthew Wilcox <willy@infradead.org>
    Fixes: 439e17576eb4 init, x86: Move mem_encrypt_init() into arch_cpu_finalize_init()
    Cc: Thomas Gleixner <tglx@linutronix.de>
    Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov 6f71a3f365 x86/mem_encrypt: Unbreak the AMD_MEM_ENCRYPT=n build
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 0a9567ac5e6a40cdd9c8cd15b19a62a15250f450
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Fri Jun 16 22:15:31 2023 +0200

    x86/mem_encrypt: Unbreak the AMD_MEM_ENCRYPT=n build

    Moving mem_encrypt_init() broke the AMD_MEM_ENCRYPT=n because the
    declaration of that function was under #ifdef CONFIG_AMD_MEM_ENCRYPT and
    the obvious placement for the inline stub was the #else path.

    This is a leftover of commit 20f07a044a76 ("x86/sev: Move common memory
    encryption code to mem_encrypt.c") which made mem_encrypt_init() depend on
    X86_MEM_ENCRYPT without moving the prototype. That did not fail back then
    because there was no stub inline as the core init code had a weak function.

    Move both the declaration and the stub out of the CONFIG_AMD_MEM_ENCRYPT
    section and guard it with CONFIG_X86_MEM_ENCRYPT.

    Fixes: 439e17576eb4 ("init, x86: Move mem_encrypt_init() into arch_cpu_finalize_init()")
    Reported-by: kernel test robot <lkp@intel.com>
    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Closes: https://lore.kernel.org/oe-kbuild-all/202306170247.eQtCJPE8-lkp@intel.com/

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov b9e4c5b5d5 init, x86: Move mem_encrypt_init() into arch_cpu_finalize_init()
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 439e17576eb47f26b78c5bbc72e344d4206d2327
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Wed Jun 14 01:39:41 2023 +0200

    init, x86: Move mem_encrypt_init() into arch_cpu_finalize_init()

    Invoke the X86ism mem_encrypt_init() from X86 arch_cpu_finalize_init() and
    remove the weak fallback from the core code.

    No functional change.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Link: https://lore.kernel.org/r/20230613224545.670360645@linutronix.de

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov 3856dae45f x86/fpu: Mark init functions __init
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 1703db2b90c91b2eb2d699519fc505fe431dde0e
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Wed Jun 14 01:39:45 2023 +0200

    x86/fpu: Mark init functions __init

    No point in keeping them around.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Link: https://lore.kernel.org/r/20230613224545.841685728@linutronix.de

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov 8cbd84d163 x86/fpu: Set X86_FEATURE_OSXSAVE feature after enabling OSXSAVE in CR4
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 2c66ca3949dc701da7f4c9407f2140ae425683a5
Author: Feng Tang <feng.tang@intel.com>
Date:   Wed Aug 23 14:57:47 2023 +0800

    x86/fpu: Set X86_FEATURE_OSXSAVE feature after enabling OSXSAVE in CR4

    0-Day found a 34.6% regression in stress-ng's 'af-alg' test case, and
    bisected it to commit b81fac906a8f ("x86/fpu: Move FPU initialization into
    arch_cpu_finalize_init()"), which optimizes the FPU init order, and moves
    the CR4_OSXSAVE enabling into a later place:

       arch_cpu_finalize_init
           identify_boot_cpu
               identify_cpu
                   generic_identify
                       get_cpu_cap --> setup cpu capability
           ...
           fpu__init_cpu
               fpu__init_cpu_xstate
                   cr4_set_bits(X86_CR4_OSXSAVE);

    As the FPU is not yet initialized the CPU capability setup fails to set
    X86_FEATURE_OSXSAVE. Many security module like 'camellia_aesni_avx_x86_64'
    depend on this feature and therefore fail to load, causing the regression.

    Cure this by setting X86_FEATURE_OSXSAVE feature right after OSXSAVE
    enabling.

    [ tglx: Moved it into the actual BSP FPU initialization code and added a comment ]

    Fixes: b81fac906a8f ("x86/fpu: Move FPU initialization into arch_cpu_finalize_init()")
    Reported-by: kernel test robot <oliver.sang@intel.com>
    Signed-off-by: Feng Tang <feng.tang@intel.com>
    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Cc: stable@vger.kernel.org
    Link: https://lore.kernel.org/lkml/202307192135.203ac24e-oliver.sang@intel.com
    Link: https://lore.kernel.org/lkml/20230823065747.92257-1-feng.tang@intel.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov b0e902bcdf x86/xen: Fix secondary processors' FPU initialization
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit fe3e0a13e597c1c8617814bf9b42ab732db5c26e
Author: Juergen Gross <jgross@suse.com>
Date:   Mon Jul 3 15:00:32 2023 +0200

    x86/xen: Fix secondary processors' FPU initialization

    Moving the call of fpu__init_cpu() from cpu_init() to start_secondary()
    broke Xen PV guests, as those don't call start_secondary() for APs.

    Call fpu__init_cpu() in Xen's cpu_bringup(), which is the Xen PV
    replacement of start_secondary().

    Fixes: b81fac906a8f ("x86/fpu: Move FPU initialization into arch_cpu_finalize_init()")
    Signed-off-by: Juergen Gross <jgross@suse.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Reviewed-by: Boris Ostrovsky <boris.ostrovsky@oracle.com>
    Acked-by: Thomas Gleixner <tglx@linutronix.de>
    Link: https://lore.kernel.org/r/20230703130032.22916-1-jgross@suse.com

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov 4a9ea6c7b3 x86/efi: Make efi_set_virtual_address_map IBT safe
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 0303c9729afc4094ef53e552b7b8cff7436028d6
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Thu Jun 29 21:35:19 2023 +0200

    x86/efi: Make efi_set_virtual_address_map IBT safe

    Niklāvs reported a boot regression on an Alderlake machine and bisected it
    to commit 9df9d2f0471b ("init: Invoke arch_cpu_finalize_init() earlier").

    By moving the invocation of arch_cpu_finalize_init() further down he
    identified that efi_enter_virtual_mode() is the function which causes the
    boot hang.

    The main difference of the earlier invocation is that the boot CPU is
    already fully initialized and mitigations and alternatives are applied.

    But the only really interesting change turned out to be IBT, which is now
    enabled before efi_enter_virtual_mode(). "ibt=off" on the kernel command
    line cured the problem.

    Inspection of the involved calls in efi_enter_virtual_mode() unearthed that
    efi_set_virtual_address_map() is the only place in the kernel which invokes
    an EFI call without the IBT safe wrapper. This went obviously unnoticed so
    far as IBT was enabled later.

    Use arch_efi_call_virt() instead of efi_call() to cure that.

    Fixes: fe379fa4d199 ("x86/ibt: Disable IBT around firmware")
    Fixes: 9df9d2f0471b ("init: Invoke arch_cpu_finalize_init() earlier")
    Reported-by: Niklāvs Koļesņikovs <pinkflames.linux@gmail.com>
    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Ard Biesheuvel <ardb@kernel.org>
    Link: https://bugzilla.kernel.org/show_bug.cgi?id=217602
    Link: https://lore.kernel.org/r/87jzvm12q0.ffs@tglx

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov d2c661e8c2 x86/fpu: Move FPU initialization into arch_cpu_finalize_init()
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit b81fac906a8f9e682e513ddd95697ec7a20878d4
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Wed Jun 14 01:39:46 2023 +0200

    x86/fpu: Move FPU initialization into arch_cpu_finalize_init()

    Initializing the FPU during the early boot process is a pointless
    exercise. Early boot is convoluted and fragile enough.

    Nothing requires that the FPU is set up early. It has to be initialized
    before fork_init() because the task_struct size depends on the FPU register
    buffer size.

    Move the initialization to arch_cpu_finalize_init() which is the perfect
    place to do so.

    No functional change.

    This allows to remove quite some of the custom early command line parsing,
    but that's subject to the next installment.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Link: https://lore.kernel.org/r/20230613224545.902376621@linutronix.de

Conflicts:
	arch/x86/kernel/cpu/common.c (Upstream merge conflict, see 9244724fbf8ab)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov 1c98ccce12 init: Invoke arch_cpu_finalize_init() earlier
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 9df9d2f0471b4c4702670380b8d8a45b40b23a7d
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Wed Jun 14 01:39:39 2023 +0200

    init: Invoke arch_cpu_finalize_init() earlier

    X86 is reworking the boot process so that initializations which are not
    required during early boot can be moved into the late boot process and out
    of the fragile and restricted initial boot phase.

    arch_cpu_finalize_init() is the obvious place to do such initializations,
    but arch_cpu_finalize_init() is invoked too late in start_kernel() e.g. for
    initializing the FPU completely. fork_init() requires that the FPU is
    initialized as the size of task_struct on X86 depends on the size of the
    required FPU register buffer.

    Fortunately none of the init calls between calibrate_delay() and
    arch_cpu_finalize_init() is relevant for the functionality of
    arch_cpu_finalize_init().

    Invoke it right after calibrate_delay() where everything which is relevant
    for arch_cpu_finalize_init() has been set up already.

    No functional change intended.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Rick Edgecombe <rick.p.edgecombe@intel.com>
    Link: https://lore.kernel.org/r/20230613224545.612182854@linutronix.de

Conflicts:
	init/main.c ((out of order 7725acaa4f0c backport)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov 6ce2f42707 x86/init: Initialize signal frame size late
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 54d9a91a3d6713d1332e93be13b4eaf0fa54349d
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Wed Jun 14 01:39:42 2023 +0200

    x86/init: Initialize signal frame size late

    No point in doing this during really early boot. Move it to an early
    initcall so that it is set up before possible user mode helpers are started
    during device initialization.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Link: https://lore.kernel.org/r/20230613224545.727330699@linutronix.de

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov 2b617010d4 x86/fpu: Remove cpuinfo argument from init functions
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 1f34bb2a24643e0087652d81078e4f616562738d
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Wed Jun 14 01:39:43 2023 +0200

    x86/fpu: Remove cpuinfo argument from init functions

    Nothing in the call chain requires it

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Link: https://lore.kernel.org/r/20230613224545.783704297@linutronix.de

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov b61a76efa8 x86/mm: Initialize text poking earlier
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 5b93a83649c7cba3a15eb7e8959b250841acb1b1
Author: Peter Zijlstra <peterz@infradead.org>
Date:   Tue Oct 25 21:38:25 2022 +0200

    x86/mm: Initialize text poking earlier

    Move poking_init() up a bunch; specifically move it right after
    mm_init() which is right before ftrace_init().

    This will allow simplifying ftrace text poking which currently has
    a bunch of exceptions for early boot.

    Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lkml.kernel.org/r/20221025201057.881703081@infradead.org

Conflicts:
	init/main.c (out of order 7725acaa4f0c backport)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:27 -04:00
Vitaly Kuznetsov be72381490 x86/mm: Use mm_alloc() in poking_init()
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit 3f4c8211d982099be693be9aa7d6fc4607dff290
Author: Peter Zijlstra <peterz@infradead.org>
Date:   Tue Oct 25 21:38:21 2022 +0200

    x86/mm: Use mm_alloc() in poking_init()

    Instead of duplicating init_mm, allocate a fresh mm. The advantage is
    that mm_alloc() has much simpler dependencies. Additionally it makes
    more conceptual sense, init_mm has no (and must not have) user state
    to duplicate.

    Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lkml.kernel.org/r/20221025201057.816175235@infradead.org

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:26 -04:00
Vitaly Kuznetsov 037738e296 mm: Move mm_cachep initialization to mm_init()
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit af80602799681c78f14fbe20b6185a56020dedee
Author: Peter Zijlstra <peterz@infradead.org>
Date:   Tue Oct 25 21:38:18 2022 +0200

    mm: Move mm_cachep initialization to mm_init()

    In order to allow using mm_alloc() much earlier, move initializing
    mm_cachep into mm_init().

    Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lkml.kernel.org/r/20221025201057.751153381@infradead.org

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:26 -04:00
Vitaly Kuznetsov e3d9129d60 init: consolidate prototypes in linux/init.h
JIRA: https://issues.redhat.com/browse/RHEL-25415

commit ad1a48301f659a02df5bff0a121d4a5c0411d36b
Author: Arnd Bergmann <arnd@arndb.de>
Date:   Wed May 17 15:10:59 2023 +0200

    init: consolidate prototypes in linux/init.h

    The init/main.c file contains some extern declarations for functions
    defined in architecture code, and it defines some other functions that are
    called from architecture code with a custom prototype.  Both of those
    result in warnings with 'make W=1':

    init/calibrate.c:261:37: error: no previous prototype for 'calibrate_delay_is_known' [-Werror=missing-prototypes]
    init/main.c:790:20: error: no previous prototype for 'mem_encrypt_init' [-Werror=missing-prototypes]
    init/main.c:792:20: error: no previous prototype for 'poking_init' [-Werror=missing-prototypes]
    arch/arm64/kernel/irq.c:122:13: error: no previous prototype for 'init_IRQ' [-Werror=missing-prototypes]
    arch/arm64/kernel/time.c:55:13: error: no previous prototype for 'time_init' [-Werror=missing-prototypes]
    arch/x86/kernel/process.c:935:13: error: no previous prototype for 'arch_post_acpi_subsys_init' [-Werror=missing-prototypes]
    init/calibrate.c:261:37: error: no previous prototype for 'calibrate_delay_is_known' [-Werror=missing-prototypes]
    kernel/fork.c:991:20: error: no previous prototype for 'arch_task_cache_init' [-Werror=missing-prototypes]

    Add prototypes for all of these in include/linux/init.h or another
    appropriate header, and remove the duplicate declarations from
    architecture specific code.

    [sfr@canb.auug.org.au: declare time_init_early()]
      Link: https://lkml.kernel.org/r/20230519124311.5167221c@canb.auug.org.au
    Link: https://lkml.kernel.org/r/20230517131102.934196-12-arnd@kernel.org
    Signed-off-by: Arnd Bergmann <arnd@arndb.de>
    Signed-off-by: Stephen Rothwell <sfr@canb.auug.org.au>
    Cc: Boqun Feng <boqun.feng@gmail.com>
    Cc: Catalin Marinas <catalin.marinas@arm.com>
    Cc: Christoph Lameter <cl@linux.com>
    Cc: Dennis Zhou <dennis@kernel.org>
    Cc: Eric Paris <eparis@redhat.com>
    Cc: Heiko Carstens <hca@linux.ibm.com>
    Cc: Helge Deller <deller@gmx.de>
    Cc: Ingo Molnar <mingo@redhat.com>
    Cc: Michael Ellerman <mpe@ellerman.id.au>
    Cc: Michal Simek <monstr@monstr.eu>
    Cc: Palmer Dabbelt <palmer@dabbelt.com>
    Cc: Paul Moore <paul@paul-moore.com>
    Cc: Pavel Machek <pavel@ucw.cz>
    Cc: Peter Zijlstra <peterz@infradead.org>
    Cc: Rafael J. Wysocki <rafael@kernel.org>
    Cc: Russell King <linux@armlinux.org.uk>
    Cc: Tejun Heo <tj@kernel.org>
    Cc: Thomas Bogendoerfer <tsbogend@alpha.franken.de>
    Cc: Thomas Gleixner <tglx@linutronix.de>
    Cc: Waiman Long <longman@redhat.com>
    Cc: Will Deacon <will@kernel.org>
    Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Conflicts:
	arch/powerpc/include/asm/irq.h
	arch/riscv/include/asm/irq.h
	(unsupported in RHEL)

RHEL-only: add #include <linux/maple_tree.h> to main.c as RHEL lacks
d4af56c5c7c67 backport.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-20 09:42:26 -04:00
Vitaly Kuznetsov 79d3676c47 x86/mm: Fix memory encryption features advertisement
JIRA: https://issues.redhat.com/browse/RHEL-26662

commit 4cab62c058f5a150d9960c112362e5c76d204d9d
Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Date:   Wed Jan 24 16:02:16 2024 +0200

    x86/mm: Fix memory encryption features advertisement

    When memory encryption is enabled, the kernel prints the encryption
    flavor that the system supports.

    The check assumes that everything is AMD SME/SEV if it doesn't have
    the TDX CPU feature set.

    Hyper-V vTOM sets cc_vendor to CC_VENDOR_INTEL when it runs as L2 guest
    on top of TDX, but not X86_FEATURE_TDX_GUEST. Hyper-V only needs memory
    encryption enabled for I/O without the rest of CoCo enabling.

    To avoid confusion, check the cc_vendor directly.

      [ bp: Massage commit message. ]

    Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Reviewed-by: Jeremi Piotrowski <jpiotrowski@linux.microsoft.com>
    Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>
    Acked-by: Tom Lendacky <thomas.lendacky@amd.com>
    Acked-by: Kai Huang <kai.huang@intel.com>
    Link: https://lore.kernel.org/r/20240124140217.533748-1-kirill.shutemov@linux.intel.com

Conflicts:
	arch/x86/mm/mem_encrypt.c (preserve RHEL-specific mark_tech_preview()).
	Put it _after_ pr_cont() to not break the output.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-13 09:43:54 +01:00
Vitaly Kuznetsov faaabb64bc x86/coco: Export cc_vendor
JIRA: https://issues.redhat.com/browse/RHEL-26662

commit 3d91c537296794d5d0773f61abbe7b63f2f132d8
Author: Borislav Petkov (AMD) <bp@alien8.de>
Date:   Sat Mar 18 12:56:33 2023 +0100

    x86/coco: Export cc_vendor

    It will be used in different checks in future changes. Export it directly
    and provide accessor functions and stubs so this can be used in general
    code when CONFIG_ARCH_HAS_CC_PLATFORM is not set.

    No functional changes.

    [ tglx: Add accessor functions ]

    Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Link: https://lore.kernel.org/r/20230318115634.9392-2-bp@alien8.de

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-03-12 15:47:25 +01:00
Vitaly Kuznetsov 699a390e23 misc/pvpanic: fix set driver data
JIRA: https://issues.redhat.com/browse/RHEL-22993

commit a99009bc4f2f0b46e6c553704fda0b67e04395f5
Author: Mihai Carabas <mihai.carabas@oracle.com>
Date:   Thu Aug 19 18:12:26 2021 +0300

    misc/pvpanic: fix set driver data

    Add again dev_set_drvdata(), but this time in devm_pvpanic_probe(), in order
    for dev_get_drvdata() to not return NULL.

    Fixes: 394febc9d0 ("misc/pvpanic: Make 'pvpanic_probe()' resource managed")
    Reviewed-by: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
    Signed-off-by: Mihai Carabas <mihai.carabas@oracle.com>
    Link: https://lore.kernel.org/r/1629385946-4584-2-git-send-email-mihai.carabas@oracle.com
    Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-29 16:06:26 +01:00
Vitaly Kuznetsov d8e5da8388 redhat: Use dracut instead of objcopy for adding SBAT information to UKI
JIRA: INTERNAL

Upstream Status: RHEL only

dracut >= 057-51.git20231114.el9 supports adding SBAT information to UKI,
use the feature instead of open coding it with objcopy.

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-23 14:51:35 +01:00
Vitaly Kuznetsov 00ee9f7052 efi/unaccepted: Fix off-by-one when checking for overlapping ranges
JIRA: https://issues.redhat.com/browse/RHEL-19178

When a task needs to accept memory it will scan the accepting_list
to see if any ranges already being processed by other tasks overlap
with its range. Due to an off-by-one in the range comparisons, a task
might falsely determine that an overlapping range is being accepted,
leading to an unnecessary delay before it begins processing the range.

Fix the off-by-one in the range comparison to prevent this and slightly
improve performance.

Fixes: 50e782a86c98 ("efi/unaccepted: Fix soft lockups caused by parallel memory acceptance")
Link: https://lore.kernel.org/linux-mm/20231101004523.vseyi5bezgfaht5i@amd.com/T/#me2eceb9906fcae5fe958b3fe88e41f920f8335b6
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Signed-off-by: Michael Roth <michael.roth@amd.com>
Acked-by: Vlastimil Babka <vbabka@suse.cz>
Signed-off-by: Ard Biesheuvel <ardb@kernel.org>
(cherry picked from commit 01b1e3ca0e5ce47bbae8217d47376ad01b331b07)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:36:05 +01:00
Vitaly Kuznetsov 026eb11a37 x86/traps: Fix load_unaligned_zeropad() handling for shared TDX memory
JIRA: https://issues.redhat.com/browse/RHEL-19178

Commit c4e34dd99f2e ("x86: simplify load_unaligned_zeropad()
implementation") changes how exceptions around load_unaligned_zeropad()
handled.  The kernel now uses the fault_address in fixup_exception() to
verify the address calculations for the load_unaligned_zeropad().

It works fine for #PF, but breaks on #VE since no fault address is
passed down to fixup_exception().

Propagating ve_info.gla down to fixup_exception() resolves the issue.

See commit 1e7769653b06 ("x86/tdx: Handle load_unaligned_zeropad()
page-cross to a shared page") for more context.

Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Reported-by: Michael Kelley <mikelley@microsoft.com>
Fixes: c4e34dd99f2e ("x86: simplify load_unaligned_zeropad() implementation")
Acked-by: Dave Hansen <dave.hansen@linux.intel.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
(cherry picked from commit 9f9116406120638b4d8db3831ffbc430dd2e1e95)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:56 +01:00
Vitaly Kuznetsov 51512ad882 x86/tdx: Fix __noreturn build warning around __tdx_hypercall_failed()
JIRA: https://issues.redhat.com/browse/RHEL-19178

LKP reported below build warning:

  vmlinux.o: warning: objtool: __tdx_hypercall+0x128: __tdx_hypercall_failed() is missing a __noreturn annotation

The __tdx_hypercall_failed() function definition already has __noreturn
annotation, but it turns out the __noreturn must be annotated to the
function declaration.

PeterZ explains:

  "FWIW, the reason being that...

   The point of noreturn is that the caller should know to stop generating
   code. For that the declaration needs the attribute, because call sites
   typically do not have access to the function definition in C."

Add __noreturn annotation to the declaration of __tdx_hypercall_failed()
to fix.  It's not a bad idea to document the __noreturn nature at the
definition site either, so keep the annotation at the definition.

Note <asm/shared/tdx.h> is also included by TDX related assembly files.
Include <linux/compiler_attributes.h> only in case of !__ASSEMBLY__
otherwise compiling assembly file would trigger build error.

Also, following the objtool documentation, add __tdx_hypercall_failed()
to "tools/objtool/noreturns.h".

Fixes: c641cfb5c157 ("x86/tdx: Make TDX_HYPERCALL asm similar to TDX_MODULE_CALL")
Reported-by: kernel test robot <lkp@intel.com>
Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Link: https://lore.kernel.org/r/20230918041858.331234-1-kai.huang@intel.com
Closes: https://lore.kernel.org/oe-kbuild-all/202309140828.9RdmlH2Z-lkp@intel.com/
(cherry picked from commit 518755a7eeae77a399430eaf211a1e71f6b87d4a)

Conflicts:
	tools/objtool/noreturns.h (not in RHEL, dropped)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov 4384aac62f x86/tdx: Replace deprecated strncpy() with strtomem_pad()
JIRA: https://issues.redhat.com/browse/RHEL-19178

strncpy() works perfectly here in all cases, however, it is deprecated and
as such we should prefer more robust and less ambiguous string APIs:

    https://www.kernel.org/doc/html/latest/process/deprecated.html#strncpy-on-nul-terminated-strings

Let's use strtomem_pad() as this matches the functionality of strncpy()
and is _not_ deprecated.

Signed-off-by: Justin Stitt <justinstitt@google.com>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Kees Cook <keescook@chromium.org>
Acked-by: Dave Hansen <dave.hansen@linux.intel.com>
Link: https://github.com/KSPP/linux/issues/90
Link: https://lore.kernel.org/r/20231003-strncpy-arch-x86-coco-tdx-tdx-c-v2-1-0bd21174a217@google.com
(cherry picked from commit c9babd5d95abf3fae6e798605ce5cac98e08daf9)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov 73b9269426 x86/tdx: Remove 'struct tdx_hypercall_args'
JIRA: https://issues.redhat.com/browse/RHEL-19178

Now 'struct tdx_hypercall_args' is basically 'struct tdx_module_args'
minus RCX.  Although from __tdx_hypercall()'s perspective RCX isn't
used as shared register thus not part of input/output registers, it's
not worth to have a separate structure just due to one register.

Remove the 'struct tdx_hypercall_args' and use 'struct tdx_module_args'
instead in __tdx_hypercall() related code.  This also saves the memory
copy between the two structures within __tdx_hypercall().

Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://lore.kernel.org/all/798dad5ce24e9d745cf0e16825b75ccc433ad065.1692096753.git.kai.huang%40intel.com
(cherry picked from commit 8a8544bde858e5d62d79df6baaa387e0b6587dc7)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov ec7ecc863d x86/tdx: Reimplement __tdx_hypercall() using TDX_MODULE_CALL asm
JIRA: https://issues.redhat.com/browse/RHEL-19178

Now the TDX_HYPERCALL asm is basically identical to the TDX_MODULE_CALL
with both '\saved' and '\ret' enabled, with two minor things though:

1) The way to restore the structure pointer is different

The TDX_HYPERCALL uses RCX as spare to restore the structure pointer,
but the TDX_MODULE_CALL assumes no spare register can be used.  In other
words, TDX_MODULE_CALL already covers what TDX_HYPERCALL does.

2) TDX_MODULE_CALL only clears shared registers for TDH.VP.ENTER

For this just need to make that code available for the non-host case.

Thus, remove the TDX_HYPERCALL and reimplement the __tdx_hypercall()
using the TDX_MODULE_CALL.

Extend the TDX_MODULE_CALL to cover "clear shared registers" for
TDG.VP.VMCALL.  Introduce a new __tdcall_saved_ret() to replace the
temporary __tdcall_hypercall().

The __tdcall_saved_ret() can also be used for those new TDCALLs which
require more input/output registers than the basic TDCALLs do.

Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://lore.kernel.org/all/e68a2473fb6f5bcd78b078cae7510e9d0753b3df.1692096753.git.kai.huang%40intel.com
(cherry picked from commit 90f5ecd37faed9a59eb2788a56dac8deeee0a508)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov f881900342 x86/tdx: Make TDX_HYPERCALL asm similar to TDX_MODULE_CALL
JIRA: https://issues.redhat.com/browse/RHEL-19178

Now the 'struct tdx_hypercall_args' and 'struct tdx_module_args' are
almost the same, and the TDX_HYPERCALL and TDX_MODULE_CALL asm macro
share similar code pattern too.  The __tdx_hypercall() and __tdcall()
should be unified to use the same assembly code.

As a preparation to unify them, simplify the TDX_HYPERCALL to make it
more like the TDX_MODULE_CALL.

The TDX_HYPERCALL takes the pointer of 'struct tdx_hypercall_args' as
function call argument, and does below extra things comparing to the
TDX_MODULE_CALL:

1) It sets RAX to 0 (TDG.VP.VMCALL leaf) internally;
2) It sets RCX to the (fixed) bitmap of shared registers internally;
3) It calls __tdx_hypercall_failed() internally (and panics) when the
   TDCALL instruction itself fails;
4) After TDCALL, it moves R10 to RAX to return the return code of the
   VMCALL leaf, regardless the '\ret' asm macro argument;

Firstly, change the TDX_HYPERCALL to take the same function call
arguments as the TDX_MODULE_CALL does: TDCALL leaf ID, and the pointer
to 'struct tdx_module_args'.  Then 1) and 2) can be moved to the
caller:

 - TDG.VP.VMCALL leaf ID can be passed via the function call argument;
 - 'struct tdx_module_args' is 'struct tdx_hypercall_args' + RCX, thus
   the bitmap of shared registers can be passed via RCX in the
   structure.

Secondly, to move 3) and 4) out of assembly, make the TDX_HYPERCALL
always save output registers to the structure.  The caller then can:

 - Call __tdx_hypercall_failed() when TDX_HYPERCALL returns error;
 - Return R10 in the structure as the return code of the VMCALL leaf;

With above changes, change the asm function from __tdx_hypercall() to
__tdcall_hypercall(), and reimplement __tdx_hypercall() as the C wrapper
of it.  This avoids having to add another wrapper of __tdx_hypercall()
(_tdx_hypercall() is already taken).

The __tdcall_hypercall() will be replaced with a __tdcall() variant
using TDX_MODULE_CALL in a later commit as the final goal is to have one
assembly to handle both TDCALL and TDVMCALL.

Currently, the __tdx_hypercall() asm is in '.noinstr.text'.  To keep
this unchanged, annotate __tdx_hypercall(), which is a C function now,
as 'noinstr'.

Remove the __tdx_hypercall_ret() as __tdx_hypercall() already does so.

Implement __tdx_hypercall() in tdx-shared.c so it can be shared with the
compressed code.

Opportunistically fix a checkpatch error complaining using space around
parenthesis '(' and ')' while moving the bitmap of shared registers to
<asm/shared/tdx.h>.

[ dhansen: quash new calls of __tdx_hypercall_ret() that showed up ]

Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://lore.kernel.org/all/0cbf25e7aee3256288045023a31f65f0cef90af4.1692096753.git.kai.huang%40intel.com
(cherry picked from commit c641cfb5c157b6c3062a824fd8ba190bf06fb952)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov 2232b16a42 x86/tdx: Extend TDX_MODULE_CALL to support more TDCALL/SEAMCALL leafs
JIRA: https://issues.redhat.com/browse/RHEL-19178

The TDX guest live migration support (TDX 1.5) adds new TDCALL/SEAMCALL
leaf functions.  Those new TDCALLs/SEAMCALLs take additional registers
for input (R10-R13) and output (R12-R13).  TDG.SERVTD.RD is an example.

Also, the current TDX_MODULE_CALL doesn't aim to handle TDH.VP.ENTER
SEAMCALL, which monitors the TDG.VP.VMCALL in input/output registers
when it returns in case of VMCALL from TDX guest.

With those new TDCALLs/SEAMCALLs and the TDH.VP.ENTER covered, the
TDX_MODULE_CALL macro basically needs to handle the same input/output
registers as the TDX_HYPERCALL does.  And as a result, they also share
similar logic in the assembly, thus should be unified to use one common
assembly.

Extend the TDX_MODULE_CALL asm to support the new TDCALLs/SEAMCALLs and
also the TDH.VP.ENTER SEAMCALL.  Eventually it will be unified with the
TDX_HYPERCALL.

The new input/output registers fit with the "callee-saved" registers in
the x86 calling convention.  Add a new "saved" parameter to support
those new TDCALLs/SEAMCALLs and TDH.VP.ENTER and keep the existing
TDCALLs/SEAMCALLs minimally impacted.

For TDH.VP.ENTER, after it returns the registers shared by the guest
contain guest's values.  Explicitly clear them to prevent speculative
use of guest's values.

Note most TDX live migration related SEAMCALLs may also clobber AVX*
state ("AVX, AVX2 and AVX512 state: may be reset to the architectural
INIT state" -- see TDH.EXPORT.MEM for example).  And TDH.VP.ENTER also
clobbers XMM0-XMM15 when the corresponding bit is set in RCX.  Don't
handle them in the TDX_MODULE_CALL macro but let the caller save and
restore when needed.

This is basically based on Peter's code.

Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://lore.kernel.org/all/d4785de7c392f7c5684407f6c24a73b92148ec49.1692096753.git.kai.huang%40intel.com
(cherry picked from commit 12f34ed8622aafd3bbd9d8aa4550dcb7016ea1e6)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov 0eec180e6f x86/tdx: Pass TDCALL/SEAMCALL input/output registers via a structure
JIRA: https://issues.redhat.com/browse/RHEL-19178

Currently, the TDX_MODULE_CALL asm macro, which handles both TDCALL and
SEAMCALL, takes one parameter for each input register and an optional
'struct tdx_module_output' (a collection of output registers) as output.
This is different from the TDX_HYPERCALL macro which uses a single
'struct tdx_hypercall_args' to carry all input/output registers.

The newer TDX versions introduce more TDCALLs/SEAMCALLs which use more
input/output registers.  Also, the TDH.VP.ENTER (which isn't covered
by the current TDX_MODULE_CALL macro) basically can use all registers
that the TDX_HYPERCALL does.  The current TDX_MODULE_CALL macro isn't
extendible to cover those cases.

Similar to the TDX_HYPERCALL macro, simplify the TDX_MODULE_CALL macro
to use a single structure 'struct tdx_module_args' to carry all the
input/output registers.  Currently, R10/R11 are only used as output
register but not as input by any TDCALL/SEAMCALL.  Change to also use
R10/R11 as input register to make input/output registers symmetric.

Currently, the TDX_MODULE_CALL macro depends on the caller to pass a
non-NULL 'struct tdx_module_output' to get additional output registers.
Similar to the TDX_HYPERCALL macro, change the TDX_MODULE_CALL macro to
take a new 'ret' macro argument to indicate whether to save the output
registers to the 'struct tdx_module_args'.  Also introduce a new
__tdcall_ret() for that purpose, similar to the __tdx_hypercall_ret().

Note the tdcall(), which is a wrapper of __tdcall(), is called by three
callers: tdx_parse_tdinfo(), tdx_get_ve_info() and tdx_early_init().
The former two need the additional output but the last one doesn't.  For
simplicity, make tdcall() always call __tdcall_ret() to avoid another
"_ret()" wrapper.  The last caller tdx_early_init() isn't performance
critical anyway.

Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://lore.kernel.org/all/483616c1762d85eb3a3c3035a7de061cfacf2f14.1692096753.git.kai.huang%40intel.com
(cherry picked from commit 57a420bb8186d1d0178b857e5dd5026093641654)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov c5324e847b x86/tdx: Rename __tdx_module_call() to __tdcall()
JIRA: https://issues.redhat.com/browse/RHEL-19178

__tdx_module_call() is only used by the TDX guest to issue TDCALL to the
TDX module.  Rename it to __tdcall() to match its behaviour, e.g., it
cannot be used to make host-side SEAMCALL.

Also rename tdx_module_call() which is a wrapper of __tdx_module_call()
to tdcall().

No functional change intended.

Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://lore.kernel.org/all/785d20d99fbcd0db8262c94da6423375422d8c75.1692096753.git.kai.huang%40intel.com
(cherry picked from commit 5efb96289e581c187af1bc288ce5d26ed6181749)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov 2bf1b79bc9 x86/tdx: Make macros of TDCALLs consistent with the spec
JIRA: https://issues.redhat.com/browse/RHEL-19178

The TDX spec names all TDCALLs with prefix "TDG".  Currently, the kernel
doesn't follow such convention for the macros of those TDCALLs but uses
prefix "TDX_" for all of them.  Although it's arguable whether the TDX
spec names those TDCALLs properly, it's better for the kernel to follow
the spec when naming those macros.

Change all macros of TDCALLs to make them consistent with the spec.  As
a bonus, they get distinguished easily from the host-side SEAMCALLs,
which all have prefix "TDH".

No functional change intended.

Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://lore.kernel.org/all/516dccd0bd8fb9a0b6af30d25bb2d971aa03d598.1692096753.git.kai.huang%40intel.com
(cherry picked from commit f0024dbfc48d8814d915eb5bd5253496b9b8a6df)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov b3fa507057 x86/tdx: Skip saving output regs when SEAMCALL fails with VMFailInvalid
JIRA: https://issues.redhat.com/browse/RHEL-19178

If SEAMCALL fails with VMFailInvalid, the SEAM software (e.g., the TDX
module) won't have chance to set any output register.  Skip saving the
output registers to the structure in this case.

Also, as '.Lno_output_struct' is the very last symbol before RET, rename
it to '.Lout' to make it short.

Opportunistically make the asm directives unindented.

Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://lore.kernel.org/all/704088f5b4d72c7e24084f7f15bd1ac5005b7213.1692096753.git.kai.huang%40intel.com
(cherry picked from commit 03a423d40cb30e0e1cb77a801acb56ddb0bf6f5e)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:05 +01:00
Vitaly Kuznetsov ad63ff1e92 x86/tdx: Zero out the missing RSI in TDX_HYPERCALL macro
JIRA: https://issues.redhat.com/browse/RHEL-19178

In the TDX_HYPERCALL asm, after the TDCALL instruction returns from the
untrusted VMM, the registers that the TDX guest shares to the VMM need
to be cleared to avoid speculative execution of VMM-provided values.

RSI is specified in the bitmap of those registers, but it is missing
when zeroing out those registers in the current TDX_HYPERCALL.

It was there when it was originally added in commit 752d13305c78
("x86/tdx: Expand __tdx_hypercall() to handle more arguments"), but was
later removed in commit 1e70c680375a ("x86/tdx: Do not corrupt
frame-pointer in __tdx_hypercall()"), which was correct because %rsi is
later restored in the "pop %rsi".  However a later commit 7a3a401874be
("x86/tdx: Drop flags from __tdx_hypercall()") removed that "pop %rsi"
but forgot to add the "xor %rsi, %rsi" back.

Fix by adding it back.

Fixes: 7a3a401874be ("x86/tdx: Drop flags from __tdx_hypercall()")
Signed-off-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>
Reviewed-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://lore.kernel.org/all/e7d1157074a0b45d34564d5f17f3e0ffee8115e9.1692096753.git.kai.huang%40intel.com
(cherry picked from commit 5d092b66119d774853cc9308522620299048a662)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:04 +01:00
Vitaly Kuznetsov 6e5b721119 x86/tdx: Retry partially-completed page conversion hypercalls
JIRA: https://issues.redhat.com/browse/RHEL-19178

TDX guest memory is private by default and the VMM may not access it.
However, in cases where the guest needs to share data with the VMM,
the guest and the VMM can coordinate to make memory shared between
them.

The guest side of this protocol includes the "MapGPA" hypercall.  This
call takes a guest physical address range.  The hypercall spec (aka.
the GHCI) says that the MapGPA call is allowed to return partial
progress in mapping this range and indicate that fact with a special
error code.  A guest that sees such partial progress is expected to
retry the operation for the portion of the address range that was not
completed.

Hyper-V does this partial completion dance when set_memory_decrypted()
is called to "decrypt" swiotlb bounce buffers that can be up to 1GB
in size.  It is evidently the only VMM that does this, which is why
nobody noticed this until now.

[ dhansen: rewrite changelog ]

Signed-off-by: Dexuan Cui <decui@microsoft.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Michael Kelley <mikelley@microsoft.com>
Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>
Acked-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Link: https://lore.kernel.org/all/20230811021246.821-2-decui%40microsoft.com
(cherry picked from commit 019b383d1132e4051de0d2e43254454b86538cf4)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:04 +01:00
Vitaly Kuznetsov 16696e9762 x86/kvm: Do not try to disable kvmclock if it was not enabled
JIRA: https://issues.redhat.com/browse/RHEL-19178
Upstream Status: git://git.kernel.org/pub/scm/virt/kvm/kvm.git

kvm_guest_cpu_offline() tries to disable kvmclock regardless if it is
present in the VM. It leads to write to a MSR that doesn't exist on some
configurations, namely in TDX guest:

	unchecked MSR access error: WRMSR to 0x12 (tried to write 0x0000000000000000)
	at rIP: 0xffffffff8110687c (kvmclock_disable+0x1c/0x30)

kvmclock enabling is gated by CLOCKSOURCE and CLOCKSOURCE2 KVM paravirt
features.

Do not disable kvmclock if it was not enabled.

Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Fixes: c02027b574 ("x86/kvm: Disable kvmclock on all CPUs on shutdown")
Reviewed-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Vitaly Kuznetsov <vkuznets@redhat.com>
Cc: Paolo Bonzini <pbonzini@redhat.com>
Cc: Wanpeng Li <wanpengli@tencent.com>
Cc: stable@vger.kernel.org
Message-Id: <20231205004510.27164-6-kirill.shutemov@linux.intel.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
(cherry picked from commit 1c6d984f523f67ecfad1083bb04c55d91977bb15)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:04 +01:00
Vitaly Kuznetsov e6bb50cd4c x86/tdx: Mark TSC reliable
JIRA: https://issues.redhat.com/browse/RHEL-19178

In x86 virtualization environments, including TDX, RDTSC instruction is
handled without causing a VM exit, resulting in minimal overhead and
jitters. On the other hand, other clock sources (such as HPET, ACPI
timer, APIC, etc.) necessitate VM exits to implement, resulting in more
fluctuating measurements compared to TSC. Thus, those clock sources are
not effective for calibrating TSC.

As a foundation, the host TSC is guaranteed to be invariant on any
system which enumerates TDX support.

TDX guests and the TDX module build on that foundation by enforcing:

  - Virtual TSC is monotonously incrementing for any single VCPU;
  - Virtual TSC values are consistent among all the TD’s VCPUs at the
    level supported by the CPU:
    + VMM is required to set the same TSC_ADJUST;
    + VMM must not modify from initial value of TSC_ADJUST before
      SEAMCALL;
  - The frequency is determined by TD configuration:
    + Virtual TSC frequency is specified by VMM on TDH.MNG.INIT;
    + Virtual TSC starts counting from 0 at TDH.MNG.INIT;

The result is that a reliable TSC is a TDX architectural guarantee.

Use the TSC as the only reliable clock source in TD guests, bypassing
unstable calibration.

This is similar to what the kernel already does in some VMWare and
HyperV environments.

[ dhansen: changelog tweaks ]

Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>
Reviewed-by: Erdem Aktas <erdemaktas@google.com>
Reviewed-by: Isaku Yamahata <isaku.yamahata@intel.com>
Acked-by: Kai Huang <kai.huang@intel.com>
Link: https://lore.kernel.org/all/20231006144549.2633-1-kirill.shutemov%40linux.intel.com
(cherry picked from commit 9ee4318c157b9802589b746cc340bae3142d984c)

Conflicts:
	arch/x86/coco/tdx/tdx.c (context, skipping da86eb9611840)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:04 +01:00
Vitaly Kuznetsov a5242b6a10 x86/tdx: Allow 32-bit emulation by default
JIRA: https://issues.redhat.com/browse/RHEL-19178

32-bit emulation was disabled on TDX to prevent a possible attack by
a VMM injecting an interrupt on vector 0x80.

Now that int80_emulation() has a check for external interrupts the
limitation can be lifted.

To distinguish software interrupts from external ones, int80_emulation()
checks the APIC ISR bit relevant to the 0x80 vector. For
software interrupts, this bit will be 0.

On TDX, the VAPIC state (including ISR) is protected and cannot be
manipulated by the VMM. The ISR bit is set by the microcode flow during
the handling of posted interrupts.

[ dhansen: more changelog tweaks ]

Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Thomas Gleixner <tglx@linutronix.de>
Reviewed-by: Borislav Petkov (AMD) <bp@alien8.de>
Cc: <stable@vger.kernel.org> # v6.0+
(cherry picked from commit f4116bfc44621882556bbf70f5284fbf429a5cf6)

Conflicts:
	arch/x86/coco/tdx/tdx.c (context, skipping ff3cfcb0d46a)

Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:04 +01:00
Vitaly Kuznetsov afe93805bf x86/entry: Do not allow external 0x80 interrupts
JIRA: https://issues.redhat.com/browse/RHEL-19178

The INT 0x80 instruction is used for 32-bit x86 Linux syscalls. The
kernel expects to receive a software interrupt as a result of the INT
0x80 instruction. However, an external interrupt on the same vector
also triggers the same codepath.

An external interrupt on vector 0x80 will currently be interpreted as a
32-bit system call, and assuming that it was a user context.

Panic on external interrupts on the vector.

To distinguish software interrupts from external ones, the kernel checks
the APIC ISR bit relevant to the 0x80 vector. For software interrupts,
this bit will be 0.

Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Borislav Petkov (AMD) <bp@alien8.de>
Cc: <stable@vger.kernel.org> # v6.0+
(cherry picked from commit 55617fb991df535f953589586468612351575704)
Signed-off-by: Vitaly Kuznetsov <vkuznets@redhat.com>
2024-01-15 16:35:04 +01:00