821 Commits
Author SHA1 Message Date
Waiman Long 0aee40aa4c posix-cpu-timers: Prevent UAF caused by non-leader exec() race
JIRA: https://redhat.atlassian.net/browse/RHEL-227851
CVE: CVE-2026-64560
Conflicts:
  A context diff with the kernel/exit.c hunk due to missing upstream
  commit 7903f907a226 ("pid: perform free_pid() calls outside of
  tasklist_lock").

commit 920f893f735e92ba3a1cd9256899a186b161928d
Author: Thomas Gleixner <tglx@kernel.org>
Date:   Fri, 3 Jul 2026 12:02:38 +0200

    posix-cpu-timers: Prevent UAF caused by non-leader exec() race

    Wongi and Jungwoo decoded and reported a non-leader exec() related race
    which can result in an UAF:

     sys_timer_delete()                     exec()
       posix_cpu_timer_del()
       // Observes old leader
       p = pid_task(pid, pid_type);         de_thread()
                                              switch_leader();
                                              release_task(old_leader)
                                                __exit_signal(old_leader)
                                                  sighand = lock(old_leader, sighand);
                                                  posix_cpu_timers*_exit();
       sighand = lock_task_sighand(p)             unhash_task(old_leader);
         sh = lock(p, sighand)                    old_leader->sighand = NULL;
                                                  unlock(sighand);
         (p->sighand == NULL)
            unlock(sh)
            return NULL;

       // Returns without action
       if(!sighand)
          return 0;
       free_posix_timer();

    This is "harmless" unless the deleted timer was armed and enqueued in
    p->signal because on exec() a TGID targeted timer is inherited.

    As sys_timer_delete() freed the underlying posix timer object
    run_posix_cpu_timers() or any timerqueue related add/delete operations on
    other timers will access the freed object's timerqueue node, which results
    in an UAF.

    There is a similar problem vs. posix_cpu_timer_set(). For regular posix
    timers it just transiently returns -ESRCH to user space, but for the use
    case in do_cpu_nanosleep() it's the same UAF just that the k_itimer is
    allocated on the stack.

    Also posix_cpu_timer_rearm() fails to rearm the timer, which means it stops
    to expire.

    While debating solutions Frederic pointed out another problem:

       posix_cpu_timer_del(tmr)
                                            __exit_signal(p)
                                              posix_cpu_timers*_exit(p);
                                              unhash_task(p);
                                              p->sighand = NULL;
         sh = lock_task_sighand(p)
            sighand = p->sighand;
            if (!sighand)
                return NULL;
            lock(sighand);

         if (!sh)
            WARN_ON_ONCE(timer_queued(tmr));

    On weakly ordered architectures it is not guaranteed that
    posix_cpu_timer_del() will observe the stores in posix_cpu_timers*_exit()
    when p->sighand is observed as NULL, which means the WARN() can be a false
    positive.

    Solve these issues by:

      1) Changing the store in __exit_signal() to smp_store_release().

      2) Adding a smp_acquire__after_ctrl_dep() into the !sighand path
         of lock_task_sighand().

      3) Creating a helper function for looking up the task and locking sighand
         which does not return when sighand == NULL. Instead it retries the
         task lookup and only if that fails it gives up.

      4) Using that helper in the three affected functions.

    #1/#2 ensures that the reader side which observes sighand == NULL also
    observes all preceeding stores, i.e. the stores in posix_cpu_timers*_exit()
    and the ones in unhash_task().

    #3 ensures that the above described non-leader exec() situation is handled
    gracefully. When the task lookup returns the old leader, but sighand ==
    NULL then it retries. In the non-leader exec() case the subsequent task
    lookup will observe the new leader due to #1/#2. In normal exit() scenarios
    the subsequent lookup fails.

    When the task lookup fails, the function also checks whether the timer is
    still enqueued and issues a warning if that's the case. Unfortunately there
    is nothing which can be done about it, but as the task is already not
    longer visible the timer should not be accessed anymore. This check also
    requires memory ordering, which is not provided when the first lookup
    fails. To achieve that the check is preceeded by a smp_rmb() which pairs
    with the smp_wmb() in write_seqlock() in __exit_signal(). That ensures that
    the stores in posix_cpu_timers*_exit() are visible.

    The history of the non-leader exec() issue goes back to the early days of
    posix CPU timers, which stored a pointer to the group leader task in the
    timer. That obviously fails when a non-leader exec() switches the leader.
    commit e0a7021710 ("posix-cpu-timers: workaround to suppress the problems
    with mt exec") added a temporary workaround for that in 2010 which survived
    about 10 years. The fix for the workaround changed the task pointer to a
    pid pointer, but failed to see the subtle race described above. So the
    Fixes tag picks that commit, which seems to be halfways accurate.

    Thanks to Frederic Weissbecker, Oleg Nesterov and Peter Zijlstra for
    review, feedback and suggestions and to Wongi and Jungwoo for the excellent
    bug report and analysis!

    Fixes: 55e8c8eb2c ("posix-cpu-timers: Store a reference to a pid not a task")
    Reported-by: Wongi Lee <qw3rtyp0@gmail.com>
    Reported-by: Jungwoo Lee <jwlee2217@gmail.com>
    Signed-off-by: Thomas Gleixner <tglx@kernel.org>
    Reviewed-by: Oleg Nesterov <oleg@redhat.com>
    Cc: stable@vger.kernel.org

Signed-off-by: Waiman Long <longman@redhat.com>
2026-08-08 14:55:00 -04:00
Waiman Long 47a6d9810b treewide: const qualify ctl_tables where applicable
JIRA: https://issues.redhat.com/browse/RHEL-154147
Conflicts:
  The following hunks that are not applicable are dropped.
   - arch/riscv/kernel/process.c
   - arch/s390/mm/pgalloc.c
   - drivers/gpu/drm/i915/i915_perf.c
   - drivers/gpu/drm/xe/xe_observation.c
   - drivers/tty/tty_io.c
   - fs/fuse/sysctl.c
   - kernel/pid.c

commit 1751f872cc97f992ed5c4c72c55588db1f0021e1
Author: Joel Granados <joel.granados@kernel.org>
Date:   Tue, 28 Jan 2025 13:48:37 +0100

    treewide: const qualify ctl_tables where applicable

    Add the const qualifier to all the ctl_tables in the tree except for
    watchdog_hardlockup_sysctl, memory_allocation_profiling_sysctls,
    loadpin_sysctl_table and the ones calling register_net_sysctl (./net,
    drivers/inifiniband dirs). These are special cases as they use a
    registration function with a non-const qualified ctl_table argument or
    modify the arrays before passing them on to the registration function.

    Constifying ctl_table structs will prevent the modification of
    proc_handler function pointers as the arrays would reside in .rodata.
    This is made possible after commit 78eb4ea25c ("sysctl: treewide:
    constify the ctl_table argument of proc_handlers") constified all the
    proc_handlers.

    Created this by running an spatch followed by a sed command:
    Spatch:
        virtual patch

        @
        depends on !(file in "net")
        disable optional_qualifier
        @

        identifier table_name != {
          watchdog_hardlockup_sysctl,
          iwcm_ctl_table,
          ucma_ctl_table,
          memory_allocation_profiling_sysctls,
          loadpin_sysctl_table
        };
        @@

        + const
        struct ctl_table table_name [] = { ... };

    sed:
        sed --in-place \
          -e "s/struct ctl_table .table = &uts_kern/const struct ctl_table *table = \&uts_kern/" \
          kernel/utsname_sysctl.c

    Reviewed-by: Song Liu <song@kernel.org>
    Acked-by: Steven Rostedt (Google) <rostedt@goodmis.org> # for kernel/trace/
    Reviewed-by: Martin K. Petersen <martin.petersen@oracle.com> # SCSI
    Reviewed-by: Darrick J. Wong <djwong@kernel.org> # xfs
    Acked-by: Jani Nikula <jani.nikula@intel.com>
    Acked-by: Corey Minyard <cminyard@mvista.com>
    Acked-by: Wei Liu <wei.liu@kernel.org>
    Acked-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Bill O'Donnell <bodonnel@redhat.com>
    Acked-by: Baoquan He <bhe@redhat.com>
    Acked-by: Ashutosh Dixit <ashutosh.dixit@intel.com>
    Acked-by: Anna Schumaker <anna.schumaker@oracle.com>
    Signed-off-by: Joel Granados <joel.granados@kernel.org>

Signed-off-by: Waiman Long <longman@redhat.com>
2026-03-08 12:39:41 -04:00
Waiman Long aad80a4afa posix-timers: Rework timer removal
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 1d25bdd3f3831bb1b9512d4b5afcd2dea8a0c515
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Mon, 10 Mar 2025 09:13:32 +0100

    posix-timers: Rework timer removal

    sys_timer_delete() and the do_exit() cleanup function itimer_delete() are
    doing the same thing, but have needlessly different implementations instead
    of sharing the code.

    The other oddity of timer deletion is the fact that the timer is not
    invalidated before the actual deletion happens, which allows concurrent
    lookups to succeed.

    That's wrong because a timer which is in the process of being deleted
    should not be visible and any actions like signal queueing, delivery and
    rearming should not happen once the task, which invoked timer_delete(), has
    the timer locked.

    Rework the code so that:

       1) The signal queueing and delivery code ignore timers which are marked
          invalid

       2) The deletion implementation between sys_timer_delete() and
          itimer_delete() is shared

       3) The timer is invalidated and removed from the linked lists before
          the deletion callback of the relevant clock is invoked.

          That requires to rework timer_wait_running() as it does a lookup of
          the timer when relocking it at the end. In case of deletion this
          lookup would fail due to the preceding invalidation and the wait loop
          would terminate prematurely.

          But due to the preceding invalidation the timer cannot be accessed by
          other tasks anymore, so there is no way that the timer has been freed
          after the timer lock has been dropped.

          Move the re-validation out of timer_wait_running() and handle it at
          the only other usage site, timer_settime().

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Link: https://lore.kernel.org/all/87zfht1exf.ffs@tglx

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:31 -05:00
Waiman Long 836668654b signal/posixtimers: Handle ignore/blocked sequences correctly
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 8c4840277b6daffe09dea0338f3fce1eb4319a43
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 14 Jan 2025 18:28:44 +0100

    signal/posixtimers: Handle ignore/blocked sequences correctly

    syzbot triggered the warning in posixtimer_send_sigqueue(), which warns
    about a non-ignored signal being already queued on the ignored list.

    The warning is actually bogus, as the following sequence causes this:

        signal($SIG, SIGIGN);
        timer_settime(...);                 // arm periodic timer

          timer fires, signal is ignored and queued on ignored list

        sigprocmask(SIG_BLOCK, ...);        // block the signal
        timer_settime(...);                 // re-arm periodic timer

          timer fires, signal is not ignored because it is blocked
            ---> Warning triggers as signal is on the ignored list

    Ideally timer_settime() could remove the signal, but that's racy and
    incomplete vs. other scenarios and requires a full reevaluation of the
    pending signal list.

    Instead of adding more complexity, handle it gracefully by removing the
    warning and requeueing the signal to the pending list. That's correct
    versus:

      1) sig[timed]wait() as that does not check for SIGIGN and only relies on
         dequeue_signal() -> posixtimers_deliver_signal() to check whether the
         pending signal is still valid.

      2) Unblocking of the signal.

         - If the unblocking happens before SIGIGN is replaced by a signal
           handler, then the timer is rearmed in dequeue_signal(), but
           get_signal() will ignore it. The next timer expiry will move it back
           to the ignored list.

         - If SIGIGN was replaced before unblocking, then the signal will be
           delivered and a subsequent expiry will queue a signal on the pending
           list again.

    There is a related scenario to trigger the complementary warning in the
    signal ignored path, which does not expect the signal to be on the pending
    list when it is ignored. That can be triggered even before the above change
    via:

    task1                   task2

    signal($SIG, SIGIGN);
                            sigprocmask(SIG_BLOCK, ...);

    timer_create();         // Signal target is task2
    timer_settime(...);     // arm periodic timer

       timer fires, signal is not ignored because it is blocked
       and queued on the pending list of task2

                            syscall()
                               // Sets the pending flag
                               sigprocmask(SIG_UNBLOCK, ...);

                            -> preemption, task2 cannot dequeue the signal

    timer_settime(...);     // re-arm periodic timer

       timer fires, signal is ignored
            ---> Warning triggers as signal is on task2's pending list
                 and the thread group is not exiting

    Consequently, remove that warning too and just keep the signal on the
    pending list.

    The following attempt to deliver the signal on return to user space of
    task2 will ignore the signal and a subsequent expiry will bring it back to
    the ignored list, if it did not get blocked or un-ignored before that.

    Fixes: df7a996b4dab ("signal: Queue ignored posixtimers on ignore list")
    Reported-by: syzbot+3c2e3cc60665d71de2f7@syzkaller.appspotmail.com
    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Link: https://lore.kernel.org/all/87ikqhcnjn.ffs@tglx

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:28 -05:00
Waiman Long 0d76a70432 posix-timers: Fix spurious warning on double enqueue versus do_exit()
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit cdc905d16b07981363e53a21853ba1cf6cd8e92a
Author: Frederic Weisbecker <frederic@kernel.org>
Date:   Sun, 17 Nov 2024 00:48:23 +0100

    posix-timers: Fix spurious warning on double enqueue versus do_exit()

    A timer sigqueue may find itself already pending when it is tried to
    be enqueued. This situation can happen if the timer sigqueue is enqueued
    but then the timer is reset afterwards and fires before the pending
    signal managed to be delivered.

    However when such a double enqueue occurs while the corresponding signal
    is ignored, the sigqueue is expected to be found either on the dedicated
    ignored list if the timer was periodic or dropped if the timer was
    one-shot. In any case it is not supposed to be queued on the real signal
    queue.

    An assertion verifies the latter expectation on top of the return value
    of prepare_signal(), assuming "false" means that the signal is being
    ignored. But prepare_signal() may also fail if the target is exiting as
    the last task of its group. In this case the double enqueue observes the
    sigqueue queued, as in such a situation:

        TASK A (same group as B)                   TASK B (same group as A)
        ------------------------                   ------------------------

        // timer event
        // queue signal to TASK B
        posix_timer_queue_signal()
        // reset timer through syscall
        do_timer_settime()
        // exit, leaving task B alone
        do_exit()
                                                   do_exit()
                                                      synchronize_group_exit()
                                                          signal->flags = SIGNAL_GROUP_EXIT
                                                      // ========> <IRQ> timer event
                                                      posix_timer_queue_signal()
                                                      // return false due to SIGNAL_GROUP_EXIT
                                                      if (!prepare_signal())
                                                         WARN_ON_ONCE(!list_empty(&q->list))

    And this spuriously triggers this warning:

        WARNING: CPU: 0 PID: 5854 at kernel/signal.c:2008 posixtimer_send_sigqueue
        CPU: 0 UID: 0 PID: 5854 Comm: syz-executor139 Not tainted 6.12.0-rc6-next-20241108-syzkaller #0
        RIP: 0010:posixtimer_send_sigqueue+0x9da/0xbc0 kernel/signal.c:2008
        Call Trace:
         <IRQ>
         alarm_handle_timer
         alarmtimer_fired
         __run_hrtimer
         __hrtimer_run_queues
         hrtimer_interrupt
         local_apic_timer_interrupt
         __sysvec_apic_timer_interrupt
         instr_sysvec_apic_timer_interrupt
         sysvec_apic_timer_interrupt
         </IRQ>

    Fortunately the recovery code in that case already does the right thing:
    just exit from posixtimer_send_sigqueue() and wait for __exit_signal()
    to flush the pending signal. Just make sure to warn only the case when
    the sigqueue is queued and the signal is really ignored.

    Fixes: df7a996b4dab ("signal: Queue ignored posixtimers on ignore list")
    Reported-by: syzbot+852e935b899bde73626e@syzkaller.appspotmail.com
    Signed-off-by: Frederic Weisbecker <frederic@kernel.org>
    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Tested-by: syzbot+852e935b899bde73626e@syzkaller.appspotmail.com
    Link: https://lore.kernel.org/all/20241116234823.28497-1-frederic@kernel.org
    Closes: https://lore.kernel.org/all/673549c6.050a0220.1324f8.008c.GAE@google.com

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:27 -05:00
Waiman Long 5f45225ffb posix-timers: Cleanup SIG_IGN workaround leftovers
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 7a66f72b09bb0762360274b1fb677b3433dbaa06
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:55 +0100

    posix-timers: Cleanup SIG_IGN workaround leftovers

    Now that ignored posix timer signals are requeued and the timers are
    rearmed on signal delivery the workaround to keep such timers alive and
    self rearm them is not longer required.

    Remove the relevant hacks and the not longer required return values from
    the related functions. The alarm timer workarounds will be cleaned up in a
    separate step.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241105064214.187239060@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:27 -05:00
Waiman Long 190ca3c164 signal: Queue ignored posixtimers on ignore list
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit df7a996b4dab03c889fa86d849447b716f07b069
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:54 +0100

    signal: Queue ignored posixtimers on ignore list

    Queue posixtimers which have their signal ignored on the ignored list:

       1) When the timer fires and the signal has SIG_IGN set

       2) When SIG_IGN is installed via sigaction() and a timer signal
          is already queued

    This only happens when the signal is for a valid timer, which delivered the
    signal in periodic mode. One-shot timer signals are correctly dropped.

    Due to the lock order constraints (sighand::siglock nests inside
    timer::lock) the signal code cannot access any of the timer fields which
    are relevant to make this decision, e.g. timer::it_status.

    This is addressed by establishing a protection scheme which requires to
    lock both locks on the timer side for modifying decision fields in the
    timer struct and therefore makes it possible for the signal delivery to
    evaluate with only sighand:siglock being held:

      1) Move the NULLification of timer->it_signal into the sighand::siglock
         protected section of timer_delete() and check timer::it_signal in the
         code path which determines whether the signal is dropped or queued on
         the ignore list.

         This ensures that a deleted timer cannot be moved onto the ignore
         list, which would prevent it from being freed on exit() as it is not
         longer in the process' posix timer list.

         If the timer got moved to the ignored list before deletion then it is
         removed from the ignored list under sighand lock in timer_delete().

      2) Provide a new timer::it_sig_periodic flag, which gets set in the
         signal queue path with both timer and sighand locks held if the timer
         is actually in periodic mode at expiry time.

         The ignore list code checks this flag under sighand::siglock and drops
         the signal when it is not set.

         If it is set, then the signal is moved to the ignored list independent
         of the actual state of the timer.

         When the signal is un-ignored later then the signal is moved back to
         the signal queue. On signal delivery the posix timer side decides
         about dropping the signal if the timer was re-armed, dis-armed or
         deleted based on the signal sequence counter check.

         If the thread/process exits then not yet delivered signals are
         discarded which means the reference of the timer containing the
         sigqueue is dropped and frees the timer.

         This is way cheaper than requiring all code paths to lock
         sighand::siglock of the target thread/process on any modification of
         timer::it_status or going all the way and removing pending signals
         from the signal queues on every rearm, disarm or delete operation.

    So the protection scheme here is that on the timer side both timer::lock
    and sighand::siglock have to be held for modifying

       timer::it_signal
       timer::it_sig_periodic

    which means that on the signal side holding sighand::siglock is enough to
    evaluate these fields.

    In posixtimer_deliver_signal() holding timer::lock is sufficient to do the
    sequence validation against timer::it_signal_seq because a concurrent
    expiry is waiting on timer::lock to be released.

    This completes the SIG_IGN handling and such timers are not longer self
    rearmed which avoids pointless wakeups.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241105064214.120756416@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:27 -05:00
Waiman Long 066efb69c6 signal: Handle ignored signals in do_sigaction(action != SIG_IGN)
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit caf77435dd8a52cb39c602bdf67d35d6f782f553
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:52 +0100

    signal: Handle ignored signals in do_sigaction(action != SIG_IGN)

    When a real handler (including SIG_DFL) is installed for a signal, which
    had previously SIG_IGN set, then the list of ignored posix timers has to be
    checked for timers which are affected by this change.

    Add a list walk function which checks for the matching signal number and if
    found requeues the timers signal, so the timer is rearmed on signal
    delivery.

    Rearming the timer right away is not possible because that requires to drop
    sighand lock.

    No functional change as the counter part which queues the timers on the
    ignored list is still missing.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241105064214.054091076@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:27 -05:00
Waiman Long 4e724c4a7b posix-timers: Move sequence logic into struct k_itimer
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 647da5f709f112319c0d51e06f330d8afecb1940
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:48 +0100

    posix-timers: Move sequence logic into struct k_itimer

    The posix timer signal handling uses siginfo::si_sys_private for handling
    the sequence counter check. That indirection is not longer required and the
    sequence count value at signal queueing time can be stored in struct
    k_itimer itself.

    This removes the requirement of treating siginfo::si_sys_private special as
    it's now always zero as the kernel does not touch it anymore.

    Suggested-by: Eric W. Biederman <ebiederm@xmission.com>
    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: "Eric W. Biederman" <ebiederm@xmission.com>
    Link: https://lore.kernel.org/all/20241105064213.852619866@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:26 -05:00
Waiman Long 8e9e504331 signal: Cleanup unused posix-timer leftovers
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit c2a4796a154bb952be1106911841aab2c8c17c4d
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:46 +0100

    signal: Cleanup unused posix-timer leftovers

    Remove the leftovers of sigqueue preallocation as it's not longer used.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241105064213.786506636@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:26 -05:00
Waiman Long 30ec7c2807 posix-timers: Embed sigqueue in struct k_itimer
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 6017a158beb13b412e55a451379798aae5876514
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:45 +0100

    posix-timers: Embed sigqueue in struct k_itimer

    To cure the SIG_IGN handling for posix interval timers, the preallocated
    sigqueue needs to be embedded into struct k_itimer to prevent life time
    races of all sorts.

    Now that the prerequisites are in place, embed the sigqueue into struct
    k_itimer and fixup the relevant usage sites.

    Aside of preparing for proper SIG_IGN handling, this spares an extra
    allocation.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241105064213.719695194@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:26 -05:00
Waiman Long 4e4a1b77c3 signal: Replace resched_timer logic
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 11629b9808e5900d675fd469d19932ea48060de3
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:43 +0100

    signal: Replace resched_timer logic

    In preparation for handling ignored posix timer signals correctly and
    embedding the sigqueue struct into struct k_itimer, hand down a pointer to
    the sigqueue struct into posix_timer_deliver_signal() instead of just
    having a boolean flag.

    No functional change.

    Suggested-by: Eric W. Biederman <ebiederm@xmission.com>
    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: "Eric W. Biederman" <ebiederm@xmission.com>
    Link: https://lore.kernel.org/all/20241105064213.652658158@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:26 -05:00
Waiman Long 379de2133c signal: Refactor send_sigqueue()
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 0360ed14d9826678a50fa2b873e522a24cd3c018
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:42 +0100

    signal: Refactor send_sigqueue()

    To handle posix timers which have their signal ignored via SIG_IGN properly
    it is required to requeue a ignored signal for delivery when SIG_IGN is
    lifted so the timer gets rearmed.

    Split the required code out of send_sigqueue() so it can be reused in
    context of sigaction().

    While at it rename send_sigqueue() to posixtimer_send_sigqueue() so its
    clear what this is about.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241105064213.586453412@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:26 -05:00
Waiman Long f62e866cf6 signal: Provide posixtimer_sigqueue_init()
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 54f1dd642fd088ba969206f09e7afffad7d9db2c
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:39 +0100

    signal: Provide posixtimer_sigqueue_init()

    To cure the SIG_IGN handling for posix interval timers, the preallocated
    sigqueue needs to be embedded into struct k_itimer to prevent life time
    races of all sorts.

    Provide a new function to initialize the embedded sigqueue to prepare for
    that.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241105064213.450427515@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:26 -05:00
Waiman Long a729c59f0e signal: Split up __sigqueue_alloc()
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 5cac427f7971b0619ebbfc131ef81fcf229c3c01
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:38 +0100

    signal: Split up __sigqueue_alloc()

    To cure the SIG_IGN handling for posix interval timers, the preallocated
    sigqueue needs to be embedded into struct k_itimer to prevent life time
    races of all sorts.

    Reorganize __sigqueue_alloc() so the ucounts retrieval and the
    initialization can be used independently.

    No functional change.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241105064213.371410037@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:26 -05:00
Waiman Long af33b6d81c posix-timers: Make signal overrun accounting sensible
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit b06b0345fff3678517acd0f1837d52477ba30944
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:32 +0100

    posix-timers: Make signal overrun accounting sensible

    The handling of the timer overrun in the signal code is inconsistent as it
    takes previous overruns into account. This is just wrong as after the
    reprogramming of a timer the overrun count starts over from a clean state,
    i.e. 0.

    Don't touch info::si_overrun in send_sigqueue() and only store the overrun
    value at signal delivery time, which is computed from the timer itself
    relative to the expiry time.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241105064213.106738193@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:25 -05:00
Waiman Long 82d377d51d posix-timers: Make signal delivery consistent
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 513793bc6ab331b947111e8efaf8fcef33fb83e5
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 5 Nov 2024 09:14:31 +0100

    posix-timers: Make signal delivery consistent

    Signals of timers which are reprogammed, disarmed or deleted can deliver
    signals related to the past. The POSIX spec is blury about this:

     - "The effect of disarming or resetting a timer with pending expiration
        notifications is unspecified."

     - "The disposition of pending signals for the deleted timer is
        unspecified."

    In both cases it is reasonable to expect that pending signals are
    discarded. Especially in the reprogramming case it does not make sense to
    account for previous overruns or to deliver a signal for a timer which has
    been disarmed. This makes the behaviour consistent and understandable.

    Remove the si_sys_private check from the signal delivery code and invoke
    posix_timer_deliver_signal() unconditionally for posix timer related
    signals.

    Change posix_timer_deliver_signal() so it controls the actual signal
    delivery via the return value. It now instructs the signal code to drop the
    signal when:

      1) The timer does not longer exist in the hash table

      2) The timer signal_seq value is not the same as the si_sys_private value
         which was set when the signal was queued.

    This is also a preparatory change to embed the sigqueue into the k_itimer
    structure, which in turn allows to remove the si_sys_private magic.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Link: https://lore.kernel.org/all/20241105064213.040348644@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:25 -05:00
Waiman Long 7976a356b7 signal: Allow POSIX timer signals to be dropped
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit c775ea28d4e23f5e58b6953645ef90c1b27a8e83
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 1 Oct 2024 10:42:04 +0200

    signal: Allow POSIX timer signals to be dropped

    In case that a timer was reprogrammed or deleted an already pending signal
    is obsolete. Right now such signals are kept around and eventually
    delivered. While POSIX is blury about this:

     - "The effect of disarming or resetting a timer with pending expiration
        notifications is unspecified."

     - "The disposition of pending signals for the deleted timer is
        unspecified."

    it is reasonable in both cases to expect that pending signals are discarded
    as they have no meaning anymore.

    Prepare the signal code to allow dropping posix timer signals.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241001083835.494416923@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:23 -05:00
Waiman Long c9e3c7cf82 posix-timers: Cure si_sys_private race
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 4febce44cfebcb490b196d5d10ae9f403ca4c956
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 1 Oct 2024 10:42:03 +0200

    posix-timers: Cure si_sys_private race

    The si_sys_private member of the siginfo which is embedded in the
    preallocated sigqueue is used by the posix timer code to decide whether a
    timer must be reprogrammed on signal delivery.

    The handling of this is racy as a long standing comment in that code
    documents. It is modified with the timer lock held, but without sighand
    lock being held. The actual signal delivery code checks for it under
    sighand lock without holding the timer lock.

    Hand the new value to send_sigqueue() as argument and store it with sighand
    lock held. This is an intermediate change to address this issue.

    The arguments to this function will be cleanup in subsequent changes.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241001083835.434338954@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:23 -05:00
Waiman Long f8f20f6b9f signal: Cleanup flush_sigqueue_mask()
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit a76e1bbe879cf39952ec4b43ed653b0905635f24
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 1 Oct 2024 10:42:02 +0200

    signal: Cleanup flush_sigqueue_mask()

    Mop up the stale return value comment and add a lockdep check instead of
    commenting on the locking requirement.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Link: https://lore.kernel.org/all/20241001083835.374933959@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:23 -05:00
Waiman Long 52cbc784cc signal: Confine POSIX_TIMERS properly
JIRA: https://issues.redhat.com/browse/RHEL-114122

commit 68f99be287a59d50a9ad231d523f7e578f8bd28a
Author: Thomas Gleixner <tglx@linutronix.de>
Date:   Tue, 1 Oct 2024 10:42:00 +0200

    signal: Confine POSIX_TIMERS properly

    Move the itimer rearming out of the signal code and consolidate all posix
    timer related functions in the signal code under one ifdef.

    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
    Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
    Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
    Link: https://lore.kernel.org/all/20241001083835.314100569@linutronix.de

Signed-off-by: Waiman Long <longman@redhat.com>
2025-12-13 20:56:23 -05:00
Roman Gushchin 9e05e5c7ee signal: restore the override_rlimit logic
Prior to commit d646969055 ("Reimplement RLIMIT_SIGPENDING on top of
ucounts") UCOUNT_RLIMIT_SIGPENDING rlimit was not enforced for a class of
signals.  However now it's enforced unconditionally, even if
override_rlimit is set.  This behavior change caused production issues.  

For example, if the limit is reached and a process receives a SIGSEGV
signal, sigqueue_alloc fails to allocate the necessary resources for the
signal delivery, preventing the signal from being delivered with siginfo. 
This prevents the process from correctly identifying the fault address and
handling the error.  From the user-space perspective, applications are
unaware that the limit has been reached and that the siginfo is
effectively 'corrupted'.  This can lead to unpredictable behavior and
crashes, as we observed with java applications.

Fix this by passing override_rlimit into inc_rlimit_get_ucounts() and skip
the comparison to max there if override_rlimit is set.  This effectively
restores the old behavior.

Link: https://lkml.kernel.org/r/20241104195419.3962584-1-roman.gushchin@linux.dev
Fixes: d646969055 ("Reimplement RLIMIT_SIGPENDING on top of ucounts")
Signed-off-by: Roman Gushchin <roman.gushchin@linux.dev>
Co-developed-by: Andrei Vagin <avagin@google.com>
Signed-off-by: Andrei Vagin <avagin@google.com>
Acked-by: Oleg Nesterov <oleg@redhat.com>
Acked-by: Alexey Gladkov <legion@kernel.org>
Cc: Kees Cook <kees@kernel.org>
Cc: "Eric W. Biederman" <ebiederm@xmission.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2024-11-07 14:14:59 -08:00
Linus Torvalds a78282e2c9 Revert "binfmt_elf, coredump: Log the reason of the failed core dumps"
This reverts commit fb97d2eb54.

The logging was questionable to begin with, but it seems to actively
deadlock on the task lock.

 "On second thought, let's not log core dump failures. 'Tis a silly place"

because if you can't tell your core dump is truncated, maybe you should
just fix your debugger instead of adding bugs to the kernel.

Reported-by: Vegard Nossum <vegard.nossum@oracle.com>
Link: https://lore.kernel.org/all/d122ece6-3606-49de-ae4d-8da88846bef2@oracle.com/
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2024-09-26 11:39:02 -07:00
Linus Torvalds f8ffbc365f Merge tag 'pull-stable-struct_fd' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs
Pull 'struct fd' updates from Al Viro:
 "Just the 'struct fd' layout change, with conversion to accessor
  helpers"

* tag 'pull-stable-struct_fd' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs:
  add struct fd constructors, get rid of __to_fd()
  struct fd: representation change
  introduce fd_file(), convert all accessors to it.
2024-09-23 09:35:36 -07:00
Linus Torvalds 667495de21 Merge tag 'execve-v6.12-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux
Pull execve updates from Kees Cook:

 - binfmt_elf: Dump smaller VMAs first in ELF cores (Brian Mak)

 - binfmt_elf: mseal address zero (Jeff Xu)

 - binfmt_elf, coredump: Log the reason of the failed core dumps (Roman
   Kisel)

* tag 'execve-v6.12-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux:
  binfmt_elf: mseal address zero
  binfmt_elf: Dump smaller VMAs first in ELF cores
  binfmt_elf, coredump: Log the reason of the failed core dumps
  coredump: Standartize and fix logging
2024-09-18 11:53:31 +02:00
Al Viro 1da91ea87a introduce fd_file(), convert all accessors to it.
For any changes of struct fd representation we need to
turn existing accesses to fields into calls of wrappers.
Accesses to struct fd::flags are very few (3 in linux/file.h,
1 in net/socket.c, 3 in fs/overlayfs/file.c and 3 more in
explicit initializers).
	Those can be dealt with in the commit converting to
new layout; accesses to struct fd::file are too many for that.
	This commit converts (almost) all of f.file to
fd_file(f).  It's not entirely mechanical ('file' is used as
a member name more than just in struct fd) and it does not
even attempt to distinguish the uses in pointer context from
those in boolean context; the latter will be eventually turned
into a separate helper (fd_empty()).

	NOTE: mass conversion to fd_empty(), tempting as it
might be, is a bad idea; better do that piecewise in commit
that convert from fdget...() to CLASS(...).

[conflicts in fs/fhandle.c, kernel/bpf/syscall.c, mm/memcontrol.c
caught by git; fs/stat.c one got caught by git grep]
[fs/xattr.c conflict]

Reviewed-by: Christian Brauner <brauner@kernel.org>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
2024-08-12 22:00:43 -04:00
Roman Kisel fb97d2eb54 binfmt_elf, coredump: Log the reason of the failed core dumps
Missing, failed, or corrupted core dumps might impede crash
investigations. To improve reliability of that process and consequently
the programs themselves, one needs to trace the path from producing
a core dumpfile to analyzing it. That path starts from the core dump file
written to the disk by the kernel or to the standard input of a user
mode helper program to which the kernel streams the coredump contents.
There are cases where the kernel will interrupt writing the core out or
produce a truncated/not-well-formed core dump without leaving a note.

Add logging for the core dump collection failure paths to be able to reason
what has gone wrong when the core dump is malformed or missing.
Report the size of the data written to aid in diagnosing the user mode
helper.

Signed-off-by: Roman Kisel <romank@linux.microsoft.com>
Link: https://lore.kernel.org/r/20240718182743.1959160-3-romank@linux.microsoft.com
Signed-off-by: Kees Cook <kees@kernel.org>
2024-08-05 21:29:20 -07:00
Thomas Gleixner 7f8af7bac5 signal: Replace BUG_ON()s
These really can be handled gracefully without killing the machine.

Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Frederic Weisbecker <frederic@kernel.org>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
2024-07-29 21:57:35 +02:00
Thomas Gleixner a2b80ce87a signal: Remove task argument from dequeue_signal()
The task pointer which is handed to dequeue_signal() is always current. The
argument along with the first comment about signalfd in that function is
confusing at best. Remove it and use current internally.

Update the stale comment for dequeue_signal() while at it.

Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Frederic Weisbecker <frederic@kernel.org>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
2024-07-29 21:57:35 +02:00
Pavel Begunkov 943ad0b62e kernel: rerun task_work while freezing in get_signal()
io_uring can asynchronously add a task_work while the task is getting
freezed. TIF_NOTIFY_SIGNAL will prevent the task from sleeping in
do_freezer_trap(), and since the get_signal()'s relock loop doesn't
retry task_work, the task will spin there not being able to sleep
until the freezing is cancelled / the task is killed / etc.

Run task_works in the freezer path. Keep the patch small and simple
so it can be easily back ported, but we might need to do some cleaning
after and look if there are other places with similar problems.

Cc: stable@vger.kernel.org
Link: https://github.com/systemd/systemd/issues/33626
Fixes: 12db8b6900 ("entry: Add support for TIF_NOTIFY_SIGNAL")
Reported-by: Julian Orth <ju.orth@gmail.com>
Acked-by: Oleg Nesterov <oleg@redhat.com>
Acked-by: Tejun Heo <tj@kernel.org>
Signed-off-by: Pavel Begunkov <asml.silence@gmail.com>
Link: https://lore.kernel.org/r/89ed3a52933370deaaf61a0a620a6ac91f1e754d.1720634146.git.asml.silence@gmail.com
Signed-off-by: Jens Axboe <axboe@kernel.dk>
2024-07-11 01:51:44 -06:00
Linus Torvalds 2ef32ad224 Merge tag 'for_linus' of git://git.kernel.org/pub/scm/linux/kernel/git/mst/vhost
Pull virtio updates from Michael Tsirkin:
 "Several new features here:

   - virtio-net is finally supported in vduse

   - virtio (balloon and mem) interaction with suspend is improved

   - vhost-scsi now handles signals better/faster

  And fixes, cleanups all over the place"

* tag 'for_linus' of git://git.kernel.org/pub/scm/linux/kernel/git/mst/vhost: (48 commits)
  virtio-pci: Check if is_avq is NULL
  virtio: delete vq in vp_find_vqs_msix() when request_irq() fails
  MAINTAINERS: add Eugenio Pérez as reviewer
  vhost-vdpa: Remove usage of the deprecated ida_simple_xx() API
  vp_vdpa: don't allocate unused msix vectors
  sound: virtio: drop owner assignment
  fuse: virtio: drop owner assignment
  scsi: virtio: drop owner assignment
  rpmsg: virtio: drop owner assignment
  nvdimm: virtio_pmem: drop owner assignment
  wifi: mac80211_hwsim: drop owner assignment
  vsock/virtio: drop owner assignment
  net: 9p: virtio: drop owner assignment
  net: virtio: drop owner assignment
  net: caif: virtio: drop owner assignment
  misc: nsm: drop owner assignment
  iommu: virtio: drop owner assignment
  drm/virtio: drop owner assignment
  gpio: virtio: drop owner assignment
  firmware: arm_scmi: virtio: drop owner assignment
  ...
2024-05-23 12:04:36 -07:00
Mike Christie 240a1853b4 kernel: Remove signal hacks for vhost_tasks
This removes the signal/coredump hacks added for vhost_tasks in:

Commit f9010dbdce ("fork, vhost: Use CLONE_THREAD to fix freezer/ps regression")

When that patch was added vhost_tasks did not handle SIGKILL and would
try to ignore/clear the signal and continue on until the device's close
function was called. In the previous patches vhost_tasks and the vhost
drivers were converted to support SIGKILL by cleaning themselves up and
exiting. The hacks are no longer needed so this removes them.

Signed-off-by: Mike Christie <michael.christie@oracle.com>
Message-Id: <20240316004707.45557-10-michael.christie@oracle.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
2024-05-22 08:31:15 -04:00
Joel Granados 11a921909f kernel misc: Remove the now superfluous sentinel elements from ctl_table array
This commit comes at the tail end of a greater effort to remove the
empty elements at the end of the ctl_table arrays (sentinels) which
will reduce the overall build time size of the kernel and run time
memory bloat by ~64 bytes per sentinel (further information Link :
https://lore.kernel.org/all/ZO5Yx5JFogGi%2FcBo@bombadil.infradead.org/)

Remove the sentinel from ctl_table arrays. Reduce by one the values used
to compare the size of the adjusted arrays.

Signed-off-by: Joel Granados <j.granados@samsung.com>
2024-04-24 09:43:53 +02:00
Linus Torvalds e5eb28f6d1 Merge tag 'mm-nonmm-stable-2024-03-14-09-36' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Pull non-MM updates from Andrew Morton:

 - Kuan-Wei Chiu has developed the well-named series "lib min_heap: Min
   heap optimizations".

 - Kuan-Wei Chiu has also sped up the library sorting code in the series
   "lib/sort: Optimize the number of swaps and comparisons".

 - Alexey Gladkov has added the ability for code running within an IPC
   namespace to alter its IPC and MQ limits. The series is "Allow to
   change ipc/mq sysctls inside ipc namespace".

 - Geert Uytterhoeven has contributed some dhrystone maintenance work in
   the series "lib: dhry: miscellaneous cleanups".

 - Ryusuke Konishi continues nilfs2 maintenance work in the series

	"nilfs2: eliminate kmap and kmap_atomic calls"
	"nilfs2: fix kernel bug at submit_bh_wbc()"

 - Nathan Chancellor has updated our build tools requirements in the
   series "Bump the minimum supported version of LLVM to 13.0.1".

 - Muhammad Usama Anjum continues with the selftests maintenance work in
   the series "selftests/mm: Improve run_vmtests.sh".

 - Oleg Nesterov has done some maintenance work against the signal code
   in the series "get_signal: minor cleanups and fix".

Plus the usual shower of singleton patches in various parts of the tree.
Please see the individual changelogs for details.

* tag 'mm-nonmm-stable-2024-03-14-09-36' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: (77 commits)
  nilfs2: prevent kernel bug at submit_bh_wbc()
  nilfs2: fix failure to detect DAT corruption in btree and direct mappings
  ocfs2: enable ocfs2_listxattr for special files
  ocfs2: remove SLAB_MEM_SPREAD flag usage
  assoc_array: fix the return value in assoc_array_insert_mid_shortcut()
  buildid: use kmap_local_page()
  watchdog/core: remove sysctl handlers from public header
  nilfs2: use div64_ul() instead of do_div()
  mul_u64_u64_div_u64: increase precision by conditionally swapping a and b
  kexec: copy only happens before uchunk goes to zero
  get_signal: don't initialize ksig->info if SIGNAL_GROUP_EXIT/group_exec_task
  get_signal: hide_si_addr_tag_bits: fix the usage of uninitialized ksig
  get_signal: don't abuse ksig->info.si_signo and ksig->sig
  const_structs.checkpatch: add device_type
  Normalise "name (ad@dr)" MODULE_AUTHORs to "name <ad@dr>"
  dyndbg: replace kstrdup() + strchr() with kstrdup_and_replace()
  list: leverage list_is_head() for list_entry_is_head()
  nilfs2: MAINTAINERS: drop unreachable project mirror site
  smp: make __smp_processor_id() 0-argument macro
  fat: fix uninitialized field in nostale filehandles
  ...
2024-03-14 18:03:09 -07:00
Oleg Nesterov a436184e3b get_signal: don't initialize ksig->info if SIGNAL_GROUP_EXIT/group_exec_task
This initialization is incomplete and unnecessary, neither do_group_exit()
nor PF_USER_WORKER need ksig->info.

Link: https://lkml.kernel.org/r/20240226165653.GA20834@redhat.com
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Eric W. Biederman <ebiederm@xmission.com>
Cc: Peter Collingbourne <pcc@google.com>
Cc: Wen Yang <wenyang.linux@foxmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2024-03-06 13:07:39 -08:00
Oleg Nesterov dd69edd643 get_signal: hide_si_addr_tag_bits: fix the usage of uninitialized ksig
ksig->ka and ksig->info are not initialized if get_signal() returns 0 or
if the caller is PF_USER_WORKER.

Check signr != 0 before SA_EXPOSE_TAGBITS and move the "out" label down.

The latter means that ksig->sig won't be initialized if a PF_USER_WORKER
thread gets a fatal signal but this is fine, PF_USER_WORKER's don't use
ksig. And there is nothing new, in this case ksig->ka and ksig-info are
not initialized anyway. Add a comment.

Link: https://lkml.kernel.org/r/20240226165650.GA20829@redhat.com
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Eric W. Biederman <ebiederm@xmission.com>
Cc: Peter Collingbourne <pcc@google.com>
Cc: Wen Yang <wenyang.linux@foxmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2024-03-06 13:07:39 -08:00
Oleg Nesterov 49fd5f5ac4 get_signal: don't abuse ksig->info.si_signo and ksig->sig
Patch series "get_signal: minor cleanups and fix".

Lets remove this clear_siginfo() right now.  It is incomplete (and thus
looks confusing) and unnecessary.  Also, PF_USER_WORKER's already don't
get a fully initialized ksig anyway.


This patch (of 3):

Cleanup and preparation for the next changes.

get_signal() uses signr or ksig->info.si_signo or ksig->sig in a chaotic
way, this looks confusing. Change it to always use signr.

Link: https://lkml.kernel.org/r/20240226165612.GA20787@redhat.com
Link: https://lkml.kernel.org/r/20240226165647.GA20826@redhat.com
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Eric W. Biederman <ebiederm@xmission.com>
Cc: Peter Collingbourne <pcc@google.com>
Cc: Wen Yang <wenyang.linux@foxmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2024-03-06 13:07:39 -08:00
Christian Brauner e1fb1dc08e pidfd: allow to override signal scope in pidfd_send_signal()
Right now we determine the scope of the signal based on the type of
pidfd. There are use-cases where it's useful to override the scope of
the signal. For example in [1]. Add flags to determine the scope of the
signal:

(1) PIDFD_SIGNAL_THREAD: send signal to specific thread reference by @pidfd
(2) PIDFD_SIGNAL_THREAD_GROUP: send signal to thread-group of @pidfd
(2) PIDFD_SIGNAL_PROCESS_GROUP: send signal to process-group of @pidfd

Since we now allow specifying PIDFD_SEND_PROCESS_GROUP for
pidfd_send_signal() to send signals to process groups we need to adjust
the check restricting si_code emulation by userspace to account for
PIDTYPE_PGID.

Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Link: https://github.com/systemd/systemd/issues/31093 [1]
Link: https://lore.kernel.org/r/20240210-chihuahua-hinzog-3945b6abd44a@brauner
Link: https://lore.kernel.org/r/20240214123655.GB16265@redhat.com
Signed-off-by: Christian Brauner <brauner@kernel.org>
2024-02-21 09:46:08 +01:00
Oleg Nesterov 81b9d8ac06 pidfd: change pidfd_send_signal() to respect PIDFD_THREAD
Turn kill_pid_info() into kill_pid_info_type(), this allows to pass any
pid_type to group_send_sig_info(), despite its name it should work fine
even if type = PIDTYPE_PID.

Change pidfd_send_signal() to use PIDTYPE_PID or PIDTYPE_TGID depending
on PIDFD_THREAD.

While at it kill another TODO comment in pidfd_show_fdinfo(). As Christian
expains fdinfo reports f_flags, userspace can already detect PIDFD_THREAD.

Reviewed-by: Tycho Andersen <tandersen@netflix.com>
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/20240209130650.GA8048@redhat.com
Signed-off-by: Christian Brauner <brauner@kernel.org>
2024-02-10 22:37:23 +01:00
Oleg Nesterov c044a95026 signal: fill in si_code in prepare_kill_siginfo()
So that do_tkill() can use this helper too. This also simplifies
the next patch.

TODO: perhaps we can kill prepare_kill_siginfo() and change the
callers to use SEND_SIG_NOINFO,  but this needs some changes in
__send_signal_locked() and TP_STORE_SIGINFO().

Reviewed-by: Tycho Andersen <tandersen@netflix.com>
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/20240209130620.GA8039@redhat.com
Signed-off-by: Christian Brauner <brauner@kernel.org>
2024-02-10 22:37:18 +01:00
Oleg Nesterov 9ed52108f6 pidfd: change do_notify_pidfd() to use __wake_up(poll_to_key(EPOLLIN))
rather than wake_up_all(). This way do_notify_pidfd() won't wakeup the
POLLHUP-only waiters which wait for pid_task() == NULL.

TODO:
    - as Christian pointed out, this asks for the new wake_up_all_poll()
      helper, it can already have other users.

    - we can probably discriminate the PIDFD_THREAD and non-PIDFD_THREAD
      waiters, but this needs more work. See
      https://lore.kernel.org/all/20240205140848.GA15853@redhat.com/

Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/20240205141348.GA16539@redhat.com
Reviewed-by: Tycho Andersen <tandersen@netflix.com>
Signed-off-by: Christian Brauner <brauner@kernel.org>
2024-02-06 13:52:51 +01:00
Oleg Nesterov 64bef697d3 pidfd: implement PIDFD_THREAD flag for pidfd_open()
With this flag:

	- pidfd_open() doesn't require that the target task must be
	  a thread-group leader

	- pidfd_poll() succeeds when the task exits and becomes a
	  zombie (iow, passes exit_notify()), even if it is a leader
	  and thread-group is not empty.

	  This means that the behaviour of pidfd_poll(PIDFD_THREAD,
	  pid-of-group-leader) is not well defined if it races with
	  exec() from its sub-thread; pidfd_poll() can succeed or not
	  depending on whether pidfd_task_exited() is called before
	  or after exchange_tids().

	  Perhaps we can improve this behaviour later, pidfd_poll()
	  can probably take sig->group_exec_task into account. But
	  this doesn't really differ from the case when the leader
	  exits before other threads (so pidfd_poll() succeeds) and
	  then another thread execs and pidfd_poll() will block again.

thread_group_exited() is no longer used, perhaps it can die.

Co-developed-by: Tycho Andersen <tycho@tycho.pizza>
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/20240131132602.GA23641@redhat.com
Tested-by: Tycho Andersen <tandersen@netflix.com>
Reviewed-by: Tycho Andersen <tandersen@netflix.com>
Signed-off-by: Christian Brauner <brauner@kernel.org>
2024-02-02 13:12:28 +01:00
Oleg Nesterov 21e25205d7 pidfd: don't do_notify_pidfd() if !thread_group_empty()
do_notify_pidfd() makes no sense until the whole thread group exits, change
do_notify_parent() to check thread_group_empty().

This avoids the unnecessary do_notify_pidfd() when tsk is not a leader, or
it exits before other threads, or it has a ptraced EXIT_ZOMBIE sub-thread.

Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/20240127132407.GA29136@redhat.com
Reviewed-by: Tycho Andersen <tandersen@netflix.com>
Signed-off-by: Christian Brauner <brauner@kernel.org>
2024-02-02 13:12:28 +01:00
Oleg Nesterov b454ec2922 kernel/signal.c: simplify force_sig_info_to_task(), kill recalc_sigpending_and_wake()
The purpose of recalc_sigpending_and_wake() is not clear, it looks
"obviously unneeded" because we are going to send the signal which can't
be blocked or ignored.

Add the comment to explain why we can't rely on send_signal_locked() and
make this logic more simple/explicit.  recalc_sigpending_and_wake() has no
other users, it can die.

In fact I think we don't even need signal_wake_up(), the target task must
be either current or a TASK_TRACED child, otherwise the usage of siglock
is not safe.  But this needs another change.

Link: https://lkml.kernel.org/r/20231120151649.GA15995@redhat.com
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Cc: Eric Biederman <ebiederm@xmission.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2023-12-10 17:21:32 -08:00
Oleg Nesterov 61a7a5e25f introduce for_other_threads(p, t)
Cosmetic, but imho it makes the usage look more clear and simple, the new
helper doesn't require to initialize "t".

After this change while_each_thread() has only 3 users, and it is only
used in the do/while loops.

Link: https://lkml.kernel.org/r/20231030155710.GA9095@redhat.com
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Reviewed-by: Christian Brauner <brauner@kernel.org>
Cc: "Eric W. Biederman" <ebiederm@xmission.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2023-12-10 17:21:25 -08:00
Linus Torvalds 8f6f76a6a2 Merge tag 'mm-nonmm-stable-2023-11-02-14-08' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Pull non-MM updates from Andrew Morton:
 "As usual, lots of singleton and doubleton patches all over the tree
  and there's little I can say which isn't in the individual changelogs.

  The lengthier patch series are

   - 'kdump: use generic functions to simplify crashkernel reservation
     in arch', from Baoquan He. This is mainly cleanups and
     consolidation of the 'crashkernel=' kernel parameter handling

   - After much discussion, David Laight's 'minmax: Relax type checks in
     min() and max()' is here. Hopefully reduces some typecasting and
     the use of min_t() and max_t()

   - A group of patches from Oleg Nesterov which clean up and slightly
     fix our handling of reads from /proc/PID/task/... and which remove
     task_struct.thread_group"

* tag 'mm-nonmm-stable-2023-11-02-14-08' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: (64 commits)
  scripts/gdb/vmalloc: disable on no-MMU
  scripts/gdb: fix usage of MOD_TEXT not defined when CONFIG_MODULES=n
  .mailmap: add address mapping for Tomeu Vizoso
  mailmap: update email address for Claudiu Beznea
  tools/testing/selftests/mm/run_vmtests.sh: lower the ptrace permissions
  .mailmap: map Benjamin Poirier's address
  scripts/gdb: add lx_current support for riscv
  ocfs2: fix a spelling typo in comment
  proc: test ProtectionKey in proc-empty-vm test
  proc: fix proc-empty-vm test with vsyscall
  fs/proc/base.c: remove unneeded semicolon
  do_io_accounting: use sig->stats_lock
  do_io_accounting: use __for_each_thread()
  ocfs2: replace BUG_ON() at ocfs2_num_free_extents() with ocfs2_error()
  ocfs2: fix a typo in a comment
  scripts/show_delta: add __main__ judgement before main code
  treewide: mark stuff as __ro_after_init
  fs: ocfs2: check status values
  proc: test /proc/${pid}/statm
  compiler.h: move __is_constexpr() to compiler.h
  ...
2023-11-02 20:53:31 -10:00
Linus Torvalds 1e0c505e13 Merge tag 'asm-generic-6.7' of git://git.kernel.org/pub/scm/linux/kernel/git/arnd/asm-generic
Pull ia64 removal and asm-generic updates from Arnd Bergmann:

 - The ia64 architecture gets its well-earned retirement as planned,
   now that there is one last (mostly) working release that will be
   maintained as an LTS kernel.

 - The architecture specific system call tables are updated for the
   added map_shadow_stack() syscall and to remove references to the
   long-gone sys_lookup_dcookie() syscall.

* tag 'asm-generic-6.7' of git://git.kernel.org/pub/scm/linux/kernel/git/arnd/asm-generic:
  hexagon: Remove unusable symbols from the ptrace.h uapi
  asm-generic: Fix spelling of architecture
  arch: Reserve map_shadow_stack() syscall number for all architectures
  syscalls: Cleanup references to sys_lookup_dcookie()
  Documentation: Drop or replace remaining mentions of IA64
  lib/raid6: Drop IA64 support
  Documentation: Drop IA64 from feature descriptions
  kernel: Drop IA64 support from sig_fault handlers
  arch: Remove Itanium (IA-64) architecture
2023-11-01 15:28:33 -10:00
Li kunyu a287116af1 kernel/signal: remove unnecessary NULL values from ucounts
ucounts is assigned first, so it does not need to initialize the
assignment.

Link: https://lkml.kernel.org/r/20230926022410.4280-1-kunyu@nfschina.com
Signed-off-by: Li kunyu <kunyu@nfschina.com>
Acked-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2023-10-18 14:43:22 -07:00
Oleg Nesterov e5ecf29c50 signal: complete_signal: use __for_each_thread()
do/while_each_thread should be avoided when possible.

Link: https://lkml.kernel.org/r/20230909164537.GA11633@redhat.com
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Cc: Eric W. Biederman <ebiederm@xmission.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2023-10-04 10:41:57 -07:00
Oleg Nesterov 3983520491 __kill_pgrp_info: simplify the calculation of return value
No need to calculate/check the "success" variable, we can kill it and update
retval in the main loop unless it is zero.

Link: https://lkml.kernel.org/r/20230823171455.GA12188@redhat.com
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Suggested-by: David Laight <David.Laight@ACULAB.COM>
Cc: Eric W. Biederman <ebiederm@xmission.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2023-10-04 10:41:56 -07:00