glibc

History

Adhemerval Zanella 8a0152b61b math: New generic fmaf implementation The current implementation relies on setting the rounding mode for different calculations (FE_TOWARDZERO) to obtain correctly rounded results. For most CPUs, this adds significant performance overhead because it requires executing a typically slow instruction (to get/set the floating-point status), necessitates flushing the pipeline, and breaks some compiler assumptions/optimizations. The original implementation adds tests to handle underflow in corner cases, whereas this implementation uses a different strategy that checks both the mantissa and the result to determine whether the result is not subject to double rounding. I tested this implementation on various targets (x86_64, i686, arm, aarch64, powerpc), including some by manually disabling the compiler instructions. Performance-wise, it shows large improvements: reciprocal-throughput master patched improvement x86_64 [1] 58.09 7.96 7.33x i686 [1] 279.41 16.97 16.46x aarch64 [2] 26.09 4.10 6.35x armhf [2] 30.25 4.20 7.18x powerpc [3] 9.46 1.46 6.45x latency master patched improvement x86_64 64.50 14.25 4.53x i686 304.39 61.04 4.99x aarch64 27.71 5.74 4.82x armhf 33.46 7.34 4.55x powerpc 10.96 2.65 4.13x Checked on x86_64-linux-gnu and i686-linux-gnu with —disable-multi-arch, and on arm-linux-gnueabihf. [1] gcc 15.2.1, Zen3 [2] gcc 15.2.1, Neoverse N1 [3] gcc 15.2.1, POWER10 Signed-off-by: Szabolcs Nagy <nsz@gcc.gnu.org> Co-authored-by: Adhemerval Zanella <adhemerval.zanella@linaro.org> Co-authored-by: Wilco Dijkstra <Wilco.Dijkstra@arm.com> Reviewed-by: Wilco Dijkstra <Wilco.Dijkstra@arm.com>		2025-11-27 14:52:25 -03:00
..
dbl-64	math: New generic fmaf implementation	2025-11-27 14:52:25 -03:00
float128	math: Handle fabsf128 !__USE_EXTERN_INLINES	2025-11-17 11:17:07 -03:00
flt-32	math: Don't redirect inlined builtin math functions	2025-11-17 11:17:07 -03:00
ldbl-64-128	…
ldbl-96	stdlib: Remove longlong.h	2025-11-26 10:10:06 -03:00
ldbl-128	stdlib: Remove longlong.h	2025-11-26 10:10:06 -03:00
ldbl-128ibm	stdlib: Remove longlong.h	2025-11-26 10:10:06 -03:00
ldbl-128ibm-compat	Enable --enable-fortify-source with clang	2025-11-21 13:13:11 -03:00
ldbl-opt	Change fromfp functions to return floating types following C23 (bug 28327)	2025-11-13 00:04:21 +00:00
soft-fp	Suppress -Wmaybe-uninitialized only for gcc	2025-10-21 09:24:05 -03:00
Makefile	…
ieee754.h	…
k_standard.c	…
k_standardf.c	…
k_standardl.c	…
libm-alias-finite.h	…
s_lib_version.c	…
s_matherr.c	…
s_signgam.c	…