# Compiler option matmul speedup

**URL:** <https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518>\
**Category:** Uncategorized\
**Created:** [September 15, 2023, 4:00pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518 "2023-09-15T16:00:07Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![rbitr](https://avatars.discourse-cdn.com/v4/letter/r/f1d935/32.png) [@rbitr](https://fortran-lang.discourse.group/u/rbitr)\
**Post date:** [September 15, 2023, 4:00pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/1 "2023-09-15T16:00:07Z")

</div>

Hi, I’m looking at the impact of different compiler options on the speed of vector matrix multiplication. I compared with both Fortran and C, and got essentially the same top speed but Fortran’s matmul intrinsic was much faster with no optimization turned on (and interestingly gets slowed way down by -O3). See the gist below. I’m curious if anybody has thoughts on the analysis I did - are there other options I should try, other circumstances in which the matmuls are occurring, something I overlooked? Thanks!

> <https://gist.github.com/rbitr/3b86154f78a0f0832e8bd1716152365d>
>
> There are more than three files. show original

---

<div class="post-metadata">

**Author:** ![sblionel](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/sblionel/32/853_2.png) [@sblionel](https://fortran-lang.discourse.group/u/sblionel)\
**Post date:** [September 15, 2023, 5:13pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/2 "2023-09-15T17:13:53Z")

</div>

You don’t say which compiler you are using, but from the options I’m guessing it’s gfortran. Please note that different compilers will not necessarily show the same behavior.

---

<div class="post-metadata">

**Author:** ![PierU](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/pieru/32/1848_2.png) [@PierU](https://fortran-lang.discourse.group/u/PierU)\
**Post date:** [September 15, 2023, 5:14pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/3 "2023-09-15T17:14:40Z")

</div>

For information your Fortran OpenMP code is not correct: `j` and `val` should be `private`. In C they are declared within the parallel region, so they are implicitely private.

By the way, multithreading matrix-vector multiplications is known to not be very efficient, at least on consumer hardware.

---

<div class="post-metadata">

**Author:** ![rbitr](https://avatars.discourse-cdn.com/v4/letter/r/f1d935/32.png) [@rbitr](https://fortran-lang.discourse.group/u/rbitr)\
**Post date:** [September 15, 2023, 5:28pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/4 "2023-09-15T17:28:45Z")

</div>

Yes it’s gfortran-10

---

<div class="post-metadata">

**Author:** ![alirezagh76](https://avatars.discourse-cdn.com/v4/letter/a/838e76/32.png) [@alirezagh76](https://fortran-lang.discourse.group/u/alirezagh76)\
**Post date:** [September 15, 2023, 5:54pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/5 "2023-09-15T17:54:29Z")

</div>

Your example is matrix-vector multiplication. Perhaps the title could be clearer if it would be explicitly mentioned matrix-vector multiplication. If you try with Intel compilers, you can use `-qopt-report` to find out which optimizations are applied. On modern CPUs you may want to benefit from AVX2 or AVX512, …

---

<div class="post-metadata">

**Author:** ![rbitr](https://avatars.discourse-cdn.com/v4/letter/r/f1d935/32.png) [@rbitr](https://fortran-lang.discourse.group/u/rbitr)\
**Post date:** [September 15, 2023, 7:40pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/6 "2023-09-15T19:40:01Z")

</div>

Thanks, I made the change but it didn’t affect the performance. I am surprised (naively?) that matrix-vector multiplication is not efficient across threads because it seems like something that would be very easy to parallelize.

---

<div class="post-metadata">

**Author:** ![PierU](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/pieru/32/1848_2.png) [@PierU](https://fortran-lang.discourse.group/u/PierU)\
**Post date:** [September 15, 2023, 8:23pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/7 "2023-09-15T20:23:42Z")

</div>

It’s easy to parallelize, but it has a moderately low arithmetic intensity (AI), that is the ratio between the number of operations and the number of memory read/write is low. In these conditions, the bottleneck is the bandwidth between the CPU and the memory, and using more cores does not help.

This is particularly true on the Intel Core CPUs, which have pretty good monocore performances (all the more than the turbo boost frequency is often enabled when a single core is used) that are high enough to saturate the bandwidth to/from the memory for such simple computations. This is less true on the Xeon line, which have a higher bandwith (and no turbo boost, iirc), or on the AMD CPUs, which have more cores with lower monocore performances.

---

<div class="post-metadata">

**Author:** ![ivanpribec](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/ivanpribec/32/3290_2.png) [@ivanpribec](https://fortran-lang.discourse.group/u/ivanpribec)\
**Post date:** [September 15, 2023, 10:16pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/8 "2023-09-15T22:16:16Z")

</div>

With **gfortran** you can use the [`-fexternal-blas`](https://gcc.gnu.org/onlinedocs/gfortran/Code-Gen-Options.html#index-fexternal-blas) compiler flag, which inserts calls to BLAS in place of the intrinsic `matmul` function above a certain matrix size. In addition you have to link an optimized BLAS library:

| Library | Link flags | Comment |
| --- | --- | --- |
| [OpenBLAS](https://github.com/xianyi/OpenBLAS) | `-lopenblas` | |
| [BLIS](https://github.com/flame/blis) | `-lblis` | |
| [Intel oneMKL](https://www.intel.com/content/www/us/en/docs/onemkl/developer-reference-fortran/2023-2/blas-and-sparse-blas-routines.html) | see [Link Line Advisor for OneMKL](https://www.intel.com/content/www/us/en/developer/tools/oneapi/onemkl-link-line-advisor.html#gs.57febj) | for Intel processors |
| [Accelerate](https://developer.apple.com/documentation/accelerate?language=objc) | `-framework Accelerate` | for macOS |
| [AOCL-BLAS](https://www.amd.com/en/developer/aocl/dense.html) | `-lblis` | for AMD processors; derived from BLIS |
| [Arm Performance Libraries](https://developer.arm.com/downloads/-/arm-performance-libraries/get-started-with-armpl-free-version/single-page) | `-larmpl_lp64` | for Arm processors |
| [MATLAB BLAS](https://www.mathworks.com/help/matlab/matlab_external/calling-lapack-and-blas-functions-from-mex-files.html) | `-lmwblas` | for MATLAB MEX functions |

---

<div class="post-metadata">

**Author:** ![alirezagh76](https://avatars.discourse-cdn.com/v4/letter/a/838e76/32.png) [@alirezagh76](https://fortran-lang.discourse.group/u/alirezagh76)\
**Post date:** [September 16, 2023, 2:49am UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/9 "2023-09-16T02:49:00Z")

</div>

AMD seems to have higher single-core bandwidth:  
[https://sites.utexas.edu/jdm4372/2023/04/25/the-evolution-of-single-core-bandwidth-in-multicore-processors/](https://sites.utexas.edu/jdm4372/2023/04/25/the-evolution-of-single-core-bandwidth-in-multicore-processors/)

---

<div class="post-metadata">

**Author:** ![PierU](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/pieru/32/1848_2.png) [@PierU](https://fortran-lang.discourse.group/u/PierU)\
**Post date:** [September 16, 2023, 6:35am UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/10 "2023-09-16T06:35:13Z")

</div>

That’s true that in the last year, the bandwidth tend to grow faster than the monocore performances.

---

<div class="post-metadata">

**Author:** ![rbitr](https://avatars.discourse-cdn.com/v4/letter/r/f1d935/32.png) [@rbitr](https://fortran-lang.discourse.group/u/rbitr)\
**Post date:** [September 16, 2023, 10:32pm UTC](https://fortran-lang.discourse.group/t/compiler-option-matmul-speedup/6518/11 "2023-09-16T22:32:21Z")

</div>

Thanks, your explanation made a big difference in how I approached the problem I’m working on.
