# Does LAPACK/BLAS automatically use multi cores or threads?

**URL:** <https://fortran-lang.discourse.group/t/does-lapack-blas-automatically-use-multi-cores-or-threads/4072>\
**Category:** Uncategorized\
**Created:** [July 26, 2022, 8:22am UTC](https://fortran-lang.discourse.group/t/does-lapack-blas-automatically-use-multi-cores-or-threads/4072 "2022-07-26T08:22:42Z")\
**Posts on this page:** 1\
**Showing post:** 23

<div class="post-metadata">

**Author:** ![CRquantum](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/crquantum/32/730_2.png) [@CRquantum](https://fortran-lang.discourse.group/u/CRquantum)\
**Post date:** [July 28, 2022, 10:37pm UTC](https://fortran-lang.discourse.group/t/does-lapack-blas-automatically-use-multi-cores-or-threads/4072/23 "2022-07-28T22:37:28Z")

</div>

I found one problem, it seems

```auto
call cpu_time()

```

does not measure the time correctly when multiple threads are involved.  
When using Intel OneAPI with MKL, I use John Burkardt’s `wtime()` instead, so I run the code below,

```auto
program kxk
   integer, parameter :: wp=selected_real_kind(14), n=5000
   real(wp) :: a(n,n), b(n,n), c(n,n)
   real(wp) :: cpu1, cpu0

   call random_number( a ); a = a - 0.5_wp
   call random_number( b ); b = b - 0.5_wp
   c = 0.0_wp

   cpu0 = wtime()
   c = matmul( a, b )
   cpu1 = wtime()
   write(*,*) 'c11=', c(1,1), 'cpu_time=', (cpu1-cpu0), ' GFLOPS=', 2*real(n,kind=wp)**3/(cpu1-cpu0)/1.e9_wp

   cpu0 = wtime()
   call dgemm( 'N', 'N', n, n, n, 1.0_wp, a, n, b, n, 0.0_wp, c, n )
   cpu1 = wtime()
   write(*,*) 'c11=', c(1,1), 'cpu_time=', (cpu1-cpu0), ' GFLOPS=', 2*real(n,kind=wp)**3/(cpu1-cpu0)/1.e9_wp

contains
    function wtime ( )

! ***************************************************************************** 80
!
!! WTIME returns a reading of the wall clock time.
!
! Discussion:
!
! To get the elapsed wall clock time, call WTIME before and after a given
! operation, and subtract the first reading from the second.
!
! This function is meant to suggest the similar routines:
!
! "omp_get_wtime ( )" in OpenMP,
! "MPI_Wtime ( )" in MPI,
! and "tic" and "toc" in MATLAB.
!
! Licensing:
!
! This code is distributed under the GNU LGPL license. 
!
! Modified:
!
! 27 April 2009
!
! Author:
!
! John Burkardt
!
! Parameters:
!
! Output, real ( kind = rk ) WTIME, the wall clock reading, in seconds.
!
  implicit none

  integer, parameter :: rk = kind ( 1.0D+00 )

  integer clock_max
  integer clock_rate
  integer clock_reading
  real ( kind = rk ) wtime

  call system_clock ( clock_reading, clock_rate, clock_max )

  wtime = real ( clock_reading, kind = rk ) &
        / real ( clock_rate, kind = rk )

  return
  end function wtime  
end program kxk

```

With Intel MKL’s matmul,

 ![image](https://global.discourse-cdn.com/free1/uploads/fortran_lang/original/2X/6/67a66a2d9d288232391265848b76d16394e96246.png)

when n=5000, what I got is,

```auto
 c11= 2.90899748439952 cpu_time= 1.50899999999092 GFLOPS=
   165.672630882375
 c11= 2.90899748439952 cpu_time= 4.14900000000489 GFLOPS=
   60.2554832489046

```

Looks like MKL automatically parallelize `matmul`, but not for `dgemm`.

---

_[View the full topic](https://fortran-lang.discourse.group/t/does-lapack-blas-automatically-use-multi-cores-or-threads/4072)._
