# Why MKL's vdsin is slower than the intrinsic sin?

**URL:** <https://fortran-lang.discourse.group/t/why-mkls-vdsin-is-slower-than-the-intrinsic-sin/3108>\
**Category:** Uncategorized\
**Created:** [April 2, 2022, 7:25pm UTC](https://fortran-lang.discourse.group/t/why-mkls-vdsin-is-slower-than-the-intrinsic-sin/3108 "2022-04-02T19:25:51Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![CRquantum](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/crquantum/32/730_2.png) [@CRquantum](https://fortran-lang.discourse.group/u/CRquantum)\
**Post date:** [April 2, 2022, 7:25pm UTC](https://fortran-lang.discourse.group/t/why-mkls-vdsin-is-slower-than-the-intrinsic-sin/3108/1 "2022-04-02T19:25:51Z")

</div>

Dear all,

I notice from below thread,

> [@Simple summation 8x slower than in Julia](https://fortran-lang.discourse.group/t/simple-summation-8x-slower-than-in-julia/1171/98):
>
> Sorry to bump up this old thread. I recently ran into very slow loops (slower than the equivalent ones in MATLAB) with 500k or so iterations which used cosh, cos/sin in each iteration. Replacing these by the Intel mkl vml routines lead to a 5X speed up. My processor only has avx2 instructions and I think if I try it on a processor with avx2-512 instruction set, this can be slashed even further. I use intel oneAPI under Linux.

That using MKL may speed up the computation, so I tried MKL’s `vdsin` function and using the same example as in the above thread,

```auto
!include "_rms.fi"
program avx
implicit none
!include "mkl_vml.f90"
integer, parameter :: dp = kind(0.d0)
real(dp) :: t1, t2, r

call cpu_time(t1)
r = f(100000000)
call cpu_time(t2)

print *, "Time", t2-t1
print *, r

contains

    real(dp) function f(N) result(r)
    integer, intent(in) :: N
    integer :: i
    real(dp) :: j(N)
    !r = 0.0_dp
    
    !j = 1.0_dp
    !do while (j<=N)
    ! r = r + sin(j)
    ! j = j + 1.0_dp
    !enddo
    
    !j = [(i, i = 1, N)]
     
    
    !r = sum( sin( dble( (/(i, i = 1,N)/) ) ) )
    
    
    
    call vdsin(N,(dble([(i,i=1,N)])),j)
    
    r = sum(j)
    
    !r = sum(sin(dble([(i,i=1,N)])))
    
    !do i = 1, N
    ! r = r + sin(dble(i))
    !end do
    
    
    return
    end function

end program

```

However, before using MKL it costs 0.6s, after using MKL it cost 1.3 second. My CPU is Xeon 2186M.

So I am confused, does anyone know how to use MKL correctly?

Thank much in advance!

PS.  
I am using Intel OneAPI 2020.3 on Windows, setting is -O3 -xHost, and

 ![image](https://global.discourse-cdn.com/free1/uploads/fortran_lang/original/2X/e/e2dbd4a5a55b77a8fd4c3d5f899c94257f93d9ae.png)

The original code by @certik is

```auto
program avx
implicit none
integer, parameter :: dp = kind(0.d0)
real(dp) :: t1, t2, r

call cpu_time(t1)
r = f(100000000)
call cpu_time(t2)

print *, "Time", t2-t1
print *, r

contains

    real(dp) function f(N) result(r)
    integer, intent(in) :: N
    integer :: i
    r = 0
    do i = 1, N
        r = r + sin(real(i,dp))
    end do
    end function

end program

```

---

<div class="post-metadata">

**Author:** ![septc](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/septc/32/77_2.png) [@septc](https://fortran-lang.discourse.group/u/septc)\
**Post date:** [April 2, 2022, 7:35pm UTC](https://fortran-lang.discourse.group/t/why-mkls-vdsin-is-slower-than-the-intrinsic-sin/3108/2 "2022-04-02T19:35:21Z")

</div>

> [@CRquantum](#):
>
> `call vdsin(N,(dble([(i,i=1,N)])),j)`

Isn’t it better to create the above input array before entering this part

```auto
call cpu_time(t1)
r = f(100000000)
call cpu_time(t2)

```

and pass the array to `f()` so that it measures purely the time for evaluating `sin`?

---

<div class="post-metadata">

**Author:** ![CRquantum](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/crquantum/32/730_2.png) [@CRquantum](https://fortran-lang.discourse.group/u/CRquantum)\
**Post date:** [April 2, 2022, 7:38pm UTC](https://fortran-lang.discourse.group/t/why-mkls-vdsin-is-slower-than-the-intrinsic-sin/3108/3 "2022-04-02T19:38:46Z")

</div>

Thanks @septc ! Yeah, uhm, well I just try to do the same thing as @certik did in that thread, his code is,

```auto
program avx
implicit none
integer, parameter :: dp = kind(0.d0)
real(dp) :: t1, t2, r

call cpu_time(t1)
r = f(100000000)
call cpu_time(t2)

print *, "Time", t2-t1
print *, r

contains

    real(dp) function f(N) result(r)
    integer, intent(in) :: N
    integer :: i
    r = 0
    do i = 1, N
        r = r + sin(real(i,dp))
    end do
    end function

end program

```

If I do

```auto
r = sum(sin(dble([(i,i=1,N)])))

```

it is two times faster than using MKL’s vdsin as below

```auto
    call vdsin(N,(dble([(i,i=1,N)])),j)
    r = sum(j)

```
