# Code slower than C++ version

**URL:** <https://fortran-lang.discourse.group/t/code-slower-than-c-version/2066>\
**Category:** Help\
**Created:** [October 13, 2021, 8:50pm UTC](https://fortran-lang.discourse.group/t/code-slower-than-c-version/2066 "2021-10-13T20:50:52Z")\
**Posts on this page:** 6\
**Page:** 2

<div class="post-metadata">

**Author:** ![lmiq](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/lmiq/32/555_2.png) [@lmiq](https://fortran-lang.discourse.group/u/lmiq)\
**Post date:** [October 14, 2021, 2:56pm UTC](https://fortran-lang.discourse.group/t/code-slower-than-c-version/2066/21 "2021-10-14T14:56:59Z")

</div>

> [@certik](#):
>
> and handle nans.

This is probably deliberate to keep a safer option as the default. What I miss there is that `@fastmath` (or the compiler option in Fortran) should skip those checks, shouldn’t it? ~~(it does not)~~

---

<div class="post-metadata">

**Author:** ![oscardssmith](https://avatars.discourse-cdn.com/v4/letter/o/b9e5f3/32.png) [@oscardssmith](https://fortran-lang.discourse.group/u/oscardssmith)\
**Post date:** [October 14, 2021, 3:05pm UTC](https://fortran-lang.discourse.group/t/code-slower-than-c-version/2066/22 "2021-10-14T15:05:01Z")

</div>

`three_median2` and `three_median1` (and associated assembly) is actually an incorrect implementation if you are following strict ieee rules. `three_median2(0.0,-0.0,0.0)==-0.0` which and wrong, and both `three_median1` and `three_median2` produce `NaN` for `three_median2(0.0,NaN,0.0)==NaN`.

If you define `three_median1` as `@fastmath max(min(a,b),min(max(a,b),c))`, Julia is able to optimize to the same assembly as `three_median2`

---

<div class="post-metadata">

**Author:** ![certik](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/certik/32/4_2.png) [@certik](https://fortran-lang.discourse.group/u/certik)\
**Post date:** [October 14, 2021, 3:25pm UTC](https://fortran-lang.discourse.group/t/code-slower-than-c-version/2066/23 "2021-10-14T15:25:25Z")

</div>

> [@oscardssmith](#):
>
> If you define `three_median1` as `@fastmath max(min(a,b),min(max(a,b),c))` , Julia is able to optimize to the same assembly as `three_median2`

Perfect. Thanks for the feedback @oscardssmith. @lmiq, do you want to try it if you now get the same performance with `three_median1`?

---

<div class="post-metadata">

**Author:** ![lmiq](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/lmiq/32/555_2.png) [@lmiq](https://fortran-lang.discourse.group/u/lmiq)\
**Post date:** [October 14, 2021, 4:11pm UTC](https://fortran-lang.discourse.group/t/code-slower-than-c-version/2066/24 "2021-10-14T16:11:10Z")

</div>

Yes, it does:

```auto
julia> function three_median1(a, b, c)
         res = @fastmath max(min(a,b),min(max(a,b),c))
         return res
       end
three_median1 (generic function with 1 method)

julia> @btime test($x,$three_median1)
  10.053 μs (0 allocations: 0 bytes)
5017.136463388665

```

(I think I had put `@fastmath` on the call to the function on the `test` function, and that didn’t work)

one curiosity, in this last version I though I would do one comparison less:

```auto
julia> function three_median3(a,b,c)
           a1, a2 = a < b ? (a, b) : (b, a)
           a3 = a2 < c ? a2 : c
           res = a3 > a1 ? a3 : a1
           return res
       end
three_median3 (generic function with 1 method)

```

but the native code is the same:

```auto
julia> @code_native three_median3(1.0,1.0,1.0)
	.text
; ┌ @ REPL[32] within `three_median3'
	vminsd	%xmm1, %xmm0, %xmm3
	vmaxsd	%xmm0, %xmm1, %xmm0
; │ @ REPL[32]:3 within `three_median3'
	vminsd	%xmm2, %xmm0, %xmm0
; │ @ REPL[32] within `three_median3'
	vmaxsd	%xmm3, %xmm0, %xmm0
; │ @ REPL[32]:5 within `three_median3'
	retq
	nopw	%cs:(%rax,%rax)
; └

```

(the compiler decided that swapping the values is not worth saving one comparison, something like that).

But, at the end, Fortran is being able to optimize the `min/max` version to optimal, isn’t it? Does it behave like Julia in this regard (having NaN checks without `fastmath` and skipping those with it)?  
(I don’t know how to check the assembly codes in these cases).

---

<div class="post-metadata">

**Author:** ![ivanpribec](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/ivanpribec/32/3290_2.png) [@ivanpribec](https://fortran-lang.discourse.group/u/ivanpribec)\
**Post date:** [October 14, 2021, 4:54pm UTC](https://fortran-lang.discourse.group/t/code-slower-than-c-version/2066/25 "2021-10-14T16:54:18Z")

</div>

> [@lmiq](#):
>
> But, at the end, Fortran is being able to optimize the `min/max` version to optimal, isn’t it? Does it behave like Julia in this regard (having NaN checks without `fastmath` and skipping those with it)?  
> (I don’t know how to check the assembly codes in these cases).

You can check them with [Godbolt](https://godbolt.org/#g:!((g:!((g:!((h:codeEditor,i:(filename:'1',fontScale:14,fontUsePx:'0',j:1,lang:fortran,selection:(endColumn:21,endLineNumber:2,positionColumn:21,positionLineNumber:2,selectionStartColumn:21,selectionStartLineNumber:2,startColumn:21,startLineNumber:2),source:'function+three_median(x,+y,+z)%0A++real(8),+intent(in)+::+x,+y,+z%0A++real(8)+::+three_median%0A%0A++three_median+%3D+max(min(x,+y),min(max(x,+y),+z))%0A++%0Aend+function%0A%0Asubroutine+med_rv(n,+x,+res)%0A++integer(8)+::+n%0A++real(8)+::+x(n),+res++++%0A++real(8)+::+param1,+param2,+param3,+param4,+pi%0A++integer(8)+::+i%0A++real(8)+::+three_median%0A%0A++param1+%3D+2d0%0A++param2+%3D+3d0%0A++param3+%3D+4d0+%0A++param4+%3D+6d0%0A++pi+%3D+acos(-1d0)%0A%0A++!!omp+parallel+do+reduction(%2B:res)%0A++do+i+%3D+1,+n+-+2%0A++++res+%3D+res+%2B+three_median(abs(x(i)),+abs(x(i+%2B+1)),+abs(x(i+%2B+2)))**param1++%0A++end+do%0A++!!omp+end+parallel+do+%0A%0A++res+%3D+res*pi/(param4+-+param3*sqrt(param2)+%2B+pi)*n/(n+-+param1)%0Aend+subroutine'),l:'5',n:'0',o:'Fortran+source+%231',t:'0')),k:50,l:'4',n:'0',o:'',s:0,t:'0'),(g:!((g:!((h:compiler,i:(compiler:ifort202130,filters:(b:'0',binary:'1',commentOnly:'0',demangle:'0',directives:'0',execute:'1',intel:'0',libraryCode:'0',trim:'1'),flagsViewOpen:'1',fontScale:14,fontUsePx:'0',j:1,lang:fortran,libs:!(),options:'-O3++-qopt-report-phase%3Dvec+-qopt-report-file%3Dstderr',selection:(endColumn:16,endLineNumber:9,positionColumn:1,positionLineNumber:1,selectionStartColumn:16,selectionStartLineNumber:9,startColumn:1,startLineNumber:1),source:1,tree:'1'),l:'5',n:'0',o:'x86-64+ifort+2021.3.0+(Fortran,+Editor+%231,+Compiler+%231)',t:'0')),k:50,l:'4',m:82.20973782771536,n:'0',o:'',s:0,t:'0'),(g:!((h:output,i:(compiler:1,editor:1,fontScale:14,fontUsePx:'0',tree:'1',wrap:'1'),l:'5',n:'0',o:'Output+of+x86-64+ifort+2021.3.0+(Compiler+%231)',t:'0')),header:(),l:'4',m:17.790262172284642,n:'0',o:'',s:0,t:'0')),k:50,l:'3',n:'0',o:'',t:'0')),l:'2',n:'0',o:'',t:'0')),version:4):

Here is the latest Intel Fortran compiler Classic with `-O3`:

```plaintext
three_median_:
        movsd xmm2, QWORD PTR [rdi] #5.3
        movsd xmm1, QWORD PTR [rsi] #5.3
        movaps xmm0, xmm2 #7.1
        maxsd xmm2, xmm1 #7.1
        minsd xmm0, xmm1 #7.1
        minsd xmm2, QWORD PTR [rdx] #7.1
        maxsd xmm0, xmm2 #7.1
        ret   

```

And here is the `gfortran` generated assembly from the link in @lkedward’s reply:

```plaintext
three_median_:
        movsd xmm1, QWORD PTR [rdi]
        movsd xmm2, QWORD PTR [rsi]
        movapd xmm0, xmm1
        minsd xmm1, xmm2
        maxsd xmm0, xmm2
        minsd xmm0, QWORD PTR [rdx]
        maxsd xmm0, xmm1
        ret

```

---

<div class="post-metadata">

**Author:** ![lmiq](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/lmiq/32/555_2.png) [@lmiq](https://fortran-lang.discourse.group/u/lmiq)\
**Post date:** [October 14, 2021, 5:37pm UTC](https://fortran-lang.discourse.group/t/code-slower-than-c-version/2066/26 "2021-10-14T17:37:38Z")

</div>

Nice tool. The version with the conditionals only seems to generate less instructions: [Compiler Explorer](https://godbolt.org/z/TP8ePE7Yx)

With conditionals only probably there is a small performance gain, although as far as I understood from the documentation the `max/min` function is dealing with NaNs (I was expecting that `--fast-math` or some other flag made the assemblies converge, but I couldn’t find that flag, if it exists).

[Previous page](https://fortran-lang.discourse.group/t/code-slower-than-c-version/2066.md?page=1)
