# Nvfortran comparison of do concurrent vs OpenMP code

**URL:** <https://fortran-lang.discourse.group/t/nvfortran-comparison-of-do-concurrent-vs-openmp-code/8552>\
**Category:** Help\
**Created:** [September 3, 2024, 1:55am UTC](https://fortran-lang.discourse.group/t/nvfortran-comparison-of-do-concurrent-vs-openmp-code/8552 "2024-09-03T01:55:08Z")\
**Posts on this page:** 5\
**Page:** 2

<div class="post-metadata">

**Author:** ![hakostra](https://avatars.discourse-cdn.com/v4/letter/h/a3d4f5/32.png) [@hakostra](https://fortran-lang.discourse.group/u/hakostra)\
**Post date:** [September 6, 2024, 11:38am UTC](https://fortran-lang.discourse.group/t/nvfortran-comparison-of-do-concurrent-vs-openmp-code/8552/21 "2024-09-06T11:38:06Z")

</div>

I cleaned up the program a bit further and merged the OpenMP and DO CONCURRENT variants in one program to make it simpler to test different variants. This forum’s inline code is not too god so I put my version on Github Gits

> <https://gist.github.com/hakostra/ac5ed9279136fe1e7f6f217f2561f08e>

Then I executed this program with all the compilers I had access to on my workstation. My workstation is an Intel i9-13900k. This CPU has 8 P-cores (performance) and 16 E-cores (efficiency). The Operating system is linux Mint 21.2 (built on top of Ubuntu 22.04). The compilers I tested was:

- NAG `nagfor` 7.2 build 7214
- GNU `gfortran` 12.3
- Intel `ifx` 2024.2.1
- Intel `ifort` 2021.13.1
- Nvidia `nvfortran` 24.7
- LLVM `flang-new`, Github main branch commit 0c1500ef (yesterday)

The results became a quite big document, and I also uploaded this as a second Gist:

> <https://gist.github.com/hakostra/87a043e8436efb898b63b192638392d2>

The results have three columns: the first is the total time spent in the time loop, the second is the time spent in the first nested loop (`seca`) and the third is the time in the second nested loop (`secb`).

I will not do too much interpretations of the results here, but rather make some remarks:

- Only `nvfortran` can parallelize the DO CONCURRENT loops - nether of the other compilers make use of more than one thread for this variant. Therefore I would not use DO CONCURRENT in any program where I want to use threads.
- `gfortran` will not compile the DO CONCURRENT since it does not understand the locality specification ([bugzilla](https://gcc.gnu.org/bugzilla/show_bug.cgi?id=101602)).
- `nvfortran`, `ifx` and `ifort` is creating incredibly fast executables when not using threads/OpenMP. None of the other compilers are even close in single-thread performance.
- `ifort` creates the fastest non-threaded executable
- Most compilers create a slower running program with OpenMP (or DO CONCURRENT) when using that program for a single thread (`OMP_NUM_THREADS=1` or `ACC_NUM_CORES=1`)
- Nealy all compilers “converge” at the same runtime (except `ifort`), between 20 and 25 seconds for the entire time-loop, with enough cores. I think this means that I have reached a state where the CPU is no longer the bottleneck, but the memory transfer is.

---

<div class="post-metadata">

**Author:** ![Shahid](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/shahid/32/3672_2.png) [@Shahid](https://fortran-lang.discourse.group/u/Shahid)\
**Post date:** [September 6, 2024, 4:15pm UTC](https://fortran-lang.discourse.group/t/nvfortran-comparison-of-do-concurrent-vs-openmp-code/8552/22 "2024-09-06T16:15:23Z")

</div>

ifx did not produce correct results when i tested.

---

<div class="post-metadata">

**Author:** ![hakostra](https://avatars.discourse-cdn.com/v4/letter/h/a3d4f5/32.png) [@hakostra](https://fortran-lang.discourse.group/u/hakostra)\
**Post date:** [September 9, 2024, 6:46am UTC](https://fortran-lang.discourse.group/t/nvfortran-comparison-of-do-concurrent-vs-openmp-code/8552/23 "2024-09-09T06:46:37Z")

</div>

I checked the results of `ifx`, `gfortran` and `nvfortran`. The final value of the phi array is plotted below, for these three compilers with and without using OpenMP/DO CONCURRENT.

 ![plot](https://global.discourse-cdn.com/free1/uploads/fortran_lang/original/2X/7/7b4a136e3255db0b7fe70f3c20a5d0d66aa295eb.png)

The right column, with OpenMP/DO CONCURRENT, are all results produced running with 8 threads. To me these results seems to be identical. If there are differences I haven’t spotted, feel free to shout out.

---

<div class="post-metadata">

**Author:** ![Shahid](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/shahid/32/3672_2.png) [@Shahid](https://fortran-lang.discourse.group/u/Shahid)\
**Post date:** [September 9, 2024, 9:25am UTC](https://fortran-lang.discourse.group/t/nvfortran-comparison-of-do-concurrent-vs-openmp-code/8552/24 "2024-09-09T09:25:15Z")

</div>

I just used Windows 10 with ifx. The code does not run in parallel. Maybe, running on Linux shows parallel performance. I did not check it there!

---

<div class="post-metadata">

**Author:** ![Shahid](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/shahid/32/3672_2.png) [@Shahid](https://fortran-lang.discourse.group/u/Shahid)\
**Post date:** [September 9, 2024, 9:26am UTC](https://fortran-lang.discourse.group/t/nvfortran-comparison-of-do-concurrent-vs-openmp-code/8552/25 "2024-09-09T09:26:16Z")

</div>

A few months back, using ifx did not produce correct results. Now it just does not run in parallel on Windows 10.

[Previous page](https://fortran-lang.discourse.group/t/nvfortran-comparison-of-do-concurrent-vs-openmp-code/8552.md?page=1)
