# DO CONCURRENT: compiler flags to enable parallelization

**URL:** <https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300>\
**Category:** Help\
**Created:** [September 12, 2022, 9:09pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300 "2022-09-12T21:09:52Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![HugoMVale](https://avatars.discourse-cdn.com/v4/letter/h/e19adc/32.png) [@HugoMVale](https://fortran-lang.discourse.group/u/HugoMVale)\
**Post date:** [September 12, 2022, 9:09pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/1 "2022-09-12T21:09:52Z")

</div>

If I understand correctly, a `do concurrent` construct does not necessarily imply that the code inside the block will run in parallel, because (for instance) the compiler might estimate that the compute task does not justify the overhead of parallelization.  
On the other hand, I have doubts about what must be done to _allow_ the compiler to consider a possible parallelization. More specifically, my questions are:

1. Is it correct that parallelization of `do` and `do concurrent` loops is _deactivated by default_ unless a specific compiler flag is used?
2. With `ifort`, according to [this page](https://www.intel.com/content/www/us/en/developer/articles/technical/automatic-parallelization-with-intel-compilers.html), it seems that parallelization of `do concurrent` requires compilation with `--parallel` or `-qopenmp`. In this manner, if the compute work justifies it, it will be (automatically) distributed among the number of available threads at runtime. Is this correct?
3. With `gfortran`, according to [this paper](https://arxiv.org/pdf/2110.10151.pdf), parallelization of `do concurrent` requires compilation with `-ftree-parallelize-loops=N`, meaning that N at runtime is fixed by the value chosen at compile time. Is this correct?

What is your opinion and experience regarding this matter?

---

<div class="post-metadata">

**Author:** ![pcosta](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/pcosta/32/5504_2.png) [@pcosta](https://fortran-lang.discourse.group/u/pcosta)\
**Post date:** [September 12, 2022, 9:20pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/2 "2022-09-12T21:20:31Z")

</div>

`nvfortran` also supports parallelization on CPUs / GPU offloading GPU offloading using `DO CONCURRENT` (they even implemented reduction clauses from the upcoming `2023` standard): [Using Fortran Standard Parallel Programming for GPU Acceleration | NVIDIA Technical Blog](https://developer.nvidia.com/blog/using-fortran-standard-parallel-programming-for-gpu-acceleration).

---

<div class="post-metadata">

**Author:** ![HugoMVale](https://avatars.discourse-cdn.com/v4/letter/h/e19adc/32.png) [@HugoMVale](https://fortran-lang.discourse.group/u/HugoMVale)\
**Post date:** [September 13, 2022, 7:40pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/4 "2022-09-13T19:40:24Z")

</div>

Thanks for the hint. Yes, the paper that I cited in my first post has a detailed comparison of `ifort`, `nvfortran` and `gfortran`, and NVIDIA’s compiler does indeed do a good job at parallelizing `do concurrent` constructs. By default, I use `gfortran`, so (more or less implicitly) I am looking for the appropriate flags for this compiler.

---

<div class="post-metadata">

**Author:** ![HugoMVale](https://avatars.discourse-cdn.com/v4/letter/h/e19adc/32.png) [@HugoMVale](https://fortran-lang.discourse.group/u/HugoMVale)\
**Post date:** [September 13, 2022, 8:02pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/5 "2022-09-13T20:02:54Z")

</div>

Thanks, very helpfull suggestion. I have just started playing with the `-fopt-info` flag. 🙂

---

<div class="post-metadata">

**Author:** ![ivanpribec](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/ivanpribec/32/3290_2.png) [@ivanpribec](https://fortran-lang.discourse.group/u/ivanpribec)\
**Post date:** [January 3, 2024, 12:53am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/6 "2024-01-03T00:53:27Z")

</div>

I summarized some do concurrent related information in [this thread](https://fortran-lang.discourse.group/t/simplify-loop-on-an-array-of-derived-type/7095/9) and thought it is worth reposting here:

## Multi-threaded do concurrent (CPU)

| Compiler | Parallel flag | Information | Number of threads | Underlying implementation |
| --- | --- | --- | --- | --- |
| `gfortran` | `-ftree-parallelize-loops=n` | `-fopt-info-loop` | using the parallel flag | OpenMP/pthreads |
| `nvfortran` | `-stdpar=multicore` | `-Minfo=stdpar,accel` | `ACC_NUM_CORES` | OpenACC |
| `ifort` (deprecated) | `-parallel` | `-qopt-report -qopt-report-phase=par` | `OMP_NUM_THREADS`, `-par-num-threads=n` | OpenMP |
| `ifx` | `-qopenmp` | `-qopt-report` | `OMP_NUM_THREADS` | OpenMP |
| CCE `ftn` (Cray/HPE) | `-h thread_do_concurrent` | ? | ? | ? |
| AMD `flang` | `-fopenmp` | ? | `OMP_NUM_THREADS` | OpenMP |

The OpenMP [environment variables](https://www.openmp.org/spec-html/5.0/openmpch6.html) can also be used to control processor affinity. This is also the case for nvfortran, which responds to `OMP_PROC_BIND` and `OMP_PLACES`, because OpenACC doesn’t have variables for thread-to-core binding.

## Resources

- [Number of threads in `do concurrent` loops | NVIDIA](https://forums.developer.nvidia.com/t/number-of-threads-in-do-concurrent-loops/252639)
- [Accelerating Fortran DO CONCURRENT with GPUs and the NVIDIA HPC SDK | NVIDIA](https://developer.nvidia.com/blog/accelerating-fortran-do-concurrent-with-gpus-and-the-nvidia-hpc-sdk/)
- [Does gfortran take advantage of DO CONCURRENT? | Stack Overflow](https://stackoverflow.com/questions/29928293/does-gfortran-take-advantage-of-do-concurrent)
- [When should I use DO CONCURRENT and when OpenMP? | Stack Overflow](https://stackoverflow.com/questions/38549666/when-should-i-use-do-concurrent-and-when-openmp?noredirect=1&lq=1)
- [Using Fortran DO CONCURRENT for Accelerator Offload | Intel](https://www.intel.com/content/www/us/en/developer/articles/technical/using-fortran-do-current-for-accelerator-offload.html#gs.2h91al)
- [The Case for OpenMP\* Target Offloading: Why ISO Fortran Is Not Enough for Heterogeneous Computing | Intel](https://www.intel.com/content/www/us/en/developer/articles/technical/the-case-for-openmp-target-offloading.html#gs.2ha5eo)
- [Transition to the Intel (R) Fortran Compiler | Intel](https://www.openmp.org/wp-content/uploads/IFX_update_apr2023.pdf)
- [DO CONCURRENT isn’t necessarily concurrent | LLVM (flang)](https://flang.llvm.org/docs/DoConcurrent.html)
- [Benchmarking Fortran DO CONCURRENT on CPUs and GPUs Using BabelStream | SC22](https://www.dcs.warwick.ac.uk/pmbs/pmbs22/PMBS/talk12.pdf)
- [Can Fortran’s `do concurrent’ replace directives for accelerated computing? | SC21](https://waccpd.org/wp-content/uploads/2021/11/ws_waccpd102s3-file5.pdf)
- [Clarification on DO CONCURRENT | Fortran Discourse](https://fortran-lang.discourse.group/t/clarification-on-do-concurrent/1647/29)

PS: I’ve made this a Wiki post, so feel free to add missing information.

---

<div class="post-metadata">

**Author:** ![Shahid](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/shahid/32/3672_2.png) [@Shahid](https://fortran-lang.discourse.group/u/Shahid)\
**Post date:** [January 3, 2024, 6:53am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/7 "2024-01-03T06:53:42Z")

</div>

**1.** The do concurrent has nothing to do with parallel flags for intel.

It uses OpenMP under the hood and the compiler flag is always

> /Qopenmp

for windows for both **ifort** and **ifx**.

**2.** I think it is same for gfortran and **-fopenmp** compiler flag is used for do concurrent.

**3**.A few months ago i asked the same question in the intel community. The moderator reply was

```auto
Why are you using /Qparallel? That turns on the auto-parallelizer. I'm not sure what that does if anything with DO CONCURRENT.

As I just posted on another thread the DO CONCURRENT / openmp combination uses OMP SIMD. 

```

---

<div class="post-metadata">

**Author:** ![ivanpribec](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/ivanpribec/32/3290_2.png) [@ivanpribec](https://fortran-lang.discourse.group/u/ivanpribec)\
**Post date:** [January 3, 2024, 8:30am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/8 "2024-01-03T08:30:53Z")

</div>

> [@Shahid](#):
>
> **1.** The do concurrent has nothing to do with parallel flags for intel.

Concerning the Intel Fortran Compiler Classic (`ifort`), [this Intel thread](http://community.intel.com/t5/Intel-Fortran-Compiler/Questions-about-DO-CONCURRENT/m-p/1028988#M110081) from 2015 stated:

> DO CONCURRENT allows the compiler to ignore any potential dependencies between iterations and to execute the loop in parallel. This can mean either SIMD parallelism (vectorization), which is enabled by default, or thread parallelism (auto-parallelization), which is enabled only by /Qparallel. This is independent of /Qopenmp, which does not enable auto-parallelization, it only enables parallelism through OpenMP directives. However, auto-parallelization with /Qparallel uses the same underlying OpenMP runtime library as /Qopenmp. The overhead for setting up and entering a parallel region is typically thousands of clock cycles, so auto-parallelization is usually worthwhile only for loops with a sufficiently large amount of work to amortize this overhead.

And in [this Intel thread](http://community.intel.com/t5/Intel-Fortran-Compiler/Advantage-DO-CONCURRENT-against-DO/m-p/964190#M95373), @sblionel stated:

> DO CONCURRENT does not “demand parallel” - it allows/requests it. As others have said, the semantics of DO CONCURRENT make it more likely that the loop can be parallelized correctly. If you’re not enabling auto-parallel, there is no benefit to DO CONCURRENT.

With the new Intel LLVM compiler (`ifx`), this has changed, again [quoting](https://fortran-lang.discourse.group/t/clarification-on-do-concurrent/1647/29) @sblionel:

> Just as a followup to my March 2022 reply, Intel’s LLVM-based ifx compiler does not support -parallel at all. It **will** (attempt to) parallelize `DO CONCURRENT` if you enable OpenMP, even if you don’t use OpenMP otherwise.

* * *

> [@Shahid](#):
>
> **2.** I think it is same for gfortran and **-fopenmp** compiler flag is used for do concurrent.

I have verified that the `-fopenmp` flag is not needed and inspected the compiler reports to verify parallelization occurs. The executable produced on Linux has a dependency on OpenMP (GOMP) and pthreads, as stated by GCC documentation for [`-ftree-parallelize-loops`](https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html#index-ftree-parallelize-loops):

> This option implies -pthread, and thus is only supported on targets that have support for -pthread.

* * *

> [@Shahid](#):
>
> **3.** A few months ago i asked the same question in the intel community. The moderator reply was
> 
> > Why are you using /Qparallel? That turns on the auto-parallelizer. […]

I’m guessing they were referring to the new Intel LLVM compiler, as ifort was “end-of-life” already.

What is worth noting is that in both ifort and gfortran, the respective parallel flags also work on regular do loops, if the compiler heuristic determines this would be profitable. Using `do concurrent` instead of `do` is about **intent** , and letting the compiler know the loop can be executed concurrently, meaning there are no data dependencies, and it can be safely parallelized.

The flang documentation captured this well when it says,

> The best option seems to be the one that assumes that users who write `DO CONCURRENT` constructs are doing so with the intent to write parallel code.

on the topic of “how to convey to a compiler that a loop is safely parallelizable”

---

<div class="post-metadata">

**Author:** ![Shahid](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/shahid/32/3672_2.png) [@Shahid](https://fortran-lang.discourse.group/u/Shahid)\
**Post date:** [January 3, 2024, 2:46pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/9 "2024-01-03T14:46:58Z")

</div>

This is very interesting.

I will check my codes again and possibly get back.

---

<div class="post-metadata">

**Author:** ![jorgeg](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jorgeg/32/6835_2.png) [@jorgeg](https://fortran-lang.discourse.group/u/jorgeg)\
**Post date:** [October 15, 2025, 12:28am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/10 "2025-10-15T00:28:58Z")

</div>

I just verified on amdflang for AMDGPUs that you can use `-fopenmp --offload-arch=gfx90a -fdo-concurrent-to-openmp=device` to offload do concurrent to devices!

---

<div class="post-metadata">

**Author:** ![RJaBi](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/rjabi/32/5161_2.png) [@RJaBi](https://fortran-lang.discourse.group/u/RJaBi)\
**Post date:** [October 15, 2025, 9:04am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/11 "2025-10-15T09:04:28Z")

</div>

When you say amdflang, what exact compiler are you talking about? The [AMD next gen fortran compiler](https://github.com/amd/InfinityHub-CI/blob/main/fortran/README.md)? If so, which ‘drop’. I tried a pre-release version of drop 6.0.0 back in April which seriously struggled with openMP offload (even a basic reduction). I was too scared to try do-concurrent at the time.

---

<div class="post-metadata">

**Author:** ![jorgeg](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jorgeg/32/6835_2.png) [@jorgeg](https://fortran-lang.discourse.group/u/jorgeg)\
**Post date:** [October 15, 2025, 1:01pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/12 "2025-10-15T13:01:46Z")

</div>

[https://repo.radeon.com/rocm/misc/flang/rocm-afar-8248-drop-7.0.0-sles.tar.bz2](https://repo.radeon.com/rocm/misc/flang/rocm-afar-8248-drop-7.0.0-sles.tar.bz2)

This drop

---

<div class="post-metadata">

**Author:** ![mklemm](https://avatars.discourse-cdn.com/v4/letter/m/e95f7d/32.png) [@mklemm](https://fortran-lang.discourse.group/u/mklemm)\
**Post date:** [October 17, 2025, 5:28pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/13 "2025-10-17T17:28:13Z")

</div>

We have added first support this in that drop. Please note, that not everything is working yet, e.g., locality specifiers might not be supported yet. Also, polymorphic types cannot be used, because the OpenMP API does not yet support to transfer and use them on a target device.

If you’re playing with and you find a bug, please send a direct message to me and we will have a look.

PS: It also supports to use host threads.

---

<div class="post-metadata">

**Author:** ![jorgeg](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jorgeg/32/6835_2.png) [@jorgeg](https://fortran-lang.discourse.group/u/jorgeg)\
**Post date:** [October 19, 2025, 10:15am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/14 "2025-10-19T10:15:43Z")

</div>

so basically as long as I’m not using polymorphism and I rely on plain old vanilla data I’ll be fine?

I’ll definitely test it more out throughout the week and I’ll report any bugs ! thanks for this

---

<div class="post-metadata">

**Author:** ![mklemm](https://avatars.discourse-cdn.com/v4/letter/m/e95f7d/32.png) [@mklemm](https://fortran-lang.discourse.group/u/mklemm)\
**Post date:** [October 25, 2025, 2:49pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/15 "2025-10-25T14:49:26Z")

</div>

Yes. The thing is because `DO CONCURRENT` is translated to `!$omp target teams distribute parallel do` under the hood, Flang is bound to what the OpenMP API supports. We are currently looking at this from two angles. First, extend Flang to allow more than what the OpenMP API permits, so that we have more flexibility to do code-gen. Second, lift the restriction in the OpenMP API.

One thing that you might have to do: If you are using function calls in the `DO CONCURRENT` construct that are defined outside of the current source file, you will have to manually tag them with `!$omp declare target` so there’s a linkable symbol and binary code for the GPU.

---

<div class="post-metadata">

**Author:** ![jorgeg](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jorgeg/32/6835_2.png) [@jorgeg](https://fortran-lang.discourse.group/u/jorgeg)\
**Post date:** [November 7, 2025, 12:16am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/16 "2025-11-07T00:16:28Z")

</div>

An interesting question to ask here, I have this code: [learning\_tools/fortran/benchmarks/validate.f90 at main · JorgeG94/learning\_tools · GitHub](https://github.com/JorgeG94/learning_tools/blob/main/fortran/benchmarks/validate.f90)

Which you can compile on AMD GPUs using: `amdflang -fopenmp --offload-arch=gfx90a -fdo-concurrent-to-openmp=device -O3 validate.f90`

And then run it with: `OMPX_FORCE_SYNC_REGIONS=1 ./a.out 1024 1024` you’ll run a sweep over a benchmark that should produce something like this:

```auto
 Nz,vertical->i->j,i->j->vertical,j->vertical->i,vertical->j->i,j->i->vertical
   10, 0.016757, 0.007373, 0.022108, 0.023841, 0.000463
   25, 0.047945, 0.021695, 0.062656, 0.068325, 0.001239
   50, 0.100081, 0.042457, 0.130965, 0.142799, 0.002703
  100, 0.203929, 0.091109, 0.274001, 0.292300, 0.005823
  200, 0.412056, 0.200474, 0.548268, 0.588676, 0.012267
  400, 0.827845, 0.394302, 1.115088, 1.183742, 0.025872

```

These are timings to evaluate a loop using multiple orderings to traverse the loop. On an MI250X I can measure the FLOPs:

```auto
 vertical->i->j elapsed: 0.8278 s
 vertical->i->j flop rate: 7.5618 GFLOP/s
 i->j->vertical elapsed: 0.3943 s
 i->j->vertical flop rate: 15.8762 GFLOP/s
 j->vertical->i elapsed: 1.1151 s
 j->vertical->i flop rate: 5.6139 GFLOP/s
 vertical->j->i elapsed: 1.1837 s
 vertical->j->i flop rate: 5.2883 GFLOP/s
 j->i->vertical elapsed: 0.0259 s
 j->i->vertical flop rate: 241.9604 GFLOP/s

```

So you can see that there’s a severe imbalance of FLOP rates and timings, whereas on a V100:

```auto
 vertical->i->j elapsed: 0.0635 s
 vertical->i->j flop rate: 98.5513 GFLOP/s
 i->j->vertical elapsed: 0.0281 s
 i->j->vertical flop rate: 222.3964 GFLOP/s
 j->vertical->i elapsed: 0.0554 s
 j->vertical->i flop rate: 112.9883 GFLOP/s
 vertical->j->i elapsed: 0.0631 s
 vertical->j->i flop rate: 99.2611 GFLOP/s
 j->i->vertical elapsed: 0.0283 s
 j->i->vertical flop rate: 220.8568 GFLOP/s

```

I wonder if there’s a problem with how the compiler is mapping the do concurrent to the threads and the GPU…

The code doesn’t do any physics or anything, just some operations that are a similar to a loop I have in a much larger application.

---

<div class="post-metadata">

**Author:** ![mklemm](https://avatars.discourse-cdn.com/v4/letter/m/e95f7d/32.png) [@mklemm](https://fortran-lang.discourse.group/u/mklemm)\
**Post date:** [November 7, 2025, 8:06am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/17 "2025-11-07T08:06:11Z")

</div>

Thanks for testing the implementation. This is still WIP, so there might indeed be something odd going on.

We will take a look at the code and see what the compiler does with it. Stay tuned!

---

<div class="post-metadata">

**Author:** ![jorgeg](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jorgeg/32/6835_2.png) [@jorgeg](https://fortran-lang.discourse.group/u/jorgeg)\
**Post date:** [November 7, 2025, 8:44am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/18 "2025-11-07T08:44:44Z")

</div>

Thanks so much. I had a plan to rewrite the do concurrent into openmp to see if it is a translation problem from do-concurrent to openmp. But alas, it is 19:44 in Canberra on a Friday 🙂

---

<div class="post-metadata">

**Author:** ![rouson](https://avatars.discourse-cdn.com/v4/letter/r/ce7236/32.png) [@rouson](https://fortran-lang.discourse.group/u/rouson)\
**Post date:** [November 8, 2025, 4:22am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/19 "2025-11-08T04:22:17Z")

</div>

@jorgeg although `flang` can’t yet automatically parallelize `do concurrent` constructs that leverage dynamic polymorphism, there’s one exception: it’s ok to invoke `non_overridable` type-bound procedures with a polymorphic passed-object dummy argument. We demonstrated this for batch inference on deep-neural networks in a recent paper on which @mklemm was a co-author. This works with LLVM `flang` 21, which also supports all locality specifiers.

---

<div class="post-metadata">

**Author:** ![jorgeg](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jorgeg/32/6835_2.png) [@jorgeg](https://fortran-lang.discourse.group/u/jorgeg)\
**Post date:** [November 10, 2025, 5:30am UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/20 "2025-11-10T05:30:32Z")

</div>

on a separate (yet similar) note…has anyone gotten gfortran to parallelize do-concurrents? My execution time seems to always be the same…

---

<div class="post-metadata">

**Author:** ![cmaapic](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/cmaapic/32/659_2.png) [@cmaapic](https://fortran-lang.discourse.group/u/cmaapic)\
**Post date:** [November 10, 2025, 4:59pm UTC](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300/22 "2025-11-10T16:59:30Z")

</div>

Here are some timings runs I did eariler this year comparing explicit do loops, whole array syntax, do concurrent and openmp.

| | Nag | Nag | | Intel | Intel | | Intel | Intel |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| | nagfor | nagfor | | ifort | ifort | | ifx | ifx |
| | Windows | Linux | | Windows | Linux | | Windows | Linux |
| | 7.2-7225 | 7.2-7225 | | 2021.13.0 | 2021.13.1 | | 2025.1.0 | 2025.0.4 |
| | | | | | | | | |
| Whole array | 4.289280 | 3.238440 | | 1.458056 | 1.458056 | | 1.862359 | 1.862359 |
| Do loop | 2.088499 | 1.872361 | | 1.862638 | 1.862638 | | 1.861140 | 1.861140 |
| Do concurrent | 1.847248 | 1.872329 | | 0.409299 | 0.409299 | | 0.498400 | 0.498400 |
| openmp | 0.499684 | 0.487578 | | 0.408716 | 0.408716 | | 0.498171 | 0.498171 |
| | | | | | | | | |
| | | | | | | | | |
| | gfortran | gfortran | | nvidia | | | amd | |
| | gfortran | gfortran | | nvfortran | | | flang | |
| | Windows | Linux | | Linux | | | Linux | |
| | 14.2.0 | 14.2.1 | | 24.9 | | | 5.0.0 | |
| | | | | | | | | |
| Whole array | 1.950938 | 1.950938 | | 1.852759 | | | 1.859273 | |
| Do loop | 1.950196 | 1.950196 | | 1.854272 | | | 1.860085 | |
| Do concurrent | 1.872747 | 1.872747 | | 1.856842 | | | 1.860175 | |
| openmp | 0.498283 | 0.498283 | | 0.500433 | | | 0.498708 | |

[Next page](https://fortran-lang.discourse.group/t/do-concurrent-compiler-flags-to-enable-parallelization/4300.md?page=2)
