# Fortran coarrays on GPU

**URL:** <https://fortran-lang.discourse.group/t/fortran-coarrays-on-gpu/3978>\
**Category:** Uncategorized\
**Created:** [July 11, 2022, 7:36pm UTC](https://fortran-lang.discourse.group/t/fortran-coarrays-on-gpu/3978 "2022-07-11T19:36:55Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![korobkin](https://avatars.discourse-cdn.com/v4/letter/k/e8c25b/32.png) [@korobkin](https://fortran-lang.discourse.group/u/korobkin)\
**Post date:** [July 11, 2022, 7:36pm UTC](https://fortran-lang.discourse.group/t/fortran-coarrays-on-gpu/3978/1 "2022-07-11T19:36:55Z")

</div>

Is it possible to run Fortran coarrays on GPUs? There were some mentions using GASnet, is it successful or advancing in any way?  
Overall, are Fortran coarrays an actively pursued direction? What happened to coarrays v.2, developed at Rice University?

---

<div class="post-metadata">

**Author:** ![shahmoradi](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/shahmoradi/32/3151_2.png) [@shahmoradi](https://fortran-lang.discourse.group/u/shahmoradi)\
**Post date:** [July 13, 2022, 7:35pm UTC](https://fortran-lang.discourse.group/t/fortran-coarrays-on-gpu/3978/2 "2022-07-13T19:35:44Z")

</div>

I think the NVIDIA Fortran compiler team is working on it. My knowledge of their work is old and dates back to 2018 at the SC conference. There, they told me they were working on GPU-parallelization of Fortran syntax like `do concurrent`, and they did release this feature in a subsequent release of their Fortran compiler. But I do not think they have a working Coarray GPU implementation yet. That is just my guess and it may be wrong. Maybe this year, I will stop by their booth again to inquire. The developers of the [OpenCoarrays](https://github.com/sourceryinstitute/OpenCoarrays) may have a much better answer to your question if you ask them on GitHub.

---

<div class="post-metadata">

**Author:** ![wyphan](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/wyphan/32/632_2.png) [@wyphan](https://fortran-lang.discourse.group/u/wyphan)\
**Post date:** [July 14, 2022, 4:12pm UTC](https://fortran-lang.discourse.group/t/fortran-coarrays-on-gpu/3978/3 "2022-07-14T16:12:08Z")

</div>

AFAIK NVIDIA Fortran compiler is based on classic Flang, which doesn’t support Fortran coarrays at all, so it’s unlikely this combination of coarrays + `do concurrent` on GPU will happen soon. Also, I think @rouson is currently in the process of reimplementing OpenCoarrays as [Caffeine](https://github.com/BerkeleyLab/Caffeine).

---

<div class="post-metadata">

**Author:** ![ivanpribec](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/ivanpribec/32/3290_2.png) [@ivanpribec](https://fortran-lang.discourse.group/u/ivanpribec)\
**Post date:** [July 14, 2022, 9:57pm UTC](https://fortran-lang.discourse.group/t/fortran-coarrays-on-gpu/3978/4 "2022-07-14T21:57:31Z")

</div>

I think NVIDIA has some of the components which would be needed in a co-array implementation. See for example:

- [NVSHMEM | NVIDIA](https://developer.nvidia.com/nvshmem)
- [NVIDIA Collective Communications Library (NCCL) | NVIDIA Developer](https://developer.nvidia.com/nccl)

The libraries were mentioned in a recent training at ISC22: [GitHub - FZJ-JSC/tutorial-multi-gpu: Efficient Distributed GPU Programming for Exascale, an SC/ISC Tutorial](https://github.com/FZJ-JSC/tutorial-multi-gpu). Unfortunately, I don’t know if coarrays are on their radar. They also have their [CUDA unified memory](https://developer.nvidia.com/blog/unified-memory-cuda-beginners/), which seems like a global address space to me. It’s just the partitioning part which is missing.

It would be good to know if Intel plans to support hybrid coarrays + OpenMP target offloading in their new LLVM compiler. After-all [Aurora](https://en.wikipedia.org/wiki/Aurora_(supercomputer)) is supposed to be on its way.
