# Fortran and MPI

**URL:** <https://fortran-lang.discourse.group/t/fortran-and-mpi/7288>\
**Category:** Advocacy\
**Created:** [January 28, 2024, 7:18pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288 "2024-01-28T19:18:30Z")\
**Posts on this page:** 13\
**Page:** 2

<div class="post-metadata">

**Author:** ![PierU](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/pieru/32/1848_2.png) [@PierU](https://fortran-lang.discourse.group/u/PierU)\
**Post date:** [January 31, 2024, 7:42am UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/21 "2024-01-31T07:42:54Z")

</div>

> [@JeffH](#):
>
> As for Fortran module standardization, that’s never going to happen. It has been discussed but it’s a nonstarter. It requires breaking backwards compatibility in every compiler

I wouldn’t bet anyway on binary compatibilty between different versions of a given compiler. e.g. I avoid using a module compiled with a version _n_ in sources compiled in a version _m /= n_.

---

<div class="post-metadata">

**Author:** ![VladimirF](https://avatars.discourse-cdn.com/v4/letter/v/e9a140/32.png) [@VladimirF](https://fortran-lang.discourse.group/u/VladimirF)\
**Post date:** [January 31, 2024, 12:22pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/22 "2024-01-31T12:22:19Z")

</div>

I use something similar in my main code. I try to have my code with computational logic without bare MPI calls and I use generic wrappers that look similar to coarray collective calls. Where point to point communication is needed, e.g. for halo regions exchange, it is in separate modules and still uses several layers of wrappers that make the calls easier.

---

<div class="post-metadata">

**Author:** ![JeffH](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jeffh/32/595_2.png) [@JeffH](https://fortran-lang.discourse.group/u/JeffH)\
**Post date:** [January 31, 2024, 1:58pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/23 "2024-01-31T13:58:00Z")

</div>

> [@RonShepard](#):
>
> > [@aerosayan](#):
> >
> > The time and manpower that went into creating coarrays, would’ve been far better utilized in improving GPU support for Fortran.
> 
> I think this is a false choice. These are orthogonal programming models, and both can be pursued independently.

First, CUDA was only introduced in 2007, by which time coarrays were well on their way to standardization in Fortran 2008, so there was no competition between coarrays and GPU support at the time. As we know, `DO CONCURRENT` was also introduced in Fortran 2008, but has been available as a GPU programming model since [2020](https://developer.nvidia.com/blog/accelerating-fortran-do-concurrent-with-gpus-and-the-nvidia-hpc-sdk/). More recent efforts to add features to the Fortran 2008 coarray model have not been in competition with anything GPU-related.

Second, while there is no intellectual competition between the two, as a practical matter, the committee is quite small and it is difficult for the relevant subcommittee to pursue too many features at once. Even if they (or we, since I am part of said subcommittee) could, there is also concern about adding too many features to Fortran, which already challenges implementers. For example, the BITS proposal was dropped from Fortran 2008 because the committee believed that it was already too much change for one release of the standard.

I hold weakly the opinion that coarrays were a mistake only because it so straightforward to implement that functionality as a library, and that the committee might have been able to make it possible to implement coarrays as a library more straightforward using other language changes, which would have been easier for compilers to adopt.

We can look to UPC++, which reimplements UPC as a C++ library, thanks to the differences between C and C++. I haven’t tried to implement PGAS as a library natively in Fortran in a minimal syntax, but I have more than 15 years of experience using Global Arrays in this context, and have not seen any compelling reason to use coarrays instead, and many practical reasons not to.

If someone wants to implement coarrays as a library, consider that one can use MPI-3 RMA and `C_F_POINTER` to get the equivalent of an allocatable coarray ([example](https://github.com/ParRes/Kernels/blob/default/FORTRAN/transpose-get-mpi.F90#L120). I made no attempt there to compress the syntax, but I imagine there is a way to create special types and use user-defined operators and/or type-bound procedures to creates something that is close enough to coarrays to solve real problems.

---

<div class="post-metadata">

**Author:** ![JeffH](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jeffh/32/595_2.png) [@JeffH](https://fortran-lang.discourse.group/u/JeffH)\
**Post date:** [January 31, 2024, 1:59pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/24 "2024-01-31T13:59:37Z")

</div>

> [@ivanpribec](#):
>
> This of course doesn’t help with the Fortran module part, but those are usually part of the Fortran compiler environment anyways. In Section 7 the authors mention the idea of a `mpi_f08_abi` which would be implemented on top of the standard-ABI version of MPI. If I understand correctly, this way you could switch MPI implementations just by loading different shared libraries into the enviroment.

The current plan is to not define a new module but to make it so that an `MPI_F08` implementation like VAPAA solves the problem we want to solve.

---

<div class="post-metadata">

**Author:** ![rwmsu](https://avatars.discourse-cdn.com/v4/letter/r/48db29/32.png) [@rwmsu](https://fortran-lang.discourse.group/u/rwmsu)\
**Post date:** [January 31, 2024, 2:08pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/25 "2024-01-31T14:08:22Z")

</div>

Jeff, I’m not advocating that compiler developers ditch their current formats. What I would like to see is a second option to generate a transportable module in a format all compilers can read generated by some compiler option (something like --enable iso\_standard\_modules). I find it hard to believe that the compiler development community can’t find some common ground on a format of some kind. There are a lot of options if people would just take the time to explore them. A markup language like implementation or just compiling to some intermediate internal representation come to mind. I’m sorry but the “can’t break backwards compatability” mantra is beginning to sound to me like a “dog ate my homework” excuse for not being willing to take the time to think outside the box. As long as the people who are tasked with defining what Fortran is are more focused on the past than the future (and to be frank trying to retain whatever competitive advantage they have by forcing lock in to their particular compiler) Fortran is doomed to extinction.

---

<div class="post-metadata">

**Author:** ![JeffH](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jeffh/32/595_2.png) [@JeffH](https://fortran-lang.discourse.group/u/JeffH)\
**Post date:** [January 31, 2024, 2:30pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/26 "2024-01-31T14:30:24Z")

</div>

> [@rwmsu](#):
>
> I’m sorry but the “can’t break backwards compatability” mantra is beginning to sound to me like a “dog ate my homework” excuse for not being willing to take the time to think outside the box. As long as the people who are tasked with defining what Fortran is are more focused on the past than the future (and to be frank trying to retain whatever competitive advantage they have by forcing lock in to their particular compiler) Fortran is doomed to extinction.

You seem to think this is a technical problem. Backwards-compatibility is not a technical problem for the implementers. The issue is users/customers don’t like it. You don’t have to convince 10 implementation teams. You need to convince 1000s of Fortran users, including the ones at the nuclear weapons labs and the commercial engineering firms, that they need to recompile 100% of their code when the change happens, for no observable benefit to them, since their code is already working and they have higher priorities that module ABI standardization.

There are famous examples showing why breaking backwards-compatibility causes problems, even when the goal is supposedly to make things better. See e.g. [https://bugzilla.redhat.com/show\_bug.cgi?id=638477](https://bugzilla.redhat.com/show_bug.cgi?id=638477).

---

<div class="post-metadata">

**Author:** ![Beliavsky](https://avatars.discourse-cdn.com/v4/letter/b/ba8739/32.png) [@Beliavsky](https://fortran-lang.discourse.group/u/Beliavsky)\
**Post date:** [January 31, 2024, 2:49pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/27 "2024-01-31T14:49:57Z")

</div>

A new arXiv preprint by some people at Intel and Argonne is

[Generating Bindings in MPICH](https://arxiv.org/abs/2401.16547)  
by Hui Zhou, Ken Raffenetti, Wesley Bland, Yanfei Guo

> The MPI Forum has recently adopted a Python scripting engine for generating the API text in the standard document. As a by-product, it made available reliable and rich descriptions of all MPI functions that are suited for scripting tools. Using these extracted API information, we developed a Python code generation toolbox to generate the language binding layers in MPICH. The toolbox replaces nearly 70,000 lines of manually maintained C and Fortran 2008 binding code with around 5,000 lines of Python scripts plus some simple configuration. In addition to completely eliminating code duplication in the binding layer and avoiding bugs from manual code copying , the code generation also minimizes the effort for API extension and code instrumentation. This is demonstrated in our implementation of MPI-4 large count functions and the prototyping of a next generation MPI profiling interface, QMPI.

---

<div class="post-metadata">

**Author:** ![rwmsu](https://avatars.discourse-cdn.com/v4/letter/r/48db29/32.png) [@rwmsu](https://fortran-lang.discourse.group/u/rwmsu)\
**Post date:** [January 31, 2024, 3:02pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/28 "2024-01-31T15:02:36Z")

</div>

And how is adding a new feature that didn’t exist and users are free use (or not use) at their own choosing breaking backwards compatability. As long as the compilers support their old format in addition to a new transportable one there is very little chance of “breaking” existing code. Its all about giving users a choice beyond “I have to support four different compilers” and waste unecessary time and money.

---

<div class="post-metadata">

**Author:** ![jeff](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/jeff/32/7201_2.png) [@jeff](https://fortran-lang.discourse.group/u/jeff)\
**Post date:** [January 31, 2024, 10:18pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/29 "2024-01-31T22:18:43Z")

</div>

> [@rwmsu](#):
>
> Other than NAG which uses their own shared memory implementation, all other compilers (except for Cray) use MPI as the transport layer so the only advantage co-arrays have is a slightly easier to use syntax.

I’m not sure about other compilers, but GNU Fortran can be directed to use an arbitrary library for a coarray implementation. OpenCoarrays using MPI just happens to be the only widely used implementation. Simply Fortran for Windows, as a counter-example, does _not_ use OpenCoarrays; instead, it ships with a [custom library that implements the GNU Fortran runtime library’s coarray API](https://simplyfortran.com/blog/6/) (that blog post is a bit old as the library now uses [named shared memory](https://learn.microsoft.com/en-us/windows/win32/memory/creating-named-shared-memory)). The whole system makes heavy use of a database shared between processes and mish-mash of Windows synchronization and memory mapping API calls to implement coarrays. It is admittedly not particularly fast when there are numerous intermittent transfers between images, but it can be performant if transfers are optimized.

The above paragraph reads like a plug, which I guess it is, but someone can just implement a functioning coarray library that doesn’t use MPI. It’s not impossible.

---

<div class="post-metadata">

**Author:** ![Federchen](https://avatars.discourse-cdn.com/v4/letter/f/ce7236/32.png) [@Federchen](https://fortran-lang.discourse.group/u/Federchen)\
**Post date:** [January 31, 2024, 11:41pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/30 "2024-01-31T23:41:30Z")

</div>

Scientists should never negate technologies or topics before they can properly understand and correctly explain them.

From _Modern Fortran explained_ , it’s introduction on coarrays, we have already learned about the distinction between _work distribution_ and _data distribution_ as a foundation of Coarray Fortran.

_Work distribution_ in Coarray Fortran adopts the SPMD programming model. Since Fortran 2018 the language allows the programmer to create multiple, even hierarchical (nested) SPMD environments. The SPMD model is a universal programming model to create parallelism for many types of devices, including CPU, GPU, FPGA, etc. The DPC++ book puts it this way:

> **[Expressing Parallelism](https://link.springer.com/chapter/10.1007/978-1-4842-9691-2_4)**
>
> Chapter 4 marks the transition from simple teaching examples toward real-world parallel code and expands upon details of the code samples we have casually shown in prior chapters. Chapter 4 explains the strengths and weaknesses of the different ways...

“One of the greatest strengths of a SPMD programming model is that it allows the same “program” to be mapped to multiple levels and types of parallelism, without any explicit direction from us. Instances of the same program could be pipelined, packed together and executed with SIMD instructions, distributed across multiple hardware threads, or a mix of all three.”

_Data distribution_ in Coarray Fortran can be expressed through (1) coarrays (since Fortran 2008) or through (2) collective subroutines (i.e. not necessarily using any coarrays, since Fortran 2018). Thus, we can also do coarray programming without even a single coarray declared in our codes. This should help to provide codes for a variety of devices in the future, including GPUs. My personal focus is currently on kernels utilizing coarrays but have also done a starting with another kernel type to utilize collective subroutines with them.  
Another important topic with coarrays is symmetric memory, that we already have with the Intel compilers.

When using coarrays for the data distribution, I am using coreRMA functionality yet not only to bulletproof check the (customized) synchronization but also to regularly check the (non-atomic) data transfer through the (theoretical) network in my programming. This may sound complicated but the coding size of individual parts can be very small (e.g. a simple basic version of a coarray-based channel implementation with non-blocking synchronization). Here, Fortran can make complicated things simple. Such codes are very easy to maintain. Main topics on top are, non-blocking synchronizations, parallel loops, pairwise independent forward progress for (not only spatial) kernels, device-wide synchronizations, hierarchical parallelism, and certainly other topics as well. With some aspects of the topics, my Fortran programming (preparation) may already be further than the DPC++ book yet is. As we are just entering the exascale era, also with new types of devices, this is certainly not the right time to negate (Coarray) Fortran.

---

<div class="post-metadata">

**Author:** ![rwmsu](https://avatars.discourse-cdn.com/v4/letter/r/48db29/32.png) [@rwmsu](https://fortran-lang.discourse.group/u/rwmsu)\
**Post date:** [February 1, 2024, 12:16am UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/31 "2024-02-01T00:16:39Z")

</div>

Well since I haven’t tried to run anything on a Windows system other than Office products since 1993, I’ll plead ignorance about anything targeted at Windows. At one time (with ifort not sure about ifx) you could supposedly try to replace Intel’s MPI with another distribution (at least I remember something on an Intel web site that implied you could). I tried it with openMPI but it didn’t work. I know there was an experimental implementation of openCoarrays based on openSHMEM but I don’t think it got very far. For small core (\< 64 core) workstations, using shared memory directly instead of MPI (which may or may not have been compiled to use shared memory instead of TCP/IP on a multi-core processor) make more sense. I’ll add Simply Fortran to my list of compilers supporting co-arrays. I also just found out that Fujitsu’s compiler also supports them but I’m not sure what the transport layer is based on.

---

<div class="post-metadata">

**Author:** ![VladimirF](https://avatars.discourse-cdn.com/v4/letter/v/e9a140/32.png) [@VladimirF](https://fortran-lang.discourse.group/u/VladimirF)\
**Post date:** [February 5, 2024, 1:51pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/32 "2024-02-05T13:51:12Z")

</div>

You certainly can use the Intel compiler with other MPI libraries and it is, in fact, quite commonly done on supercomputers and HPC clusters.

---

<div class="post-metadata">

**Author:** ![rwmsu](https://avatars.discourse-cdn.com/v4/letter/r/48db29/32.png) [@rwmsu](https://fortran-lang.discourse.group/u/rwmsu)\
**Post date:** [February 5, 2024, 2:06pm UTC](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288/33 "2024-02-05T14:06:47Z")

</div>

Yes, I’ve done that for MPI only runs for a couple of decades. The question though is can a non-Intel MPI implementation be used in place of Intel’s MPI as the transport layer for co-arrays. Again, I read a comment from an Intel person either here or on the Intel Fortran forum that implied you could. My one attempt at using openMPI instead of Intel’s MPI for co-arrays failed.

Edit. You still have to build any MPI distribution on any system with the Intel compilers if you want to use the mpi\_f08 module instead of mpif.h in your Fortran applications compiled with ifort or ifx.

[Previous page](https://fortran-lang.discourse.group/t/fortran-and-mpi/7288.md?page=1)
