# Fortran Arrays in C++

**URL:** <https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856>\
**Category:** Announcements\
**Created:** [November 19, 2024, 5:21pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856 "2024-11-19T17:21:41Z")\
**Posts on this page:** 11\
**Page:** 2

<div class="post-metadata">

**Author:** ![Beliavsky](https://avatars.discourse-cdn.com/v4/letter/b/ba8739/32.png) [@Beliavsky](https://fortran-lang.discourse.group/u/Beliavsky)\
**Post date:** [November 28, 2024, 12:08pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/21 "2024-11-28T12:08:02Z")

</div>

> [@gronki](#):
>
> Encouraged by this article, I wrote my own very simple wrapper for arrays in C++. My convolution benchmark run 15% **faster** in C++ than in Fortran, despite using objects with overloaded `()` operator to emulate arrays. Of course, this got all nicely inlined using a modern compiler and vectorized efficiently. One more nail in the coffin I guess 🙂

I think the codes are in

> **[GitHub - gronki/fastconv](https://github.com/gronki/fastconv)**
>
> Contribute to gronki/fastconv development by creating an account on GitHub.

and

> **[GitHub - gronki/fastconv-cpp](https://github.com/gronki/fastconv-cpp)**
>
> Contribute to gronki/fastconv-cpp development by creating an account on GitHub.

---

<div class="post-metadata">

**Author:** ![PierU](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/pieru/32/1848_2.png) [@PierU](https://fortran-lang.discourse.group/u/PierU)\
**Post date:** [November 28, 2024, 2:09pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/22 "2024-11-28T14:09:08Z")

</div>

> [@Beliavsky](#):
>
> I think the codes are in
> 
> [GitHub - gronki/fastconv](https://github.com/gronki/fastconv)

What is the need of such a convoluted (pun intended 🙂 ) code to perform something as simple as a convolution? What about the OOP runtime overheads?

Here is my code for 1D convolution:

```fortran
! *******************************************************************************
subroutine sconv1D &
          (d,dfirst,dlast, &
           e,efirst,elast, &
           x,xfirst,xlast)
! *******************************************************************************! x = x + d * e
! *******************************************************************************
implicit none

integer, intent(in) :: dfirst, dlast
integer, intent(in) :: efirst, elast
integer, intent(in) :: xfirst, xlast
real, intent(in) :: d(dfirst:dlast), e(efirst:elast)
real, intent(inout) :: x(xfirst:xlast)

integer :: id, ixmin, ixmax

do id = dfirst, dlast
   ixmin = max( xfirst, efirst+id)
   ixmax = min( xlast, elast +id)
   x(ixmin:ixmax) = x(ixmin:ixmax) + d(id)*e(ixmin-id:ixmax-id)
end do

end subroutine sconv1D

```

KISS… no OOP, explicit shape arguments because they have less overheads than assumed shape… The 2D and 3D versions are essentially the same, just with more nested loops.

---

<div class="post-metadata">

**Author:** ![ivanpribec](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/ivanpribec/32/3290_2.png) [@ivanpribec](https://fortran-lang.discourse.group/u/ivanpribec)\
**Post date:** [November 28, 2024, 2:19pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/23 "2024-11-28T14:19:36Z")

</div>

> [@PierU](#):
>
> ```fortran
> x(ixmin:ixmax) = x(ixmin:ixmax) + d(id)*e(ixmin-id:ixmax-id)
> 
> ```

Not to mention this (auto-)vectorizes nicely. Say with `gfortran -O3 -march=skylake-avx512 -mprefer-vector-width=512`, the bulk of the work gets done in the hot loop:

```txt
.L5:
        vmovups zmm0, ZMMWORD PTR [r15+rax]
        vfmadd213ps zmm0, zmm1, ZMMWORD PTR [rdx+rax]
        vmovups ZMMWORD PTR [rdx+rax], zmm0
        add rax, 64
        cmp rdi, rax
        jne .L5

```

---

<div class="post-metadata">

**Author:** ![fedebenelli](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/fedebenelli/32/1823_2.png) [@fedebenelli](https://fortran-lang.discourse.group/u/fedebenelli)\
**Post date:** [November 28, 2024, 2:27pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/24 "2024-11-28T14:27:00Z")

</div>

I’m just getting aware that there is an overhead when using explicit shape instead of assumed shape, I always thought that it did not matter. Does that overhead becomes more noticeable as the arrays get larger?  
All my codes are using assumed shape, with pretty small arrays (30x30 is a relatively huge dimension for my cases). But the routines are called millions of times so there might be a room for improvement by using explicit shape? I might do some tests later

---

<div class="post-metadata">

**Author:** ![certik](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/certik/32/4_2.png) [@certik](https://fortran-lang.discourse.group/u/certik)\
**Post date:** [November 28, 2024, 3:08pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/25 "2024-11-28T15:08:30Z")

</div>

> [@Brady](#):
>
> think that is the thing, as jkd2022 says there is more to Fortran than pure speed.

Yes, the main advantage of Fortran is that it’s easy for a domain expert (say a physicist) to write fast code.

That being said, Fortran should absolutely give the best performance. I looked at @gronki’s examples, but they are layers and layers of OOP, so it would be a long investigation why the Fortran version is slower, which I don’t have time for right now (you should compare gfortran with the same version of gcc, and clang with the same version of flang, etc.). I recommend writing Fortran code like @PierU did above (or how @jkd2022 recommends). @PierU’s convolution code should look simpler than the correspodning C++ code, and it should be as fast or faster — otherwise the Fortran compilers must improve.

---

<div class="post-metadata">

**Author:** ![gronki](https://avatars.discourse-cdn.com/v4/letter/g/a88e4f/32.png) [@gronki](https://fortran-lang.discourse.group/u/gronki)\
**Post date:** [November 28, 2024, 3:50pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/26 "2024-11-28T15:50:10Z")

</div>

In my version, `kernel` is an `allocatable`, therefore array is guaranteed to be contiguous. Actually, it can be guaranteed by `contiguous` attribute that I use in my procedural part of the code. Anyway, most compilers nowadays do generate contigous/non-contigous branches. So I disagree that OOP introduces any overhead here while providing much cleaner interface. (From my tests, there was not much difference caused by that, but I should provide such tests for comparison on the repo). Subroutine with 10 arguments feels like 70’s coding, but it must be a matter of preference, since I know many people who hate OOP interfaces! I think the art is to use them in non-critical parts of the code (configuring the computation) and stick to procedural in the critical parts (where we perform the computation).

To be fair, Fortran version is performing faster with `gcc/gfortran` combo compared to C/C++. Intel optimizations seem to be much superior, but perhaps a bit better for C/C++ compiler. I want to analyze the code with a profiler today and see the curlpit. When analyzing the assembly, both versions seem to be vectorized correctly, but Fortran version for some reason is just a little bit slower.

Anyway, a new thread about “OOP overhead in Fortran” could be a better place for the further discussion, I do not want to derail the topic of FAR++. 🙂

---

<div class="post-metadata">

**Author:** ![PierU](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/pieru/32/1848_2.png) [@PierU](https://fortran-lang.discourse.group/u/PierU)\
**Post date:** [November 28, 2024, 4:42pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/27 "2024-11-28T16:42:38Z")

</div>

> [@gronki](#):
>
> In my version, `kernel` is an `allocatable`, therefore array is guaranteed to be contiguous. Actually, it can be guaranteed by `contiguous` attribute that I use in my procedural part of the code. Anyway, most compilers nowadays do generate contigous/non-contigous branches. So I disagree that OOP introduces any overhead here while providing much cleaner interface.

Contiguity and OOP are two orthogonal problems. Explicit shape dummy arrays are also guaranteed to be contiguous. And the compiler has just to pass an address in all cases, whereas with assumed shape the compiler possibly has to create a full descriptor, depending on what is the actual argument.

> [@gronki](#):
>
> Subroutine with 10 arguments feels like 70’s coding, but it must be a matter of preference, since I know many people who hate OOP interfaces!

Yes, explicit shapes requires more arguments, but beyond the preferences it’s really a choice to ensure the least possible overheads in the calls. I do also use assumed shapes whenever performances are not an issue. And although I do not often use OOP, I don’t hate it and sometimes use it. But, frankly, the 70’s coding style requires here just 20 lines of code, is easy to understand, to maintain (assuming there is something to maintain), and to use. What else?

---

<div class="post-metadata">

**Author:** ![RonShepard](https://avatars.discourse-cdn.com/v4/letter/r/a3d4f5/32.png) [@RonShepard](https://fortran-lang.discourse.group/u/RonShepard)\
**Post date:** [November 28, 2024, 6:33pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/28 "2024-11-28T18:33:10Z")

</div>

> [@PierU](#):
>
> Yes, explicit shapes requires more arguments, […]

I find that it is the use of derived types that is mostly responsible for reducing the number of subroutine arguments. Closely related variables (scalars and arrays) are grouped together into just a few derived types, and then those derived types are passed as single arguments. Those derived types themselves can, of course, be scalars, assumed shape arrays, or explicit shape arrays. Fortran is very flexible in this regard. The explicit shape vs. assumed shape is a separate issue, the only thing here really is whether the bounds and dimensions are passed as separate arguments or as part of the assumed shape declaration. Unless the call is in a tight loop with varying size arrays, the overhead associated with the array constructor for the argument association is trivial.

---

<div class="post-metadata">

**Author:** ![PierU](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/pieru/32/1848_2.png) [@PierU](https://fortran-lang.discourse.group/u/PierU)\
**Post date:** [November 28, 2024, 7:04pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/29 "2024-11-28T19:04:41Z")

</div>

> [@RonShepard](#):
>
> Unless the call is in a tight loop with varying size arrays, the overhead associated with the array constructor for the argument association is trivial.

It depends. For instance if such a 1D convolution routine, but with assumed shapes, is called on each column of a 2D array, the compiler may have to generate array descriptors on each call:

```auto
do j = 1, size(e,2)
   call sconv1d( f, e(:,i), x(:,i) )
end do

```

Moreover, passing the lower and upper bounds also makes the routine much more versatile.

---

<div class="post-metadata">

**Author:** ![Beliavsky](https://avatars.discourse-cdn.com/v4/letter/b/ba8739/32.png) [@Beliavsky](https://fortran-lang.discourse.group/u/Beliavsky)\
**Post date:** [November 28, 2024, 10:22pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/30 "2024-11-28T22:22:04Z")

</div>

> [@gronki](#):
>
> Subroutine with 10 arguments feels like 70’s coding, but it must be a matter of preference, since I know many people who hate OOP interfaces!

Subroutines with many arguments can be easily usable if only a few of the arguments are required and they appear before the `optional` arguments. The alternative may be an inflexible subroutine with various parameters hard-coded internally.

---

<div class="post-metadata">

**Author:** ![gronki](https://avatars.discourse-cdn.com/v4/letter/g/a88e4f/32.png) [@gronki](https://fortran-lang.discourse.group/u/gronki)\
**Post date:** [November 28, 2024, 10:51pm UTC](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856/31 "2024-11-28T22:51:12Z")

</div>

I created a topic to continue this side discussion: [Discussion about performance of OOP in Fortran](https://fortran-lang.discourse.group/t/discussion-about-performance-of-oop-in-fortran/8900/1)

[Previous page](https://fortran-lang.discourse.group/t/fortran-arrays-in-c/8856.md?page=1)
