# Introducing PyTorch Monarch – PyTorch

**URL:** <https://fortran-lang.discourse.group/t/introducing-pytorch-monarch-pytorch/10454>\
**Category:** Humor\
**Created:** [October 23, 2025, 12:23pm UTC](https://fortran-lang.discourse.group/t/introducing-pytorch-monarch-pytorch/10454 "2025-10-23T12:23:07Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![milancurcic](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/milancurcic/32/2_2.png) [@milancurcic](https://fortran-lang.discourse.group/u/milancurcic)\
**Post date:** [October 23, 2025, 12:23pm UTC](https://fortran-lang.discourse.group/t/introducing-pytorch-monarch-pytorch/10454/1 "2025-10-23T12:23:07Z")

</div>

PyTorch invented coarrays:

> **[Introducing PyTorch Monarch – PyTorch](https://pytorch.org/blog/introducing-pytorch-monarch/)**

(but seriously now, looks like a very interesting and potentially powerful approach for distributing work in Python)

---

<div class="post-metadata">

**Author:** ![certik](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/certik/32/4_2.png) [@certik](https://fortran-lang.discourse.group/u/certik)\
**Post date:** [October 24, 2025, 2:43pm UTC](https://fortran-lang.discourse.group/t/introducing-pytorch-monarch-pytorch/10454/2 "2025-10-24T14:43:17Z")

</div>

What are the missing things in Fortran so that it can be used for LLM training and inference?

---

<div class="post-metadata">

**Author:** ![loiseaujc](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/loiseaujc/32/4582_2.png) [@loiseaujc](https://fortran-lang.discourse.group/u/loiseaujc)\
**Post date:** [October 24, 2025, 3:44pm UTC](https://fortran-lang.discourse.group/t/introducing-pytorch-monarch-pytorch/10454/3 "2025-10-24T15:44:50Z")

</div>

I’m not familiar enough with `neural-fortran` by @milancurcic or `fiats` by @rouson but I’d say _easy-to-use automatic differentiation_ and a package implementing the standard optimizers used for training such networks. In a nutshell, I’d say it rather is that the ecosystem is still lacking some building blocks rather than an actual limitation of the language.

A good playground may be to port as much as possible of [`nanochat`](https://github.com/karpathy/nanochat) in Fortran.

---

<div class="post-metadata">

**Author:** ![jeremie.vandenplas](https://avatars.discourse-cdn.com/v4/letter/j/76d3ee/32.png) [@jeremie.vandenplas](https://fortran-lang.discourse.group/u/jeremie.vandenplas)\
**Post date:** [October 25, 2025, 9:02pm UTC](https://fortran-lang.discourse.group/t/introducing-pytorch-monarch-pytorch/10454/4 "2025-10-25T21:02:52Z")

</div>

I agree with @loiseaujc that the language is not the limiting factor: a first implementation of transformers was added in `neural-fortran` but I never tested it. The library `athena` seems to be quite complete too (but I didn’t tested it either).

---

<div class="post-metadata">

**Author:** ![nedanator](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/nedanator/32/4416_2.png) [@nedanator](https://fortran-lang.discourse.group/u/nedanator)\
**Post date:** [October 30, 2025, 9:11am UTC](https://fortran-lang.discourse.group/t/introducing-pytorch-monarch-pytorch/10454/5 "2025-10-30T09:11:47Z")

</div>

@loiseaujc and @jeremie.vandenplas are right, it’s not a language issue, it’s just building out the libraries to do the work and make it easy for people to build their own neural networks and perform automated hyperparameter optimisation using available tools like Weights and Biases (and optimising the codes so that they run at comparable speeds to PyTorch).

Most of the libraries such as `neural-fortran`, `fiats`, `athena` already have the standard optimisers implemented (such as adam and sgd). @loiseaujc, interesting you mention that, I am currently in the process of implementing automatic differentiation into `athena`. It works for some of the layers (such as fully connected), I’m just battling with fixing memory issues.
