# fortbench: a benchmark for agentic coding of Fortran

**URL:** <https://fortran-lang.discourse.group/t/fortbench-a-benchmark-for-agentic-coding-of-fortran/10785>\
**Category:** Announcements\
**Created:** [March 14, 2026, 11:05am UTC](https://fortran-lang.discourse.group/t/fortbench-a-benchmark-for-agentic-coding-of-fortran/10785 "2026-03-14T11:05:56Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![krystophny](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/krystophny/32/6650_2.png) [@krystophny](https://fortran-lang.discourse.group/u/krystophny)\
**Post date:** [March 14, 2026, 11:05am UTC](https://fortran-lang.discourse.group/t/fortbench-a-benchmark-for-agentic-coding-of-fortran/10785/1 "2026-03-14T11:05:56Z")

</div>

I was playing with the new Qwen 3.5 family of models and benchmarked them inspired by SWE-Bench on [GitHub - lazy-fortran/fortbench: Real-world Fortran coding benchmark for agent CLIs · GitHub](https://github.com/lazy-fortran/fortbench) . Qwen is still a bit worse than Claude and GPT, but can solve more than 50% of my tasks. Would be curious about input or feedback how to expand it.

PS: Changed my user from @ert to @krystophny to be consistent with GitHub.

---

<div class="post-metadata">

**Author:** ![certik](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/certik/32/4_2.png) [@certik](https://fortran-lang.discourse.group/u/certik)\
**Post date:** [March 14, 2026, 2:00pm UTC](https://fortran-lang.discourse.group/t/fortbench-a-benchmark-for-agentic-coding-of-fortran/10785/2 "2026-03-14T14:00:23Z")

</div>

Are you running Qwen 3.5 locally? I tried it using the qwen-code, but it wasn’t able to fix a simple Fortran problem. But as a chat it works really well, probably the best local model I tried.

---

<div class="post-metadata">

**Author:** ![krystophny](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/krystophny/32/6650_2.png) [@krystophny](https://fortran-lang.discourse.group/u/krystophny)\
**Post date:** [March 14, 2026, 2:11pm UTC](https://fortran-lang.discourse.group/t/fortbench-a-benchmark-for-agentic-coding-of-fortran/10785/3 "2026-03-14T14:11:49Z")

</div>

Yes! I am using opencode with the qwens, not qwen-code. I am now also trying to wire it to codex as a local model. For this, llama.cpp needed some modifications because of unknown tool names but it runs now. How good I cannot tell yet. I also did only benchmarks but no practical work yet, and seems like for benchmarks the larger qwens (27B, 35B-A3B, 122B-A10B) are barely usable.

---

<div class="post-metadata">

**Author:** ![certik](https://yyz2.discourse-cdn.com/free1/user_avatar/fortran-lang.discourse.group/certik/32/4_2.png) [@certik](https://fortran-lang.discourse.group/u/certik)\
**Post date:** [March 14, 2026, 2:29pm UTC](https://fortran-lang.discourse.group/t/fortbench-a-benchmark-for-agentic-coding-of-fortran/10785/4 "2026-03-14T14:29:24Z")

</div>

Yes, I used Qwen3.5-35B-A3B-8bit, and once it runs, it has about 70 tokens/s on my laptop, so very usable. But in qwen-code, it would load the whole conversation over and over, so it would take 5-10 minutes to load the prompt, then quickly generate a response in a few seconds, then qwen-code would load again for every request, and for any task you need, say, 20 requests, so in practice it was unusable. Given that the task continues, I would think you don’t need to reload the prompt from scratch. I am sure this will get figured out in the coming years. As a chat, it is very good.
