Current Fortran codebases: a vision of the future for AI-generated codebases?

I’ve been thinking informally about how current long-lived Fortran codebases might be used to provide a vision of the future for software engineering, especially in relation to Agentic AI. I’ll set out the emerging ‘argument’ below. I am interested in people’s thoughts. For those who identify weaknesses in the following (which is fine), I’d very-much also appreciate suggestions on how to tackle those weaknesses.

TL;DR

This is speculative at the moment. Accepting that, the emerging argument leads to the following three hypotheses:
H1: Agentic AI is accelerating the comprehension debt; and therefore
H2: Agentic AI is accelerating codebases to legacy status; and
H3: The challenges confronting Fortran software developers – e.g., with their two-decade (or longer-lived) Fortran codebases – provide an insight into (a vision of, or a metaphor for) the impact of Agentic AI on codebases. What took 20 years (240 months) of human coding can now, with the aid of Agentic AI, take 2 years (24 months) to achieve; but with a commensurate impact on comprehension debt and a commensurate shift to legacy status.

The emerging ‘argument’

I am attending the 20th Empirical Software Engineering and Measurement (ESEM) conference, at Garching, near Munich, in Germany. Today’s keynote talk was by Professor Emeritus Markku Oivo. Markku spoke about the relationship of AI for Software Engineering and Software Engineering for AI. It was a very provocative (in a good way) and thought provoking talk. For example, Markku reported analysis which suggested that 80% of the papers being presented at the conference this year, which reported results relating to LLMs, were already out of date because of the speed with which LLM models are being released and replaced.

Markku was citing examples, some from research and some from the grey literature, where practitioners are reporting an order of magnitude increase in the productivity of software development. A software change that might have taken 10 months to develop can now, with the appropriate agentic AI tooling, take a 1 month. These figures, of 10 months and 1 month, are, of course, convenient approximations. But we can use them to think through some implications.

Markku also raised concerns about “comprehension debt”, also referred to as “knowledge debt”. Essentially, as human software engineers offload the actual coding onto Agentic AI tooling, so the human’s understanding of the code decreases. Again, using approximations, if there is an order of magnitude increase in productivity using Agentic AI, one might also - simplistically - infer an order of magnitude decrease in human understanding. Or, phrased differently, an order of magnitude increase in comprehension debt. (I am of course using abstract units of measurement here.)

In our recent international survey of the Fortran community (also to be reported at the ESEM conference), participants reported challenges relating to comprehension debt, though they don’t necessarily use that term. Fortran software engineers and computational scientists said they struggled to understand both what the existing code does, and/or why the code was written the way it was, because (for example) the code was written a long time ago and the knowledge of what and why the coding is written the way it is, is no longer available, e.g., the person who wrote the code is no longer around to explain the code (that person left, retired, or worse) and/or there is no suitable documentation or related artefacts, e.g., test cases.

Long-lived Fortran codebases of, say, a couple of decades will have accumulated layers of code written by different developers, with different coding styles, potentially to different Fortran language standards, perhaps also with different software engineering practices. These factors make it (much) harder for the current Fortran software developers to understand the code. Or, put simply, there has been a decrease in human understanding and an increase in comprehension debt.

Bringing the insights on Agentic AI together with the insights on Fortran codebases, I suggest three hypotheses:
H1: Agentic AI is accelerating the comprehension debt; and therefore
H2: Agentic AI is accelerating codebases to legacy status; and
H3: The challenges confronting Fortran software developers – e.g., with their two-decade (or longer-lived) Fortran codebases – provide an insight into (a vision of, or a metaphor for) the impact of Agentic AI on codebases. What took 20 years (240 months) of human coding can now, with the aid of Agentic AI, take a mere 2 years (24 months) to achieve; but with a commensurate impact on comprehension debt and a commensurate shift to legacy status.

I’m interested to know what people think about this logic.

Thanks

Austen

What I would like to see is a study that compares time, cost and return on investment of using Agentic AI to modernize old existing code bases versus just starting from scratch and writing new code from the ground up using AI. A case study might be LAPACK. Would it be better to write a replacement for LAPACK from first principles using AI and make it aware of GPU’s, parallelism etc or just try to add that capability to a refactored LAPACK.