There is no shortage of opinions in this world and nowhere is this more evident than in the current debate over AI. Some claims are more convincing than others. AI is undoubtedly a technological revolution. But its economic impacts are much less certain, while claims that it poses an existential threat to humanity are more speculative still. I am not qualified to opine on the latter while the former is the subject of much coverage elsewhere. But over the last year or so, my focus has been on how we can use AI in the realm of economics – less a case of what AI means for the economy and more what it means for economics.
The current state of play
My starting point is that as a practitioner of model-based forecasting, I am sceptical of many of the claims made for representative agent (RANK) models, epitomised by the dominance of DSGE models. Indeed, I have shared these concerns on this blog over the years (here and here, for example). One of the biggest weaknesses of representative agent models is that they assume away the heterogeneity in behaviour that can drive aggregate economic outcomes. In the wake of the GFC, the economics profession realised that household balance sheets, debt, liquidity constraints and differences in marginal propensities to consume (MPC) were issues that could not be assumed away. Accordingly, many economists increasingly questioned whether a representative household could adequately describe the transmission of monetary and fiscal policy. Thus was the HANK (Heterogeneous Agent New Keynesian) research paradigm born.
In a conventional RANK model, households are effectively treated as having substantial access to financial markets. An interest rate cut stimulates consumption largely through intertemporal substitution: consumption today becomes cheaper relative to consumption tomorrow. In HANK, many households have high MPCs and relatively little liquid wealth, implying limited ability to smooth consumption. Households with a high MPC thus spend much of the income freed up by the rate cut, which stimulates firms’ demand and creates additional jobs. The important insight from HANK models is that monetary policy can have a substantial effect through redistribution and changes in household income, rather than solely because everyone optimises consumption across time.
This is undoubtedly a more useful representation of how the
economy works in practice. But the computational aspects of HANK models are
significantly more onerous than in RANK models (see these
notes by Ben Moll – one of the originators of the research paradigm). In a
RANK model, the behaviour of a representative household can generally be
characterised by a relatively small number of equations. In HANK, by contrast,
the modeller has to solve the optimisation problem of heterogeneous households
and track the resulting distribution of households across income, wealth and
other state variables. The computational burden therefore rises substantially.
However, this raises a more fundamental question: How much
do we gain from this additional complexity? If a HANK model produces materially
better forecasts or a more accurate description of how the economy responds to
shocks and policy, then the additional computational sophistication may be
justified. But if its empirical performance is little better than that of much
simpler models, we can question whether greater theoretical and computational
complexity necessarily represents progress. Moreover, despite their
differences, RANK and HANK models share the common characteristic that behaviour
is governed by specified rules. Given the state of the economy and the relevant
shocks, the model determines how agents respond. Furthermore, such models rely
on the assumption of rational expectations. But as
Ben Moll argues in this accessible essay, the computational requirements of
this assumption are enormous, which limits how far heterogeneous‑agent
models can be pushed, especially for big questions involving uncertainty and
crises.
Can AI provide a solution?
My contention is that the use of LLMs which act as agents
can cut through this Gordian knot of complexity. Rather than attempting to
capture heterogeneity by adding ever more complicated equations to an
essentially deterministic behavioural framework, we can envisage an economy
populated by large numbers of heterogeneous agents, each with their own
characteristics, circumstances, information and history.
The attraction is that we may be able to model behavioural
heterogeneity without having to specify in advance a separate mathematical
decision rule for every possible state of the economy. The complexity is still
there, but it is handled differently. Instead of solving an increasingly
elaborate system of equations designed to describe how agents should behave, we
allow the LLMs to do the heavy lifting. They can process the information
available to each agent, take account of its particular circumstances and
generate a response that is not necessarily constrained by a single,
pre-specified behavioural rule. In principle, this allows the aggregate behaviour
of the economy to emerge from the interaction of heterogeneous agents, rather
than being imposed from the outset by the structure of the model.
This is not a costless exercise and many challenges remain. One of the most challenging aspects is to determine whether the behaviour generated by the agents is economically credible, empirically realistic and sufficiently stable to allow for forecasting or policy analysis. But in my view, this is a more interesting challenge than simply adding another layer of complexity to an already highly demanding optimisation framework.
What this does and does not mean
It is important to stress at the outset that this does not
mean outsourcing all of our analysis to the machine. Economists have a critical
role to play in knowing what questions to ask, recognising when the answer is
implausible, and understanding the economic mechanisms behind the results. Some
tasks will remain fundamentally human: exercising judgement, balancing
competing policy objectives and deciding which evidence should carry the
greatest weight.
From an economic modelling perspective, it appears as though
a hybrid system which allows for agentic AI to be embedded into a structural
model may be a promising line of research. The structural model provides the
economic framework and the constraints within which the agents operate, while
the AI provides a more flexible representation of behaviour and interaction.
This could potentially combine some of the strengths of traditional economic
models with the ability of AI to represent heterogeneous and adaptive
behaviour.
There is also a practical reason for retaining a structural
framework. Policymakers require a system against which forecasts can be
assessed and alternative scenarios explored. A forecast on its own – as
produced by a neural network model, for instance – may tell us what the model
expects to happen, but provides relatively little information about what
generates particular outcomes or how the outcome might change if the underlying
assumptions are altered. A structural model, by contrast, provides a framework
for asking counterfactual questions: what happens if interest rates remain
higher for longer or if energy prices rise? The attraction of a hybrid approach
is therefore that it could combine the behavioural flexibility of agentic AI
with the interpretability and policy relevance of a structural economic model.
The evidence so far
Space considerations preclude a more in-depth analysis of
the current state of play in the AI and economics research field. However, it
is worthwhile pointing interested readers in the direction of some of the work
that is currently being conducted.
One reason for not trusting LLMs as pure forecasting models
is an empirical problem. LLMs are trained on enormous quantities of historical
data, including economic and financial information. This creates a particular
problem when testing their forecasting ability: a model may appear to forecast
a historical observation successfully simply because it has memorised the
observation rather than because it has learned the underlying economic
relationships. Moreover, prompt instructions or masking techniques do not
reliably prevent this recall because prompts cannot change what information is
encoded in the model’s parameters (Lopez-Lira
et al, 2025). Evaluations of LLMs’ forecasting capabilities should thus focus
exclusively on data beyond their training cutoff, where memorization is
impossible.
The evidence in terms of AI agents, however, is more
promising. del Rio-Chanona et al
(2025) suggest that LLMs do not generate results strictly in line with rational
expectations, but instead display results consistent with bounded rationality,
similar to human participants in price setting experiments. This is an
interesting result which suggests that LLM-based agents may be capable of
reproducing some of the behavioural characteristics observed in actual economic
agents, rather than simply imposing the rational expectations assumption that
underpins much of conventional macroeconomic modelling. This could provide a
basis for modelling heterogeneity in expectations and behaviour without having
to specify these behavioural responses explicitly in advance.
Last word
Although the results of the agentic AI approach are
promising, many issues need to be resolved. Naïve inference from synthetic LLM
outputs may produce unreliable results unless we can be certain that the model
is not hallucinating, a known problem with LLMs. There are also questions about
reproducibility, calibration and validation. Outputs can vary between runs, thanks
to the stochastic nature of LLMs, while results may also be sensitive to the
precise instructions given to the model or to the particular LLM employed.
There is also a version of the Lucas critique to contend
with. Agents built on LLMs have learned from text describing how people behaved
in the past, so they may reproduce historical behaviour well but respond poorly
to a genuinely new policy regime, which is precisely the situation in which
policymakers most need a model. The structural framework only partly helps,
since it constrains what agents can do but not how they choose within those
constraints. The test, therefore, is whether agents' responses plausibly shift
when the policy environment changes, and not just whether they fit the data
from the regime in which they were trained.
More fundamentally, there is a difficult aggregation
problem. Even if individual agents display plausible behaviour, it does not
follow that their interaction will generate realistic aggregate economic
outcomes. The model therefore needs to be tested against established empirical
regularities at both the micro and macro levels. Finally, the possibility of
training-data contamination raises a further concern: apparent success in
reproducing historical outcomes may reflect memorisation rather than an ability
to generate genuinely out-of-sample economic behaviour.
Nonetheless, it is a fascinating area to be involved with –
assuming of course that our silicon masters do not extinguish us before we get
a chance to see the results.



No comments:
Post a Comment