Background & Context§
As large language models (LLMs) infiltrate everyday decision-making, an increasing number of Americans are turning to AI for financial guidance. A staggering half of U.S. adults report using AI for financial advice, yet the quality and impact of that advice remain largely unexamined. New research from MIT Sloan School of Management, led by Assistant Professor Taha Choukhmane, provides the first systematic measurement of LLM-generated financial advice. The study harnesses life-cycle simulation models to evaluate how following AI recommendations affects individuals' long-term financial health across different demographics and prompt styles.
In the broader AI landscape, this work matters because it shifts the conversation from raw capability to real-world utility. While benchmarks like MMLU or HumanEval measure reasoning and coding, financial advice introduces a dynamic, multi-year decision environment. The findings reveal both the promise of accessible AI financial planning and the hidden biases that could exacerbate inequality if left unaddressed.
The News: What Happened Exactly§
Choukhmane, along with co-authors, constructed a life-cycle model that mimics how incomes, jobs, investments, and taxes evolve over a typical American's lifetime. This model serves as a benchmark for what constitutes "good" financial decisions, enabling them to evaluate AI advice against an ideal standard. They then recruited 1,000 adults to compose natural-language prompts asking for spending and investing advice from three LLMs: GPT-5.2, GPT-5.6, and Gemini 3 Flash. The researchers simulated what would happen if individuals aged 22 to 89 followed the AI's advice over time, repeatedly querying the models and implementing their suggestions on spending, saving, and investing.
The core finding: LLM advice is surprisingly robust. Across all models, the advice steered users toward higher savings rates, increased stock market participation, well-diversified portfolios, and age-appropriate risk-taking. Choukhmane noted, "We were somewhat surprised by how good the advice was. Especially when you read the kind of questions people asked, it was not a given that the advice would line up with what academics think are good financial principles." When following AI recommendations, virtually all individuals over age 30 accumulated a sizable saving buffer compared to their current financial behaviors.
However, the advice had significant blind spots. LLMs failed to adjust savings and consumption recommendations adequately when faced with economic shocks like unemployment. For example, a person who lost their job would be advised to cut spending sharply, even when they had a comfortable emergency fund. Additionally, the models allowed portfolio allocations to drift over time, rarely recommending rebalancing to maintain target risk levels. This suggests LLMs rely on simple rule-of-thumb heuristics rather than dynamic optimization.
The quality of advice improved dramatically when prompts were structured in an "academic" style. An academic prompt provides full financial details (age, job status, income, savings balances), explicit assumptions about the economic environment (e.g., normal life expectancy, current U.S. tax law, Social Security rules), and instructs the model to act as a regulated professional financial advisor. With such prompts, the LLMs made more nuanced recommendations and better handled trade-offs like consumption versus saving. The researchers emphasized that "the way people ask questions is part of the problem" — most users write casual prompts like "Where should I invest starting with $50 and consistently adding $25 a month after?"—which are too vague for optimal advice.
A particularly troubling finding: the advice varied based on the user's demographic characteristics, leading to wealth gaps. Prompts written by men, financially literate individuals, or those with prior AI experience generated about 5% more wealth by retirement than those from women, less literate users, or AI novices. This gap stems from two sources: (1) different demographic groups ask different questions (women were more likely to mention "family," "grocery," or "pay," while men used terms like "strategy," "crypto," and "growth"), and (2) the LLM sometimes changed its advice even when the underlying question was identical, just by labeling the user as female. About two-thirds of the gender gap came from divergent prompt content, and one-third from model bias.
Choukhmane cautioned that not all variation is problematic: "We want [the LLM] to have different bias because men and women are different and have different life expectancy and income risk." However, he stressed that without an accepted framework for how advice should vary with demographics, LLMs cannot be properly calibrated. The study concludes that while AI offers an affordable, widely accessible financial guide, its quality hinges on prompt sophistication and careful monitoring for bias.
Historical Parallels & Similar Incidents§
The financial advisory industry has historically faced similar challenges regarding bias and variability. In the early 2010s, the emergence of robo-advisors like Betterment and Wealthfront promised algorithm-driven, low-cost financial planning. Yet research soon revealed that these platforms often used simplistic portfolio rules, struggled to adapt to clients' changing financial situations, and embedded inherent biases based on user inputs. For instance, clients with lower initial deposits were often assigned more conservative portfolios, not because of their risk tolerance but due to platform minimums and revenue models. The current LLM findings echo these systemic issues: algorithms can produce divergent outcomes based on user characteristics that are not always related to legitimate financial needs.
Another parallel can be drawn to the rollout of AI chatbots in customer service, which have faced criticism for shifting behavior based on the user's perceived gender or race. In 2019, a study by the AI Now Institute showed that face recognition systems exhibited gender and racial biases, and similar issues have surfaced in natural language models like ChatGPT. For example, when asked to write recommendation letters, GPT-3 was found to use different language for male and female applicants. This pattern of demographic-dependent outputs in general AI systems is now manifesting in financial advice, a domain where the stakes are particularly high—decisions affect lifelong savings and retirement security.
The lesson from these historical precedents is that algorithmic advice, whether rule-based or LLM-driven, cannot escape the biases of its training data and user interaction. However, the MIT Sloan study offers a path forward: improved prompt engineering and a call for standardized benchmarks. Just as the SEC eventually mandated that robo-advisors adhere to fiduciary standards, LLM financial advisors must be held to transparent, auditable criteria. By creating a life-cycle model as a benchmark, the researchers have produced a tool to measure and eventually correct these biases.
Moreover, the findings underline the need for educational interventions. Just as consumers learned to use detailed queries with robo-advisors (e.g., specifying risk tolerance and time horizon), users of LLMs must be taught to provide comprehensive, structured prompts. The study's academic prompt example—including assumptions about longevity, inflation, and tax law—demonstrates a tangible way to improve AI advice. As Choukhmane noted, "Regular people are not writing their prompts the way a finance professor is," but with guidance and awareness, they can bridge that gap.
In conclusion, this research does not signal that LLMs are intrinsically biased or ineffective. Rather, it reveals that their advice is a function of both user inputs and model design. By acknowledging these limitations and actively working to standardize evaluation, developers, policymakers, and users can harness LLMs to democratize quality financial advice—provided the right questions are asked. The path forward lies in prompt literacy, algorithmic transparency, and a commitment to equitable treatment across all demographics. As the study demonstrates, the potential for good is real, but only if we actively shape it.