Large language models can steer users toward broadly sound saving and investing habits, but their advice remains sensitive to how questions are asked and can break down when circumstances change, according to research described by the MIT Sloan School of Management.

The researchers created a model of how income, employment, investments and taxes commonly develop over a person's life, giving them a benchmark against which to assess financial decisions. They then asked 1,000 adults to write prompts seeking spending and investment guidance from GPT-5.2, GPT-5.6 or Gemini 3 Flash. The team simulated outcomes for people aged 22 to 89 who repeatedly followed the resulting advice over time.

A second set of tests used structured academic prompts containing details such as age, job status, income and account balances, alongside assumptions about taxes, Social Security, longevity and economic risks. Advice improved when models received that fuller context, the researchers reported.

Across the simulations, the systems consistently encouraged saving during working years, drawing down savings in retirement and holding diversified stock funds. They also generally recommended reducing equity exposure after age 45. Following the guidance produced meaningful savings buffers for nearly all simulated people older than 30, while increasing stock-market participation and encouraging age-related risk adjustments.

The results were not uniformly strong. The models relied heavily on simple rules and reacted poorly to some financial shocks. After a job loss, for example, they sometimes recommended spending cuts that were too severe even when the user had savings available. Portfolios were also allowed to drift instead of being actively rebalanced.

Outcomes varied with users' gender, financial literacy and previous experience with AI. Prompts written by men, people with greater financial knowledge and experienced AI users produced about 5% more wealth near retirement in the simulations. For the gender difference, researchers attributed roughly two-thirds of the gap to differences in prompt wording and one-third to models changing their answer when an otherwise identical prompt was labeled as coming from a woman.

Those findings do not establish that a chatbot is a substitute for regulated, individualized financial advice. They arise from simulations built around selected assumptions, not from years of observed household outcomes. The models' weaknesses around unemployment, portfolio maintenance and demographic variation also concern decisions where generic guidance can carry substantial consequences.

The research nevertheless suggests that AI can widen access to basic financial guidance if systems collect the information needed for a useful answer. Taha Choukhmane, an MIT Sloan assistant professor and co-author, framed the central design challenge as making such advice work for people who lack financial expertise and do not know how to construct a technically complete prompt.