Mathematician Timothy Gowers has reported that ChatGPT 5.5 Pro produced what he regarded as PhD-level work on additive number theory in about an hour, prompting him to revise upward his assessment of large language models' mathematical abilities. The results described in his account were reviewed by mathematicians connected to the underlying problems, but they were not presented as formally verified proofs or peer-reviewed publications.

Gowers selected questions from a paper by Mel Nathanson concerning the possible sizes of sumsets and the diameter needed to realize them. On one problem, Nathanson had established an exponential upper bound. Gowers said ChatGPT reasoned for 17 minutes and five seconds before proposing a construction with a quadratic upper bound, which he considered optimal. The model then produced a preprint-style version after an additional request, and Gowers spent time checking the argument.

The construction combined arithmetic progressions with more efficient Sidon sets, whose sums have a maximal distinctness property. Gowers characterized the improvement as arising from a better component within a reframing of Nathanson's argument. The model also adapted the method to a related restricted-sumset question.

A more demanding test concerned work by Isaac Rajagopal on general repeated sumsets. ChatGPT first proposed changing an exponential dependence into a different exponential bound. After further prompting, it produced a polynomial bound for fixed parameters. Rajagopal, who had developed the framework being improved, examined the output and described it as almost certainly correct. In a guest evaluation included in Gowers's post, he called the main idea original and clever and said it was the kind of contribution he would have been proud to find after extended work.

The sequence still depended on people at several stages. Gowers chose the problems, requested reformulations and checked one argument; Nathanson passed material to Rajagopal; and Rajagopal assessed the ideas against his own research. Gowers also raised unresolved questions about where AI-written mathematical results should live, how they should be moderated and whether human certification or proof-assistant formalization should be required.

The report is a detailed case study rather than a broad benchmark. It does not show that the model can reliably solve arbitrary open problems, and the account distinguishes improvements within an existing framework from solving a field's central obstacles unaided. Even with those limits, the experience suggests that models may rapidly clear some research questions previously considered suitable entry points for graduate students. Gowers expects that shift to affect both research practice and the way newcomers are introduced to mathematics, while emphasizing that present assessments may change quickly. The episode also leaves reproducibility open: the source reports a particular interaction, not repeated trials across accounts or independent model runs.