# A recent experience with ChatGPT 5.5 Pro

# A recent experience with ChatGPT 5.5 Pro

**Author:** Tim Gowers (gowers.wordpress.com)
**Published:** 2026-05-08
**URL:** https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/
**Co-authored sections:** Isaac Rajagopal (MIT undergrad)

## Context

Gowers is a Fields medalist (1998) in combinatorics. He has been given access to ChatGPT 5.5 Pro and recounts an experience over "a little over a week" where the model produced a piece of PhD-level research in additive combinatorics in under two hours, with **no serious mathematical input** from him.

## The setup

The problem space: Mel Nathanson's paper *Diversity, Equity and Inclusion for Problems in Additive Number Theory* (arXiv:2603.15556). For a set $A$ of integers, the $h$-fold sumset $hA = \{a_1 + \dots + a_h : a_i \in A\}$. Define $\mathcal{R}(h,k)$ as the set of all $t$ such that some $|A|=k$ achieves $|hA|=t$. The question Gowers fed to ChatGPT: how large a diameter (i.e. max element minus min) does $A$ need if you want $|A|=k$ and $|hA|$ of a prescribed size? Nathanson had given a $2^k - 1$ upper bound and asked whether it could be improved.

## What ChatGPT did

**Round 1 ($h=2$).** ChatGPT thought for 17:05 and returned a construction yielding a quadratic upper bound (clearly best possible), then wrote it up as a LaTeX preprint in 2:23. The improvement: use a more efficient Sidon set than powers of 2.

**Round 2 (restricted sumset).** No trouble. Gowers got both written up in a single note.

**Round 3 (general $h$).** Gowers was less optimistic. The problem leans on Isaac Rajagopal's earlier work. ChatGPT was asked to tighten Rajagopal's argument:

1. **16:41** — improved the bound from exponential in $k$ to exponential in $k^{1/2 + \varepsilon}$ (routine extension of Rajagopal's work).
2. **47:39** — wrote it up in preprint form. Sent to Rajagopal via Nathanson; Rajagopal said it looked correct.
3. ChatGPT speculated on pushing to polynomial; Gowers asked it to try.
4. **13:33** — felt optimistic, but flagged technical statements needing checking.
5. **9:12** — checked them.
6. **31:40** — final preprint with $N(h,k) \leq O(k^{10h^3})$ — polynomial in $k$.
7. Rajagopal: "almost certainly correct, not just at line-by-line level but at the level of ideas."

Total: **under 2 hours of model time** for a result Rajagopal said would have made him "very proud" after a week or two of pondering.

## Rajagopal's evaluation (guest section)

The key idea — original to ChatGPT — was the use of $h^2$-dissociated sets to recreate the sumset structure of the geometric series $S = \{0, 1, m, m^2, \dots\}$ and $T = \{1, m, \dots\}$ but with all elements of polynomial size in $\ell$ rather than exponential.

Rajagopal asked ChatGPT (via Gowers) whether such polynomial-sized sets could exist; he had no idea. ChatGPT returned a construction:

$$G = \{0, u_1, \dots, u_r, mu_1, \dots, mu_r\}, \quad H = \{u_1, \dots, u_r, mu_1, \dots, mu_r\}$$

where $\{u_1, \dots, u_r\}$ is an $h^2$-dissociated set built from finite-field generators (Bose–Chowla 1963 style). The intuition Rajagopal reconstructs in retrospect: $S$ and $T$ contain ~$\ell$ relations of the form $mx = y$; $G$ and $H$ contain ~$\ell/2$ of them while having few low-order relations because $U$ is $h^2$-dissociated. ChatGPT preserved the four key properties ($B_{m-1}$/$B_m$ membership, linear/quadratic deficits) while collapsing the diameter from exponential to polynomial.

Rajagopal: *"As far as I can tell, this idea is completely original."*

The full proof carries through as in Rajagopal's original, with $G, H$ replacing $S, T$. Final bound:

$$N(h,k) \leq 2qM(2hM)^{2q-1} \leq k^{10h^3}$$

Rajagopal contributed three appendices working through the dissociated-set construction, the parameter bookkeeping for the disjoint-union construction, and a section-by-section correspondence between his paper and the ChatGPT preprint.

## Gowers' read

> "I would judge the level of the result that ChatGPT found in under two hours to be that of a perfectly reasonable chapter in a combinatorics PhD."

The result leans heavily on Isaac's framework but the extension is non-trivial — reading the paper, looking for non-optimal places, familiarizing oneself with the algebraic techniques. For a beginning PhD student, that's weeks of work.

**The bar for PhD-trainable problems just rose.** Mentors traditionally hand new students "a problem that looks as though it might be a relatively gentle one." If LLMs can solve gentle problems, that's no longer an option. The lower bound for a meaningful contribution is now: prove something *in collaboration with LLMs* that LLMs can't do alone.

Two qualifications:

1. The PhD student can use LLMs too. The task is collaborative LLM-prove, not LLM-can't-prove.
2. Gowers doesn't know how this generalizes outside combinatorics. Combinatorics is problem-driven (start with a question and reason backward). Forward-reasoning fields — picking which observations are interesting — may be different.

## What it means for mathematical research

Gowers' personal answer to "should one still do mathematical research":

- The era of having one's name attached to a particular theorem may be ending — not just for you, but for anybody.
- A thought experiment: a mathematician solves a major problem via a long LLM exchange where the human played a guiding role but the LLM did the technical work and had the main ideas. Would we count that as a major achievement *of the mathematician*? "I don't think we would."
- Reason to still do it: solving hard problems gives you insight into the *problem-solving process itself*. People who have solved difficult problems are likely better at AI-assisted problem solving — same way good coders are better at vibe coding, or people good at arithmetic are better at noticing when a calculator gives an off-feeling answer.

> "By doing research in mathematics, you may not get the same rewards as your equivalents a generation ago, but there is a good chance that you will be equipping yourself very well for the world we are about to experience."

## On where the result should live

> "Had the result been produced by a human mathematician, it would definitely have been publishable, so I think it would be wrong to describe it as AI slop."

But journal publication seems pointless when the result can just be hosted as a PDF. arXiv has an anti-AI policy, which Gowers endorses. He suggests a separate moderated repository — moderation requiring human-mathematician certification, ideally proof-assistant verification, ideally answering a question in a human-written paper. He warns against AI-assisted moderation for obvious reasons.

## Preprint links

- The $h=2$ note: https://drive.google.com/file/d/11r-ggU__GMmHIrgEHQVULUIR1VxKSwmi/view
- The general-$h$ preprint: https://drive.google.com/file/d/1IkJBcWYz_3J_QGsESBmMa-jrEHAJDcJB/view
- Rajagopal's original: https://arxiv.org/pdf/2510.23022
