After GPT-6 Astra posted near-perfect results on ARC-AGI 3, a 4B small model that can run on a phone also delivered a remarkable score on the same benchmark.

S

trictly speaking, this is a combination of 4B Qwen-3.5 and 753B GLM-5.2: GLM-5.2 stays in the cloud, Qwen-3.5 runs on the phone, the large model never outputs a single word, and the