While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose Prim, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce Absorb, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that Absorb consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.
A mathematical solution often begins before the first line of a proof, when an opaque problem is recognized through the structure that makes it solvable. For a mathematical problem, a primitive is the essential conceptual observation that reveals why the problem can be solved. It captures the latent mathematical structure shared by the problem and its solution, such as an invariant, a theorem condition, a representation, a reduction, or a reformulation. A primitive should be concise, problem-specific, and explanatory: it must make clear what property is being exploited and how that property unlocks the solution. It is not a routine calculation, a piece of generic advice, or a full proof.
Problem
Let \(L := \mathbb{Q}\left(\sqrt{(2+\sqrt{2})(3+\sqrt{3})},\ \sqrt{2},\ \sqrt{3}\right)\). What is the Galois group of \(L/\mathbb{Q}\)?
Primitive (core concept)
The key idea is to track the added radical by its square class over the biquadratic base; preservation of that square class gives the lifts, and the orders of those lifts together with whether they commute are what pin down the isomorphism type.
The full solution establishes the degree of the extension, constructs the lifted automorphisms explicitly, and computes their relations. The primitive is the structural idea that makes that derivation possible, not a shortened proof.
Taking the mathematical primitive as the unit of structural understanding, Prim decomposes mathematical reasoning into four complementary dimensions. Let \(x\) be the problem, \(p\) its primitive, \(y\) a complete solution, and \(\pi\) the model.
Discovery
\(\pi(x) \rightarrow \hat{p}\)
Find the primitive from the problem alone.
Generation
\(\pi(x) \rightarrow \hat{y}\)
Solve the problem directly.
Digestion
\(\pi(x, y) \rightarrow \hat{p}\)
Recover the primitive from a correct solution.
Execution
\(\pi(x, p) \rightarrow \hat{y}\)
Solve the problem once the primitive is given.
Data curation. We build on Humanity's Last Exam, focusing on its text-only, free-form mathematics subset with exact-match evaluation. To reduce annotation noise, we draw from the Gold and Revision subsets of HLE-Verified and randomly sample 200 problems. For each problem, GPT-5.4 drafts a primitive from the problem, reference answer, and gold rationale. Three human experts with graduate-level training in mathematics independently review the primitive and verify the validity of the problem, answer, and rationale. This excludes 18 unsuitable problems, yielding a final evaluation set of 182 problems.
Primitive scoring. Each predicted primitive is judged against the expert-verified gold primitive, with the reference solution used only to recognize mathematically equivalent formulations. The judge decides whether the response is a valid primitive (\(V\)), whether it identifies the essential structure (\(\sigma_{\text{gate}}\)), and whether it explains how that structure enables the solution (\(\sigma_{\text{mech}}\)):
\[ \mathrm{Score} = V \cdot \sigma_{\text{gate}} \cdot \big(0.6 + 0.4\,\sigma_{\text{mech}}\big), \qquad \mathrm{PrimitiveAcc} = \mathbf{1}\left[\mathrm{Score} \ge 0.8\right]. \]Generation and Execution are scored by final-answer accuracy with the official HLE judge.
We evaluate 12 open- and closed-source models from three lineages on all four dimensions of Prim.
| Model | Discovery | Generation | Digestion (Dig. − Disc.) | Execution (Exec. − Gen.) |
|---|---|---|---|---|
| OpenAI | ||||
| gpt-5.4 | 82.42 | 67.03 | 100.00 (+17.58) | 86.81 (+19.78) |
| gpt-5.4-mini | 59.34 | 50.55 | 97.80 (+38.46) | 71.98 (+21.43) |
| gpt-5.4-nano | 41.76 | 44.51 | 97.25 (+55.49) | 68.68 (+24.17) |
| gpt-oss-20B | 34.62 | 43.41 | 91.21 (+56.59) | 60.99 (+17.58) |
| Qwen | ||||
| Qwen3.6-27B | 24.73 | 52.75 | 92.31 (+67.58) | 78.57 (+25.82) |
| Qwen3.5-27B | 28.57 | 47.80 | 93.41 (+64.84) | 71.98 (+24.18) |
| Qwen3.5-9B | 13.19 | 38.46 | 79.67 (+66.48) | 61.54 (+23.08) |
| Qwen3.5-4B | 6.59 | 26.37 | 70.88 (+64.29) | 56.04 (+29.67) |
| DeepSeek-R1 Distill | ||||
| R1-0528-8B | 6.59 | 17.03 | 68.68 (+62.09) | 39.56 (+22.53) |
| R1-Distill-32B | 6.04 | 18.13 | 65.38 (+59.34) | 43.41 (+25.28) |
| R1-Distill-14B | 4.40 | 15.93 | 67.03 (+62.63) | 39.01 (+23.08) |
| R1-Distill-7B | 7.14 | 13.19 | 48.90 (+41.76) | 32.42 (+19.23) |
Absorb is an on-policy self-distillation method built on two design choices motivated by the diagnosis.
With \(\mathcal{S}_t\) the teacher's top-\(K\) token support at step \(t\), and \(\bar{\pi}_S\), \(\bar{\pi}_T\) the student and primitive-conditioned teacher distributions renormalized over \(\mathcal{S}_t\), Absorb minimizes a reverse KL with a one-sided clamp:
\[ \mathcal{L}_{\text{Absorb}} = \frac{1}{T}\sum_{t=1}^{T} \sum_{v\in\mathcal{S}_t} \min\left\{ \bar{\pi}_S(v \mid \hat{y}_{<t}, x)\, \log \frac{\bar{\pi}_S(v \mid \hat{y}_{<t}, x)}{\bar{\pi}_T(v \mid \hat{y}_{<t}, x, p)},\ \tau \right\}. \]The student never sees a primitive at inference time. Training uses 709 mathematics Ph.D. qualifying-exam proof problems, each paired with a human-written proof and a one-sentence primitive.
Absorb consistently outperforms SFT and OPSD across Qwen3.5 model scales and benchmarks. It improves Generation at all three scales, including a 5.49-point gain at 9B, where SFT and OPSD lower Generation by 6.04 and 4.95 points. The gains concentrate in downstream reasoning rather than explicit Discovery: the structural guidance of the primitive is reflected in the model's unassisted reasoning.
| Model | Prim | HLE Math | HMMT25 | Omni-MATH | Avg. | |
|---|---|---|---|---|---|---|
| Discovery | Generation | |||||
| Qwen3.5-4B | 6.59 | 26.37 | 26.70 | 76.67 | 78.00 | 51.94 |
| + SFT | 5.49 (−1.10) | 25.27 (−1.10) | 27.36 (+0.65) | 83.33 (+6.66) | 79.33 (+1.33) | 53.82 (+1.89) |
| + OPSD | 8.79 (+2.20) | 30.22 (+3.85) | 29.58 (+2.88) | 83.33 (+6.66) | 79.33 (+1.33) | 55.62 (+3.68) |
| + Absorb | 9.89 (+3.30) | 31.32 (+4.95) | 31.02 (+4.32) | 83.33 (+6.66) | 80.00 (+2.00) | 56.42 (+4.48) |
| Qwen3.5-9B | 13.19 | 38.46 | 35.60 | 90.00 | 78.67 | 60.68 |
| + SFT | 12.64 (−0.55) | 32.42 (−6.04) | 34.55 (−1.05) | 90.00 (+0.00) | 80.67 (+2.00) | 59.41 (−1.27) |
| + OPSD | 10.99 (−2.20) | 33.52 (−4.95) | 35.08 (−0.52) | 90.00 (+0.00) | 81.33 (+2.67) | 59.98 (−0.70) |
| + Absorb | 14.29 (+1.10) | 43.96 (+5.49) | 38.35 (+2.75) | 93.33 (+3.33) | 82.00 (+3.33) | 64.41 (+3.73) |
| Qwen3.5-27B | 28.57 | 47.80 | 42.67 | 96.67 | 87.33 | 68.62 |
| + SFT | 28.02 (−0.55) | 48.35 (+0.55) | 43.59 (+0.92) | 96.67 (+0.00) | 86.67 (−0.67) | 68.82 (+0.20) |
| + OPSD | 28.57 (+0.00) | 44.51 (−3.30) | 42.80 (+0.13) | 96.67 (+0.00) | 88.00 (+0.67) | 67.99 (−0.62) |
| + Absorb | 28.02 (−0.55) | 48.90 (+1.10) | 46.60 (+3.93) | 100.00 (+3.33) | 88.67 (+1.33) | 71.04 (+2.42) |
| Release | Description |
|---|---|
| 🤗 Prim | 182 research-level problems with gold answers, reference solutions, and expert-verified primitives |
| 🤗 math-phd-qual-709 | 709 Ph.D. qualifying-exam proof problems with proofs and primitives, the Absorb training set |
| 🤗 qwen-3.5-4b-absorb | Qwen3.5-4B trained with Absorb |
| 🤗 qwen-3.5-9b-absorb | Qwen3.5-9B trained with Absorb |
| 🤗 qwen-3.5-27b-absorb | Qwen3.5-27B trained with Absorb |
| Math-Primitive | Prim inference, judging, and scoring, and Absorb training code |
Evaluate a model on all four dimensions of Prim, then score it:
git clone https://github.com/taco-group/Math-Primitive && cd Math-Primitive
pip install -r requirements.txt
bash eval/run_all.sh shuoxing/qwen-3.5-9b-absorb qwen-3.5-9b-absorb
export OPENAI_API_KEY=... # the judges run on the OpenAI API
bash eval/judge_all.sh results/qwen-3.5-9b-absorb
We introduced the notion of Mathematical Primitive and Prim, a benchmark for systematically evaluating structural mathematical understanding in LLMs across Discovery, Generation, Digestion, and Execution. Our diagnosis reveals that similar solution accuracy can mask markedly different capability profiles, correct primitives can unlock substantial latent execution capacity, and independent Discovery constitutes a dominant bottleneck in mathematical reasoning. We further show that discovery-limited failures are substantially more amenable to post-training, while existing methods can introduce regressions on problems the model already solves. Building on these findings, we proposed Absorb, a primitive-privileged post-training paradigm that selectively transfers primitive-guided reasoning through a bounded override mechanism. Experiments on challenging mathematical reasoning benchmarks show that Absorb consistently improves over strong post-training baselines. Overall, our results highlight the value of explicitly diagnosing structural mathematical understanding and leveraging this structure to improve LLM reasoning.
@article{xing2026missing,
title = {The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models},
author = {Xing, Shuo and Dai, Zilin and Qian, Chengyuan and Lin, Fangzhou and Chen, Wenjing and He, Ping and Lu, Pan and Velasquez, Alvaro and Bansal, Mohit and Tu, Zhengzhong},
journal = {arXiv:2610.02191},
year = {2026}
}