The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

1Texas A&M University, 2Harvard University, 3Vanderbilt University,
4Stanford University, 5DARPA, 6University of North Carolina at Chapel Hill
*Corresponding author
Diagnosing and internalizing mathematical primitives

Diagnosing and internalizing mathematical primitives. We introduce Mathematical Primitives to probe structural mathematical understanding in LLMs across four dimensions: Discovery, Generation, Digestion, and Execution. Our diagnosis further motivates Absorb, which selectively transfers primitive-guided reasoning into the student without requiring primitives at inference time.

Abstract

While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose Prim, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce Absorb, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that Absorb consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.

Mathematical Primitive

A mathematical solution often begins before the first line of a proof, when an opaque problem is recognized through the structure that makes it solvable. For a mathematical problem, a primitive is the essential conceptual observation that reveals why the problem can be solved. It captures the latent mathematical structure shared by the problem and its solution, such as an invariant, a theorem condition, a representation, a reduction, or a reformulation. A primitive should be concise, problem-specific, and explanatory: it must make clear what property is being exploited and how that property unlocks the solution. It is not a routine calculation, a piece of generic advice, or a full proof.

Problem

Let \(L := \mathbb{Q}\left(\sqrt{(2+\sqrt{2})(3+\sqrt{3})},\ \sqrt{2},\ \sqrt{3}\right)\). What is the Galois group of \(L/\mathbb{Q}\)?

Primitive (core concept)

The key idea is to track the added radical by its square class over the biquadratic base; preservation of that square class gives the lifts, and the orders of those lifts together with whether they commute are what pin down the isomorphism type.

The full solution establishes the degree of the extension, constructs the lifted automorphisms explicitly, and computes their relations. The primitive is the structural idea that makes that derivation possible, not a shortened proof.

The Prim Benchmark

Taking the mathematical primitive as the unit of structural understanding, Prim decomposes mathematical reasoning into four complementary dimensions. Let \(x\) be the problem, \(p\) its primitive, \(y\) a complete solution, and \(\pi\) the model.

Discovery

\(\pi(x) \rightarrow \hat{p}\)

Find the primitive from the problem alone.

Generation

\(\pi(x) \rightarrow \hat{y}\)

Solve the problem directly.

Digestion

\(\pi(x, y) \rightarrow \hat{p}\)

Recover the primitive from a correct solution.

Execution

\(\pi(x, p) \rightarrow \hat{y}\)

Solve the problem once the primitive is given.

Data curation. We build on Humanity's Last Exam, focusing on its text-only, free-form mathematics subset with exact-match evaluation. To reduce annotation noise, we draw from the Gold and Revision subsets of HLE-Verified and randomly sample 200 problems. For each problem, GPT-5.4 drafts a primitive from the problem, reference answer, and gold rationale. Three human experts with graduate-level training in mathematics independently review the primitive and verify the validity of the problem, answer, and rationale. This excludes 18 unsuitable problems, yielding a final evaluation set of 182 problems.

Primitive scoring. Each predicted primitive is judged against the expert-verified gold primitive, with the reference solution used only to recognize mathematically equivalent formulations. The judge decides whether the response is a valid primitive (\(V\)), whether it identifies the essential structure (\(\sigma_{\text{gate}}\)), and whether it explains how that structure enables the solution (\(\sigma_{\text{mech}}\)):

\[ \mathrm{Score} = V \cdot \sigma_{\text{gate}} \cdot \big(0.6 + 0.4\,\sigma_{\text{mech}}\big), \qquad \mathrm{PrimitiveAcc} = \mathbf{1}\left[\mathrm{Score} \ge 0.8\right]. \]

Generation and Execution are scored by final-answer accuracy with the official HLE judge.

Diagnosing Mathematical Reasoning

We evaluate 12 open- and closed-source models from three lineages on all four dimensions of Prim.

Model Discovery Generation Digestion (Dig. − Disc.) Execution (Exec. − Gen.)
OpenAI
gpt-5.482.4267.03100.00 (+17.58)86.81 (+19.78)
gpt-5.4-mini59.3450.5597.80 (+38.46)71.98 (+21.43)
gpt-5.4-nano41.7644.5197.25 (+55.49)68.68 (+24.17)
gpt-oss-20B34.6243.4191.21 (+56.59)60.99 (+17.58)
Qwen
Qwen3.6-27B24.7352.7592.31 (+67.58)78.57 (+25.82)
Qwen3.5-27B28.5747.8093.41 (+64.84)71.98 (+24.18)
Qwen3.5-9B13.1938.4679.67 (+66.48)61.54 (+23.08)
Qwen3.5-4B6.5926.3770.88 (+64.29)56.04 (+29.67)
DeepSeek-R1 Distill
R1-0528-8B6.5917.0368.68 (+62.09)39.56 (+22.53)
R1-Distill-32B6.0418.1365.38 (+59.34)43.41 (+25.28)
R1-Distill-14B4.4015.9367.03 (+62.63)39.01 (+23.08)
R1-Distill-7B7.1413.1948.90 (+41.76)32.42 (+19.23)
Table 1. Overall performance on the four dimensions of Prim.
  • Answer accuracy masks distinct capability profiles. Models with similar Generation accuracy can differ sharply in Discovery: Qwen3.6-27B and gpt-5.4-mini solve problems at a comparable rate yet differ by over 30 points in Discovery.
  • Correct primitives unlock latent execution capacity. Providing the gold primitive improves accuracy by 17.58 to 29.67 points across all 12 models. A teacher-written primitive also helps far more than a step-by-step teacher plan, so the gain comes from exposing the load-bearing structure, not from task decomposition.
  • Independent Discovery is the dominant bottleneck. Models recover the primitive from a correct solution far better than they find it alone; for example, Qwen3.6-27B rises from 24.73% Discovery to 92.31% Digestion. 83.6% of all Generation failures come with a failed Discovery.
  • Discovery-limited failures are repairable. After post-training, SFT and OPSD repair 20.4% and 21.1% of failures where the model cannot find the primitive but can execute it once given, about three times their repair rates on failures where it can do neither.

Absorb: Internalizing Primitive-Guided Reasoning

Absorb is an on-policy self-distillation method built on two design choices motivated by the diagnosis.

  • Primitives as privileged information. The primitive conditions the teacher rather than serving as a prediction target for the student. Teacher and student share the same base model; the student generates its own solution from the problem alone, and the primitive-conditioned teacher provides token-level supervision along that trajectory. Compared with full reference proofs, primitives supply the missing structural information without prescribing execution steps the student can already perform.
  • Bounded override. When the privileged teacher prefers a choice more than the student does, this positive guidance is transferred in full. When the teacher strongly suppresses a choice the student favors, the pressure is bounded, since that disagreement may come from information the student cannot access.

With \(\mathcal{S}_t\) the teacher's top-\(K\) token support at step \(t\), and \(\bar{\pi}_S\), \(\bar{\pi}_T\) the student and primitive-conditioned teacher distributions renormalized over \(\mathcal{S}_t\), Absorb minimizes a reverse KL with a one-sided clamp:

\[ \mathcal{L}_{\text{Absorb}} = \frac{1}{T}\sum_{t=1}^{T} \sum_{v\in\mathcal{S}_t} \min\left\{ \bar{\pi}_S(v \mid \hat{y}_{<t}, x)\, \log \frac{\bar{\pi}_S(v \mid \hat{y}_{<t}, x)}{\bar{\pi}_T(v \mid \hat{y}_{<t}, x, p)},\ \tau \right\}. \]

The student never sees a primitive at inference time. Training uses 709 mathematics Ph.D. qualifying-exam proof problems, each paired with a human-written proof and a one-sentence primitive.

Results

Absorb consistently outperforms SFT and OPSD across Qwen3.5 model scales and benchmarks. It improves Generation at all three scales, including a 5.49-point gain at 9B, where SFT and OPSD lower Generation by 6.04 and 4.95 points. The gains concentrate in downstream reasoning rather than explicit Discovery: the structural guidance of the primitive is reflected in the model's unassisted reasoning.

Model Prim HLE Math HMMT25 Omni-MATH Avg.
DiscoveryGeneration
Qwen3.5-4B6.5926.3726.7076.6778.0051.94
+ SFT5.49 (−1.10)25.27 (−1.10)27.36 (+0.65)83.33 (+6.66)79.33 (+1.33)53.82 (+1.89)
+ OPSD8.79 (+2.20)30.22 (+3.85)29.58 (+2.88)83.33 (+6.66)79.33 (+1.33)55.62 (+3.68)
+ Absorb9.89 (+3.30)31.32 (+4.95)31.02 (+4.32)83.33 (+6.66)80.00 (+2.00)56.42 (+4.48)
Qwen3.5-9B13.1938.4635.6090.0078.6760.68
+ SFT12.64 (−0.55)32.42 (−6.04)34.55 (−1.05)90.00 (+0.00)80.67 (+2.00)59.41 (−1.27)
+ OPSD10.99 (−2.20)33.52 (−4.95)35.08 (−0.52)90.00 (+0.00)81.33 (+2.67)59.98 (−0.70)
+ Absorb14.29 (+1.10)43.96 (+5.49)38.35 (+2.75)93.33 (+3.33)82.00 (+3.33)64.41 (+3.73)
Qwen3.5-27B28.5747.8042.6796.6787.3368.62
+ SFT28.02 (−0.55)48.35 (+0.55)43.59 (+0.92)96.67 (+0.00)86.67 (−0.67)68.82 (+0.20)
+ OPSD28.57 (+0.00)44.51 (−3.30)42.80 (+0.13)96.67 (+0.00)88.00 (+0.67)67.99 (−0.62)
+ Absorb28.02 (−0.55)48.90 (+1.10)46.60 (+3.93)100.00 (+3.33)88.67 (+1.33)71.04 (+2.42)
Table 2. Comparison of SFT, OPSD, and Absorb across three Qwen3.5 model scales.

Models and Data

ReleaseDescription
🤗 Prim182 research-level problems with gold answers, reference solutions, and expert-verified primitives
🤗 math-phd-qual-709709 Ph.D. qualifying-exam proof problems with proofs and primitives, the Absorb training set
🤗 qwen-3.5-4b-absorbQwen3.5-4B trained with Absorb
🤗 qwen-3.5-9b-absorbQwen3.5-9B trained with Absorb
🤗 qwen-3.5-27b-absorbQwen3.5-27B trained with Absorb
Math-PrimitivePrim inference, judging, and scoring, and Absorb training code

Evaluate a model on all four dimensions of Prim, then score it:

git clone https://github.com/taco-group/Math-Primitive && cd Math-Primitive
pip install -r requirements.txt
bash eval/run_all.sh shuoxing/qwen-3.5-9b-absorb qwen-3.5-9b-absorb
export OPENAI_API_KEY=...        # the judges run on the OpenAI API
bash eval/judge_all.sh results/qwen-3.5-9b-absorb

Conclusion

We introduced the notion of Mathematical Primitive and Prim, a benchmark for systematically evaluating structural mathematical understanding in LLMs across Discovery, Generation, Digestion, and Execution. Our diagnosis reveals that similar solution accuracy can mask markedly different capability profiles, correct primitives can unlock substantial latent execution capacity, and independent Discovery constitutes a dominant bottleneck in mathematical reasoning. We further show that discovery-limited failures are substantially more amenable to post-training, while existing methods can introduce regressions on problems the model already solves. Building on these findings, we proposed Absorb, a primitive-privileged post-training paradigm that selectively transfers primitive-guided reasoning through a bounded override mechanism. Experiments on challenging mathematical reasoning benchmarks show that Absorb consistently improves over strong post-training baselines. Overall, our results highlight the value of explicitly diagnosing structural mathematical understanding and leveraging this structure to improve LLM reasoning.

BibTeX

@article{xing2026missing,
  title   = {The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models},
  author  = {Xing, Shuo and Dai, Zilin and Qian, Chengyuan and Lin, Fangzhou and Chen, Wenjing and He, Ping and Lu, Pan and Velasquez, Alvaro and Bansal, Mohit and Tu, Zhengzhong},
  journal = {arXiv:2610.02191},
  year    = {2026}
}