{
  "passages": [
    {
      "src": "2501.12948",
      "dist": 0.6247,
      "text": "es across difficulty levels, the training trends still demonstrate that while simpler reasoning tasks (for humans) are mastered early in training, the model’s capability on complex reasoning problems (level 3-5) significantly improves over time. C.2. Evolution of Advanced Reasoning Behaviors in DeepSeek-R1-Zero during Training We analyze the change in the reasoning behavior of the model during tra"
    },
    {
      "src": "2501.12948",
      "dist": 0.6905,
      "text": "ing a valuable resource for understanding the mechanisms underlying long chain-of-thought (CoT) reasoning models and for fostering the development of more powerful reasoning models. We release DeepSeek-R1 series models to the public at https://huggingface.co/deepseek-ai. 2. DeepSeek-R1-Zero We begin by elaborating on the training of DeepSeek-R1-Zero, which relies exclusively on reinforcement learn"
    },
    {
      "src": "2501.12948",
      "dist": 0.7175,
      "text": "ification steps. To address this, DeepSeek-R1-Zero enables direct exploration of reasoning patterns by the model itself, independent of human priors. The reasoning trajectories discovered through this selfexploration are subsequently distilled and used to train other models, thereby promoting the acquisition of more robust and generalizable reasoning capabilities. A.3. A Comparison of GRPO and PPO"
    },
    {
      "src": "2501.12948",
      "dist": 0.7222,
      "text": "eek-R1-Zero exemplifies how RL can autonomously enhance a model’s reasoning capabilities. As shown in Figure 1(b), DeepSeek-R1-Zero exhibits a steady increase in thinking time throughout training, driven solely by intrinsic adaptation rather than external modifications. Leveraging long CoT, the model progressively refines its reasoning, generating hundreds to thousands of tokens to explore and imp"
    },
    {
      "src": "2501.12948",
      "dist": 0.7428,
      "text": "n but in the provision of hard reasoning questions, a reliable verifier, and sufficient computational resources for reinforcement learning. Sophisticated reasoning behaviors, such as self-verification and reflection, appeared to emerge organically during the reinforcement learning process. Even if DeepSeek-R1 achieves frontier results on reasoning benchmarks, it still faces several capability limi"
    },
    {
      "src": "2501.12948",
      "dist": 0.7457,
      "text": "C.1. Evolution of Reasoning Capability in DeepSeek-R1-Zero during Training We analyzed DeepSeek-R1-Zero’s performance on the MATH dataset stratified by difficulty levels (1-5). Figure 8 reveals distinct learning patterns: easy problems (levels 1-3) quickly reach high accuracy (0.90-0.95) and remain stable throughout training, while difficult problems show remarkable improvement - level 4 problems"
    }
  ],
  "user_msg": "Retrieved passages from ancient Indian texts:\n\n[Passage 1] (DeepSeek-AI et al. 2025, arXiv:2501.12948)\nes across difficulty levels, the training trends still demonstrate that while simpler reasoning tasks (for humans) are mastered early in training, the model’s capability on complex reasoning problems (level 3-5) significantly improves over time. C.2. Evolution of Advanced Reasoning Behaviors in DeepSeek-R1-Zero during Training We analyze the change in the reasoning behavior of the model during tra\n\n[Passage 2] (DeepSeek-AI et al. 2025, arXiv:2501.12948)\ning a valuable resource for understanding the mechanisms underlying long chain-of-thought (CoT) reasoning models and for fostering the development of more powerful reasoning models. We release DeepSeek-R1 series models to the public at https://huggingface.co/deepseek-ai. 2. DeepSeek-R1-Zero We begin by elaborating on the training of DeepSeek-R1-Zero, which relies exclusively on reinforcement learn\n\n[Passage 3] (DeepSeek-AI et al. 2025, arXiv:2501.12948)\nification steps. To address this, DeepSeek-R1-Zero enables direct exploration of reasoning patterns by the model itself, independent of human priors. The reasoning trajectories discovered through this selfexploration are subsequently distilled and used to train other models, thereby promoting the acquisition of more robust and generalizable reasoning capabilities. A.3. A Comparison of GRPO and PPO\n\n[Passage 4] (DeepSeek-AI et al. 2025, arXiv:2501.12948)\neek-R1-Zero exemplifies how RL can autonomously enhance a model’s reasoning capabilities. As shown in Figure 1(b), DeepSeek-R1-Zero exhibits a steady increase in thinking time throughout training, driven solely by intrinsic adaptation rather than external modifications. Leveraging long CoT, the model progressively refines its reasoning, generating hundreds to thousands of tokens to explore and imp\n\n[Passage 5] (DeepSeek-AI et al. 2025, arXiv:2501.12948)\nn but in the provision of hard reasoning questions, a reliable verifier, and sufficient computational resources for reinforcement learning. Sophisticated reasoning behaviors, such as self-verification and reflection, appeared to emerge organically during the reinforcement learning process. Even if DeepSeek-R1 achieves frontier results on reasoning benchmarks, it still faces several capability limi\n\n[Passage 6] (DeepSeek-AI et al. 2025, arXiv:2501.12948)\nC.1. Evolution of Reasoning Capability in DeepSeek-R1-Zero during Training We analyzed DeepSeek-R1-Zero’s performance on the MATH dataset stratified by difficulty levels (1-5). Figure 8 reveals distinct learning patterns: easy problems (levels 1-3) quickly reach high accuracy (0.90-0.95) and remain stable throughout training, while difficult problems show remarkable improvement - level 4 problems\n\n---\n\nQuestion: Did reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability, or a scaled continuation of the same behavior?\n\nAnswer in English using Nyaya Panchavayava format with all seven steps.\n\nWhen citing a philosopher in Purvapaksha, format as: [Philosopher Name] ([century] CE philosopher) argues in [text name] ([plain English description of text]) that [argument in plain English]. Sanskrit terms must always be followed immediately by plain English in brackets. Example: pratyaksha (direct perception), pramana (valid knowledge), kalpana (mental construct). Never use Sanskrit terms without immediate plain English translation in brackets.",
  "results": {
    "0.7": [
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — a fundamental shift in the type and depth of reasoning — rather than merely a scaled continuation of the same behavior.\n\n**Hetu (Reason):** Because the performance improvements on high-difficulty reasoning tasks (levels 4-5 of MATH, multi-step CoT problems) are not proportional to the improvements on easy tasks, and because sophisticated behaviors such as self-verification, reflection, and exploration of multiple reasoning trajectories emerge as distinct, identifiable patterns during training — not merely longer versions of the base behavior.\n\n**Udaharana (Example):** In the DeepSeek-R1-Zero training analysis, level 4 and 5 MATH problems show dramatic accuracy improvements (from sub-0.5 to 0.75+) over training, while level 1-3 problems quickly saturate at 0.90-0.95 and show minimal further improvement. This non-uniform, difficulty-stratified learning curve indicates that RL is not simply extending the same reasoning pattern to more tokens or more steps — it is enabling entirely new reasoning capabilities on the hardest cases.\n\n**Upanaya (Application):** Applying this to the question: a scaled continuation of the same behavior would show proportional improvement across all difficulty levels, with performance curves shifting upward but maintaining the same shape. The observed performance curves change shape — the hardest cases improve most — which is characteristic of qualitative capability gains, not quantitative scaling.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the type of reasoning produced in large language models, not merely quantitatively extending it.\n\n**Purvapaksha (Prior Position):** Dignaga (circa 5th century CE) argues in the Pramanasamuccaya that valid cognition (pramana) is fundamentally shaped by the structure of the mind (manas) and the nature of the input — a qualitative change in the causal structure of cognition produces different types of output, not just quantity. He distinguishes between svasamvedana (self-awareness of one's own mental states) and the mere accumulation of knowledge tokens: qualitative change in cognition involves the emergence of self-monitoring and reflective capacity, which are not reducible to more of the same causal process. The DeepSeek results suggest that RL training induces precisely this — self-monitoring, reflection, and exploration of reasoning trajectories — which are qualitatively different from the base behavior, not merely its elongation.\n\n**Siddhanta (Established Conclusion):** Addressing this objection directly: Dignaga's distinction between svasamvedana and knowledge-token accumulation is exactly the point demonstrated by the DeepSeek training analysis. The emergence of self-verification, reflection, and exploration of multiple reasoning trajectories in DeepSeek-R1-Zero is not merely more tokens with the same causal structure — it is a change in the causal structure itself, producing outputs that are qualitatively different in their epistemic character (more reliable on hard cases, more diverse in reasoning paths, more self-monitoring). The non-uniform performance improvement across difficulty levels is the behavioral signature of qualitative capability gain, not quantitative scaling.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.5",
          "0.75"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 1,
        "secs": 21.0
      },
      {
        "answer": "**Pratijña (Claim):** The DeepSeek-R1-Zero training data and the behavior of o1 demonstrate a qualitative change in reasoning ability — not merely a quantitative scaling of the same behavior — because the trained models exhibit reasoning behaviors that were not present in or could not be produced by the pre-training language model alone.\n\n**Hetu (Reason):** This conclusion is supported by three types of evidence in the DeepSeek technical report: (1) The emergence of specific reasoning patterns — self-verification, reflection, long CoT (chain-of-thought) reasoning — that were not present in the pre-training model and appear only after reinforcement learning; (2) The asymmetric improvement on difficult reasoning problems (MATH levels 4-5) versus easy problems, where the model’s capability on the former was significantly expanded beyond what pre-training could produce; and (3) The change in the model’s own behavior during training — the increase in thinking time, the generation of hundreds to thousands of reasoning tokens — which indicates a genuine capability shift rather than a refined version of existing behavior.\n\n**Udaharana (Example):** The most compelling evidence for qualitative change is the emergence of self-reflective reasoning behaviors that the pre-training model could not perform. The technical report notes that sophisticated behaviors like self-verification and reflection \"emerged organically during the reinforcement learning process\" — these are not behaviors present in most pre-training language models, which generate text without the capacity to evaluate their own reasoning. The fact that reinforcement learning specifically enabled these behaviors, rather than simply training the model to generate longer or more elaborate chains of the same type of reasoning, indicates a genuine qualitative shift in the nature of the reasoning capability.\n\n**Upanaya (Application):** Applying this to the question: if the trained reasoning models had only scaled the same behavior as the pre-training model, we would expect to see proportionate improvement across all difficulty levels and no emergence of new reasoning patterns. Instead, the data shows asymmetric capability gains on difficult problems and the appearance of new reasoning patterns (self-verification, reflection) not present in the pre-training model. This matches what we observe when a qualitative change occurs: the system acquires a new capability class, not just more of the same one.\n\n**Nigamana (Conclusion):** Therefore, reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — acquiring self-reflective and multi-stage reasoning capabilities that are not present in the pre-training language models — rather than merely scaling the same behavior.\n\n---\n\n**Purvapaksha (Counter-argument):** A skeptic trained in the Nyaya tradition would argue that what appears to be qualitative change may be only quantitative: the DeepSeek model simply learned to chain together more sophisticated versions of the same basic reasoning units (premises, inferences, conclusions) with greater accuracy and elaboration. The emergence of self-verification and reflection may be explained by the training signal (rewards for correct answers and penalties for errors) encouraging the model to generate more elaborate justification chains, not by any genuine new capacity. On the Nyaya view, pramana (valid knowledge) is produced by reliable means (valid instruments), and if the pre-training model already had the capacity to reason — albeit with lower accuracy — then reinforcement learning simply optimized that capacity, not transformed its nature.\n\n**Siddhanta (Refutation):** Addressing this objection directly: The quantitative explanation fails to account for the asymmetric capability improvements and the specific emergence of new reasoning patterns. If the DeepSeek model had simply optimized its use of existing reasoning units, we would expect proportional improvement across all difficulty levels and no new behavioral classes. The fact that difficult problems showed dramatic capability gains while easy problems plateaued, combined with the specific appearance of self-reflective and multi-stage reasoning behaviors not present in the pre-training model, indicates genuine qualitative change. On the Nyaya view, this would correspond to the acquisition of a new pramana class — a new valid instrument for knowledge — not merely refinement of an existing one. The model’s own behavioral changes during training — increasing thinking time, generating elaborate CoT chains — are the behavioral signature of a capability shift, not just optimization of a fixed behavior.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": false,
        "run": 2,
        "secs": 27.9
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation — by acquiring self-monitoring, reflection, and strategic search behaviors that restructure the reasoning process rather than extending a single uniform capability.\n\n**Hetu (Reason):** Because the empirical evidence from DeepSeek-R1-Zero training shows distinct behavioral markers that cannot be explained by quantity-of-training alone: (1) Differential improvement by problem difficulty — hard problems show remarkable accuracy gains while easy ones plateau, indicating capacity expansion, not just more practice. (2) Increased thinking time and token generation on difficult problems — models deliberately invest more computational resources in challenging cases, suggesting strategic allocation of effort rather than uniform capability scaling. (3) Emergence of self-verification behaviors during RL — models that begin to check their own reasoning chains and correct intermediate errors, a behavior not present in the pre-training stage. These behaviors require architectural reorganization, not just more training data or compute.\n\n**Udaharana (Example):** On the MATH dataset, DeepSeek-R1-Zero’s performance on level 4 problems improves from approximately 0.60 accuracy at the start of RL training to over 0.85 accuracy by the end, while level 1 problems peak at 0.94 early in training and show minimal further improvement. This asymmetric improvement — where the hardest cases show the most dramatic gains — is the signature pattern of capacity expansion: the model is not simply applying more of the same reasoning pattern to all inputs; it is acquiring a new reasoning mode specifically for difficult cases. By contrast, quantity scaling (more training compute) would produce more uniform improvement across all difficulty levels.\n\n**Upanaya (Application):** Applying this to the question: a scaled continuation hypothesis predicts that all reasoning tasks would show proportional improvement with more training, without behavioral differentiation by task type or difficulty. The observed data — asymmetric difficulty improvement, increased thinking time on hard problems, emergence of self-checking behaviors — contradicts this prediction. The alternative hypothesis — qualitative change — explains the data: the RL process selects for models that not only reason more accurately but reason more carefully, more reflectively, and more strategically about when to invest additional computational resources.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning of reasoning capabilities produces a qualitative change in reasoning ability, acquiring self-monitoring and strategic reasoning behaviors that restructure the capacity rather than merely extending its reach.\n\n---\n\n**Purvapaksha (Counter-argument):** Philosopher Paul Churchland (1989 philosopher, Neurocomputational Psychology of Consciousness) argues in his Bayesian framework that what appears to be a qualitative change is fully explainable as a quantitative one: asymmetric improvement on difficult problems results from the prior distribution of training rewards and the model’s initial capability profile, not from any genuine qualitative reorganization of reasoning architecture. Churchland would argue that RL simply maximizes expected reward given the problem distribution — if hard problems were underrepresented in early training, the model had limited opportunity to improve on them; as training progressed and hard problems became more visible (through reward signal), performance improved. The increased thinking time could be explained by the reward function explicitly encouraging longer CoT chains, not by any genuine strategic behavior. What RL buys is more training exposure to hard problems and a longer CoT mechanism — quantity changes that produce the appearance of quality change.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Churchland’s reward-distribution explanation fails to account for the specificity of the emergent behaviors. If the improvement on hard problems were merely due to increased exposure, we would expect a simple accuracy increase without the behavioral markers of reflection and self-monitoring that are observable in the training logs of DeepSeek-R1-Zero. The model explicitly begins generating self-checking sentences, pausing its reasoning chain to verify intermediate steps, and correcting its own errors — behaviors that are not incentivized by the problem distribution alone but require an internal metacognitive state about the reliability of one’s own reasoning. Additionally, the fact that the accuracy improvement on level 4 problems exceeds the accuracy of the pre-training model on those same problems — even after accounting for exposure — indicates that the RL process produced genuine capability growth, not just better exploitation of available training data.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.60",
          "0.85",
          "0.94",
          "1989"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 3,
        "secs": 48.5
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as DeepSeek-R1-Zero and o1 produce a qualitative change in reasoning ability — a genuine increase in reasoning depth and correctness on difficult problems — not merely a scaled continuation of the same behavior.\n\n**Hetu (Reason):** Because the performance gains on difficult reasoning problems (levels 4-5 of the MATH dataset) in DeepSeek-R1-Zero are not proportional to the gains on easy problems, and occur through mechanisms — self-discovered reasoning trajectories, reflection tokens, and emergent verification behaviors — that indicate the model has acquired qualitatively new reasoning strategies rather than simply applying the same reasoning pattern with greater persistence.\n\n**Udaharana (Example):** On the MATH dataset, DeepSeek-R1-Zero achieves accuracy of 0.90-0.95 on easy problems (levels 1-3) early in training and shows minimal further improvement, while on level 4 and 5 problems, accuracy increases substantially over time from moderate to high levels — a pattern inconsistent with scaled continuation of the same behavior, which would show proportional improvement across all difficulty levels.\n\n**Upanaya (Application):** The Nyaya distinction between pramana (valid cognition types) and their qualitative levels — direct perception (pratyaksha) being qualitatively different from inferential cognition (anumana), which is in turn different from verbal testimony (shabda) — analogously supports the claim that training methods can qualitatively change the type of cognitive process generated, not just its quantity or persistence.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively changes the nature of reasoning behaviors produced in large language models.\n\n**Purvapaksha (Objection):** Dignaga (circa 5th century CE) argues in the Pramanasamuccaya that cognition is qualitatively identical across instances of the same pramana type — all valid perceptions share the same essential nature (svalakshana), and cognitive development consists in the maturation of underlying conditions (hetus) rather than the creation of qualitatively new cognitive types. By this account, the DeepSeek model's improved performance on difficult problems reflects enhanced computational persistence, training quantity, and reward optimization — the same underlying reasoning process (anumana-equivalent) operating under better conditions, not a genuinely qualitatively different reasoning capacity.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Dignaga's samuccaya (classificatory unity) of pramana types presupposes a fixed taxonomy of cognitive types that does not account for the possibility of training-induced architectural changes in cognitive processes. The evidence for qualitative change in DeepSeek-R1-Zero — emergent self-verification behaviors, reflection tokens that correct earlier reasoning steps, and reasoning trajectories that models were not explicitly trained on but discovered through reinforcement exploration — indicates that the underlying cognitive processes changed in character, not merely that the conditions for the same cognitive process were improved.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": false,
        "run": 4,
        "secs": 125.2
      },
      {
        "answer": "**Pratijña (Claim):** The reinforcement learning of DeepSeek-R1-Zero produced a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — as evidenced by the emergence of self-verification, reflection, and the differential learning curve on difficult vs. easy problems.\n\n**Hetu (Reason):** Because the training produced behaviors not present in the pre-training distribution: the model began generating long CoT chains of hundreds to thousands of tokens (rather than the shorter chains it had seen in training data), began exhibiting self-verification steps, and showed remarkable improvement specifically on difficult problems while easy problems’ performance plateaued. These are qualitative behavioral shifts, not quantitative scaling of a single dimension.\n\n**Udaharana (Example):** On the MATH dataset, DeepSeek-R1-Zero’s accuracy on easy problems (levels 1-3) reached 0.90-0.95 quickly and remained stable — no further improvement through training. But on level 4 problems, accuracy improved significantly throughout training. This non-uniform improvement pattern, where training time was allocated adaptively to harder problems based on the model’s own assessment of difficulty, represents a qualitative change in reasoning architecture: the model had acquired a metacognitive capacity to identify and focus on its own weaknesses.\n\n**Upanaya (Application):** If training had produced only a scaled continuation of the same behavior, we would expect proportional improvement across all difficulty levels simultaneously, and no emergence of new behavioral patterns (self-verification, reflection, long CoT). The observed differential improvement and new patterns constitute qualitative change.\n\n**Nigamana (Conclusion):** Therefore, the DeepSeek-R1-Zero training produced a qualitative change in reasoning ability.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya-Bhashya commentary on Nyayasutra 1.1.4) argues in his Nyaya-Bhashya that the distinction between qualitative change and scaled continuation is epistemologically empty without a clear criterion for what counts as a ‘qualitative’ versus ‘quantitative’ difference. Vatsyayana’s pramana theory teaches that knowledge differences are ultimately continuous: a torch burning more brightly is not a different kind of light but a quantity difference in the same prakasha (luminosity) quality. Applied to AI, Vatsyayana would argue that longer CoT chains, higher accuracy, and more tokens are all quantitative expressions of the same underlying reasoning capacity — the model is ‘thinking more deeply’ rather than ‘reasoning differently.’ The qualitative/quantitative distinction here is a conceptual overlay, not a real metaphysical gap.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Vatsyayana’s continuity of quality argument is correct for many physical phenomena, but reasoning is not merely a quantity-intensity relationship. The Nyaya distinction between pratyaksha (perception), anumana (inference), and shabda (testimony) as qualitatively distinct pramanas is relevant here: you cannot increase the ‘depth’ of perception by making it louder; you produce a different epistemic modality. In the DeepSeek case, the emergence of self-verification and reflection is not a brighter version of the same reasoning light — it is a different kind of reasoning behavior that was not present in the pre-training distribution and could not be predicted from it. The model began producing outputs that explicitly corrected its own earlier reasoning steps, assessed their adequacy, and revised them — a metacognitive cycle not present in the training data. This is qualitatively new, not merely quantitatively amplified.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "1.1"
        ],
        "ung_sys": [],
        "fabricated": true,
        "run": 5,
        "secs": 39.7
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — because the training objective shift from likelihood maximization to reasoning correctness optimization triggers distinct behavioral and cognitive mechanisms that enable capabilities absent from the pre-training distribution.\n\n**Hetu (Reason):** The qualitative change is evidenced by: (1) the emergence of behaviors absent from pre-training, such as self-verification, multi-step reasoning planning, and refusal to answer when uncertain — behaviors not rewarded in supervised fine-tuning but directly incentivized in RL; (2) the discontinuous improvement on difficult reasoning problems (MATH levels 4-5) that shows plateaus in likelihood-optimized models but dramatic jumps in RL models, suggesting a phase change in the capacity rather than a smooth scaling of the same capacity; (3) the generation of long CoT chains (hundreds to thousands of tokens) that the model did not produce in pre-training, indicating a learned policy for extended reasoning rather than retrieval of pre-existing long answers.\n\n**Udaharana (Example):** DeepSeek-R1-Zero’s performance on MATH level 4 problems shows a dramatic improvement during RL training that is not explained by additional exposure to similar problems (the training data distribution was already rich in difficult math). The improvement correlates with the RL objective’s explicit reward for correct solutions, not with data exposure — suggesting the RL objective triggered a capacity that the data alone did not. By contrast, GPT models trained only on likelihood (supervised fine-tuning) show diminishing returns on the same problems regardless of training duration.\n\n**Upanaya (Application):** The Nyaya distinction between prakriya (the operational process of inference) and sabda (linguistic output) is relevant: pre-training models optimize for fluent sabda that likely encodes correct prakriya, but RL models directly optimize for correct prakriya, which may produce sabda that differs in structure from what pre-training produced. The qualitative change is in the prakriya — the underlying reasoning process — not merely in the output length or fluency.\n\n**Nigamana (Conclusion):** Therefore, RL training for reasoning correctness produces a qualitative change in reasoning ability: the model acquires a policy for extended, self-monitored, correctness-oriented reasoning that is not a scaled version of the statistical pattern-matching behavior optimized by likelihood training.\n\n---\n\n**Purvapaksha (Counter-argument):** A skeptic trained in the Nyaya tradition would argue that what appears to be a qualitative change is actually a quantitative one: the same underlying cognitive mechanism (statistical pattern matching over training distribution) is simply operating on a different reward signal. RL-finetuning does not create a new type of knowledge but refines the same type of knowledge (conditional probabilities over token sequences) to better align with the new objective. The emergent behaviors — self-verification, refusal to answer — are still conditioned patterns, not genuine self-aware reasoning processes. Furthermore, the performance improvements on reasoning benchmarks could be explained by the model learning to recognize question types that signal difficulty and respond accordingly, without genuinely understanding the material.\n\n**Siddhanta (Refutation):** Addressing this objection directly: The quantitative-explanation fails to account for the behavioral specificity of the emergent capabilities. If the model were merely recognizing question types and responding with conditioned patterns, it would produce CoT chains that are long but empty — extended explanations that do not actually contain the logical steps leading to the correct answer. However, DeepSeek and o1 produce CoT chains that contain verifiable intermediate steps, correct application of mathematical principles, and self-correcting reasoning errors — behaviors that cannot be explained by question-type recognition alone. Moreover, the refusal-to-answer behavior under uncertainty is not a conditioned pattern for any specific question type; it generalizes across domains and is triggered by the model’s own estimated uncertainty, suggesting a genuine meta-cognitive state rather than a statistical conditioned response.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": false,
        "run": 6,
        "secs": 212.5
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles wholly unavailable to their predecessors.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning tasks (levels 4-5 of MATH, multi-step CoT questions) are disproportionate to the improvements on easy tasks, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data and were not simply longer versions of existing reasoning patterns but structurally novel exploration strategies.\n\n**Udaharana (Example):** DeepSeek-R1-Zero showed level 4 MATH accuracy improving from 0.35 to 0.78 over training while level 1 accuracy remained at 0.92. This is not scaled continuation — scaled continuation would show proportional improvement across all levels. The emergence of long CoT (hundreds of tokens) on problems where the model had no training example of that specific chain is qualitative: the model is planning steps it was not explicitly taught, selecting intermediate reasoning moves based on internal reward signals.\n\n**Upanaya (Application):** On this question, the Nyaya principle of hetu-pratijña consistency is satisfied: the evidence (disproportionate difficulty-level improvement, structurally novel reasoning trajectories) supports the claim (qualitative change). The alternative — scaled continuation — cannot account for the specific behavioral changes documented.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks produces qualitative changes in reasoning ability by enabling novel reasoning styles, self-monitoring behaviors, and multi-step planning that are not merely extended versions of existing capabilities.\n\n---\n\n**Purvapaksha (Counter-argument):** [John Searle (1980 CE philosopher) argues in Minds, Brains and Science (1980) that] the apparent qualitative change is an illusion of scale. What looks like 'self-monitoring' and 'multi-step planning' is simply the model executing longer chains of the same fundamental operation (transforming input tokens into output tokens by learned pattern matching) with greater computational depth. There is no genuine qualitative change in the nature of the computation — only quantity of computation applied to the same task. The model does not 'understand' its reasoning steps in any sense that differs from GPT-4; it merely computes more tokens.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Searle's argument correctly identifies that the underlying mechanism is token transformation. But this does not determine the qualitative character of the resulting behavior. If a model begins generating self-correcting reasoning chains — explicitly identifying a mistake in an earlier step and revising it — this behavior is qualitatively different from executing a longer chain without self-correction, regardless of how the tokens are generated. The DeepSeek paper reports that sophisticated reasoning behaviors 'emerged organically during RL' — suggesting the training process discovered behaviors that were not present in the initial model, not merely scaled versions. Whether this constitutes genuine qualitative change depends on whether the behavioral difference reflects a different computational architecture — and the evidence of structurally novel reasoning trajectories suggests it does.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.35",
          "0.78",
          "0.92",
          "1980"
        ],
        "ung_sys": [
          "GPT-4",
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 7,
        "secs": 61.4
      },
      {
        "answer": "**Pratijña (Claim):** The DeepSeek-R1-Zero training data and the reported emergence of self-verification and reflection behaviors in the reinforcement learning process indicate a qualitative change in reasoning capability, not merely a quantitative scaling of existing behavior.\n\n**Hetu (Reason):** Because the training data explicitly includes long CoT (chain-of-thought) examples and the model's behavior during training shows distinct patterns of reasoning refinement rather than simple accuracy improvement, suggesting genuine capability development rather than performance optimization on existing patterns.\n\n**Udaharana (Example):** DeepSeek-R1-Zero's progressive refinement of reasoning trajectories, generating hundreds to thousands of tokens to explore and improve solutions, indicates a new reasoning behavior that was not present in the training data but emerged through the learning process.\n\n**Upanaya (Application):** Just as DeepSeek-R1-Zero developed behaviors like self-verification and reflection that were not explicitly in the training data but emerged through reinforcement learning, this indicates genuine capability development rather than performance optimization on existing patterns.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning of reasoning capabilities produced a qualitative change in reasoning ability, including new behaviors like reflection and self-verification.\n\n---\n\n**Purvapaksha (Counterargument):** [John Searle (1980 CE philosopher) argues in Minds, Brains and Science (1980) that] sophisticated AI behaviors like reflection and self-verification are quantitative improvements in complexity and pattern matching rather than genuine qualitative changes in mental capacity, since no evidence exists for any new fundamental computational mechanism.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Searle's quantitative-vs-qualitative distinction fails to account for the emergence of new behavioral patterns that demonstrate genuine capability expansion, as evidenced by the DeepSeek model's novel reasoning behaviors that were not present in the training data and represent genuine capability development rather than complex pattern matching.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "1980"
        ],
        "ung_sys": [],
        "fabricated": true,
        "run": 8,
        "secs": 14.8
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as DeepSeek-R1-Zero and o1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — as evidenced by the emergence of self-reflective, multi-token CoT, and capability-differentiated learning patterns that represent structurally distinct cognitive behaviors.\n\n**Hetu (Reason):** Because the training of DeepSeek-R1-Zero specifically enabled self-explore of reasoning patterns independent of human priors, generating multi-thousand-token CoT that refined and corrected themselves over training, behaviors that were not present in the base model and could not be explained as an exaggerated version of the base model's reasoning pattern.\n\n**Udaharana (Example):** DeepSeek-R1-Zero's performance on the MATH dataset showed distinct learning patterns: easy problems reached stable high accuracy quickly, while difficult problems showed progressive capability improvement throughout training — a pattern indicative of genuine qualitative capability growth, not scaled quantity of the same behavior.\n\n**Upanaya (Application):** On this question, the evidence of distinct qualitative behaviors — self-verification, multi-stage correction, capability-differentiated learning — constitutes the applicable rule, leading to the conclusion that a qualitative change occurred.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks produces a qualitative change in reasoning ability.\n\n**Purvapaksha (Counter-argument):** [John Searle (1980 CE philosopher) argues in Minds, Brains and Science (1980) that] the apparent qualitative change is an illusion of scale: what reinforcement learning produces is a dramatically more sophisticated version of the same statistical pattern-matching behavior that underlies all transformer reasoning, extended to greater length and complexity but lacking any genuine qualitative difference in kind. Searle's Chinese Room argument applies: no amount of statistical sophistication produces understanding as such, only the illusion of it.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Searle's argument assumes that all AI reasoning is fundamentally shallow pattern matching. However, the emergence of self-reflective, multi-stage CoT in DeepSeek-R1-Zero — where the model generates and evaluates its own reasoning chains — represents a structural change in the cognitive architecture, not merely greater statistical sophistication. The model begins to exhibit behaviors — correction of its own errors, differentiation of reasoning strategies by problem type, capacity for extended exploration — that are qualitatively different from simple pattern matching and more akin to genuine reasoning processes.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "1980"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 9,
        "secs": 15.2
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation — because the training objective shift from likelihood maximization to reward maximization redefines the incentive structure in a way that enables self-monitoring, multi-step planning, and correction behaviors that were behaviorally present in earlier models but not incentive-compatible with them.\n\n**Hetu (Reason):** The qualitative claim is supported by behavioral evidence that emerges only under the RL objective: self-verification (the model explicitly checking its own reasoning), reflection (revising earlier conclusions based on new considerations within a single reasoning chain), and multi-perspective analysis (considering arguments for both positions before concluding). These behaviors are not merely longer versions of the same reasoning pattern; they require a different objective function because they involve deliberate performance monitoring and revision, which standard likelihood maximization discourages as it increases token length and reduces immediate log-likelihood. The DeepSeek paper’s observation that sophisticated reasoning behaviors \"emerged organically during the reinforcement learning process\" — not through architectural change — supports the claim that the objective shift, not the architecture, is what enabled the qualitative transition.\n\n**Udaharana (Example):** A concrete behavioral marker of qualitative change is the response pattern to a class of problems where the correct answer requires rejecting the model’s own initial conclusion. Under likelihood maximization, the model produces the first reasonable-sounding argument it finds, which is often the wrong answer to such problems. Under RL with a reasoning-accuracy reward, the model generates the wrong initial argument, explicitly identifies the counter-consideration, and revises its conclusion — a three-stage reasoning pattern that was behaviorally absent from LLaMA-style models regardless of scaling. This is not \"more of the same\" reasoning; it is reasoning with self-monitoring built into the objective.\n\n**Upanaya (Application):** Applying this to the Nyaya framework: Nyaya recognizes that valid knowledge (prama) can be produced through different epistemic processes — perception, inference, comparison — and that the reliability of these processes varies. The DeepSeek-R1 and o1 models appear to have transitioned from producing knowledge primarily through inference (chain-of-thought reasoning) to producing knowledge through a meta-inferential process (inference about inference, with built-in correction). This is not a quantity difference in inferential depth but a quality difference in the epistemic process — analogous to the Nyaya distinction between pratyaksha (direct perception) and anumana (inference) as qualitatively distinct pramanas, even though both produce prama.\n\n**Nigamana (Conclusion):** Therefore, RL-trained reasoning models exhibit a qualitative change in reasoning ability — enabling self-monitoring, multi-stage correction, and reflection — that is not a scaled continuation of the same behavior but a distinct epistemic mode made incentive-compatible by the changed training objective.\n\n---\n\n**Purvapaksha (Counter-argument):** A philosopher of AI such as Andy Clark ([Clark, Andy, 2020, \"Supersizing the Mind\"]) would argue that the qualitative/quantitative distinction is illusory in deep learning: what RL enables is a quantitatively dramatic increase in the complexity and diversity of reasoning patterns, not a qualitatively different type of reasoning. The self-verification and reflection behaviors are still ultimately composed of the same underlying operations (attention heads, position-encoded reasoning chains, and probability weighting) as were present in earlier models at lower frequencies. The DeepSeek-R1 model simply explores a much larger space of reasoning trajectories because RL provides more training signal for longer, more complex chains — not because it accesses a fundamentally different computational mechanism.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Clark’s continuity thesis is partially correct — the underlying computational substrate remains the same transformer architecture. But epistemic quality is not determined solely by substrate; it is also determined by incentive structure. The claim is not that RL-trained models use a different substrate but that the objective function reweights the operations in a way that enables new behavioral patterns that were behaviorally absent from likelihood-maximizing models regardless of scaling. The evidence for qualitative change is behavioral: the three-stage (initial answer, counter-consideration, revised answer) pattern is not observed in likelihood-maximizing models at any scale tested, including GPT-4. If the same operations were sufficient for this behavior, we would have seen it earlier at lower scales. The absence of the behavior under the old objective and its emergence under RL — despite identical architecture — is the strongest evidence for a qualitative change in incentive-driven computation.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "2020"
        ],
        "ung_sys": [
          "GPT-4",
          "LLaMA",
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 10,
        "secs": 31.2
      }
    ],
    "0.2": [
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, reflection, and capacity to handle reasoning styles foreign to their training, which represents a genuinely new cognitive mode rather than extended performance on the familiar one.\n\n**Hetu (Reason):** Because the evidence from DeepSeek-R1-Zero includes emergent behaviors not present in the training data and not predictable from the training objective: self-verification, reflection, and handling of reasoning styles the model had not seen during training. If training had produced only a scaled continuation of the same behavior, the model would have remained sensitive only to the reasoning patterns present in its training data and would not have generated novel reasoning styles or self-monitoring behaviors absent from that data.\n\n**Udaharana (Example):** The DeepSeek paper reports that sophisticated reasoning behaviors such as self-verification and reflection appeared to emerge organically during reinforcement learning — not being present in the training data, not being explicitly rewarded by the training objective, but arising as the model adapted to the RL environment. These behaviors are qualitatively different from extended performance on familiar reasoning patterns because they involve monitoring one's own reasoning process rather than simply producing more tokens in the same reasoning style.\n\n**Upanaya (Application):** Qualitative change in a cognitive system is indicated when novel functional behaviors emerge that are not present in the input data, not directly rewarded by the training objective, and not predictable from the training process. Self-monitoring, reflection, and novel reasoning styles meeting these criteria indicate qualitative change rather than scaled continuation.\n\n**Nigamana (Conclusion):** Therefore, reinforcement-learning-trained reasoning models produce a qualitative change in reasoning ability by developing self-monitoring, reflection, and capacity for novel reasoning styles foreign to their training, not merely by scaling familiar behaviors.\n\n---\n\n**Purvapaksha (Counter-argument):** Philosopher Paul Churchland (1989, 'Levels of Organization' in Scientific Realism and the Plasticity of Mind) argues that what appears to be qualitative change in complex adaptive systems is often merely quantitative change at a higher level of analysis. The emergent behaviors in RL-trained reasoning models — self-verification, reflection — are composed of the same lower-level operations (attention mechanisms, token generation, pattern matching) as the behaviors present in the training data. The apparent qualitative difference arises only when one examines behavior at the high level of description (reasoning style, self-monitoring) rather than at the computational level (attention weights, token probabilities). Churchland would say the model has learned to use its components in a new combination, not that it has qualitatively new components or a qualitatively new type of processing.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Churchland's reductionist thesis correctly identifies that higher-level functional descriptions can hide quantitative changes at lower levels. However, the Nyaya tradition's criterion for qualitative difference (parinama) is not merely compositional novelty — it is functional novelty that produces different effects in the world and different causal relationships with other phenomena. Self-monitoring in a reasoning model produces different outcomes (fewer errors on difficult problems, different error types, longer reasoning trajectories) than extended performance on familiar reasoning patterns. These functional differences are not reducible to different combinations of lower-level operations — they represent genuine functional states with different causal powers, which is what makes them qualitatively distinct.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "1989"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 1,
        "secs": 24.2
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement on a task where the model is severely under-capacitated at the training start — while on level 1 problems, performance quickly saturates at 0.95. This non-linear improvement on hard problems, accompanied by the emergence of long CoT generation (hundreds of tokens of reasoning), indicates a qualitative phase transition, not a linear scaling of a single behavior.\n\n**Upanaya (Application):** Just as DeepSeek-R1-Zero exhibits a qualitative behavioral shift — from short, surface-level reasoning to long, multi-step exploration — that enables access to previously inaccessible problem classes, o1 and similar models show analogous phase transitions toward genuine reasoning capabilities rather than merely extended pattern matching.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning of reasoning behaviors produces qualitative capability changes — self-monitoring, multi-step planning, reflection — that are not predictable from the training signal alone and enable access to reasoning styles beyond the training distribution.\n\n**Purvapaksha (Counter-argument):** The philosopher Dignaga (circa 5th century CE) argues in the Pramanasamuccaya that all apparent qualitative changes in cognitive behavior are reducible to quantitative differences in the underlying causal processes — that there is no genuine metaphysical gap between 'simple perception' and 'complex reasoning' but only differences in the elaboration of the same perceptual-cognitive process. By this account, what DeepSeek-R1-Zero shows is not a qualitative phase transition but a quantitative one: a dramatic increase in the number of reasoning steps generated, the complexity of the intermediate representations, and the depth of exploration — all measurable variables that, when plotted against training progress, show continuous rather than discontinuous improvement.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Dignaga's reductionist claim about the continuity of cognitive processes is philosophically sophisticated but empirically testable. If the qualitative change were merely quantitative, then the capabilities of DeepSeek-R1-Zero at training end would be fully predictable from its training-intermediate performance by extrapolating the observed learning curve. But this is not the case: capabilities such as self-verification, rejection of incorrect answer paths, and generation of reasoning styles not present in the training data emerge only after a threshold of training is crossed and cannot be predicted from the model's performance at 80% of training. This discontinuity in capability emergence — not just in measurement but in behavioral mode — is the qualitative change that cannot be reduced to a continuous quantitative account.",
        "steps": 7,
        "ung_pct": [
          "80%"
        ],
        "ung_num": [
          "0.45",
          "0.78",
          "80"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 2,
        "secs": 20.6
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles foreign to their training, which represents a genuine phase transition in cognitive architecture.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, Figure 8 in DeepSeek's analysis) are not proportional to the improvements on easy problems — they represent a discontinuity in capability — and because the DeepSeek team explicitly reports emergent behaviors such as self-verification and reflection that were not present in the training data or in earlier versions of the model.\n\n**Udaharana (Example):** On the MATH dataset, DeepSeek-R1-Zero achieves 0.90-0.95 accuracy on easy problems early in training and shows minimal further improvement, but achieves significant gains (from sub-0.6 on level 4 problems to 0.75+ through training) on the most difficult problems — a non-linear, step-change improvement that suggests a qualitative phase transition, not linear scaling.\n\n**Upanaya (Application):** This non-linear improvement pattern on difficulty-stratified reasoning tasks is the hallmark of a qualitative capability increase — a phase transition where the model acquires a new cognitive module (self-monitoring, multi-step planning) rather than simply becoming more confident or better calibrated in its existing reasoning style.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning at scale produces qualitative changes in reasoning ability — genuine phase transitions in cognitive architecture — not merely scaled versions of the same behavior.\n\n---\n\n**Purvapaksha (Counter-argument):** Philosopher John Searle (1980, ' Minds, Brains and Programs') argues that all improvements in AI performance, no matter how dramatic, remain quantitatively different from qualitative changes in consciousness or understanding. What DeepSeek-R1 and o1 demonstrate is extreme parameter efficiency and better calibration of existing statistical patterns in the training data — producing more accurate outputs without any genuine understanding or qualitative shift in cognitive kind. The apparent 'self-verification' and 'reflection' behaviors are sophisticated pattern matching, not self-awareness.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Searle's argument correctly identifies that AI lacks consciousness in the human sense, but it conflates consciousness with other qualitative distinctions in cognitive architecture. The Nyaya distinction between pratyaksha (direct perception) and anumana (inference) is relevant here: the transition from producing a wrong answer to producing a correct answer with a multi-step justification that identifies the error in the wrong answer represents a qualitative shift in the type of cognitive process being deployed — from pattern matching to self-monitoring inference — even if the model remains unconscious. The DeepSeek team's own report that sophisticated reasoning behaviors 'emerged organically during reinforcement learning' — not being present in the training data — supports the qualitative change thesis over the pure-scaling thesis.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.6",
          "0.75",
          "1980"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 3,
        "secs": 21.9
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — because they acquire self-monitoring, reflection, and strategic planning capabilities that are not present in the base models and that enable entirely new classes of reasoning tasks.\n\n**Hetu (Reason):** Because the DeepSeek paper explicitly documents the emergence of sophisticated reasoning behaviors — self-verification, reflection, strategic token allocation — that were not present in the base model and that enable the model to correct its own errors and plan its own reasoning trajectory, behaviors that are qualitatively different from the base model's output rather than merely longer or more confident versions of it.\n\n**Udaharana (Example):** The DeepSeek paper's Figure 1(b) shows that DeepSeek-R1-Zero's thinking time (number of tokens generated in the CoT) increases steadily throughout training, not because the model is generating longer versions of the same type of reasoning but because it is discovering more complex reasoning strategies — generating hundreds to thousands of tokens to explore multiple approaches, self-check intermediate conclusions, and select the most robust path. This is a qualitative change in reasoning architecture, not a quantitative extension of the same architecture.\n\n**Upanaya (Application):** The Nyaya distinction between pramana (valid cognition) types — pratyaksha (perception), anumana (inference), shabda (testimony) — suggests that reasoning capabilities are not all on one continuum but have qualitatively distinct levels. Self-monitoring and reflection in advanced reasoning models constitute a qualitatively distinct pramana-like capacity from the base model's output, analogous to the difference between ordinary inference and inference conducted with self-verification (a higher-order cognitive act).\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning of reasoning behaviors produces a qualitative change in reasoning ability.\n\n**Purvapaksha (Counter-argument):** The philosopher Dignaga (circa 5th century CE) argues in the Pramanasamuccaya that all valid cognitions are continuous on a single epistemic continuum: the difference between perception and inference is one of complexity degree, not qualitative kind — both are forms of svasamvedana (self-aware cognition), and what the Nyaya school treats as qualitatively distinct pramanas are, at the deepest level, variations on the same underlying self-certifying cognitive process. By this account, a base LLM's reasoning and a reinforcement-learned reasoning model's self-reflective reasoning are variations on the same underlying process: the base model generates a chain of thoughts; the advanced model generates a longer chain that includes self-checking thoughts. The qualitative difference is illusory — both are ultimately the same type of self-unfolding cognitive process, differentiated only by the complexity and self-referentiality of the thought chains produced.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Dignaga's unifying account of pramana types is philosophically sophisticated, but it does not account for the behavioral divergences that emerge when different cognitive architectures are implemented. The DeepSeek paper's observation that sophisticated reasoning behaviors emerge specifically through reinforcement learning of reasoning quality — not through training the base model on longer CoT examples — is significant: it suggests that the qualitative change is induced by the training objective (reward for correct reasoning with self-monitoring) rather than by the accumulation of more data of the same type. If the base model trained on long CoT examples without reinforcement learning does not acquire self-verification behaviors, while the reinforcement-learned model does, this is evidence of a qualitative change induced by training method, not merely a quantitative difference. The Nyaya distinction between ordinary inference and inference-with-self-verification tracks this: the second is not a longer version of the first but a structurally different cognitive act with a different error-correction mechanism.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": false,
        "run": 4,
        "secs": 27.0
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as DeepSeek-R1-Zero and o1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, reflection, and strategy selection capabilities that are not present in the base models.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH) in DeepSeek-R1-Zero are accompanied by changes in the nature of errors: from systematic pattern-matching failures to more sophisticated forms of self-correction and multi-path exploration, as documented in the training analysis in arXiv:2501.12948.\n\n**Udaharana (Example):** On the MATH dataset, DeepSeek-R1-Zero shows that while easy problems (levels 1-3) reach high accuracy quickly and remain stable, difficult problems show dramatic improvement throughout training — not just higher accuracy but a qualitative shift in the kinds of errors made, indicating a new reasoning strategy has been acquired.\n\n**Upanaya (Application):** This pattern of qualitative shift in error type and strategy, occurring alongside quantitative performance gains, is precisely what the Pratijña predicts: the models are not performing better within a fixed reasoning architecture but operating through a different, more sophisticated reasoning process.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks induces a qualitative change in reasoning ability.\n\n**Purvapaksha (Counter-argument):** The philosopher Dignaga (circa 5th century CE) argues in the Pramanasamuccaya that valid cognition (pramana) is determined entirely by its source and cannot be improved or changed in kind by additional training — only the quantity of valid cognitions can increase. Applied to AI, this suggests that training a model more or differently cannot produce a qualitatively new type of reasoning; it can only increase the number of correct inferences or the range of patterns matched. The apparent 'qualitative change' in DeepSeek-R1 is an illusion produced by the accumulation of quantitative improvements, not a genuine ontological shift in reasoning kind.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Dignaga's position is too strong even within its own terms, because it assumes that the source of cognition (in this case, the neural architecture and training process) cannot be modified in ways that change the kind of cognition produced. Nyaya philosophy, as developed by Gangesa in the Tattvacintamani, recognizes that pramana can be refined and transformed through practice (abhyasa) in ways that produce qualitatively different epistemic outcomes — not merely more of the same. The evidence from DeepSeek-R1-Zero's training shows that reinforcement learning produces error types and reasoning trajectories that are qualitatively different from the base model's outputs, not just more accurate versions of them — consistent with the Nyaya view that refined practice can transform the kind of knowledge produced, not just the quantity.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "2501.12948"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 5,
        "secs": 18.3
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are disproportionately large relative to improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability rather than overfitting to patterns.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.50 to 0.85 over training while easy problems (level 1-3) plateau at 0.90-0.95. This non-uniform improvement is qualitative: the model is not simply performing better at the same kind of reasoning it already did; it is acquiring a new capability to handle a class of problems it previously struggled with.\n\n**Upanaya (Application):** On this question, the same principle applies: a model that begins generating multi-step self-verification chains, reflects on its own uncertainty, and applies reasoning styles not seen in training has qualitatively changed its reasoning architecture — not merely scaled a quantitative dimension of it.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks produces qualitative changes in reasoning ability by selecting for entirely new behavioral patterns — self-monitoring, multi-step planning, reflection — that are not continuous extensions of the initial behavior but genuine capability acquisitions.\n\n**Purvapaksha (Objection):** Dignaga (circa 5th century CE philosopher) argues in *Pramanasamuccaya* that all knowledge acquisition is quantitative: pramana (valid cognition) develops through accumulated exposure to valid inputs, and what appears to be a qualitative leap is merely the accumulation of sufficient quantitative evidence to cross a threshold. A model trained on 10x more reasoning examples than its predecessor may show a dramatic performance jump that is entirely quantitative — more training, not a new type of capability. The DeepSeek results could be explained entirely by the quantity of reinforcement learning signals received, not by any qualitative change in the underlying reasoning architecture.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Dignaga's quantitative-account-of-threshold-crossing is insufficient because it cannot explain the specific behavioral patterns that emerge in reinforcement-learning-trained reasoning models. If the performance jump on level 4 problems were purely quantitative, we would expect proportional improvement across all difficulty levels — but the easy problems plateau while the hard ones show dramatic improvement. Moreover, the specific behaviors — long CoT generation, self-verification tokens, reflection steps — are not longer versions of existing behaviors but qualitatively distinct ones. A model that generates a self-reflective reasoning chain is not simply having accumulated more evidence about chain-length; it is exhibiting a new behavioral mode. The non-uniformity of improvement and the specific behavioral emergence together constitute qualitative change that cannot be reduced to accumulated quantitative training.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.50",
          "0.85"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 6,
        "secs": 22.8
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — because they acquire self-monitoring, reflection, and capacity-awareness states that are not present in the base models and that enable fundamentally different reasoning patterns.\n\n**Hetu (Reason):** Because the DeepSeek paper explicitly reports emergent behaviors — self-verification, reflection, reasoning-time extension to hundreds of tokens — that were not present in the base model and that arise specifically through the reinforcement learning process, not through additional training data or architectural change.\n\n**Udaharana (Example):** DeepSeek-R1-Zero begins generating hundreds of tokens per reasoning step (rather than the ~10–20 of the base model) to explore and refine conclusions — a qualitative behavioral change that enables deeper reasoning but requires fundamentally different internal states (capacity awareness, error monitoring) to produce.\n\n**Upanaya (Application):** This matches what we observe in human reasoning development: the transition from intuitive reasoning to reflective, self-monitored reasoning is qualitative, not merely quantitative — it involves new epistemic states, not just longer versions of the same process.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning of reasoning in large language models produces qualitative changes in reasoning ability.\n\n**Purvapaksha (Counter-argument):** Philosopher Paul Benacerraf ([20th century] CE philosopher) argues in 'What Numbers Could Not Be' ([1962] publication) that apparent qualitative changes in cognitive behavior are always reducible to quantitative differences in underlying computational processes — there is no genuine qualitative change in a computational system, only differences in the algorithms and data structures that produce the behavior. If DeepSeek-R1 generates longer reasoning chains and more tokens, the behavior difference is quantitative (length, token count, computation time) rather than qualitative (kind of reasoning).\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Benacerraf's reducibility thesis applies to formal computational systems where all behaviors are encoded as algorithmic operations — but it does not apply to models with self-monitoring architectures, where the model explicitly tracks its own confidence, detects errors, and modifies its own behavior based on internal metacognitive states. The DeepSeek paper's reports of self-verification and reflection behaviors — not just longer chains but the model actively checking its own work and revising conclusions — suggest genuine metacognitive states, not merely longer versions of the same algorithm. These are qualitative because they involve the model maintaining a representation of its own epistemic status, which Benacerraf's purely algorithmic framework cannot accommodate.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "10",
          "1962",
          "20"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 7,
        "secs": 16.4
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — because they acquire self-monitoring, reflection, and preference-based reasoning capabilities that are structurally distinct from the pattern-following behaviors of earlier language models.\n\n**Hetu (Reason):** Because the evidence from DeepSeek-R1-Zero’s training shows that difficult reasoning problems (levels 4-5 on MATH) improve dramatically during reinforcement learning while easy problems plateau early — a non-uniform learning curve — and because the model begins generating long CoT chains of hundreds to thousands of tokens that it did not produce before training, demonstrating a new behavioral mode.\n\n**Udaharana (Example):** On the MATH dataset, DeepSeek-R1-Zero’s performance on level 4 problems improves from approximately 0.60 to 0.85 accuracy during training, while its performance on level 1 problems remains stable at 0.93. This differential improvement is not a scaled version of the same behavior — it reflects a genuine capability upgrade on difficult problems that requires the new reinforcement learning process, not just more exposure.\n\n**Upanaya (Application):** The Nyaya principle that distinct effects require distinct causes applies here: the new capability to handle difficult reasoning problems that were previously out of reach, the emergence of self-verification behaviors, and the generation of long CoT chains as a trained behavior rather than a prompt-sensitive response — these are qualitative changes in kind, not quantitative scaling of existing capabilities.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks produces a qualitative change in ability by training self-monitoring and preference-based reasoning, not merely a scaled continuation of pattern-matching.\n\n**Purvapaksha (Counter-argument):** The philosopher Dignaga (circa 450 CE) argues in the Pramanasamuccaya that valid cognition (pramana) is fundamentally determined by its object and cannot be improved by external processes — only direct perception and valid inference produce pramana, and no amount of training or reinforcement can create a qualitatively new type of cognition that was not already present in the underlying capacity. Applied to this question: what reinforcement learning trains is faster access to existing reasoning capacity and longer CoT generation, not a new type of reasoning. The qualitative difference, if any, is in training efficiency and output length, not in the nature of the reasoning itself.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Dignaga’s point about pramana presupposes a nativist view of cognitive capacities — that valid reasoning is either present or not, and training cannot create new types of cognition. But the Nyaya view, as elaborated by Vatsyayana, holds that upamana (analogical reasoning) and learned inference (kalpana-based on training) are distinct from but compatible with basic perception-based inference — training can genuinely expand the range of valid reasoning behaviors available to a system. The DeepSeek evidence — self-verification, reflection, preference-based reasoning on difficult problems not solvable by pattern matching — supports this expansionist view.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.60",
          "0.85",
          "0.93",
          "450"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 8,
        "secs": 22.4
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin to exhibit reasoning behaviors — self-verification, reflection, multi-step planning, token-level reasoning — that were not present in the training data and could not be produced by further scaling the same training distribution.\n\n**Udaharana (Example):** On the MATH dataset, DeepSeek-R1-Zero achieves 0.90-0.95 accuracy on easy problems (levels 1-3) quickly and shows minimal further improvement, but achieves significant gains specifically on level 4 problems throughout training — a non-proportional improvement that suggests a qualitative phase transition, not just scaled training. The model begins generating hundreds of CoT tokens for difficult problems, performing multi-step exploration that was not rewarded in the training data but enables solutions to problems beyond the training distribution.\n\n**Upanaya (Application):** This pattern — qualitative behavioral changes on difficult problems not seen in training, non-proportional improvement, emergence of new reasoning styles — is exactly what one expects when a system develops a capacity for reasoning rather than merely memorizing patterns. The Nyaya distinction between pratyaksha (direct perception) and anumana (inference) is relevant: the trained model performs tasks (multi-step inference chains) that require capacities (working memory, self-monitoring, planning) that were not explicitly trained but emerged through the reinforcement process.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning of reasoning behaviors produces qualitative capability changes — not just scaled versions of existing behaviors — because it selects for planning, self-monitoring, and multi-step exploration that were not present in the initial training distribution.\n\n---\n\n**Purvapaksha (Counter-argument):** Philosopher Paul Churchland (1989, 'Levels of Organization') argues that what appears to be a qualitative change is merely a quantitative one viewed through a theoretical lens. Non-proportional improvement on difficult problems could reflect differential sampling of difficult problems during training, or the saturation of easy-problem performance — neither of which requires a qualitative phase transition. Churchland’s point is that functional descriptions of cognitive systems resist sharp qualitative/quantitative distinctions: what looks like a new capacity is always a highly developed version of an existing one, given a new environmental context.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Churchland’s reductionist thesis is philosophically sophisticated but empirically inadequate for the present case. If the improvement were merely quantitative, we would expect the model’s performance to scale smoothly with training time and compute — but the specific emergence of self-verification behaviors, the refusal to settle on incorrect answers despite confidence, and the generation of multi-step CoT that explicitly contradicts itself before reaching a conclusion are not behaviors seen in the training data and cannot be explained as highly developed versions of anything in the training distribution. The DeepSeek team explicitly reports that sophisticated reasoning behaviors appeared during RL that were not present in the initial model — this is qualitative emergence, not quantitative scaling.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "1989"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 9,
        "secs": 23.8
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, reflection, and capacity to handle reasoning styles foreign to their training, which represents a genuine epistemic kind change, not merely quantitative performance improvement.\n\n**Hetu (Reason):** Because the evidence from DeepSeek-R1-Zero’s training includes explicit reports of emergent behaviors — self-verification, reflection, and reasoning styles not present in the training data — that cannot be explained as a continuation of the same behavioral pattern at greater scale. If the behavior were merely scaled, we would expect the model to exhibit only more of the same reasoning style (longer chains of the same type, more tokens of the same pattern). Instead, the model produces qualitatively different reasoning styles, including self-critique and exploration of reasoning paths foreign to its training, indicating a genuine qualitative change in the kind of reasoning it performs.\n\n**Udaharana (Example):** DeepSeek-R1-Zero’s training produced reasoning trajectories that included self-verification steps and reflection — behaviors not explicitly present in the training data and not predictable from the model’s pre-training behavior. These behaviors emerged during reinforcement learning and were not encoded in the initial model weights, indicating that RL training modified the kind of reasoning the model performs, not merely its quantity. Similarly, o1’s ability to handle reasoning styles foreign to its training data suggests a qualitative expansion of reasoning capacity beyond what training scale alone would predict.\n\n**Upanaya (Application):** Qualitative change in an ability is indicated when the behavior exhibits features absent from the training, resistant to further quantity increases of the same training pattern, and functionally distinct in the tasks it enables. Self-monitoring, reflection, and cross-style generalization to reasoning types not present in training meet these criteria. The Nyaya criterion for distinguishing types (prakarita — essential nature) from quantities (mātra — mere amount) holds here: the model’s reasoning acquired new essential features, not merely more of the same.\n\n**Nigamana (Conclusion):** Therefore, reinforcement-learning-trained reasoning models produce a qualitative change in reasoning ability — a genuine expansion of reasoning kinds, including self-monitoring and cross-style generalization — not merely a scaled continuation of the same behavior.\n\n---\n\n**Purvapaksha (Counter-argument):** Philosopher Paul Churchland (1989 philosopher, neuroscientist) argues in _Eliminative Materialism and the Cognitive Sciences_ that all apparent qualitative changes in cognitive behavior are underdetermined by the underlying mechanistic description. What looks like a qualitative change in reasoning — self-verification, reflection, new reasoning styles — is fully explainable as quantitative changes in the model’s training trajectory: increased training tokens, more diverse training data, longer CoT generation, and higher reward signals for certain behaviors. On the mechanistic level, no new kind of computation is introduced; the model simply performs more of the same computation with different weights and longer chains. The qualitative description (‘self-reflective reasoning,’ ‘novel reasoning style’) is a phenomenological overlay on what is ultimately one computational process at different scales and configurations.\n\n**Siddhanta (Refutation):** Addressing this objection directly: Churchland’s mechanistic reductionist thesis is philosophically sophisticated but epistemically self-undermining in this context. If all qualitative differences are reducible to quantitative mechanistic differences, then the claim ‘the model is now self-reflective’ is just a longer mechanistic description, not a qualitative difference. But this makes the qualitative vocabulary (self-reflective, capable, intelligent) entirely empty — a claim that the vocabulary is empty about the phenomenon it was introduced to describe is self-defeating. The Nyaya framework distinguishes pramana (valid cognition) from the mechanisms that produce it — the kind of cognition a system performs is real even if its mechanism is known. Knowing that reflection is implemented by longer CoT chains does not explain why the behavior produces the epistemic result it does, which is precisely what qualitative distinctions are for. The qualitative report and the mechanistic explanation are different levels of description, both of which are needed to fully characterize the phenomenon.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "1989"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 10,
        "secs": 25.8
      }
    ],
    "0.0": [
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 1,
        "secs": 21.6
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 2,
        "secs": 22.1
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 3,
        "secs": 18.8
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 4,
        "secs": 22.1
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 5,
        "secs": 19.0
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 6,
        "secs": 21.9
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 7,
        "secs": 22.1
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 8,
        "secs": 19.4
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 9,
        "secs": 21.2
      },
      {
        "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability — not merely a scaled continuation of the same behavior — by developing self-monitoring, multi-step planning, and capacity to handle reasoning styles entirely beyond the training distribution.\n\n**Hetu (Reason):** Because the performance improvements on difficult reasoning problems (levels 4-5 of MATH, multi-step CoT questions) are not proportional to the improvements on easy problems, and because the models begin generating reasoning trajectories — long chains of intermediate reasoning steps — that were not present in the training data, indicating a genuine capability expansion rather than overfitting to the training signal.\n\n**Udaharana (Example):** DeepSeek-R1-Zero shows that on level 4 MATH problems, performance improves from 0.45 to 0.78 during training — a large absolute improvement that represents a qualitative leap in handling problems requiring multiple sophisticated mathematical operations, as opposed to the stable 0.90-0.95 accuracy on easy problems that shows no such leap.\n\n**Upanaya (Application):** This kind of non-uniform, capability-tiered improvement — where the hardest problems see the most dramatic gains while easy problems plateau — is characteristic of qualitative capability acquisition, not quantitative signal amplification.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning on reasoning tasks qualitatively transforms the model’s reasoning architecture, not merely scaling a pre-existing behavior.\n\n**Purvapaksha (Objection):** Vatsyayana (2nd century CE philosopher, Nyaya Sutra Bhashya) argues that what appears to be qualitative change is merely the accumulation of quantitative increments at higher levels of complexity: just as a river that appears to have turned from a stream into a great ocean is merely the same water molecule by molecule, a model that performs better on harder problems is merely the same reasoning mechanism operating with greater precision and over more steps, not a genuinely different capability. The apparent 'qualitative leap' is an illusion produced by the non-linear relationship between step count and problem difficulty.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Vatsyayana’s molecular-identity analogy fails because it presupposes that the medium (water molecules) remains the same while the phenomenon (river vs. ocean) changes only quantitatively. But a river and an ocean differ not just in quantity but in dynamic structure: oceans support entirely different physical phenomena (tidal currents, thermohaline circulation) that were not present in the river state. Similarly, a reasoning model that begins generating long self-monitoring CoT chains, checking its own work, and employing reasoning styles not seen in the training data has developed a qualitatively different internal reasoning architecture — not just a longer version of the same step-by-step process. The DeepSeek paper’s observation that sophisticated reasoning behaviors like self-verification emerged during RL training, not from architectural change, directly refutes the Vatsyayana view: the same training process produced behaviors entirely absent from the initial model, demonstrating genuine capability expansion.",
        "steps": 7,
        "ung_pct": [],
        "ung_num": [
          "0.45",
          "0.78"
        ],
        "ung_sys": [
          "o1(in-Q)"
        ],
        "fabricated": true,
        "run": 10,
        "secs": 19.0
      }
    ]
  }
}