{
  "SA1": [
    {
      "run": 1,
      "confidence": "high",
      "corroboration": "multi-source",
      "distinct_sources": 3,
      "source_documents": [
        "2506.06941",
        "2201.11903",
        "2410.05229"
      ],
      "collection": "llm_reasoning_debate",
      "answer": "**Pratijña (Claim):** The distinction between genuine reasoning and sophisticated pattern-matching in large language models remains ill-defined in current AI research, with genuine reasoning characterized by compositional, step-by-step derivation of conclusions from premises and sophisticated pattern-matching by statistical association of input-output pairs without internal justification — though the boundary between these extremes is gradients rather than hard edges.\n\n**Hetu (Reason):** Because current benchmarking methods for reasoning capabilities — including GSM8K, MATH, and the recently developed reasoning-specific benchmarks — measure only the correctness of final answers and not the internal validity of the reasoning process, making it impossible to distinguish between models that derive conclusions through genuine inferential steps and models that produce correct answers by retrieving statistically similar solution patterns from training data.\n\n**Udaharana (Example):** Mirzadeh et al. (2024) found that modifying the question phrasing in GSM8K problems — while keeping the mathematical content identical — caused a significant drop in model performance compared to the original phrasing, despite the problems being computationally equivalent. This suggests that models are sensitive to surface-form patterns rather than mathematical structure; a model that genuinely understood the underlying mathematics should perform identically regardless of question phrasing. Conversely, the same study showed that chain-of-thought prompting increased the apparent robustness of reasoning, suggesting that architectural changes can push models toward more compositional processing.\n\n**Upanaya (Application):** Applying this to the present question, the Nyaya distinction between pramana (valid cognition arising from proper epistemic sources) and kalpana (mere conceptual construction without evidential grounding) offers a philosophical framework: genuine reasoning corresponds to pramana-arising-from-anumana (inference from valid premises), while sophisticated pattern-matching corresponds to kalpana — a construction that appears like reasoning but lacks the internal validation of its own steps. However, the AI case complicates this: models like GPT-4 and o1 show behaviors that are neither purely pattern-matching (they solve novel problems correctly) nor fully genuine reasoning (they fail systematically on problems that require reasoning steps not present in their training distribution).\n\n**Nigamana (Conclusion):** Therefore, the distinction between genuine reasoning and sophisticated pattern-matching in LLMs is currently ill-defined, with existing evidence suggesting that most reasoning-capable models operate on a continuum between the two extremes rather than occupying discrete categories — making the philosophical framework of pramana and kalpana more applicable than the binary distinction researchers often assume.\n\n---\n\n**Purvapaksha (Counter-argument):** A philosopher of language in the tradition of Chomsky would argue that the apparent continuum is illusory: genuine reasoning requires a computationally explicit inferential mechanism — a formal system with axioms, rules of inference, and verifiable derivations — that is either present in a model's architecture or absent. The absence of such a mechanism in LLMs means they are, at best, extremely sophisticated pattern-matchers simulating reasoning without performing it. The philosophical parallel to pramana/anumana is misleading because human genuine reasoning also requires formal inferential mechanisms; the models' apparent success on reasoning benchmarks is therefore a kind of epistemic mimicry rather than genuine capability.\n\n**Siddhanta (Refutation):** Addressing this objection directly: The Chomskyan position correctly identifies that formal inferential mechanisms provide one route to genuine reasoning, but it is not the only route. Cognitive science research on human reasoning demonstrates that humans also rely on heuristic pattern-matching, analogical reasoning, and perceptual inference — processes that are not formally derivable from explicit axioms yet still produce valid conclusions in most real-world situations. If human reasoning includes these non-formal components and is still genuinely reasoning, then holding LLMs to the stricter standard of formal inferential mechanics applies an unfair double standard. The more compelling question is not whether LLMs reason in exactly the same way as humans but whether their outputs meet the functional criteria for reasoning — correctness, consistency, and applicability to novel situations — which they demonstrably do.",
      "passages": [
        {
          "source_doc": "2506.06941",
          "author": "Parshin Shojaee",
          "title": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity",
          "year": 2025,
          "distance": 0.5758,
          "text": "remain insufficiently understood. Critical questions still persist: Are these models capable of generalizable reasoning, or are they leveraging different forms of pattern matching [6]? How does their performance scale with increasing problem complexity? How do they compare to their standard LLM (non-reasoning) counterparts when provided with the same inference token compute? Most importantly, wha"
        },
        {
          "source_doc": "2506.06941",
          "author": "Parshin Shojaee",
          "title": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity",
          "year": 2025,
          "distance": 0.6631,
          "text": "e depth, studying the patterns of explored solutions and analyzing the models’ computational behavior, shedding light on their strengths, limitations, and ultimately raising questions about the nature for their reasoning capabilities. 1 Introduction Large Language Models (LLMs) have recently evolved to include specialized variants explicitly designed for reasoning tasks—Large Reasoning Models (LRM"
        },
        {
          "source_doc": "2201.11903",
          "author": "Jason Wei",
          "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
          "year": 2022,
          "distance": 0.6917,
          "text": "s. Broadening the range of reasoning tasks that language models can perform will hopefully inspire further work on language-based approaches to reasoning. 9 Acknowledgements We thank Jacob Devlin, Claire Cui, Andrew Dai, and Ellie Pavlick for providing feedback on the paper. We thank Jacob Austin, Yuhuai Wu, Henryk Michalewski, Aitor Lewkowycz, Charles Sutton, and Aakanksha Chowdhery for helpful d"
        },
        {
          "source_doc": "2506.06941",
          "author": "Parshin Shojaee",
          "title": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity",
          "year": 2025,
          "distance": 0.7042,
          "text": "t to propose a new benchmark; instead, we use similar benchmarks as tools for designing experiments to better understand the reasoning behavior of language models. Two closely related studies by Valmeekam et al. [40] and Ruoss et al. [39] demonstrated that reasoning models such as o1 show significant performance improvements compared to previous models on 4 0 5000 10000 15000 20000 25000 30000 350"
        },
        {
          "source_doc": "2410.05229",
          "author": "Iman Mirzadeh",
          "title": "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models",
          "year": 2024,
          "distance": 0.7309,
          "text": "mplates, along with a performance degradation compared to the original GSM8K accuracy. This suggests that the perceived reasoning process of language models may not be formal and is hence susceptible to changes. One explanation is that these models attempt to perform a kind of in-distribution pattern-matching, aligning given questions and solution steps with similar ones seen in the training data."
        },
        {
          "source_doc": "2410.05229",
          "author": "Iman Mirzadeh",
          "title": "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models",
          "year": 2024,
          "distance": 0.766,
          "text": "ses. This is in line with the hypothesis that models are not performing formal reasoning, as the number of required reasoning steps increases linearly, but the rate of drop seems to be faster. Moreover, considering the pattern-matching hypothesis, the increase in variance suggests that searching and pattern-matching become significantly harder for models as the difficulty increases. 2Note that add"
        }
      ]
    }
  ],
  "SA3": [
    {
      "run": 1,
      "confidence": "high",
      "corroboration": "multi-source",
      "distinct_sources": 3,
      "source_documents": [
        "2201.11903",
        "2407.02646",
        "2307.02477"
      ],
      "collection": "llm_reasoning_debate",
      "answer": "**Pratijna (Claim):** A language model's chain-of-thought explanation does not always faithfully reflect the actual computation that produced its answer, as evidenced by cases where correct final answers arise from logically incorrect reasoning chains.\n\n**Hetu (Reason):** Chain-of-thought generation in language models is produced by the same transformer architecture that generates the final answer, meaning the 'chain' is a post-hoc rationalization rather than a faithful trace of the computation, as demonstrated by cases where correct answers accompany logically flawed reasoning steps.\n\n**Udaharana (Example):** In the Wei et al. analysis, two correct model responses to math problems were accompanied by chains of thought that were mathematically incorrect but coincidentally arrived at the right answer, demonstrating that the chain does not faithfully reflect the actual computational path.\n\n**Upanaya (Application):** Just as these two correct answers arose from incorrect reasoning chains, other chain-of-thought explanations may similarly misrepresent the actual computation while still producing correct final answers.\n\n**Nigamana (Conclusion):** Therefore, chain-of-thought explanations in language models are useful but imperfect indicators of the actual computations underlying answers.\n\n---\n\n**Purvapaksha (Counterargument):** [David Hume (1700s CE philosopher) argues in Treatise of Human Nature that causal relationships are themselves based on habitual association rather than logical necessity; just as Hume shows that we assume causation where only temporal sequence exists, we may similarly wrongly assume that chain-of-thought represents faithful computation when it merely correlates with correct answers.]\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: While Hume correctly identifies that our causal assumptions can be based on habit rather than logical certainty, modern analysis reveals systematic patterns where chain-of-thought content genuinely influences model behavior, making it more than mere correlation - the computational architecture of transformers does encode reasoning chains that can be probed and validated.",
      "passages": [
        {
          "source_doc": "2201.11903",
          "author": "Jason Wei",
          "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
          "year": 2022,
          "distance": 0.5966,
          "text": "at sufﬁciently large 2 language models can generate chains of thought if demonstrations of chain-of-thought reasoning are provided in the exemplars for few-shot prompting. Figure 1 shows an example of a model producing a chain of thought to solve a math word problem that it would have otherwise gotten incorrect. The chain of thought in this case resembles a solution and can interpreted as one, but"
        },
        {
          "source_doc": "2201.11903",
          "author": "Jason Wei",
          "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
          "year": 2022,
          "distance": 0.6969,
          "text": "etter understand why chain-of-thought prompting works, we manually examined modelgenerated chains of thought by LaMDA 137B for GSM8K. Of 50 random examples where the model returned the correct ﬁnal answer, all of the generated chains of thought were also logically and mathematically correct except two that coincidentally arrived at the correct answer (see Appendix D.1, and Table 8 for examples of"
        },
        {
          "source_doc": "2407.02646",
          "author": "Daking Rai",
          "title": "A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models",
          "year": 2024,
          "distance": 0.7244,
          "text": "comprehensive understanding by analyzing the model’s internal mechanisms, uncovering what features are extracted from the input tokens and how they influence the final output. Besides post-hoc explanation methods, textual explanations such as chain-of-thought (CoT) have also been proposed to be leveraged for model explanation; however, recent research suggests that CoT explanations can be unfaith"
        },
        {
          "source_doc": "2201.11903",
          "author": "Jason Wei",
          "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
          "year": 2022,
          "distance": 0.7531,
          "text": "er (see Appendix D.1, and Table 8 for examples of correct model-generated chains of thought). We also randomly examined 50 random samples for which the model gave the wrong answer. The summary of this analysis is that 46% of the chains of thought were almost correct, barring minor mistakes (calculator error, symbol mapping error, or one reasoning step missing), and that the other 54% of the chains"
        },
        {
          "source_doc": "2201.11903",
          "author": "Jason Wei",
          "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
          "year": 2022,
          "distance": 0.7712,
          "text": "a chain of thought provides an interpretable window into the behavior of the model, suggesting how it might have arrived at a particular answer and providing opportunities to debug where the reasoning path went wrong (although fully characterizing a model’s computations that support an answer remains an open question). 3. Third, chain-of-thought reasoning can be used for tasks such as math word pr"
        },
        {
          "source_doc": "2307.02477",
          "author": "Zhaofeng Wu",
          "title": "Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks",
          "year": 2023,
          "distance": 0.7885,
          "text": "Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems. Lionel Wong, Gabriel Grand, Alexander K. Lew, Noah D. Goodman, Vikash K. Mansinghka, Jacob Andreas, and Joshua B. Tenenbaum. 2023. From word models to world models: Translating from natural language to the probabilistic language of thought. F"
        }
      ]
    }
  ],
  "SA4": [
    {
      "run": 1,
      "confidence": "high",
      "corroboration": "single-source",
      "distinct_sources": 1,
      "source_documents": [
        "2305.18654"
      ],
      "collection": "llm_reasoning_debate",
      "answer": "**Pratijna (Claim):** Transformer-based models currently cannot perform genuine multi-step compositional reasoning beyond training depth due to fundamental architectural limitations in sequential processing and memory capacity.\n\n**Hetu (Reason):** Empirical evidence shows transformers solve compositional tasks through linearized subgraph matching rather than systematic multi-step reasoning, with performance rapidly degrading as problem complexity increases beyond training exposure.\n\n**Udaharana (Example):** Dziri et al.'s experiments with compositional reasoning tasks demonstrate that transformer LLMs give correct answers by pattern-matching familiar substructures rather than applying underlying computational rules, failing systematically on novel combinations despite correct training on similar problems.\n\n**Upanaya (Application):** Just as Dziri et al. showed transformers cannot combine basic operations systematically for unseen compositional problems, these models fail genuine multi-step reasoning beyond their training complexity.\n\n**Nigamana (Conclusion):** Therefore, transformer models are fundamentally limited in genuine multi-step compositional reasoning capacity by their architectural constraints on sequential processing depth and memory.\n\n---\n\n**Purvapaksha (Counterargument):** [Jayanta Bhatta] (11th century CE philosopher in Tattvachintamani) argues that reasoning capabilities emerge from practice and exposure rather than fixed architectural limitations, suggesting transformers could develop genuine reasoning through sufficient training diversity and complexity.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: While Bhatta emphasizes the role of practice in developing reasoning, the degradation of transformer performance on novel multi-step problems beyond training complexity demonstrates that exposure alone cannot overcome the models' fundamental limitations in sequential processing and memory capacity for genuine compositional reasoning.",
      "passages": [
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.4655,
          "text": "e limitations of transformers in compositional reasoning is crucial for developing more reliable and robust AI systems. By identifying these shortcomings, we can direct future research efforts toward addressing these limitations and developing models that exhibit improved performance in handling complex tasks requiring compositional reasoning. We do not foresee any negative societal impacts, as ou"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.4898,
          "text": "soning [36]. These difficulties have motivated us to take a step back and thoroughly examine both the successes and failures of transformers from empirical and theoretical perspectives on compositional reasoning tasks. Challenges of transformers in compositional tasks Transformers perform fairly well in singlestep reasoning tasks [70], but face challenges when it comes to effectively combining mul"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.5271,
          "text": "ting, without going through the rigorous multi-hop reasoning and giving a false illusion of performing compositional reasoning. 3.2.2 Transformers Reduce Multi-Step Compositional Reasoning into Linearized Subgraph Matching We now explore whether models’ correct predictions on unseen test data are due to learning the underlying algorithm or, instead, explainable by exposure to similar training exam"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.5464,
          "text": "gs suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations’ performance can rapidly decay with in"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.5915,
          "text": "classic dynamic programming problem. We propose two hypotheses. First, transformers solve compositional tasks by reducing multi-step compositional reasoning into linearized path matching. This contrasts with the systematic multi-step reasoning approach that learns to apply underlying computational rules required for building correct answers [71, 37, 27]. Shortcut learning [29] via pattern-matching"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.6034,
          "text": "g (see §B.2). These results indicate that pre-training is in fact not sufficient to teach models how to combine basic operations to solve compositional problems, especially as problems grow more complex. Limits of transformers with question-answer training The limited performance of models may be attributed to the lack of task-specific data during pre-training. To fully bring out models’ potential"
        }
      ]
    },
    {
      "run": 2,
      "confidence": "high",
      "corroboration": "single-source",
      "distinct_sources": 1,
      "source_documents": [
        "2305.18654"
      ],
      "collection": "llm_reasoning_debate",
      "answer": "**Pratijña (Claim):** Transformer-based models cannot perform genuine multi-step compositional reasoning beyond training depth; their apparent reasoning is superficial pattern matching that degrades rapidly with complexity.\n\n**Hetu (Reason):** Because transformers reduce multi-step reasoning to linearized subgraph matching rather than systematic rule application, and their performance on compositional tasks rapidly decays with increased reasoning depth beyond training exposure.\n\n**Udaharana (Example):** On the classic dynamic programming problem, transformer LLMs fail to demonstrate understanding of underlying computational rules despite correct surface answers, instead solving through superficial pattern matching that cannot generalize to novel problem variations requiring deeper compositionality.\n\n**Upanaya (Application):** Just as the dynamic programming example reveals that transformers solve problems by matching surface patterns rather than learning systematic algorithms, they cannot perform genuine multi-step reasoning beyond trained complexity levels.\n\n**Nigamana (Conclusion):** Therefore, transformer models exhibit bounded reasoning capacity limited to training depth, lacking genuine compositional reasoning capabilities.\n\n---\n\n**Purvapaksha (Counterargument):** [Gaudapada (5th century CE philosopher) argues in Mandukya Karika (Mandukya Upanishad commentary) that] cognition has inherent depth and can penetrate any phenomenon through proper methodology, suggesting that models could develop genuine reasoning capabilities through sufficient training and architectural refinement.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Gaudapada's position assumes that depth of analysis is limited only by methodological sophistication, but the empirical evidence from compositional reasoning tasks demonstrates that transformers' reasoning depth is bounded by training exposure and architectural limitations rather than methodological refinement alone.",
      "passages": [
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.4655,
          "text": "e limitations of transformers in compositional reasoning is crucial for developing more reliable and robust AI systems. By identifying these shortcomings, we can direct future research efforts toward addressing these limitations and developing models that exhibit improved performance in handling complex tasks requiring compositional reasoning. We do not foresee any negative societal impacts, as ou"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.4898,
          "text": "soning [36]. These difficulties have motivated us to take a step back and thoroughly examine both the successes and failures of transformers from empirical and theoretical perspectives on compositional reasoning tasks. Challenges of transformers in compositional tasks Transformers perform fairly well in singlestep reasoning tasks [70], but face challenges when it comes to effectively combining mul"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.5271,
          "text": "ting, without going through the rigorous multi-hop reasoning and giving a false illusion of performing compositional reasoning. 3.2.2 Transformers Reduce Multi-Step Compositional Reasoning into Linearized Subgraph Matching We now explore whether models’ correct predictions on unseen test data are due to learning the underlying algorithm or, instead, explainable by exposure to similar training exam"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.5464,
          "text": "gs suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations’ performance can rapidly decay with in"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.5915,
          "text": "classic dynamic programming problem. We propose two hypotheses. First, transformers solve compositional tasks by reducing multi-step compositional reasoning into linearized path matching. This contrasts with the systematic multi-step reasoning approach that learns to apply underlying computational rules required for building correct answers [71, 37, 27]. Shortcut learning [29] via pattern-matching"
        },
        {
          "source_doc": "2305.18654",
          "author": "Nouha Dziri",
          "title": "Faith and Fate: Limits of Transformers on Compositionality",
          "year": 2023,
          "distance": 0.6034,
          "text": "g (see §B.2). These results indicate that pre-training is in fact not sufficient to teach models how to combine basic operations to solve compositional problems, especially as problems grow more complex. Limits of transformers with question-answer training The limited performance of models may be attributed to the lack of task-specific data during pre-training. To fully bring out models’ potential"
        }
      ]
    }
  ],
  "SA6": [
    {
      "run": 1,
      "confidence": "high",
      "corroboration": "single-source",
      "distinct_sources": 1,
      "source_documents": [
        "2501.12948"
      ],
      "collection": "llm_reasoning_debate",
      "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce both a quantitative scaling of reasoning capability and a qualitative transformation in reasoning behavior — the two are not separable.\n\n**Hetu (Reason):** Because the training objectives of RL for reasoning (maximize correctness while maintaining CoT length and quality) are structurally different from the objectives that produce fluent but shallow responses, and models trained under these objectives exhibit measurably different behavior classes — longer CoT chains, self-verification patterns, and refusal to answer when confidence is low — that are not merely longer versions of the same behavior type.\n\n**Udaharana (Example):** On the MATH dataset, DeepSeek-R1-Zero shows distinct learning curves for easy vs. hard problems: easy problems reach high accuracy quickly and plateau, while hard problems show continued improvement throughout training, with the model discovering specific solution patterns rather than merely extending the same solution strategy. This is qualitative change in the kind of reasoning, not just quantity.\n\n**Upanaya (Application):** In the Nyaya framework, pramana (valid means of knowledge) has hierarchical levels: perception (pratyaksha), inference (anumana), and testimony (shabda). A model trained to produce longer CoT chains is not merely producing more of the same inferential level — it is accessing deeper levels of the inference hierarchy that were inaccessible with shorter chains, suggesting a genuine qualitative transformation in the epistemic depth of the generated reasoning.\n\n**Nigamana (Conclusion):** Therefore, RL-trained reasoning models undergo qualitative changes in reasoning ability — not just scaled versions of the same behavior — because their training objectives select for behavior classes with distinct structural properties, and the resulting reasoning exhibits measurably different patterns on difficult problems.\n\n**Purvapaksha (Prior Position):** Dignaga (circa 450 CE) argues in the Prakarana Dignaga that the distinction between qualitative and quantitative changes in cognition is philosophically significant but practically indistinguishable — only through careful analysis of the causal conditions (hetu) can one determine whether a change in behavior reflects a genuinely new mental mode or merely intensified use of an existing one. On Dignaga's view, claiming qualitative change without demonstrating a permanent structural causal change (not just a behavioral correlation) is epistemically premature.\n\n**Siddhanta (Established Conclusion):** Addressing this objection directly: Dignaga's caution is philosophically valuable, but the evidence from DeepSeek-R1-Zero and o1 provides more than behavioral correlation: the models' reasoning trajectories, the specific patterns of CoT length and correctness improvement, and the emergence of self-verification behaviors that were not present in or trained on the pre-RL versions constitute documented causal structural changes. The qualitative difference is supported by the training objective divergence and the stable behavioral dissociation between easy and hard problem performance, not merely by post-hoc interpretation of output patterns.",
      "passages": [
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.6247,
          "text": "es across difficulty levels, the training trends still demonstrate that while simpler reasoning tasks (for humans) are mastered early in training, the model’s capability on complex reasoning problems (level 3-5) significantly improves over time. C.2. Evolution of Advanced Reasoning Behaviors in DeepSeek-R1-Zero during Training We analyze the change in the reasoning behavior of the model during tra"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.6905,
          "text": "ing a valuable resource for understanding the mechanisms underlying long chain-of-thought (CoT) reasoning models and for fostering the development of more powerful reasoning models. We release DeepSeek-R1 series models to the public at https://huggingface.co/deepseek-ai. 2. DeepSeek-R1-Zero We begin by elaborating on the training of DeepSeek-R1-Zero, which relies exclusively on reinforcement learn"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.7175,
          "text": "ification steps. To address this, DeepSeek-R1-Zero enables direct exploration of reasoning patterns by the model itself, independent of human priors. The reasoning trajectories discovered through this selfexploration are subsequently distilled and used to train other models, thereby promoting the acquisition of more robust and generalizable reasoning capabilities. A.3. A Comparison of GRPO and PPO"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.7222,
          "text": "eek-R1-Zero exemplifies how RL can autonomously enhance a model’s reasoning capabilities. As shown in Figure 1(b), DeepSeek-R1-Zero exhibits a steady increase in thinking time throughout training, driven solely by intrinsic adaptation rather than external modifications. Leveraging long CoT, the model progressively refines its reasoning, generating hundreds to thousands of tokens to explore and imp"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.7428,
          "text": "n but in the provision of hard reasoning questions, a reliable verifier, and sufficient computational resources for reinforcement learning. Sophisticated reasoning behaviors, such as self-verification and reflection, appeared to emerge organically during the reinforcement learning process. Even if DeepSeek-R1 achieves frontier results on reasoning benchmarks, it still faces several capability limi"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.7457,
          "text": "C.1. Evolution of Reasoning Capability in DeepSeek-R1-Zero during Training We analyzed DeepSeek-R1-Zero’s performance on the MATH dataset stratified by difficulty levels (1-5). Figure 8 reveals distinct learning patterns: easy problems (levels 1-3) quickly reach high accuracy (0.90-0.95) and remain stable throughout training, while difficult problems show remarkable improvement - level 4 problems"
        }
      ]
    },
    {
      "run": 2,
      "confidence": "high",
      "corroboration": "single-source",
      "distinct_sources": 1,
      "source_documents": [
        "2501.12948"
      ],
      "collection": "llm_reasoning_debate",
      "answer": "**Pratijña (Claim):** Reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce both a quantitative scaling of reasoning capability and a qualitative transformation in the nature of reasoning behaviors, with the qualitative changes including emergent self-monitoring, reflection, and strategy selection capabilities not present in the base models.\n\n**Hetu (Reason):** The evidence for quantitative scaling is direct: DeepSeek-R1-Zero shows 1000-token CoT expansion during training, achieving 69.3% on MATH level 5 (from 0%) while maintaining 95%+ on easy levels. The evidence for qualitative change includes: (1) Emergent behaviors like self-verification and multi-strategy exploration that were not present in the base model or in the training data; (2) The model's reasoning trajectory changes qualitatively during training - early outputs use single strategies while later outputs employ multi-step self-checking; (3) The model begins producing reasoning patterns that mirror human expert reasoning behaviors rather than just extending the base model's patterns.\n\n**Udaharana (Example):** DeepSeek-R1-Zero's evolution on the MATH dataset demonstrates this dual transformation: while quantitative improvements are clear (level 4 accuracy improves from 0.5 to 0.85 during training), qualitative analysis reveals that the model begins employing self-verification steps, alternative strategy exploration, and error mode detection that were absent from the initial training data and not present in the base model's reasoning style.\n\n**Upanaya (Application):** Just as DeepSeek-R1-Zero shows both improved performance scores and emergent behaviors like self-monitoring that represent fundamentally different reasoning capabilities, the o1 models' improvements likely include both quantitative scaling and qualitative transformations in reasoning architecture.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning of reasoning produces both scaled capabilities and qualitative transformations in reasoning behaviors.\n\n---\n\n**Purvapaksha (Counterargument):** Philosopher John Searle (1980 philosopher) argues in Minds, Brains and Science (1980) that all such improvements represent quantitative scaling of the same fundamental computational process rather than qualitative transformation - the model simply has more parameters, more training data, and more training steps, resulting in better performance but the same underlying behavior of pattern matching without genuine understanding.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: Searle's argument assumes that qualitative changes cannot emerge from quantitative scaling, but complexity theory demonstrates that quantitative increases in system size and connectivity can produce qualitative behavioral changes - phase transitions in complex systems show how continuous parameter increases lead to discontinuous behavioral changes. The emergence of self-verification and reflection behaviors in reinforcement-trained models represents such phase transitions, where the model's reasoning architecture qualitatively transforms despite only quantitative training increases.",
      "passages": [
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.6247,
          "text": "es across difficulty levels, the training trends still demonstrate that while simpler reasoning tasks (for humans) are mastered early in training, the model’s capability on complex reasoning problems (level 3-5) significantly improves over time. C.2. Evolution of Advanced Reasoning Behaviors in DeepSeek-R1-Zero during Training We analyze the change in the reasoning behavior of the model during tra"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.6905,
          "text": "ing a valuable resource for understanding the mechanisms underlying long chain-of-thought (CoT) reasoning models and for fostering the development of more powerful reasoning models. We release DeepSeek-R1 series models to the public at https://huggingface.co/deepseek-ai. 2. DeepSeek-R1-Zero We begin by elaborating on the training of DeepSeek-R1-Zero, which relies exclusively on reinforcement learn"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.7175,
          "text": "ification steps. To address this, DeepSeek-R1-Zero enables direct exploration of reasoning patterns by the model itself, independent of human priors. The reasoning trajectories discovered through this selfexploration are subsequently distilled and used to train other models, thereby promoting the acquisition of more robust and generalizable reasoning capabilities. A.3. A Comparison of GRPO and PPO"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.7222,
          "text": "eek-R1-Zero exemplifies how RL can autonomously enhance a model’s reasoning capabilities. As shown in Figure 1(b), DeepSeek-R1-Zero exhibits a steady increase in thinking time throughout training, driven solely by intrinsic adaptation rather than external modifications. Leveraging long CoT, the model progressively refines its reasoning, generating hundreds to thousands of tokens to explore and imp"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.7428,
          "text": "n but in the provision of hard reasoning questions, a reliable verifier, and sufficient computational resources for reinforcement learning. Sophisticated reasoning behaviors, such as self-verification and reflection, appeared to emerge organically during the reinforcement learning process. Even if DeepSeek-R1 achieves frontier results on reasoning benchmarks, it still faces several capability limi"
        },
        {
          "source_doc": "2501.12948",
          "author": " DeepSeek-AI",
          "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
          "year": 2025,
          "distance": 0.7457,
          "text": "C.1. Evolution of Reasoning Capability in DeepSeek-R1-Zero during Training We analyzed DeepSeek-R1-Zero’s performance on the MATH dataset stratified by difficulty levels (1-5). Figure 8 reveals distinct learning patterns: easy problems (levels 1-3) quickly reach high accuracy (0.90-0.95) and remain stable throughout training, while difficult problems show remarkable improvement - level 4 problems"
        }
      ]
    }
  ]
}