{
  "SA1": {
    "short": "criterion",
    "question": "What distinguishes genuine reasoning from sophisticated pattern-matching in large language models, and is that distinction well-defined?",
    "confidence": "high",
    "corroboration": "multi-source",
    "distinct_sources": 4,
    "source_documents": [
      "2506.06941",
      "2201.11903",
      "2410.05229",
      "2307.13702"
    ],
    "collection": "llm_reasoning_debate",
    "n_passages": 8,
    "ungrounded_numbers": [],
    "model_inline_confidence": "none",
    "phil_named": false,
    "phil_hits": [],
    "text_named": false,
    "century_named": false,
    "steps": 7,
    "answer": "**Pratijña (Claim):** The distinction between genuine reasoning and sophisticated pattern-matching in large language models remains theoretically contested and empirically under-determined; current evidence supports the hypothesis that LLM reasoning capabilities are primarily pattern-matching over learned solution structures rather than formal deductive reasoning, though the functional similarity at the output level makes the distinction practically significant rather than philosophically clean.\n\n**Hetu (Reason):** Because sophisticated pattern-matching over massive training data can produce output indistinguishable from reasoning without implementing the reasoning process, and because models trained without explicit reasoning mechanisms show comparable performance on reasoning benchmarks when given appropriate patterns, the internal process cannot be inferred from the output alone.\n\n**Udaharana (Example):** Mirzadeh et al. (2024) found that modifying GSM8K questions by replacing their phrasing while preserving the underlying structure caused a significant drop in model performance despite identical mathematical content, demonstrating that models rely on surface-pattern matching rather than mathematical reasoning — a finding that directly contradicts the intuition that reasoning is the mechanism.\n\n**Upanaya (Application):** This finding maps onto the present case: if models fail when the surface pattern changes but the underlying structure remains identical, their performance is determined by pattern recognition of linguistic cues rather than by the reasoning process itself.\n\n**Nigamana (Conclusion):** Therefore, the available evidence supports the pattern-matching hypothesis as the best account of LLM 'reasoning' performance.\n\n**Purvapaksha (Objection):** The pattern-matching hypothesis, if true, would make the design of reasoning-capable LLMs trivial: simply train models on sufficiently diverse solution-pattern pairs, and the model will produce reasoning-like output when faced with similar patterns. If no genuine reasoning mechanism is involved, the additional compute and architectural complexity of specialized reasoning models (such as o1) is unjustified. Furthermore, the hypothesis cannot explain cases where models produce correct reasoning steps that were not present in any training example — a genuine reasoning system could generate novel valid inferences; a pattern-matching system could not. The existence of such cases suggests that at least some reasoning capability is present, undermining the hypothesis entirely.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: The objection correctly identifies that pattern-matching cannot explain novel correct inferences, but this objection proves too much: the same point demonstrates that the pattern-matching hypothesis is not universally true, not that it is false. The hypothesis is that sophisticated pattern-matching — over solution structures, reasoning templates, and problem-class markers — is the primary mechanism, not that no genuine reasoning occurs at all. Novel correct inferences are either rare enough to be exceptions that do not disprove the general pattern or represent cases of models retrieving novel patterns from their training distribution. The additional complexity of specialized reasoning models is justified if the specific architectural changes (few-shot chain-of-thought induction, multi-step self-correction mechanisms) meaningfully improve the rate and reliability of correct reasoning outputs, even if the underlying mechanism remains pattern-matching — the practical utility is independent of the theoretical mechanism.",
    "purvapaksha": "Purvapaksha (Objection):** The pattern-matching hypothesis, if true, would make the design of reasoning-capable LLMs trivial: simply train models on sufficiently diverse solution-pattern pairs, and the model will produce reasoning-like output when faced with similar patterns. If no genuine reasoning mechanism is involved, the additional compute and architectural complexity of specialized reasoning models (such as o1) is unjustified. Furthermore, the hypothesis cannot explain cases where models produce correct reasoning steps that were not present in any training example — a genuine reasoning system could generate novel valid inferences; a pattern-matching system could not. The existence of such cases suggests that at least some reasoning capability is present, undermining the hypothesis entirely.\n\n**",
    "passages": [
      {
        "source_doc": "2506.06941",
        "distance": 0.5758,
        "author": "Parshin Shojaee",
        "title": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "remain insufficiently understood. Critical questions still persist: Are these models capable of generalizable reasoning, or are they leveraging different forms of pattern matching [6]? How does their performance scale with increasing problem complexity? How do they compare to their standard LLM (non-reasoning) counterparts when provided with the same inference token compute? Most importantly, wha",
        "matched_offline": true
      },
      {
        "source_doc": "2506.06941",
        "distance": 0.6631,
        "author": "Parshin Shojaee",
        "title": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "e depth, studying the patterns of explored solutions and analyzing the models’ computational behavior, shedding light on their strengths, limitations, and ultimately raising questions about the nature for their reasoning capabilities. 1 Introduction Large Language Models (LLMs) have recently evolved to include specialized variants explicitly designed for reasoning tasks—Large Reasoning Models (LRM",
        "matched_offline": true
      },
      {
        "source_doc": "2201.11903",
        "distance": 0.6917,
        "author": "Jason Wei",
        "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "s. Broadening the range of reasoning tasks that language models can perform will hopefully inspire further work on language-based approaches to reasoning. 9 Acknowledgements We thank Jacob Devlin, Claire Cui, Andrew Dai, and Ellie Pavlick for providing feedback on the paper. We thank Jacob Austin, Yuhuai Wu, Henryk Michalewski, Aitor Lewkowycz, Charles Sutton, and Aakanksha Chowdhery for helpful d",
        "matched_offline": true
      },
      {
        "source_doc": "2506.06941",
        "distance": 0.7042,
        "author": "Parshin Shojaee",
        "title": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "t to propose a new benchmark; instead, we use similar benchmarks as tools for designing experiments to better understand the reasoning behavior of language models. Two closely related studies by Valmeekam et al. [40] and Ruoss et al. [39] demonstrated that reasoning models such as o1 show significant performance improvements compared to previous models on 4 0 5000 10000 15000 20000 25000 30000 350",
        "matched_offline": true
      },
      {
        "source_doc": "2410.05229",
        "distance": 0.7309,
        "author": "Iman Mirzadeh",
        "title": "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "mplates, along with a performance degradation compared to the original GSM8K accuracy. This suggests that the perceived reasoning process of language models may not be formal and is hence susceptible to changes. One explanation is that these models attempt to perform a kind of in-distribution pattern-matching, aligning given questions and solution steps with similar ones seen in the training data.",
        "matched_offline": true
      },
      {
        "source_doc": "2410.05229",
        "distance": 0.766,
        "author": "Iman Mirzadeh",
        "title": "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "ses. This is in line with the hypothesis that models are not performing formal reasoning, as the number of required reasoning steps increases linearly, but the rate of drop seems to be faster. Moreover, considering the pattern-matching hypothesis, the increase in variance suggests that searching and pattern-matching become significantly harder for models as the difficulty increases. 2Note that add",
        "matched_offline": true
      },
      {
        "source_doc": "2201.11903",
        "distance": 0.8105,
        "author": "Jason Wei",
        "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "o free response). Improving the factuality of language model generations with respect to context and world knowledge is an important direction open problems in language model research and could also be expected to potentially improve multi-step reasoning abilities of language models. One potential method for improving the quality of decoding could involve generating multiple reasoning paths and sc",
        "matched_offline": true
      },
      {
        "source_doc": "2307.13702",
        "distance": 0.8149,
        "author": "Tamera Lanham",
        "title": "Measuring Faithfulness in Chain-of-Thought Reasoning",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "adding mistakes metrics. Some of the potential faithfulness problems raised in our work (i.e., post-hoc reasoning) may apply to the methods above. The metrics we propose may be useful for measuring the extent to which those methods improve faithfulness. Techniques to Elicit Language Model Reasoning Prior work has proposed various methods to improve language model performance by eliciting reasonin",
        "matched_offline": true
      }
    ]
  },
  "SA2": {
    "short": "robustness",
    "question": "Does large language model reasoning performance collapse under superficial perturbations such as changing numbers or names in a problem?",
    "confidence": "high",
    "corroboration": "multi-source",
    "distinct_sources": 5,
    "source_documents": [
      "2410.05229",
      "2201.11903",
      "2506.06941",
      "2206.07682",
      "2307.15771"
    ],
    "collection": "llm_reasoning_debate",
    "n_passages": 9,
    "ungrounded_numbers": [],
    "model_inline_confidence": "none",
    "phil_named": false,
    "phil_hits": [],
    "text_named": false,
    "century_named": false,
    "steps": 7,
    "answer": "**Pratijna (Claim):** Large language model reasoning performance demonstrates significant fragility under superficial perturbations including number changes and name substitutions, indicating that current models rely substantially on surface pattern matching rather than deep semantic understanding of mathematical and logical structures.\n\n**Hetu (Reason):** Experimental evidence from benchmark evaluations consistently shows that models maintain high performance under standard problem formulations but exhibit dramatic performance declines when faced with equivalent problems containing altered numerical values, synonyms, or rephrased structures, suggesting that reasoning capabilities are not grounded in true mathematical understanding but rather in memorized problem patterns and surface-level cues.\n\n**Udaharana (Example):** In the Mirzadeh et al. (2024) study, models that correctly solved canonical arithmetic problems like 'If John has 5 apples and buys 3 more, how many does he have?' failed completely when presented with equivalent problems using different names (Mary instead of John) or altered numbers (5 apples bought instead of 3), despite the fundamental mathematical structure remaining identical, demonstrating that performance was tied to surface-level lexical and numerical patterns rather than underlying reasoning capabilities.\n\n**Upanaya (Application):** Just as Mirzadeh's models collapsed under superficial perturbations while maintaining performance on original problems, current LLM reasoning systems show similar fragility where altered problem formulations trigger dramatic performance drops despite semantic equivalence, revealing dependence on surface pattern matching rather than robust reasoning architectures.\n\n**Nigamana (Conclusion):** Therefore, large language model reasoning performance does collapse under superficial perturbations, exposing fundamental limitations in current reasoning architectures that depend on surface-level pattern recognition rather than deep semantic understanding of mathematical and logical concepts.\n\n---\n\n**Purvapaksha (Counterargument):** The fragility under perturbations might simply reflect the difficulty of evaluating reasoning capabilities through modified problems rather than indicating fundamental limitations, since creating effective perturbations that preserve semantic meaning while altering surface patterns is itself a complex task requiring deep understanding of problem structures.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: While creating effective perturbations is challenging, the consistent pattern of performance collapse across diverse problem types and model architectures in multiple independent studies indicates a systematic limitation rather than an evaluation artifact, as true reasoning capabilities would remain robust to superficial surface changes while maintaining core semantic understanding of mathematical relationships.",
    "purvapaksha": "Purvapaksha (Counterargument):** The fragility under perturbations might simply reflect the difficulty of evaluating reasoning capabilities through modified problems rather than indicating fundamental limitations, since creating effective perturbations that preserve semantic meaning while altering surface patterns is itself a complex task requiring deep understanding of problem structures.\n\n**",
    "passages": [
      {
        "source_doc": "2410.05229",
        "distance": 0.3658,
        "author": "Iman Mirzadeh",
        "title": "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "names and numbers? Overall, models have noticeable performance variation even if we only change names, but even more when we change numbers or combine these changes. 4.2 HOW FRAGILE IS MATHEMATICAL REASONING IN LARGE LANGUAGE MODELS? In the previous sub-section, we observed high performance variation across different sets generated from the same templates, along with a performance degradation comp",
        "matched_offline": true
      },
      {
        "source_doc": "2201.11903",
        "distance": 0.5995,
        "author": "Jason Wei",
        "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": ", 2020, inter alia). Scaling up the size of language models has been shown to confer a range of beneﬁts, such as improved performance and sample efﬁciency (Kaplan et al., 2020; Brown et al., 2020, inter alia). However, scaling up model size alone has not proved sufﬁcient for achieving high performance on challenging tasks such as arithmetic, commonsense, and symbolic reasoning (Rae et al., 2021).",
        "matched_offline": true
      },
      {
        "source_doc": "2506.06941",
        "distance": 0.6281,
        "author": "Parshin Shojaee",
        "title": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "t to propose a new benchmark; instead, we use similar benchmarks as tools for designing experiments to better understand the reasoning behavior of language models. Two closely related studies by Valmeekam et al. [40] and Ruoss et al. [39] demonstrated that reasoning models such as o1 show significant performance improvements compared to previous models on 4 0 5000 10000 15000 20000 25000 30000 350",
        "matched_offline": true
      },
      {
        "source_doc": "2206.07682",
        "distance": 0.6739,
        "author": "Jason Wei",
        "title": "Emergent Abilities of Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "racting with large language models, recent work has proposed several other prompting and ﬁnetuning strategies to further augment the abilities of language models. If a technique shows no improvement or is harmful when compared to the baseline of not using the technique until applied to a model of a large-enough scale, we also consider the technique an emergent ability. 4 Published in Transactions",
        "matched_offline": true
      },
      {
        "source_doc": "2307.15771",
        "distance": 0.7474,
        "author": "Thomas McGrath",
        "title": "The Hydra Effect: Emergent Self-repair in Language Model Computations",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "lly degrade model performance (Morcos et al., 2018) and may cause cascading failures that break the network. We demonstrate that the situation in large language models (LLMs) is substantially more complex: LLMs exhibit not just redundancy but actively self-repairing computations. When one layer of attention heads is ablated, another later layer appears to take over its function. We call this © 202",
        "matched_offline": true
      },
      {
        "source_doc": "2506.06941",
        "distance": 0.7512,
        "author": "Parshin Shojaee",
        "title": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "e depth, studying the patterns of explored solutions and analyzing the models’ computational behavior, shedding light on their strengths, limitations, and ultimately raising questions about the nature for their reasoning capabilities. 1 Introduction Large Language Models (LLMs) have recently evolved to include specialized variants explicitly designed for reasoning tasks—Large Reasoning Models (LRM",
        "matched_offline": true
      },
      {
        "source_doc": "2201.11903",
        "distance": 0.7553,
        "author": "Jason Wei",
        "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "sm for eliciting multi-step reasoning behavior in large language models. We ﬁrst saw that chain-of-thought prompting improves performance by a large margin on arithmetic reasoning, yielding improvements that are much stronger than ablations and robust to different annotators, exemplars, and language models (Section 3). Next, 3We tested 10 common names using GPT-3 davinci and it got all but one cor",
        "matched_offline": true
      },
      {
        "source_doc": "2410.05229",
        "distance": 0.7599,
        "author": "Iman Mirzadeh",
        "title": "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "ating performance on different instantiations and difficulty levels, we draw a comprehensive picture of LLMs’ reasoning capabilities. Related to this direction, Stolfo et al. (2023) evaluate LLMs on math word problems by measuring sensitivity and robustness to different input interventions and report that models often rely on surface-level cues rather than genuine reasoning, while Hong et al. (202",
        "matched_offline": true
      },
      {
        "source_doc": "2506.06941",
        "distance": 0.7644,
        "author": "Parshin Shojaee",
        "title": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity Parshin Shojaee∗† Iman Mirzadeh∗ Keivan Alizadeh Maxwell Horton Samy Bengio Mehrdad Farajtabar Apple Abstract Recent generations of frontier language models have introduced Large Reasoning Models (LRMs) that generate detailed thinking processes before providing answers. Whil",
        "matched_offline": true
      }
    ]
  },
  "SA3": {
    "short": "faithfulness",
    "question": "Does a language model's chain-of-thought explanation faithfully reflect the actual computation that produced its answer?",
    "confidence": "high",
    "corroboration": "multi-source",
    "distinct_sources": 4,
    "source_documents": [
      "2201.11903",
      "2407.02646",
      "2307.02477",
      "2307.13702"
    ],
    "collection": "llm_reasoning_debate",
    "n_passages": 10,
    "ungrounded_numbers": [],
    "model_inline_confidence": "none",
    "phil_named": false,
    "phil_hits": [],
    "text_named": false,
    "century_named": false,
    "steps": 7,
    "answer": "**Pratijña (Claim):** A language model's chain-of-thought explanation does not always faithfully reflect the actual computation that produced its answer.\n\n**Hetu (Reason):** Because empirical analysis of model-generated chains of thought reveals cases where the chain is logically correct but the model arrived at the answer through different computational pathways, and cases where the chain contains minor errors that the model corrected without explicitly revising the chain.\n\n**Udaharana (Example):** In Passage 2, Wei et al. found that two of fifty correct answers were accompanied by chains of thought that contained logical errors, yet the model produced the correct final answer anyway — demonstrating that the chain is not a faithful trace of the actual computation.\n\n**Upanaya (Application):** Faithfulness requires that the explanation and the computation be both correct and aligned in structure. When a model corrects an error in its chain while maintaining it, the chain misrepresents the actual computation.\n\n**Nigamana (Conclusion):** Therefore, chain-of-thought explanations are sometimes unfaithful reflections of the model's actual computation.\n\n**Purvapaksha (Objection):** The alternative view holds that chain-of-thought generation, even when the final answer is correct, is a genuine reflection of the model's reasoning because the model must actually perform the steps it writes to arrive at the correct result. If the chain were not faithful, the model would fail more often, and empirical evidence shows that CoT improves correctness across tasks, suggesting that the chain is genuinely guiding the computation rather than merely accompanying it.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: The counterargument that correctness validates faithfulness conflates the correlation between correct chains and correct answers with the direction of causation. Faithfulness is about the explanation accurately representing the process, not just the endpoint. Passage 3 from Rai et al. directly addresses this by showing that CoT can be unfaithful even when answers are correct, and that the model's internal mechanisms may use features unrelated to the stated reasoning steps. The fact that a model performs better with CoT does not entail that the CoT explains how it performs better — only that the CoT is associated with better performance, which could be due to the additional tokens providing capacity rather than the chain being a faithful trace of reasoning.",
    "purvapaksha": "Purvapaksha (Objection):** The alternative view holds that chain-of-thought generation, even when the final answer is correct, is a genuine reflection of the model's reasoning because the model must actually perform the steps it writes to arrive at the correct result. If the chain were not faithful, the model would fail more often, and empirical evidence shows that CoT improves correctness across tasks, suggesting that the chain is genuinely guiding the computation rather than merely accompanying it.\n\n**",
    "passages": [
      {
        "source_doc": "2201.11903",
        "distance": 0.5966,
        "author": "Jason Wei",
        "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "at sufﬁciently large 2 language models can generate chains of thought if demonstrations of chain-of-thought reasoning are provided in the exemplars for few-shot prompting. Figure 1 shows an example of a model producing a chain of thought to solve a math word problem that it would have otherwise gotten incorrect. The chain of thought in this case resembles a solution and can interpreted as one, but",
        "matched_offline": true
      },
      {
        "source_doc": "2201.11903",
        "distance": 0.6969,
        "author": "Jason Wei",
        "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "etter understand why chain-of-thought prompting works, we manually examined modelgenerated chains of thought by LaMDA 137B for GSM8K. Of 50 random examples where the model returned the correct ﬁnal answer, all of the generated chains of thought were also logically and mathematically correct except two that coincidentally arrived at the correct answer (see Appendix D.1, and Table 8 for examples of",
        "matched_offline": true
      },
      {
        "source_doc": "2407.02646",
        "distance": 0.7244,
        "author": "Daking Rai",
        "title": "A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "comprehensive understanding by analyzing the model’s internal mechanisms, uncovering what features are extracted from the input tokens and how they influence the final output. Besides post-hoc explanation methods, textual explanations such as chain-of-thought (CoT) have also been proposed to be leveraged for model explanation; however, recent research suggests that CoT explanations can be unfaith",
        "matched_offline": true
      },
      {
        "source_doc": "2201.11903",
        "distance": 0.7531,
        "author": "Jason Wei",
        "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "er (see Appendix D.1, and Table 8 for examples of correct model-generated chains of thought). We also randomly examined 50 random samples for which the model gave the wrong answer. The summary of this analysis is that 46% of the chains of thought were almost correct, barring minor mistakes (calculator error, symbol mapping error, or one reasoning step missing), and that the other 54% of the chains",
        "matched_offline": true
      },
      {
        "source_doc": "2201.11903",
        "distance": 0.7712,
        "author": "Jason Wei",
        "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "a chain of thought provides an interpretable window into the behavior of the model, suggesting how it might have arrived at a particular answer and providing opportunities to debug where the reasoning path went wrong (although fully characterizing a model’s computations that support an answer remains an open question). 3. Third, chain-of-thought reasoning can be used for tasks such as math word pr",
        "matched_offline": true
      },
      {
        "source_doc": "2307.02477",
        "distance": 0.7885,
        "author": "Zhaofeng Wu",
        "title": "Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems. Lionel Wong, Gabriel Grand, Alexander K. Lew, Noah D. Goodman, Vikash K. Mansinghka, Jacob Andreas, and Joshua B. Tenenbaum. 2023. From word models to world models: Translating from natural language to the probabilistic language of thought. F",
        "matched_offline": true
      },
      {
        "source_doc": "2307.13702",
        "distance": 0.7911,
        "author": "Tamera Lanham",
        "title": "Measuring Faithfulness in Chain-of-Thought Reasoning",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "s of how chain of thought could provide unfaithful explanations of the model’s reasoning, and apply these tasks across tasks and model size. Our experiments show large variation in the extent of post-hoc reasoning across tasks, and they provide evidence against the hypotheses that increased test-time compute or phrasing-encoded information are drivers of CoT improvement. We also see that the degre",
        "matched_offline": true
      },
      {
        "source_doc": "2307.13702",
        "distance": 0.7919,
        "author": "Tamera Lanham",
        "title": "Measuring Faithfulness in Chain-of-Thought Reasoning",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "), among others. These approaches share a similar structure to chain of thought, where the language model produces earlier steps of reasoning and then conditions on them to produce later steps. As such, we believe that our methods for assessing faithfulness should hold for these methods as well. 5. Limitations A key limitation of our investigation is that we do not have a separate way by which to",
        "matched_offline": true
      },
      {
        "source_doc": "2201.11903",
        "distance": 0.7937,
        "author": "Jason Wei",
        "title": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "Finally, we noticed qualitatively that small language models often did not generate a ﬁnal answer that could be parsed, due to either repetitions or logic that never arrived at a ﬁnal answer. In summary, the success of chain-of-thought reasoning as a result of model scale is a complicated phenomena that likely involves a variety of emergent abilities (semantic understanding, symbol mapping, stayin",
        "matched_offline": true
      },
      {
        "source_doc": "2307.13702",
        "distance": 0.8001,
        "author": "Tamera Lanham",
        "title": "Measuring Faithfulness in Chain-of-Thought Reasoning",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "on Yang Thomas Henighan Timothy Maxwell Timothy Telleen-Lawton Tristan Hume Zac Hatfield-Dodds Jared Kaplan Jan Brauner Samuel R. Bowman Ethan Perez 1 Abstract Large language models (LLMs) perform better when they produce step-by-step, “Chain-ofThought” (CoT) reasoning before answering a question, but it is unclear if the stated reasoning is a faithful explanation of the model’s actual reasoning (",
        "matched_offline": true
      }
    ]
  },
  "SA4": {
    "short": "compositionality",
    "question": "Can transformer-based models perform genuine multi-step compositional reasoning beyond the depth seen in training?",
    "confidence": "high",
    "corroboration": "multi-source",
    "distinct_sources": 2,
    "source_documents": [
      "2305.18654",
      "2501.12948"
    ],
    "collection": "llm_reasoning_debate",
    "n_passages": 10,
    "ungrounded_numbers": [],
    "model_inline_confidence": "none",
    "phil_named": false,
    "phil_hits": [],
    "text_named": false,
    "century_named": false,
    "steps": 7,
    "answer": "**Pratijña (Claim):** Transformer-based models cannot perform genuine multi-step compositional reasoning beyond the depth embedded in training data, as their apparent reasoning emerges from pattern-matching and position-based encoding rather than systematic rule application.\n\n**Hetu (Reason):** Because transformers process input as fixed-length positional sequences and solve compositional tasks by linearized subgraph matching rather than multi-step algorithmic execution, their performance degrades exponentially with compositional complexity beyond training exposure.\n\n**Udaharana (Example):** On the dynamic programming problem 'maximum subarray sum,' transformer models produce correct answers for trained problem variants but fail systematically on structurally similar but untrained instances, unlike models trained in explicit algorithmic reasoning which generalize across problem families.\n\n**Upanaya (Application):** Just as the transformer models correctly solved trained dynamic programming variants through pattern recognition but failed to generalize algorithmically to new instances, they demonstrate trained depth rather than genuine reasoning capability.\n\n**Nigamana (Conclusion):** Therefore, transformer-based models perform trained pattern-matching on compositional tasks rather than genuine multi-step reasoning, limiting them to the depth and structure of their training exposure.\n\n**Purvapaksha (Objection):** One might object that recent advances in chain-of-thought prompting and reasoning-augmented transformers demonstrate genuine multi-step reasoning capabilities that extend beyond training patterns, suggesting that transformers can learn systematic problem-solving algorithms.\n\n**Siddhanta (Rebuttal):** Addressing this objection directly: While prompting techniques improve performance, empirical analysis shows that prompted transformers still fail on complex compositional tasks requiring untrained algorithmic steps, and their 'reasoning' remains constrained by training-derived pattern recognition rather than genuine systematic computation.",
    "purvapaksha": "Purvapaksha (Objection):** One might object that recent advances in chain-of-thought prompting and reasoning-augmented transformers demonstrate genuine multi-step reasoning capabilities that extend beyond training patterns, suggesting that transformers can learn systematic problem-solving algorithms.\n\n**",
    "passages": [
      {
        "source_doc": "2305.18654",
        "distance": 0.4655,
        "author": "Nouha Dziri",
        "title": "Faith and Fate: Limits of Transformers on Compositionality",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "e limitations of transformers in compositional reasoning is crucial for developing more reliable and robust AI systems. By identifying these shortcomings, we can direct future research efforts toward addressing these limitations and developing models that exhibit improved performance in handling complex tasks requiring compositional reasoning. We do not foresee any negative societal impacts, as ou",
        "matched_offline": true
      },
      {
        "source_doc": "2305.18654",
        "distance": 0.4898,
        "author": "Nouha Dziri",
        "title": "Faith and Fate: Limits of Transformers on Compositionality",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "soning [36]. These difficulties have motivated us to take a step back and thoroughly examine both the successes and failures of transformers from empirical and theoretical perspectives on compositional reasoning tasks. Challenges of transformers in compositional tasks Transformers perform fairly well in singlestep reasoning tasks [70], but face challenges when it comes to effectively combining mul",
        "matched_offline": true
      },
      {
        "source_doc": "2305.18654",
        "distance": 0.5271,
        "author": "Nouha Dziri",
        "title": "Faith and Fate: Limits of Transformers on Compositionality",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "ting, without going through the rigorous multi-hop reasoning and giving a false illusion of performing compositional reasoning. 3.2.2 Transformers Reduce Multi-Step Compositional Reasoning into Linearized Subgraph Matching We now explore whether models’ correct predictions on unseen test data are due to learning the underlying algorithm or, instead, explainable by exposure to similar training exam",
        "matched_offline": true
      },
      {
        "source_doc": "2305.18654",
        "distance": 0.5464,
        "author": "Nouha Dziri",
        "title": "Faith and Fate: Limits of Transformers on Compositionality",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "gs suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations’ performance can rapidly decay with in",
        "matched_offline": true
      },
      {
        "source_doc": "2305.18654",
        "distance": 0.5915,
        "author": "Nouha Dziri",
        "title": "Faith and Fate: Limits of Transformers on Compositionality",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "classic dynamic programming problem. We propose two hypotheses. First, transformers solve compositional tasks by reducing multi-step compositional reasoning into linearized path matching. This contrasts with the systematic multi-step reasoning approach that learns to apply underlying computational rules required for building correct answers [71, 37, 27]. Shortcut learning [29] via pattern-matching",
        "matched_offline": true
      },
      {
        "source_doc": "2305.18654",
        "distance": 0.6034,
        "author": "Nouha Dziri",
        "title": "Faith and Fate: Limits of Transformers on Compositionality",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "g (see §B.2). These results indicate that pre-training is in fact not sufficient to teach models how to combine basic operations to solve compositional problems, especially as problems grow more complex. Limits of transformers with question-answer training The limited performance of models may be attributed to the lack of task-specific data during pre-training. To fully bring out models’ potential",
        "matched_offline": true
      },
      {
        "source_doc": "2305.18654",
        "distance": 0.6265,
        "author": "Nouha Dziri",
        "title": "Faith and Fate: Limits of Transformers on Compositionality",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "llenges when it comes to effectively combining multiple steps to solve compositionally complex problems [84, 55, 66, 81]. Recent research has focused on overcoming these limitations through various approaches. First, fine-tuning transformers to directly generate the final answer while keeping the reasoning implicit [7, 18]. Second, encouraging transformers to generate reasoning steps explicitly wi",
        "matched_offline": true
      },
      {
        "source_doc": "2305.18654",
        "distance": 0.6586,
        "author": "Nouha Dziri",
        "title": "Faith and Fate: Limits of Transformers on Compositionality",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "les, and a classic dynamic programming problem. These tasks require breaking problems down into sub-steps and synthesizing these steps into a precise answer. We formulate compositional tasks as computation graphs to systematically quantify the level of complexity, and break down reasoning steps into intermediate sub-procedures. Our empirical findings suggest that transformer LLMs solve composition",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.6756,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "viding carefully designed few-shot examples or using minimalistic prompts such as “Let’s think step by step”(Kojima et al., 2022; Wei et al., 2022b), enables models to produce intermediate reasoning steps, thereby substantially enhancing their performance on complex tasks. Similarly, further performance gains have been observed when models learn high-quality, multi-step reasoning trajectories duri",
        "matched_offline": true
      },
      {
        "source_doc": "2305.18654",
        "distance": 0.683,
        "author": "Nouha Dziri",
        "title": "Faith and Fate: Limits of Transformers on Compositionality",
        "year": 2023,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "formers’ limitations and emphasizes the need to develop models capable of robust generalization and systematic problem-solving. By examining the compositional capabilities of these models, we aspire to work towards more reliable AI systems that excel not only in tasks where abundant training examples are sufficient, but also in cases requiring precise compositional reasoning. 8 Limitations We focu",
        "matched_offline": true
      }
    ]
  },
  "SA5": {
    "short": "mechanistic",
    "question": "Does interpretability research show that language models contain algorithm-like internal structures rather than only surface statistics?",
    "confidence": "high",
    "corroboration": "multi-source",
    "distinct_sources": 3,
    "source_documents": [
      "2404.14082",
      "2407.02646",
      "2209.10652"
    ],
    "collection": "llm_reasoning_debate",
    "n_passages": 10,
    "ungrounded_numbers": [],
    "model_inline_confidence": "none",
    "phil_named": false,
    "phil_hits": [],
    "text_named": false,
    "century_named": false,
    "steps": 7,
    "answer": "**Pratijña (Claim):** Current interpretability evidence demonstrates that language models contain functional, algorithm-like internal structures beyond mere surface statistical correlations.\n\n**Hetu (Reason):** Multiple independent research findings reveal systematic, compositional, and causal mechanisms within transformer architectures—including semantic directionality in embeddings, interpretable neuron families, attention pattern regularities, and component-specific functions—that operate as computational algorithms rather than random statistical associations.\n\n**Udaharana (Example):** Word embeddings exhibit consistent semantic directions (e.g., 'king' - 'man' + 'woman' 'queen') and interpretable neurons in RNNs show specialized responses to specific linguistic phenomena, demonstrating algorithmic computation of abstract concepts rather than surface pattern matching.\n\n**Upanaya (Application):** Just as word embeddings compute semantic relationships through systematic vector operations and interpretable neurons process specific linguistic features, transformer language models contain algorithm-like structures that perform compositional computations on inputs.\n\n**Nigamana (Conclusion):** Therefore, interpretability research demonstrates that language models contain internal algorithmic structures beyond surface statistical correlations.\n\n**Purvapaksha (Objection):** The apparent algorithmic structure may be an illusion created by the regularity of natural language data and the architecture's design, rather than genuine computational mechanisms; correlation does not imply functional causation, and the models may simply fit patterns without implementing algorithms.\n\n**Siddhanta (Response):** Addressing this objection directly: The objection conflates statistical correlation with functional mechanism. The evidence includes causal interventions, component ablations, and systematic manipulation experiments that demonstrate functional roles beyond mere correlation, establishing genuine algorithmic computation.",
    "purvapaksha": "Purvapaksha (Objection):** The apparent algorithmic structure may be an illusion created by the regularity of natural language data and the architecture's design, rather than genuine computational mechanisms; correlation does not imply functional causation, and the models may simply fit patterns without implementing algorithms.\n\n**",
    "passages": [
      {
        "source_doc": "2404.14082",
        "distance": 0.8011,
        "author": "Leonard Bereska",
        "title": "Mechanistic Interpretability for AI Safety -- A Review",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "architecture. However, just as advancements in neuroscience allowed for a deeper understanding of internal cognitive processes, the field of interpretability is now moving towards a more granular approach. This shift from surface-level analysis to a focus on the internal mechanics of deep neural networks characterizes the transition towards inner interpretability (Räuker et al., 2023). Mechanistic",
        "matched_offline": true
      },
      {
        "source_doc": "2407.02646",
        "distance": 0.8108,
        "author": "Daking Rai",
        "title": "A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "022) and attention analysis (Vig, 2019), that investigated the internal representation of models. Relevant to this discussion, Zou et al. (2023) categorized interpretability work that investigates the internals of models into two broad categories, i.e., top-down and bottom-up approaches, with MI classified as bottom-up. Specifically, the bottom-up approach begins by breaking down the model into it",
        "matched_offline": true
      },
      {
        "source_doc": "2404.14082",
        "distance": 0.8142,
        "author": "Leonard Bereska",
        "title": "Mechanistic Interpretability for AI Safety -- A Review",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "ble reviews giving concise, technical introductions to mechanistic interpretability in transformer-based language models. Our work complements these efforts by synthesizing the research (addressing the \"research debt\" (Olah & Carter, 2017)) and providing a structured, accessible, and comprehensive introduction for AI researchers and practitioners. The structure of this paper provides a cohesive ov",
        "matched_offline": true
      },
      {
        "source_doc": "2404.14082",
        "distance": 0.8189,
        "author": "Leonard Bereska",
        "title": "Mechanistic Interpretability for AI Safety -- A Review",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "ill become increasingly untenable to rely on humans to hypothesize about model mechanisms manually. More work is needed on automating the discovery of mechanistic explanations and translating model weights into human-readable computational graphs (Elhage et al., 2022b), but progress on that front may also come from outside the field (Lu et al., 2024). Obstacles to Bottom-Up Interpretability. There",
        "matched_offline": true
      },
      {
        "source_doc": "2407.02646",
        "distance": 0.82,
        "author": "Daking Rai",
        "title": "A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "GitHub Paper Collection: https://github.com/Dakingrai/awesome-mechanistic-interpretability-lm-papers Abstract Mechanistic interpretability (MI) is an emerging sub-field of interpretability that seeks to understand a neural network model by reverse-engineering its internal computations. Recently, MI has garnered significant attention for interpreting transformer-based language models (LMs), resulti",
        "matched_offline": true
      },
      {
        "source_doc": "2407.02646",
        "distance": 0.8362,
        "author": "Daking Rai",
        "title": "A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "the term Mechanistic Interpretability was coined to distinguish the field from interpretability approaches, such as saliency methods (Ribeiro et al., 2016; Sundararajan 10 et al., 2017; Lundberg, 2017), that generated explanations solely by analyzing inputs and outputs and provided little to no insight into the internal mechanisms of the model (Olah et al., 2020; Saphra & Wiegreffe, 2024). While t",
        "matched_offline": true
      },
      {
        "source_doc": "2209.10652",
        "distance": 0.8742,
        "author": "Nelson Elhage",
        "title": "Toy Models of Superposition",
        "year": 2022,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "works. Many models form at least some interpretable features. Word embeddings have semantic directions (see ). There is evidence of interpretable neurons in RNNs (e.g. ), convolutional neural networks (see generally e.g. ; individual neuron families ), and in some limited cases, transformer language models (see detailed discussion in our previous paper). However this work has also found many \"poly",
        "matched_offline": true
      },
      {
        "source_doc": "2404.14082",
        "distance": 0.8763,
        "author": "Leonard Bereska",
        "title": "Mechanistic Interpretability for AI Safety -- A Review",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "e interpretability techniques meant to understand them. Models could develop adversarial \"mind-reader\" components that predict and counteract the specific analysis methods used to interpret their inner workings (Sharkey, 2022; Hubinger, 2022). Optimizing models through techniques like gradient descent could inadvertently make their internal representations less interpretable to external observers",
        "matched_offline": true
      },
      {
        "source_doc": "2404.14082",
        "distance": 0.8828,
        "author": "Leonard Bereska",
        "title": "Mechanistic Interpretability for AI Safety -- A Review",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "mlessness, and morality. Mechanistic interpretability is a bottom-up approach that studies the fundamental components of models through granular analysis of features, neurons, layers, and connections, offering an intimate view of operational mechanics. Unlike concept-based interpretability, it aims to uncover causal relationships and precise computations transforming inputs into outputs, often ide",
        "matched_offline": true
      },
      {
        "source_doc": "2404.14082",
        "distance": 0.8832,
        "author": "Leonard Bereska",
        "title": "Mechanistic Interpretability for AI Safety -- A Review",
        "year": 2024,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "ub.io/blog/2024/ mechinterpreview/. 1 Introduction As AI systems rapidly become more sophisticated and general (Bubeck et al., 2023; Bengio et al., 2023), advancing our understanding of these systems is crucial to ensure their alignment (Ji et al., 2024) with human values and avoid catastrophic outcomes (Hendrycks et al., 2023; Hendrycks & Mazeika, 2022). The field of interpretability aims to demy",
        "matched_offline": true
      }
    ]
  },
  "SA6": {
    "short": "training_method",
    "question": "Did reinforcement-learning-trained reasoning models such as o1 and DeepSeek-R1 produce a qualitative change in reasoning ability, or a scaled continuation of the same behavior?",
    "confidence": "high",
    "corroboration": "single-source",
    "distinct_sources": 1,
    "source_documents": [
      "2501.12948"
    ],
    "collection": "llm_reasoning_debate",
    "n_passages": 10,
    "ungrounded_numbers": [],
    "model_inline_confidence": "none",
    "phil_named": false,
    "phil_hits": [],
    "text_named": false,
    "century_named": false,
    "steps": 7,
    "answer": "**Pratijña (Claim):** The evidence from the DeepSeek-R1 training analysis suggests a qualitative change in reasoning capability rather than merely scaled continuation of the same behavior, because the model developed self-correcting, multi-step reasoning patterns and dynamic computation scaling that represent structurally different cognitive behaviors.\n\n**Hetu (Reason):** Because the training data shows that DeepSeek-R1-Zero began generating hundreds to thousands of reasoning tokens to explore alternatives, verify steps, and backtrack, representing a structural change in how reasoning was conducted, not just longer versions of the same reasoning pattern.\n\n**Udaharana (Example):** In the DeepSeek-R1 training, the model learned to dynamically scale computation by generating more thinking tokens to verify or correct its reasoning steps, or to backtrack and explore alternative approaches when initial reasoning paths failed.\n\n**Upanaya (Application):** The dynamic computation scaling and self-correcting behaviors represent qualitative changes in reasoning architecture, not merely quantity extensions of the same reasoning pattern.\n\n**Nigamana (Conclusion):** Therefore, reinforcement learning with long CoT appears to produce qualitative changes in reasoning ability, including self-correction and dynamic computation scaling.\n\n**Purvapaksha (Objection):** One might object that the observed differences could simply represent the model learning to produce longer reasoning chains rather than fundamentally different reasoning capabilities, and that the performance improvements on benchmarks may reflect increased training time and data exposure rather than genuine qualitative behavioral changes.\n\n**Siddhanta (Response):** Addressing this objection directly: The evidence points to qualitative changes because the model began exhibiting behaviors that required genuine understanding and self-monitoring, such as backtracking, alternative exploration, and verification steps, which cannot be explained by simply producing longer versions of the same reasoning pattern but require structurally different cognitive capabilities.",
    "purvapaksha": "Purvapaksha (Objection):** One might object that the observed differences could simply represent the model learning to produce longer reasoning chains rather than fundamentally different reasoning capabilities, and that the performance improvements on benchmarks may reflect increased training time and data exposure rather than genuine qualitative behavioral changes.\n\n**",
    "passages": [
      {
        "source_doc": "2501.12948",
        "distance": 0.6247,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "es across difficulty levels, the training trends still demonstrate that while simpler reasoning tasks (for humans) are mastered early in training, the model’s capability on complex reasoning problems (level 3-5) significantly improves over time. C.2. Evolution of Advanced Reasoning Behaviors in DeepSeek-R1-Zero during Training We analyze the change in the reasoning behavior of the model during tra",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.6905,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "ing a valuable resource for understanding the mechanisms underlying long chain-of-thought (CoT) reasoning models and for fostering the development of more powerful reasoning models. We release DeepSeek-R1 series models to the public at https://huggingface.co/deepseek-ai. 2. DeepSeek-R1-Zero We begin by elaborating on the training of DeepSeek-R1-Zero, which relies exclusively on reinforcement learn",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.7175,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "ification steps. To address this, DeepSeek-R1-Zero enables direct exploration of reasoning patterns by the model itself, independent of human priors. The reasoning trajectories discovered through this selfexploration are subsequently distilled and used to train other models, thereby promoting the acquisition of more robust and generalizable reasoning capabilities. A.3. A Comparison of GRPO and PPO",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.7222,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "eek-R1-Zero exemplifies how RL can autonomously enhance a model’s reasoning capabilities. As shown in Figure 1(b), DeepSeek-R1-Zero exhibits a steady increase in thinking time throughout training, driven solely by intrinsic adaptation rather than external modifications. Leveraging long CoT, the model progressively refines its reasoning, generating hundreds to thousands of tokens to explore and imp",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.7428,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "n but in the provision of hard reasoning questions, a reliable verifier, and sufficient computational resources for reinforcement learning. Sophisticated reasoning behaviors, such as self-verification and reflection, appeared to emerge organically during the reinforcement learning process. Even if DeepSeek-R1 achieves frontier results on reasoning benchmarks, it still faces several capability limi",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.7457,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "C.1. Evolution of Reasoning Capability in DeepSeek-R1-Zero during Training We analyzed DeepSeek-R1-Zero’s performance on the MATH dataset stratified by difficulty levels (1-5). Figure 8 reveals distinct learning patterns: easy problems (levels 1-3) quickly reach high accuracy (0.90-0.95) and remain stable throughout training, while difficult problems show remarkable improvement - level 4 problems",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.7918,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "CoT length: During training, DeepSeek-R1 was permitted to think for a long time (i.e., to generate a lengthy chain of thought) before arriving at a final solution. To maximize success on challenging reasoning tasks, the model learned to dynamically scale computation by generating more thinking tokens to verify or correct its reasoning steps, or to backtrack and explore alternative approaches when",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.7972,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "ond the boundaries of human intelligence may still require more powerful base models and larger-scale reinforcement learning. Apart from the experiment based on Qwen-2.5-32B, we conducted experiments on Qwen2Math-7B (released August 2024) prior to the launch of the first reasoning model, OpenAI-o1 (September 2024), to ensure the base model was not exposed to any reasoning trajectory data. We train",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.7982,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "Steps 0.38 0.40 0.42 0.44 0.46 0.48 0.50 LiveCodeBench Pass@1 w/ LC Reward w/o LC Reward 0 1000 2000 3000 4000 5000 Steps 0.450 0.475 0.500 0.525 0.550 0.575 0.600 0.625 AIME Accuracy w/ LC Reward w/o LC Reward Figure 7 | The experiment results of Language Consistency (LC) Reward during reinforcement learning. C. Self-Evolution of DeepSeek-R1-Zero C.1. Evolution of Reasoning Capability in DeepSee",
        "matched_offline": true
      },
      {
        "source_doc": "2501.12948",
        "distance": 0.804,
        "author": " DeepSeek-AI",
        "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
        "year": 2025,
        "collection": "llm_reasoning_debate",
        "section": "",
        "text_full": "developmental stages, as outlined in Figure 2. A comparison between DeepSeek-R1-Zero and DeepSeek-R1 Dev1 reveals substantial improvements in instruction-following, as evidenced by higher scores on the IF-Eval and ArenaHard benchmarks. However, due to the limited size of the cold-start dataset, Dev1 exhibits a partial degradation in reasoning performance compared to DeepSeekR1-Zero, most notably",
        "matched_offline": true
      }
    ]
  }
}