Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Oliver Stanley

REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards

May 30, 2025

Zafir Stojanovski, Oliver Stanley, Joe Sharratt, Richard Jones, Abdulhakeem Adefioye, Jean Kaddour, Andreas Köpf

Figure 1 for REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards

Figure 2 for REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards

Figure 3 for REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards

Figure 4 for REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards

Abstract:We introduce Reasoning Gym (RG), a library of reasoning environments for reinforcement learning with verifiable rewards. It provides over 100 data generators and verifiers spanning multiple domains including algebra, arithmetic, computation, cognition, geometry, graph theory, logic, and various common games. Its key innovation is the ability to generate virtually infinite training data with adjustable complexity, unlike most previous reasoning datasets, which are typically fixed. This procedural generation approach allows for continuous evaluation across varying difficulty levels. Our experimental results demonstrate the efficacy of RG in both evaluating and reinforcement learning of reasoning models.

* For code, see https://github.com/open-thought/reasoning-gym

Via

Access Paper or Ask Questions

OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Apr 14, 2023

Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi(+8 more)

Figure 1 for OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Figure 2 for OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Figure 3 for OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Figure 4 for OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Abstract:Aligning large language models (LLMs) with human preferences has proven to drastically improve usability and has driven rapid adoption as demonstrated by ChatGPT. Alignment techniques such as supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) greatly reduce the required skill and domain knowledge to effectively harness the capabilities of LLMs, increasing their accessibility and utility across various domains. However, state-of-the-art alignment techniques like RLHF rely on high-quality human feedback data, which is expensive to create and often remains proprietary. In an effort to democratize research on large-scale alignment, we release OpenAssistant Conversations, a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages distributed across 66,497 conversation trees, in 35 different languages, annotated with 461,292 quality ratings. The corpus is a product of a worldwide crowd-sourcing effort involving over 13,500 volunteers. To demonstrate the OpenAssistant Conversations dataset's effectiveness, we present OpenAssistant, the first fully open-source large-scale instruction-tuned model to be trained on human data. A preference study revealed that OpenAssistant replies are comparably preferred to GPT-3.5-turbo (ChatGPT) with a relative winrate of 48.3% vs. 51.7% respectively. We release our code and data under fully permissive licenses.

* 44 pages, 39 figures

Via

Access Paper or Ask Questions