-
Can Large Language Models Understand Context?
Paper • 2402.00858 • Published • 23 -
OLMo: Accelerating the Science of Language Models
Paper • 2402.00838 • Published • 85 -
Self-Rewarding Language Models
Paper • 2401.10020 • Published • 151 -
SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
Paper • 2401.17072 • Published • 25
Collections
Discover the best community collections!
Collections including paper arxiv:2510.08558
-
Towards General-Purpose Model-Free Reinforcement Learning
Paper • 2501.16142 • Published • 30 -
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 142 -
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Paper • 2504.13837 • Published • 138 -
Learning to Reason under Off-Policy Guidance
Paper • 2504.14945 • Published • 88
-
Agent Learning via Early Experience
Paper • 2510.08558 • Published • 266 -
Regret Bounds for Markov Decision Processes with Recursive Optimized Certainty Equivalents
Paper • 2301.12601 • Published -
Bayesian Risk Markov Decision Processes
Paper • 2106.02558 • Published -
Metrics for Markov Decision Processes with Infinite State Spaces
Paper • 1207.1386 • Published
-
ExGRPO: Learning to Reason from Experience
Paper • 2510.02245 • Published • 80 -
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
Paper • 2510.01132 • Published • 5 -
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Paper • 2510.04618 • Published • 123 -
MixReasoning: Switching Modes to Think
Paper • 2510.06052 • Published • 21
-
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
Paper • 2509.26507 • Published • 535 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 493 -
Agent Learning via Early Experience
Paper • 2510.08558 • Published • 266 -
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Paper • 2510.04618 • Published • 123
-
MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
Paper • 2510.04363 • Published -
Control Plane as a Tool: A Scalable Design Pattern for Agentic AI Systems
Paper • 2505.06817 • Published -
Agentic Web: Weaving the Next Web with AI Agents
Paper • 2507.21206 • Published -
Improving Autonomous AI Agents with Reflective Tree Search and Self-Learning
Paper • 2410.02052 • Published • 9
-
Can Large Language Models Understand Context?
Paper • 2402.00858 • Published • 23 -
OLMo: Accelerating the Science of Language Models
Paper • 2402.00838 • Published • 85 -
Self-Rewarding Language Models
Paper • 2401.10020 • Published • 151 -
SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
Paper • 2401.17072 • Published • 25
-
ExGRPO: Learning to Reason from Experience
Paper • 2510.02245 • Published • 80 -
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
Paper • 2510.01132 • Published • 5 -
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Paper • 2510.04618 • Published • 123 -
MixReasoning: Switching Modes to Think
Paper • 2510.06052 • Published • 21
-
Towards General-Purpose Model-Free Reinforcement Learning
Paper • 2501.16142 • Published • 30 -
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 142 -
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Paper • 2504.13837 • Published • 138 -
Learning to Reason under Off-Policy Guidance
Paper • 2504.14945 • Published • 88
-
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
Paper • 2509.26507 • Published • 535 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 493 -
Agent Learning via Early Experience
Paper • 2510.08558 • Published • 266 -
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Paper • 2510.04618 • Published • 123
-
MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
Paper • 2510.04363 • Published -
Control Plane as a Tool: A Scalable Design Pattern for Agentic AI Systems
Paper • 2505.06817 • Published -
Agentic Web: Weaving the Next Web with AI Agents
Paper • 2507.21206 • Published -
Improving Autonomous AI Agents with Reflective Tree Search and Self-Learning
Paper • 2410.02052 • Published • 9
-
Agent Learning via Early Experience
Paper • 2510.08558 • Published • 266 -
Regret Bounds for Markov Decision Processes with Recursive Optimized Certainty Equivalents
Paper • 2301.12601 • Published -
Bayesian Risk Markov Decision Processes
Paper • 2106.02558 • Published -
Metrics for Markov Decision Processes with Infinite State Spaces
Paper • 1207.1386 • Published