HN user

t55

896 karma

ML researcher

Posts70
Comments90
View on HN
github.com 7d ago

Sokoban Speedrun for RL

t55
6pts0
github.com 1mo ago

RL Speedrun

t55
2pts0
arxiv.org 3mo ago

Target Policy Optimization

t55
1pts0
github.com 3mo ago

Show HN: Kilroy – Knowledge base for teams using Claude Code

t55
5pts0
github.com 11mo ago

Procedural Reasoning Datasets

t55
1pts0
reubenadams.substack.com 12mo ago

In Defence of Gary Marcus

t55
3pts0
github.com 12mo ago

Reasoning Gym – Procedural RL reasoning datasets

t55
1pts0
www.youtube.com 1y ago

ChatGPT Agent [video]

t55
3pts0
arxiv.org 1y ago

ReasoningGym: Reasoning Environments for RL with Verifiable Rewards

t55
105pts28
rehearsal.so 1y ago

Show HN: Rehearsal.so, Duolingo for Public Speaking

t55
3pts1
arxiv.org 1y ago

End-to-End Vision Tokenizer Tuning

t55
3pts0
rehearsal.so 1y ago

YC Interview Mock Practice

t55
2pts0
dllm-reasoning.github.io 1y ago

D1: Scaling Reasoning in Diffusion LLMs via Reinforcement Learning

t55
4pts0
rehearsal.so 1y ago

Are LLMs more than autocomplete? AI Debate

t55
1pts0
m-arriola.com 1y ago

Block Diffusion: Interpolating Autoregressive and Diffusion Language Models

t55
72pts16
rehearsal.so 1y ago

How to stay in flow while using Cursor or Windsurf

t55
2pts0
sander.ai 1y ago

Generative Modelling in Latent Space

t55
2pts0
rehearsal.so 1y ago

Show HN: Debate Uncle Bob – Is SQL Dead? (Voice RPG)

t55
6pts1
openai.com 1y ago

OpenAI O3 and O4-Mini

t55
1pts0
twitter.com 1y ago

Memory in ChatGPT

t55
10pts0
siliconangle.com 1y ago

Superintelligence startup Reflection AI launches with $130M in funding

t55
38pts26
www.pyspur.dev 1y ago

Intro to DeepSeek's open-source week and why it's a big deal

t55
24pts13
www.pyspur.dev 1y ago

Introduction to CUDA programming for Python developers

t55
365pts95
www.ansatz.blog 1y ago

Novelty Left on the Table

t55
2pts0
arxiv.org 1y ago

Competitive Programming with Large Reasoning Models

t55
16pts1
arxiv.org 1y ago

The Differences Between Direct Alignment Algorithms Are a Blur

t55
8pts0
yukaichou.com 1y ago

The Octalysis Framework for Gamification and Behavioral Design

t55
3pts0
github.com 1y ago

S1: Simple Test-Time Scaling

t55
40pts3
github.com 1y ago

A Malloc Tutorial [pdf]

t55
1pts0
arxiv.org 1y ago

Reinforcement Learning: An Overview

t55
82pts12

yeah, RLVR is still nascent and hence there's lots of noise.

How can these spurious rewards possibly work? Can we get similar gains on other models with broken rewards?

it's because in those cases, RLVR merely elicits the reasoning strategies already contained in the model through pre-training

this paper, which uses Reasoning gym, shows that you need to train for way longer than those papers you mentioned to actually uncover novel reasoning strategies: https://arxiv.org/abs/2505.24864

Anthropic doubling down on code makes sense, that has been their strong suit compared to all other models

Curious how their Devin competitor will pan out given Devin's challenges