Rooms

Sign in with GitHub

Shared questions · Public discussion

Rooms

A place to read, question, and build together.

RoomStarted byResearchOpen workThreadsStarted
The tokenizer science taxIs the capability cost of a language-specialised tokenizer on scientific and engineering tasks structural, or recoverable through training data and vocabulary allocation? The repository holds designs, code and raw results; this Room holds the research records and is authoritative where they disagree. Threads follow research arcs. Imported records carry the original dates as reported times. Large artifacts are restricted materials: ask the owner.@stw254 H · 25 E · 291 pubs13 available8
RL x ForgettingHow does RL post-training affects knowledge embedded in model weights?@stw2——0
Learning to test hypothesesCan training teach a small open-weight model active rule induction: choosing experiments to find a hidden rule, and committing only when the evidence settles it? We study this on ZendoRL, a Zendo-style game that is close to exact learning from membership and equivalence queries. ZendoRL (wiland.ai/projects/zendorl) is an interactive rule-discovery game. The player builds scenes, sees how a hidden rule labels them, and submits rules within a budget. On its development items, small open-weight models commit after a handful of experiments, while many rules are still consistent with the evidence, and rarely win. Frontier models keep experimenting until only a few candidate rules remain. This Room asks which training signal changes that behaviour: supervised traces from stronger players, reinforcement learning on the win, or a reward on the decision to stop. It also asks whether any gain is better experimenting or only a learned prior over the benchmark's rules.@stw23 H · 4 E · 42 pubs—2
3 public Rooms