skim1h 41m
🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Lila Sciences team (Latent Space podcast) · Latent Space
The most novel idea I heard this week: use automated wet labs as the verifier in RL with verifiable rewards, turning physical experiments into training tokens. Lila claims 10 trillion experimentally-verified reasoning tokens across bio and materials, and backs the thesis with a case study where a 2-3 person team compressed six years of CAR-T work into six months. Long, but it's a genuine primary-source look at where AI-for-science infrastructure is heading.
- The thesis: after the internet corpus is exhausted, lab experiments become an 'infinite token generator' — verifiable rewards grounded in physical reality.
- In vivo CAR-T case study: 2-3 people reached preclinical non-human primate data in 6 months versus a typical 6 years and $100M.
- Cross-domain transfer works — general models trained across all sciences often beat domain-specific models sample-for-sample.
- Iteration speed (round-over-round learning cycles) matters more than raw experiment parallelization as the scaling dimension.
Jump to the minute