AI Researchers Discuss Counterarguments to RSI, Chinese Lab Progress, Long‑Horizon RL and More (Dwarkesh Podcast)
| Source: Techmeme | Original article
Researchers John Schulman, Beren Millidge and Charlie O'Neill discuss steelmanning the case against RSI, Chinese labs' progress and long‑horizon reinforcement learning on the Dwarkesh Podcast.
A new episode of the Dwarkesh Podcast brings together three leading AI researchers—John Schulman, Beren Millidge and Charlie O’Neill—for a deep‑dive Q&A on several hot‑button topics in advanced AI. The trio “steel‑mans” the case against recursive self‑improvement (RSI), evaluates the pace of work emerging from Chinese laboratories, and examines the challenges of long‑horizon reinforcement learning (RL). Their discussion, captured in a YouTube release, is punctuated by the refrain that the field is “nowhere near the ceiling,” underscoring a belief that current capabilities remain far below any theoretical limits.
The conversation matters because it surfaces expert perspectives on the safety and strategic dimensions of next‑generation AI. By articulating a rigorous critique of RSI—a scenario where an AI system could autonomously accelerate its own intelligence—the guests highlight potential failure modes that policymakers and developers must anticipate. Their assessment of Chinese labs adds a geopolitical layer, suggesting that rapid progress abroad could compress the timeline for alignment research. Meanwhile, the focus on long‑horizon RL points to concrete technical hurdles that must be solved before agents can reliably pursue complex, temporally extended goals without unintended side effects.
As we reported on 12 September 2026, AI researchers were already debating how close the community is to recursive self‑improvement. This podcast builds on that dialogue, offering a more granular view of the arguments and the research directions that could shape the next wave of breakthroughs. Going forward, observers will watch for any new publications or open‑source releases from the discussed Chinese groups, for updates from Schulman’s Thinking Machines Lab and Millidge’s alignment team, and for subsequent Dwarkesh episodes that may track how the community’s stance on RSI and long‑horizon RL evolves in response to emerging data.
Sources
Back to AIPULSEN