
Free Daily Podcast Summary
by Machine Learning Street Talk (MLST)
We engage in fascinating discussions with pre-eminent figures in the AI field. Our flagship show covers current affairs in AI, cognitive science, neuroscience and philosophy of mind with in-depth analysis. Our approach is unrivalled in terms of scope and rigour – we believe in intellectual diversity in AI, and we touch on all of the main ideas in the field with the hype surgically removed. MLST is run by Tim Scarfe, Ph.D and features regular appearances from MIT Doctor of Philosophy Keith Duggar.
The most recent episodes — sign up to get AI-powered summaries of each one.
Tsung-Hsien (Shawn) Wen, CTO of PolyAI, tells Tim Scarfe why voice agents are harder than text agents. Voice adds time, and a good conversation depends on adapting to the person on the line, not just on reasoning to the best answer. Shawn describes an audio-native model (Dialog-RSN-1) that first predicts a turn-taking signal, then replies in text with citations, and writes the transcript last so enterprises can audit it.Along the way: training on real, noisy calls with synthetic noise added, and why over-cleaned audio made the new model worse. Latency, and what a voice agent should do while it thinks. Why a voice with a hint of regional accent beats a generic one. Why public benchmarks fall short for voice, why enterprises want to own their agent harness, and whether behaviour belongs in the harness or in the weights.The last stretch is about working with agents: cognitive debt, the shift from producing content to checking it, Wispr Flow, building tools that agents can use, and whether slop is in the eye of the reader.This episode was produced in partnership with PolyAI.https://poly.ai CHAPTERS0:00 Why voice agents are harder than text4:27 What enterprises want, and why PolyAI built its own model9:04 How an audio-native voice model works15:21 Training data, spectrograms and synthetic noise20:36 The cocktail party problem and the future of turn-taking25:21 Latency, adaptive reasoning and keeping callers' trust31:33 Voices, personality and the uncanny valley36:32 How do you benchmark a voice agent?42:00 Harness engineering and owning the intelligence45:09 Well-specified problems and auditable agents49:46 Weight adaptation and cognitive debt56:31 Agents at work: Wispr Flow, voice and tool building1:02:50 The next decade of voice, and what counts as slopREFERENCESThe Bitter Lesson: http://www.incompleteideas.net/IncIdeas/BitterLesson.html [9:05]Retrieval-augmented generation: https://arxiv.org/abs/2005.11401 [13:23]Mel scale: https://en.wikipedia.org/wiki/Mel_scale [17:16]Victor Zue: https://en.wikipedia.org/wiki/Victor_Zue [17:45]Cocktail party effect: https://en.wikipedia.org/wiki/Cocktail_party_effect [20:39]Speaker diarisation: https://en.wikipedia.org/wiki/Speaker_diarisation [21:16]
Leonardo de Moura created Lean and co-created Z3. ---This episode is sponsored by Parallel.Parallel, where agents find answers: web search, extraction and deep research APIs built for AI agents.Start free with the Parallel MCP server and $5 of credits every month: https://parallel.ai/mlst?utm_source=creator&utm_medium=podcast&utm_content=MLST---Tim Scarfe talks with Leo about how Lean escaped its original audience, why dependent types and Mathlib made it useful to working mathematicians, and what happens when formal verification leaves the lab. De Moura explains the small trusted kernel and independent checkers, and gives his account of the recent Collatz incident, in which a purported proof was accepted by both Lean's official kernel and Nanoda, apparently by exploiting a different bug in each.---TIMESTAMPS:00:00:00 Cold open: the green checkmark can lie00:00:59 Cathedral or bazaar: who controls Lean's core?00:04:55 Why Lean's core stays small and protected00:08:07 The Slack purge, Brandolini's law and the Lean FRO00:11:12 The Collatz exploit: two kernels, two bugs00:16:44 More kernels, reward hacking and safety by transparency00:21:04 Sponsor: Parallel00:21:59 Kim Morrison, Claude and the zlib proof00:25:25 Can we specify complex systems?00:28:04 Specs change: proofs are cheaper to redo with AI00:31:27 From Lean 1 to Lean 400:34:57 Dependent types in plain terms00:37:25 Lean 4's extensibility and Mathlib's growth00:41:44 Mathlib as infrastructure: Formal Frontiers00:44:15 Creativity, abstraction and nut-sniping00:48:46 Breadcrumbs, not learning: what AI agents lack00:53:03 Competence without comprehension, and verified guardrails00:56:21 Is the human still the author?01:01:44 AlphaProof, LLMs and why certificates still matter01:06:38 What's next for Lean, and its legacy01:11:42 How to start learning Lean---REFERENCES:tool:[00:00:48] Leanhttps://lean-lang.org/[00:01:24] Mathlibhttps://github.com/leanprover-community/mathlib4[00:12:09] nanoda_libhttps://github.com/ammkrn/nanoda_lib[00:12:19] CollatzLeanhttps://github.com/xrchz/CollatzLean/blob/a79357462a33d2a6babd4cf6c8d8bcd25425d653/README.md[00:13:06] Lean issue 14576https://github.com/leanprover/lean4/issues/14576[00:13:21] Lean pull request 14577https://github.com/leanprover/lean4/pull/14577[00:13:45] nanoda_lib pull request 22https://github.com/ammkrn/nanoda_lib/pull/22[00:15:33] Lean comparatorhttps://github.com/leanprover/comparator[00:18:02] Lean4Leanhttps://github.com/digama0/lean4lean[00:19:50] ARC-AGI-3https://arcprize.org/arc-agi/3[00:21:59] lean-ziphttps://github.com/kim-em/lean-zip[00:29:45] CompCerthttps://compcert.org/[00:29:45] seL4https://www.sel4.org/[00:30:23] Z3https://github.com/Z3Prover/z3[00:38:10] Veilhttps://github.com/verse-lab/veil[00:38:10] Velvethttps://github.com/verse-lab/velvetorganization:[00:00:52] Lean FROhttps://lean-lang.org/fro/[00:42:39] Mathlib Initiativehttps://mathlib-initiative.org/about/person:[00:05:11] Ilya Sergeyhttps://ilyasergey.net/[00:10:31] Joachim Breitnerhttps://www.joachim-breitner.de/[00:32:58] Adam Chlipalahttps://adam.chlipala.net/[00:38:42] Kevin Buzzardhttps://www.ma.imperial.ac.uk/~buzzard/[01:10:16] Terence Taohttps://terrytao.wordpress.com/book:[00:08:19] The Proof in the Codehttps://us.macmillan.com/books/9780374620059/theproofinthecode/other:[00:10:05] Brandolini's lawhttps://en.wikipedia.org/wiki/Brandolini%27s_law[00:45:33] A new result on unit distanceshttps://openai.com/index/model-disproves-discrete-geometry-conjecture/[01:01:08] Fermat's Last Theorem formalisationhttps://imperialcollegelondon.github.io/FLT/paper:[00:34:36] The Lean 4 theorem prover and programming languagehttps://doi.org/10.1007/978-3-030-79876-5_37---RESCRIPT:https://app.rescript.info/share/7d3d4a0059443236a01f6c9acbf4db58https://app.rescript.info/api/public/sessions/b007264c0ce89047/pdf
Weco let an AI coding agent rewrite the harness around another agent for eight days: its code, prompts and tools, while the underlying language model stayed fixed. Tim Scarfe asks Weco co-founder Zhengyao Jiang what the reported gains over two years of human engineering actually demonstrate.The discussion examines AIDE 85's generated code, held-out evaluation and the difficulty of separating useful discoveries from reward hacking. Jiang explains Weco's four levels of recursive self-improvement and compares the experiment with AlphaEvolve and the Darwin Gödel Machine.The limits matter as much as the gains. Jiang explains why the experiment did not establish that the system had become a better improver. The conversation closes with open-ended search, human-designed primitives and Parameter Golf: where does the next useful idea come from when the agent is searching inside a space that people designed?---TIMESTAMPS:00:00:00 Eight days of self-improvement: what counts?00:03:25 AIDE and the puzzle of useful spaghetti code00:08:38 Four levels of recursive self-improvement00:12:02 What AIDE 85 changed and how it was tested00:20:04 AlphaEvolve, Darwin Gödel Machine and the RSI claim00:26:21 Reward hacking and the limits of detection00:33:09 Open-ended search, harness tuning and creativity00:39:43 Parameter Golf and the limits of self-improvement---REFERENCES:organization:[00:00:30] Weco AIhttps://www.weco.ai/other:[00:00:33] AIDE²: The First Evidence of Recursive Self-Improvementhttps://www.weco.ai/blog/first-evidence-of-recursive-self-improvement[00:14:11] Faulty reward functions in the wildhttps://openai.com/index/faulty-reward-functions/[00:29:59] The Hugging Face incident and the road aheadhttps://openai.com/index/hugging-face-incident-and-the-road-ahead/tool:[00:03:29] AIDEhttps://github.com/WecoAI/aideml[00:04:29] MLE-benchhttps://github.com/openai/mle-bench[00:04:33] ALE-Benchhttps://github.com/SakanaAI/ALE-Bench[00:04:52] WeatherBench 2https://github.com/google-research/weatherbench2[00:08:18] ReActhttps://react-lm.github.io/[00:39:43] Parameter Golfhttps://github.com/openai/parameter-golfpaper:[00:20:08] AlphaEvolve: A coding agent for scientific and algorithmic discoveryhttps://arxiv.org/abs/2506.13131v1[00:21:35] Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agentshttps://arxiv.org/abs/2505.22954v3[00:23:45] Hyperagentshttps://arxiv.org/abs/2603.19461v1[00:27:01] SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agentshttps://arxiv.org/abs/2605.21384book:[00:33:14] Why Greatness Cannot Be Planned: The Myth of the Objectivehttps://link.springer.com/book/10.1007/978-3-319-15524-1---LINKS:https://app.rescript.info/share/3a9dc6189cb539c6a05fcc4f75c101b3PDF:https://app.rescript.info/api/public/sessions/9eda60ede2b31c92/pdf
Frank Hutter, co-founder of Prior Labs, talks about TabPFN, a tabular foundation model that makes predictions in a single forward pass, and the research behind it.TabPFN is pre-trained on synthetic datasets drawn from a prior over structural causal models, rather than on real data. At prediction time it takes the whole training table as context and outputs an approximation of the Bayesian posterior predictive distribution, without per-dataset training or hyperparameter search. Frank explains how this grew out of his earlier work on AutoML and neural architecture search, how the priors are built and revised, and why tabular data was hard for deep learning for so long.The conversation also covers the TabArena benchmark, how the architecture changed from TabPFN v1 to v3, scaling to larger tables, using the model with coding agents, test-time compute, Google's TabFM, causal inference and interventions, and relational data. At the end, a short update Frank recorded after the interview covers the TabPFN-3.5 release.TOC:00:00 Introduction00:44 Welcome and Frank's background02:05 Why tabular data was hard for deep learning10:17 Pre-training on synthetic data12:52 The TabArena benchmark19:28 From AutoML to neural architecture search26:34 TabPFN as a learned algorithm30:50 Bayesian prediction in one forward pass39:37 Scaling to larger tables47:48 Using TabPFN with coding agents57:47 Output heads and architecture from v1 to v31:05:29 Test-time compute and adaptation1:13:32 Google's TabFM1:16:53 How the priors are designed1:18:40 Correlation, causation and interventions1:35:22 Relational and multimodal data1:38:31 Use in organisations1:46:38 The open research arm1:50:21 Update: TabPFN-3.5REFS:TabPFN v2, Nature (Hollmann et al., 2025)https://www.nature.com/articles/s41586-024-08328-6Transformers Can Do Bayesian Inference (Müller et al.)https://arxiv.org/abs/2112.10510TabArena (Erickson et al.)https://arxiv.org/abs/2506.16791AutoGluon-Tabular (Erickson et al.)https://arxiv.org/abs/2003.06505Beyond IID: How General Are Tabular Foundation Models, Really?https://arxiv.org/abs/2606.30410Neural Architecture Search: A Survey (Elsken, Metzen & Hutter)https://arxiv.org/abs/1808.05377Auto-WEKA (Thornton et al.)https://www.cs.ubc.ca/~hutter/papers/AutoWEKA-KDD2013.pdfTabPFN v1 (Hollmann et al., 2022)https://arxiv.org/abs/2207.01848TabPFN-3 technical reporthttps://arxiv.org/abs/2605.13986TabPFN-2.5 reporthttps://arxiv.org/abs/2511.08667CAAFE (Hollmann et al.)https://arxiv.org/abs/2305.03403TabICL (Qu et al.)https://arxiv.org/abs/2502.05564TabICLv2 (Qu et al.)https://arxiv.org/abs/2602.11139Google TabFMhttps://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/TALENT benchmark (Ye et al.)https://arxiv.org/abs/2407.00956Do-PFN (Robertson et al.)https://arxiv.org/abs/2506.06039CausalPFN (Balazadeh et al.)https://arxiv.org/abs/2506.07918Causal Foundation Models with Partial Graphs (Reuter et al.)https://arxiv.org/abs/2602.14972RelBench (Robinson et al.)https://arxiv.org/abs/2407.20060RelArena-α, TabPFN-Rel and RPIhttps://arxiv.org/abs/2608.16319TabPFN on GitHubhttps://github.com/PriorLabs/TabPFNTabPFN-3.5 technical reporthttps://arxiv.org/abs/2609.17895Otto Group Product Classification Challenge (Kaggle, 2015)https://www.kaggle.com/competitions/otto-group-product-classification-challengePrior Labs:TabPFN-3.5: https://priorlabs.ai/tabpfn-3-5Careers at Prior Labs: https://priorlabs.ai/careers
Alexander Mattick is a researcher at Fraunhofer IIS and a PhD researcher at the University of Technology Nuremberg (UTN), and a regular on Yannic Kilcher's Discord. He first came on MLST in 2022, after helping research the Yann LeCun and Randall Balestriero episode on interpolation.SPONSOR:---Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Applications for Batch 1 are now open.Apply now: https://cyber.fund---Alexander treats inference as the thread running through modern machine learning: once you have a model, what does it cost to get an answer out of it? He works through Monte Carlo, GFlowNets, energy-based models, diffusion, normalising flows and flow matching, with four short explainers he recorded himself. He is blunt about energy-based models: you can sample from them in principle, but it is rarely worth the compute. JEPA and "world model", he says, are closer to branding than to technical categories.Next: theories of deep learning, none of which he thinks predicts enough yet to guide practice, then reinforcement learning. ---0:00 Cold open: information is expensive0:51 Welcome back, Alexander Mattic2:08 Alexander's research background2:50 Inference: densities, sampling and Monte Carlo6:42 GFlowNets, energy functions and MCMC9:45 Explainer: energy-based models11:03 Why model a density at all?17:30 From learned energies to flow matching25:08 Explainers: diffusion and normalising flows28:33 Are energy-based models generative?33:22 JEPA, contrastive learning and collapse41:13 Why non-language modalities need flows44:51 Inference as search: branch and bound49:43 Q-learning and delayed consequences55:14 Flow matching, optimal transport, Fokker-Planck1:00:03 Explainer: flow matching1:01:49 AlphaFold, latents and scale versus architecture1:07:52 Two families of deep learning theory1:15:04 What a good theory would predict1:23:53 The manifold hypothesis and compression1:28:25 Is reward enough?1:32:01 Control theory versus reinforcement learning1:37:22 The Bitter Lesson and expensive information1:42:08 Constrained RL: the constrained MDP toolbox1:50:12 Creativity as constrained search1:55:44 Reality is protean: when abstractions hold2:00:32 What is a world model?2:04:38 Prediction is not control2:08:13 Robot demos, MPC and reliability---REFERENCES:[6:55] GFlowNets (Bengio et al., 2021)https://arxiv.org/abs/2106.04399[38:46] Contrastive Self-Supervised Learning (Anand, 2020)https://ankeshanand.com/blog/2020/01/26/contrative-self-supervised-learning.html[38:56] LeJEPA (Balestriero and LeCun, 2025)https://arxiv.org/abs/2511.08544v3[47:10] RL for Node Selection in Branch-and-Bound (Mattick)https://openreview.net/forum?id=0ez68a5UqI[56:20] Flow Matching for Generative Modeling https://arxiv.org/abs/2210.02747v2[1:12:41] Disentangling feature and lazy training in deep neural networkshttps://arxiv.org/abs/1906.08034v4[1:31:05] Reward is enough (Silver)https://doi.org/10.1016/j.artint.2021.103535[1:35:12] Learning ReLU networks to high uniform accuracy is intractable (Berner et al.)https://arxiv.org/abs/2205.13531v2[1:40:20] Dota 2 with Large Scale Deep RL https://arxiv.org/abs/1912.06680v1[1:45:41] Constrained Update Projection for Safe Policy Optimization (Yang et al., 2022)https://arxiv.org/abs/2209.07089[1:46:11] SafeMPO (ICLR 2026)https://openreview.net/forum?id=1m0EU6QXj6[1:50:17] Why Creativity Cannot Be Interpolatedhttps://archive.mlst.ai/paper/why-creativity-cannot-be-interpolated/[1:51:39] Invalid Action Masking (Huang and Ontañón)https://arxiv.org/abs/2006.14171[2:00:04] Probability Theory: The Logic of Science (Jaynes, 2003)https://www.cambridge.org/core/books/probability-theory/9CA08E224FF30123304E6D8935CF1A99[2:01:53] Training Agents Inside of Scalable World Models (Hafner et al., 2025)https://arxiv.org/abs/2509.24527v1[2:03:43] World Models (Ha and Schmidhuber, 2018)https://arxiv.org/abs/1803.10122v4
The car making a left turn at the start of this episode was never filmed. Cosmos 3 generated it. Ming-Yu Liu, who leads the Cosmos research at NVIDIA, explains how one model can describe a video, generate one, and produce robot actions.He walks Tim through the architecture. A vision language model reasons one token at a time; its weights then initialise a bidirectional diffusion generator for video, audio and action, and a shared temporal position scheme lines up signals that run at different rates. Ming-Yu treats "world model" as a set of tools, not one definition: forward dynamics, inverse dynamics and policy, trained together under a capacity limit so that each helps the others. He also explains why plentiful first-person human video carries over to robots, which have far less data of their own, and why a Cosmos model post-trained on the DROID dataset is a good starting point for pick-and-place policies.The most practical thread is testing. A neural simulator does not need accurate success rates. It only needs to rank policy A above policy B the way the real world would, so a team can narrow down which checkpoints deserve a real trial. Cosmos Dreams applies that closed-loop idea to driving and robotics, and Ming-Yu argues that humanoids around children and pets make safety matter even more than it does for cars. The conversation ends on the Super, Nano and Edge sizes (Edge targets Jetson Thor, Orin and DGX Spark) and where to find the open weights, code and data.This episode is a paid partnership with NVIDIA.Learn more about Cosmos: https://nvda.ws/4cJoY1SExplore Cosmos Lab: https://research.nvidia.com/labs/cosmos-lab/cosmos3/---TIMESTAMPS:00:00:00 A road that was never filmed00:02:28 Inside Cosmos 3: reasoning and generator towers00:05:02 World models: dynamics, policy and one clock00:08:59 Learning robot skills from human video00:11:06 Ambiguous tasks and system 2 planning00:12:53 Neural simulators for policy verification00:16:41 Cosmos as a starting point for robot policies00:19:00 Cosmos Dreams and robot safety00:22:04 Super, Nano and Edge model sizes00:24:24 Open models, the Cosmos repo and feedback---REFERENCES:tool:[00:00:13] Cosmos 3 (NVIDIA Cosmos Lab project page)https://research.nvidia.com/labs/cosmos-lab/cosmos3/[00:18:27] NVIDIA Cosmos GitHub repositoryhttps://github.com/NVIDIA/cosmos[00:22:05] Cosmos3-Edge model cardhttps://huggingface.co/nvidia/Cosmos3-Edge[00:22:15] Cosmos3-Super model cardhttps://huggingface.co/nvidia/Cosmos3-Super[00:22:16] Cosmos3-Nano model cardhttps://huggingface.co/nvidia/Cosmos3-Nano[00:22:50] NVIDIA Jetson Thorhttps://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/[00:22:52] NVIDIA Jetson Orinhttps://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/[00:22:53] NVIDIA DGX Sparkhttps://www.nvidia.com/en-us/products/workstations/dgx-spark/[00:24:42] Cosmos 3 collection on Hugging Facehttps://huggingface.co/collections/nvidia/cosmos3other:[00:01:07] Cosmos-Dreams closed-loop simulators (NVIDIA SIGGRAPH 2026 blog)https://blogs.nvidia.com/blog/siggraph-news-2026/paper:[00:08:54] Cosmos 3: Omnimodal World Models for Physical AIhttps://arxiv.org/abs/2606.02800[00:17:43] DROID: A Large-Scale In-The-Wild Robot Manipulation Datasethttps://arxiv.org/abs/2403.12945---RESCRIPT: https://app.rescript.info/share/e2385948cf465f0d6a2c0930150fc3ab
Pavankumar Reddy Muddireddy leads audio research at Mistral AI. He joins Tim Scarfe for a deep technical tour of Voxtral — and explains why the frontier of deployed voice is still a cascade of specialised models rather than one end-to-end system.IN PARTNERSHIP WITH MISTRAL AI:---This episode was produced in partnership with Mistral AI.Mistral AI: https://mistral.ai/---The conversation opens on architecture. Voxtral Chat feeds a 3B Ministral text trunk with continuous embeddings from an audio encoder, passed to the decoder as direct token input rather than through cross-attention as in Whisper, so the model can answer questions about emotion, timing and who spoke when without an intermediate transcript to lose them. The real-time model becomes a dual-stream decoder that consumes audio and emits text at once, at a target delay down to 160ms, with slower streams in parallel for anything that can wait for more context.On generation, Pavan explains why Voxtral TTS predicts continuous latents rather than discrete codec tokens, traces the lineage from SoundStream through EnCodec to Mimi's split of semantic and acoustic codebooks, and places FSQ and flow matching in it. Tim presses on the priors underneath: why a mel spectrogram instead of raw waveform, what noise augmentation buys, and when acoustic overfitting becomes somebody's fine-tuning problem. Then the failure modes. Diarisation is emitted autoregressively inside the transcript rather than by a separate head, which makes streaming diarisation fragile — less context, late speaker changes, invented extra speakers. And because the architecture commits to what it has already predicted, one out-of-distribution mistake compounds into looping or skipped segments, which is what DPO corrects: the negative supervision pre-training and SFT cannot give.The last third is the argument Tim keeps returning to. Customers running voice agents over millions of sessions describe scaffolding, not a solved problem, with a sharp drop outside the top few languages. Cascades survive because each component stays separately adaptable, observable and constrainable. And voice alone is cognitive debt: absorbing information and deciding in one serial stream is harder than glancing at a menu. Voice becomes ubiquitous beside a screen, not instead of one.---TIMESTAMPS:00:00:00 Cold open00:00:46 Why Mistral moved into audio00:09:27 Inside Voxtral: trunk, encoder, dual streams00:20:22 Speech that works in real time00:30:52 How a voice becomes tokens00:39:59 Flow matching, FSQ and the new codec00:52:51 When speech models lose the speaker01:03:23 Correcting hallucinations with preferences01:12:12 Controlling synthetic speech01:20:06 Why cascades still win01:29:25 Speech in the wild01:33:46 Audio models as interfaces01:37:54 Why voice still needs a screen---REFERENCES:paper:[00:01:42] Mistral 7Bhttps://arxiv.org/abs/2310.06825[00:09:38] Voxtralhttps://arxiv.org/abs/2507.13264[00:14:41] Whisper: Robust Speech Recognitionhttps://arxiv.org/abs/2212.04356[00:19:11] Voxtral Realtimehttps://arxiv.org/abs/2602.11298[00:21:52] Delayed Streams Modeling (Kyutai)https://arxiv.org/abs/2509.08753[00:30:52] Voxtral TTShttps://arxiv.org/abs/2603.25551[00:32:38] SoundStream neural audio codechttps://arxiv.org/abs/2107.03312[00:34:59] Flow Matching for Generative Modelinghttps://arxiv.org/abs/2210.02747[00:37:03] EnCodec: High Fidelity Neural Audio Compressionhttps://arxiv.org/abs/2210.13438[00:37:42] Moshi and the Mimi codechttps://arxiv.org/abs/2410.00037[00:39:05] Finite Scalar Quantization (FSQ)https://arxiv.org/abs/2309.15505[01:03:33] Direct Preference Optimization (DPO)https://arxiv.org/abs/2305.18290dataset:[00:46:14] Mozilla Common Voicehttps://commonvoice.mozilla.org/en/datasetsorganization:[00:50:47] Hugging Facehttps://huggingface.co/
Can a machine learn the judgement that separates a plausible-looking result from a faithful experiment? Edward Hughes, Chief Scientist and co-founder of Inherent, joins Tim Scarfe to argue that creativity is not optimisation, and that the missing capability in AI is choosing which questions are worth asking.SPONSOR:---Cyber Fund built the Monastery to help founders ship products that were impossible a year ago.Apply now: https://cyber.fund---Edward makes the case that Move 37 was innovative rather than creative, and that the field, not the individual, decides what counts as a discovery. That reframing runs through Csikszentmihalyi, Deutsch and exaptation into open-endedness, where deceptive goals and imperfect world models turn out to be the point rather than the problem. The second half turns to the paper: Replica, a task space built by redacting figures from real papers, and Faraday, a 27-billion-parameter model trained to steer a frontier coding agent that then beats the frontier on held-out replications.---TIMESTAMPS:00:00:00 Cold open: Move 37, Faraday and collective intelligence00:01:08 Sponsor: CyberFund00:01:46 Inherent's $50M raise and the road from string theory00:09:14 Three timescales of learning: weights, context, culture00:13:47 Move 37 was innovative, not creative: the field decides00:20:39 Creativity as satisficing: the urinal and evolution00:25:06 Exaptation and the Tristan chord: creativity in context00:30:56 Coherence for whom? Deutsch's hard-to-vary explanations00:35:53 Why copying is creative: Deutsch and the constraint engineer00:42:27 Societies of agents and the strong Moravec paradox00:45:51 Evaluate in hindsight: from Lean proofs to climate change00:51:56 Picbreeder, local goals and why discovery needs deception00:57:21 Spaghetti proofs, translation layers and superhuman Go01:00:37 Does nature compress? Naturalness and real patterns01:07:36 Why replicate? Replica's redacted figures and Faraday01:12:31 Faraday beats Codex, Claude and GLM 5.2 on held-out tasks01:15:31 Replication to innovation: how the Transformer happened01:18:26 Deep replication: what Faraday learns from Voyager and GNoME01:23:37 Can the AI scientist cheat? Goodharting the judge01:29:09 Inside Replica: scale-down, 8xB300 runs, per-task rubrics01:34:11 The RL crisis: getting GRPO to work with per-turn credit01:39:43 Weights vs harnesses: AlphaEvolve, DGM and EvoTune01:45:45 The recursive company: agents cross a phase transition01:50:35 Collective intelligence and the electric dynamo01:55:46 What replaces OKRs? Incumbents and the burden of knowledge---REFERENCES:organization:[00:01:47] Inherenthttps://inherentlabs.ai/other:[00:20:51] Marcel Duchamp, Fountain (1917)https://www.tate.org.uk/art/artworks/duchamp-fountain-t07573[00:57:33] OpenAI unit distancehttps://openai.com/index/model-disproves-discrete-geometry-conjecture/[00:05:19] Human-Timescale Adaptation in an Open-Ended Task Space (Adaptive Agent)https://arxiv.org/abs/2301.07608[00:06:05] The AI Scientisthttps://arxiv.org/abs/2408.06292[00:12:13] Training AI Scientists to Replicate Research (Replica and Faraday)https://arxiv.org/abs/2608.13331[01:44:46] Evolutionary Principles in Self-Referential Learninghttps://people.idsia.ch/~juergen/diploma.html[01:59:33] Are Ideas Getting Harder to Find?https://www.nber.org/papers/w23782book:[00:16:04] Creativity: Flow and the Psychology of Discovery and Inventionhttps://search.worldcat.org/title/254487436[00:26:22] Why Greatness Cannot Be Plannedhttps://link.springer.com/book/10.1007/978-3-319-15524-1[00:33:03] The Beginning of Infinityhttps://www.penguinrandomhouse.com/books/293575/the-beginning-of-infinity-by-david-deutsch/[01:55:47] Laws of Knowledgehttps://www.penguin.co.nz/books/the-infinite-alphabet-9780241655672(Full list refs on YT/rescript)---RESCRIPT:https://app.rescript.info/session/670296ba913761d0?share=6281911cac9bdbff637f10819d4d1e5c
Free AI-powered daily recaps. Key takeaways, quotes, and mentions — in a 5-minute read.
Get Free Summaries →Free forever for up to 3 podcasts. No credit card required.
Listeners also like.

Latent Space: The AI Engineer Podcast
Explores AI engineering breakthroughs in foundation models, code generation, and AI agents through interviews with researchers and developers.

Everyday AI Podcast – An AI and ChatGPT Podcast
Practical AI and ChatGPT tips for professionals to improve productivity and grow their careers.

Artificial Intelligence Masterclass
Examines the impact of advanced artificial intelligence on ethics, society, and the future of human labor.

Intelligent Machines (Audio)
Explores the impact of artificial intelligence through conversations with pioneers shaping the future of intelligent technology.

"The Cognitive Revolution"
Interviews with AI developers and researchers exploring the transformative impact of artificial intelligence on society and technology.

The AI Daily Brief: Artificial Intelligence News and Analysis
A daily analysis of artificial intelligence news, exploring its creative potential, industry impacts, and ethical challenges.

80,000 Hours Podcast
Discusses artificial intelligence and global catastrophic risks with experts and researchers.

Me, Myself, and AI
AI leaders from top companies share real-world strategies for turning artificial intelligence into measurable business results.

Primary Technology
Tech news covering consumer gadgets, AI, and major industry stories explained for a general audience.

How I AI
A practical guide to using AI tools in work and life, featuring guests who share specific, actionable techniques and workflows.

Unsupervised Learning with Jacob Effron
Conversations with leading AI experts to understand current breakthroughs and future implications for technology and business.

OpenAI Podcast
Conversations with OpenAI researchers and builders exploring how frontier AI models are developed and used in practice.
We engage in fascinating discussions with pre-eminent figures in the AI field. Our flagship show covers current affairs in AI, cognitive science, neuroscience and philosophy of mind with in-depth analysis. Our approach is unrivalled in terms of scope and rigour – we believe in intellectual diversity in AI, and we touch on all of the main ideas in the field with the hype surgically removed. MLST is run by Tim Scarfe, Ph.D and features regular appearances from MIT Doctor of Philosophy Keith Duggar.
AI-powered recaps with compact key takeaways, quotes, and insights.
Get key takeaways from Machine Learning Street Talk (MLST) in a 5-minute read.
Stay current on your favorite podcasts without falling behind.
It's a free AI-powered email that summarizes new episodes of Machine Learning Street Talk (MLST) as soon as they're published. You get the key takeaways, notable quotes, and links & mentions — all in a quick read.
When a new episode drops, our AI transcribes and analyzes it, then generates a personalized summary tailored to your interests and profession. It's delivered to your inbox every morning.
No. Podzilla is an independent service that summarizes publicly available podcast content. We're not affiliated with or endorsed by Machine Learning Street Talk (MLST).
Absolutely! The free plan covers up to 3 podcasts. Upgrade to Pro for 15, or Premium for 50. Browse our full catalog at /podcasts.
Machine Learning Street Talk (MLST) publishes weekly. Our AI generates a summary within hours of each new episode.
Machine Learning Street Talk (MLST) covers topics including Technology. Our AI identifies the specific themes in each episode and highlights what matters most to you.
Free forever for up to 3 podcasts. No credit card required.
Free forever for up to 3 podcasts. No credit card required.