Last Week in AI podcast | Listen online for free

Available Episodes

5 of 245

#206 - Llama 4, Nova Act, xAI buys X, PaperBench
Our 206th episode with a summary and discussion of last week's big AI news! Recorded on 04/07/2025 Try out the Astrocade demo here! https://www.astrocade.com/ Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. Join our Discord here! https://discord.gg/nTyezGSKwP In this episode: Meta releases LlAMA-4, a series of advanced large language models, sparking debate on performance and release timing, with models featuring up to 2 trillion parameters for different configurations and applications. Amazon's AGI Lab debuts NOVA Act, an AI agent for web browser control, boasting competitive benchmarking against OpenAI's and Anthropic's best agents. OpenAI's image generation capabilities and ongoing financing developments, notably a $40 billion funding round led by SoftBank, highlight significant advancements and strategic shifts in the tech giant’s operations. Timestamps + Links: (00:00:00) Intro / Banter Tools & Apps (00:01:46) Meta releases Llama 4, a new crop of flagship AI models (00:13:55) Amazon unveils Nova Act, an AI agent that can control a web browser (00:17:06) Alibaba Preparing for Flagship AI Model Release as Soon as April (00:17:59) Runway releases an impressive new video-generating AI model (00:19:10) Adobe launches Premiere Pro’s generative AI video extender (00:20:54) OpenAI prepares reasoning slider and memory update for ChatGPT users Applications & Business (00:21:28) Nvidia H20 Chips: $16 Billion Orders from ByteDance, Alibaba, and Tencent (00:24:45) Elon Musk sells X for $33 billion to his own AI startup company xAI (00:28:00) SoftBank dethroned Microsoft as OpenAI's largest investor, pushing the ChatGPT maker's market cap to $300 billion — but reportedly buried itself in debt (00:30:48) DeepMind is holding back release of AI research to give Google an edge (00:34:06) SMIC Is Rumored To Complete 5nm Chip Development By 2025; Costs Could Be Up To 50 Percent Higher Than TSMC’s Version Due To The Use Of Older-Generation Equipment (00:36:04) Google-backed Isomorphic Labs raises $600m to advance AI drug discovery Research & Advancements (00:38:03) PaperBench: Evaluating AI's Ability to Replicate AI Research (00:43:50) Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains (00:48:39) Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead (00:54:34) Overtrained Language Models Are Harder to Fine-Tune Policy & Safety (00:58:28) Taking a responsible path to AGI (01:02:32) This A.I. Forecast Predicts Storms Ahead (01:06:24) The Secrets and Misdirection Behind Sam Altman’s Firing From OpenAI
--------
1:13:44
#205 - Gemini 2.5, ChatGPT Image Gen, Thoughts of LLMs
Our 205th episode with a summary and discussion of last week's big AI news! Recorded on 03/28/2025 Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. Join our Discord here! https://discord.gg/nTyezGSKwP In this episode: OpenAI's new image generation capabilities represent significant advancements in AI tools, showcasing impressive benchmarks and multimodal functionalities. OpenAI is finalizing a historic $40 billion funding round led by SoftBank, and Sam Altman shifts focus to technical direction while COO Brad Lightcap takes on more operational responsibilities., Anthropic unveils groundbreaking interpretability research, introducing cross-layer tracers and showcasing deep insights into model reasoning through applications on Claude 3.5. New challenging benchmarks such as ARC AGI 2 and complex Sudoku variations aim to push the boundaries of reasoning and problem-solving capabilities in AI models. Timestamps + Links: (00:00:00) Intro / Banter (00:01:01) News Preview Tools & Apps (00:02:46) Gemini 2.5: Our most intelligent AI model (00:08:41) OpenAI rolls out image generation powered by GPT-4o to ChatGPT (00:16:14) Ideogram presents version 3.0 of its AI image generation system (00:19:20) New Reve Image Generator Beats AI Art Heavyweights MidJourney and Flux at a Penny Per Image (00:21:56) Alibaba Releases Qwen2.5 Omni, Adds Voice and Video Modes to Qwen Chat (00:23:58) The official version of Tencent's Hunyuan Deep Thinking Model T1 is here, with fast articulation, instant responses, and a decoding speed increase of 2 times Applications & Business (00:25:45) OpenAI Close to Finalizing $40 Billion SoftBank-Led Funding (00:29:26) OpenAI reshuffles leadership as Sam Altman pivots to technical focus (00:33:23) Nvidia shows off Rubin Ultra with 600,000-Watt Kyber racks and infrastructure, coming in 2027 (00:35:23) China's SiCarrier emerges as challenger to ASML, other chip tool titans (00:38:24) Pony.ai wins first permit for fully driverless taxi operation in the center of China’s Silicon Valley Projects & Open Source (00:40:27) A new, challenging AGI test stumps most AI models (00:45:16) Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models (00:48:13) Wan: Open and Advanced Large-Scale Video Generative Models (00:50:38) DeepSeek V3-0324 tops non-reasoning AI models in open-source first (00:54:46) OpenAI adopts rival Anthropic’s standard for connecting AI models to data Research & Advancements (00:55:56) Anthropic can now track the bizarre inner workings of a large language model (01:06:00) Chain-of-Tools: Utilizing Massive Unseen Tools in the CoT Reasoning of Frozen Language Models (01:11:50) Inside-Out: Hidden Factual Knowledge in LLMs (01:15:14) Sakana AI super-powers AI reasoning using Japan’s own Sudoku Puzzles Policy & Safety (01:18:38) Senator Wiener Introduces Legislation to Protect AI Whistleblowers & Boost Responsible AI Development (01:21:50) NVIDIA & Other Tech Giants Demand Trump Administration To Reconsider “AI Diffusion” Policy Which Is Set To Be Effective By May 15 (01:23:17) U.S. blacklists over 50 Chinese companies in bid to curb Beijing's AI, chip capabilities (01:26:44) Netflix’s Reed Hastings Gives $50 Million to Bowdoin for A.I. Program (01:27:55) Judge allows 'New York Times' copyright case against OpenAI to go forward (01:29:48) Judge rules that AI can continue training on copyrighted lyrics, for now
--------
1:34:18
#204 - OpenAI Audio, Rubin GPUs, MCP, Zochi
Our 204th episode with a summary and discussion of last week's big AI news! Recorded on 03/21/2025 Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. Join our Discord here! https://discord.gg/nTyezGSKwP In this episode: Baidu launched two new multimodal models, Ernie 4.5 and Ernie X1, boasting competitive pricing and capabilities compared to Western counterparts like GPT-4.5 and DeepSeek R1. OpenAI introduced new audio models, including impressive speech-to-text and text-to-speech systems, and added O1 Pro to their developer API at high costs, reflecting efforts for more profitability. Nvidia and Apple announced significant hardware advancements, including Nvidia's future GPU plans and Apple's new Mac Studio offering that can run DeepSeek R1. DeepSeek employees are facing travel restrictions, suggesting China is treating its AI development with increased secrecy and urgency, emphasizing a wartime footing in AI competition. Timestamps + Links: (00:00:00) Intro / Banter (00:01:36) News Preview Tools & Apps (00:02:50) Baidu launches two new versions of its AI model Ernie (00:10:46) OpenAI Unveils New Audio Models to Make AI Agents Sound More Human Than Ever (00:16:41) OpenAI’s o1-pro is the company’s most expensive AI model yet (00:20:53) Google brings a ‘canvas’ feature to Gemini, plus Audio Overview (00:22:18) Anthropic adds web search to its Claude chatbot (00:23:55) xAI launches an API for generating images Applications & Business (00:26:28) Nvidia announces Rubin GPUs in 2026, Rubin Ultra in 2027, Feynman also added to roadmap (00:36:25) M3 Ultra Runs DeepSeek R1 With 671 Billion Parameters Using 448GB Of Unified Memory, Delivering High Bandwidth Performance At Under 200W Power Consumption, With No Need For A Multi-GPU Setup (00:40:07) Intel reaches 'exciting milestone' for 18A 1.8nm-class wafers with first run at Arizona fab (00:42:45) Elon Musk’s AI company, xAI, acquires a generative AI video startup (00:44:44) Tencent Reportedly Makes Massive NVIDIA H20 Chip Purchase for WeChat’s DeepSeek Integration Projects & Open Source (00:46:32) Anthropic’s Not-So-Secret Weapon That’s Giving Agents a Boost (00:50:50) Mistral AI drops new open-source model that outperforms GPT-4o Mini with fraction of parameters (00:53:30) EXAONE Deep: Reasoning Enhanced Language Models Research & Advancements (00:55:58) Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification (01:07:44) Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models (01:12:27) Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo (01:18:46) Transformers without Normalization (01:19:52) Measuring AI Ability to Complete Long Tasks (01:26:12) HCAST: Human-Calibrated Autonomy Software Tasks Policy & Safety (01:26:45) Announcing Zochi, an Intology Project (01:32:46) DeepSeek, a National Treasure in China, is Now Being Closely Guarded (01:37:02) Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations Synthetic Media & Art (01:42:27) US appeals court rejects copyrights for AI-generated art lacking 'human' creator (01:45:10) Trump urged by Ben Stiller, Paul McCartney and hundreds of stars to protect AI copyright rules
--------
1:49:03
#203 - Gemini Image Gen, Ascend 910C, Gemma 3, Gemini Robotics
Our 203rd episode with a summary and discussion of last week's big AI news! Recorded on 03/14/2025 Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. Join our Discord here! https://discord.gg/nTyezGSKwP In this episode: OpenAI's new 'deep research' feature has raised concerns about cybersecurity and the potential misuse of AI models for bio-weapons and autonomous capabilities, prompting new safety and governance measures. Google's extensive $3 billion investment in Anthropic is revealed, aligning with their AI strategy and reinforcing the importance of multiple technology partnerships. Huawei's advancements in the AI chip industry are highlighted, with significant progress in producing chips comparable to Nvidia's H100, despite export control challenges. China's recent directive discourages AI executives from traveling to the US, reflecting heightened security concerns and potentially signaling a more adversarial stance in the AI race. Timestamps + Links: (00:00:00) Intro / Banter (00:01:30) News Preview Tools & Apps (00:02:30) OpenAI launches new tools to help businesses build AI agents (00:08:50) You can now test Gemini 2.0 Flash’s native image output (00:13:32) Waymo is now offering 24/7 robotaxi rides in Silicon Valley (00:17:19) Moonvalley releases a video generator it claims was trained on licensed content (00:21:11) Snap introduces AI Video Lenses powered by its in-house generative model (00:23:37) Sudowrite Launches Muse AI Model That Can Generate Narrative-Driven Fiction Applications & Business (00:27:48) In another chess move with Microsoft, OpenAI is pouring $12B into CoreWeave (00:30:54) Huawei’s Ascend 910C Takes on NVIDIA as China’s AI Race Heats Up: More Alleged Details (00:36:26) Huawei reportedly acquired two million Ascend 910 AI chips from TSMC last year through shell companies (00:40:27) Inside Google’s Investment in the A.I. Start-Up Anthropic (00:43:26) Meta is reportedly testing in-house chips for AI training (00:46:48) Elon Musk's xAI buys 1 million sq ft site for second Memphis data center (00:50:02) Superintelligence startup Reflection AI launches with $130M in funding Projects & Open Source (00:53:11) Google calls Gemma 3 the most powerful AI model you can run on one GPU (00:58:18) Sesame, the startup behind the viral virtual assistant Maya, releases its base AI model (01:01:13) Reka AI Open Sourced Reka Flash 3: A 21B General-Purpose Reasoning Model that was Trained from Scratch (01:04:19) Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k Research & Advancements (01:06:25) Google’s Gemini Robotics AI Model Reaches Into the Physical World (01:14:33) Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning (01:23:29) Deep Research System Card (01:29:50) Claude 3.7 Sonnet System Card Policy & Safety (01:33:24) Detecting misbehavior in frontier reasoning models (01:39:30) China tells its AI leaders to avoid US travel over security concerns, WSJ reports (01:43:48) Outro
--------
1:46:23
#202 - Qwen-32B, Anthropic's $3.5 billion, LLM Cognitive Behaviors
Our 202nd episode with a summary and discussion of last week's big AI news! Recorded on 03/07/2025 Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. Join our Discord here! https://discord.gg/nTyezGSKwP In this episode: Alibaba released Qwen-32B, their latest reasoning model, on par with leading models like DeepMind’s R1. Anthropic raised $3.5 billion in a funding round, valuing the company at $61.5 billion, solidifying its position as a key competitor to OpenAI. DeepMind introduced BigBench Extra Hard, a more challenging benchmark to evaluate the reasoning capabilities of large language models. Reinforcement Learning pioneers Andrew Bartow and Rich Sutton were awarded the prestigious Turing Award for their contributions to the field. Timestamps + Links: cle picks: (00:00:00) Intro / Banter (00:01:41) Episode Preview (00:02:50) GPT-4.5 Discussion (00:14:13) Alibaba’s New QwQ 32B Model is as Good as DeepSeek-R1 ; Outperforms OpenAI’s o1-mini (00:21:29) With Alexa Plus, Amazon finally reinvents its best product (00:26:08) Another DeepSeek moment? General AI agent Manus shows ability to handle complex tasks (00:29:14) Microsoft’s new Dragon Copilot is an AI assistant for healthcare (00:32:24) Mistral’s new OCR API turns any PDF document into an AI-ready Markdown file (00:33:19) A.I. Start-Up Anthropic Closes Deal That Values It at $61.5 Billion (00:35:49) Nvidia-Backed CoreWeave Files for IPO, Shows Growing Revenue (00:38:05) Waymo and Uber's Austin robotaxi expansion begins today (00:38:54) UK competition watchdog drops Microsoft-OpenAI probe (00:41:17) Scale AI announces multimillion-dollar defense deal, a major step in U.S. military automation (00:44:43) DeepSeek Open Source Week: A Complete Summary (00:45:25) DeepSeek AI Releases DualPipe: A Bidirectional Pipeline Parallelism Algorithm for Computation-Communication Overlap in V3/R1 Training (00:53:00) Physical Intelligence open-sources Pi0 robotics foundation model (00:54:23) BIG-Bench Extra Hard (00:56:10) Cognitive Behaviors that Enable Self-Improving Reasoners (01:01:49) The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems (01:05:32) Pioneers of Reinforcement Learning Win the Turing Award (01:06:56) OpenAI launches $50M grant program to help fund academic research (01:07:25) The Nuclear-Level Risk of Superintelligent AI (01:13:34) METR’s GPT-4.5 pre-deployment evaluations (01:17:16) Chinese buyers are getting Nvidia Blackwell chips despite US export controls
--------
1:19:52