『Interconnects』のカバーアート

Interconnects

Interconnects

著者: Nathan Lambert
無料で聴く

【Amazonプライム会員限定】今ならプレミアムプランが4か月 月額99円。

10月19日まで。※適用条件あり
Audio essays about the latest developments in AI and interviews with leading scientists in the field. Breaking the hype, understanding what's under the hood, and telling stories.

www.interconnects.aiInterconnects AI, LLC
科学
エピソード
  • Debating RSI, the US-China Gap, and Jaggedness with JS Denain of Epoch AI
    2026/09/22
    This episode is with Jean-Stanislas “JS” Denain of Epoch AI, who leads their Insights Team and is one of the people I find myself debating the state and trajectory of AI with more and more. We’ve had follow-on discussions of many of my favorite recent posts online and/or in private, so I wanted to dig into the nuance in a public episode.A big takeaway of this podcast is how JS and I both have so much uncertainty with exactly where we are heading, and this was our best effort at stating our observations today. Chapters / topics include:* 00:00 Predictions for RSI* 18:15 The role of robotics in an AI acceleration* 24:20 How far behind are Chinese models?* 27:39 Does distillation explain the gap?* 40:58 What Chinese job postings reveal about their labs* 48:13 Are open or closed models safer?* 58:10 How Epoch AI ticks* 1:00:55 What a frontier post-training recipe looks likeEnjoy!More from JS: Epoch AI profile and writing, X, LinkedInListen on Apple Podcasts, Spotify, and where ever you get your podcasts. For other Interconnects interviews, go here.Transcript00:00:00 Nathan Lambert: I’m here with JS Denain, who is a senior researcher at Epoch AI. He leads the insights team. He is one of the people who I feel like I get the best feedback on my writing from, whether it’s from US-China AI capabilities, now RSI. And I just wanted to open this discussion and honestly go deeper with him, trying to understand how he thinks about these various things. And I think you have a very useful, moderate point of view, which I feel like you’re probably a step further into what would be called faster scenarios for AI progress. But let’s get into this, and it’s like, what measurements do you think OpenAI and Anthropic are seeing when we get all these proclamations on RSI happening very imminently?00:00:47 JS Denain: Yeah. I think, so there’s the measurements they’ve published, right? So, OpenAI and Anthropic both had blog posts, I mean, Anthropic two at least, on the effect AI has on accelerating AI progress. I think at least the things they publish, I don’t think are super strong evidence of imminent self-sustaining acceleration AI capabilities, or full automation of the job of AI researcher. But I think the kinds of things we see are, I think probably the most striking thing I saw in the OpenAI blog post was increasing usage of AI systems in model deployment, like the increase in spending on Codex that we saw. And it’s kind of unclear how exactly to interpret this, because maybe it’s a measurement artifact where they’re only looking at Codex, but in fact, there was a bunch of ChatGPT usage before from the researchers. But overall, that plot, for example, just shows a 2X a month increase in Codex spending by researchers, and that does seem to me to be some evidence of they’re getting a lot of value out of this probably. I don’t think this is strong evidence that in six months we have a software intelligence explosion.00:01:57 Nathan Lambert: Do you think this is the same? So, what is the information they have internally relative to what we have? And this is obviously hypothetical. We don’t have this internal information. Because I get the sense that a lot of people are more scared in their updates from the labs than the information we have. And I try to take this very seriously of, what will they be seeing that is making the acceleration of risk comments go faster, and how much of this is material evidence versus how cultures evolve over time? And I’m much more interested in evidence.00:02:29 JS Denain: Yeah. So two things. I think, first of all, I guess I don’t think, I don’t know, right? I don’t have full information here. I don’t currently think that either there’s some specific thing that people at OpenAI or Anthropic are seeing right now that we don’t have access to that warrants being way more freaked out about this. I also don’t think that... I think the public evidence we have right now, more general on AI progress and just a priori case for this being an important dynamic, I think is enough. I think to care about this particular dynamic of AI accelerating AI progress, that being a big deal and worth tracking. And then it’s kind of unclear what the urgency is of when the feedback loop really kicks in. So, basically on the what is there on the inside that people have access to, I could give examples of kinds of metrics, right, they could be looking at. It’s plausible that we have access to the capabilities of AI systems, but internal teams have their KPIs, and maybe they’re seeing compute multipliers in the pre-training team or other kinds of metrics that people are tracking going crazy. And then the combination of this plus some intuitions of how the different outputs of different teams combine yields a prediction on the trend in actual performance of the end AI systems. So that could be an early warning sign. It’s unclear to me that the recent discourse we’...
    続きを読む 一部表示
    1 時間 5 分
  • The current balance of power in open models
    2026/09/21
    I was recently invited to brief a group of Congressional members and staff on the state of open-weight models in the lens of U.S.-China competition. I’m sharing my prepared remarks as a state of the union on open models that is accessible to a broader audience.Interconnects AI is a reader-supported publication. Consider becoming a subscriber.Recap: What is an open source v. open-weight vs. closed model?Open language models are AI models where their weights are publicly available for inspection or downstream use. These are most often contrasted to so-called “closed” AI models. Closed models offer access only through Application Programming Interfaces (APIs) that developers can use to directly query a model, like GPT-4 or Claude Opus 4.5, or through products, like ChatGPT and Claude Code.Open language models primarily are bucketed into two categories, open-weight and open-source models. Open-weight models are the most common form, such as popular models like Meta’s Llama, Alibaba’s Qwen, Google’s Gemma, or DeepSeek’s models. These models are governed by licenses, governing documents dictating what is allowed with downstream use, and are often accompanied by inference code in libraries such as Transformers, VLLM, SGLANG, etc. Since about April 2025, Chinese AI companies have been the clear leader in open-weight models.True “open-source” models are similar to these, as they include the weights, licenses, and inference code, but they also include the complete information needed to reproduce the model – the training code and training data. The most prominent open-source models have been built in the United States, led recently by the Allen Institute for AI’s Olmo models that I helped build in my recent 2.5 years there. The other prominent open-source models are also built by American non-profit organizations, including OpenAthena’s Marin models and EleutherAI’s Pythia models.Open-weight, open-source, and every other label for a model – including closed models primarily offered via an API – exist on a spectrum. For example, Nvidia’s Nemotron models are far more open than most open-weight models, releasing large quantities of their training data under permissive licenses, but they’re not fully open-source because they do not release all of the data. Closed models also exist on a spectrum based on what information the API reveals and the terms of use.The state of competition between American and Chinese open-weight models (unit economics, technical capabilities, etc.)We are living in a world where GLM-5.2 and Kimi K3, some of the latest, leading Chinese models, have enacted a step change in the commercial viability of open models — crossing a similar threshold in agentic capabilities that Anthropic’s Claude Code crossed in December of 2025.America was the early leader in open language models, primarily through Meta’s Llama models, which were used extensively across research and commercial tasks. Chinese open-weight models surpassed American open-weight models in these two key areas about 18 months ago. The simple metric showing this is Hugging Face Downloads, where China took the lead in July of 2025 primarily through the success of Alibaba’s Qwen models. I personally maintain tools to track this data, and since I first published the American Truly Open Models (ATOM) Project in August of 2025, China’s download lead has grown to about 1.6B – with a total of 3.2B downloads, twice that of America’s total.On popular capabilities benchmarks, such as the Artificial Analysis Intelligence Index (AAII), the Chinese open-weight models have a clear lead over American counterparts. The top three Chinese models as of writing this on September 14, 2026 are Z.ai’s GLM-5.3 and GLM-5.3-Flash and Moonshot AI’s Kimi K3 with scores of 45, 42, and 44 respectively. By comparison, the leading American models are Thinking Machines’ Inkling and Inkling Small, both with a score of 26, and Nvidia’s Nemotron 3 Ultra, with a score of 23. The top American models were released in June and July of 2026, and are updated less frequently than their Chinese counterparts. For example, Chinese labs released models with scores above these American models 2-6 months before the American companies got there (e.g. GLM-5 or DeepSeek V4 Pro). There is a trend of more American companies releasing models, including names like Arcee AI, Poolside and IBM, but they are not rapidly closing this performance gap. Other benchmarks tell a similar story.Together, Chinese open-weight models are approximately 2-5 months behind the closed American frontier, with the open-weight American models being approximately 6-9 months behind the likes of OpenAI and Anthropic. The Chinese labs are closest in tasks with clear user demand, such as agentic coding, and further behind on more open-ended scientific tasks, such as physics or biology.The reasons why Chinese labs can produce these strong models, despite having fewer ...
    続きを読む 一部表示
    18 分
  • Why I still haven’t bought into true RSI
    2026/09/19
    We’re in an era where a few organizations are using thousands of concurrent agents to improve their processes and output. These organizations happen to be just the frontier AI labs, in particular OpenAI and Anthropic. In the last few weeks, I’ve been pondering what it means for so many employees across these organizations to rapidly update their expectations for the pace of AI progress and associated risks.A core perspective I have is that the frontier labs and broader frenetic, competitive culture in the San Francisco AI scene set up an environment that amplifies any AI concern. This has some benefits in causing more general audience awareness of AI, as fear sells, but exaggerating risk timelines or severity will have negative second-order effects. I remember many loud AI safety debates, and their associated clouds over the viability of open-source AI, in 2023 and 2024 — the primary risks then did not arrive in the forecasted timelines.The general populace of these two key labs was very anxious about AI risks and the rate of progress even a year ago, and especially as agents got stronger product-market fit at the start of 2026. This cultural precondition, when exposed to the reality that thousands of agents will constantly be working fairly productively in your business, will only increase this anxiety. The step from this anxiety, and incidents like OpenAI-HuggingFace, to extinction risks feels very religious.Richard Ngo had an apt summary of the situation:Now a large proportion of the AI safety community is implicitly or explicitly orienting to futures where an intelligence explosion occurs within a few years. My default expectation (absent an extensive pause) is that a similar thing will happen: they’ll turn out to be directionally correct (relative to the expectations of almost anyone not linked to the community) but factually wrong. Specifically, we won’t have superintelligence within the next 8 years, but things will still be moving so fast that it’ll *feel* like the people who argued for short timelines were right.… I wanted to say something now because it feels like the level of bandwagoning towards “singularity soon” is getting pretty wild.Personally, I think this view aligns closely to what I outlined in my alternate scenario to true recursive self-improvement (RSI), which I called lossy self-improvement. A summary of this view is that:* Automatable research is too narrow to achieve a massive net acceleration in progress, in the face of scaling laws’ exponential costs,* Diminishing returns of more AI agents in parallel are real, &* Resource bottlenecks and politics are a major factor in building strong LLMs (and AI can do much less to accelerate this).So, I’m left balancing the above, latent increase in the cultural temperature with the potential that the labs have seen genuinely scary, specific breakthroughs that are not public yet. My expectation is that more of the current AI safety concern is on the former – scaled agents working – but I hold high levels of uncertainty here. Foundational, imagination-based AI breakthroughs are the sort of thing that would make me update my RSI timelines from closer to a tool to sustain progress in the face of exponential costs (scaling laws), to something more unpredictable and/or unstable.Interconnects AI is a reader-supported publication. Consider becoming a subscriber.Some of the best recent resources on RSI have been Dwarkesh’s podcasts with Noam Brown and the trio of John Schulman, Beren Millidge and Charlie O’Neill. I have a few important reflections from both of them.First, the podcast with Noam Brown made me internalize how big of a short-term acceleration mass inference capacity is. These labs will throw thousands of agents at important, measurable problems. At the same time, compute capacity available to them is going to continue to scale. I have my doubts that the labs can afford to spend a constant portion of this compute on internal R&D as the total volume goes up, especially with plans to IPO, as they face increased scrutiny on basic economics. It is important to not confuse massive steps in inference-time scaling, a dynamic which should be fairly predictable, with being the outputs of RSI, which is highly uncertain.Second, the trio podcast debating the state of the art in technical capacities induced more of a surprising reaction that I haven’t fully settled. Through the first hour or so of this podcast, where they debate the role of RL, distillation, scaling, inference-time compute, etc., I found myself strongly agreeing with the distribution of claims. A TLDR would be that our current techniques work and let us solve problems we know how to state, but they don’t result in a magical level of generalization to unknown, harder problems in most partially verifiable domains (i.e. progress in math is an exception, rather than a rule).The surprise of this podcast was the end, where they were predicting timelines...
    続きを読む 一部表示
    10 分
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません