PICK YOUR SUPPORT STYLE
MONTHLY SUPPORT
Reader
$5/mo
Contributor
$15/mo
Architect
$50/mo
Recurring subscriptions auto-bill monthly via Stripe Checkout. Cancel anytime from the receipt email.
What was talked about: Anthropic’s Opus 5 arrived with strong benchmark-level intelligence but weak real-world usability. It often ignores project instructions, misses relevant context, uses inefficient tools, edits files through unnecessary Python scripts, and changes more code than requested. Its code style can be good, but many users are returning to Opus 4.6 or 4.8.
Takeaway: Do not assume Opus 5 is an upgrade for coding work. Test it against an older Opus model before using it on a full codebase.
Links shown: Opus 5 usability report, Theo on Opus 5 over-editing
What was talked about: Frontier Code measures correctness, mergeability, code style, and unnecessary changes. Opus 5 performs worse at higher reasoning levels because it expands the scope of tasks and modifies too many files. Medium reasoning often gives better coding results because it stays closer to the requested change.
Takeaway: Use medium reasoning with Opus 5 when precise, limited code changes matter more than broad autonomous work.
Links shown: Opus 5 reasoning-effort analysis
What was talked about: Practical tests found that Opus 5 forgets visible context, misses instructions in CLAUDE.md, refuses routine development actions, and can lose track of large plans. Similar problems also appear in OpenCode, which suggests that both the model and its surrounding harness can contribute. The model appears more useful for planning and research than for implementation.
Takeaway: Use Opus 5 as a planning partner, but use a more reliable model for implementation and verification.
Links shown: Opus 5 research and coding review
What was talked about: Sonnet 5 and GLM 5.2 were compared for cybersecurity work, with GLM 5.2 offering similar capability at a lower price and with fewer guardrails. Opus 4.6 remains the most reliable older Opus option, while 4.8 offers more high-level intelligence. Anthropic models still degrade in very large contexts, and Claude Code and Codex differ in how well they handle compaction and sub-agents. Small Cisco cybersecurity models also appear overfit to benchmarks and weak in real workflows.
Takeaway: Prefer reliable models, shorter active contexts, and deliberate compaction. For cybersecurity tasks, GLM 5.2 can offer better practical value than specialized or heavily restricted alternatives.
Links shown: Opus 5 research and coding review, GLM-5.2 cybersecurity use, Cisco Antares security models
What was talked about: Kimi K3 received an open-weights release and technical report, but its license adds revenue-based conditions for inference providers and large businesses. These terms can help recover the high cost of training frontier models, but they also make lower hosted prices unlikely. More providers have improved inference speed, while pricing remains close to other frontier reasoning models.
Takeaway: Kimi K3 is more accessible and faster than before, but its license and hosted price limit the benefits of the open-weights release.
Links shown: Kimi K3 release post, Kimi K3 model weights, Kimi K3 repository and report, Kimi K3 technical blog, Kimi K3 on OpenRouter
What was talked about: Frontier Bench replaces Terminal Bench V2 with 74 carefully vetted coding tasks intended for frontier models. Slop Code Bench adds a different measure by tracking whether agents improve or erode a codebase over time. Current models can solve individual tasks while still adding comments, complexity, and maintenance problems during extended work.
Takeaway: Judge coding agents on maintainability and long-term codebase health, not only on whether they pass isolated tasks.
Links shown: Frontier-Bench announcement, Why Software Factories Fail, SlopCodeBench, DeepSWE, FrontierCode 1.1
What was talked about: A highly constrained agent workflow can reduce the need to inspect every generated line. Lint rules can prohibit risky patterns, while separate agents and clear rubrics can produce and review meaningful tests. Functional testing then becomes the main verification layer, although excessive comments still waste tokens and pollute future context.
Takeaway: Strong constraints, independent tests, and functional verification make autonomous coding safer and more useful.
Links shown: Why Software Factories Fail, Uncle Bob on constrained coding agents
What was talked about: SGLang and vLLM provide similar model-serving capabilities but use different inference techniques and have historically competed through performance benchmarks. vLLM supports more features and model configurations, while SGLang can be faster, especially for some mixture-of-experts models.
Takeaway: Choose vLLM for broad compatibility and SGLang when workload-specific performance tests show a clear advantage.
Links shown: SGLang repository, SGLang website
What was talked about: A new EQ-Bench evaluates how models respond to emotions without becoming sycophantic. Fable and Kimi rank highly, while GPT 5.6 falls behind earlier GPT versions and often produces longer, harder-to-understand explanations. Prompts based on ASD-STE100 Simplified Technical English and an I Have ADHD skill can make responses easier to read, though they do not fix incorrect answers.
Takeaway: Use explicit language and formatting rules to improve clarity, but verify the substance separately.
Links shown: EQ-Bench 4 announcement, EQ-Bench 4 leaderboard, i-have-adhd discussion, i-have-adhd coding-agent skill
What was talked about: The Flux 3 series expands beyond image generation into video, audio, and robotic action prediction. Early-access video examples show strong motion, creativity, human rendering, and scene consistency. A public or open release could provide a needed local alternative to closed frontier video models.
Takeaway: Flux 3 is a promising multimodal system, but its quality and local usability need confirmation after public release.
Links shown: FLUX 3 announcement
What was talked about: Anthropic is pushing for stronger action against companies that train on outputs from its models. This form of off-policy distillation normally happens after pre-training and provides structured assistant traces, but it does not replace reinforcement learning or the other work needed to build a frontier model. The dispute also raises questions because AI labs commonly learn from model-generated data and Anthropic customers pay for the outputs being collected.
Takeaway: Distillation can improve later training stages, but it is only one input to frontier model development and is difficult to treat as simple model theft.
Links shown: Michael Kratsios on model distillation
What was talked about: Proposed US policy could restrict very large training runs and discourage companies from using Chinese open models. Chinese token resellers can distribute Anthropic access across many normal users, retain real-world prompts and outputs, and later use that data for training. Because the traffic looks like ordinary usage, prompt classifiers cannot reliably identify it as distillation, leading Anthropic to use signals such as location and time zone.
Takeaway: Distributed collection makes model-output distillation hard to detect, while new regulation risks affecting legitimate open-model use as well as industrial training.
Links shown: Michael Kratsios on model distillation
What was talked about: NVIDIA organized an American open-weights initiative supported by major US AI companies, but Anthropic declined to join. Anthropic accepts smaller open models while arguing that frontier systems should remain closed and carefully controlled. This leaves it isolated from other major labs, all of which have released some form of open model.
Takeaway: Anthropic’s consistent opposition to frontier open weights creates a clear policy divide within the US AI industry.
Links shown: Jensen Huang on open-weight models, Open Weights and American AI Leadership, Anthropic’s position on open-weight models
What was talked about: Safe Superintelligence received a large NVIDIA investment after two years of limited public information. The funding appears intended to scale training for a safe, highly capable model that could help train later AI systems. The team has deep frontier-model experience, but there is no confirmed near-term model release or clear public product plan.
Takeaway: Treat release rumors cautiously. The available evidence points to a new training phase, not an imminent public frontier model.
Links shown: SSI on X, Safe Superintelligence, NVIDIA investment in SSI
0:20 All right, let’s get started here. Welcome back, everyone, to AI Tools Club.
0:26 If you’re joining us here on YouTube, thank you for subscribing and coming back.
0:31 If you want to come and join the discussion, feel free to join the Sunday Club Discord here and chime in.
0:36 I’ll also be reading the YouTube chat if you want to leave anything there.
0:41 So yeah, let’s get into it this week, starting with Opus V.
0:45 So this was sort of a surprise release from Anthropic.
0:49 Last week, they dropped this on Friday afternoon, which is very late and very annoying for me writing the AI news, because this released, I think, an hour or two before I actually sat down and started writing it.
1:01 So did not have too much time to go and gather feedback about this model.
1:07 The initial vibe that I had been seeing for Opus was that it was a fairly decent, fairly solid model.
1:15 But now as the week has sort of gone on, like my original review, I’d said like it’s around the same tier as like GPT 5.6, Sol, and like Fable 5, which intelligence-wise, it probably is.
1:29 But in terms of actual real-world usability, I’ve seen quite a few reports of people saying that they struggled to go and use this model, and that a lot of people are actually reverting all the way back to things like Opus
1:40 4.6 instead of using Opus 5. So one of the things that Anthropic highlighted with this model when it got released was how little prompting and instructing that it needed to have to
1:54 be able to go and accomplish tasks effectively.
1:57 So they went into Claude Code and they went and ripped out a whole bunch of the prompting that they have in there because they said this model doesn’t need it anymore.
2:06 And what people have been noticing is that the model’s instruction following capabilities have declined quite a bit because of this.
2:15 It seems to not really want to follow the various like clod.md files that you have within your code base.
2:25 And it also seems to not use the tools that it has very well.
2:31 So you can see here this user, which is just some person on Twitter, I don’t actually know them at all.
2:38 But yeah, it ignores regular read and text edit tool.
2:43 And it will go, I’ve seen multiple instances of people saying that it’s just using Python to go and modify different files.
2:50 Writing a Python script to go and edit the text in there instead of just using its text editing tools that it has directly.
2:58 And so, yeah, it also, yeah, its ability to read context is also declining.
3:05 Like it is regressing back to some of the older Claude models where they aren’t very thorough in the work that they do.
3:14 And so yeah, it doesn’t look across the entire code base and finds sort of like all the instances where, oh, I want to go or it should go and replace this logic there.
3:25 So yeah, it also is overly ambitious.
3:28 It doesn’t do enough and it also does too much, where if it goes and detects a small bug, like it will go and rewrite a lot of your code base.
3:34 I believe Theo was talking about this, yeah.
3:38 Where it’ll go and find the bug, and even if you told it to fix it or not, it will just go and of its own volition find these issues and start ripping through them all for you.
3:48 So I know a lot of people have talked about it going and using a lot of tokens and being very expensive because of this.
3:54 Interestingly, we also actually see this on benchmarks.
3:58 And this is sort of a good point of design your benchmarks well, is it actually catches this sort of over-eagerness behavior from Opus 5.
4:06 So this is frontier code. So this is from Cognition AI, the owners of Devon and Windsurf.
4:15 They have this benchmark, and not only does it test for correctness in its code, but it also tests for things like mergeability.
4:23 And they try and make the tests be based on the code style.
4:28 And would a open source maintainer merge this into their code base?
4:33 And so what they found is that the model actually I’m trying to find a good showcase of it.
4:44 But the yeah, I think it was Frontier Best.
4:47 Did they show it well? But yeah, basically, after medium reasoning, the model starts to perform worse.
4:55 And that is because it tries to go and modify too many things.
5:00 So the benchmark actually has built in questions where it says, okay, did it modify these files?
5:05 And did it go the extra mile and do too much?
5:08 Did it do more than I asked it to? And it catches that behavior and it penalizes Opus for it.
5:14 And so actually, I’ve seen a lot of people say that medium reasoning is actually the smartest version of Opus that you can go and use for coding because it’s not overly eager.
5:23 It’s not trying to do too much. It will follow your instructions and not take it too far.
5:27 So if you are using Opus, be aware of that.
5:30 That higher reasoning is probably not better unless you want it to go and write and overhaul your entire code base for you.
5:37 But yeah, it’s bad enough. I’m thinking about potentially even sending out a correction, basically, for the AI news, where this model is probably not actually Fable Level or GPT 5.6 in terms of usability in the real world
5:50 and the type of model that you would want to go and unleash on your code base.
5:55 The one good thing I’ve seen from it is that people really like the code that it writes.
5:59 And its code style is better than GPT 5.6, but otherwise it’s not actually that great of a model.
6:07 And a lot of people are reverting back to Opus 4.8 or 4.6 instead.
6:13 So yeah, has anybody here in the community been able to use Opus?
6:16 I know we have a lot of cloud users here.
6:18 What is your sort of like initial vibe from the model?
6:22 I’ve been tempted to go back to Opus 4.8.
6:30 I don’t know what it is, but I’ve tried a couple of intelligence levels, and it keeps on missing things that is right there.
6:41 Like 4.8, I was a big fan of. It wasn’t great at first, but today I’ve already tweaked claw.md a few times, just trying to figure out what the heck is going
6:56 on. Like, why am I correcting you? You’re supposed to be the smart one between the two of us.
7:03 And it’s just missing things. I completely concur with that.
7:07 I’m using your trip right now, and it is literally, I can fit the entire trip into a context window, and it is literally forgetting half of the trip at a given time.
7:21 Yeah, so I wonder if they have if some of their prompt cleanup, they need to put it back, give it some more system prompt.
7:32 I can’t believe that they would train something dumb, but something about this is not good.
7:41 I’ve been a clan for a while, but I’ve been using it all day, becoming a lesson less happy, where I probably should have.
7:52 Has anybody tried it in not cloud code?
7:54 Like, does anybody, I guess it’s basically impossible nowadays to use it not in cloud code, right?
7:59 Because they have it locked out. I was going to say, has anyone used it in open code?
8:02 Or I guess cursor, but I don’t think we have many cursor users here in open code.
8:05 I don’t think you can use your Anthropic subscription in open code anymore.
8:10 So yeah, we can’t tell if it’s like a harness level issue with clawed code where what they removed is problematic or if it’s the model itself.
8:20 Yeah, I’ve encountered some similar difficulties with it, but I am actually using it for open code now.
8:27 Oh, you are, okay. There’s an extension for it.
8:31 And how is it working in there? Is it usable, would you say?
8:35 Or is it for getting things there as well?
8:40 It still drops things from time to time, but I’d say it’s…
8:45 I’m trying to compare it to Opus, the older one, 4.8.
8:48 But I haven’t actually used 4.8 and OpenCode before.
8:51 And I’ve also just kind of switched to OpenCode the past few days.
8:56 Okay, cool. And I’ve also been using OpenCode a little bit more now on the side to check it out.
9:01 So hopefully in the coming weeks, you can tell us more about it and what you feel about it.
9:06 Because I definitely want to learn more about it.
9:09 Brandon, I think you were going to say something as well.
9:13 Oh, yeah. It’s just a bad model. And it says no to way too many things.
9:19 Like, you’re creating a developmental system and you want to iterate quickly.
9:23 And it’s like, I can’t do that. No, you have to approve this.
9:26 No, you have to explicitly tell me to get merged.
9:29 Oh, I won’t get merged. It’s just a bad system.
9:33 Anthropics doesn’t know too much for dumb stuff.
9:39 I know in their, I think in their technical report, they mention how refusals are way down versus versus fable.
9:50 But in the real world, it seems like they’re about the same.
9:55 Where it’s still saying no to most biology questions and any cybersecurity questions, anything like that, it still has super high false positive rates for their classifier for it.
10:05 Yeah, I’m not even talking about those kind of refusals.
10:07 I’m just saying, like, I want it to get commit and push all on its own.
10:11 And it’s like, no, how about not? It’s, yeah.
10:18 So it seems, does anybody here have something nice to say about it?
10:21 Or is it just universally this model isn’t that good?
10:24 Yeah, I’d say, I mean, I don’t have like a fair comparison, I guess, because I only have like the 20 bucks version, but it actually feels like it feels like it’s a cleaned up version of Fable.
10:39 It’s a little bit, at least when I’ve been using it, I kind of just default into, oh, I’m going to use it for planning.
10:46 So I didn’t use it as much for code reviews or just building.
10:53 And I think it’s a lot less frustrating than Fable.
10:57 And burns a lot less tokens, so I was actually able to use it in conjunction.
11:02 I mean, my main workhorse right now is just with the ChatGPT.
11:07 And I would still stick with chat over Opus.
11:12 So yeah, I kind of actually mildly positive.
11:15 But yeah, I mean, I agree that there’s a lot of kind of quite a bit slower than 4.8, which I
11:29 thought was pretty good. Interestingly, your sort of like take there is this is one of the few positive ones I saw.
11:36 This is one that came out on Friday, which is one of the ones I’m based on my review on XJDR.
11:40 I trust his opinion a decent amount.
11:43 He was saying, yeah, using it as sort of like a research peer or like planning mode.
11:49 The model, like, its intelligence is, or does seem to be able to shine through much more.
11:54 But it seems that on the implementation level, the model still really struggles a lot.
12:01 Yeah, I would like just to add to that is like chat, I mean, the reason why I even switched to chat was because it tends to, like, okay, it’s not as kind of clever, and there’s still a lot of nice design harness components
12:14 that are nice and clawed that chat doesn’t have, and it’s kind of, it takes a while to get used to.
12:20 Like, it’s just more like kind of a, you know, simple, it just goes, right?
12:24 So it can make kind of silly decisions sometimes.
12:27 Whereas I think with anthropic models, they just have a lot of mental gymnastics now.
12:32 But the flip side is now, especially with the latest release, I don’t know if you guys have seen this, but for me, it’s like, I just get exhausted sometimes.
12:43 It’ll like use this really specific sounding language, especially, for example, in code reviews, but it doesn’t actually mean anything.
12:52 It’s just, I’m not sure where they’re pulling it from, but then you ask, okay, can you rephrase this and like, you know, for a dumbass like me or something?
12:59 And it just doesn’t, it can’t really capture the it’ll rephrase it, but then you still, it doesn’t make any sense.
13:06 And Claude, I think they do a better job of actually having, like, the reiteration is actually legible, which is interesting.
13:15 And yeah, something else. No, I agree with that.
13:18 Yeah. No, you go quick. I found that I could get more intelligent, more something that I can understand better out of Claude in general than Codex.
13:28 And with Codex, it’s like, oh, come on, man.
13:30 Are you showing off? I have to say, I just down, when I just downgraded just now and I took my entire trip plan and threw it into 4.8 and I was like, hey, can you tell me how badly Opus 5 messed up?
13:43 It was like, oh, I can, and it found like five errors right off the bat that were all true.
13:49 And I was like, oh boy, here we go.
13:51 Oh. That’s a great A-B test. It just hasn’t been the same since you kind of find all these reports that Fable was basically undermining your work when you actually were trying to work in the right direction.
14:06 It’s just like, okay. So how many things did let’s say that they actually cleaned it up, but maybe they missed some things and now it’s like partially being obtuse.
14:18 I have a trust issue now. It’s yeah.
14:22 Because I know this is like similar issues and like similar vibes to Sonnet 5 as well, it feels like with Opus 5.
14:28 So maybe it’s just the fifth generation Quad models are all cursed in some way.
14:33 Fable is the only good one. But iterating on what you guys say, Will.
14:39 I do want to comment on that because I’ve been using Sonnet 5 a bit for some cyber tasks that, like, just because it’s a little bit easier and it doesn’t have as many guardrails and stuff, it’s
14:54 doing pretty good. So if you wanted to hear some positivity about at least one of the five level five models, Sonnet 5 is pretty good.
15:05 It’s my one issue with Sonnet 5 is that I feel like capability-wise, it fits very similar to GLM 5.2.
15:14 And you get a lot more usage and much better pricing with GLM 5.2 than you do with Sonnet 5.
15:20 Because capability-wise, they’re roughly the same.
15:22 I actually have, yeah, something somewhere.
15:27 Oh yeah, this. Talking about how this is a former Anthropic employee who works in the cyberspace.
15:34 And he’s talking about how most of the big groups doing cybersecurity stuff right now are using GLM 5.2 because there’s none of the guardrails to deal with.
15:42 And it’s a good enough model to actually go and use with it.
15:46 I concur I’m also using it for a lot of cyber stuff.
15:48 But yeah, the Sonnet stuff is good for like, let’s say your employer pays for Glot or something.
15:55 That’s true. Yeah. Yeah. Naturally Stupid says, what’s the most recommended Opus model now?
16:03 Many people said 4.6 is still best.
16:04 Yeah, it’s 4.6 or 4.8 are the two, I would say.
16:09 4.8 has more high-level intelligence, but I think 4.6 is still probably the most reliable Opus model.
16:16 But I think compared to the GPT 5.6 models, it will feel a bit dated and a little bit worse than them in the real world.
16:25 Also, one thing I’ll say, Will, with the model you’re using, I do not recommend using the 1 million context length version of any of the Anthropic models.
16:34 Anything over a quarter million context length, you start to see pretty noticeable performance degradation, even still.
16:43 So yeah, I think actually if they’re dropping, I think they’ve gotten at around 400K now.
16:47 It’s about good before the model starts becoming really dumb.
16:50 But it’s still not able to go and utilize the entire 1 million context window.
16:54 So just be careful when using that.
16:56 Yeah, I found that, and I had that in the back of my head while I’m using it, but like I found that 250 is just too low and that like around 500 to 600, it’ll, yeah,
17:11 it’ll definitely start to fall off.
17:13 But then that’s when I just, that’s when you should be compacting anyways, because eventually your next tool call is just going to take like 100,000 tokens.
17:21 It’s a question. It’s how is the Anthropic compacting now?
17:25 Because I know Codex’s is basically seamless.
17:29 And clawed code for a while, it was pretty bad.
17:32 Have they been able to improve that at this point?
17:36 I don’t notice it as much as I used to.
17:38 Okay. Like most of the time, it’s not noticeable when I’m working.
17:42 Some days I have too many clubs working, so I’m not always watching.
17:48 But today, with my other complaints, compaction has not been an issue.
17:54 It’s been a few weeks since I’ve talked about that.
17:59 I found that Codex seems to do just fine.
18:04 If you tell it to not spawn agents and to just keep it in a single thread, it does compaction very well over a very long task.
18:13 I’ve had it going for like five, six, seven hours at a time and it’ll just go.
18:19 I’ve found that Claud tends to actually understand how to use sub-agents very well and it’ll actually go through and spawn them, correctly use the ephemeral context,
18:33 grab what it needs, and spawn. Whereas Codex, I don’t know if you guys saw the nightmare that was the initial Codex ultra launch, but it spawned like 100 sub-agents and was
18:47 just a mess. So yeah, I think the Frontier Labs are starting to experiment with their harnesses a bit.
18:57 And quite honestly, I think Anthropic is still winning.
19:00 But it spawned like a hundred sub-agents.
19:05 Cool. Yeah, good to know there. Because I know, yeah, for a long while, yeah, Claude, I’ve been the innovators on using sub-agents.
19:13 And then also their context compaction has been pretty poor, at least when I was a Claude code main back in the day.
19:22 But glad to hear it’s a bit better now.
19:26 Brandon says, what of latest Cisco foundation models for cyber?
19:30 Yeah, I almost talked about those in the AI news, but the little third-party use I saw of them was fairly poor.
19:40 I want to say it’s like Cisco and Ataris.
19:45 Yeah. So these are allegedly super small models that are good at catching cybersecurity issues within your code base.
19:56 It’s yeah, I don’t have them on hand immediately, but it’s yeah, I saw one or two people talking about it, how they were using it and like running it through some of their internal benchmarks and workflows, and they said
20:07 that these are overfit, it seems, to the benchmarks that they’re doing well on, and that they’re not that great in the real world, and that you’ll get better use out of like GLM 5.2 for cyber stuff than these,
20:22 which is a bit sad. I really wanted these to be good because it’d be cool to have a 1 billion parameter model that’s really good at cybersecurity that you can run locally to just go and check through all your code.
20:33 But it seems like this is not the case.
20:39 But yeah. Anything else about Opus that people want to discuss before we move on?
20:53 All right then. We will move on then to Kimmy K3.
21:00 So a bit of an update here. The model got officially released as an open source model this week.
21:07 Not as nicely. Actually, they’ll say they also released the technical report.
21:12 But the license for it is not an open source license.
21:16 So what they have gone and done is that if you are surfing the model as an inference provider and you do over 20 million a year in revenue, then you have to go and sign an agreement with the Kimmy team for profit
21:30 sharing. And if you are a business and you do, I believe, over $100 million a year in revenue and you’re just using the Kimmy model at all, you have to sort of like say that you are somewhere on your website.
21:43 But because of this license, we’re not going to see a decrease in cost for this model, most likely.
21:52 I think probably Kevi, as a part of their licensing agreement, are basically telling providers that, yeah, you need to match our price so that we’re not getting price gouged on our own stuff and we can go and recoup some
22:05 of our costs. Quinton says, open quote unquote source.
22:09 It’s yeah, it’s honestly, I think this is correct.
22:13 I think was it GeoHots and the Tiny Grad team, they talked about this where it’s not, models aren’t like software anymore because open source software, it’s very cheap
22:27 to go and like run and execute it. And creating it is just human hours.
22:31 Whereas these models, basically the compiling of the model takes tens of millions, if not hundreds of millions of dollars.
22:40 And so these licenses allow these companies to go and recoup their costs in some way to be able to keep training these models.
22:47 So the model is still open source for you to go and run if you want to go and host it yourself.
22:53 You completely can do that. And then you can get your cheap tokens that way if you’re generating enough tokens.
23:00 But yeah, I think this is the Mini Max team.
23:02 They also did this as well with their M3 model.
23:05 They got a decent amount of flak for it, but I agree that it’s actually good, I think.
23:12 And this is much more sustainable for the community to have a model like this.
23:17 I think with the super small models, like, you know, sub 100 billion parameter models, I don’t think this license works as well.
23:24 But for these frontier models, or should be allowed to, or we should be fine with them having these not as permissive licenses so that they can make their money back.
23:38 But yeah, if you’re trying to use this model, previously I’d said like, don’t use it, it’s very slow.
23:45 Now that we have a bunch of providers, hopefully it should be picking up.
23:49 I think we’re about 50% to 100% faster, depending on what provider you’re using in terms of speed.
23:56 So hopefully it should be a bit better in the real world, but pricing will be the same as it was before, which means that it’s about the same price as using GPG 5.6 or most reasoning heavy tasks like coding.
24:22 Moving on, I have some benchmarking stuff to go and talk about and the benchmarks that we’re going to like, that I’ll be using for the AI news.
24:32 So Frontier Bench is a new benchmark, but it’s not really a new benchmark.
24:37 So Terminal Bench is one of the more popular coding benchmarks.
24:41 I think of the Opus 5 release. They included, it’s usually in pretty much every single one.
24:47 Oh, no way they have it. Oh, did they already replace it with the new one?
24:53 Frontier bench and yeah, Frontier Bench.
24:55 Okay, so they’ve already switched. But previously, most models had used terminal bench.
24:59 That was usually their headlining coding score.
25:02 That team, or well rather, that benchmark had started to get saturated, wasn’t really that useful or meaningful.
25:08 So the team went and made a new version.
25:10 So this is Terminal Bench V3, which is now called Frontier Bench, meant to measure these Frontier model capabilities.
25:18 And yeah, same idea, I think, as before, where the models are just tasked with doing something and trying to implement it well.
25:25 So yeah, it’s 74 different tasks. And yeah, this team does a very good job of vetting everything.
25:33 This is one of the big things we’ve seen in a lot of software engineering benchmarks specifically, where we find that models can only get up to 60 or 70% on these benchmarks because 30% of the questions are false.
25:47 So yeah, this team does a good job of preventing that.
25:52 But you can see initial rankings here.
25:54 Once again, you can tell it’s a good benchmark because it roughly follows how I would rank these models.
25:59 With Sol and Fable at the top, roughly tied, and then everybody else below to varying degrees.
26:06 So yeah. What are the hot takes on decks and slop code bench?
26:14 I don’t know too much about that. Do
26:31 you have a link to this at all? Because I have not actually heard about this.
26:34 Do you want to talk about this, Brandon?
26:37 Yeah. So he’s, Dex is like, software factories don’t fail.
26:41 Use my software factory. That’s kind of his like, the king is dead.
26:45 Long live the king. And essentially, he has his own software factory called Human Layer.
26:51 But essentially, he just points to the models suck at code bases over time.
26:57 And the biggest example is Slop Codebench, that Frontier models still perform very poorly on Slop Code Bench.
27:04 So it’s kind of like that long-term bench we looked at like maybe a month or two ago that looked at code bases over time.
27:14 And that’s his problem. It does not maintain code bases well.
27:17 Any of these models connects this reason against we’ve talked about it in the past, I guess.
27:22 Yeah. Etsy, this was a minute ago, I think.
27:25 Yeah, back in, I guess, yeah, three months ago.
27:28 But yes, okay, this is the benchmark that it’s talking about.
27:31 Etsy, my remembering, yeah, this benchmark basically shows, yeah, what Brandon just said, where as time goes on, these models make the code bases worse and worse.
27:39 Whereas humans, yeah, you can see agents make things worse over time.
27:44 They start worse and then continue to make things worse.
27:46 Whereas humans, we plateau and keep things about the same.
27:49 This is actually another criticism of Opus that I’ve seen is that Opus comments a ton on your code base.
27:54 Like integrating stuff that you just offhandedly mentioned in your prompt, it will go and add as a comment in your code base.
28:01 And yeah, like 50% of the tokens it’s using are just on comments.
28:06 So yeah, if that’s his argument, at least in terms of like, it’s, I guess, I don’t know, has slop code bench been updated?
28:12 I completely forgot about this benchmark.
28:15 But I like the idea of it. Let’s see.
28:18 How are the new models doing? Oh, they keep it mildly updated.
28:22 Okay, so it seems the GPT models seem decent at it.
28:26 The quad models, they tend to erode.
28:30 I think the other similar one would be the Deep Suite.
28:33 I think that was the other one we looked at.
28:34 Deep Suite? For long-term maintainability?
28:39 I don’t think Deep Suite. It’s Frontier Code.
28:44 Yes, Frontier Code. This is the one we were just talking about a little bit ago.
28:47 But this one is measuring, they have direct measures of code quality in here based on what maintainers are looking for in their code bases.
28:58 So I know that these guys have a score.
29:00 Any other hot take I have is in the complete opposite direction of Dex.
29:03 Dex is like, software factories don’t work.
29:05 Use my software factory. Uncle Bob Martin is like, don’t ever read code at all.
29:12 The guy who wrote clean code originally, the guy in the bathrobe who does rants, the old guy, this guy’s great.
29:18 And he’s like, I never read my code.
29:20 My agents are so constrained, they only ever write good code.
29:24 And it’s fire. Such a good hot take.
29:27 I did see a bunch of people responding to that.
29:32 Yes, yeah, I think it’s this tweet.
29:35 And yeah, extreme constraints. This is sort of the direction I’m going, I think, where I actually updated my linting rules, where I got rid of the ability for my agents to use useEffect in the front end because it’s sort
29:48 of like non-deterministic re-renders in your React apps that it goes and adds in.
29:53 So there are ways to not use it. But yeah, I think, yeah, having a lot of this sort of stuff and like defining, I think a lot of it is defining what a good test is.
30:01 I find models make tests that are meaningless or self-referential.
30:05 So like having it go and spawn a sub-agent to go and write tests and having a rubric of what a good test is is very important.
30:11 And we’ll be able to catch a lot of this.
30:14 Versus, yeah, like yeah, software factories I’m not as much interested in.
30:20 Yeah. When it says Uncle Bob, make me feel young again.
30:24 And I do know Uncle Bob. I feel like a lot of his takes of late, or like in the last like year, I’ve seen have been getting a lot of flack and I haven’t agreed with, but I think this one I do agree with.
30:35 And yeah, and I’ve seen a decent number of people also saying it.
30:38 It’s very weird though that he starts out with, I’m significantly older than you.
30:43 It’s a bit of a… Yeah, he’s a great hot take guy.
30:46 The other thing he’s saying is he doesn’t read his code and he doesn’t care if agents leave comments in the code because he has these constraints.
30:53 So everyone’s like, I hate the comments in my code base.
30:55 They’re making it worse. The team’s not maintainable.
30:58 And he’s like, why are you even reading agent code?
30:59 The whole point of agent code is to be faster, but you need to have it properly constrained.
31:04 So I think actually, if you follow the Uncle Bob methodology, you do have a software factory, but you’re much more active as an overseer in the factory itself.
31:13 But you don’t actually look at the code, if that makes sense.
31:16 It’s yeah. It’s I guess my issue with comments is that it’s just wasteful of tokens.
31:19 Like I said, with Opus using, you know, like 20 or 30% of its tokens just on comments.
31:25 Yeah, seems a bit bizarre and unnecessary, especially when it’s just introducing irrelevant context to go and mess up future models.
31:33 But yeah, I’d say that’s definitely where I’ve been.
31:35 I feel like with my usage where the vast majority of my code, I don’t actually go and write unless it’s sort of like research code.
31:43 If it’s just like web app stuff, I basically ignore all of the actual output and just test functionality myself.
31:50 And that’s about all I do to interact with the code.
31:59 Yeah, cool. Frido says, someone from SG Lang reached out Sunday.
32:03 Would people be interested if one of their developers presented SG Lang here for 20 minutes at AI Tool Club at some point?
32:08 Yeah, I’d love to have SG Lang guys on here as I’d love to talk to them.
32:13 I’ve been an SG Lang user for a few years now, so would love to hear what they have to say.
32:21 You can put me in touch with them if you want, Frito.
32:24 Yes. Sounds good. I didn’t understand.
32:27 How are they different than VLLM? Are they not kind of…
32:31 For me, they look the same. This is actually sort of a bit of a controversial thing, but yes, they are very similar to VLLM.
32:39 And a lot of people have said that, oh, why are they wasting their efforts doing basically the same thing that VLLM is doing, but in their own way?
32:48 So I think architecturally, they do things a little bit differently.
32:51 As I know they used Radix Attention for a while, which I don’t think VLM was using.
32:56 I think they were using page detention.
32:59 But yeah, I believe VLLM and SG-Lang are actually both from Berkeley, and they were like two competing labs there.
33:05 And they actually kind of hate each other, from what I’ve seen.
33:10 I remember when they were both still relatively new, they’d keep publishing benchmarks showing that one was better than the other.
33:15 And then the other one would come and be like, no, you did the benchmarks for my inference engine wrong.
33:20 Here’s how you do it properly to get a more fair eval.
33:22 And so there was a lot of drama back and forth between them early on.
33:25 I don’t think they’ve gotten big enough where they don’t have that publicly anymore.
33:30 But yeah, no, they’re basically just VLM.
33:32 They don’t have as many features, though, as VLLM and as wide a range of support of all the different miscellaneous things and bells and whistles on the inference engine.
33:42 But in my experience, they tend to be anywhere like between 5% and 30% faster a lot of the time, especially for mixture of experts models.
33:50 I don’t know if this is still the case, but they were for a while the fastest inference engine for mixture of experts models by a noticeable margin.
34:00 But yeah. Moving on, slot
34:14 code bench to the opposite side of things for emotional intelligence.
34:18 So this is Sam Peach. This is EQ Bench.
34:20 This is a benchmark I’ve referenced a number of times, I think, at the AI News now for sort of like writing capabilities of models.
34:30 They have a lot of good writing evals on here and sort of measuring soft skills using LLM as a judge.
34:36 So they released a new updated version.
34:39 This is how well a model can sort of respond to a user’s emotions without getting syncophantic on them and sort of echoing and like acting the way it should behave to
34:53 sort of like keep the user engaged and interested without going too far.
34:58 So yeah, thought this was a cool benchmark.
35:02 Fable and Kimmy, both at the top, sort of as expected.
35:06 Interestingly, the GPT 5.6 series actually falls down a bit compared to even like GPT 5.4 or 5.5.
35:14 So I know OpenAI, they claim to have tried to work on the writing style of GPT 5.6, but as Sasha was mentioning earlier, I feel like they’ve actually regressed in a lot of ways, where the model feels very incomprehensible
35:28 now with a lot of its writing. It also feels much more verbose also.
35:31 As I’ve noticed with ENCODEX, when it’s going and designing and explaining things to me, it feels like it’s two to three times longer than it had been previously.
35:40 And I feel like I’m not getting as much out of it.
35:42 And a lot of the wording is sort of complex.
35:45 I’m like, what the heck are you actually saying?
35:49 So yeah, but this benchmark seems to reflect that a bit as well that these models are available.
35:55 Somebody tweeted this and I found it to be useful, which is to basically, you know, the byte is the ask code as to HSID, I just have it so that when
36:10 I ask it for layman explanation, I have that code it as can you give it to me in what is it, ASD ASD 100 simplified technical English.
36:22 And actually like it helps. I mean it’s still it’s not it’s not great.
36:27 I think OPA is is still better somehow, but that mitigates it.
36:32 Okay. I think we mentioned this last week or I think two weeks ago maybe to go and use this.
36:39 But yeah, I’ve been using this. It’s helped a bit.
36:41 I feel like it’s the in terms of like crazy wording, it tends to lighten up a bit, but it’s still very verbose and it’s still not correct.
36:52 So I might go and try it. I think the other one was I have ADHD prompt.
36:58 So this was another prompt that people were talking about.
37:04 Do we have it? No. Let’s see if I can get it here.
37:11 Oh yeah, this is another pump we’re talking about to try and improve the model’s outputs.
37:16 Yeah, this one. So I’m going to go try this this week.
37:28 I’m going to try this week. Is this like a skill that you plug in or something you should have on your rules file?
37:40 This one is a skill, I believe. So you just say slash I have ADHD and then ask a question and it’ll respond to this format.
37:55 I think it’s meant to plug in directly to the cloud ecosystem, but you can also pretty easily slam it into Codex as well.
38:02 Oh yeah, they added Codex. Nice. I did when I first checked this out, it did not add Codex, but now it appears they do.
38:12 All right. Continuing onwards, wanted to mention the Flux team.
38:17 So Flux, they’re like the big open source image generation team.
38:23 They also have closed source bigger models as well.
38:26 They’ve been quiet though for the last, I think, probably close to like a year now.
38:29 But they recently announced their Flux 3 series of models.
38:33 And previously, Flux had only really done image generation.
38:37 But now this model is now an Omni model where it can do image, video, audio, and then also, interestingly, for robotics, action prediction.
38:47 And so I’ve been seeing a lot of example generations from this.
38:51 And it seems, at least from the people that they’ve given access to, that their video generation model is very strong and very good, very creative, handles people and motion, all of that very well, while maintaining cohesion
39:03 throughout an entire scene. This hasn’t actually been released publicly yet, though.
39:09 This is just early access, so I have not been able to get access to it or anything, and I don’t know too many people that have.
39:16 I think it’s mostly sort of like industry insiders that have.
39:20 But once this does come out, I’ll be sure to go and talk about it and mention if it’s good or not.
39:27 Interestingly, like I said, they have usually historically open source models that would be really cool to have sort of like a good video generation model that you can run at home because we don’t really have anything like
39:41 that. The last one was the WAN 2.2 model from Alibaba.
39:46 And that was once again, I think about a year ago now that that was released.
39:49 And it’s far below the likes of VO3 and the other Frontier video generation models.
39:55 So Flux has been the torchbearer for that, and hopefully they continue.
40:03 Up next, yeah, we’ll talk about this a little bit.
40:07 So circling back to the cybersecurity stuff that we were talking about before, this time in the context of national security, this has been a big debate and something that Anthropic really doesn’t like, and that is distillation
40:21 from frontier models. So now Anthropic has basically drummed up enough sort of like hoo-ha about this sort of thing that the US government is now talking about it and looking like they’re
40:35 going to step in or say something or do something about it.
40:40 But I find this is a bit silly from Anthropic.
40:45 Because like, so first off, they literally got caught pirating books online to go and train their model, taking sort of like from other people’s, you know, what is it?
40:57 Not patents or copyright, but their IP and using it to go and train their models.
41:03 And now they’re getting mad that other people are doing the same thing for their models.
41:08 And so this is also not a very good form.
41:10 Like distilling does like a little bit, but you need to do a lot of RL on top of it because this is what would be considered off-policy distillation, which isn’t that strong.
41:19 It only can get you so far. But they’ve, yeah, been talking about how they’ve caught, I think, Alibaba recently.
41:27 They also talked about how the Kimmy team has gone and distilled a bunch of tokens in Minimax from them, which, by the way, they’re all paying customers.
41:36 It’s not like they’re not paying for this.
41:37 They’re not stealing anything. They’re just using the model and then taking those outputs and training on them.
41:41 But Anthropic gets very upset about this because, as we’ve mentioned before, Anthropic believes they’re the only ones who should be allowed to build LLMs.
41:48 So the fact that someone’s taking their LLMs and building other LLMs with it makes them very unhappy, which if I were Anthropic shoes, I feel like I would actually want them, if they’re going to be distilling any model, to
41:58 be distilling theirs. Because they have all the safety guardrails and stuff built in and they’ve made the model to be helpful, harmless, and honest that they put all this work into.
42:09 And so by distilling from a model like that, the model that gets trained downstream should also hopefully be helpful, honest, and harmless as well.
42:19 And so, and yeah, I don’t really understand their argument here.
42:23 Yeah, they yeah, but they’re getting the U.S.
42:27 government on their side, although we’ll see how well the U.S.
42:30 government comes on their side, because the U.S.
42:31 government and Anthropic historically have not been on the best terms.
42:35 You know, like the military has blacklisted Anthropic, and they had the whole Fable drama as well, where they banned Fable from existing.
42:43 So yeah, it’s interesting to see that now the government’s on their side of trying to protect their IP.
42:50 It’s a quote, it says the industry is based on distillation.
42:52 Yeah, it’s like I think we’ve mentioned this in previous AI Tools Club, but if you there are previous versions of Cloud where if you asked what model it was in Chinese, it would say deep seek.
43:04 So they all distill from each other in one way or another.
43:08 But like I said, this is not where the sauce is.
43:10 This is not how you make frontier models at all.
43:13 It helps a little bit in the mid-training phases of the model, but this is not what makes models good.
43:20 Will says there are Chinese resellers of Anthropic accounts, and it’s a big problem.
43:24 It’s yeah, there’s a whole black market token distribution network in China.
43:29 From my understanding, a lot of it is actually these big labs to go and get realistic prompts and data.
43:36 Instead of trying to synthetically generate them all themselves, they go and buy, you know, like $1,000 $200 a month anthropic subscriptions.
43:43 And then they basically build an API that goes and just calls all of those.
43:47 And then they go and sell that to students or workers in China who want to go and use AI.
43:52 So they sell it at massive discounts.
43:54 And then they get all of this real world sort of like, here’s what people are actually trying to use AI for data in their realistic code bases.
44:02 And then they go and capture all that data and they use that to train.
44:05 And that also helps them sort of like the IP addresses that are using all of the accounts are distributed all across the, you know, China instead of just being all one lab, which is one of the reasons why it’s been hard for
44:18 Anthropic to catch these guys. Because yeah, they’re coming from sources all over the country.
44:26 See, I don’t know if it’s necessarily a big problem.
44:29 Because once again, they’re all still paying for Anthropic.
44:33 They’re just sort of like violating the terms of service a little bit.
44:36 But that’s about it. But yeah, it’ll be interesting.
44:48 It’s I know there is legislation that has been brought forward recently on restricting training runs that use over 10 to the 26 flops, I believe.
45:00 I don’t have a good measure. I think that’s less than like DeepSeek V4 used.
45:05 So still like very large models. But definitely if you want to be training Frontier, I assume Kimmy K3 probably used more than that.
45:13 Anthropic and OpenAI and Google, they all probably train models beyond that flop count.
45:18 But yeah, the US government is looking to add restrictions for them.
45:23 They’ve also talked about that’s sort of like their plan is that they’re going to very negatively incentivize companies from running these open source Chinese models instead of trying to block or ban them entirely.
45:36 They’re just going to try and raise up enough of a stink or like scare people away from using them, that sort of thing, as well to try and prevent these Chinese companies from getting into the US.
45:46 So yeah, AI is definitely getting very geopolitical now.
45:50 And sadly, I think the politicians do not have as deep of an understanding as we do about these sorts of things.
45:58 So it’ll be interesting to see how they go and navigate all of this.
46:03 Andrew, but there’s something that I don’t quite understand.
46:06 Like if, you know, Entropic is Entropic, shouldn’t they be able to just train like a crazy good classifier model that could determine whether a prompt is a distillation attack kernel?
46:20 Or is that just something that is impossible to do?
46:24 The way these distillation attacks work is it’s not like a preset curation of prompts from like, let’s just say like the Alibaba team that they’re just going and like throwing at the API.
46:34 Because if they do that, then Anthropic will be able to very easily catch them because it’s all coming from the same IP or roughly the same IP.
46:42 What they’re doing instead to get these distillation attacks is they’re just going at giving a whole bunch of people across China access to super cheap anthropic credits.
46:51 And then they’re just basically sitting as a middleman similar to the way OpenRouter does, except unlike OpenRouter, they are saving all of the data that’s coming through.
47:00 So it’s literally just real world clawed code usage that they are capturing and saving it for themselves is what’s happening here.
47:08 So there is no way for Anthropic to go and catch this unless they go and try and detect everyone in China that’s using it, which is something that they actually go and do.
47:20 They added in, I think I mentioned this in the AI news, but nope, not that one.
47:29 But basically, Anthropic embeds into clawed code an invisible addition to the prompt if it detects you’re within a Chinese time zone and it sends it back to their server.
47:41 So that’s another way that they sort of try and detect who’s using these models and why.
47:48 But yeah, let’s see. Was it a while ago?
47:52 Or did I just not cover the news? I might not have covered in the news.
47:54 But yeah. So you’re saying that like they’re not like it’s impossible to distinguish between them because they’re really just like normal prompts.
48:03 The only difference is that they’re like saving the outputs.
48:06 Yeah, exactly. There is no such thing as sort of like a distillation prompt that they’re sending in.
48:12 Okay. It’s just downloading outputs from the model to use for training.
48:17 Got it. Moving
48:36 on, I guess we’ll actually keep with the Anthropic news.
48:40 So NVIDIA announced sort of like this, yeah, Open Weights and American AI leadership proposal that basically everyone has gone and signed.
48:51 They got Google, they have, I believe, Amazon, they have OpenAI, all the open source labs here in the US, they’ve all signed this thing, except for Anthropic.
49:02 Because Anthropic, once again, they hate open source AI.
49:05 They don’t think anybody should be working about this.
49:07 So everyone has been talking about how Anthropic isn’t doing it.
49:12 They’ve outlined a whole thing, or yeah, Dario wrote a whole thing on why they aren’t supporting it, which basically boils down to they are fine with open source models as long as they’re
49:27 not good, which basically means they don’t want any open source models.
49:31 But yeah, they basically believe that the Frontier should be closed source and highly vetted.
49:36 And then, yeah, all of the smaller open source models should be handled by other people.
49:43 Will South Dario is an idiot. It’s, yeah, they’ve been very firm on their stance, which, you know, I’ll at least give them, they’ve not flip-flopped back and forth on open source or closed source.
49:52 They have been very vehemently the same ideology of we hate closed source the whole time.
49:58 As they’re also, every single other major AI lab across the board has released some form of open source model, except Anthropic.
50:07 So yeah, they are a strong proponent of this, but it’s funny because literally everybody has gone and signed this proposal except Anthropic.
50:17 So yeah, it’s a bit silly from them.
50:21 Shandu asks, may I know what you mean by distillation?
50:24 Is it during pre-training or after that?
50:26 Also for the prompt changing, they check Chinese time zone as you told.
50:28 Does that mean people in China can’t use it?
50:31 So okay, yeah, so for distillation, that usually will come after the pre-training phase as a part of the supervised fine-tuning phase, most likely, which is now also sort of like
50:45 SFT has been combined in with, what is it, mid-training, which is just training on a bunch of agentic traces.
50:56 So yeah, that’s, yeah, that’s an after-pre-training step where, because the data is a bit more structured, right?
51:01 It is an assistant chat instead of normal pre-training.
51:03 That’s just general documents. It’s not assistant style training that you’re doing during pre-training.
51:09 So it’s the step after that, though.
51:11 And then for the prompt changing, people in China can use it.
51:15 They are just being identified that they are a Chinese cloud code user when they’re using it.
51:21 But they don’t block China from being able to go and use it.
51:26 Will says, are we going to talk about how Ilia and Safe Super Intelligence is rumored to be releasing a Frontier model next month?
51:33 They’re not. I haven’t heard that rumor at all.
51:37 They announced, the only thing I know about Safe Super Intelligence is that they announced that NVIDIA is investing $5 billion in cash into them because they said that in the last two years they have done all of the sort
51:51 of initial research that they need to.
51:52 Now they’re going to try and scale things up and make safe super intelligence.
51:57 So yeah, for those that don’t know, Ilya Suskevar, I guess I pronounce it.
52:03 He is one of the best LLM researchers in the world.
52:09 He previously worked at OpenAI, but he was actually the one, he didn’t think OpenAI was taking AI safety seriously enough.
52:16 And he tried to get Sam Altman kicked out.
52:19 That was the whole Sam Altman step down as CEO saga that we had.
52:23 And Ilya basically got the boot because of it.
52:25 And he went and founded Safe Superintelligence, where I think they raised like $2 billion just off rip.
52:31 And then they’ve been in radio silence for the last two years.
52:34 Now they come out and say this, essentially, that we’re getting a lot more compute from NVIDIA.
52:40 And now, yeah, they’re going to be scaling that up, whatever that means.
52:44 And they vaguely mentioned that they’re studying models that are similar to the human mind, which is pretty vague.
52:51 You could argue that all LLM architectures are based on the human mind.
52:56 I mean, literally the foundational unit is called the neuron in LLMs.
53:01 So, yeah. I took the safe superintelligence and video post to mean that they need money to run models at scale now.
53:14 They need the cash, not just to build, but to run.
53:18 That’s the kind of scale I was thinking of.
53:21 I have no information other than the one text, the one tweet that he sent out.
53:28 Maybe it’s for training. Maybe it’s not inference right now.
53:32 But I thought it was inference. Yeah, my understanding is that it is for training.
53:36 Ilya’s whole idea is that, at least last I knew, I haven’t talked to them since, but the general idea is you make a safe, you know, ASI model, like a super intelligent AI
53:50 model that is deemed safe, and you use that to go and train all future AIs.
53:54 So you basically make an LLM that is known to be like super smart, but also like follows guidelines, has humans’ best intentions in mind.
54:03 And then you go and build that first.
54:05 And then you use that to go and make all other subsequent AI models in the future.
54:10 So that’s the general idea. So my guess is that they’re just trying to scale up and get to that same level and train a model that is good enough to be able to go and do that more than anything else.
54:24 Well, I’m glad that there’s still other labs that are coming up trying different things so we don’t just get two winners out of this.
54:35 I hope not anyway. I don’t even know.
54:39 They’ve done or announced nothing at Safe Superintelligence.
54:45 But yeah, I don’t even know if they would have a model that’s directly given to people.
54:51 I think Ilya cares. He does not care about the money and selling a model per se.
54:57 He very much believes he’s building the most powerful entity in the world and that it needs to be done very carefully and very correctly.
55:05 And he cares about doing that. And beyond that, I don’t know if he necessarily has too much of a business plan other than I’m the guy that can create ASI or AGI or something like that.
55:19 So we’ll see. Just my take on this, like all labs who release Frontier Model, they release it over stepwise.
55:28 Like Kimi, we saw over the future, and nobody starts on the top.
55:34 There’s a lot of work and a lot of good engineering OpenAIDS that have the right data, the right GPUs, the right people.
55:40 There’s a lot of million little things you have to do right to be a frontier lab.
55:44 And then you need to be in the game for a long time.
55:46 I believe rather they are good in creating four more.
55:51 I cannot imagine they drop a world-class model.
55:54 So, I mean, like, Ilya is one of the people that worked on the original, like, 01 reasoning model.
56:02 Like, he has went and made, and I believe a lot of the team that came with him, like, a lot of people also left OpenAI to go and join him from, I believe, is the research group at OpenAI because he was the head of research
56:13 there. But yeah, no, like, they all have plenty of experience of training and managing these frontier models.
56:20 I don’t think that is the issue. Like, I don’t think they spent two years sort of like building up the infrastructure that they had at OpenAI.
56:26 My guess is they carried a lot of that over and were building on top of that.
56:31 I don’t think they slowed down too much, like, maybe like three to six months, but they all knew how to go and train frontier models when they started.
56:43 But yeah, like I said, we’ll see. My guess is that it’ll probably be another year or two at least before we hear from them again in any meaningful way.
56:50 So yeah, cool. And with that, we have reached the end of this AI Tools Club session.
57:02 Thank you all for joining this week, and hopefully talk to you all next week.
Subscribe to get the latest AI news in your inbox every week!