PICK YOUR SUPPORT STYLE
MONTHLY SUPPORT
Reader
$5/mo
Contributor
$15/mo
Architect
$50/mo
Recurring subscriptions auto-bill monthly via Stripe Checkout. Cancel anytime from the receipt email.
What was talked about: Darkbloom connects Apple silicon computers to a distributed inference cluster available through OpenRouter. Mac owners can earn money by serving approved models, but current participation requires at least 48 GB of unified memory. Estimated earnings depend on hardware, demand, duty cycle, token throughput, and electricity costs, so the published figures need careful validation.
Takeaway: A high-memory Mac can earn money from unused inference capacity, but owners should verify real throughput, demand, power costs, and payback time before buying hardware for it.
Links shown: Darkbloom announcement, Darkbloom nodes
What was talked about: Distributed inference services have difficulty proving that a remote machine is running the declared model and software. Darkbloom uses Apple hardware and operating-system integrity mechanisms to attest the machine configuration and execution environment. This can give customers stronger assurances about model identity, quantization, software integrity, and data handling than an arbitrary Windows, Linux, NVIDIA, or AMD host can provide.
Takeaway: Apple’s controlled hardware and software stack makes verifiable third-party inference practical, although attestation still depends on the integrity of the audited software and its checks.
Links shown: Compute Community, Darkbloom nodes, Qwen 3.6 providers on OpenRouter
What was talked about: Low-priority inference workloads may tolerate only a few tokens per second if the price is low. SSD-based inference is limited by drive endurance, while used DDR4 server memory is designed for repeated access and can provide large capacity at a reasonable price. A proposed server design combines large amounts of RAM, multiple inexpensive RTX 3060 GPUs, wide PCIe connectivity, and NUMA-aware CPUs to treat system memory as an extended model store.
Takeaway: For large models that do not need interactive speed, used DDR4 and several low-cost GPUs could provide better durability and capacity than SSD-backed inference, but PCIe bandwidth and parallelism remain important constraints.
Links shown: Qwen 3.8 27B on OpenRouter, KTransformers, DDR4 server motherboard listing
What was talked about: FreeToken is a new inference engine that uses system RAM with a small GPU to run models that do not fit fully in VRAM. Early results show substantially higher generation speeds than Ollama, llama.cpp, and KTransformers for Qwen models on systems with fast DDR5 memory, including results near 40 tokens per second with 8 GB of GPU memory. Model support is still uncertain because the project is new.
Takeaway: FreeToken appears promising for fast hybrid CPU-memory and GPU inference, especially with DDR5, but users should confirm hardware requirements and supported models.
Links shown: FreeToken paper, FreeToken repository
What was talked about: Apple’s M5 Ultra restores the Ultra tier with high memory bandwidth and configurations expected to reach 512 GB of unified memory, which could hold extremely large models. The first M6 systems offer a smaller performance increase but arrive with substantial price increases. High-end Mac Studio configurations are expensive, yet their memory capacity and bandwidth make them unusually accessible machines for very large local inference workloads.
Takeaway: The M5 Ultra is the more important AI upgrade because its large unified memory and bandwidth can support models that do not fit on most individual GPUs.
Links shown: MKBHD on M5 Ultra and M6, Darkbloom nodes, Apple M6 and M5 Ultra announcement, Apple Mac mini, Configure Mac Studio
What was talked about: The RTX Pro 6000 offers more compute, stronger parallel performance, and higher memory bandwidth than a Mac, but its 96 GB of memory limits model size and it consumes more electricity. A Mac Studio includes the full computer and much more unified memory, while DGX Spark and NVIDIA GPUs provide a better environment for training and fine-tuning. Macs can train models through MLX, but their architecture is more favorable for memory-bound inference than compute-heavy training.
Takeaway: Own high-memory hardware for frequent inference, but use cloud GPUs for occasional training unless the training workload will run continuously for several months.
Links shown: Configure Mac Studio, NVIDIA RTX PRO 6000 family, Apple MLX
What was talked about: Artificial Analysis compared small models on mobile-relevant intelligence, speed, tool use, instruction following, and hallucination measures. Liquid AI’s LFM 2.5 family performed strongly, with the 2.6-billion-parameter model offering high intelligence and the 8-billion-parameter mixture-of-experts model offering a strong speed tradeoff. Function calling, instruction following, and calibrated refusal provided useful signals, while GPQA Diamond and Math 500 were treated as overfit or saturated benchmarks.
Takeaway: LFM 2.5 models are strong current choices for phone deployments, and selection should emphasize latency, tool reliability, instruction following, and hallucination control instead of old aggregate benchmarks.
Links shown: Artificial Analysis mobile benchmark announcement, Artificial Analysis benchmark charts, Mobile benchmark launch article, Live mobile inference results
What was talked about: A personalized phone assistant could improve through a memory system that stores and retrieves relevant user context. Continuous fine-tuning on a phone is less practical because useful training normally requires many high-quality examples and substantial compute. Small on-device models also have limited capability for dangerous autonomous behavior compared with frontier systems that have escaped evaluation sandboxes to find answers.
Takeaway: Build personalization around persistent memory and retrieval before attempting on-device fine-tuning; current small models are better suited to bounded assistant tasks.
Links shown: LFM 2.5 on an old Kindle Fire, Artificial Analysis benchmark charts, Mobile benchmark launch article, Live mobile inference results
What was talked about: The Pixel’s Rambler voice-typing system provides low-latency transcription, edits earlier dictated text, and translates a long English passage into another language on request. Its responsiveness suggests significant on-device processing, possibly using a Google model designed for mobile deployment. Similar capabilities could be extracted or recreated as a cross-platform alternative to products such as Wispr Flow.
Takeaway: Real-time voice interfaces can go beyond transcription by applying contextual edits and translations, creating an opportunity for capable on-device, cross-platform dictation tools.
Links shown: Artificial Analysis benchmark charts, Pixel Rambler discussion
What was talked about: AI coding tools make it practical to replace narrow SaaS products with small applications tailored to one workflow. Examples included marketing lead discovery, monitoring, LLM tracing, meeting transcription, and automatic summaries. These products often combine search, speech recognition, model calls, and a simple interface, although mature hosted services still provide convenience, maintenance, and operational reliability.
Takeaway: Replace a SaaS product when its core workflow is simple, stable, and specific to your needs, but account for the maintenance burden that the subscription previously covered.
Links shown: swyx: another SaaS to kill, Langtrace
What was talked about: Aux Alpha is a stealth model available at no cost through several coding-agent platforms with unusually large reported capacity. Shared tokenizer details suggest that it belongs to or derives from the GLM family, and speculation identifies it as GLM 5.3 Flash. Early agentic coding results appear competitive with larger models while using tokens efficiently, but its true architecture, size, release date, and future price remain unknown.
Takeaway: Aux Alpha is useful for cost-free coding experiments now, but decisions about local deployment should wait for the official model identity, weights, size, and license.
Links shown: X search for Ox Alpha, OpenCode Ox Alpha announcement, OpenCode capacity discussion, Ox Alpha DeepSWE results
What was talked about: Jalapeño is an internal OpenAI accelerator designed for high-speed, energy-efficient inference. Reported comparisons emphasize tokens per second for each user and total throughput per kilowatt, which directly affects data-center operating costs. External analysis suggests that it may outperform current NVIDIA systems and upcoming Vera Rubin hardware for its targeted inference workloads, although it is not expected to become a consumer product.
Takeaway: Custom inference silicon can reduce OpenAI’s latency and electricity costs even if users never gain direct access to the hardware.
Links shown: Jalapeño first results
What was talked about: Hugging Face serves as a central repository for open models, datasets, quantizations, and lightweight compute services. Reports indicate that the company is considering offers around a $13 billion valuation. Its storage-heavy open ecosystem has raised questions about monetization, although the company says that enterprise plans and compute services support a profitable business.
Takeaway: A sale could materially affect open-model distribution, so the important issue is whether new ownership preserves Hugging Face’s accessible hosting and community focus.
Links shown: Jalapeño first results, Reuters on a possible Hugging Face sale, AMead10 models on Hugging Face
What was talked about: A humanoid robot competition in China featured running, jumping, fighting, tennis, and dexterity events. Tennis highlighted the difficulty of autonomous perception, motion prediction, balance, and real-time control, while running events demonstrated both rapid progress and the physical fragility of current machines.
Takeaway: Humanoid robots are improving quickly in dynamic autonomous tasks, but reliability and physical durability still lag behind their best demonstrations.
Links shown: Reuters World Robot Games post, Reuters World Robot Games report, Kyle Chan on autonomous robot tennis, China AI News high-jump clip, Digital Trends robot sprint crash
0:01 Let’s get started then. Thank you, everyone, for coming back to this week’s edition of AI Tools Club.
0:08 This is fully community-driven. So, even though I will be leading and I have a bunch of topics prepared, I also want to hear from you as well.
0:17 So, if you have any questions or comments or anything that you want to chime in with or whole topics that you want to discuss, feel free to jump on in and we can do that.
0:27 This week, though, I wanted to get started, not with models.
0:31 Usually we start by covering the different models that have come out, but there hasn’t been anything too interesting.
0:36 I know Quentin has mentioned OxAlpha in the chat.
0:39 We’ll get to that in a little bit. Well, since that’s rumors right now, so I usually don’t like talking about sort of rumors of models or things that aren’t fully released.
0:49 But anyway, I want to talk about something else this week, which is a platform that a lot of us could potentially be using.
0:55 Because I know a lot of us here probably have M-Series Macs that are sitting around.
1:01 So, this platform is called Darkbloom.
1:04 So, this platform, what they allow you to go and do is you get to use your Mac as a machine that can go and do inference as a part of a big cluster.
1:15 And so they are actually now a part of OpenRouter, and they turned on people paying for these models.
1:22 So you, the sort of node owner, get paid for the tokens that your Mac is generating at home.
1:31 I believe, I don’t actually know where they’re at right now in terms of token usage, but I think they were doing like, I think they’ve done a couple billion tokens.
1:40 Let’s say dark blue. Let’s see. So yeah, they’re doing, yeah, about 10 billion tokens between Gemma 4, GPT OSS, and Quinn 3.6.
1:52 So yeah, it’s very interesting, I thought.
1:55 You stand to make a decent amount of money if you have a strong enough chip.
1:59 So I think I have an M4 with 16 gigabytes of memory.
2:02 I can’t actually go and use this platform because I can’t run any of these models, like this, like Quin3, 6, Gemma, 4, or GPT-OSS.
2:12 You need more memory than that to be able to run those models.
2:14 So it looks like, actually, interestingly, when I checked this last week, you could do this with 24 gigs because technically you can run GPT-OSS on there, but now they’re requiring at least 48 gigs of memory.
2:26 So if we go down, oh, this is, oh, this is Mac Mini.
2:29 We’ll go to a Mac Studio instead then and look at the prices.
2:32 We’ll say you have one of the nice ones.
2:36 Interesting. Okay, they’ve decreased the earnings that they show for this.
2:39 Oh, this is because duty cycle is slow.
2:41 So yeah, let’s say you have a nice Mac Studio, an M3 Ultra.
2:44 This is, to be fair, fairly expensive.
2:46 And you were running at 75% of the time.
2:48 You would be making, according to them, about just shy of $400 per month, depending on, yeah, as long as they can get the customer’s worth.
2:59 10 billion tokens a day is probably enough to saturate it.
3:02 I know they have gone a bit viral. I think they have about like 200 Macs on the network right now.
3:07 I’d also sort of run these models yourself and see how fast it goes.
3:13 Because yeah, I saw some of their numbers, at least for the low-end stuff.
3:17 I think that’s probably why they took it away.
3:18 But for those low-end models, they were not doing the math correctly, essentially, for how fast the models were running.
3:25 And so they were fastly overestimating how much money you’d make.
3:29 But yeah, you technically can make, what is this, $4,700 a year?
3:33 So that definitely, you know, could pay for the entire unit.
3:37 Yeah. The nice thing, because Will says, be very careful of your homes kilowatt hours.
3:41 Macs are very low energy usage usually.
3:44 So even if you’re running it full tilt, I believe it only uses like 100-ish watts to 200 watts.
3:50 So it’s yeah, it’s definitely something to be aware of, but also not something crazy expensive either.
3:55 I believe also these calculations that they show here factor in the price of electricity as well being subtracted from that.
4:02 But yeah, I thought that was cool. Like I said, if you want to go and use the giant cluster of Macs to run local models, you can go and check it out here on OpenRouter.
4:12 I looked into this a bit, so for those that don’t know, I have a platform that I’m eternally chipping away at called Compute Community that is sort of similar to this, where you are able to go and say if you’re running a
4:24 model locally, you can go and share it with a community here.
4:29 Yeah, this platform works with Macs.
4:31 It also works with GPUs like NVIDIA or AMD, or if you have a DGX Spark like Nate here, you can host models here.
4:40 This is not the same as that. So basically, there’s something very special with Macs that you are able to go and do.
4:48 So one of the issues that I ran into with compute community is that there’s no way to go and verify what model is running on a given piece of hardware in a general sense.
4:58 You’re unable to do it with NVIDIA GPUs or AMD GPUs or anything like that.
5:04 There’s no way for me, a remote observer, to go and verify that as a compute community admin.
5:12 If you say you’re hosting Quen 3.6, I have no way of verifying that other than it looks like the right shape.
5:18 But Macs, on the other hand, they have a very interesting architecture where they have a bunch of OS level internal integrity mechanisms that Darkbloom is able to go and backpack, piggyback
5:33 off of to be able to go and sort of verify what is actually running on your machine and that you are also, you know, your machine is what you say your machine is.
5:43 You know, like I tried, very first thing I did, as I saw, I can’t come on here with my M4 Mac Mini.
5:51 But I saw, like, can I go and spoof this somehow?
5:53 And sort of convince them that I do have, you know, a stronger Mac Mini that would be able to run the models they say they want.
6:03 But everything seems to be sort of like in boxes in the Apple ecosystem and signed properly, where, at least through my non-cybersecurity lens, there’s no easy way into this box to break this box.
6:17 You probably have to go and develop some sort of zero-day exploit, essentially, to be able to go and infiltrate the system.
6:24 So I thought that was pretty cool that somebody figured out how to do this with Max.
6:27 I was vaguely aware of this, but it seemed like a big hassle.
6:30 But I know that this has actually been around for a couple of months now, and I think most of their work has actually been going around this privacy and security, both for you, the model
6:44 host, and then also the people using your server.
6:49 So any questions about that? Does anybody here have a Mac mini that they would try and put on here?
6:55 I know Quinton says he just tried to have, I don’t know if you actually have an M5 Pro, MacBook Pro, Quentin, but yeah, it’d be cool to see if you’re actually able to make as much money as they claim to on their page.
7:10 Yeah, I have the M5 Mac Pro 48 gig.
7:15 So at full tilt, they said it’s $192 a month.
7:19 So I have of that that you just had up there, I guess.
7:24 I’m definitely not ever going to do this.
7:27 Not a chance. But it’s very cool. Why would you not do this because of the privacy stuff?
7:36 Unnecessary complexity in my life. I just don’t want to deal with it.
7:43 I’m nothing against it. And I believe the Apple hardware may be a really secure path.
7:51 My understanding is that all of that privacy stuff is abstracted away from you.
7:55 And the install, from what I saw, what they claimed.
7:59 And yeah, it’s very direct, where I think, yeah, you literally just install something like you normally would, and then you just set it up and then log into your account and then just run a setup command and pick a model
8:09 to use. And it just gets added in. Oh, I might use it.
8:14 Oh, I thought you meant would I host other people’s models on my Mac?
8:20 It’s well, it’s, well, I guess the other people’s models are, you know, big name models.
8:24 These aren’t like random models. Like I said, there’s three models to actually go and choose from right now for you to host.
8:30 Right. That’s cool. But yeah, so is it that they also
8:45 guarantee crypto, like, because it’s a trust execution environment, like, well, the guarantee of the user safety.
8:54 So it’s like actually the only platform on open router that can basically guarantee that the provider cannot read your data because it’s on a Mac and you like they can leverage that.
9:05 Because I’m guessing this thing is not only for like the Mac, the trust execution environments are not only for verifying which model you’re running, but also for ensuring that the data that it’s processing cannot leave and
9:16 you cannot physically read it without compromising your, like, without making it obvious for the server.
9:26 So is it like, are they, are they saying that also that, like, are they also using that as an advertising?
9:31 That like your data is actually not being able to read?
9:36 Like, I don’t know if they’re, how much they’re leaning on that side of things.
9:41 I think they definitely could. I think theoretically, for H100s, I believe they also have a secure execution environment on them.
9:50 I remember looking into that. I don’t know if the newer B200s do.
9:54 So technically, other providers on OpenRouter could theoretically have the same thing where they have sort of like a way for a third party to attest to the fact that they are, like, what you’re specifically attesting to is
10:05 that they are running a specific version of a piece of software.
10:08 Like, you basically have a hash of the piece of software, and you can go and verify that they are the same.
10:13 So then you’d be able to audit that piece of software and be able to say, like, oh yeah, nowhere in this software are they saving any data or anything like that.
10:20 My guess is that they might not be fully confident in their internal stack.
10:26 Like I said, part of me wants to go and just throw an agent at it and see, can GLM 5.3 go and hack into this thing and go and mess with it.
10:34 I think that the closest thing or the easiest potential way to go and mess with it is to run a different model than what you say you are.
10:43 I think that is a little bit weird because that’s a live execution environment, but they still have checks every couple of minutes that you’d have to dodge.
10:52 Yeah. But yeah, technically, I think based on that, this is probably the securest way.
11:01 Or yeah, like the, I won’t say securest, but you know what model is being run with a higher chance of correctness than any of the other providers.
11:11 Like, you know, other providers in OpenRouter, they can quantitize the model, and you wouldn’t necessarily know.
11:15 But here, you know pretty well that whatever model they say they’re running, whatever quantitization of Quantum Nurgemma, that that is specifically what they are running.
11:25 You have better guarantees, I’d say, than a random provider.
11:32 Natural Stupid says, is there a way to verify in Windows which model it is using?
11:35 Some verification from Lama C
11:38 It’s like I looked into this, like the more general case of can you verify on arbitrary hardware or arbitrary operating system.
11:47 To my knowledge, there’s no good way of doing it on existing operating systems.
11:53 Mac is a bit special because Apple sort of like has it clamped down a little bit more and it’s a little bit more sort of like privacy and security focused than things like Windows or Linux are in terms of like can a third
12:06 party go and mess with it. It’s Apple makes it a bit harder to do so.
12:11 And like I said, yeah, they have a bunch of like hardware level like I forget what they call it, but they have yeah like OS integrity stuff basically on Mac that no other operating system has.
12:24 So because Windows runs on like any hardware you give to it and Mac, like Apple can control the exact hardware and firmware and
12:38 software that they’re putting on there.
12:40 So they can I think you can still turn it off in some like Mac BIO settings, but but then I think the app can request that it’s not turned off in order to run or something, which I’m sure they’re doing.
12:54 So yeah, that’s a pretty nice way to leverage that because for now that system integrity protection was kind of annoying to deal with.
13:02 Like I’d turn it off always and it’s just kind of not useful for anything.
13:06 But this feels like the only like the actual use case then you know the banking app telling you you can’t use it because you have it turned off or something in my experience.
13:16 So this is pretty nice use case for it.
13:23 Anything else about Darkloom anybody wants to discuss?
13:29 Do you know what their prices are for like Quen models for example?
13:33 Because I think Quen was really expensive last time I checked like them 70 cents.
13:38 But this is the mixture of experts one.
13:41 The expensive one that you’re thinking of is the dense model.
13:46 So yeah, this is about $3 per million output.
13:49 The issue, this is not a model that would run well on Macs because of Macs low memory bandwidth and also low compute intensity.
13:58 Because this model, I think, on a 3090 at batch size like two, it’s already saturated all of the compute and the 3090 has like 50x the compute that a Mac has.
14:09 So yeah, this you need more proper compute for that.
14:12 Macs just don’t have the horsepower to run this at decent speeds.
14:18 I see. Is there a reason they’re not offering it at like a slow snail speed at a discount?
14:25 It feels like there would be a market for that.
14:27 Is that a thing on open router or something?
14:29 It’s maybe. I feel like, I mean, I guess what we got this provider.
14:33 Oh, this is for the MOE model. They’re running at eight tokens per second.
14:37 Okay. But yeah, I think it’s not a very good user experience to go and get that.
14:44 So who, I guess you got 13 tokens per second as well per sale for the 27.
14:50 And still as expensive, so it’s a bit…
14:52 Yeah. It’s a strange proposition. It feels like a lot of background tasks can run at like, you know, five tokens per second, and you don’t care because it’s so slow and you don’t think it’s not an urgent thing.
15:02 It can just be running for as long as you’re saving money on it.
15:06 Like off-peak and stuff. But then you’re also competing with all the off-peak pricing and flex pricing of the likes of OpenAI and stuff.
15:21 That was at least the premise for my SSD idea that if you have your SSDs running your model at like three tokens per second, that’s good enough.
15:32 But I’m switching now to, because SSDs are expensive, so I’m switching to a DDR4 idea with like just spam cheap DDR4 from servers.
15:41 And that’s now the same price, but it’s not going to burn after 80 million tokens.
15:47 And Zach, so remember, you had said previously that, yeah, if you’re trying to, you know, run LLM inference like off an SSD, essentially, the SSD gets worn out super quick and you run through all of the effectively cycles
15:58 that the SSD has for reading. Yeah, you get like 100 million tokens for the entire lifetime on SSD, or at least it’s warranted terabytes that it’s supposed to be able to deal
16:13 with. But RAM is actually designed to do that.
16:17 And RAM is pretty, it’s not that cheap, but it’s pretty cheap.
16:21 That is DDR4. Like if I found some like terabyte of RAM from DDR4 for like $2,000, which is like you can run the Kimik3 on that already.
16:36 It’s funny you mentioned I just bought actually 128 gigabytes of DDR4 for a new server I’m building right now.
16:43 I also think if you’re looking at doing, I assume like DDR4 inference with or without GPU, I know K-Transformers is sort of that hybrid environment where you’re using both DDR4 memory and the GPU moving stuff back and forth.
16:56 I believe last I knew this is the most optimized for it.
17:00 I think Lama C
17:05 Cool. Yeah, I was thinking to get like one of those server like plates with a lot of PCIe links and just spend 3060s in there because they’re super cheap.
17:13 And each one gets you a 32 gigabytes per second bandwidth when you plug it in.
17:18 So you basically, with the 3060s for like $150, you get 32 gigabytes of memory bandwidth.
17:24 And that’s your, you’re using your RAM as VRAM.
17:27 And then you have like eight of them.
17:29 And you get memory bandwidth of DGX Spark with terabytes of RAM for the same price-ish, at least in theory.
17:39 I said, that would be a very interesting server ID.
17:41 Yeah, 256 gigabytes a second of memory bandwidth is, yeah, that is.
17:47 Through PCIe lanes, entirely through like, you know, moving stuff into and out of GPUs.
17:52 Yeah, I’d say, I guess you would also have to distribute, because I think PCIe lanes are also bottlenecked at a couple of gigabytes per second, right?
17:59 32. 32, okay. It’s not bad. So then you need eight 3060s.
18:06 Yeah. Yeah. And there are server plates where you can fit all eight at the same bandwidth, and they will get the same access to the RAM.
18:12 So you basically use RAM as the VRAM.
18:14 It’s just that you still do the token sensor parallelism stuff.
18:17 Yeah, stuff like that, actually. This is the motherboard I bought, which is basically tablets.
18:22 You already bought it. Okay. Yeah, yeah, yeah.
18:25 Okay, cool. Yeah, I was looking a little into one like that or a similar one.
18:30 Yeah. But I wanted to get double CPU one because then it’s easier to do token parallelism if there’s two CPUs and they’re actually linked by the NUMA.
18:44 Because I would expand it eventually.
18:47 Yeah. Or at least that’s a very interesting project.
18:51 Hope to hear more from it in the future.
18:54 And Pavel in the chat also mentions free token running at small models of free token.
19:04 And so yes, that is also an inference engine that came out this week.
19:09 This is, I guess, also talk about sort of inference using RAM and your GPU.
19:17 Yeah, free token. This, I believe, is actually, yeah, I forgot about this.
19:19 This came out, yeah, this week. And I believe they’ve claimed it’s like two to four times faster than regular sort of like Olama way around it, which Olama, not very good per
19:34 se, but Antioch Pavelo says, I’ll just open up this image.
19:42 Yeah. It’s basically from the same paper you’ve been just showing.
19:44 Yeah, yeah, yeah. It’s an easy way to find the graph.
19:48 But yeah, using, I believe, what, yeah, eight gigabytes of memory and then just like some fast RAM, they’re able to run the model much faster than you would conventionally.
19:57 So you can see like, yeah, K transformers, Lama C
20:07 Yeah, they show Codex median at 33 tokens per second.
20:10 So fast enough to go and do reasonable things on here.
20:13 So I haven’t gotten the transport. I believe in their paper, they are using DDR5 memory, which is the memory that has gone up hugely in cost.
20:22 Like Daniel Zadi was talking about DDR4 memory, which is a bit slower.
20:27 So you will definitely won’t see as good performance on lesser systems.
20:32 But yeah, if you have DDR5 memory in a small-ish GPU, this is actually probably the best way to go and run it now, not K-transformers anymore.
20:40 I think, do they have wide-range model support?
20:43 I’m not sure what models they support yet.
20:46 Because this is relatively new. But yeah, once again, if you’re interested in that, go check this out.
20:51 I’ll toss a link in here for you. Cool.
20:56 All right. Moving along now, we’ll stay in Mac land and hardware land.
21:03 We have new chips from Apple coming out.
21:06 So first, the M5 Ultra. So the Ultra series of M series chips are the really good ones.
21:12 These are the ones that you want to go and do inference with usually.
21:16 Like if we go to Dark Bloom, the highest earning ones will be the Mac Studios with the Ultra chips because they have the best memory bandwidth speeds.
21:27 So yeah, those will go and make the most and they can serve the most tokens.
21:31 So we haven’t actually had one with the M4 series of models.
21:36 And so we’re getting that back though now with M5.
21:39 So that’s very nice to see. I believe it goes up to 512 gigabytes of memory.
21:45 So you can run pretty large models up to a trillion parameter models theoretically on there.
21:53 So there’s that. There’s also the new M6 chip that Apple announced.
21:57 I’ve heard rumors that this chip sort of series isn’t going to have too much interesting things about it because M7 is a much bigger upgrade and they’re focusing a lot of their development plans on M7.
22:09 But I believe M6, they’ve only released the base chip.
22:12 It’s about 10 to 15% better than the M5 chip.
22:16 They’ve also increased pricing across the board.
22:19 I wonder if we can go and see it. Yeah, we can.
22:23 So Mac Minis, they used to be $600 base price.
22:26 Now it’s $900 with the new M6 chip.
22:29 And then I believe, because Will has mentioned, would you like a new car or a Mac Studio?
22:34 That is the correct assessment because the new Mac Studios, I believe the high end is about $20,000 for a Mac Studio.
22:43 So if we go to pre-order, Lesgo’s Mac spec outside M5 Ultra for the good that, we’ll do that.
22:49 And then yeah, you can’t even buy the max amount of memory version with, like I said, 512 gigabytes of memory.
22:56 But yeah, base model is $10,800 for 256 gigabytes of memory.
23:03 So these things are not cheap, especially if you get storage upgrades or whatever.
23:06 Yeah, you can get it up to 18.3. But yeah, these are probably the easiest way to run very large models at home and also relatively cost effective.
23:20 Like I said, this has, I believe it is five times the memory bandwidth of a DGX Spark.
23:26 And two DGX Sparks are roughly the same price.
23:29 I believe two DGX Sparks right now will put you back about $9,000.
23:33 And this will be able to run it about five times faster for a given model.
23:38 And yeah, DGX Sparks have the same amount of memory as 256 gigabytes.
23:43 So yeah. So if you are super into local LM inference, you like running the big models, this is probably the best way to go.
23:51 I believe the previous M3 Ultra, which was the previous Ultra chip that they had, that max memory bandwidth was 800 gigabytes a second.
23:59 So this one, not counting any of the CPU speed and performance increases, will just be 50% faster because of that faster memory.
24:08 And yeah, Wilson’s cost comparison I saw was the RTX Pro 6000 versus M5 Ultra.
24:14 Yeah, that’s another RTX 6000 Pro. This is sort of the sort of in between if you can’t afford a B200 GPU from NVIDIA, but you want something more powerful than a 5090.
24:28 Yeah, the 6000 Pro is what you get.
24:31 These actually have gone up in price a ton.
24:33 When they originally came out, they were $8,000 each, I believe.
24:36 Now they are, I believe they go from about $1,200 to $13,000.
24:40 But these have 96 gigabytes of memory.
24:42 These will be, if you can fit the model into memory, this will be faster than the Mac.
24:47 I believe its memory bandwidth speeds is about 1.8 terabytes of memory bandwidth.
24:55 So it will be a bit faster. It also will handle parallelization much better.
24:58 Like it has a lot more flops. So it can do a lot more compute than the Mac can.
25:04 But the issue is that, yeah, it has, was that, one-third roughly the amount of memory as the Mac.
25:10 So you can’t run as large of a model on here.
25:12 It also uses, I believe, about three times the amount of electricity as well.
25:17 So sort of cost of ownership over time will also be higher.
25:21 And yeah, and you also do need a PC to plug this into.
25:24 I actually saw that figure, Will, on Twitter of someone saying, oh, it needs like a 3 to 5k computer to run it on.
25:30 You don’t need to have the 3 to 5k.
25:32 You can actually like, because for LM inference, everything matters about the GPU.
25:36 The rest of it doesn’t matter. You can plug this into like a $500 PC and it will work just fine.
25:44 So yeah, you don’t need necessarily a very expensive computer around this, but you need the rest of the computer, unlike the Mac, which comes with a computer, along with a very fancy LM inference chip.
25:55 So yeah, I thought that was interesting for those looking to get into the Mac space.
25:59 We have some good options. Yeah. And sticking now with…
26:07 I just wanted to add, I guess, real quick, that like, I think another trade-off is that with a DGX Spark, you can plausibly train something, like a little Quinn or to some LoRas or something,
26:21 but on a Mac, as far as I can tell, like all you can do is inference.
26:25 So you’re kind of, right? And like same with, you know, the RTX 6000.
26:30 I think a lot of companies are buying them to fine-tune their own little models to do stuff.
26:36 And you can’t really do anything. A Mac is only an inference device at this point, as far as I can tell.
26:42 Unless it’s changed recently. I mean, I know people do train stuff.
26:45 I believe Brandon, Snyder, has trained stuff on his Mac.
26:50 So MLX is sort of the Apple version of PyTorch for like running and writing models and training them.
26:58 So you can definitely train on Macs.
27:00 Training happens to be a paradigm where more compute matters versus memory bandwidth.
27:06 So the Macs will be slower. But yeah, usually for training, actually, the math makes sense usually to go and rent out GPUs in the cloud.
27:17 We’ll just be more cost effective usually.
27:19 Especially if you’re just doing sporadic training, like if you’re doing like incessant, you know, training, training, training over the course of, you know, like, I think the break-even is usually between like three to six
27:28 months straight of training, like 24-7.
27:31 Then it makes sense to buy GPU. But if you’re doing anything less than that, it’s probably easier to just go and rent a GPU in the cloud and use that to train instead.
27:43 Whereas inference, it makes more sense to, you know, inference, you are much more likely to be running 24-7 or, you know, you can batch stuff and, you know, do it later or whatever.
27:52 So inference is a workload that favors actually owning the hardware a bit more than training does.
28:00 So yeah, that’s why usually when I’m talking about local hardware, it’s all about inference.
28:05 And then when you want to train something, go to the cloud.
28:08 It’s training stuff locally. You really realize how little memory all these devices have for training.
28:14 Like B200s start running out of memory.
28:16 And I think they have like close to 200 gigabytes on there during training.
28:21 So yeah, use the cloud, GPUs, it makes life much easier if you want to train a model.
28:27 Yeah, any other questions or comments about the new Macs or Macs in general?
28:33 If not, we can move along to the other form factor from Apple, the iPhone.
28:38 So this is a question I believe that has come up a number of times here at AI Tools Club, and that is, I want to go and run a model on my phone.
28:48 What model should I go and pick? And so it’s always been sort of like, go look at like the liquid AI models are probably a good option or some of like the smaller Quenn models.
28:57 But now artificial analysis has put together a benchmark for us, gathering like a couple sort of decent benchmarks that are representative of tasks that you would want to run on your phone.
29:11 So yeah, they go and find that as per the recommendations here, the liquid foundation models are all really good.
29:20 We can see LFM 2.5 is sort of dominating the cost frontier, or rather, I guess it’s time frontier.
29:28 So how fast it takes to generate on a phone versus its intelligence.
29:33 So my takeaway from this is that this LFM 2.5, 2.6b model, this seems to be the best in terms of raw intelligence.
29:42 And then their 8 billion parameter mixture of experts model is the, if you want the best speed to cost trade-off, that’s the one to go and look at.
29:52 And then everything else you can go and ignore all the quens and demos and stuff.
29:55 They are not as good. So talking about these benchmarks, I sort of mentioned that these benchmarks are okay.
30:02 So a lot of them probably aren’t measuring everything you’d care about.
30:07 Let’s see. Is this the, do they show everything here?
30:10 I thought they had, oh yeah, they break down here.
30:12 So the, I think like three of these are good.
30:16 The other two are meaningless. So I’ll start with the good ones, which is Berkeley function calling leaderboard.
30:21 This is tool use for LLMs and how reliably they are able to go and use tools.
30:28 It is not meant to be a very complex reasoning benchmark outside of like picking the right tool or set of tools and then going and using them in the correct format.
30:37 So yeah, this is a good benchmark. We can see the LFM models are here near the top at around 75-ish percent.
30:43 I don’t know if you can see that on stream or not, but they’re near the top.
30:47 This is one definitely where the top end I wouldn’t read too much into.
30:50 You mostly just want to see any of the models down here, you don’t want to use.
30:53 That’s the big takeaway I’d say from this graph.
30:55 But I think up in like these, you know, upper 70s percents, there’s it’s mostly just evaluation noise between these models.
31:03 So I wouldn’t go and say like, ah, the NVIDIA model here is so much better than the liquid models.
31:07 Like, no. Next up is instruction following bench.
31:11 This does what it says on the CAN, where it’s how good is the model at following instructions.
31:15 We see actually the liquid model here at the top.
31:18 So very nice to see there, that’s something expected.
31:21 And then finally, and this is the real reason why I think the liquid model is really good, is hallucination rate.
31:26 One of the main things that we have learned from doing these benchmarks is that the smaller the model is, the less world knowledge it has.
31:34 And because of this, it is much more likely to hallucinate by being a small model, just in terms of number of parameters.
31:41 And so you need to go and train in the model to not hallucinate.
31:45 That’s not something that the model inherently learns itself.
31:47 You have to tell the model, no, don’t hallucinate on this answer.
31:50 Go do this instead. And so we can actually see, once again, the LFM 26B model is near the top here of not hallucinating, but it also does well of actually still answering questions.
32:01 I believe it’s here at 7%, whereas these other two models are like super low on accuracy because they just refuse most questions.
32:09 So the LFM model does a good job of balancing, you know, answering the few questions that it does know and then ignoring or not hallucinating on the rest of them and just saying, I don’t know the answer to that.
32:20 So they seem to have done a good job with that.
32:23 Then they have GPQA Diamond at Math 500.
32:25 These benchmarks are meaningless. These are the ones that I say you could just ignore.
32:28 I think these are adding noise to it.
32:30 GPQA Diamond has been out for a gazillion years now, is definitely overfit on.
32:35 And I think the scores that we see here are roughly just in order of parameters, essentially, because it’s just how much knowledge about the benchmark can the model hold.
32:44 And the larger the model is, the more information it can hold about these benchmarks.
32:49 So I’d ignore that. Math 500 as well.
32:51 This is a pretty saturated benchmark.
32:52 We see like at the high end, like 96, 97% for these models.
32:57 I think it’s useful for the low end.
32:59 Like if models struggles on this, don’t use that model.
33:02 But in terms of like if the model is good, it doesn’t really mean much being at the top here because this is another benchmark that’s been around for a million years.
33:09 Very famously, overfit on. There used to be the Hugging Face LLM leaderboard where people’s fine tunes would all get benchmarked and stuff.
33:17 And this is one of the benchmarks that got overfit on there.
33:21 So there’s not any meaningful signal there.
33:24 We have people in the chat cheering for Liquid AI.
33:27 Yes, they are based in Boston here around the MIT ecosystem.
33:32 Have a lot if we have any Liquid AI folks at Sunday.
33:35 I think we’re actually planning an upcoming hack with them in the future.
33:38 I don’t know how much more I can say than that.
33:40 But yeah, I think there’s stuff in the works.
33:42 We’ve been talking with them for a while.
33:46 And then, oh, Will says he’s been working on this as well.
33:50 We’ll open this link up. Natural Stupid says, is it possible that the model keeps evolving on the iPhone and everyone gets personalized model eventually?
33:57 I think somebody actually did a hack on that.
34:01 I think the, I mean, that’s a whole discussion around memory.
34:05 I think we have Pavel in here. Yeah, we do.
34:08 But there’s a whole discussion around, yeah, do you train the model to get better or do you build a good memory system for the model to use, for it to evolve and work with?
34:18 I think building a proper memory system around the model versus training it is something that we could see.
34:24 But I don’t think we’re going to be training the models on our phones.
34:26 At least not at this point, because that’s a little bit too much for essentially.
34:31 It’s very hard to go and fine-tune models off small amounts of data.
34:35 Usually you want thousands and thousands of examples, which is not something you can really train on your phone.
34:43 So yeah, Will, I assume… Oh, he, yeah, just got it running.
34:48 Yeah. Oh, one of the liquid models on what is this?
34:51 Oh, an old Kindle fire. That’s super cool to see.
34:56 We should start including Felony Bench in these comparisons.
34:58 It’s the nice thing is, I believe, yeah, all of these small models are not smart enough to go and hack into other systems.
35:05 So for those that don’t know, Felony Bench is the models hacking out of their sandboxes during reinforcement learning.
35:12 So that’s what we saw during the OpenAI and HuggingFace incident.
35:16 That was a couple weeks ago, where one of OpenAI’s models that they were evaluating broke out of its sandbox and hacked the HuggingFace servers to go and look for the answer to the question it was being asked.
35:28 So yeah, those big models, like I think we’ve seen it from Anthropic, Meta, I believe XAI, OpenAI, and I believe Kimmy also mentioned that that happened on their end.
35:41 But that’s it. And those are like basically all the biggest models out there.
35:45 So yeah, but I don’t think we need to worry about that for these small models.
35:49 They will all score zero on that benchmark for a while now.
35:57 But yeah, if you or anyone else in the future, oh, actually, I guess one other thing.
36:01 They also show speed for all of these models, because that matters as well for, you know, if you’re deploying it on a phone, is the model can be slow a lot of the times.
36:14 And so there needs to be a balance of, yeah, intelligence and then also number of output tokens and also the model architecture for how fast it is.
36:22 And that’s where you can see the LFM models once again doing very well.
36:25 Basically everything in the bottom of the chart here being LFM models.
36:28 And then the 2.6B model right here at eight seconds for responses.
36:33 Respectable, I would say. I think this is just average generation time across the entire benchmark suite.
36:40 So yeah, that seems to be the sweet spot for now.
36:43 So yeah, if any new models, small models come out and you’re wondering how they do, I would go and check out this.
36:50 And yeah, hopefully we can have some more hacks in the future at Sunday with people running these models on their phone.
36:56 Because I think it is a super cool form factor.
36:58 And maybe we can start doing like small agentic things with it.
37:01 You know, like someone can make more proper Siri, finally.
37:04 Siri that actually works. But yeah, any questions or comments about on-device AI?
37:17 Just a quick comment. So I’ve got Pixel 11 for myself and it has this crazy fantastic voice typing system called Rambler.
37:29 I think that’s kind of the only good thing about the phone actually.
37:32 But it’s just unbelievably. It’s literally agentic.
37:36 It can go and change things that you’ve said before.
37:39 It can translate to another language in the real time.
37:46 It seems there is a real-time LLM or something running inside of it.
37:53 So I don’t think anybody has built anything similar that runs not on a Google phone.
38:01 But I think it’s just an unbelievable opportunity, seriously.
38:06 Quick comment from myself. Does that run on the phone?
38:09 I think it does because latency is extremely fast.
38:13 Like to us, I don’t know for sure. That’s amazing.
38:16 Yeah. It’d be interesting to go and reverse engineer it to figure out what it’s actually…
38:24 Like, I wonder if this is being built on top of the Google or the JAMA 4 E4B models.
38:31 So no, Google trained these for on-device deployments.
38:34 So I wonder if they sort of like fine-tuned it and add, I think they actually have audio understanding out of the box for them.
38:40 So it’d be interesting to see if Google use that model or if they have something else under the hood that they’re using for it.
38:46 And then, yeah, can we go and pull it out and use it somewhere else potentially?
38:50 And just one more thing. So I was just completely, my mind was blown out when I was yapping to it for a couple of minutes.
38:59 And then I was like, oh, actually, send this message in Russian.
39:04 And then all the things I was yapping in English, it immediately put as a block of Russian text.
39:10 Like in a few seconds. Just unbelievable.
39:15 That’s super cool. Shikan asked, does this mean Whisperflow is dead?
39:19 I don’t think so. Just because this is isolated to the pixel, whereas Whisperflow, I believe, is Whisperflow also works on non-Mac devices, but it’s primarily used on Mac.
39:28 And they also are very big, have a lot of momentum.
39:31 But I think, I mean, if this system is that good, you could probably definitely disrupt Whisperflow by building something like this yourself.
39:40 Because I know Whisperflow is actually fairly bare bones in terms of functionality and features.
39:45 You can go and Vibe code Whisperflow in like a day or two.
39:48 It’s not that complex of an app. It’s just basic automatic speech recognition with a few extra like bells and whistles for a nice user experience.
40:07 Sorry, that’s my dog whining. Cool.
40:15 And sticking, actually, what do we want to talk about?
40:18 Yeah, well, sticking with the idea of killing Whisper Flow, I saw this where this guy, he’s a pretty big name.
40:25 He runs the AI Engineer Conference, pretty good YouTube series.
40:29 They have a bunch of people come and give talks, upload them all to YouTube.
40:31 I would definitely go and recommend checking those out.
40:34 But he’s been talking about how he’s been sort of systematically going and removing a whole bunch of different SaaS services that he goes and relies on and building new ones and better ones using AI for his use cases.
40:48 So I wanted to talk with the community here and see if anybody is doing this, what kind of success they had or failures with replacing different SaaS products in their local stack, if they’re doing that, or why you’re not
41:00 doing that for certain cases. Has anybody tried this or done this before?
41:06 I’ve done something very simple. Some of the marketing apps, they seemed like they were doing very simple things for me.
41:16 And so I sort of one-shotted it with GPT-56Sol.
41:22 And it’s doing okay. I mean, it’s doing as well as the paid services, at least for my case.
41:30 That’s my only direct attempt to replace Assass.
41:37 And essentially, it’s got search strings and it’s looking for potential customers on places like Reddit or forums or whatever.
41:49 We’ll see how it works out. Nice. Is anybody else?
41:57 How many people here actually pay for different SaaS products that you would even consider replacing with AI?
42:01 You know, I assume you’re not going to replace your banking app with AI.
42:05 But does anyone even have good candidates for something that they could replace?
42:13 Sentry is one for monitoring. I don’t think that it has to be that special.
42:21 And what’s the other one that I was thinking of?
42:26 Oh, the tracing, Lang Trace. I’m not sure that that’s all that special, really.
42:33 But it’s easier to use. I mean, that’s all.
42:36 I’d say no, I’ve had beef with all of the Lang chain or Lang whatever products.
42:42 I’ve used them since they first came out, since ChatGPT released.
42:46 I’ve tried them out, and I’ve never really liked them or, you know, thought that the abstractions that they gave were very useful.
42:55 So, yeah, I would agree that this is one that you definitely like, whenever I’m doing and building my own sort of like prompt or like agentic things, a lot of the times if they’re small and simple, I will go and roll it myself
43:07 or use something that isn’t Langchain or one of their associated products.
43:11 Because yeah, I’m just not a fan of their ecosystem.
43:15 Pavel says granola, following along with like Whisperflow and stuff.
43:19 Yeah, granola is just automatic speech recognition just running in the background for you, like at all times, I think.
43:26 And then just spinning that into an LM to get a summary.
43:28 So that’s something definitely very feasible to go and replace with either a full local stack or just something that you send up to the cloud or fairly cheap.
43:43 Interesting. Okay. So it seems we don’t have too many SaaS users, at least people who are interested in replacing them or anything like that.
43:50 Cool, cool. We will touch briefly on what Quentin brought up earlier, which is the Ox Alpha model.
44:01 I won’t give too much time for this because it is a stealth model, but many people have been seeing, it’s been making a splash.
44:07 You can use it for free across a bunch of different platforms.
44:10 I think Hermes Agent, you can use it for free there, which is their open claw competitor.
44:16 You can also use it for free in OpenCode.
44:18 I think also Klein, you can use it for free.
44:20 And they’ve been also saying that they have a ton of capacity to go and run this model, which interestingly, where is it?
44:29 From the, this is from the Open Code team.
44:31 They basically say that 10 trillion tokens of deep seek flash is, I believe they say, yes, 1,000 B300s.
44:38 So a ton of GPUs to serve 10 trillion tokens.
44:41 And whoever is behind Aux Alpha is offering 10x that amount for free.
44:46 So that’d be, you know, like 10,000 B300 equivalents, which a B300 is like $6 an hour.
44:51 So that’s $60,000 an hour if this model is the same size as Deep CD V4 Flash.
44:58 But yeah, this model, I believe it’s like 95% assumed, I think that’s what the current betting markets are said, that this is a GLM model.
45:06 It actually shares the same tokenizer as the GLM series of models.
45:10 So we at least know at the very, yeah, that it is at least a fine-tune of a GLM model.
45:16 But yeah, this is suspected to be GLM 5.3 flash is what people have been calling it.
45:21 So probably a small model. My guess at these token rates, it’s probably like a 30 billion parameter model.
45:28 The thing is, is that, and the reason people are getting so hyped for it, is that it is a free or cheap model that is doing fairly well.
45:35 So yeah, I think, is this like a cohesive, this is like a decent-ish wrap.
45:40 But on DeepSuite, which is right now, I would say, the best benchmark for agentic coding applications, we can see Aux Alpha is up here above, you know, Opus 4.5.
45:53 It’s better than Luma as well. And then also, usually, these smaller models are very token hungry.
46:00 I don’t think, yeah, they have like Quenn over here.
46:02 This isn’t the smaller Quenn, but I think the smaller Quenn uses even more tokens.
46:06 But yeah, the small models tend to use a ton of tokens.
46:09 But we see here that this is actually a pretty efficient model as well.
46:13 So yeah, very interesting to go and see.
46:19 Yeah, so yeah, there’s been a lot of hype around this model.
46:21 Don’t know when this is actually going to be released.
46:23 I guess you hear people in the comments saying it’s GLM 5.3 Flash.
46:27 But yeah, we don’t know when this is going to be released, but this should be a pretty interesting model because hopefully it will either be very cheap or a good model that you can run locally.
46:35 Because yeah, if this is a 30 billion parameter model, I’m definitely going to be running this locally on my 3090s.
46:41 There’s also rumors that it’s a bit bigger than that, but that it can fit into your DGX Spark as well.
46:46 So we’ll have to wait and see for that.
46:50 Will saying, what billionaire has that much compute to give away?
46:54 That’s what I said. It’s very weird.
46:56 I think this 100 trillion token number is a bit of a sort of marketing thing because using 100 trillion tokens is insane per day.
47:08 All the feverish around the deep CP4 flash release, that only peaked at, I think, around like 20 trillion tokens per day.
47:17 So this would be 5x that. It’s also, because of that, probably a pretty efficient model.
47:23 Interestingly, GLM, they actually just brought online a bunch of compute from Huawei.
47:28 So they are starting to bring online Chinese GPUs, essentially, to go and serve this stuff.
47:33 So I think they mentioned that they have a bunch of new superpods is what they’re being called, which I think is like a couple thousand Huawei 950s, I believe, which are roughly like H100 capability level.
47:46 And that’s fully like local, as in local to China hardware, where yeah, they control that entire supply chain.
47:54 So that’s probably where the compute for this is coming from.
47:59 But yeah. But yeah, people have been saying that like, oh, this, like, only Elon could be doing this or something like that to be giving this much away.
48:06 But I don’t think this is the case.
48:08 This is actually GLM. They famously did this for their GLM5 model.
48:12 I think it’s called like Pink Pony or something like that, that model.
48:17 And that also garnered a lot of hype.
48:19 And that, you know, yeah, that sort of like helped launch that model and, you know, gain a little bit of aura because these stealth models, like Nano Banana, that was the code name for the Gemini image generation model
48:33 that was used. And everybody liked it so much that that’s why it’s still called that.
48:37 It was not meant to be called Nano Banana when it actually released, but that’s the name that stuck and that everybody liked.
48:43 So yeah. Daniel asks if I got another 3090.
48:46 You’re putting the pieces together here, Daniel.
48:49 That is why I’m getting all these fancy server components and all of these, all this memory and all this stuff.
48:55 So yeah, I have a second 309D and it does not fit into my desktop at home.
48:59 I thought I had a second PCIe 16X slot, but I don’t.
49:03 So yeah, now I have to go and build a new fancy computer.
49:08 Oh no. So yeah. So computer community will have.
49:13 Can you just put it in the other small PCIe and then get an NV link?
49:17 Wouldn’t that be enough? I could if I wanted to.
49:22 But I’ve been building this sort of like more proper LLM server because this is just my desktop, what I’m streaming off this right now.
49:32 But I’ve had plans on building an actual proper server for running LLMs.
49:37 And so yeah, I just decided like instead of trying to jury-rig it into this desktop and spend a bunch of money on parts that won’t transfer to the server in the future, just pull the trigger now on building the server and
49:48 just use that as my reason for spending all that money.
49:52 And then yeah, like NV-Link, it’s I think for inference, if you’re on like PCI 3.0, like 16X, NV-Link only actually increases the speeds by like 5 to 10%.
50:03 And that’s also assuming if you’re doing like Tensor Parallel or something like that, like using both GPUs to run the model.
50:10 And Will says, yeah, it’s dicey on consumer hardware.
50:14 And yeah, they’re also, it’s expensive, I think.
50:16 NVLink is like $300 or $400 for 3090s.
50:19 So it’s like one-third of the cost of the 3090 just to connect them to each other.
50:25 But yeah. Any questions about Aux Alpha or comments?
50:30 Has anybody actually gone and used it?
50:31 Like I said, it is free tokens, basically unlimited free tokens to go and do stuff, and it’s a pretty smart model.
50:36 So yeah, I would go, if for no other reason, go check it out for that.
50:41 You can go and run whatever silly experiments you want with this, and it won’t cost you anything.
50:46 Yeah, basically, yeah, unlimited rate limits as well.
50:52 It’s Natural Stupid says OpenAI releases their chip.
50:56 Oh yeah, we can talk about this briefly if we want.
51:00 Yeah, so I think we also covered this momentarily.
51:06 But yeah, this is Jalapenio. This is Opening Eyes internal chip.
51:10 So you can imagine this is similar to the Cerebris chips or the chips from Grok with a Q that NVIDIA bought as well, where these are chips meant specifically for super fast inference.
51:25 And so they compare to a bunch of NVIDIA deployments, and they basically show that they can do much faster tokens per second per user, which is sort of like the Cerebrus metric, and that also they deliver more
51:39 throughput per kilowatt, so tokens per kilowatt, which is the main thing that data center designers look for when purchasing and sort of like getting these chips, because electricity is your big cost at these data centers.
51:54 So the fact that, yeah, you have higher throughput per watt is very good.
51:58 So yeah, they actually show, I don’t know, they don’t have the direct, but Semi-Analysis, one of the sort of like silicon manufacturer commentators, they ran through their benchmarks with
52:12 these and they actually saw that this outdoes the new Vera Rubin GPUs from NVIDIA that are coming out as well, not only the B200s.
52:21 So they seem to already be ahead of sort of like the latest and greatest from NVIDIA, which seems to be very good.
52:28 This is definitely probably not something that we’re going to be able to ever buy, but OpenAI will be relying on this in their new massive data centers that they are building out.
52:36 And this will be saving them a lot of money.
52:38 So that’s pretty cool. Then Shrikan says, Hugging Face Explorer Cell.
52:46 Yes, I saw this as well. We’re Hugging Face, for those that don’t know, this is where all of, this is like the GitHub for LLM models.
53:00 So all of the models and data sets, they get hosted on Hugging Face.
53:04 So you can see here, actually, I have a bunch of models that I go and like, if you want to go and download these things, I think the most popular ones were like the old Llama 3.2 models.
53:13 I went and quantitized them. And so, yeah, like you can go and download it from here.
53:19 Same with data sets. And they also have some compute stuff if you just want to go and test out models really quick.
53:24 They’re a really big name in the open source space.
53:27 And yeah, they are fielding offers or sort of like, yeah, looking around to sell to, I believe it’s not like another company.
53:34 It’s basically sort of like not private equity, but something along those lines at around $13 billion, which would be very interesting.
53:42 And hopefully the company doesn’t, you know, change its ways at all.
53:46 Because I know there’s been a lot of people saying that like, oh, they don’t make enough money or, you know, their business model is bad where they’re just spending all this money on storage.
53:53 And there’s no real, I think they only make money really on enterprise plans and like some of their compute stuff.
53:59 But they keep saying that they are profitable.
54:01 It’s not an issue. So yeah. But yeah, interesting that they might be getting sold in the future here.
54:09 It is six o’clock, but I will just leave you with some things to go and look at, which is the humanoid robot Olympics that just happened in China.
54:20 There’s a bunch of fun and silly clips of robots doing all the different types of Olympic sports.
54:25 So you can see here, this is like, I think the like 400 meter, but then it’s also doing like the 100 meter dash where it broke the human world record from Usain Bolt.
54:34 They also did long jump and fighting and then like dexterity stuff.
54:38 I think there’s some videos of them playing tennis, which I think that was actually the most interesting demonstration.
54:44 Like this is by far the most complex task that these models have to go and do where they have to like move around and calculate where the ball is going to be.
54:51 Like this is a fully autonomous demo.
54:53 So I thought this was really cool. So yeah, during the week, I’d say, yeah, go check out a bunch of these videos, a bunch of super cool stuff, to see what the future for robotics look like.
55:02 There’s also a bunch of really cool explosions and stuff, especially from the running ones, where the robots just sort of fold on themselves and blow up.
55:10 So yeah, go enjoy a look at those. But yeah, we are out of time today.
55:17 Are there any final questions, comments, or concerns before we all depart?
55:26 If not, thank you all for coming this week and hope to see you all again next week.
55:30 Thank you.
Subscribe to get the latest AI news in your inbox every week!