I Gave Two AI Supercomputers a Real Job
Read full transcript 16 segments
-
So, I recently covered the Camino Grando So, I recently covered the Camino Grando and the DJX station, but I have Wendell and the DJX station, but I have Wendell and the DJX station, but I have Wendell here today visiting me and he's going to here today visiting me and he's going to here today visiting me and he's going to show us Turnstone, which is how to use show us Turnstone, which is how to use show us Turnstone, which is how to use these machines in a real scenario. these machines in a real scenario. these machines in a real scenario. >> I like doing developer and developer >> I like doing developer and developer >> I like doing developer and developer things. It's fun. things. It's fun. things. It's fun. >> Let's talk about the hardware a little >> Let's talk about the hardware a little >> Let's talk about the hardware a little bit. You haven't had a chance to test bit. You haven't had a chance to test bit. You haven't had a chance to test the Grand yet, but you will. I'm going the Grand yet, but you will. I'm going the Grand yet, but you will. I'm going to give it to Wendell, so watch his to give it to Wendell, so watch his to give it to Wendell, so watch his channel for some updates. He's going to channel for some updates. He's going to channel for some updates. He's going to go deep. go deep. go deep. >> It's eight RTX Pro 6000s. And this is >> It's eight RTX Pro 6000s. And this is >> It's eight RTX Pro 6000s. And this is kind of the question, isn't it? eight kind of the question, isn't it? eight kind of the question, isn't it? eight RTX Pro 6000s versus DGX station, RTX Pro 6000s versus DGX station, RTX Pro 6000s versus DGX station, >> right? There's a reason why you don't >> right? There's a reason why you don't >> right? There's a reason why you don't hear them right now, but they are hear them right now, but they are hear them right now, but they are actually running and they're busy. It's actually running and they're busy. It's actually running and they're busy. It's getting toasty in here. We have 6,000 getting toasty in here. We have 6,000 getting toasty in here. We have 6,000 watts of heat generation in here. This watts of heat generation in here. This watts of heat generation in here. This is nuts. I've got the Grandondo. It's is nuts. I've got the Grandondo. It's is nuts. I've got the Grandondo. It's using four plugs as I showed in my other using four plugs as I showed in my other using four plugs as I showed in my other video and the DJX station is running as video and the DJX station is running as video and the DJX station is running as well off of the Jackary. well off of the Jackary. well off of the Jackary. It's hilarious, I know, but it works. It's hilarious, I know, but it works. It's hilarious, I know, but it works. We're down to 86% and it's right now We're down to 86% and it's right now We're down to 86% and it's right now idling at 319 watts. There's so many idling at 319 watts. There's so many idling at 319 watts. There's so many watts. What are you doing with all watts. What are you doing with all watts. What are you doing with all these? And that's down to 86% in just these? And that's down to 86% in just these? And that's down to 86% in just about 45 to 50 minutes. You can see why about 45 to 50 minutes. You can see why about 45 to 50 minutes. You can see why this door is closed. It's kind of loud this door is closed. It's kind of loud this door is closed. It's kind of loud in there. By the way, these machines are in there. By the way, these machines are in there. By the way, these machines are running headless. I got the DJX station running headless. I got the DJX station running headless. I got the DJX station up and running remotely through its up and running remotely through its up and running remotely through its management port. This machine is management port. This machine is management port. This machine is completely off right now, but I could completely off right now, but I could completely off right now, but I could still get into it. How? Well, that's still get into it. How? Well, that's still get into it. How? Well, that's through the management port. system 10G through the management port. system 10G through the management port. system 10G Ethernet port goes here. And this is the Ethernet port goes here. And this is the Ethernet port goes here. And this is the management port. And here I can control management port. And here I can control management port. And here I can control it. I can check the sensors, the LEDs, it. I can check the sensors, the LEDs, it. I can check the sensors, the LEDs, the logs, update the BIOS. And here's
-
the logs, update the BIOS. And here's the logs, update the BIOS. And here's the KVM. So I can power the machine on the KVM. So I can power the machine on the KVM. So I can power the machine on from here. And there it is. It's turning from here. And there it is. It's turning from here. And there it is. It's turning on on on just like any other desktop. You can just like any other desktop. You can just like any other desktop. You can still hear me over it. And this is still hear me over it. And this is still hear me over it. And this is pretty much what it sounds like all the pretty much what it sounds like all the pretty much what it sounds like all the time, even when it's cranking away. I time, even when it's cranking away. I time, even when it's cranking away. I don't expect one AI prompt to build an don't expect one AI prompt to build an don't expect one AI prompt to build an entire project for me. In reality, I'm entire project for me. In reality, I'm entire project for me. In reality, I'm constantly moving between models constantly moving between models constantly moving between models depending on the job. GPT for research, depending on the job. GPT for research, depending on the job. GPT for research, claude for coding, Gemini for massive claude for coding, Gemini for massive claude for coding, Gemini for massive context, nano banana, midjourney, flux context, nano banana, midjourney, flux context, nano banana, midjourney, flux for images, and then I've got seed dance for images, and then I've got seed dance for images, and then I've got seed dance and cling for video. That's why chat lm and cling for video. That's why chat lm and cling for video. That's why chat lm by Abacus Aai makes sense. It brings day by Abacus Aai makes sense. It brings day by Abacus Aai makes sense. It brings day one support for the latest GPT, Claw, one support for the latest GPT, Claw, one support for the latest GPT, Claw, Gemini, Grock, Deepseek, and more in one Gemini, Grock, Deepseek, and more in one Gemini, Grock, Deepseek, and more in one place the moment they drop. Pick any place the moment they drop. Pick any place the moment they drop. Pick any model from the interface or let route model from the interface or let route model from the interface or let route LLM automatically choose the best model LLM automatically choose the best model LLM automatically choose the best model for each prompt. Create professional for each prompt. Create professional for each prompt. Create professional presentations with graphs and charts and presentations with graphs and charts and presentations with graphs and charts and deep research detailed content. Need deep research detailed content. Need deep research detailed content. Need human sounding copy? Human eyes rewrites human sounding copy? Human eyes rewrites human sounding copy? Human eyes rewrites text to defeat AI detectors. Need text to defeat AI detectors. Need text to defeat AI detectors. Need visuals? Pick frontier or open- source visuals? Pick frontier or open- source visuals? Pick frontier or open- source models. And when you need more than models. And when you need more than models. And when you need more than chat, Abacus AI agent can help build chat, Abacus AI agent can help build chat, Abacus AI agent can help build complex apps and websites, connect complex apps and websites, connect complex apps and websites, connect payments, or run 247 agents that keep payments, or run 247 agents that keep payments, or run 247 agents that keep working through longer tasks. The best working through longer tasks. The best working through longer tasks. The best part is app hosting, back-end database, part is app hosting, back-end database, part is app hosting, back-end database, and off support comes with the and off support comes with the and off support comes with the subscription. All that starts at just subscription. All that starts at just subscription. All that starts at just $10 a month. Way cheaper than paying for $10 a month. Way cheaper than paying for $10 a month. Way cheaper than paying for all those subscriptions separately.
-
all those subscriptions separately. all those subscriptions separately. Check out chatlm.abacus.ai Check out chatlm.abacus.ai Check out chatlm.abacus.ai or click the link below. We're going to or click the link below. We're going to or click the link below. We're going to turn to Turnstone and we're actually turn to Turnstone and we're actually turn to Turnstone and we're actually going to give it a couple of projects going to give it a couple of projects going to give it a couple of projects that I want done. Cloning the repo. In that I want done. Cloning the repo. In that I want done. Cloning the repo. In this current working directory, you'll this current working directory, you'll this current working directory, you'll find Turnstone. Read the instructions find Turnstone. Read the instructions find Turnstone. Read the instructions and set this up on this Mac. Boom. This and set this up on this Mac. Boom. This and set this up on this Mac. Boom. This is how I work now. is how I work now. is how I work now. >> I worry about the future. >> I worry about the future. >> I worry about the future. >> Wait, it just deleted my home directory. >> Wait, it just deleted my home directory. >> Wait, it just deleted my home directory. >> Woo! Dangerously doing things. >> Woo! Dangerously doing things. >> Woo! Dangerously doing things. >> I didn't get it, did I? No, I didn't. >> I didn't get it, did I? No, I didn't. >> I didn't get it, did I? No, I didn't. Okay, I have a scenario. I run sessions Okay, I have a scenario. I run sessions Okay, I have a scenario. I run sessions with cloud code and I run it on multiple with cloud code and I run it on multiple with cloud code and I run it on multiple DJX sparks just to kind of distribute DJX sparks just to kind of distribute DJX sparks just to kind of distribute the load, but I want that automated. I the load, but I want that automated. I the load, but I want that automated. I want to go to one thing and tell it go want to go to one thing and tell it go want to go to one thing and tell it go start a session. You pick what's the start a session. You pick what's the start a session. You pick what's the best place to run it. This is where best place to run it. This is where best place to run it. This is where Turnstone will come in, right? I thought Turnstone will come in, right? I thought Turnstone will come in, right? I thought so. Oh man, that's going to be nice. I'm so. Oh man, that's going to be nice. I'm so. Oh man, that's going to be nice. I'm going to have a fleet. going to have a fleet. going to have a fleet. >> So, the first thing we want to do is set >> So, the first thing we want to do is set >> So, the first thing we want to do is set up a model. If you put in the base URL, up a model. If you put in the base URL, up a model. If you put in the base URL, if you got something local, if you know if you got something local, if you know if you got something local, if you know the IP address off top of your head, the IP address off top of your head, the IP address off top of your head, >> you could load up two models. Actually, >> you could load up two models. Actually, >> you could load up two models. Actually, that's not a bad idea.
-
that's not a bad idea. that's not a bad idea. >> Let's see if this model is loaded. We're >> Let's see if this model is loaded. We're >> Let's see if this model is loaded. We're connected and it's enabled. connected and it's enabled. connected and it's enabled. >> And you can add more models. You can >> And you can add more models. You can >> And you can add more models. You can have a mix of models that are good at have a mix of models that are good at have a mix of models that are good at really different things, but you can really different things, but you can really different things, but you can also configure Turnstone. Like in also configure Turnstone. Like in also configure Turnstone. Like in general, this kind of request goes to general, this kind of request goes to general, this kind of request goes to this model. These are the personas that this model. These are the personas that this model. These are the personas that it ships with, but you can build your it ships with, but you can build your it ships with, but you can build your own. Some of them have access to MCP own. Some of them have access to MCP own. Some of them have access to MCP model context protocol. So like if you model context protocol. So like if you model context protocol. So like if you want to run MCP servers, it can want to run MCP servers, it can want to run MCP servers, it can inventory them and use them. So you have inventory them and use them. So you have inventory them and use them. So you have engineer, executive, manager, engineer, executive, manager, engineer, executive, manager, orchestrator, researcher, scribe, and orchestrator, researcher, scribe, and orchestrator, researcher, scribe, and writer. Like depending on what your writer. Like depending on what your writer. Like depending on what your writing preferences are, you can teach writing preferences are, you can teach writing preferences are, you can teach it. It's like no m dashes or no extra it. It's like no m dashes or no extra it. It's like no m dashes or no extra bits for you. Um, which is bits for you. Um, which is bits for you. Um, which is >> this is scribe. >> this is scribe. >> this is scribe. >> Mhm. >> Mhm. >> Mhm. >> I need that. >> I need that. >> I need that. >> Yeah. You can also save memories. So you >> Yeah. You can also save memories. So you >> Yeah. You can also save memories. So you can say whenever I tell you to do this, can say whenever I tell you to do this, can say whenever I tell you to do this, remember to do that and it'll just build remember to do that and it'll just build remember to do that and it'll just build memory that over time. memory that over time. memory that over time. >> That's really cool. Can you tell me >> That's really cool. Can you tell me >> That's really cool. Can you tell me about these nodes? What are these? Are about these nodes? What are these? Are about these nodes? What are these? Are these basically just virtual machines these basically just virtual machines these basically just virtual machines that are running here inside Docker or that are running here inside Docker or that are running here inside Docker or what? what? what? >> Yes. And they're very limited right now >> Yes. And they're very limited right now >> Yes. And they're very limited right now because we haven't mounted a workspace because we haven't mounted a workspace because we haven't mounted a workspace or anything like that. So if you had a or anything like that. So if you had a or anything like that. So if you had a shared workspace, it's mounted at shared workspace, it's mounted at shared workspace, it's mounted at SLworkspace on the nodes by default. And SLworkspace on the nodes by default. And SLworkspace on the nodes by default. And this manager orchestrator is going to this manager orchestrator is going to this manager orchestrator is going to create workers with different personas create workers with different personas create workers with different personas for that kind of a task, including a for that kind of a task, including a for that kind of a task, including a manager and an evaluator. And that will manager and an evaluator. And that will manager and an evaluator. And that will help bring the project to completion. I help bring the project to completion. I help bring the project to completion. I have a bunch of repositories on my have a bunch of repositories on my have a bunch of repositories on my GitHub. And one of the things I like to GitHub. And one of the things I like to GitHub. And one of the things I like to do is I have to set up a lot of machines do is I have to set up a lot of machines do is I have to set up a lot of machines here, Linux, Macs, Windows, and I here, Linux, Macs, Windows, and I here, Linux, Macs, Windows, and I created these repositories. I don't have created these repositories. I don't have created these repositories. I don't have one for Mac yet. Is uh to bootstrap a one for Mac yet. Is uh to bootstrap a one for Mac yet. Is uh to bootstrap a bunch of developer tools and set it up bunch of developer tools and set it up bunch of developer tools and set it up as a developer machine. So I have this as a developer machine. So I have this as a developer machine. So I have this Windows one. Problem is this works on Windows one. Problem is this works on Windows one. Problem is this works on x86 machines but not on armbbased x86 machines but not on armbbased x86 machines but not on armbbased machines. And I ran into this recently machines. And I ran into this recently machines. And I ran into this recently in a recent video. This was a problem
-
in a recent video. This was a problem in a recent video. This was a problem for me. So I want turn stone to turn for me. So I want turn stone to turn for me. So I want turn stone to turn every stone overh. >> Okay. Okay. >> Okay. Okay. >> Been hanging out too much. >> Been hanging out too much. >> Been hanging out too much. >> And figure out how to create another >> And figure out how to create another >> And figure out how to create another branch for me to do a setup on arm. branch for me to do a setup on arm. branch for me to do a setup on arm. Let's turn over a leaf and stone and not Let's turn over a leaf and stone and not Let's turn over a leaf and stone and not leave any stones unturned. leave any stones unturned. leave any stones unturned. This is getting bad. This is getting bad. This is getting bad. >> Click dashboard for the workspace. And >> Click dashboard for the workspace. And >> Click dashboard for the workspace. And so paste it in, but then you got to so paste it in, but then you got to so paste it in, but then you got to explain what you want. You can't just explain what you want. You can't just explain what you want. You can't just sort of grunt in the general direction sort of grunt in the general direction sort of grunt in the general direction of the uh repository. of the uh repository. of the uh repository. >> I want you to look at this repository >> I want you to look at this repository >> I want you to look at this repository and create a separate branch that'll and create a separate branch that'll and create a separate branch that'll handle Windows on ARM situations because handle Windows on ARM situations because handle Windows on ARM situations because there's certain uh software that this there's certain uh software that this there's certain uh software that this installs that are only compatible with installs that are only compatible with installs that are only compatible with x86. Python, node are being some of the x86. Python, node are being some of the x86. Python, node are being some of the examples. I want you to go through and examples. I want you to go through and examples. I want you to go through and figure out every single example of a figure out every single example of a figure out every single example of a software that's installed and create the software that's installed and create the software that's installed and create the ARM equivalent of this. So before we run ARM equivalent of this. So before we run ARM equivalent of this. So before we run this, where is this running this, where is this running this, where is this running >> on those nodes? >> on those nodes? >> on those nodes? >> On the nodes, does it matter which node?
-
>> On the nodes, does it matter which node? >> On the nodes, does it matter which node? >> It'll schedule it itself. >> It'll schedule it itself. >> It'll schedule it itself. >> And the nodes are using an LLM. >> And the nodes are using an LLM. >> And the nodes are using an LLM. >> The nodes are connected through the >> The nodes are connected through the >> The nodes are connected through the model service model service model service >> which is going to use the Grando. >> which is going to use the Grando. >> which is going to use the Grando. >> Mhm. >> Mhm. >> Mhm. >> And also create an auditor that'll check >> And also create an auditor that'll check >> And also create an auditor that'll check the work. I'll start by exploring the the work. I'll start by exploring the the work. I'll start by exploring the repository to understand what it does. repository to understand what it does. repository to understand what it does. And there it goes. And there it goes. And there it goes. >> All right. While that's working, go and >> All right. While that's working, go and >> All right. While that's working, go and let's go and create another task and let's go and create another task and let's go and create another task and another task and another task cuz didn't another task and another task cuz didn't another task and another task cuz didn't you have some other tasks you wanted to you have some other tasks you wanted to you have some other tasks you wanted to do like find a bug or something? do like find a bug or something? do like find a bug or something? >> I do. Where's the new task button? >> I do. Where's the new task button? >> I do. Where's the new task button? >> Uh you just >> Uh you just >> Uh you just I think uh I think uh I think uh >> issue issue issue number one. I have >> issue issue issue number one. I have >> issue issue issue number one. I have this repository called code needle which this repository called code needle which this repository called code needle which is um actually a benchmark for code for is um actually a benchmark for code for is um actually a benchmark for code for LLMs. So we're going to do a little bit LLMs. So we're going to do a little bit LLMs. So we're going to do a little bit of a meta thing and I'm going to have it of a meta thing and I'm going to have it of a meta thing and I'm going to have it evaluate this and look for bugs. evaluate this and look for bugs. evaluate this and look for bugs. specifically. There's one bug that it specifically. There's one bug that it specifically. There's one bug that it has that I want fixed. Fix the generated has that I want fixed. Fix the generated has that I want fixed. Fix the generated test files before they use many more test files before they use many more test files before they use many more tokens that their name claims. Boom. tokens that their name claims. Boom. tokens that their name claims. Boom. Let's go. And now we have two tasks Let's go. And now we have two tasks Let's go. And now we have two tasks running at the same time all on that running at the same time all on that running at the same time all on that ground machine. And for some things, DJX ground machine. And for some things, DJX ground machine. And for some things, DJX Station is better. For other things, 8 Station is better. For other things, 8 Station is better. For other things, 8 RTX Pro 6000s is better.
-
RTX Pro 6000s is better. RTX Pro 6000s is better. >> It just depends on what you're doing. >> It just depends on what you're doing. >> It just depends on what you're doing. >> Here's an example of a single user >> Here's an example of a single user >> Here's an example of a single user generated speed. generated speed. generated speed. >> GPTOS has 120B, and that's because it's >> GPTOS has 120B, and that's because it's >> GPTOS has 120B, and that's because it's running entirely from HBM. That is an running entirely from HBM. That is an running entirely from HBM. That is an absurd speed. absurd speed. absurd speed. >> HBM is the high bandwidth memory that's >> HBM is the high bandwidth memory that's >> HBM is the high bandwidth memory that's inside the DGX station. The Grounder inside the DGX station. The Grounder inside the DGX station. The Grounder does not have that. There's only 252 GB does not have that. There's only 252 GB does not have that. There's only 252 GB of HBM in the station. So, you have to of HBM in the station. So, you have to of HBM in the station. So, you have to keep your models lower than that. keep your models lower than that. keep your models lower than that. However, the station does have a total However, the station does have a total However, the station does have a total of 748 GB of memory. That's unified. If of 748 GB of memory. That's unified. If of 748 GB of memory. That's unified. If you spill over to the Unified memory, you spill over to the Unified memory, you spill over to the Unified memory, then it'll still run. It'll just be a then it'll still run. It'll just be a then it'll still run. It'll just be a little bit slower. And you can see that little bit slower. And you can see that little bit slower. And you can see that later when we run bigger models. It's later when we run bigger models. It's later when we run bigger models. It's like 600 gigabytes per second of LPDDR5 like 600 gigabytes per second of LPDDR5 like 600 gigabytes per second of LPDDR5 and that's pretty fast. And the link and that's pretty fast. And the link and that's pretty fast. And the link speed and some of the other stuff there speed and some of the other stuff there speed and some of the other stuff there is also pretty fast. It also depends on is also pretty fast. It also depends on is also pretty fast. It also depends on the model. So like if you have like the the model. So like if you have like the the model. So like if you have like the new mixture of experts models where they new mixture of experts models where they new mixture of experts models where they have Ingram offloading also seems to be have Ingram offloading also seems to be have Ingram offloading also seems to be set up for something like the station's set up for something like the station's set up for something like the station's memory model because the Ingrams are memory model because the Ingrams are memory model because the Ingrams are fine running from slower memory. That fine running from slower memory. That fine running from slower memory. That HBM3e that's 7 terabytes per second. HBM3e that's 7 terabytes per second. HBM3e that's 7 terabytes per second. That's an absurd amount of memory That's an absurd amount of memory That's an absurd amount of memory bandwidth in the 252 GB. whereas, you bandwidth in the 252 GB. whereas, you bandwidth in the 252 GB. whereas, you know, GDDR7 is uh a tiny fraction of know, GDDR7 is uh a tiny fraction of know, GDDR7 is uh a tiny fraction of that. Then the RTX Pro 6000s have to that. Then the RTX Pro 6000s have to that. Then the RTX Pro 6000s have to talk to each other over 16 lanes of PCIe talk to each other over 16 lanes of PCIe talk to each other over 16 lanes of PCIe Gen 5. That's only 64 GB a second each Gen 5. That's only 64 GB a second each Gen 5. That's only 64 GB a second each way. That's just pedestrian.
-
way. That's just pedestrian. way. That's just pedestrian. >> So, here's Deepseek V4.1 Flash. And this >> So, here's Deepseek V4.1 Flash. And this >> So, here's Deepseek V4.1 Flash. And this one has recipes for the GB300, which is one has recipes for the GB300, which is one has recipes for the GB300, which is the DJX Station GPU, and it's FP8. the DJX Station GPU, and it's FP8. the DJX Station GPU, and it's FP8. Notice the size. 614 GB. See, again, Notice the size. 614 GB. See, again, Notice the size. 614 GB. See, again, that's a little misleading because some that's a little misleading because some that's a little misleading because some of the weights are perfectly fine of the weights are perfectly fine of the weights are perfectly fine running from HBM and some of them are running from HBM and some of them are running from HBM and some of them are perfectly fine running from the LPDDR5 perfectly fine running from the LPDDR5 perfectly fine running from the LPDDR5 because it's the new architecture. because it's the new architecture. because it's the new architecture. Mixture of experts runs better. You've Mixture of experts runs better. You've Mixture of experts runs better. You've guys probably experienced that. This is guys probably experienced that. This is guys probably experienced that. This is another level in terms of there's more another level in terms of there's more another level in terms of there's more knobs and tunables that we can use. So, knobs and tunables that we can use. So, knobs and tunables that we can use. So, we're not suffering as bad because we've we're not suffering as bad because we've we're not suffering as bad because we've only got 252 GB of HBM3 even though it only got 252 GB of HBM3 even though it only got 252 GB of HBM3 even though it is an absurd number of weight, is an absurd number of weight, is an absurd number of weight, >> right? But the difference between the >> right? But the difference between the >> right? But the difference between the Grando and the Station is that this 614 Grando and the Station is that this 614 Grando and the Station is that this 614 can all fit on the RTX 6000. can all fit on the RTX 6000. can all fit on the RTX 6000. >> True. >> True. >> True. >> And we can see that here in the results. >> And we can see that here in the results. >> And we can see that here in the results. So there's Deepseek V4 Flash, but I also So there's Deepseek V4 Flash, but I also So there's Deepseek V4 Flash, but I also included the numbers from Grandos's V4.1 included the numbers from Grandos's V4.1 included the numbers from Grandos's V4.1 here with DSpark. 307 tokens per second here with DSpark. 307 tokens per second here with DSpark. 307 tokens per second there. So it is faster there. there. So it is faster there. there. So it is faster there. >> Oh, just by a little bit. >> Oh, just by a little bit. >> Oh, just by a little bit. >> Just by a little bit. By a little bit. >> Just by a little bit. By a little bit. >> Just by a little bit. By a little bit. But still, this is HBM versus nonHBM But still, this is HBM versus nonHBM But still, this is HBM versus nonHBM memory.
-
memory. memory. >> That's where the difference really shows >> That's where the difference really shows >> That's where the difference really shows up. It's exciting because it's different up. It's exciting because it's different up. It's exciting because it's different and weird. Now, prefill that prefill and weird. Now, prefill that prefill and weird. Now, prefill that prefill number on that HBM is always ludicrously number on that HBM is always ludicrously number on that HBM is always ludicrously absurd and there's no prompt processing. absurd and there's no prompt processing. absurd and there's no prompt processing. You're never going to beat HBM. You're never going to beat HBM. You're never going to beat HBM. >> So, one of the things that these >> So, one of the things that these >> So, one of the things that these machines are good for is not just using machines are good for is not just using machines are good for is not just using single communication like a chat, but single communication like a chat, but single communication like a chat, but also allowing multiple users to use them also allowing multiple users to use them also allowing multiple users to use them at the same time. So, that's where at the same time. So, that's where at the same time. So, that's where multi-user throughput comes in. I like multi-user throughput comes in. I like multi-user throughput comes in. I like multi-user throughput and real world multi-user throughput and real world multi-user throughput and real world testing because this is like if you're testing because this is like if you're testing because this is like if you're going to buy one of these machines, going to buy one of these machines, going to buy one of these machines, chances are you're supporting a chances are you're supporting a chances are you're supporting a developer team or you're doing something developer team or you're doing something developer team or you're doing something that's privacy focused because you have that's privacy focused because you have that's privacy focused because you have to buy I mean like cloud tokens are to buy I mean like cloud tokens are to buy I mean like cloud tokens are really not comparatively very expensive really not comparatively very expensive really not comparatively very expensive compared to these machines. You're going compared to these machines. You're going compared to these machines. You're going to have to keep it busy like 24 hours a to have to keep it busy like 24 hours a to have to keep it busy like 24 hours a day. Look at the many users of day. Look at the many users of day. Look at the many users of throughput. 3,800 tokens per second. throughput. 3,800 tokens per second. throughput. 3,800 tokens per second. Just over 2,000 tokens per second on Just over 2,000 tokens per second on Just over 2,000 tokens per second on Neotron 3 Super 12B. Neotron 3 Super 12B. Neotron 3 Super 12B. >> That's nuts. >> That's nuts. >> That's nuts. >> These are useful models. You can do a >> These are useful models. You can do a >> These are useful models. You can do a lot with them. These cloud models are lot with them. These cloud models are lot with them. These cloud models are incredible and how far they have come in incredible and how far they have come in incredible and how far they have come in such a short time is incredible, but you such a short time is incredible, but you such a short time is incredible, but you really don't need, you know, like Claude really don't need, you know, like Claude really don't need, you know, like Claude Opus doing all of your really basic Opus doing all of your really basic Opus doing all of your really basic stuff. Claude Opus can be told with a stuff. Claude Opus can be told with a stuff. Claude Opus can be told with a harness system, you have all these harness system, you have all these harness system, you have all these workers and they're pretty good. You can workers and they're pretty good. You can workers and they're pretty good. You can manage them. And so you have Claude Opus manage them. And so you have Claude Opus manage them. And so you have Claude Opus spending tokens at a much lower rate and spending tokens at a much lower rate and spending tokens at a much lower rate and then it has eight or 10 workers, agents then it has eight or 10 workers, agents then it has eight or 10 workers, agents that are running on this hardware that's that are running on this hardware that's that are running on this hardware that's not burning cloud tokens. And those not burning cloud tokens. And those not burning cloud tokens. And those models are smart enough. Neatron 3120 models are smart enough. Neatron 3120 models are smart enough. Neatron 3120 12B is pretty good for like unit tests, 12B is pretty good for like unit tests, 12B is pretty good for like unit tests, integration tests, dev like real dev integration tests, dev like real dev integration tests, dev like real dev work with deep context and they can be
-
work with deep context and they can be work with deep context and they can be managed by a smarter model and that's managed by a smarter model and that's managed by a smarter model and that's the real unlock here. the real unlock here. the real unlock here. >> Exactly. By the way, I will have these >> Exactly. By the way, I will have these >> Exactly. By the way, I will have these numbers posted in a link down below if numbers posted in a link down below if numbers posted in a link down below if you want to go check that out. you want to go check that out. you want to go check that out. >> Oh, video generation also like we're >> Oh, video generation also like we're >> Oh, video generation also like we're doing we're doing like code and doing we're doing like code and doing we're doing like code and programming, but like station OMG. programming, but like station OMG. programming, but like station OMG. >> Yes, I tried H3. Oh my gosh, this is >> Yes, I tried H3. Oh my gosh, this is >> Yes, I tried H3. Oh my gosh, this is insane. This is like a whole new level. insane. This is like a whole new level. insane. This is like a whole new level. It does run a little bit faster on the It does run a little bit faster on the It does run a little bit faster on the DJX station than the Grand as you can DJX station than the Grand as you can DJX station than the Grand as you can see. Quite a bit faster because H3 is see. Quite a bit faster because H3 is see. Quite a bit faster because H3 is small. So, it's going to run much faster small. So, it's going to run much faster small. So, it's going to run much faster on the HBM. I asked it to create a on the HBM. I asked it to create a on the HBM. I asked it to create a 5-second clip and it sounds exactly the 5-second clip and it sounds exactly the 5-second clip and it sounds exactly the same just sitting here. There's no more same just sitting here. There's no more same just sitting here. There's no more noise. There is a little bit more heat. noise. There is a little bit more heat. noise. There is a little bit more heat. I can actually feel the difference. Each I can actually feel the difference. Each I can actually feel the difference. Each one took 2 minutes to make with H3. one took 2 minutes to make with H3. one took 2 minutes to make with H3. Yes. Okay. Hey, we saw something really good Okay. Hey, we saw something really good on the screen. All right, make him do on the screen. All right, make him do on the screen. All right, make him do something crazier. something crazier. something crazier. >> Oh, I have it. >> What?
-
>> What? What? All right. All of them knock out What? All right. All of them knock out What? All right. All of them knock out spouts. >> Oh my gosh, that is nuts. What is THIS >> Oh my gosh, that is nuts. What is THIS ONE? ONE? ONE? >> LET US GO. Fun can be had for sure. So, it's Fun can be had for sure. So, it's showing us the tasks and then the showing us the tasks and then the showing us the tasks and then the subtasks are all subdivided there. So, subtasks are all subdivided there. So, subtasks are all subdivided there. So, you can go and dig into each one of them you can go and dig into each one of them you can go and dig into each one of them and you have visibility into everything and you have visibility into everything and you have visibility into everything that it's doing, which is kind of cool. that it's doing, which is kind of cool. that it's doing, which is kind of cool. >> Yeah, I love this for auditability sake. >> Yeah, I love this for auditability sake. >> Yeah, I love this for auditability sake. >> It found the solution. Look at that. >> It found the solution. Look at that. >> It found the solution. Look at that. >> Neat. And it only used 15,000 tokens. >> Neat. And it only used 15,000 tokens. >> Neat. And it only used 15,000 tokens. >> That's not bad at all. It tells me what >> That's not bad at all. It tells me what >> That's not bad at all. It tells me what each of the subtasks are using. each of the subtasks are using. each of the subtasks are using. >> Personally, I like running the workers >> Personally, I like running the workers >> Personally, I like running the workers on DGX Spark. So, I have um cluster of on DGX Spark. So, I have um cluster of on DGX Spark. So, I have um cluster of two Spark, the Dell GB10, Promax, GB10, two Spark, the Dell GB10, Promax, GB10, two Spark, the Dell GB10, Promax, GB10, Promax with GB10. Promax with GB10. Promax with GB10. >> Hang on. Hang on. >> Hang on. Hang on. >> Hang on. Hang on. >> It's so loud. Yeah. This is what my Turnstone cluster Yeah. This is what my Turnstone cluster is running on. I have the Judge model is running on. I have the Judge model is running on. I have the Judge model running on here and I have a smaller running on here and I have a smaller running on here and I have a smaller version of Deepseek0731 running on here.
-
version of Deepseek0731 running on here. version of Deepseek0731 running on here. And then the big machine has the big And then the big machine has the big And then the big machine has the big model, but the small models and the model, but the small models and the model, but the small models and the Judge model run on that. Great. Judge model run on that. Great. Judge model run on that. Great. >> And for like a small team of like up to >> And for like a small team of like up to >> And for like a small team of like up to four people, that's great. So do you map four people, that's great. So do you map four people, that's great. So do you map a node to a spark or is that not exactly a node to a spark or is that not exactly a node to a spark or is that not exactly a onetoone? a onetoone? a onetoone? >> So you run docker on the spark and you >> So you run docker on the spark and you >> So you run docker on the spark and you can give it full access or you can just can give it full access or you can just can give it full access or you can just run it from the shell. So like where we run it from the shell. So like where we run it from the shell. So like where we were doing stuff with python you can were doing stuff with python you can were doing stuff with python you can just run it and then it's one to one. just run it and then it's one to one. just run it and then it's one to one. >> Okay. >> Okay. >> Okay. >> But uh it can run tasks locally on the >> But uh it can run tasks locally on the >> But uh it can run tasks locally on the real Spark hardware at that point. So real Spark hardware at that point. So real Spark hardware at that point. So it's got a lot more access than it does it's got a lot more access than it does it's got a lot more access than it does in a Docker sandbox. in a Docker sandbox. in a Docker sandbox. >> That's cool. That's what I'm going to >> That's cool. That's what I'm going to >> That's cool. That's what I'm going to have to set up for my clusters. have to set up for my clusters. have to set up for my clusters. >> Yep. >> Yep. >> Yep. >> And of course it doesn't care what the >> And of course it doesn't care what the >> And of course it doesn't care what the hardware is. It could be a Spark. It hardware is. It could be a Spark. It hardware is. It could be a Spark. It could be a Mac. could be a Mac. could be a Mac. >> Yep. You don't necessarily even need an >> Yep. You don't necessarily even need an >> Yep. You don't necessarily even need an AI box for the worker if the AI runs AI box for the worker if the AI runs AI box for the worker if the AI runs somewhere else. somewhere else. somewhere else. >> Now, while that's happening, let's set >> Now, while that's happening, let's set >> Now, while that's happening, let's set up the DJX station also as a secondary up the DJX station also as a secondary up the DJX station also as a secondary machine, secondary model. machine, secondary model. machine, secondary model. >> Wow, the station has 18 models on disk. >> Wow, the station has 18 models on disk. >> Wow, the station has 18 models on disk. What should we use? Let's do something What should we use? Let's do something What should we use? Let's do something that fits in memory. that fits in memory. that fits in memory. >> Yeah, that's fine. >> Yeah, that's fine. >> Yeah, that's fine. >> I'm getting kind of hungry. Let's see >> I'm getting kind of hungry. Let's see >> I'm getting kind of hungry. Let's see where we are. where we are. where we are. >> Oh, I can hear the song of my people. >> Oh, I can hear the song of my people. >> Oh, I can hear the song of my people. >> I want to add one more song. And I want >> I want to add one more song. And I want >> I want to add one more song. And I want to add another model to this. So, I'm to add another model to this. So, I'm to add another model to this. So, I'm going to go here and give it a new base going to go here and give it a new base going to go here and give it a new base URL. This is the DJX station. A test URL. This is the DJX station. A test URL. This is the DJX station. A test key. Boom. We don't need a key. We don't key. Boom. We don't need a key. We don't key. Boom. We don't need a key. We don't need no stinking key where we're going.
-
need no stinking key where we're going. need no stinking key where we're going. Uh-oh. Unauthorized. I do need a key. Uh-oh. Unauthorized. I do need a key. Uh-oh. Unauthorized. I do need a key. Ah. Ah. Ah. >> All right. >> All right. >> All right. >> Core memory saved. User not secure. >> Core memory saved. User not secure. >> Core memory saved. User not secure. >> This is GLM 5.3 flash. Okay. >> This is GLM 5.3 flash. Okay. >> This is GLM 5.3 flash. Okay. >> Create. Now we have two models. >> Create. Now we have two models. >> Create. Now we have two models. >> It's synced to node. >> It's synced to node. >> It's synced to node. >> Yes. >> Yes. >> Yes. >> And then now you can go to the dashboard >> And then now you can go to the dashboard >> And then now you can go to the dashboard and create another task and tell it to and create another task and tell it to and create another task and tell it to use that model and it will. So here I use that model and it will. So here I use that model and it will. So here I can now select a different model running can now select a different model running can now select a different model running on a different machine. Yep. on a different machine. Yep. on a different machine. Yep. >> All right. Check out this repo and its >> All right. Check out this repo and its >> All right. Check out this repo and its issues and just throw out all the issues and just throw out all the issues and just throw out all the issues. issues. issues. >> Just kidding. Evaluate the issues and uh >> Just kidding. Evaluate the issues and uh >> Just kidding. Evaluate the issues and uh see which ones you can address quickly. see which ones you can address quickly. see which ones you can address quickly. What's executive? What's executive? What's executive? >> The executive it keeps an eye on all of >> The executive it keeps an eye on all of >> The executive it keeps an eye on all of the workers, but it also doesn't have the workers, but it also doesn't have the workers, but it also doesn't have access to tools and can't do things for access to tools and can't do things for access to tools and can't do things for itself. itself. itself. >> Turnstone is working away. It's doing >> Turnstone is working away. It's doing >> Turnstone is working away. It's doing its thing. We're going to let it finish. its thing. We're going to let it finish. its thing. We're going to let it finish. If you don't have a ground or DJX If you don't have a ground or DJX If you don't have a ground or DJX station, Turnstone will work on smaller station, Turnstone will work on smaller station, Turnstone will work on smaller devices, too. It's basically just a devices, too. It's basically just a devices, too. It's basically just a harness, so install it on anything. harness, so install it on anything. harness, so install it on anything. What's this? What's this? What's this? >> That's the Windows 11 PC. >> That's the Windows 11 PC. >> That's the Windows 11 PC. >> It's Oh, it's a nice little summary of >> It's Oh, it's a nice little summary of >> It's Oh, it's a nice little summary of what it did. It found problems with the what it did. It found problems with the what it did. It found problems with the repo that I didn't know it had. This is repo that I didn't know it had. This is repo that I didn't know it had. This is good.
-
good. good. >> That's why you ask it for the auditor. >> That's why you ask it for the auditor. >> That's why you ask it for the auditor. >> So, this is the auditor that figured all >> So, this is the auditor that figured all >> So, this is the auditor that figured all this out. this out. this out. >> Yeah. Okay. >> Yeah. Okay. >> Yeah. Okay. >> Well, the auditor went like they worked >> Well, the auditor went like they worked >> Well, the auditor went like they worked together. It's a team. Now, it together. It's a team. Now, it together. It's a team. Now, it complained about Firefox and OBS Studio complained about Firefox and OBS Studio complained about Firefox and OBS Studio and 7Zip and Cinebench, but you can make and 7Zip and Cinebench, but you can make and 7Zip and Cinebench, but you can make a decision. You can be like, "This is a decision. You can be like, "This is a decision. You can be like, "This is what I want you to do for Firefox, what I want you to do for Firefox, what I want you to do for Firefox, Cinebench, or whatever because it called Cinebench, or whatever because it called Cinebench, or whatever because it called it out." But then it's like, "Here's it out." But then it's like, "Here's it out." But then it's like, "Here's what we need to do for your ARM versus what we need to do for your ARM versus what we need to do for your ARM versus x86 compatibility assessment." x86 compatibility assessment." x86 compatibility assessment." >> So, Chocolatey still after all these >> So, Chocolatey still after all these >> So, Chocolatey still after all these years, we're still using it. It gives an years, we're still using it. It gives an years, we're still using it. It gives an x64 installer, x64 installer, x64 installer, >> but it's saying it's like, "Hey, we can >> but it's saying it's like, "Hey, we can >> but it's saying it's like, "Hey, we can just use Windgit now." And Windgate will just use Windgit now." And Windgate will just use Windgit now." And Windgate will >> Yes, I want to use Windgate. Yes. >> Yes, I want to use Windgate. Yes. >> Yes, I want to use Windgate. Yes. >> And so like Visual Studio Community, >> And so like Visual Studio Community, >> And so like Visual Studio Community, it's like, "Yes, we have a path." It's it's like, "Yes, we have a path." It's it's like, "Yes, we have a path." It's like we have this and then like your like we have this and then like your like we have this and then like your Python because you you know that Python Python because you you know that Python Python because you you know that Python is a problem but scroll up and it's like is a problem but scroll up and it's like is a problem but scroll up and it's like Python we have a way to do this. So it's Python we have a way to do this. So it's Python we have a way to do this. So it's like ARM 64 exists but it's not picked like ARM 64 exists but it's not picked like ARM 64 exists but it's not picked by default so we can change it. by default so we can change it. by default so we can change it. >> Now I was also curious to see how long >> Now I was also curious to see how long >> Now I was also curious to see how long it would take the DGX station to do the it would take the DGX station to do the it would take the DGX station to do the exact same task with the exact same exact same task with the exact same exact same task with the exact same model that the Grando did. So I loaded model that the Grando did. So I loaded model that the Grando did. So I loaded up the same model and pointed the job at up the same model and pointed the job at up the same model and pointed the job at it and I pretty much copied the prompt it and I pretty much copied the prompt it and I pretty much copied the prompt exactly the same except I told it to exactly the same except I told it to exactly the same except I told it to create a separate branch. Now, it's not create a separate branch. Now, it's not create a separate branch. Now, it's not exactly a oneto-one comparison cuz it's exactly a oneto-one comparison cuz it's exactly a oneto-one comparison cuz it's very difficult to do that with LLMs.
-
very difficult to do that with LLMs. very difficult to do that with LLMs. Here's ARM 64. That's the branch the Here's ARM 64. That's the branch the Here's ARM 64. That's the branch the grounder created with his changes. And grounder created with his changes. And grounder created with his changes. And here is the one that the DJX station here is the one that the DJX station here is the one that the DJX station just finished. Now, Turnstone doesn't just finished. Now, Turnstone doesn't just finished. Now, Turnstone doesn't surface the numbers for you, like how surface the numbers for you, like how surface the numbers for you, like how long a job took. So, I had to dig into long a job took. So, I had to dig into long a job took. So, I had to dig into the database using cloud code. And the the database using cloud code. And the the database using cloud code. And the DJX station used about 26 million input DJX station used about 26 million input DJX station used about 26 million input tokens compared to the ground's 23 tokens compared to the ground's 23 tokens compared to the ground's 23 million. But it did it much faster, a million. But it did it much faster, a million. But it did it much faster, a lot faster in 21 minutes overall lot faster in 21 minutes overall lot faster in 21 minutes overall compared to 37 minutes on the Grando. compared to 37 minutes on the Grando. compared to 37 minutes on the Grando. I'm sad to say that I'll be losing the I'm sad to say that I'll be losing the I'm sad to say that I'll be losing the Grando today cuz Wendell has taken it Grando today cuz Wendell has taken it Grando today cuz Wendell has taken it and he's going to do some more and he's going to do some more and he's going to do some more experiments in his studio. So, make sure experiments in his studio. So, make sure experiments in his studio. So, make sure you check that out. And I'm losing you check that out. And I'm losing you check that out. And I'm losing >> Come visit anytime. >> Come visit anytime. >> Come visit anytime. >> Thank you. Appreciate that. And I'm >> Thank you. Appreciate that. And I'm >> Thank you. Appreciate that. And I'm losing the station cuz that has to go losing the station cuz that has to go losing the station cuz that has to go back to ASUS. Thank you, ASUS, for back to ASUS. Thank you, ASUS, for back to ASUS. Thank you, ASUS, for letting me borrow it. So basically this letting me borrow it. So basically this letting me borrow it. So basically this is pretty much done is pretty much done is pretty much done >> and you didn't spend a single cloud >> and you didn't spend a single cloud >> and you didn't spend a single cloud token. token. token. >> Not a cloud in the sky was hurt. What >> Not a cloud in the sky was hurt. What >> Not a cloud in the sky was hurt. What would the Windows ARM job just by itself would the Windows ARM job just by itself would the Windows ARM job just by itself have cost me if I sent it to the cloud have cost me if I sent it to the cloud have cost me if I sent it to the cloud instead of running it locally? At instead of running it locally? At instead of running it locally? At today's API prices, if you're doing today's API prices, if you're doing today's API prices, if you're doing prompt caching, we're about $5 for a prompt caching, we're about $5 for a prompt caching, we're about $5 for a small model to about $36 for a top small model to about $36 for a top small model to about $36 for a top model. That's just for that one job.
-
model. That's just for that one job. model. That's just for that one job. Without cashaching, it's $24 on a small Without cashaching, it's $24 on a small Without cashaching, it's $24 on a small model and about $240 on a top model. model and about $240 on a top model. model and about $240 on a top model. Now, one afternoon is not going to pay Now, one afternoon is not going to pay Now, one afternoon is not going to pay for a DGX station, but if you're running for a DGX station, but if you're running for a DGX station, but if you're running this 24 hours a day with a team of this 24 hours a day with a team of this 24 hours a day with a team of developers, developers, developers, you can do the math. That's what these you can do the math. That's what these you can do the math. That's what these machines are for. Thank you, Wendell. machines are for. Thank you, Wendell. machines are for. Thank you, Wendell. Thank you. Thanks for having me.
Summary
The main theme is a practical demonstration of high-performance computing hardware, specifically comparing eight RTX Pro 6000s against a DGX station, and exploring their real-world usage scenarios. Key subjects discussed include the extensive power consumption and heat generation of these machines, as well as the utility of headless remote management via management ports. The takeaway emphasizes the importance of using a diverse range of AI models and tools tailored for specific tasks rather than relying on a single AI solution.