{"text": " So in this video, I want to continue the series I started on refurbished data center GPUs. I've been working on this video for the past two months and I want to start with an important premise. I know some people will feel compelled to use the comment section to vent and compare the hardware that we discuss here with stuff that's two to five times more expensive. First of all, this video comes out of a collaboration with Bargain Hardware. This is not a sponsored video. If you don't know them, Bargain Hardware are a major reseller of refurbished hardware. I am not getting paid for this. This collaboration simply gives me access to some hardware so I can learn an experiment, which is the main objective of this channel. Now, of course, I know that people that want to get into local AI are after alternatives to more expensive hardware. For reference, in August last year, one could get a 128GB Strixello machine for just over $2000. Now, we are looking at twice as much for the same machine. The landscape has totally changed. This is also why I am experimenting with refurbished data center hardware to see what kind of performance we can achieve, what the cost is, and most importantly, the trade-offs you need to be aware of if you decide to go this way. On top of all of this, Bargain Hardware is offering a 10% discount on the GPUs listed in the description of this video. Just use the code DONATO10 at checkout. This is not an affiliate link and I get zero if you buy or don't buy. We could have done a 5% split where you get 5% discount and I get 5%. But I am trying to see what I can do to make some hardware more affordable for my viewers, since it looks like that the memory and GPU situation that we have now will last for a while. So I'd rather get the full 10% discount to you. That's the only thing that's in my power. Now, the first consideration is that when you want to build a multi-GPU system, it's not just the GPUs that cost money. The issue is that unless you already have a rig to place your cards in, you also have to build a system to host them. If you watched my Dual Radio 9700 video, you'll have an idea of what that looks like. But that setup with 64GB of RAM would cost around $4,000 to $5,000 with current prices, maybe even more, I really cannot keep up myself. And if you want to scale up to four dual-slot GPUs, you need a motherboard that physically can accommodate them, a GPU with enough PCIe lanes to avoid lane starvation, a high capacity power supply, and a chassis with proper cooling, let alone system RAM. If you try to build that base system from scratch using currently available hardware, it can easily cost you around $4,000 before you even buy the GPUs. That is where refurbished data center servers start making sense in 2026. They offer great value by giving you a base system for a fraction of the cost. Now, Bargain Hardware is a simple online builder where you can select all the components for the system. For a quad GPU configuration, one of the cheapest options is to get a super micro server. I'll get this 12th generation one. The motherboard in these servers typically hosts two Intel Xeon CPUs. So we can now configure these. You can see there is a base price, of course. So the first thing we want to do is to select the CPUs. And here I'm going for two of these Xeon 18 core CPUs that you can see here. They are incredibly cheap. Then, of course, the heat sinks are already included. And for the RAM, as you can see here, there are a lot of options. Actually, Bargain Hardware stocks right now, a lot of DDR4 at a very, very good price. Here I'm getting four of these 16 GB sticks. Now we want to select the storage. And for the storage, I went with SATA. And I selected two of these 480 GB SSDs, which gives me almost one terabyte, which is what you want for LLMs, at least one terabyte, because you're going to be downloading a lot of weights. And moving on, now we select the GPU accelerators. Now for this build, I have selected four P100s, 16 GB each. So we get 64 GB. But in this video and in the benchmarks, I've also looked at the V100s and also the AMD MI25. And again, each of these options ultimately gives you four GPUs with 64 GB of RAM. But of course, the V100 is much more expensive because it's a more modern architecture. But check the website because more GPUs are coming. And probably by the time this video is uploaded, you will see some more options. I know that there are some other GPUs coming. RTX 3080, 3090, Quadro RTX 5000, Quadro M6000. But just check the stock because they are updating it and they will also make some A100s and RTX 8000 available. So there is quite a lot that's coming to their stock. But for this build, I will keep the P100s in there and we can add to the cart. Oh, actually, we forgot to select the Super Macro NVMe Enablement Kit. We add to the cart and now we can go and check out. And I just want to show you the different options and costs for delivery. That's actually at least in the UK for domestic delivery. When you look at this, it brings us to a total of just over 2000 pounds. But this is not including the 10% discount on the GPUs that you can get with the code Donato 10. Now, I thought it'd be interesting to actually go in person to their warehouse, see how an order like this is put together and generally get a look at the process they use to test and refurbish hardware. And a thank you to Jack Moyers from the Bargain Hardware team, who was kind enough to give me a tour. The first process at Bargain Hardware is inbound. So pallets of servers, workstations, components, they arrive on trucks. So we buy equipment from all kinds of places all around Europe and some more worldwide. And then before the servers and workstations and any equipment gets processed, it gets stored in our rack in here. So the first part of the refurbishment process is cleaning. So this is one of our cleaning rooms. So in here, the first part of the process is the servers, they get blown out with our air compressor machine. You can see here, Michael's doing some reskinning on a server. So certain generations of servers get reskinned just so we have a nicer finish. So it just gets rid of all these scuffs and scrapes off the lids. So sometimes we have orders which are just a batch of drives or it could be CPUs or GPUs. So not necessarily every component comes out of a machine. Sometimes we buy them in batch, but still every single component needs testing. So then we come to component testing here. So when a machine has been tested, the next thing that we do, because we stock our servers and workstations as configured to order, the next thing we need to do is take the components out and that's called disassembly. So this is the next section. So here we've got Harry who's working on a workstation. So he'll be taking out the CPU, the RAM, any GPUs and other cards and components so that when we stock it and we offer it as configured to order on our website, it's basically a chassis with a PSU and other base components. As we were walking around, my attention was immediately captured by a stock of GPUs that they got in and that they were testing and getting ready to refurbish. So then we come to picking. So if someone orders on our website, they've configured a server. We've stocked it as a chassis and components. So then our picking team will they'll receive the order. They'll have a list of all the components that need to be picked. And then it's just the case of them assembling it all onto a trolley ready for the next part of the process. It's actually very well organized. Yeah. Wherever I look, it's tidy. Yeah. Like look around me. There is nothing left to have. Yeah. This is where we keep all the RAM and CPUs. So there's a lot of value behind a closed door here. So we just need to keep it safe. Given the current situation with memory, it's refreshing to see how much DDR4 memory they've been able to stock at a pretty competitive price. So when a member of the picking team has picked an order, it goes onto a trolley and then the trolley is moved through to assembly so that that particular order can be assembled. So this is the assembly area. So you can see a member of staff is assembling an order that's been configured online. And then when the machine is assembled, it then needs to be tested. It's clear that they put a lot of effort in testing the builds before sending them out. And of course, they give people warranty, but also the care they put in packaging to make sure that stuff gets shipped in the best possible way. We've invested a lot in the cardboard boxes we use, in the foam inserts, making sure that every server is catered for and that it's really secure inside. Yeah. So once the order has been packed, then we've got our outbound department here. So we've got our outbound department here. At the end of the warehouse tour, I spent some time with Toby Sheriff, who's the lead engineer that actually put together the server build you see me configure at the beginning of this video. What we've got here is a Supermicro DGQ in CSE 118 chassis. It's got space for four double wide full height cards. It takes the scalable CPUs, Intel first and second generation. So currently we've got a gold 6150 or two gold 6150s in there. So I believe they're 18 core CPUs, but they're also threaded. So obviously you've got twice as many of them than that. We've also got some DDR4 registered DIMMs. So it can take up to six per CPU. Currently we've got 64 gigs in there in 16 gig DIMMs. That's running at 2666 megahertz mega transfers. And that's just a limitation of the CPU. So there's not really anything you can do about that. If you go to a second generation CPU, you'd get slightly faster at 2933. So it's not fully utilising the speed, but it shouldn't really be a problem. We've also got on board 10 gig NICs. So that's integrated into the motherboard. We've also got two spare PCI slots at the rear. If you wanted to add a network card, maybe a PCIe storage drive or something. Half of these three slots are CPU2 dependent. So you could, if you felt so inclined, just put two cards in it initially with one CPU. And expand in the future if that was, you know, that was what you wanted to do. Keep the cost down initially. Storage wise, at the front, you've got two NVMe ports built into a backplane. So that would do U.2 drives. So two and a quarter inch NVMe drives. And then we've also got two SATA slots just behind that backplane internally. So because these are data center graphics cards, they're not designed like a consumer card with a fan. As you said, they're passively cooled. So that has got to travel through the the the fin stack as it, you know, travels through the server. So you've got all these fans across the front here. I won't lie. It is loud. But it's got to be to be able to push enough air through to keep all these cards and your CPU and even your power supplies or network cards at the back. Cool. So right now we've got four P100s in here, but obviously we can swap any card that would fit in there. In the video, I'm going to show you the performance with the, I think I have it here, with one of the MI-50s, so MI-25 from AMD. And then maybe we are also going to try, is this the... It's the V100, that one. The V100, which is going to perform so much better. And in fact, we'll probably do some benchmarks with that. So Toby is now closing the chassis, so we can power this on. And mostly I want to give you an idea of the noise levels. When you first power it on, this is very noisy. And then it becomes a little bit quieter, but obviously it is still quite noisy. And we need to be aware of that. So I wanted to also show you quickly the specs of these three GPUs side by side, along with more modern GPUs like the R9700 AI Pro, that I'm sure you've already seen on my channel, and also the RTX 5090. Now, during the comparison, keep in mind the considerable price difference. I have put it here, and obviously for the 5090 and the R9700, I've had to estimate some brackets, because the price is very variable. But for the other GPUs that you can get on bargain hardware, I have put the price that you would get with the 10% discount. One of the most important things to consider is memory bandwidth. During inference, generating each token requires reading a lot of data, such as model weights, back and forth from memory. So the bandwidth here directly limits how many tokens per second you can generate and process. So the V100 stands out at 900 gigabytes per second, which is actually higher than the more modern R9700, which has a memory bandwidth of 640 gigabytes per second. And the MI25 is significantly lower at 484 gigabytes per second. And that shows in the benchmarks. The other important factor is the data format support. Modern GPU architectures support formats like BF16 and FP8 in a native way. And these are the standard precisions used by current LLMs. None of these three older cards support these formats natively. They do support FP16, which is also a 16-bit format and uses the same amount of memory, but FP16 cannot represent as wide a range of values as BF16. When an inference framework converts BF16 model weights to FP16, some values can overflow or lose accuracy. And depending on the model, this can downgrade output quality. One last difference to keep in mind is that the V100 is more expensive because it has Tensor Cores, which are essentially hardware parts optimized for matrix operations used in machine learning. And you'd find such support also on all the modern GPUs. The P100 and MI25 do not have any equivalent, and this will be clearly reflected in the benchmarks. So very quickly, I want to show you how to set up this server. Right now, I've got the four V100s installed. So this is the repository I am going to use with the V100 AI toolboxes. As usual, with all of my repositories, you are going to need to create a toolbox. And this is essentially a Docker container. I'm not going to repeat myself. I explain about toolboxes, Podman and Docker containers in a lot of the other videos. You can choose between the CUDA backend and the Vulkan backend. And CUDA is almost always going to work better for Nvidia cards. So we're going to take that one and we're going to create the toolbox, which is essentially going to connect to Docker Hub and pull this toolbox that I have pre-built with Lama CPP. So let us do that. And actually, I have already created this toolbox, so I can enter it. Lama V100 CUDA. Now, once you enter the toolbox, you can run Lama CLI, list devices. And this essentially confirms that Lama via the CUDA backend can see the four V100s, each of them with 16 gigabytes of VRAM. So now, if you want to run a model, the easiest way to do that is to download the model weights. I've already downloaded some of these in GGUF format from Hugging Face. And let's run, for example, one of the best models that we can have today for local agentic workflows, which is QAN 3.6, 27 billion parameters. And I think I have these in Q4 KXL quantization. So I'm just going to say Lama server. I'm going to pass the model. Actually, I also have it in Q8 quantization, for a matter of fact. And we are going to enable flash attention, and we can pass a context size. All right. So the server is up and running on port 8080. And now I need to forward that port to my actual laptop. So I can SSH into the box. I've called it BH for bargain hardware. And I'm going to say that I want to forward that port to 8081 on my host, because 8080 is already used by something else. And there we go. This is Lama CPP web UI. So I can type a prompt, such as write a CUDA kernel to multiply to As you can see now, it's thinking about it. When models spend a lot of time thinking, and you can see it's going at around 31 tokens per second. So for being a 27 billion parameter dense model, this is really good. But in a bit, we'll take a look at the proper benchmarks and what happens when the context size grows, and obviously what happens with other models and quantizations. But I just thought I'd show you how to get one of these up and running. Now, Lama CPP is my recommendation, especially for these older GPUs, not just the V100, but also the P100 and the MI25. And I've got toolboxes and links to all of those. And the reason for that is that Lama CPP has a much more uniform ecosystem, and it works pretty much everywhere. If your hardware is supported, you can run any model on it. And it's got pretty much any model and any quantization. And quantizations are going to be really important because, of course, here, you don't have that much memory available. I mean, in this case, you've got 64 gigabytes, which is quite a lot. But again, even if you want to run the QEM model that we just saw 27 billion parameters is a lot, then you are going to need a quantization. However, I know that a lot of people really, really like VLLM. So I've also created a VLLM toolbox that you can get up and running like this. So here, I've already created the toolbox so I can simply enter it. And you will see that, obviously, on this server, I have also the VLLM toolbox for the P100 and the one for the AMD MI25. But obviously, in this case, we enter this particular one. And what I do in my toolboxes for VLLM, I always give you a start VLLM script with a list of models that I have tested. So at least you have a starting point. I use LAMA 3.1 just as a benchmark. But then if you want to run some proper models here, you can see the QEM 3.6 family, the 27 billion parameter. Now, this is in GPTQ 4-bit quantization. You cannot run, or at least I haven't been able to find a way on VLLM to run AWQ quants, which are a little bit better activation aware. But again, this is why I tell people that LAMA CPP is 90% of the times better for most people, especially on older hardware. But anyway, I've got four GPUs. So Tensor Parallelism 4, I can set concurrent requests, size of the context GPU utilization. I don't remember why I set it this low. You should probably set it a little bit higher. And then you can launch the server. And this should show you exactly how we are running that particular model. And it will take a little bit. And then VLLM will come up and you'll be able to use the model. So here we can see the model loading. It's taking quite a bit of time. And that's basically one of the caveats with PCI 3.0. Especially model loading is going to be fairly slow. And you can see here, it's also looking for different attention backhand that it can use. And ultimately, I think it's falling back to PyTorch, SDPA attention, which is okay, but it's not as good as some of the things you can get on modern hardware, of course. But again, this will work and you'll be able to run some models, even with VLLM, if you use these tool boxes. So here we go. This is now up and running. And you can start using it, for example, with a coding agent. But again, I would not recommend to use perhaps VLLM. Just stick to Lama CPP and you're going to get probably the best performance here. Let's now take a look at the benchmarks, which are arguably one of the most important factor in deciding whether or not some of these cards might be good for you. Here, I've got the individual repositories for all the cards that I tested. I should have the MI25 here as well. And I have put all of the benchmark results in the readme. So you can probably scroll and find them for Lama CPP and VLLM, token generation and prompt processing. So this is the V100 and obviously the P100. But obviously for this video to make things a little bit easier to compare, I just put everything together. You can see the comparison here of the three cards on different models. Now, these are modern LLMs that I recommend you run on these cards, mostly for agentic workflows and coding. The QAN 3.5 and 3.6 families are right now some of the best you can run with 32 to 64 gigabytes of RAM. And you can see obviously that the V100, the blue one here, it's the best performer at prompt processing and token generation. Even when the context goes to 32,000 tokens per second. And that's because this is just a more modern architecture and it's got tensor cores, which essentially allow to do matrix multiplication much better. And that's the core operation in LLMs. Just to give you an idea for a model like 27 billion parameters, you get on the V100, 852 tokens per second in prompt processing with the Q4 quantization, which is a very good quantization. And that's a very good prompt processing speed. And you get around 34 tokens per second. And even when you scale up the context, you see that you're still getting around 622 tokens per second on this 27 billion parameter model, which is a dense model. So these are the hardest model to run. And the performance is really good for these. And the tokens per second that you get in token generation is around 28 tokens per second. Again, this would be perfectly useful if you were using this model, let's say, in the PI coding agent. Actually, check out the video I've done on coding agents and it's going to give you an idea of the type of performance you can get. So you will also see that the MI25 is missing from some of these. It's just because it's the first card that I tested. And back then, I didn't include all the models. But the ones that I included, you can see that on LLAMA CPP, it does perform close to the P100, but just below it. So be aware of that. The MI25 is the one that's going to give you the least performance. So you can see that the other ones that you can see in the LLAMA CPP, but this is what you get on LLAMA CPP. You can also run VLLM, but with some caveats, VLLM is much more sensible to the different GPU architectures. You require specific kernels for specific architectures, and some of them are just not available for a lot of these cards. So you are not really able to pick and choose like you do with LLAMA CPP and run anything you want. But you can see that the MI25, when there are kernels available, actually performs better than the P100 on VLLM. So that's something to keep in mind. Again, I do not recommend using VLLM with these cards. A lot of models you just cannot easily run. But on the V100, at least, you can run some quantization of the QAN 3.6 27 billion parameter model. This is the GPT-Q 4-bit quantization. And you can see this throughput that you get. And you can also run the QAN 3.5 9 billion parameter model. But again, I do not recommend running any of these. I did try these, and I have all the toolboxes if you want to experiment, but probably stick to LLAMA CPP for these older GPUs. So we've seen the benchmarks. Now let's discuss the caveats and trade-offs you need to be aware of. First, the GPU architectures. The Pascal-based P100 and the Vega-based MI25 do not have hardware paths optimized for modern machine learning. The V100 does have Tensor Cores, but it is still a legacy architecture compared to modern GPUs. Because of this, although these cards will be able to run modern LLMs, they just cannot match the throughput that a modern GPU gives you. And that's reflected in the price you pay. Second, these setups are noisy. Because these passive cards require high airflow fans to stay cool, the system is loud. You will not want this server under your desk. This more likely belongs in a garage, a basement, or a dedicated room. Third, the system bus. These older servers use PCIe Gen 3 slots. A PCIe Gen 3 slot has a theoretical bandwidth of around 16GB per second, whereas a modern PCIe Gen 5 slot provides 64GB, four times the speed. For single GPU setups, the difference is not that noticeable, only adding a few seconds when you load the model weights from storage into the GPU memory. But if you run multi-GPU workloads, the GPUs must constantly exchange information and synchronize their activity. So, over a Gen 3 bus, this exchange of information will be slower and it will limit the overall throughput and token generation performance. The same is true for the memory. The service runs on DDR4-ACC memory, which usually operates at around 2400 to 3200 megatransfer per second. In comparison, DDR5 starts at 4800 megatransfer and goes much higher, effectively doubling the bandwidth. This does not affect you much if you choose models that fully fit inside the GPU memory. But if you want to do some offloading to the CPU and to the system memory, this will be a bottleneck. As I was editing this video, I realized that I was talking about PCIe 3.0 and DDR4. And I chose those for the workstations that we built here because I wanted to show you the cheapest available options to get a viable setup. But if you go on bargain hardware, for example, you can find servers and workstations with DDR5 and PCIe 4 and 5. Just to give you an example, these Dell Precision 7960, well, of course, more expensive, but these will come with DDR5 memory. And it will also come with PCIe 5.0. So these are the main trade-offs you make, which allow you to keep the total cost at this price point. Even with these caveats, the value is clear. If you have a budget of 2,000 to 3,000 pounds and you need 64 gigabytes of VRAM to run certain models, a refurbished enterprise server like the one I showed you is one of the real options available to you today. You get a pre-built base system that can host four cards for a fraction of the cost of a DIY workstation. You can even start with cheap cards like the MI25 and the P100 and then upgrade to faster GPUs when you are ready for that switch. As I mentioned at the start of the video, bargain hardware is offering a 10% discount on all of the GPUs that they offer at the link in the description. So if you want to use that discount, just type Donato 10 at checkout. Again, this is not an affiliate link and I get zero commission. My goal is just to see what I can do to make hardware a bit more affordable for anyone looking to get started with local inference. So in the description of this video, you will also find all of the GitHub links to the toolboxes and configurations that I put together for the GPUs that I tested in this video. So to wrap up, as usual, remember this channel is a hobby project and a significant amount of time goes into making these videos. If you find the work useful and want to support my research and me maintaining all of the different toolboxes and containers, the link to support the channel is in the description. You can use Buy Me A Coffee to make a donation. Thanks for watching and I'll see you in the next one. Thank you.", "segments": [{"id": 0, "seek": 0, "start": 0.4, "end": 7.2, "text": " So in this video, I want to continue the series I started on refurbished data center GPUs.", "tokens": [50385, 407, 294, 341, 960, 11, 286, 528, 281, 2354, 264, 2638, 286, 1409, 322, 1895, 16659, 4729, 1412, 3056, 18407, 82, 13, 50725], "temperature": 0, "avg_logprob": -0.14777266842195358, "compression_ratio": 1.6079295154185023, "no_speech_prob": 1.4838562236579866e-12}, {"id": 1, "seek": 0, "start": 7.2, "end": 13.6, "text": " I've been working on this video for the past two months and I want to start with an important premise.", "tokens": [50725, 286, 600, 668, 1364, 322, 341, 960, 337, 264, 1791, 732, 2493, 293, 286, 528, 281, 722, 365, 364, 1021, 22045, 13, 51045], "temperature": 0, "avg_logprob": -0.14777266842195358, "compression_ratio": 1.6079295154185023, "no_speech_prob": 1.4838562236579866e-12}, {"id": 2, "seek": 0, "start": 13.6, "end": 21.76, "text": " I know some people will feel compelled to use the comment section to vent and compare the hardware that we discuss here", "tokens": [51045, 286, 458, 512, 561, 486, 841, 40021, 281, 764, 264, 2871, 3541, 281, 6931, 293, 6794, 264, 8837, 300, 321, 2248, 510, 51453], "temperature": 0, "avg_logprob": -0.14777266842195358, "compression_ratio": 1.6079295154185023, "no_speech_prob": 1.4838562236579866e-12}, {"id": 3, "seek": 0, "start": 21.76, "end": 26.400000000000002, "text": " with stuff that's two to five times more expensive.", "tokens": [51453, 365, 1507, 300, 311, 732, 281, 1732, 1413, 544, 5124, 13, 51685], "temperature": 0, "avg_logprob": -0.14777266842195358, "compression_ratio": 1.6079295154185023, "no_speech_prob": 1.4838562236579866e-12}, {"id": 4, "seek": 2640, "start": 26.4, "end": 31.2, "text": " First of all, this video comes out of a collaboration with Bargain Hardware.", "tokens": [50365, 2386, 295, 439, 11, 341, 960, 1487, 484, 295, 257, 9363, 365, 4156, 70, 491, 11817, 3039, 13, 50605], "temperature": 0, "avg_logprob": -0.041000396326968544, "compression_ratio": 1.6071428571428572, "no_speech_prob": 1.0731596817789568e-12}, {"id": 5, "seek": 2640, "start": 31.2, "end": 34.16, "text": " This is not a sponsored video.", "tokens": [50605, 639, 307, 406, 257, 16621, 960, 13, 50753], "temperature": 0, "avg_logprob": -0.041000396326968544, "compression_ratio": 1.6071428571428572, "no_speech_prob": 1.0731596817789568e-12}, {"id": 6, "seek": 2640, "start": 34.16, "end": 40.239999999999995, "text": " If you don't know them, Bargain Hardware are a major reseller of refurbished hardware.", "tokens": [50753, 759, 291, 500, 380, 458, 552, 11, 4156, 70, 491, 11817, 3039, 366, 257, 2563, 2025, 4658, 295, 1895, 16659, 4729, 8837, 13, 51057], "temperature": 0, "avg_logprob": -0.041000396326968544, "compression_ratio": 1.6071428571428572, "no_speech_prob": 1.0731596817789568e-12}, {"id": 7, "seek": 2640, "start": 40.239999999999995, "end": 42.959999999999994, "text": " I am not getting paid for this.", "tokens": [51057, 286, 669, 406, 1242, 4835, 337, 341, 13, 51193], "temperature": 0, "avg_logprob": -0.041000396326968544, "compression_ratio": 1.6071428571428572, "no_speech_prob": 1.0731596817789568e-12}, {"id": 8, "seek": 2640, "start": 43.519999999999996, "end": 49.28, "text": " This collaboration simply gives me access to some hardware so I can learn an experiment,", "tokens": [51221, 639, 9363, 2935, 2709, 385, 2105, 281, 512, 8837, 370, 286, 393, 1466, 364, 5120, 11, 51509], "temperature": 0, "avg_logprob": -0.041000396326968544, "compression_ratio": 1.6071428571428572, "no_speech_prob": 1.0731596817789568e-12}, {"id": 9, "seek": 2640, "start": 49.28, "end": 52.480000000000004, "text": " which is the main objective of this channel.", "tokens": [51509, 597, 307, 264, 2135, 10024, 295, 341, 2269, 13, 51669], "temperature": 0, "avg_logprob": -0.041000396326968544, "compression_ratio": 1.6071428571428572, "no_speech_prob": 1.0731596817789568e-12}, {"id": 10, "seek": 5248, "start": 52.48, "end": 59.279999999999994, "text": " Now, of course, I know that people that want to get into local AI are after alternatives to more", "tokens": [50365, 823, 11, 295, 1164, 11, 286, 458, 300, 561, 300, 528, 281, 483, 666, 2654, 7318, 366, 934, 20478, 281, 544, 50705], "temperature": 0, "avg_logprob": -0.1518343554602729, "compression_ratio": 1.3877551020408163, "no_speech_prob": 1.9661678968968532e-12}, {"id": 11, "seek": 5248, "start": 59.279999999999994, "end": 60.72, "text": " expensive hardware.", "tokens": [50705, 5124, 8837, 13, 50777], "temperature": 0, "avg_logprob": -0.1518343554602729, "compression_ratio": 1.3877551020408163, "no_speech_prob": 1.9661678968968532e-12}, {"id": 12, "seek": 5248, "start": 60.72, "end": 69.92, "text": " For reference, in August last year, one could get a 128GB Strixello machine for just over $2000.", "tokens": [50777, 1171, 6408, 11, 294, 6897, 1036, 1064, 11, 472, 727, 483, 257, 29810, 8769, 745, 6579, 11216, 3479, 337, 445, 670, 1848, 25743, 13, 51237], "temperature": 0, "avg_logprob": -0.1518343554602729, "compression_ratio": 1.3877551020408163, "no_speech_prob": 1.9661678968968532e-12}, {"id": 13, "seek": 5248, "start": 70.64, "end": 74.8, "text": " Now, we are looking at twice as much for the same machine.", "tokens": [51273, 823, 11, 321, 366, 1237, 412, 6091, 382, 709, 337, 264, 912, 3479, 13, 51481], "temperature": 0, "avg_logprob": -0.1518343554602729, "compression_ratio": 1.3877551020408163, "no_speech_prob": 1.9661678968968532e-12}, {"id": 14, "seek": 7480, "start": 74.8, "end": 77.92, "text": " The landscape has totally changed.", "tokens": [50365, 440, 9661, 575, 3879, 3105, 13, 50521], "temperature": 0, "avg_logprob": -0.08101023236910503, "compression_ratio": 1.5101214574898785, "no_speech_prob": 2.3076035440133813e-12}, {"id": 15, "seek": 7480, "start": 77.92, "end": 86.08, "text": " This is also why I am experimenting with refurbished data center hardware to see what kind of performance", "tokens": [50521, 639, 307, 611, 983, 286, 669, 29070, 365, 1895, 16659, 4729, 1412, 3056, 8837, 281, 536, 437, 733, 295, 3389, 50929], "temperature": 0, "avg_logprob": -0.08101023236910503, "compression_ratio": 1.5101214574898785, "no_speech_prob": 2.3076035440133813e-12}, {"id": 16, "seek": 7480, "start": 86.08, "end": 93.92, "text": " we can achieve, what the cost is, and most importantly, the trade-offs you need to be aware of if you decide", "tokens": [50929, 321, 393, 4584, 11, 437, 264, 2063, 307, 11, 293, 881, 8906, 11, 264, 4923, 12, 19231, 291, 643, 281, 312, 3650, 295, 498, 291, 4536, 51321], "temperature": 0, "avg_logprob": -0.08101023236910503, "compression_ratio": 1.5101214574898785, "no_speech_prob": 2.3076035440133813e-12}, {"id": 17, "seek": 7480, "start": 93.92, "end": 95.6, "text": " to go this way.", "tokens": [51321, 281, 352, 341, 636, 13, 51405], "temperature": 0, "avg_logprob": -0.08101023236910503, "compression_ratio": 1.5101214574898785, "no_speech_prob": 2.3076035440133813e-12}, {"id": 18, "seek": 7480, "start": 95.6, "end": 103.92, "text": " On top of all of this, Bargain Hardware is offering a 10% discount on the GPUs listed in the description of", "tokens": [51405, 1282, 1192, 295, 439, 295, 341, 11, 4156, 70, 491, 11817, 3039, 307, 8745, 257, 1266, 4, 11635, 322, 264, 18407, 82, 10052, 294, 264, 3855, 295, 51821], "temperature": 0, "avg_logprob": -0.08101023236910503, "compression_ratio": 1.5101214574898785, "no_speech_prob": 2.3076035440133813e-12}, {"id": 19, "seek": 10392, "start": 103.92, "end": 104.8, "text": " this video.", "tokens": [50365, 341, 960, 13, 50409], "temperature": 0, "avg_logprob": -0.07606174285153308, "compression_ratio": 1.3823529411764706, "no_speech_prob": 2.826617629195227e-12}, {"id": 20, "seek": 10392, "start": 104.8, "end": 108.72, "text": " Just use the code DONATO10 at checkout.", "tokens": [50409, 1449, 764, 264, 3089, 20403, 2218, 46, 3279, 412, 37153, 13, 50605], "temperature": 0, "avg_logprob": -0.07606174285153308, "compression_ratio": 1.3823529411764706, "no_speech_prob": 2.826617629195227e-12}, {"id": 21, "seek": 10392, "start": 108.72, "end": 114.64, "text": " This is not an affiliate link and I get zero if you buy or don't buy.", "tokens": [50605, 639, 307, 406, 364, 23975, 2113, 293, 286, 483, 4018, 498, 291, 2256, 420, 500, 380, 2256, 13, 50901], "temperature": 0, "avg_logprob": -0.07606174285153308, "compression_ratio": 1.3823529411764706, "no_speech_prob": 2.826617629195227e-12}, {"id": 22, "seek": 10392, "start": 114.64, "end": 120.96000000000001, "text": " We could have done a 5% split where you get 5% discount and I get 5%.", "tokens": [50901, 492, 727, 362, 1096, 257, 1025, 4, 7472, 689, 291, 483, 1025, 4, 11635, 293, 286, 483, 1025, 6856, 51217], "temperature": 0, "avg_logprob": -0.07606174285153308, "compression_ratio": 1.3823529411764706, "no_speech_prob": 2.826617629195227e-12}, {"id": 23, "seek": 10392, "start": 121.92, "end": 128.88, "text": " But I am trying to see what I can do to make some hardware more affordable for my viewers,", "tokens": [51265, 583, 286, 669, 1382, 281, 536, 437, 286, 393, 360, 281, 652, 512, 8837, 544, 12028, 337, 452, 8499, 11, 51613], "temperature": 0, "avg_logprob": -0.07606174285153308, "compression_ratio": 1.3823529411764706, "no_speech_prob": 2.826617629195227e-12}, {"id": 24, "seek": 12888, "start": 128.88, "end": 135.12, "text": " since it looks like that the memory and GPU situation that we have now will last for a while.", "tokens": [50365, 1670, 309, 1542, 411, 300, 264, 4675, 293, 18407, 2590, 300, 321, 362, 586, 486, 1036, 337, 257, 1339, 13, 50677], "temperature": 0, "avg_logprob": -0.05877799766008244, "compression_ratio": 1.4852941176470589, "no_speech_prob": 3.491189325480204e-12}, {"id": 25, "seek": 12888, "start": 135.12, "end": 139.68, "text": " So I'd rather get the full 10% discount to you.", "tokens": [50677, 407, 286, 1116, 2831, 483, 264, 1577, 1266, 4, 11635, 281, 291, 13, 50905], "temperature": 0, "avg_logprob": -0.05877799766008244, "compression_ratio": 1.4852941176470589, "no_speech_prob": 3.491189325480204e-12}, {"id": 26, "seek": 12888, "start": 139.68, "end": 142.96, "text": " That's the only thing that's in my power.", "tokens": [50905, 663, 311, 264, 787, 551, 300, 311, 294, 452, 1347, 13, 51069], "temperature": 0, "avg_logprob": -0.05877799766008244, "compression_ratio": 1.4852941176470589, "no_speech_prob": 3.491189325480204e-12}, {"id": 27, "seek": 12888, "start": 142.96, "end": 148.72, "text": " Now, the first consideration is that when you want to build a multi-GPU system,", "tokens": [51069, 823, 11, 264, 700, 12381, 307, 300, 562, 291, 528, 281, 1322, 257, 4825, 12, 38, 8115, 1185, 11, 51357], "temperature": 0, "avg_logprob": -0.05877799766008244, "compression_ratio": 1.4852941176470589, "no_speech_prob": 3.491189325480204e-12}, {"id": 28, "seek": 12888, "start": 148.72, "end": 152.16, "text": " it's not just the GPUs that cost money.", "tokens": [51357, 309, 311, 406, 445, 264, 18407, 82, 300, 2063, 1460, 13, 51529], "temperature": 0, "avg_logprob": -0.05877799766008244, "compression_ratio": 1.4852941176470589, "no_speech_prob": 3.491189325480204e-12}, {"id": 29, "seek": 15216, "start": 152.16, "end": 159.04, "text": " The issue is that unless you already have a rig to place your cards in, you also have to build a", "tokens": [50365, 440, 2734, 307, 300, 5969, 291, 1217, 362, 257, 8329, 281, 1081, 428, 5632, 294, 11, 291, 611, 362, 281, 1322, 257, 50709], "temperature": 0, "avg_logprob": -0.08862997845905583, "compression_ratio": 1.412621359223301, "no_speech_prob": 2.316044057926181e-12}, {"id": 30, "seek": 15216, "start": 159.04, "end": 160.56, "text": " system to host them.", "tokens": [50709, 1185, 281, 3975, 552, 13, 50785], "temperature": 0, "avg_logprob": -0.08862997845905583, "compression_ratio": 1.412621359223301, "no_speech_prob": 2.316044057926181e-12}, {"id": 31, "seek": 15216, "start": 160.56, "end": 167.44, "text": " If you watched my Dual Radio 9700 video, you'll have an idea of what that looks like.", "tokens": [50785, 759, 291, 6337, 452, 37625, 17296, 1722, 18197, 960, 11, 291, 603, 362, 364, 1558, 295, 437, 300, 1542, 411, 13, 51129], "temperature": 0, "avg_logprob": -0.08862997845905583, "compression_ratio": 1.412621359223301, "no_speech_prob": 2.316044057926181e-12}, {"id": 32, "seek": 15216, "start": 167.44, "end": 176.56, "text": " But that setup with 64GB of RAM would cost around $4,000 to $5,000 with current prices,", "tokens": [51129, 583, 300, 8657, 365, 12145, 8769, 295, 14561, 576, 2063, 926, 1848, 19, 11, 1360, 281, 1848, 20, 11, 1360, 365, 2190, 7901, 11, 51585], "temperature": 0, "avg_logprob": -0.08862997845905583, "compression_ratio": 1.412621359223301, "no_speech_prob": 2.316044057926181e-12}, {"id": 33, "seek": 17656, "start": 176.56, "end": 180.64000000000001, "text": " maybe even more, I really cannot keep up myself.", "tokens": [50365, 1310, 754, 544, 11, 286, 534, 2644, 1066, 493, 2059, 13, 50569], "temperature": 0, "avg_logprob": -0.09813350967214077, "compression_ratio": 1.4360189573459716, "no_speech_prob": 2.28050217598863e-12}, {"id": 34, "seek": 17656, "start": 180.64000000000001, "end": 188.08, "text": " And if you want to scale up to four dual-slot GPUs, you need a motherboard that physically can", "tokens": [50569, 400, 498, 291, 528, 281, 4373, 493, 281, 1451, 11848, 12, 10418, 310, 18407, 82, 11, 291, 643, 257, 32916, 300, 9762, 393, 50941], "temperature": 0, "avg_logprob": -0.09813350967214077, "compression_ratio": 1.4360189573459716, "no_speech_prob": 2.28050217598863e-12}, {"id": 35, "seek": 17656, "start": 188.08, "end": 197.2, "text": " accommodate them, a GPU with enough PCIe lanes to avoid lane starvation, a high capacity power supply,", "tokens": [50941, 21410, 552, 11, 257, 18407, 365, 1547, 6465, 40, 68, 25397, 281, 5042, 12705, 3543, 11116, 11, 257, 1090, 6042, 1347, 5847, 11, 51397], "temperature": 0, "avg_logprob": -0.09813350967214077, "compression_ratio": 1.4360189573459716, "no_speech_prob": 2.28050217598863e-12}, {"id": 36, "seek": 17656, "start": 197.2, "end": 201.6, "text": " and a chassis with proper cooling, let alone system RAM.", "tokens": [51397, 293, 257, 28262, 365, 2296, 14785, 11, 718, 3312, 1185, 14561, 13, 51617], "temperature": 0, "avg_logprob": -0.09813350967214077, "compression_ratio": 1.4360189573459716, "no_speech_prob": 2.28050217598863e-12}, {"id": 37, "seek": 20160, "start": 201.6, "end": 207.44, "text": " If you try to build that base system from scratch using currently available hardware,", "tokens": [50365, 759, 291, 853, 281, 1322, 300, 3096, 1185, 490, 8459, 1228, 4362, 2435, 8837, 11, 50657], "temperature": 0, "avg_logprob": -0.0369260311126709, "compression_ratio": 1.4663461538461537, "no_speech_prob": 2.9167740271673903e-12}, {"id": 38, "seek": 20160, "start": 207.44, "end": 214.32, "text": " it can easily cost you around $4,000 before you even buy the GPUs.", "tokens": [50657, 309, 393, 3612, 2063, 291, 926, 1848, 19, 11, 1360, 949, 291, 754, 2256, 264, 18407, 82, 13, 51001], "temperature": 0, "avg_logprob": -0.0369260311126709, "compression_ratio": 1.4663461538461537, "no_speech_prob": 2.9167740271673903e-12}, {"id": 39, "seek": 20160, "start": 214.32, "end": 221.12, "text": " That is where refurbished data center servers start making sense in 2026.", "tokens": [51001, 663, 307, 689, 1895, 16659, 4729, 1412, 3056, 15909, 722, 1455, 2020, 294, 945, 10880, 13, 51341], "temperature": 0, "avg_logprob": -0.0369260311126709, "compression_ratio": 1.4663461538461537, "no_speech_prob": 2.9167740271673903e-12}, {"id": 40, "seek": 20160, "start": 221.12, "end": 228.32, "text": " They offer great value by giving you a base system for a fraction of the cost.", "tokens": [51341, 814, 2626, 869, 2158, 538, 2902, 291, 257, 3096, 1185, 337, 257, 14135, 295, 264, 2063, 13, 51701], "temperature": 0, "avg_logprob": -0.0369260311126709, "compression_ratio": 1.4663461538461537, "no_speech_prob": 2.9167740271673903e-12}, {"id": 41, "seek": 22832, "start": 228.32, "end": 234.07999999999998, "text": " Now, Bargain Hardware is a simple online builder where you can select all the components for the", "tokens": [50365, 823, 11, 4156, 70, 491, 11817, 3039, 307, 257, 2199, 2950, 27377, 689, 291, 393, 3048, 439, 264, 6677, 337, 264, 50653], "temperature": 0, "avg_logprob": -0.10150429725646973, "compression_ratio": 1.5222672064777327, "no_speech_prob": 5.489657773499745e-12}, {"id": 42, "seek": 22832, "start": 234.07999999999998, "end": 234.79999999999998, "text": " system.", "tokens": [50653, 1185, 13, 50689], "temperature": 0, "avg_logprob": -0.10150429725646973, "compression_ratio": 1.5222672064777327, "no_speech_prob": 5.489657773499745e-12}, {"id": 43, "seek": 22832, "start": 234.79999999999998, "end": 242.16, "text": " For a quad GPU configuration, one of the cheapest options is to get a super micro server.", "tokens": [50689, 1171, 257, 10787, 18407, 11694, 11, 472, 295, 264, 29167, 3956, 307, 281, 483, 257, 1687, 4532, 7154, 13, 51057], "temperature": 0, "avg_logprob": -0.10150429725646973, "compression_ratio": 1.5222672064777327, "no_speech_prob": 5.489657773499745e-12}, {"id": 44, "seek": 22832, "start": 242.16, "end": 244.72, "text": " I'll get this 12th generation one.", "tokens": [51057, 286, 603, 483, 341, 2272, 392, 5125, 472, 13, 51185], "temperature": 0, "avg_logprob": -0.10150429725646973, "compression_ratio": 1.5222672064777327, "no_speech_prob": 5.489657773499745e-12}, {"id": 45, "seek": 22832, "start": 244.72, "end": 249.92, "text": " The motherboard in these servers typically hosts two Intel Xeon CPUs.", "tokens": [51185, 440, 32916, 294, 613, 15909, 5850, 21573, 732, 19762, 1783, 27015, 13199, 82, 13, 51445], "temperature": 0, "avg_logprob": -0.10150429725646973, "compression_ratio": 1.5222672064777327, "no_speech_prob": 5.489657773499745e-12}, {"id": 46, "seek": 22832, "start": 249.92, "end": 252.4, "text": " So we can now configure these.", "tokens": [51445, 407, 321, 393, 586, 22162, 613, 13, 51569], "temperature": 0, "avg_logprob": -0.10150429725646973, "compression_ratio": 1.5222672064777327, "no_speech_prob": 5.489657773499745e-12}, {"id": 47, "seek": 22832, "start": 252.4, "end": 255.35999999999999, "text": " You can see there is a base price, of course.", "tokens": [51569, 509, 393, 536, 456, 307, 257, 3096, 3218, 11, 295, 1164, 13, 51717], "temperature": 0, "avg_logprob": -0.10150429725646973, "compression_ratio": 1.5222672064777327, "no_speech_prob": 5.489657773499745e-12}, {"id": 48, "seek": 25536, "start": 255.36, "end": 260.88, "text": " So the first thing we want to do is to select the CPUs.", "tokens": [50365, 407, 264, 700, 551, 321, 528, 281, 360, 307, 281, 3048, 264, 13199, 82, 13, 50641], "temperature": 0, "avg_logprob": -0.1100536748903607, "compression_ratio": 1.5443037974683544, "no_speech_prob": 3.800423951233478e-12}, {"id": 49, "seek": 25536, "start": 260.88, "end": 268.40000000000003, "text": " And here I'm going for two of these Xeon 18 core CPUs that you can see here.", "tokens": [50641, 400, 510, 286, 478, 516, 337, 732, 295, 613, 1783, 27015, 2443, 4965, 13199, 82, 300, 291, 393, 536, 510, 13, 51017], "temperature": 0, "avg_logprob": -0.1100536748903607, "compression_ratio": 1.5443037974683544, "no_speech_prob": 3.800423951233478e-12}, {"id": 50, "seek": 25536, "start": 268.40000000000003, "end": 270.64, "text": " They are incredibly cheap.", "tokens": [51017, 814, 366, 6252, 7084, 13, 51129], "temperature": 0, "avg_logprob": -0.1100536748903607, "compression_ratio": 1.5443037974683544, "no_speech_prob": 3.800423951233478e-12}, {"id": 51, "seek": 25536, "start": 270.64, "end": 273.28000000000003, "text": " Then, of course, the heat sinks are already included.", "tokens": [51129, 1396, 11, 295, 1164, 11, 264, 3738, 43162, 366, 1217, 5556, 13, 51261], "temperature": 0, "avg_logprob": -0.1100536748903607, "compression_ratio": 1.5443037974683544, "no_speech_prob": 3.800423951233478e-12}, {"id": 52, "seek": 25536, "start": 273.28000000000003, "end": 277.12, "text": " And for the RAM, as you can see here, there are a lot of options.", "tokens": [51261, 400, 337, 264, 14561, 11, 382, 291, 393, 536, 510, 11, 456, 366, 257, 688, 295, 3956, 13, 51453], "temperature": 0, "avg_logprob": -0.1100536748903607, "compression_ratio": 1.5443037974683544, "no_speech_prob": 3.800423951233478e-12}, {"id": 53, "seek": 25536, "start": 277.12, "end": 284.64, "text": " Actually, Bargain Hardware stocks right now, a lot of DDR4 at a very, very good price.", "tokens": [51453, 5135, 11, 4156, 70, 491, 11817, 3039, 12966, 558, 586, 11, 257, 688, 295, 49272, 19, 412, 257, 588, 11, 588, 665, 3218, 13, 51829], "temperature": 0, "avg_logprob": -0.1100536748903607, "compression_ratio": 1.5443037974683544, "no_speech_prob": 3.800423951233478e-12}, {"id": 54, "seek": 28464, "start": 284.64, "end": 289.36, "text": " Here I'm getting four of these 16 GB sticks.", "tokens": [50365, 1692, 286, 478, 1242, 1451, 295, 613, 3165, 26809, 12518, 13, 50601], "temperature": 0, "avg_logprob": -0.13911705017089843, "compression_ratio": 1.5172413793103448, "no_speech_prob": 3.1281257965865006e-12}, {"id": 55, "seek": 28464, "start": 289.36, "end": 292.0, "text": " Now we want to select the storage.", "tokens": [50601, 823, 321, 528, 281, 3048, 264, 6725, 13, 50733], "temperature": 0, "avg_logprob": -0.13911705017089843, "compression_ratio": 1.5172413793103448, "no_speech_prob": 3.1281257965865006e-12}, {"id": 56, "seek": 28464, "start": 292.0, "end": 294.88, "text": " And for the storage, I went with SATA.", "tokens": [50733, 400, 337, 264, 6725, 11, 286, 1437, 365, 31536, 32, 13, 50877], "temperature": 0, "avg_logprob": -0.13911705017089843, "compression_ratio": 1.5172413793103448, "no_speech_prob": 3.1281257965865006e-12}, {"id": 57, "seek": 28464, "start": 294.88, "end": 303.44, "text": " And I selected two of these 480 GB SSDs, which gives me almost one terabyte, which is what you want", "tokens": [50877, 400, 286, 8209, 732, 295, 613, 1017, 4702, 26809, 30262, 82, 11, 597, 2709, 385, 1920, 472, 1796, 34529, 11, 597, 307, 437, 291, 528, 51305], "temperature": 0, "avg_logprob": -0.13911705017089843, "compression_ratio": 1.5172413793103448, "no_speech_prob": 3.1281257965865006e-12}, {"id": 58, "seek": 28464, "start": 304.08, "end": 308.64, "text": " for LLMs, at least one terabyte, because you're going to be downloading a lot of weights.", "tokens": [51337, 337, 441, 43, 26386, 11, 412, 1935, 472, 1796, 34529, 11, 570, 291, 434, 516, 281, 312, 32529, 257, 688, 295, 17443, 13, 51565], "temperature": 0, "avg_logprob": -0.13911705017089843, "compression_ratio": 1.5172413793103448, "no_speech_prob": 3.1281257965865006e-12}, {"id": 59, "seek": 30864, "start": 308.64, "end": 313.52, "text": " And moving on, now we select the GPU accelerators.", "tokens": [50365, 400, 2684, 322, 11, 586, 321, 3048, 264, 18407, 10172, 3391, 13, 50609], "temperature": 0, "avg_logprob": -0.14954196384974888, "compression_ratio": 1.3313253012048192, "no_speech_prob": 3.1889196144829768e-12}, {"id": 60, "seek": 30864, "start": 313.52, "end": 320.47999999999996, "text": " Now for this build, I have selected four P100s, 16 GB each.", "tokens": [50609, 823, 337, 341, 1322, 11, 286, 362, 8209, 1451, 430, 6879, 82, 11, 3165, 26809, 1184, 13, 50957], "temperature": 0, "avg_logprob": -0.14954196384974888, "compression_ratio": 1.3313253012048192, "no_speech_prob": 3.1889196144829768e-12}, {"id": 61, "seek": 30864, "start": 320.47999999999996, "end": 322.56, "text": " So we get 64 GB.", "tokens": [50957, 407, 321, 483, 12145, 26809, 13, 51061], "temperature": 0, "avg_logprob": -0.14954196384974888, "compression_ratio": 1.3313253012048192, "no_speech_prob": 3.1889196144829768e-12}, {"id": 62, "seek": 30864, "start": 322.56, "end": 332.08, "text": " But in this video and in the benchmarks, I've also looked at the V100s and also the AMD MI25.", "tokens": [51061, 583, 294, 341, 960, 293, 294, 264, 43751, 11, 286, 600, 611, 2956, 412, 264, 691, 6879, 82, 293, 611, 264, 34808, 13696, 6074, 13, 51537], "temperature": 0, "avg_logprob": -0.14954196384974888, "compression_ratio": 1.3313253012048192, "no_speech_prob": 3.1889196144829768e-12}, {"id": 63, "seek": 33208, "start": 332.08, "end": 337.91999999999996, "text": " And again, each of these options ultimately gives you four GPUs with 64 GB of RAM.", "tokens": [50365, 400, 797, 11, 1184, 295, 613, 3956, 6284, 2709, 291, 1451, 18407, 82, 365, 12145, 26809, 295, 14561, 13, 50657], "temperature": 0, "avg_logprob": -0.08525167431747704, "compression_ratio": 1.3454545454545455, "no_speech_prob": 3.2661074365891718e-12}, {"id": 64, "seek": 33208, "start": 337.91999999999996, "end": 344.4, "text": " But of course, the V100 is much more expensive because it's a more modern architecture.", "tokens": [50657, 583, 295, 1164, 11, 264, 691, 6879, 307, 709, 544, 5124, 570, 309, 311, 257, 544, 4363, 9482, 13, 50981], "temperature": 0, "avg_logprob": -0.08525167431747704, "compression_ratio": 1.3454545454545455, "no_speech_prob": 3.2661074365891718e-12}, {"id": 65, "seek": 33208, "start": 344.4, "end": 349.2, "text": " But check the website because more GPUs are coming.", "tokens": [50981, 583, 1520, 264, 3144, 570, 544, 18407, 82, 366, 1348, 13, 51221], "temperature": 0, "avg_logprob": -0.08525167431747704, "compression_ratio": 1.3454545454545455, "no_speech_prob": 3.2661074365891718e-12}, {"id": 66, "seek": 34920, "start": 349.2, "end": 355.03999999999996, "text": " And probably by the time this video is uploaded, you will see some more options.", "tokens": [50365, 400, 1391, 538, 264, 565, 341, 960, 307, 17135, 11, 291, 486, 536, 512, 544, 3956, 13, 50657], "temperature": 0, "avg_logprob": -0.07702954610188802, "compression_ratio": 1.2446043165467626, "no_speech_prob": 4.079707058290971e-12}, {"id": 67, "seek": 34920, "start": 355.03999999999996, "end": 357.84, "text": " I know that there are some other GPUs coming.", "tokens": [50657, 286, 458, 300, 456, 366, 512, 661, 18407, 82, 1348, 13, 50797], "temperature": 0, "avg_logprob": -0.07702954610188802, "compression_ratio": 1.2446043165467626, "no_speech_prob": 4.079707058290971e-12}, {"id": 68, "seek": 34920, "start": 357.84, "end": 365.36, "text": " RTX 3080, 3090, Quadro RTX 5000, Quadro M6000.", "tokens": [50797, 44573, 2217, 4702, 11, 2217, 7771, 11, 29619, 340, 44573, 23777, 11, 29619, 340, 376, 21, 1360, 13, 51173], "temperature": 0, "avg_logprob": -0.07702954610188802, "compression_ratio": 1.2446043165467626, "no_speech_prob": 4.079707058290971e-12}, {"id": 69, "seek": 36536, "start": 365.36, "end": 375.6, "text": " But just check the stock because they are updating it and they will also make some A100s and RTX 8000 available.", "tokens": [50365, 583, 445, 1520, 264, 4127, 570, 436, 366, 25113, 309, 293, 436, 486, 611, 652, 512, 316, 6879, 82, 293, 44573, 1649, 1360, 2435, 13, 50877], "temperature": 0, "avg_logprob": -0.12873342457939596, "compression_ratio": 1.432748538011696, "no_speech_prob": 4.533474724094377e-12}, {"id": 70, "seek": 36536, "start": 375.6, "end": 379.04, "text": " So there is quite a lot that's coming to their stock.", "tokens": [50877, 407, 456, 307, 1596, 257, 688, 300, 311, 1348, 281, 641, 4127, 13, 51049], "temperature": 0, "avg_logprob": -0.12873342457939596, "compression_ratio": 1.432748538011696, "no_speech_prob": 4.533474724094377e-12}, {"id": 71, "seek": 36536, "start": 379.76, "end": 386.88, "text": " But for this build, I will keep the P100s in there and we can add to the cart.", "tokens": [51085, 583, 337, 341, 1322, 11, 286, 486, 1066, 264, 430, 6879, 82, 294, 456, 293, 321, 393, 909, 281, 264, 5467, 13, 51441], "temperature": 0, "avg_logprob": -0.12873342457939596, "compression_ratio": 1.432748538011696, "no_speech_prob": 4.533474724094377e-12}, {"id": 72, "seek": 38688, "start": 386.88, "end": 394.15999999999997, "text": " Oh, actually, we forgot to select the Super Macro NVMe Enablement Kit.", "tokens": [50365, 876, 11, 767, 11, 321, 5298, 281, 3048, 264, 4548, 5707, 340, 46512, 12671, 2193, 712, 518, 23037, 13, 50729], "temperature": 0, "avg_logprob": -0.1080839474995931, "compression_ratio": 1.4796380090497738, "no_speech_prob": 4.19370311047218e-12}, {"id": 73, "seek": 38688, "start": 394.15999999999997, "end": 398.32, "text": " We add to the cart and now we can go and check out.", "tokens": [50729, 492, 909, 281, 264, 5467, 293, 586, 321, 393, 352, 293, 1520, 484, 13, 50937], "temperature": 0, "avg_logprob": -0.1080839474995931, "compression_ratio": 1.4796380090497738, "no_speech_prob": 4.19370311047218e-12}, {"id": 74, "seek": 38688, "start": 398.32, "end": 402.96, "text": " And I just want to show you the different options and costs for delivery.", "tokens": [50937, 400, 286, 445, 528, 281, 855, 291, 264, 819, 3956, 293, 5497, 337, 8982, 13, 51169], "temperature": 0, "avg_logprob": -0.1080839474995931, "compression_ratio": 1.4796380090497738, "no_speech_prob": 4.19370311047218e-12}, {"id": 75, "seek": 38688, "start": 402.96, "end": 408.56, "text": " That's actually at least in the UK for domestic delivery.", "tokens": [51169, 663, 311, 767, 412, 1935, 294, 264, 7051, 337, 10939, 8982, 13, 51449], "temperature": 0, "avg_logprob": -0.1080839474995931, "compression_ratio": 1.4796380090497738, "no_speech_prob": 4.19370311047218e-12}, {"id": 76, "seek": 38688, "start": 408.56, "end": 415.68, "text": " When you look at this, it brings us to a total of just over 2000 pounds.", "tokens": [51449, 1133, 291, 574, 412, 341, 11, 309, 5607, 505, 281, 257, 3217, 295, 445, 670, 8132, 8319, 13, 51805], "temperature": 0, "avg_logprob": -0.1080839474995931, "compression_ratio": 1.4796380090497738, "no_speech_prob": 4.19370311047218e-12}, {"id": 77, "seek": 41568, "start": 415.68, "end": 424.56, "text": " But this is not including the 10% discount on the GPUs that you can get with the code Donato 10.", "tokens": [50365, 583, 341, 307, 406, 3009, 264, 1266, 4, 11635, 322, 264, 18407, 82, 300, 291, 393, 483, 365, 264, 3089, 1468, 2513, 1266, 13, 50809], "temperature": 0, "avg_logprob": -0.09611383739270662, "compression_ratio": 1.5188284518828452, "no_speech_prob": 3.2781479353954923e-12}, {"id": 78, "seek": 41568, "start": 426.48, "end": 432.0, "text": " Now, I thought it'd be interesting to actually go in person to their warehouse,", "tokens": [50905, 823, 11, 286, 1194, 309, 1116, 312, 1880, 281, 767, 352, 294, 954, 281, 641, 22244, 11, 51181], "temperature": 0, "avg_logprob": -0.09611383739270662, "compression_ratio": 1.5188284518828452, "no_speech_prob": 3.2781479353954923e-12}, {"id": 79, "seek": 41568, "start": 432.0, "end": 440.0, "text": " see how an order like this is put together and generally get a look at the process they use to test and refurbish hardware.", "tokens": [51181, 536, 577, 364, 1668, 411, 341, 307, 829, 1214, 293, 5101, 483, 257, 574, 412, 264, 1399, 436, 764, 281, 1500, 293, 1895, 16659, 742, 8837, 13, 51581], "temperature": 0, "avg_logprob": -0.09611383739270662, "compression_ratio": 1.5188284518828452, "no_speech_prob": 3.2781479353954923e-12}, {"id": 80, "seek": 41568, "start": 440.0, "end": 444.64, "text": " And a thank you to Jack Moyers from the Bargain Hardware team,", "tokens": [51581, 400, 257, 1309, 291, 281, 4718, 47254, 433, 490, 264, 4156, 70, 491, 11817, 3039, 1469, 11, 51813], "temperature": 0, "avg_logprob": -0.09611383739270662, "compression_ratio": 1.5188284518828452, "no_speech_prob": 3.2781479353954923e-12}, {"id": 81, "seek": 44464, "start": 444.64, "end": 447.52, "text": " who was kind enough to give me a tour.", "tokens": [50365, 567, 390, 733, 1547, 281, 976, 385, 257, 3512, 13, 50509], "temperature": 0, "avg_logprob": -0.08470794132777623, "compression_ratio": 1.5721153846153846, "no_speech_prob": 5.199047355824993e-12}, {"id": 82, "seek": 44464, "start": 447.52, "end": 450.64, "text": " The first process at Bargain Hardware is inbound.", "tokens": [50509, 440, 700, 1399, 412, 4156, 70, 491, 11817, 3039, 307, 294, 18767, 13, 50665], "temperature": 0, "avg_logprob": -0.08470794132777623, "compression_ratio": 1.5721153846153846, "no_speech_prob": 5.199047355824993e-12}, {"id": 83, "seek": 44464, "start": 451.2, "end": 456.71999999999997, "text": " So pallets of servers, workstations, components, they arrive on trucks.", "tokens": [50693, 407, 24075, 1385, 295, 15909, 11, 589, 372, 763, 11, 6677, 11, 436, 8881, 322, 16156, 13, 50969], "temperature": 0, "avg_logprob": -0.08470794132777623, "compression_ratio": 1.5721153846153846, "no_speech_prob": 5.199047355824993e-12}, {"id": 84, "seek": 44464, "start": 456.71999999999997, "end": 462.24, "text": " So we buy equipment from all kinds of places all around Europe and some more worldwide.", "tokens": [50969, 407, 321, 2256, 5927, 490, 439, 3685, 295, 3190, 439, 926, 3315, 293, 512, 544, 13485, 13, 51245], "temperature": 0, "avg_logprob": -0.08470794132777623, "compression_ratio": 1.5721153846153846, "no_speech_prob": 5.199047355824993e-12}, {"id": 85, "seek": 44464, "start": 462.71999999999997, "end": 466.64, "text": " And then before the servers and workstations and any equipment gets processed,", "tokens": [51269, 400, 550, 949, 264, 15909, 293, 589, 372, 763, 293, 604, 5927, 2170, 18846, 11, 51465], "temperature": 0, "avg_logprob": -0.08470794132777623, "compression_ratio": 1.5721153846153846, "no_speech_prob": 5.199047355824993e-12}, {"id": 86, "seek": 46664, "start": 466.64, "end": 469.59999999999997, "text": " it gets stored in our rack in here.", "tokens": [50365, 309, 2170, 12187, 294, 527, 14788, 294, 510, 13, 50513], "temperature": 0, "avg_logprob": -0.11035630336174598, "compression_ratio": 1.7706422018348624, "no_speech_prob": 3.815307011295621e-12}, {"id": 87, "seek": 46664, "start": 469.59999999999997, "end": 473.2, "text": " So the first part of the refurbishment process is cleaning.", "tokens": [50513, 407, 264, 700, 644, 295, 264, 1895, 16659, 30273, 1399, 307, 8924, 13, 50693], "temperature": 0, "avg_logprob": -0.11035630336174598, "compression_ratio": 1.7706422018348624, "no_speech_prob": 3.815307011295621e-12}, {"id": 88, "seek": 46664, "start": 473.2, "end": 475.2, "text": " So this is one of our cleaning rooms.", "tokens": [50693, 407, 341, 307, 472, 295, 527, 8924, 9396, 13, 50793], "temperature": 0, "avg_logprob": -0.11035630336174598, "compression_ratio": 1.7706422018348624, "no_speech_prob": 3.815307011295621e-12}, {"id": 89, "seek": 46664, "start": 475.2, "end": 478.0, "text": " So in here, the first part of the process is the servers,", "tokens": [50793, 407, 294, 510, 11, 264, 700, 644, 295, 264, 1399, 307, 264, 15909, 11, 50933], "temperature": 0, "avg_logprob": -0.11035630336174598, "compression_ratio": 1.7706422018348624, "no_speech_prob": 3.815307011295621e-12}, {"id": 90, "seek": 46664, "start": 478.56, "end": 481.03999999999996, "text": " they get blown out with our air compressor machine.", "tokens": [50961, 436, 483, 16479, 484, 365, 527, 1988, 28765, 3479, 13, 51085], "temperature": 0, "avg_logprob": -0.11035630336174598, "compression_ratio": 1.7706422018348624, "no_speech_prob": 3.815307011295621e-12}, {"id": 91, "seek": 46664, "start": 481.91999999999996, "end": 486.4, "text": " You can see here, Michael's doing some reskinning on a server.", "tokens": [51129, 509, 393, 536, 510, 11, 5116, 311, 884, 512, 725, 5843, 773, 322, 257, 7154, 13, 51353], "temperature": 0, "avg_logprob": -0.11035630336174598, "compression_ratio": 1.7706422018348624, "no_speech_prob": 3.815307011295621e-12}, {"id": 92, "seek": 46664, "start": 486.4, "end": 491.84, "text": " So certain generations of servers get reskinned just so we have a nicer finish.", "tokens": [51353, 407, 1629, 10593, 295, 15909, 483, 725, 5843, 9232, 445, 370, 321, 362, 257, 22842, 2413, 13, 51625], "temperature": 0, "avg_logprob": -0.11035630336174598, "compression_ratio": 1.7706422018348624, "no_speech_prob": 3.815307011295621e-12}, {"id": 93, "seek": 49184, "start": 491.84, "end": 494.71999999999997, "text": " So it just gets rid of all these scuffs and scrapes off the lids.", "tokens": [50365, 407, 309, 445, 2170, 3973, 295, 439, 613, 795, 28296, 293, 23138, 279, 766, 264, 287, 3742, 13, 50509], "temperature": 0, "avg_logprob": -0.08403902304799933, "compression_ratio": 1.5555555555555556, "no_speech_prob": 4.272842930169718e-12}, {"id": 94, "seek": 49184, "start": 494.71999999999997, "end": 502.79999999999995, "text": " So sometimes we have orders which are just a batch of drives or it could be CPUs or GPUs.", "tokens": [50509, 407, 2171, 321, 362, 9470, 597, 366, 445, 257, 15245, 295, 11754, 420, 309, 727, 312, 13199, 82, 420, 18407, 82, 13, 50913], "temperature": 0, "avg_logprob": -0.08403902304799933, "compression_ratio": 1.5555555555555556, "no_speech_prob": 4.272842930169718e-12}, {"id": 95, "seek": 49184, "start": 502.79999999999995, "end": 506.08, "text": " So not necessarily every component comes out of a machine.", "tokens": [50913, 407, 406, 4725, 633, 6542, 1487, 484, 295, 257, 3479, 13, 51077], "temperature": 0, "avg_logprob": -0.08403902304799933, "compression_ratio": 1.5555555555555556, "no_speech_prob": 4.272842930169718e-12}, {"id": 96, "seek": 49184, "start": 506.08, "end": 510.4, "text": " Sometimes we buy them in batch, but still every single component needs testing.", "tokens": [51077, 4803, 321, 2256, 552, 294, 15245, 11, 457, 920, 633, 2167, 6542, 2203, 4997, 13, 51293], "temperature": 0, "avg_logprob": -0.08403902304799933, "compression_ratio": 1.5555555555555556, "no_speech_prob": 4.272842930169718e-12}, {"id": 97, "seek": 51040, "start": 510.4, "end": 513.12, "text": " So then we come to component testing here.", "tokens": [50365, 407, 550, 321, 808, 281, 6542, 4997, 510, 13, 50501], "temperature": 0, "avg_logprob": -0.06403636932373047, "compression_ratio": 1.696078431372549, "no_speech_prob": 4.1576160916823035e-12}, {"id": 98, "seek": 51040, "start": 518.3199999999999, "end": 522.8, "text": " So when a machine has been tested, the next thing that we do,", "tokens": [50761, 407, 562, 257, 3479, 575, 668, 8246, 11, 264, 958, 551, 300, 321, 360, 11, 50985], "temperature": 0, "avg_logprob": -0.06403636932373047, "compression_ratio": 1.696078431372549, "no_speech_prob": 4.1576160916823035e-12}, {"id": 99, "seek": 51040, "start": 522.8, "end": 527.36, "text": " because we stock our servers and workstations as configured to order,", "tokens": [50985, 570, 321, 4127, 527, 15909, 293, 589, 372, 763, 382, 30538, 281, 1668, 11, 51213], "temperature": 0, "avg_logprob": -0.06403636932373047, "compression_ratio": 1.696078431372549, "no_speech_prob": 4.1576160916823035e-12}, {"id": 100, "seek": 51040, "start": 527.36, "end": 531.28, "text": " the next thing we need to do is take the components out and that's called disassembly.", "tokens": [51213, 264, 958, 551, 321, 643, 281, 360, 307, 747, 264, 6677, 484, 293, 300, 311, 1219, 717, 29386, 356, 13, 51409], "temperature": 0, "avg_logprob": -0.06403636932373047, "compression_ratio": 1.696078431372549, "no_speech_prob": 4.1576160916823035e-12}, {"id": 101, "seek": 51040, "start": 531.28, "end": 532.9599999999999, "text": " So this is the next section.", "tokens": [51409, 407, 341, 307, 264, 958, 3541, 13, 51493], "temperature": 0, "avg_logprob": -0.06403636932373047, "compression_ratio": 1.696078431372549, "no_speech_prob": 4.1576160916823035e-12}, {"id": 102, "seek": 51040, "start": 532.9599999999999, "end": 536.8, "text": " So here we've got Harry who's working on a workstation.", "tokens": [51493, 407, 510, 321, 600, 658, 9378, 567, 311, 1364, 322, 257, 589, 19159, 13, 51685], "temperature": 0, "avg_logprob": -0.06403636932373047, "compression_ratio": 1.696078431372549, "no_speech_prob": 4.1576160916823035e-12}, {"id": 103, "seek": 53680, "start": 536.8, "end": 543.28, "text": " So he'll be taking out the CPU, the RAM, any GPUs and other cards and components so that", "tokens": [50365, 407, 415, 603, 312, 1940, 484, 264, 13199, 11, 264, 14561, 11, 604, 18407, 82, 293, 661, 5632, 293, 6677, 370, 300, 50689], "temperature": 0, "avg_logprob": -0.0722759485244751, "compression_ratio": 1.5145631067961165, "no_speech_prob": 7.893499222311195e-12}, {"id": 104, "seek": 53680, "start": 544.0, "end": 547.3599999999999, "text": " when we stock it and we offer it as configured to order on our website,", "tokens": [50725, 562, 321, 4127, 309, 293, 321, 2626, 309, 382, 30538, 281, 1668, 322, 527, 3144, 11, 50893], "temperature": 0, "avg_logprob": -0.0722759485244751, "compression_ratio": 1.5145631067961165, "no_speech_prob": 7.893499222311195e-12}, {"id": 105, "seek": 53680, "start": 547.3599999999999, "end": 552.0, "text": " it's basically a chassis with a PSU and other base components.", "tokens": [50893, 309, 311, 1936, 257, 28262, 365, 257, 8168, 52, 293, 661, 3096, 6677, 13, 51125], "temperature": 0, "avg_logprob": -0.0722759485244751, "compression_ratio": 1.5145631067961165, "no_speech_prob": 7.893499222311195e-12}, {"id": 106, "seek": 53680, "start": 555.04, "end": 562.56, "text": " As we were walking around, my attention was immediately captured by a stock of GPUs that", "tokens": [51277, 1018, 321, 645, 4494, 926, 11, 452, 3202, 390, 4258, 11828, 538, 257, 4127, 295, 18407, 82, 300, 51653], "temperature": 0, "avg_logprob": -0.0722759485244751, "compression_ratio": 1.5145631067961165, "no_speech_prob": 7.893499222311195e-12}, {"id": 107, "seek": 56256, "start": 562.56, "end": 567.8399999999999, "text": " they got in and that they were testing and getting ready to refurbish.", "tokens": [50365, 436, 658, 294, 293, 300, 436, 645, 4997, 293, 1242, 1919, 281, 1895, 16659, 742, 13, 50629], "temperature": 0, "avg_logprob": -0.09107277311127761, "compression_ratio": 1.7333333333333334, "no_speech_prob": 8.910949235441112e-12}, {"id": 108, "seek": 56256, "start": 569.8399999999999, "end": 571.3599999999999, "text": " So then we come to picking.", "tokens": [50729, 407, 550, 321, 808, 281, 8867, 13, 50805], "temperature": 0, "avg_logprob": -0.09107277311127761, "compression_ratio": 1.7333333333333334, "no_speech_prob": 8.910949235441112e-12}, {"id": 109, "seek": 56256, "start": 571.3599999999999, "end": 575.5999999999999, "text": " So if someone orders on our website, they've configured a server.", "tokens": [50805, 407, 498, 1580, 9470, 322, 527, 3144, 11, 436, 600, 30538, 257, 7154, 13, 51017], "temperature": 0, "avg_logprob": -0.09107277311127761, "compression_ratio": 1.7333333333333334, "no_speech_prob": 8.910949235441112e-12}, {"id": 110, "seek": 56256, "start": 576.3199999999999, "end": 578.2399999999999, "text": " We've stocked it as a chassis and components.", "tokens": [51053, 492, 600, 4127, 292, 309, 382, 257, 28262, 293, 6677, 13, 51149], "temperature": 0, "avg_logprob": -0.09107277311127761, "compression_ratio": 1.7333333333333334, "no_speech_prob": 8.910949235441112e-12}, {"id": 111, "seek": 56256, "start": 578.2399999999999, "end": 581.76, "text": " So then our picking team will they'll receive the order.", "tokens": [51149, 407, 550, 527, 8867, 1469, 486, 436, 603, 4774, 264, 1668, 13, 51325], "temperature": 0, "avg_logprob": -0.09107277311127761, "compression_ratio": 1.7333333333333334, "no_speech_prob": 8.910949235441112e-12}, {"id": 112, "seek": 56256, "start": 581.76, "end": 584.64, "text": " They'll have a list of all the components that need to be picked.", "tokens": [51325, 814, 603, 362, 257, 1329, 295, 439, 264, 6677, 300, 643, 281, 312, 6183, 13, 51469], "temperature": 0, "avg_logprob": -0.09107277311127761, "compression_ratio": 1.7333333333333334, "no_speech_prob": 8.910949235441112e-12}, {"id": 113, "seek": 56256, "start": 585.5999999999999, "end": 591.76, "text": " And then it's just the case of them assembling it all onto a trolley ready for the next part of the process.", "tokens": [51517, 400, 550, 309, 311, 445, 264, 1389, 295, 552, 43867, 309, 439, 3911, 257, 20680, 2030, 1919, 337, 264, 958, 644, 295, 264, 1399, 13, 51825], "temperature": 0, "avg_logprob": -0.09107277311127761, "compression_ratio": 1.7333333333333334, "no_speech_prob": 8.910949235441112e-12}, {"id": 114, "seek": 59176, "start": 591.76, "end": 594.4, "text": " It's actually very well organized.", "tokens": [50365, 467, 311, 767, 588, 731, 9983, 13, 50497], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 115, "seek": 59176, "start": 594.4, "end": 594.8, "text": " Yeah.", "tokens": [50497, 865, 13, 50517], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 116, "seek": 59176, "start": 594.8, "end": 597.76, "text": " Wherever I look, it's tidy.", "tokens": [50517, 30903, 286, 574, 11, 309, 311, 34646, 13, 50665], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 117, "seek": 59176, "start": 597.76, "end": 598.16, "text": " Yeah.", "tokens": [50665, 865, 13, 50685], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 118, "seek": 59176, "start": 598.16, "end": 599.92, "text": " Like look around me.", "tokens": [50685, 1743, 574, 926, 385, 13, 50773], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 119, "seek": 59176, "start": 599.92, "end": 602.8, "text": " There is nothing left to have.", "tokens": [50773, 821, 307, 1825, 1411, 281, 362, 13, 50917], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 120, "seek": 59176, "start": 602.8, "end": 603.68, "text": " Yeah.", "tokens": [50917, 865, 13, 50961], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 121, "seek": 59176, "start": 603.68, "end": 606.08, "text": " This is where we keep all the RAM and CPUs.", "tokens": [50961, 639, 307, 689, 321, 1066, 439, 264, 14561, 293, 13199, 82, 13, 51081], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 122, "seek": 59176, "start": 606.08, "end": 610.08, "text": " So there's a lot of value behind a closed door here.", "tokens": [51081, 407, 456, 311, 257, 688, 295, 2158, 2261, 257, 5395, 2853, 510, 13, 51281], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 123, "seek": 59176, "start": 610.08, "end": 611.4399999999999, "text": " So we just need to keep it safe.", "tokens": [51281, 407, 321, 445, 643, 281, 1066, 309, 3273, 13, 51349], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 124, "seek": 59176, "start": 612.64, "end": 618.88, "text": " Given the current situation with memory, it's refreshing to see how much DDR4 memory", "tokens": [51409, 18600, 264, 2190, 2590, 365, 4675, 11, 309, 311, 19772, 281, 536, 577, 709, 49272, 19, 4675, 51721], "temperature": 0, "avg_logprob": -0.17945052959300853, "compression_ratio": 1.508695652173913, "no_speech_prob": 9.67063061574347e-12}, {"id": 125, "seek": 61888, "start": 618.88, "end": 623.68, "text": " they've been able to stock at a pretty competitive price.", "tokens": [50365, 436, 600, 668, 1075, 281, 4127, 412, 257, 1238, 10043, 3218, 13, 50605], "temperature": 0, "avg_logprob": -0.07608597314179834, "compression_ratio": 1.5647058823529412, "no_speech_prob": 1.222252681010172e-11}, {"id": 126, "seek": 61888, "start": 624.96, "end": 630.72, "text": " So when a member of the picking team has picked an order, it goes onto a trolley", "tokens": [50669, 407, 562, 257, 4006, 295, 264, 8867, 1469, 575, 6183, 364, 1668, 11, 309, 1709, 3911, 257, 20680, 2030, 50957], "temperature": 0, "avg_logprob": -0.07608597314179834, "compression_ratio": 1.5647058823529412, "no_speech_prob": 1.222252681010172e-11}, {"id": 127, "seek": 61888, "start": 631.6, "end": 636.08, "text": " and then the trolley is moved through to assembly so that that particular order can be assembled.", "tokens": [51001, 293, 550, 264, 20680, 2030, 307, 4259, 807, 281, 12103, 370, 300, 300, 1729, 1668, 393, 312, 24204, 13, 51225], "temperature": 0, "avg_logprob": -0.07608597314179834, "compression_ratio": 1.5647058823529412, "no_speech_prob": 1.222252681010172e-11}, {"id": 128, "seek": 61888, "start": 637.2, "end": 638.8, "text": " So this is the assembly area.", "tokens": [51281, 407, 341, 307, 264, 12103, 1859, 13, 51361], "temperature": 0, "avg_logprob": -0.07608597314179834, "compression_ratio": 1.5647058823529412, "no_speech_prob": 1.222252681010172e-11}, {"id": 129, "seek": 63880, "start": 638.8, "end": 644.64, "text": " So you can see a member of staff is assembling an order that's been configured online.", "tokens": [50365, 407, 291, 393, 536, 257, 4006, 295, 3525, 307, 43867, 364, 1668, 300, 311, 668, 30538, 2950, 13, 50657], "temperature": 0, "avg_logprob": -0.08748803138732911, "compression_ratio": 1.5031055900621118, "no_speech_prob": 5.709337516646151e-12}, {"id": 130, "seek": 63880, "start": 644.64, "end": 647.8399999999999, "text": " And then when the machine is assembled, it then needs to be tested.", "tokens": [50657, 400, 550, 562, 264, 3479, 307, 24204, 11, 309, 550, 2203, 281, 312, 8246, 13, 50817], "temperature": 0, "avg_logprob": -0.08748803138732911, "compression_ratio": 1.5031055900621118, "no_speech_prob": 5.709337516646151e-12}, {"id": 131, "seek": 63880, "start": 650.24, "end": 655.4399999999999, "text": " It's clear that they put a lot of effort in testing the builds before sending them out.", "tokens": [50937, 467, 311, 1850, 300, 436, 829, 257, 688, 295, 4630, 294, 4997, 264, 15182, 949, 7750, 552, 484, 13, 51197], "temperature": 0, "avg_logprob": -0.08748803138732911, "compression_ratio": 1.5031055900621118, "no_speech_prob": 5.709337516646151e-12}, {"id": 132, "seek": 65544, "start": 655.44, "end": 661.6, "text": " And of course, they give people warranty, but also the care they put in packaging", "tokens": [50365, 400, 295, 1164, 11, 436, 976, 561, 26852, 11, 457, 611, 264, 1127, 436, 829, 294, 16836, 50673], "temperature": 0, "avg_logprob": -0.08504342541252215, "compression_ratio": 1.6, "no_speech_prob": 7.563223484996495e-12}, {"id": 133, "seek": 65544, "start": 661.6, "end": 665.6800000000001, "text": " to make sure that stuff gets shipped in the best possible way.", "tokens": [50673, 281, 652, 988, 300, 1507, 2170, 25312, 294, 264, 1151, 1944, 636, 13, 50877], "temperature": 0, "avg_logprob": -0.08504342541252215, "compression_ratio": 1.6, "no_speech_prob": 7.563223484996495e-12}, {"id": 134, "seek": 65544, "start": 665.6800000000001, "end": 669.44, "text": " We've invested a lot in the cardboard boxes we use, in the foam inserts,", "tokens": [50877, 492, 600, 13104, 257, 688, 294, 264, 22248, 9002, 321, 764, 11, 294, 264, 12958, 49163, 11, 51065], "temperature": 0, "avg_logprob": -0.08504342541252215, "compression_ratio": 1.6, "no_speech_prob": 7.563223484996495e-12}, {"id": 135, "seek": 65544, "start": 669.44, "end": 673.6800000000001, "text": " making sure that every server is catered for and that it's really secure inside.", "tokens": [51065, 1455, 988, 300, 633, 7154, 307, 21557, 292, 337, 293, 300, 309, 311, 534, 7144, 1854, 13, 51277], "temperature": 0, "avg_logprob": -0.08504342541252215, "compression_ratio": 1.6, "no_speech_prob": 7.563223484996495e-12}, {"id": 136, "seek": 65544, "start": 675.5200000000001, "end": 675.7600000000001, "text": " Yeah.", "tokens": [51369, 865, 13, 51381], "temperature": 0, "avg_logprob": -0.08504342541252215, "compression_ratio": 1.6, "no_speech_prob": 7.563223484996495e-12}, {"id": 137, "seek": 65544, "start": 675.7600000000001, "end": 681.7600000000001, "text": " So once the order has been packed, then we've got our outbound department here.", "tokens": [51381, 407, 1564, 264, 1668, 575, 668, 13265, 11, 550, 321, 600, 658, 527, 484, 18767, 5882, 510, 13, 51681], "temperature": 0, "avg_logprob": -0.08504342541252215, "compression_ratio": 1.6, "no_speech_prob": 7.563223484996495e-12}, {"id": 138, "seek": 68544, "start": 685.44, "end": 687.2800000000001, "text": " So we've got our outbound department here.", "tokens": [50365, 407, 321, 600, 658, 527, 484, 18767, 5882, 510, 13, 50457], "temperature": 0, "avg_logprob": -0.2732240576493113, "compression_ratio": 1.4663865546218486, "no_speech_prob": 6.031074511331225e-12}, {"id": 139, "seek": 68544, "start": 687.2800000000001, "end": 691.7600000000001, "text": " At the end of the warehouse tour, I spent some time with Toby Sheriff,", "tokens": [50457, 1711, 264, 917, 295, 264, 22244, 3512, 11, 286, 4418, 512, 565, 365, 40223, 32492, 11, 50681], "temperature": 0, "avg_logprob": -0.2732240576493113, "compression_ratio": 1.4663865546218486, "no_speech_prob": 6.031074511331225e-12}, {"id": 140, "seek": 68544, "start": 691.7600000000001, "end": 697.44, "text": " who's the lead engineer that actually put together the server build you see me configure", "tokens": [50681, 567, 311, 264, 1477, 11403, 300, 767, 829, 1214, 264, 7154, 1322, 291, 536, 385, 22162, 50965], "temperature": 0, "avg_logprob": -0.2732240576493113, "compression_ratio": 1.4663865546218486, "no_speech_prob": 6.031074511331225e-12}, {"id": 141, "seek": 68544, "start": 697.44, "end": 699.6800000000001, "text": " at the beginning of this video.", "tokens": [50965, 412, 264, 2863, 295, 341, 960, 13, 51077], "temperature": 0, "avg_logprob": -0.2732240576493113, "compression_ratio": 1.4663865546218486, "no_speech_prob": 6.031074511331225e-12}, {"id": 142, "seek": 68544, "start": 699.6800000000001, "end": 707.0400000000001, "text": " What we've got here is a Supermicro DGQ in CSE 118 chassis.", "tokens": [51077, 708, 321, 600, 658, 510, 307, 257, 4548, 13195, 340, 413, 38, 48, 294, 383, 5879, 2975, 23, 28262, 13, 51445], "temperature": 0, "avg_logprob": -0.2732240576493113, "compression_ratio": 1.4663865546218486, "no_speech_prob": 6.031074511331225e-12}, {"id": 143, "seek": 68544, "start": 707.0400000000001, "end": 711.7600000000001, "text": " It's got space for four double wide full height cards.", "tokens": [51445, 467, 311, 658, 1901, 337, 1451, 3834, 4874, 1577, 6681, 5632, 13, 51681], "temperature": 0, "avg_logprob": -0.2732240576493113, "compression_ratio": 1.4663865546218486, "no_speech_prob": 6.031074511331225e-12}, {"id": 144, "seek": 71176, "start": 711.76, "end": 716.88, "text": " It takes the scalable CPUs, Intel first and second generation.", "tokens": [50365, 467, 2516, 264, 38481, 13199, 82, 11, 19762, 700, 293, 1150, 5125, 13, 50621], "temperature": 0, "avg_logprob": -0.11921961961594303, "compression_ratio": 1.5957446808510638, "no_speech_prob": 2.553641109334648e-12}, {"id": 145, "seek": 71176, "start": 716.88, "end": 722.3199999999999, "text": " So currently we've got a gold 6150 or two gold 6150s in there.", "tokens": [50621, 407, 4362, 321, 600, 658, 257, 3821, 1386, 20120, 420, 732, 3821, 1386, 20120, 82, 294, 456, 13, 50893], "temperature": 0, "avg_logprob": -0.11921961961594303, "compression_ratio": 1.5957446808510638, "no_speech_prob": 2.553641109334648e-12}, {"id": 146, "seek": 71176, "start": 722.3199999999999, "end": 726.08, "text": " So I believe they're 18 core CPUs, but they're also threaded.", "tokens": [50893, 407, 286, 1697, 436, 434, 2443, 4965, 13199, 82, 11, 457, 436, 434, 611, 47493, 13, 51081], "temperature": 0, "avg_logprob": -0.11921961961594303, "compression_ratio": 1.5957446808510638, "no_speech_prob": 2.553641109334648e-12}, {"id": 147, "seek": 71176, "start": 726.08, "end": 728.4, "text": " So obviously you've got twice as many of them than that.", "tokens": [51081, 407, 2745, 291, 600, 658, 6091, 382, 867, 295, 552, 813, 300, 13, 51197], "temperature": 0, "avg_logprob": -0.11921961961594303, "compression_ratio": 1.5957446808510638, "no_speech_prob": 2.553641109334648e-12}, {"id": 148, "seek": 71176, "start": 728.4, "end": 732.56, "text": " We've also got some DDR4 registered DIMMs.", "tokens": [51197, 492, 600, 611, 658, 512, 49272, 19, 13968, 413, 6324, 26386, 13, 51405], "temperature": 0, "avg_logprob": -0.11921961961594303, "compression_ratio": 1.5957446808510638, "no_speech_prob": 2.553641109334648e-12}, {"id": 149, "seek": 71176, "start": 732.56, "end": 735.68, "text": " So it can take up to six per CPU.", "tokens": [51405, 407, 309, 393, 747, 493, 281, 2309, 680, 13199, 13, 51561], "temperature": 0, "avg_logprob": -0.11921961961594303, "compression_ratio": 1.5957446808510638, "no_speech_prob": 2.553641109334648e-12}, {"id": 150, "seek": 71176, "start": 735.68, "end": 741.2, "text": " Currently we've got 64 gigs in there in 16 gig DIMMs.", "tokens": [51561, 19964, 321, 600, 658, 12145, 34586, 294, 456, 294, 3165, 8741, 413, 6324, 26386, 13, 51837], "temperature": 0, "avg_logprob": -0.11921961961594303, "compression_ratio": 1.5957446808510638, "no_speech_prob": 2.553641109334648e-12}, {"id": 151, "seek": 74120, "start": 741.2, "end": 745.84, "text": " That's running at 2666 megahertz mega transfers.", "tokens": [50365, 663, 311, 2614, 412, 7551, 15237, 17986, 35655, 17986, 29137, 13, 50597], "temperature": 0, "avg_logprob": -0.09620900424021595, "compression_ratio": 1.516260162601626, "no_speech_prob": 3.613440709843152e-12}, {"id": 152, "seek": 74120, "start": 746.88, "end": 748.5600000000001, "text": " And that's just a limitation of the CPU.", "tokens": [50649, 400, 300, 311, 445, 257, 27432, 295, 264, 13199, 13, 50733], "temperature": 0, "avg_logprob": -0.09620900424021595, "compression_ratio": 1.516260162601626, "no_speech_prob": 3.613440709843152e-12}, {"id": 153, "seek": 74120, "start": 748.5600000000001, "end": 750.96, "text": " So there's not really anything you can do about that.", "tokens": [50733, 407, 456, 311, 406, 534, 1340, 291, 393, 360, 466, 300, 13, 50853], "temperature": 0, "avg_logprob": -0.09620900424021595, "compression_ratio": 1.516260162601626, "no_speech_prob": 3.613440709843152e-12}, {"id": 154, "seek": 74120, "start": 750.96, "end": 755.76, "text": " If you go to a second generation CPU, you'd get slightly faster at 2933.", "tokens": [50853, 759, 291, 352, 281, 257, 1150, 5125, 13199, 11, 291, 1116, 483, 4748, 4663, 412, 9413, 10191, 13, 51093], "temperature": 0, "avg_logprob": -0.09620900424021595, "compression_ratio": 1.516260162601626, "no_speech_prob": 3.613440709843152e-12}, {"id": 155, "seek": 74120, "start": 755.76, "end": 760.0, "text": " So it's not fully utilising the speed, but it shouldn't really be a problem.", "tokens": [51093, 407, 309, 311, 406, 4498, 4976, 3436, 264, 3073, 11, 457, 309, 4659, 380, 534, 312, 257, 1154, 13, 51305], "temperature": 0, "avg_logprob": -0.09620900424021595, "compression_ratio": 1.516260162601626, "no_speech_prob": 3.613440709843152e-12}, {"id": 156, "seek": 74120, "start": 760.0, "end": 763.2, "text": " We've also got on board 10 gig NICs.", "tokens": [51305, 492, 600, 611, 658, 322, 3150, 1266, 8741, 426, 2532, 82, 13, 51465], "temperature": 0, "avg_logprob": -0.09620900424021595, "compression_ratio": 1.516260162601626, "no_speech_prob": 3.613440709843152e-12}, {"id": 157, "seek": 74120, "start": 763.2, "end": 765.2800000000001, "text": " So that's integrated into the motherboard.", "tokens": [51465, 407, 300, 311, 10919, 666, 264, 32916, 13, 51569], "temperature": 0, "avg_logprob": -0.09620900424021595, "compression_ratio": 1.516260162601626, "no_speech_prob": 3.613440709843152e-12}, {"id": 158, "seek": 76528, "start": 765.28, "end": 767.68, "text": " We've also got two spare PCI slots at the rear.", "tokens": [50365, 492, 600, 611, 658, 732, 13798, 6465, 40, 24266, 412, 264, 8250, 13, 50485], "temperature": 0, "avg_logprob": -0.12150571117662404, "compression_ratio": 1.4076086956521738, "no_speech_prob": 5.687150403388408e-12}, {"id": 159, "seek": 76528, "start": 768.24, "end": 773.4399999999999, "text": " If you wanted to add a network card, maybe a PCIe storage drive or something.", "tokens": [50513, 759, 291, 1415, 281, 909, 257, 3209, 2920, 11, 1310, 257, 6465, 40, 68, 6725, 3332, 420, 746, 13, 50773], "temperature": 0, "avg_logprob": -0.12150571117662404, "compression_ratio": 1.4076086956521738, "no_speech_prob": 5.687150403388408e-12}, {"id": 160, "seek": 76528, "start": 774.0, "end": 778.3199999999999, "text": " Half of these three slots are CPU2 dependent.", "tokens": [50801, 15917, 295, 613, 1045, 24266, 366, 13199, 17, 12334, 13, 51017], "temperature": 0, "avg_logprob": -0.12150571117662404, "compression_ratio": 1.4076086956521738, "no_speech_prob": 5.687150403388408e-12}, {"id": 161, "seek": 76528, "start": 778.3199999999999, "end": 784.56, "text": " So you could, if you felt so inclined, just put two cards in it initially with one CPU.", "tokens": [51017, 407, 291, 727, 11, 498, 291, 2762, 370, 28173, 11, 445, 829, 732, 5632, 294, 309, 9105, 365, 472, 13199, 13, 51329], "temperature": 0, "avg_logprob": -0.12150571117662404, "compression_ratio": 1.4076086956521738, "no_speech_prob": 5.687150403388408e-12}, {"id": 162, "seek": 78456, "start": 784.56, "end": 789.5999999999999, "text": " And expand in the future if that was, you know, that was what you wanted to do.", "tokens": [50365, 400, 5268, 294, 264, 2027, 498, 300, 390, 11, 291, 458, 11, 300, 390, 437, 291, 1415, 281, 360, 13, 50617], "temperature": 0, "avg_logprob": -0.1562567376471185, "compression_ratio": 1.4655172413793103, "no_speech_prob": 3.1897783026035853e-12}, {"id": 163, "seek": 78456, "start": 789.5999999999999, "end": 791.04, "text": " Keep the cost down initially.", "tokens": [50617, 5527, 264, 2063, 760, 9105, 13, 50689], "temperature": 0, "avg_logprob": -0.1562567376471185, "compression_ratio": 1.4655172413793103, "no_speech_prob": 3.1897783026035853e-12}, {"id": 164, "seek": 78456, "start": 791.04, "end": 796.16, "text": " Storage wise, at the front, you've got two NVMe ports built into a backplane.", "tokens": [50689, 36308, 10829, 11, 412, 264, 1868, 11, 291, 600, 658, 732, 46512, 12671, 18160, 3094, 666, 257, 646, 36390, 13, 50945], "temperature": 0, "avg_logprob": -0.1562567376471185, "compression_ratio": 1.4655172413793103, "no_speech_prob": 3.1897783026035853e-12}, {"id": 165, "seek": 78456, "start": 796.16, "end": 798.64, "text": " So that would do U.2 drives.", "tokens": [50945, 407, 300, 576, 360, 624, 13, 17, 11754, 13, 51069], "temperature": 0, "avg_logprob": -0.1562567376471185, "compression_ratio": 1.4655172413793103, "no_speech_prob": 3.1897783026035853e-12}, {"id": 166, "seek": 78456, "start": 798.64, "end": 802.16, "text": " So two and a quarter inch NVMe drives.", "tokens": [51069, 407, 732, 293, 257, 6555, 7227, 46512, 12671, 11754, 13, 51245], "temperature": 0, "avg_logprob": -0.1562567376471185, "compression_ratio": 1.4655172413793103, "no_speech_prob": 3.1897783026035853e-12}, {"id": 167, "seek": 80216, "start": 802.16, "end": 807.12, "text": " And then we've also got two SATA slots just behind that backplane internally.", "tokens": [50365, 400, 550, 321, 600, 611, 658, 732, 31536, 32, 24266, 445, 2261, 300, 646, 36390, 19501, 13, 50613], "temperature": 0, "avg_logprob": -0.12749209227385344, "compression_ratio": 1.640495867768595, "no_speech_prob": 4.3267338496744134e-12}, {"id": 168, "seek": 80216, "start": 809.04, "end": 814.7199999999999, "text": " So because these are data center graphics cards, they're not designed like a consumer card with a fan.", "tokens": [50709, 407, 570, 613, 366, 1412, 3056, 11837, 5632, 11, 436, 434, 406, 4761, 411, 257, 9711, 2920, 365, 257, 3429, 13, 50993], "temperature": 0, "avg_logprob": -0.12749209227385344, "compression_ratio": 1.640495867768595, "no_speech_prob": 4.3267338496744134e-12}, {"id": 169, "seek": 80216, "start": 814.7199999999999, "end": 816.0799999999999, "text": " As you said, they're passively cooled.", "tokens": [50993, 1018, 291, 848, 11, 436, 434, 1320, 3413, 27491, 13, 51061], "temperature": 0, "avg_logprob": -0.12749209227385344, "compression_ratio": 1.640495867768595, "no_speech_prob": 4.3267338496744134e-12}, {"id": 170, "seek": 80216, "start": 816.8, "end": 822.4, "text": " So that has got to travel through the the the fin stack as it, you know, travels through the server.", "tokens": [51097, 407, 300, 575, 658, 281, 3147, 807, 264, 264, 264, 962, 8630, 382, 309, 11, 291, 458, 11, 19863, 807, 264, 7154, 13, 51377], "temperature": 0, "avg_logprob": -0.12749209227385344, "compression_ratio": 1.640495867768595, "no_speech_prob": 4.3267338496744134e-12}, {"id": 171, "seek": 80216, "start": 822.9599999999999, "end": 825.6, "text": " So you've got all these fans across the front here.", "tokens": [51405, 407, 291, 600, 658, 439, 613, 4499, 2108, 264, 1868, 510, 13, 51537], "temperature": 0, "avg_logprob": -0.12749209227385344, "compression_ratio": 1.640495867768595, "no_speech_prob": 4.3267338496744134e-12}, {"id": 172, "seek": 80216, "start": 826.48, "end": 827.12, "text": " I won't lie.", "tokens": [51581, 286, 1582, 380, 4544, 13, 51613], "temperature": 0, "avg_logprob": -0.12749209227385344, "compression_ratio": 1.640495867768595, "no_speech_prob": 4.3267338496744134e-12}, {"id": 173, "seek": 80216, "start": 827.8399999999999, "end": 828.48, "text": " It is loud.", "tokens": [51649, 467, 307, 6588, 13, 51681], "temperature": 0, "avg_logprob": -0.12749209227385344, "compression_ratio": 1.640495867768595, "no_speech_prob": 4.3267338496744134e-12}, {"id": 174, "seek": 82848, "start": 828.48, "end": 834.24, "text": " But it's got to be to be able to push enough air through to keep all these cards and your CPU", "tokens": [50365, 583, 309, 311, 658, 281, 312, 281, 312, 1075, 281, 2944, 1547, 1988, 807, 281, 1066, 439, 613, 5632, 293, 428, 13199, 50653], "temperature": 0, "avg_logprob": -0.12157415350278218, "compression_ratio": 1.52863436123348, "no_speech_prob": 7.271976857486928e-12}, {"id": 175, "seek": 82848, "start": 834.8000000000001, "end": 837.2, "text": " and even your power supplies or network cards at the back.", "tokens": [50681, 293, 754, 428, 1347, 11768, 420, 3209, 5632, 412, 264, 646, 13, 50801], "temperature": 0, "avg_logprob": -0.12157415350278218, "compression_ratio": 1.52863436123348, "no_speech_prob": 7.271976857486928e-12}, {"id": 176, "seek": 82848, "start": 837.2, "end": 837.44, "text": " Cool.", "tokens": [50801, 8561, 13, 50813], "temperature": 0, "avg_logprob": -0.12157415350278218, "compression_ratio": 1.52863436123348, "no_speech_prob": 7.271976857486928e-12}, {"id": 177, "seek": 82848, "start": 838.0, "end": 845.36, "text": " So right now we've got four P100s in here, but obviously we can swap any card that would fit in there.", "tokens": [50841, 407, 558, 586, 321, 600, 658, 1451, 430, 6879, 82, 294, 510, 11, 457, 2745, 321, 393, 18135, 604, 2920, 300, 576, 3318, 294, 456, 13, 51209], "temperature": 0, "avg_logprob": -0.12157415350278218, "compression_ratio": 1.52863436123348, "no_speech_prob": 7.271976857486928e-12}, {"id": 178, "seek": 82848, "start": 846.0, "end": 852.4, "text": " In the video, I'm going to show you the performance with the, I think I have it here,", "tokens": [51241, 682, 264, 960, 11, 286, 478, 516, 281, 855, 291, 264, 3389, 365, 264, 11, 286, 519, 286, 362, 309, 510, 11, 51561], "temperature": 0, "avg_logprob": -0.12157415350278218, "compression_ratio": 1.52863436123348, "no_speech_prob": 7.271976857486928e-12}, {"id": 179, "seek": 85240, "start": 852.4, "end": 859.68, "text": " with one of the MI-50s, so MI-25 from AMD.", "tokens": [50365, 365, 472, 295, 264, 13696, 12, 2803, 82, 11, 370, 13696, 12, 6074, 490, 34808, 13, 50729], "temperature": 0, "avg_logprob": -0.2798888466574929, "compression_ratio": 1.3705882352941177, "no_speech_prob": 5.7765658575958945e-12}, {"id": 180, "seek": 85240, "start": 861.4399999999999, "end": 865.12, "text": " And then maybe we are also going to try, is this the...", "tokens": [50817, 400, 550, 1310, 321, 366, 611, 516, 281, 853, 11, 307, 341, 264, 485, 51001], "temperature": 0, "avg_logprob": -0.2798888466574929, "compression_ratio": 1.3705882352941177, "no_speech_prob": 5.7765658575958945e-12}, {"id": 181, "seek": 85240, "start": 865.12, "end": 867.6, "text": " It's the V100, that one.", "tokens": [51001, 467, 311, 264, 691, 6879, 11, 300, 472, 13, 51125], "temperature": 0, "avg_logprob": -0.2798888466574929, "compression_ratio": 1.3705882352941177, "no_speech_prob": 5.7765658575958945e-12}, {"id": 182, "seek": 85240, "start": 867.6, "end": 870.72, "text": " The V100, which is going to perform so much better.", "tokens": [51125, 440, 691, 6879, 11, 597, 307, 516, 281, 2042, 370, 709, 1101, 13, 51281], "temperature": 0, "avg_logprob": -0.2798888466574929, "compression_ratio": 1.3705882352941177, "no_speech_prob": 5.7765658575958945e-12}, {"id": 183, "seek": 85240, "start": 871.1999999999999, "end": 873.84, "text": " And in fact, we'll probably do some benchmarks with that.", "tokens": [51305, 400, 294, 1186, 11, 321, 603, 1391, 360, 512, 43751, 365, 300, 13, 51437], "temperature": 0, "avg_logprob": -0.2798888466574929, "compression_ratio": 1.3705882352941177, "no_speech_prob": 5.7765658575958945e-12}, {"id": 184, "seek": 87384, "start": 873.84, "end": 878.1600000000001, "text": " So Toby is now closing the chassis, so we can power this on.", "tokens": [50365, 407, 40223, 307, 586, 10377, 264, 28262, 11, 370, 321, 393, 1347, 341, 322, 13, 50581], "temperature": 0, "avg_logprob": -0.18983727467210987, "compression_ratio": 1.5271739130434783, "no_speech_prob": 6.887511394548795e-12}, {"id": 185, "seek": 87384, "start": 878.1600000000001, "end": 882.24, "text": " And mostly I want to give you an idea of the noise levels.", "tokens": [50581, 400, 5240, 286, 528, 281, 976, 291, 364, 1558, 295, 264, 5658, 4358, 13, 50785], "temperature": 0, "avg_logprob": -0.18983727467210987, "compression_ratio": 1.5271739130434783, "no_speech_prob": 6.887511394548795e-12}, {"id": 186, "seek": 87384, "start": 882.24, "end": 885.6800000000001, "text": " When you first power it on, this is very noisy.", "tokens": [50785, 1133, 291, 700, 1347, 309, 322, 11, 341, 307, 588, 24518, 13, 50957], "temperature": 0, "avg_logprob": -0.18983727467210987, "compression_ratio": 1.5271739130434783, "no_speech_prob": 6.887511394548795e-12}, {"id": 187, "seek": 87384, "start": 885.6800000000001, "end": 891.52, "text": " And then it becomes a little bit quieter, but obviously it is still quite noisy.", "tokens": [50957, 400, 550, 309, 3643, 257, 707, 857, 43339, 11, 457, 2745, 309, 307, 920, 1596, 24518, 13, 51249], "temperature": 0, "avg_logprob": -0.18983727467210987, "compression_ratio": 1.5271739130434783, "no_speech_prob": 6.887511394548795e-12}, {"id": 188, "seek": 87384, "start": 891.52, "end": 893.6800000000001, "text": " And we need to be aware of that.", "tokens": [51249, 400, 321, 643, 281, 312, 3650, 295, 300, 13, 51357], "temperature": 0, "avg_logprob": -0.18983727467210987, "compression_ratio": 1.5271739130434783, "no_speech_prob": 6.887511394548795e-12}, {"id": 189, "seek": 89368, "start": 893.68, "end": 910.8, "text": " So I wanted to also show you quickly the specs of these three GPUs side by side,", "tokens": [50365, 407, 286, 1415, 281, 611, 855, 291, 2661, 264, 27911, 295, 613, 1045, 18407, 82, 1252, 538, 1252, 11, 51221], "temperature": 0, "avg_logprob": -0.16394289802102482, "compression_ratio": 1.2587412587412588, "no_speech_prob": 4.096966689515202e-12}, {"id": 190, "seek": 89368, "start": 910.8, "end": 918.4799999999999, "text": " along with more modern GPUs like the R9700 AI Pro, that I'm sure you've already seen on my channel,", "tokens": [51221, 2051, 365, 544, 4363, 18407, 82, 411, 264, 497, 24, 18197, 7318, 1705, 11, 300, 286, 478, 988, 291, 600, 1217, 1612, 322, 452, 2269, 11, 51605], "temperature": 0, "avg_logprob": -0.16394289802102482, "compression_ratio": 1.2587412587412588, "no_speech_prob": 4.096966689515202e-12}, {"id": 191, "seek": 91848, "start": 918.48, "end": 920.96, "text": " and also the RTX 5090.", "tokens": [50365, 293, 611, 264, 44573, 2625, 7771, 13, 50489], "temperature": 0, "avg_logprob": -0.06811688609958924, "compression_ratio": 1.5084745762711864, "no_speech_prob": 2.5543380344911215e-12}, {"id": 192, "seek": 91848, "start": 920.96, "end": 926.64, "text": " Now, during the comparison, keep in mind the considerable price difference.", "tokens": [50489, 823, 11, 1830, 264, 9660, 11, 1066, 294, 1575, 264, 24167, 3218, 2649, 13, 50773], "temperature": 0, "avg_logprob": -0.06811688609958924, "compression_ratio": 1.5084745762711864, "no_speech_prob": 2.5543380344911215e-12}, {"id": 193, "seek": 91848, "start": 926.64, "end": 935.12, "text": " I have put it here, and obviously for the 5090 and the R9700, I've had to estimate some brackets,", "tokens": [50773, 286, 362, 829, 309, 510, 11, 293, 2745, 337, 264, 2625, 7771, 293, 264, 497, 24, 18197, 11, 286, 600, 632, 281, 12539, 512, 26179, 11, 51197], "temperature": 0, "avg_logprob": -0.06811688609958924, "compression_ratio": 1.5084745762711864, "no_speech_prob": 2.5543380344911215e-12}, {"id": 194, "seek": 91848, "start": 935.12, "end": 937.52, "text": " because the price is very variable.", "tokens": [51197, 570, 264, 3218, 307, 588, 7006, 13, 51317], "temperature": 0, "avg_logprob": -0.06811688609958924, "compression_ratio": 1.5084745762711864, "no_speech_prob": 2.5543380344911215e-12}, {"id": 195, "seek": 91848, "start": 937.52, "end": 941.04, "text": " But for the other GPUs that you can get on bargain hardware,", "tokens": [51317, 583, 337, 264, 661, 18407, 82, 300, 291, 393, 483, 322, 34302, 8837, 11, 51493], "temperature": 0, "avg_logprob": -0.06811688609958924, "compression_ratio": 1.5084745762711864, "no_speech_prob": 2.5543380344911215e-12}, {"id": 196, "seek": 91848, "start": 941.04, "end": 945.6, "text": " I have put the price that you would get with the 10% discount.", "tokens": [51493, 286, 362, 829, 264, 3218, 300, 291, 576, 483, 365, 264, 1266, 4, 11635, 13, 51721], "temperature": 0, "avg_logprob": -0.06811688609958924, "compression_ratio": 1.5084745762711864, "no_speech_prob": 2.5543380344911215e-12}, {"id": 197, "seek": 94560, "start": 945.6, "end": 950.48, "text": " One of the most important things to consider is memory bandwidth.", "tokens": [50365, 1485, 295, 264, 881, 1021, 721, 281, 1949, 307, 4675, 23647, 13, 50609], "temperature": 0, "avg_logprob": -0.05274926222764052, "compression_ratio": 1.5509259259259258, "no_speech_prob": 3.0088867595395863e-12}, {"id": 198, "seek": 94560, "start": 950.48, "end": 957.44, "text": " During inference, generating each token requires reading a lot of data, such as model weights,", "tokens": [50609, 6842, 38253, 11, 17746, 1184, 14862, 7029, 3760, 257, 688, 295, 1412, 11, 1270, 382, 2316, 17443, 11, 50957], "temperature": 0, "avg_logprob": -0.05274926222764052, "compression_ratio": 1.5509259259259258, "no_speech_prob": 3.0088867595395863e-12}, {"id": 199, "seek": 94560, "start": 957.44, "end": 959.76, "text": " back and forth from memory.", "tokens": [50957, 646, 293, 5220, 490, 4675, 13, 51073], "temperature": 0, "avg_logprob": -0.05274926222764052, "compression_ratio": 1.5509259259259258, "no_speech_prob": 3.0088867595395863e-12}, {"id": 200, "seek": 94560, "start": 959.76, "end": 966.32, "text": " So the bandwidth here directly limits how many tokens per second you can generate and process.", "tokens": [51073, 407, 264, 23647, 510, 3838, 10406, 577, 867, 22667, 680, 1150, 291, 393, 8460, 293, 1399, 13, 51401], "temperature": 0, "avg_logprob": -0.05274926222764052, "compression_ratio": 1.5509259259259258, "no_speech_prob": 3.0088867595395863e-12}, {"id": 201, "seek": 94560, "start": 966.32, "end": 971.84, "text": " So the V100 stands out at 900 gigabytes per second,", "tokens": [51401, 407, 264, 691, 6879, 7382, 484, 412, 22016, 42741, 680, 1150, 11, 51677], "temperature": 0, "avg_logprob": -0.05274926222764052, "compression_ratio": 1.5509259259259258, "no_speech_prob": 3.0088867595395863e-12}, {"id": 202, "seek": 97184, "start": 971.84, "end": 981.2, "text": " which is actually higher than the more modern R9700, which has a memory bandwidth of 640 gigabytes per second.", "tokens": [50365, 597, 307, 767, 2946, 813, 264, 544, 4363, 497, 24, 18197, 11, 597, 575, 257, 4675, 23647, 295, 1386, 5254, 42741, 680, 1150, 13, 50833], "temperature": 0, "avg_logprob": -0.09550587580754206, "compression_ratio": 1.4585635359116023, "no_speech_prob": 3.863471608606117e-12}, {"id": 203, "seek": 97184, "start": 981.2, "end": 988.32, "text": " And the MI25 is significantly lower at 484 gigabytes per second.", "tokens": [50833, 400, 264, 13696, 6074, 307, 10591, 3126, 412, 11174, 19, 42741, 680, 1150, 13, 51189], "temperature": 0, "avg_logprob": -0.09550587580754206, "compression_ratio": 1.4585635359116023, "no_speech_prob": 3.863471608606117e-12}, {"id": 204, "seek": 97184, "start": 988.32, "end": 990.8000000000001, "text": " And that shows in the benchmarks.", "tokens": [51189, 400, 300, 3110, 294, 264, 43751, 13, 51313], "temperature": 0, "avg_logprob": -0.09550587580754206, "compression_ratio": 1.4585635359116023, "no_speech_prob": 3.863471608606117e-12}, {"id": 205, "seek": 97184, "start": 990.8000000000001, "end": 994.88, "text": " The other important factor is the data format support.", "tokens": [51313, 440, 661, 1021, 5952, 307, 264, 1412, 7877, 1406, 13, 51517], "temperature": 0, "avg_logprob": -0.09550587580754206, "compression_ratio": 1.4585635359116023, "no_speech_prob": 3.863471608606117e-12}, {"id": 206, "seek": 99488, "start": 994.88, "end": 1001.6, "text": " Modern GPU architectures support formats like BF16 and FP8 in a native way.", "tokens": [50365, 19814, 18407, 6331, 1303, 1406, 25879, 411, 363, 37, 6866, 293, 36655, 23, 294, 257, 8470, 636, 13, 50701], "temperature": 0, "avg_logprob": -0.08690806439048365, "compression_ratio": 1.4568527918781726, "no_speech_prob": 4.177340331285029e-12}, {"id": 207, "seek": 99488, "start": 1001.6, "end": 1006.08, "text": " And these are the standard precisions used by current LLMs.", "tokens": [50701, 400, 613, 366, 264, 3832, 4346, 4252, 1143, 538, 2190, 441, 43, 26386, 13, 50925], "temperature": 0, "avg_logprob": -0.08690806439048365, "compression_ratio": 1.4568527918781726, "no_speech_prob": 4.177340331285029e-12}, {"id": 208, "seek": 99488, "start": 1006.08, "end": 1010.4, "text": " None of these three older cards support these formats natively.", "tokens": [50925, 14492, 295, 613, 1045, 4906, 5632, 1406, 613, 25879, 8470, 356, 13, 51141], "temperature": 0, "avg_logprob": -0.08690806439048365, "compression_ratio": 1.4568527918781726, "no_speech_prob": 4.177340331285029e-12}, {"id": 209, "seek": 99488, "start": 1010.4, "end": 1018.32, "text": " They do support FP16, which is also a 16-bit format and uses the same amount of memory,", "tokens": [51141, 814, 360, 1406, 36655, 6866, 11, 597, 307, 611, 257, 3165, 12, 5260, 7877, 293, 4960, 264, 912, 2372, 295, 4675, 11, 51537], "temperature": 0, "avg_logprob": -0.08690806439048365, "compression_ratio": 1.4568527918781726, "no_speech_prob": 4.177340331285029e-12}, {"id": 210, "seek": 101832, "start": 1018.32, "end": 1025.3600000000001, "text": " but FP16 cannot represent as wide a range of values as BF16.", "tokens": [50365, 457, 36655, 6866, 2644, 2906, 382, 4874, 257, 3613, 295, 4190, 382, 363, 37, 6866, 13, 50717], "temperature": 0, "avg_logprob": -0.06211990606589395, "compression_ratio": 1.3832335329341316, "no_speech_prob": 3.448999549501841e-12}, {"id": 211, "seek": 101832, "start": 1025.3600000000001, "end": 1031.04, "text": " When an inference framework converts BF16 model weights to FP16,", "tokens": [50717, 1133, 364, 38253, 8388, 38874, 363, 37, 6866, 2316, 17443, 281, 36655, 6866, 11, 51001], "temperature": 0, "avg_logprob": -0.06211990606589395, "compression_ratio": 1.3832335329341316, "no_speech_prob": 3.448999549501841e-12}, {"id": 212, "seek": 101832, "start": 1031.04, "end": 1034.72, "text": " some values can overflow or lose accuracy.", "tokens": [51001, 512, 4190, 393, 37772, 420, 3624, 14170, 13, 51185], "temperature": 0, "avg_logprob": -0.06211990606589395, "compression_ratio": 1.3832335329341316, "no_speech_prob": 3.448999549501841e-12}, {"id": 213, "seek": 101832, "start": 1034.72, "end": 1039.92, "text": " And depending on the model, this can downgrade output quality.", "tokens": [51185, 400, 5413, 322, 264, 2316, 11, 341, 393, 760, 8692, 5598, 3125, 13, 51445], "temperature": 0, "avg_logprob": -0.06211990606589395, "compression_ratio": 1.3832335329341316, "no_speech_prob": 3.448999549501841e-12}, {"id": 214, "seek": 103992, "start": 1039.92, "end": 1047.76, "text": " One last difference to keep in mind is that the V100 is more expensive because it has Tensor Cores,", "tokens": [50365, 1485, 1036, 2649, 281, 1066, 294, 1575, 307, 300, 264, 691, 6879, 307, 544, 5124, 570, 309, 575, 34306, 383, 2706, 11, 50757], "temperature": 0, "avg_logprob": -0.08889940308361519, "compression_ratio": 1.4326530612244899, "no_speech_prob": 4.176488148377455e-12}, {"id": 215, "seek": 103992, "start": 1047.76, "end": 1054.64, "text": " which are essentially hardware parts optimized for matrix operations used in machine learning.", "tokens": [50757, 597, 366, 4476, 8837, 3166, 26941, 337, 8141, 7705, 1143, 294, 3479, 2539, 13, 51101], "temperature": 0, "avg_logprob": -0.08889940308361519, "compression_ratio": 1.4326530612244899, "no_speech_prob": 4.176488148377455e-12}, {"id": 216, "seek": 103992, "start": 1054.64, "end": 1059.76, "text": " And you'd find such support also on all the modern GPUs.", "tokens": [51101, 400, 291, 1116, 915, 1270, 1406, 611, 322, 439, 264, 4363, 18407, 82, 13, 51357], "temperature": 0, "avg_logprob": -0.08889940308361519, "compression_ratio": 1.4326530612244899, "no_speech_prob": 4.176488148377455e-12}, {"id": 217, "seek": 103992, "start": 1059.76, "end": 1068.0800000000002, "text": " The P100 and MI25 do not have any equivalent, and this will be clearly reflected in the benchmarks.", "tokens": [51357, 440, 430, 6879, 293, 13696, 6074, 360, 406, 362, 604, 10344, 11, 293, 341, 486, 312, 4448, 15502, 294, 264, 43751, 13, 51773], "temperature": 0, "avg_logprob": -0.08889940308361519, "compression_ratio": 1.4326530612244899, "no_speech_prob": 4.176488148377455e-12}, {"id": 218, "seek": 106808, "start": 1068.08, "end": 1073.84, "text": " So very quickly, I want to show you how to set up this server.", "tokens": [50365, 407, 588, 2661, 11, 286, 528, 281, 855, 291, 577, 281, 992, 493, 341, 7154, 13, 50653], "temperature": 0, "avg_logprob": -0.1100140056390872, "compression_ratio": 1.5223880597014925, "no_speech_prob": 3.907096868260851e-12}, {"id": 219, "seek": 106808, "start": 1073.84, "end": 1077.6799999999998, "text": " Right now, I've got the four V100s installed.", "tokens": [50653, 1779, 586, 11, 286, 600, 658, 264, 1451, 691, 6879, 82, 8899, 13, 50845], "temperature": 0, "avg_logprob": -0.1100140056390872, "compression_ratio": 1.5223880597014925, "no_speech_prob": 3.907096868260851e-12}, {"id": 220, "seek": 106808, "start": 1077.6799999999998, "end": 1085.04, "text": " So this is the repository I am going to use with the V100 AI toolboxes.", "tokens": [50845, 407, 341, 307, 264, 25841, 286, 669, 516, 281, 764, 365, 264, 691, 6879, 7318, 44593, 279, 13, 51213], "temperature": 0, "avg_logprob": -0.1100140056390872, "compression_ratio": 1.5223880597014925, "no_speech_prob": 3.907096868260851e-12}, {"id": 221, "seek": 106808, "start": 1085.04, "end": 1092.24, "text": " As usual, with all of my repositories, you are going to need to create a toolbox.", "tokens": [51213, 1018, 7713, 11, 365, 439, 295, 452, 22283, 2083, 11, 291, 366, 516, 281, 643, 281, 1884, 257, 44593, 13, 51573], "temperature": 0, "avg_logprob": -0.1100140056390872, "compression_ratio": 1.5223880597014925, "no_speech_prob": 3.907096868260851e-12}, {"id": 222, "seek": 106808, "start": 1092.24, "end": 1094.6399999999999, "text": " And this is essentially a Docker container.", "tokens": [51573, 400, 341, 307, 4476, 257, 33772, 10129, 13, 51693], "temperature": 0, "avg_logprob": -0.1100140056390872, "compression_ratio": 1.5223880597014925, "no_speech_prob": 3.907096868260851e-12}, {"id": 223, "seek": 109464, "start": 1094.64, "end": 1099.68, "text": " I'm not going to repeat myself. I explain about toolboxes, Podman and Docker containers", "tokens": [50365, 286, 478, 406, 516, 281, 7149, 2059, 13, 286, 2903, 466, 44593, 279, 11, 12646, 1601, 293, 33772, 17089, 50617], "temperature": 0, "avg_logprob": -0.10057148706345331, "compression_ratio": 1.625984251968504, "no_speech_prob": 2.7169963327799973e-12}, {"id": 224, "seek": 109464, "start": 1100.24, "end": 1102.96, "text": " in a lot of the other videos.", "tokens": [50645, 294, 257, 688, 295, 264, 661, 2145, 13, 50781], "temperature": 0, "avg_logprob": -0.10057148706345331, "compression_ratio": 1.625984251968504, "no_speech_prob": 2.7169963327799973e-12}, {"id": 225, "seek": 109464, "start": 1102.96, "end": 1107.92, "text": " You can choose between the CUDA backend and the Vulkan backend.", "tokens": [50781, 509, 393, 2826, 1296, 264, 29777, 7509, 38087, 293, 264, 41434, 5225, 38087, 13, 51029], "temperature": 0, "avg_logprob": -0.10057148706345331, "compression_ratio": 1.625984251968504, "no_speech_prob": 2.7169963327799973e-12}, {"id": 226, "seek": 109464, "start": 1107.92, "end": 1113.5200000000002, "text": " And CUDA is almost always going to work better for Nvidia cards.", "tokens": [51029, 400, 29777, 7509, 307, 1920, 1009, 516, 281, 589, 1101, 337, 46284, 5632, 13, 51309], "temperature": 0, "avg_logprob": -0.10057148706345331, "compression_ratio": 1.625984251968504, "no_speech_prob": 2.7169963327799973e-12}, {"id": 227, "seek": 109464, "start": 1113.5200000000002, "end": 1117.68, "text": " So we're going to take that one and we're going to create the toolbox,", "tokens": [51309, 407, 321, 434, 516, 281, 747, 300, 472, 293, 321, 434, 516, 281, 1884, 264, 44593, 11, 51517], "temperature": 0, "avg_logprob": -0.10057148706345331, "compression_ratio": 1.625984251968504, "no_speech_prob": 2.7169963327799973e-12}, {"id": 228, "seek": 109464, "start": 1117.68, "end": 1123.92, "text": " which is essentially going to connect to Docker Hub and pull this toolbox that I have pre-built", "tokens": [51517, 597, 307, 4476, 516, 281, 1745, 281, 33772, 18986, 293, 2235, 341, 44593, 300, 286, 362, 659, 12, 23018, 51829], "temperature": 0, "avg_logprob": -0.10057148706345331, "compression_ratio": 1.625984251968504, "no_speech_prob": 2.7169963327799973e-12}, {"id": 229, "seek": 112392, "start": 1123.92, "end": 1127.2, "text": " with Lama CPP. So let us do that.", "tokens": [50365, 365, 441, 2404, 383, 17755, 13, 407, 718, 505, 360, 300, 13, 50529], "temperature": 0, "avg_logprob": -0.20422033965587616, "compression_ratio": 1.3055555555555556, "no_speech_prob": 3.5592290833358353e-12}, {"id": 230, "seek": 112392, "start": 1128.48, "end": 1133.6000000000001, "text": " And actually, I have already created this toolbox, so I can enter it.", "tokens": [50593, 400, 767, 11, 286, 362, 1217, 2942, 341, 44593, 11, 370, 286, 393, 3242, 309, 13, 50849], "temperature": 0, "avg_logprob": -0.20422033965587616, "compression_ratio": 1.3055555555555556, "no_speech_prob": 3.5592290833358353e-12}, {"id": 231, "seek": 112392, "start": 1133.6000000000001, "end": 1136.0, "text": " Lama V100 CUDA.", "tokens": [50849, 441, 2404, 691, 6879, 29777, 7509, 13, 50969], "temperature": 0, "avg_logprob": -0.20422033965587616, "compression_ratio": 1.3055555555555556, "no_speech_prob": 3.5592290833358353e-12}, {"id": 232, "seek": 112392, "start": 1137.2, "end": 1145.3600000000001, "text": " Now, once you enter the toolbox, you can run Lama CLI, list devices.", "tokens": [51029, 823, 11, 1564, 291, 3242, 264, 44593, 11, 291, 393, 1190, 441, 2404, 12855, 40, 11, 1329, 5759, 13, 51437], "temperature": 0, "avg_logprob": -0.20422033965587616, "compression_ratio": 1.3055555555555556, "no_speech_prob": 3.5592290833358353e-12}, {"id": 233, "seek": 114536, "start": 1145.36, "end": 1154.4799999999998, "text": " And this essentially confirms that Lama via the CUDA backend can see the four V100s, each of them with 16", "tokens": [50365, 400, 341, 4476, 39982, 300, 441, 2404, 5766, 264, 29777, 7509, 38087, 393, 536, 264, 1451, 691, 6879, 82, 11, 1184, 295, 552, 365, 3165, 50821], "temperature": 0, "avg_logprob": -0.1304252059371383, "compression_ratio": 1.4086538461538463, "no_speech_prob": 2.3437313288049433e-12}, {"id": 234, "seek": 114536, "start": 1154.4799999999998, "end": 1155.6799999999998, "text": " gigabytes of VRAM.", "tokens": [50821, 42741, 295, 13722, 2865, 13, 50881], "temperature": 0, "avg_logprob": -0.1304252059371383, "compression_ratio": 1.4086538461538463, "no_speech_prob": 2.3437313288049433e-12}, {"id": 235, "seek": 114536, "start": 1155.6799999999998, "end": 1161.4399999999998, "text": " So now, if you want to run a model, the easiest way to do that is to download the model weights.", "tokens": [50881, 407, 586, 11, 498, 291, 528, 281, 1190, 257, 2316, 11, 264, 12889, 636, 281, 360, 300, 307, 281, 5484, 264, 2316, 17443, 13, 51169], "temperature": 0, "avg_logprob": -0.1304252059371383, "compression_ratio": 1.4086538461538463, "no_speech_prob": 2.3437313288049433e-12}, {"id": 236, "seek": 114536, "start": 1161.4399999999998, "end": 1167.1999999999998, "text": " I've already downloaded some of these in GGUF format from Hugging Face.", "tokens": [51169, 286, 600, 1217, 21748, 512, 295, 613, 294, 460, 32298, 37, 7877, 490, 46892, 3249, 4047, 13, 51457], "temperature": 0, "avg_logprob": -0.1304252059371383, "compression_ratio": 1.4086538461538463, "no_speech_prob": 2.3437313288049433e-12}, {"id": 237, "seek": 116720, "start": 1167.2, "end": 1175.44, "text": " And let's run, for example, one of the best models that we can have today for local agentic workflows,", "tokens": [50365, 400, 718, 311, 1190, 11, 337, 1365, 11, 472, 295, 264, 1151, 5245, 300, 321, 393, 362, 965, 337, 2654, 9461, 299, 43461, 11, 50777], "temperature": 0, "avg_logprob": -0.1852187645144579, "compression_ratio": 1.3282051282051281, "no_speech_prob": 2.7507792051129076e-12}, {"id": 238, "seek": 116720, "start": 1175.44, "end": 1179.3600000000001, "text": " which is QAN 3.6, 27 billion parameters.", "tokens": [50777, 597, 307, 1249, 1770, 805, 13, 21, 11, 7634, 5218, 9834, 13, 50973], "temperature": 0, "avg_logprob": -0.1852187645144579, "compression_ratio": 1.3282051282051281, "no_speech_prob": 2.7507792051129076e-12}, {"id": 239, "seek": 116720, "start": 1179.92, "end": 1184.72, "text": " And I think I have these in Q4 KXL quantization.", "tokens": [51001, 400, 286, 519, 286, 362, 613, 294, 1249, 19, 591, 55, 43, 4426, 2144, 13, 51241], "temperature": 0, "avg_logprob": -0.1852187645144579, "compression_ratio": 1.3282051282051281, "no_speech_prob": 2.7507792051129076e-12}, {"id": 240, "seek": 116720, "start": 1184.72, "end": 1187.3600000000001, "text": " So I'm just going to say Lama server.", "tokens": [51241, 407, 286, 478, 445, 516, 281, 584, 441, 2404, 7154, 13, 51373], "temperature": 0, "avg_logprob": -0.1852187645144579, "compression_ratio": 1.3282051282051281, "no_speech_prob": 2.7507792051129076e-12}, {"id": 241, "seek": 116720, "start": 1187.92, "end": 1190.4, "text": " I'm going to pass the model.", "tokens": [51401, 286, 478, 516, 281, 1320, 264, 2316, 13, 51525], "temperature": 0, "avg_logprob": -0.1852187645144579, "compression_ratio": 1.3282051282051281, "no_speech_prob": 2.7507792051129076e-12}, {"id": 242, "seek": 119040, "start": 1190.4, "end": 1196.4, "text": " Actually, I also have it in Q8 quantization, for a matter of fact.", "tokens": [50365, 5135, 11, 286, 611, 362, 309, 294, 1249, 23, 4426, 2144, 11, 337, 257, 1871, 295, 1186, 13, 50665], "temperature": 0, "avg_logprob": -0.1267563756306966, "compression_ratio": 1.4382022471910112, "no_speech_prob": 2.108370986478314e-12}, {"id": 243, "seek": 119040, "start": 1196.96, "end": 1203.76, "text": " And we are going to enable flash attention, and we can pass a context size.", "tokens": [50693, 400, 321, 366, 516, 281, 9528, 7319, 3202, 11, 293, 321, 393, 1320, 257, 4319, 2744, 13, 51033], "temperature": 0, "avg_logprob": -0.1267563756306966, "compression_ratio": 1.4382022471910112, "no_speech_prob": 2.108370986478314e-12}, {"id": 244, "seek": 119040, "start": 1204.96, "end": 1205.3600000000001, "text": " All right.", "tokens": [51093, 1057, 558, 13, 51113], "temperature": 0, "avg_logprob": -0.1267563756306966, "compression_ratio": 1.4382022471910112, "no_speech_prob": 2.108370986478314e-12}, {"id": 245, "seek": 119040, "start": 1205.3600000000001, "end": 1208.24, "text": " So the server is up and running on port 8080.", "tokens": [51113, 407, 264, 7154, 307, 493, 293, 2614, 322, 2436, 4688, 4702, 13, 51257], "temperature": 0, "avg_logprob": -0.1267563756306966, "compression_ratio": 1.4382022471910112, "no_speech_prob": 2.108370986478314e-12}, {"id": 246, "seek": 119040, "start": 1208.24, "end": 1212.16, "text": " And now I need to forward that port to my actual laptop.", "tokens": [51257, 400, 586, 286, 643, 281, 2128, 300, 2436, 281, 452, 3539, 10732, 13, 51453], "temperature": 0, "avg_logprob": -0.1267563756306966, "compression_ratio": 1.4382022471910112, "no_speech_prob": 2.108370986478314e-12}, {"id": 247, "seek": 121216, "start": 1212.16, "end": 1215.0400000000002, "text": " So I can SSH into the box.", "tokens": [50365, 407, 286, 393, 12238, 39, 666, 264, 2424, 13, 50509], "temperature": 0, "avg_logprob": -0.07598002413485913, "compression_ratio": 1.3878504672897196, "no_speech_prob": 2.7931023379584863e-12}, {"id": 248, "seek": 121216, "start": 1215.0400000000002, "end": 1218.24, "text": " I've called it BH for bargain hardware.", "tokens": [50509, 286, 600, 1219, 309, 40342, 337, 34302, 8837, 13, 50669], "temperature": 0, "avg_logprob": -0.07598002413485913, "compression_ratio": 1.3878504672897196, "no_speech_prob": 2.7931023379584863e-12}, {"id": 249, "seek": 121216, "start": 1218.24, "end": 1224.88, "text": " And I'm going to say that I want to forward that port to 8081 on my host,", "tokens": [50669, 400, 286, 478, 516, 281, 584, 300, 286, 528, 281, 2128, 300, 2436, 281, 4688, 32875, 322, 452, 3975, 11, 51001], "temperature": 0, "avg_logprob": -0.07598002413485913, "compression_ratio": 1.3878504672897196, "no_speech_prob": 2.7931023379584863e-12}, {"id": 250, "seek": 121216, "start": 1224.88, "end": 1228.8000000000002, "text": " because 8080 is already used by something else.", "tokens": [51001, 570, 4688, 4702, 307, 1217, 1143, 538, 746, 1646, 13, 51197], "temperature": 0, "avg_logprob": -0.07598002413485913, "compression_ratio": 1.3878504672897196, "no_speech_prob": 2.7931023379584863e-12}, {"id": 251, "seek": 121216, "start": 1228.8000000000002, "end": 1229.52, "text": " And there we go.", "tokens": [51197, 400, 456, 321, 352, 13, 51233], "temperature": 0, "avg_logprob": -0.07598002413485913, "compression_ratio": 1.3878504672897196, "no_speech_prob": 2.7931023379584863e-12}, {"id": 252, "seek": 121216, "start": 1229.52, "end": 1231.8400000000001, "text": " This is Lama CPP web UI.", "tokens": [51233, 639, 307, 441, 2404, 383, 17755, 3670, 15682, 13, 51349], "temperature": 0, "avg_logprob": -0.07598002413485913, "compression_ratio": 1.3878504672897196, "no_speech_prob": 2.7931023379584863e-12}, {"id": 253, "seek": 121216, "start": 1231.8400000000001, "end": 1240.96, "text": " So I can type a prompt, such as write a CUDA kernel to multiply to", "tokens": [51349, 407, 286, 393, 2010, 257, 12391, 11, 1270, 382, 2464, 257, 29777, 7509, 28256, 281, 12972, 281, 51805], "temperature": 0, "avg_logprob": -0.07598002413485913, "compression_ratio": 1.3878504672897196, "no_speech_prob": 2.7931023379584863e-12}, {"id": 254, "seek": 124216, "start": 1243.1200000000001, "end": 1245.28, "text": " As you can see now, it's thinking about it.", "tokens": [50413, 1018, 291, 393, 536, 586, 11, 309, 311, 1953, 466, 309, 13, 50521], "temperature": 0, "avg_logprob": -0.11042279243469239, "compression_ratio": 1.6008403361344539, "no_speech_prob": 1.7760908846073398e-12}, {"id": 255, "seek": 124216, "start": 1245.28, "end": 1249.44, "text": " When models spend a lot of time thinking,", "tokens": [50521, 1133, 5245, 3496, 257, 688, 295, 565, 1953, 11, 50729], "temperature": 0, "avg_logprob": -0.11042279243469239, "compression_ratio": 1.6008403361344539, "no_speech_prob": 1.7760908846073398e-12}, {"id": 256, "seek": 124216, "start": 1249.44, "end": 1254.0, "text": " and you can see it's going at around 31 tokens per second.", "tokens": [50729, 293, 291, 393, 536, 309, 311, 516, 412, 926, 10353, 22667, 680, 1150, 13, 50957], "temperature": 0, "avg_logprob": -0.11042279243469239, "compression_ratio": 1.6008403361344539, "no_speech_prob": 1.7760908846073398e-12}, {"id": 257, "seek": 124216, "start": 1254.0, "end": 1259.52, "text": " So for being a 27 billion parameter dense model, this is really good.", "tokens": [50957, 407, 337, 885, 257, 7634, 5218, 13075, 18011, 2316, 11, 341, 307, 534, 665, 13, 51233], "temperature": 0, "avg_logprob": -0.11042279243469239, "compression_ratio": 1.6008403361344539, "no_speech_prob": 1.7760908846073398e-12}, {"id": 258, "seek": 124216, "start": 1259.52, "end": 1262.88, "text": " But in a bit, we'll take a look at the proper benchmarks", "tokens": [51233, 583, 294, 257, 857, 11, 321, 603, 747, 257, 574, 412, 264, 2296, 43751, 51401], "temperature": 0, "avg_logprob": -0.11042279243469239, "compression_ratio": 1.6008403361344539, "no_speech_prob": 1.7760908846073398e-12}, {"id": 259, "seek": 124216, "start": 1262.88, "end": 1265.68, "text": " and what happens when the context size grows,", "tokens": [51401, 293, 437, 2314, 562, 264, 4319, 2744, 13156, 11, 51541], "temperature": 0, "avg_logprob": -0.11042279243469239, "compression_ratio": 1.6008403361344539, "no_speech_prob": 1.7760908846073398e-12}, {"id": 260, "seek": 124216, "start": 1265.68, "end": 1269.44, "text": " and obviously what happens with other models and quantizations.", "tokens": [51541, 293, 2745, 437, 2314, 365, 661, 5245, 293, 4426, 14455, 13, 51729], "temperature": 0, "avg_logprob": -0.11042279243469239, "compression_ratio": 1.6008403361344539, "no_speech_prob": 1.7760908846073398e-12}, {"id": 261, "seek": 126944, "start": 1269.44, "end": 1274.56, "text": " But I just thought I'd show you how to get one of these up and running.", "tokens": [50365, 583, 286, 445, 1194, 286, 1116, 855, 291, 577, 281, 483, 472, 295, 613, 493, 293, 2614, 13, 50621], "temperature": 0, "avg_logprob": -0.07379197059793675, "compression_ratio": 1.4608294930875576, "no_speech_prob": 2.7396251500028113e-12}, {"id": 262, "seek": 126944, "start": 1274.56, "end": 1280.56, "text": " Now, Lama CPP is my recommendation, especially for these older GPUs,", "tokens": [50621, 823, 11, 441, 2404, 383, 17755, 307, 452, 11879, 11, 2318, 337, 613, 4906, 18407, 82, 11, 50921], "temperature": 0, "avg_logprob": -0.07379197059793675, "compression_ratio": 1.4608294930875576, "no_speech_prob": 2.7396251500028113e-12}, {"id": 263, "seek": 126944, "start": 1280.56, "end": 1286.0, "text": " not just the V100, but also the P100 and the MI25.", "tokens": [50921, 406, 445, 264, 691, 6879, 11, 457, 611, 264, 430, 6879, 293, 264, 13696, 6074, 13, 51193], "temperature": 0, "avg_logprob": -0.07379197059793675, "compression_ratio": 1.4608294930875576, "no_speech_prob": 2.7396251500028113e-12}, {"id": 264, "seek": 126944, "start": 1286.0, "end": 1290.16, "text": " And I've got toolboxes and links to all of those.", "tokens": [51193, 400, 286, 600, 658, 44593, 279, 293, 6123, 281, 439, 295, 729, 13, 51401], "temperature": 0, "avg_logprob": -0.07379197059793675, "compression_ratio": 1.4608294930875576, "no_speech_prob": 2.7396251500028113e-12}, {"id": 265, "seek": 126944, "start": 1290.16, "end": 1296.64, "text": " And the reason for that is that Lama CPP has a much more uniform ecosystem,", "tokens": [51401, 400, 264, 1778, 337, 300, 307, 300, 441, 2404, 383, 17755, 575, 257, 709, 544, 9452, 11311, 11, 51725], "temperature": 0, "avg_logprob": -0.07379197059793675, "compression_ratio": 1.4608294930875576, "no_speech_prob": 2.7396251500028113e-12}, {"id": 266, "seek": 129664, "start": 1296.64, "end": 1299.1200000000001, "text": " and it works pretty much everywhere.", "tokens": [50365, 293, 309, 1985, 1238, 709, 5315, 13, 50489], "temperature": 0, "avg_logprob": -0.07902413083795916, "compression_ratio": 1.5985130111524164, "no_speech_prob": 2.175755018860026e-12}, {"id": 267, "seek": 129664, "start": 1299.1200000000001, "end": 1303.6000000000001, "text": " If your hardware is supported, you can run any model on it.", "tokens": [50489, 759, 428, 8837, 307, 8104, 11, 291, 393, 1190, 604, 2316, 322, 309, 13, 50713], "temperature": 0, "avg_logprob": -0.07902413083795916, "compression_ratio": 1.5985130111524164, "no_speech_prob": 2.175755018860026e-12}, {"id": 268, "seek": 129664, "start": 1303.6000000000001, "end": 1307.8400000000001, "text": " And it's got pretty much any model and any quantization.", "tokens": [50713, 400, 309, 311, 658, 1238, 709, 604, 2316, 293, 604, 4426, 2144, 13, 50925], "temperature": 0, "avg_logprob": -0.07902413083795916, "compression_ratio": 1.5985130111524164, "no_speech_prob": 2.175755018860026e-12}, {"id": 269, "seek": 129664, "start": 1307.8400000000001, "end": 1311.44, "text": " And quantizations are going to be really important because, of course, here,", "tokens": [50925, 400, 4426, 14455, 366, 516, 281, 312, 534, 1021, 570, 11, 295, 1164, 11, 510, 11, 51105], "temperature": 0, "avg_logprob": -0.07902413083795916, "compression_ratio": 1.5985130111524164, "no_speech_prob": 2.175755018860026e-12}, {"id": 270, "seek": 129664, "start": 1311.44, "end": 1313.6000000000001, "text": " you don't have that much memory available.", "tokens": [51105, 291, 500, 380, 362, 300, 709, 4675, 2435, 13, 51213], "temperature": 0, "avg_logprob": -0.07902413083795916, "compression_ratio": 1.5985130111524164, "no_speech_prob": 2.175755018860026e-12}, {"id": 271, "seek": 129664, "start": 1313.6000000000001, "end": 1317.0400000000002, "text": " I mean, in this case, you've got 64 gigabytes, which is quite a lot.", "tokens": [51213, 286, 914, 11, 294, 341, 1389, 11, 291, 600, 658, 12145, 42741, 11, 597, 307, 1596, 257, 688, 13, 51385], "temperature": 0, "avg_logprob": -0.07902413083795916, "compression_ratio": 1.5985130111524164, "no_speech_prob": 2.175755018860026e-12}, {"id": 272, "seek": 129664, "start": 1317.0400000000002, "end": 1321.92, "text": " But again, even if you want to run the QEM model that we just saw 27 billion parameters", "tokens": [51385, 583, 797, 11, 754, 498, 291, 528, 281, 1190, 264, 1249, 6683, 2316, 300, 321, 445, 1866, 7634, 5218, 9834, 51629], "temperature": 0, "avg_logprob": -0.07902413083795916, "compression_ratio": 1.5985130111524164, "no_speech_prob": 2.175755018860026e-12}, {"id": 273, "seek": 132192, "start": 1321.92, "end": 1324.96, "text": " is a lot, then you are going to need a quantization.", "tokens": [50365, 307, 257, 688, 11, 550, 291, 366, 516, 281, 643, 257, 4426, 2144, 13, 50517], "temperature": 0, "avg_logprob": -0.09752774470060774, "compression_ratio": 1.6164383561643836, "no_speech_prob": 2.2448217330134357e-12}, {"id": 274, "seek": 132192, "start": 1324.96, "end": 1330.0, "text": " However, I know that a lot of people really, really like VLLM.", "tokens": [50517, 2908, 11, 286, 458, 300, 257, 688, 295, 561, 534, 11, 534, 411, 691, 24010, 44, 13, 50769], "temperature": 0, "avg_logprob": -0.09752774470060774, "compression_ratio": 1.6164383561643836, "no_speech_prob": 2.2448217330134357e-12}, {"id": 275, "seek": 132192, "start": 1330.0, "end": 1335.92, "text": " So I've also created a VLLM toolbox that you can get up and running like this.", "tokens": [50769, 407, 286, 600, 611, 2942, 257, 691, 24010, 44, 44593, 300, 291, 393, 483, 493, 293, 2614, 411, 341, 13, 51065], "temperature": 0, "avg_logprob": -0.09752774470060774, "compression_ratio": 1.6164383561643836, "no_speech_prob": 2.2448217330134357e-12}, {"id": 276, "seek": 132192, "start": 1336.48, "end": 1340.88, "text": " So here, I've already created the toolbox so I can simply enter it.", "tokens": [51093, 407, 510, 11, 286, 600, 1217, 2942, 264, 44593, 370, 286, 393, 2935, 3242, 309, 13, 51313], "temperature": 0, "avg_logprob": -0.09752774470060774, "compression_ratio": 1.6164383561643836, "no_speech_prob": 2.2448217330134357e-12}, {"id": 277, "seek": 132192, "start": 1340.88, "end": 1347.52, "text": " And you will see that, obviously, on this server, I have also the VLLM toolbox for the P100", "tokens": [51313, 400, 291, 486, 536, 300, 11, 2745, 11, 322, 341, 7154, 11, 286, 362, 611, 264, 691, 24010, 44, 44593, 337, 264, 430, 6879, 51645], "temperature": 0, "avg_logprob": -0.09752774470060774, "compression_ratio": 1.6164383561643836, "no_speech_prob": 2.2448217330134357e-12}, {"id": 278, "seek": 134752, "start": 1347.52, "end": 1351.36, "text": " and the one for the AMD MI25.", "tokens": [50365, 293, 264, 472, 337, 264, 34808, 13696, 6074, 13, 50557], "temperature": 0, "avg_logprob": -0.09103355407714844, "compression_ratio": 1.407960199004975, "no_speech_prob": 2.0439368513675005e-12}, {"id": 279, "seek": 134752, "start": 1351.36, "end": 1354.8799999999999, "text": " But obviously, in this case, we enter this particular one.", "tokens": [50557, 583, 2745, 11, 294, 341, 1389, 11, 321, 3242, 341, 1729, 472, 13, 50733], "temperature": 0, "avg_logprob": -0.09103355407714844, "compression_ratio": 1.407960199004975, "no_speech_prob": 2.0439368513675005e-12}, {"id": 280, "seek": 134752, "start": 1354.8799999999999, "end": 1361.44, "text": " And what I do in my toolboxes for VLLM, I always give you a start VLLM script", "tokens": [50733, 400, 437, 286, 360, 294, 452, 44593, 279, 337, 691, 24010, 44, 11, 286, 1009, 976, 291, 257, 722, 691, 24010, 44, 5755, 51061], "temperature": 0, "avg_logprob": -0.09103355407714844, "compression_ratio": 1.407960199004975, "no_speech_prob": 2.0439368513675005e-12}, {"id": 281, "seek": 134752, "start": 1361.44, "end": 1363.84, "text": " with a list of models that I have tested.", "tokens": [51061, 365, 257, 1329, 295, 5245, 300, 286, 362, 8246, 13, 51181], "temperature": 0, "avg_logprob": -0.09103355407714844, "compression_ratio": 1.407960199004975, "no_speech_prob": 2.0439368513675005e-12}, {"id": 282, "seek": 134752, "start": 1364.56, "end": 1366.72, "text": " So at least you have a starting point.", "tokens": [51217, 407, 412, 1935, 291, 362, 257, 2891, 935, 13, 51325], "temperature": 0, "avg_logprob": -0.09103355407714844, "compression_ratio": 1.407960199004975, "no_speech_prob": 2.0439368513675005e-12}, {"id": 283, "seek": 134752, "start": 1366.72, "end": 1370.08, "text": " I use LAMA 3.1 just as a benchmark.", "tokens": [51325, 286, 764, 441, 38136, 805, 13, 16, 445, 382, 257, 18927, 13, 51493], "temperature": 0, "avg_logprob": -0.09103355407714844, "compression_ratio": 1.407960199004975, "no_speech_prob": 2.0439368513675005e-12}, {"id": 284, "seek": 137008, "start": 1370.08, "end": 1377.1999999999998, "text": " But then if you want to run some proper models here, you can see the QEM 3.6 family, the 27 billion", "tokens": [50365, 583, 550, 498, 291, 528, 281, 1190, 512, 2296, 5245, 510, 11, 291, 393, 536, 264, 1249, 6683, 805, 13, 21, 1605, 11, 264, 7634, 5218, 50721], "temperature": 0, "avg_logprob": -0.12075480355156792, "compression_ratio": 1.3714285714285714, "no_speech_prob": 1.8108428168420176e-12}, {"id": 285, "seek": 137008, "start": 1377.1999999999998, "end": 1377.84, "text": " parameter.", "tokens": [50721, 13075, 13, 50753], "temperature": 0, "avg_logprob": -0.12075480355156792, "compression_ratio": 1.3714285714285714, "no_speech_prob": 1.8108428168420176e-12}, {"id": 286, "seek": 137008, "start": 1377.84, "end": 1381.6799999999998, "text": " Now, this is in GPTQ 4-bit quantization.", "tokens": [50753, 823, 11, 341, 307, 294, 26039, 51, 48, 1017, 12, 5260, 4426, 2144, 13, 50945], "temperature": 0, "avg_logprob": -0.12075480355156792, "compression_ratio": 1.3714285714285714, "no_speech_prob": 1.8108428168420176e-12}, {"id": 287, "seek": 137008, "start": 1382.24, "end": 1389.04, "text": " You cannot run, or at least I haven't been able to find a way on VLLM to run AWQ quants,", "tokens": [50973, 509, 2644, 1190, 11, 420, 412, 1935, 286, 2378, 380, 668, 1075, 281, 915, 257, 636, 322, 691, 24010, 44, 281, 1190, 25815, 48, 421, 1719, 11, 51313], "temperature": 0, "avg_logprob": -0.12075480355156792, "compression_ratio": 1.3714285714285714, "no_speech_prob": 1.8108428168420176e-12}, {"id": 288, "seek": 137008, "start": 1389.04, "end": 1393.12, "text": " which are a little bit better activation aware.", "tokens": [51313, 597, 366, 257, 707, 857, 1101, 24433, 3650, 13, 51517], "temperature": 0, "avg_logprob": -0.12075480355156792, "compression_ratio": 1.3714285714285714, "no_speech_prob": 1.8108428168420176e-12}, {"id": 289, "seek": 139312, "start": 1393.12, "end": 1402.4799999999998, "text": " But again, this is why I tell people that LAMA CPP is 90% of the times better for most people,", "tokens": [50365, 583, 797, 11, 341, 307, 983, 286, 980, 561, 300, 441, 38136, 383, 17755, 307, 4289, 4, 295, 264, 1413, 1101, 337, 881, 561, 11, 50833], "temperature": 0, "avg_logprob": -0.07947630683581035, "compression_ratio": 1.4100418410041842, "no_speech_prob": 2.2273602233446876e-12}, {"id": 290, "seek": 139312, "start": 1402.4799999999998, "end": 1405.6, "text": " especially on older hardware.", "tokens": [50833, 2318, 322, 4906, 8837, 13, 50989], "temperature": 0, "avg_logprob": -0.07947630683581035, "compression_ratio": 1.4100418410041842, "no_speech_prob": 2.2273602233446876e-12}, {"id": 291, "seek": 139312, "start": 1405.6, "end": 1407.4399999999998, "text": " But anyway, I've got four GPUs.", "tokens": [50989, 583, 4033, 11, 286, 600, 658, 1451, 18407, 82, 13, 51081], "temperature": 0, "avg_logprob": -0.07947630683581035, "compression_ratio": 1.4100418410041842, "no_speech_prob": 2.2273602233446876e-12}, {"id": 292, "seek": 139312, "start": 1407.4399999999998, "end": 1415.28, "text": " So Tensor Parallelism 4, I can set concurrent requests, size of the context GPU utilization.", "tokens": [51081, 407, 34306, 3457, 336, 338, 1434, 1017, 11, 286, 393, 992, 37702, 12475, 11, 2744, 295, 264, 4319, 18407, 37074, 13, 51473], "temperature": 0, "avg_logprob": -0.07947630683581035, "compression_ratio": 1.4100418410041842, "no_speech_prob": 2.2273602233446876e-12}, {"id": 293, "seek": 139312, "start": 1415.28, "end": 1417.4399999999998, "text": " I don't remember why I set it this low.", "tokens": [51473, 286, 500, 380, 1604, 983, 286, 992, 309, 341, 2295, 13, 51581], "temperature": 0, "avg_logprob": -0.07947630683581035, "compression_ratio": 1.4100418410041842, "no_speech_prob": 2.2273602233446876e-12}, {"id": 294, "seek": 139312, "start": 1417.4399999999998, "end": 1419.52, "text": " You should probably set it a little bit higher.", "tokens": [51581, 509, 820, 1391, 992, 309, 257, 707, 857, 2946, 13, 51685], "temperature": 0, "avg_logprob": -0.07947630683581035, "compression_ratio": 1.4100418410041842, "no_speech_prob": 2.2273602233446876e-12}, {"id": 295, "seek": 141952, "start": 1419.52, "end": 1422.24, "text": " And then you can launch the server.", "tokens": [50365, 400, 550, 291, 393, 4025, 264, 7154, 13, 50501], "temperature": 0, "avg_logprob": -0.06680558811534534, "compression_ratio": 1.5958333333333334, "no_speech_prob": 2.2801593512616902e-12}, {"id": 296, "seek": 141952, "start": 1422.24, "end": 1427.76, "text": " And this should show you exactly how we are running that particular model.", "tokens": [50501, 400, 341, 820, 855, 291, 2293, 577, 321, 366, 2614, 300, 1729, 2316, 13, 50777], "temperature": 0, "avg_logprob": -0.06680558811534534, "compression_ratio": 1.5958333333333334, "no_speech_prob": 2.2801593512616902e-12}, {"id": 297, "seek": 141952, "start": 1427.76, "end": 1428.96, "text": " And it will take a little bit.", "tokens": [50777, 400, 309, 486, 747, 257, 707, 857, 13, 50837], "temperature": 0, "avg_logprob": -0.06680558811534534, "compression_ratio": 1.5958333333333334, "no_speech_prob": 2.2801593512616902e-12}, {"id": 298, "seek": 141952, "start": 1428.96, "end": 1433.52, "text": " And then VLLM will come up and you'll be able to use the model.", "tokens": [50837, 400, 550, 691, 24010, 44, 486, 808, 493, 293, 291, 603, 312, 1075, 281, 764, 264, 2316, 13, 51065], "temperature": 0, "avg_logprob": -0.06680558811534534, "compression_ratio": 1.5958333333333334, "no_speech_prob": 2.2801593512616902e-12}, {"id": 299, "seek": 141952, "start": 1433.52, "end": 1436.16, "text": " So here we can see the model loading.", "tokens": [51065, 407, 510, 321, 393, 536, 264, 2316, 15114, 13, 51197], "temperature": 0, "avg_logprob": -0.06680558811534534, "compression_ratio": 1.5958333333333334, "no_speech_prob": 2.2801593512616902e-12}, {"id": 300, "seek": 141952, "start": 1436.16, "end": 1438.4, "text": " It's taking quite a bit of time.", "tokens": [51197, 467, 311, 1940, 1596, 257, 857, 295, 565, 13, 51309], "temperature": 0, "avg_logprob": -0.06680558811534534, "compression_ratio": 1.5958333333333334, "no_speech_prob": 2.2801593512616902e-12}, {"id": 301, "seek": 141952, "start": 1438.4, "end": 1442.96, "text": " And that's basically one of the caveats with PCI 3.0.", "tokens": [51309, 400, 300, 311, 1936, 472, 295, 264, 11730, 1720, 365, 6465, 40, 805, 13, 15, 13, 51537], "temperature": 0, "avg_logprob": -0.06680558811534534, "compression_ratio": 1.5958333333333334, "no_speech_prob": 2.2801593512616902e-12}, {"id": 302, "seek": 141952, "start": 1442.96, "end": 1447.36, "text": " Especially model loading is going to be fairly slow.", "tokens": [51537, 8545, 2316, 15114, 307, 516, 281, 312, 6457, 2964, 13, 51757], "temperature": 0, "avg_logprob": -0.06680558811534534, "compression_ratio": 1.5958333333333334, "no_speech_prob": 2.2801593512616902e-12}, {"id": 303, "seek": 144736, "start": 1447.36, "end": 1453.1999999999998, "text": " And you can see here, it's also looking for different attention backhand that it can use.", "tokens": [50365, 400, 291, 393, 536, 510, 11, 309, 311, 611, 1237, 337, 819, 3202, 646, 5543, 300, 309, 393, 764, 13, 50657], "temperature": 0, "avg_logprob": -0.1260614300718402, "compression_ratio": 1.5299145299145298, "no_speech_prob": 2.2095511184594407e-12}, {"id": 304, "seek": 144736, "start": 1453.1999999999998, "end": 1460.3999999999999, "text": " And ultimately, I think it's falling back to PyTorch, SDPA attention, which is okay,", "tokens": [50657, 400, 6284, 11, 286, 519, 309, 311, 7440, 646, 281, 9953, 51, 284, 339, 11, 318, 11373, 32, 3202, 11, 597, 307, 1392, 11, 51017], "temperature": 0, "avg_logprob": -0.1260614300718402, "compression_ratio": 1.5299145299145298, "no_speech_prob": 2.2095511184594407e-12}, {"id": 305, "seek": 144736, "start": 1460.3999999999999, "end": 1465.36, "text": " but it's not as good as some of the things you can get on modern hardware, of course.", "tokens": [51017, 457, 309, 311, 406, 382, 665, 382, 512, 295, 264, 721, 291, 393, 483, 322, 4363, 8837, 11, 295, 1164, 13, 51265], "temperature": 0, "avg_logprob": -0.1260614300718402, "compression_ratio": 1.5299145299145298, "no_speech_prob": 2.2095511184594407e-12}, {"id": 306, "seek": 144736, "start": 1465.36, "end": 1473.28, "text": " But again, this will work and you'll be able to run some models, even with VLLM, if you use these", "tokens": [51265, 583, 797, 11, 341, 486, 589, 293, 291, 603, 312, 1075, 281, 1190, 512, 5245, 11, 754, 365, 691, 24010, 44, 11, 498, 291, 764, 613, 51661], "temperature": 0, "avg_logprob": -0.1260614300718402, "compression_ratio": 1.5299145299145298, "no_speech_prob": 2.2095511184594407e-12}, {"id": 307, "seek": 147328, "start": 1473.28, "end": 1474.56, "text": " tool boxes.", "tokens": [50365, 2290, 9002, 13, 50429], "temperature": 0, "avg_logprob": -0.13715739202017735, "compression_ratio": 1.4602510460251046, "no_speech_prob": 2.7615509705369856e-12}, {"id": 308, "seek": 147328, "start": 1474.56, "end": 1475.6, "text": " So here we go.", "tokens": [50429, 407, 510, 321, 352, 13, 50481], "temperature": 0, "avg_logprob": -0.13715739202017735, "compression_ratio": 1.4602510460251046, "no_speech_prob": 2.7615509705369856e-12}, {"id": 309, "seek": 147328, "start": 1475.6, "end": 1478.0, "text": " This is now up and running.", "tokens": [50481, 639, 307, 586, 493, 293, 2614, 13, 50601], "temperature": 0, "avg_logprob": -0.13715739202017735, "compression_ratio": 1.4602510460251046, "no_speech_prob": 2.7615509705369856e-12}, {"id": 310, "seek": 147328, "start": 1478.0, "end": 1482.24, "text": " And you can start using it, for example, with a coding agent.", "tokens": [50601, 400, 291, 393, 722, 1228, 309, 11, 337, 1365, 11, 365, 257, 17720, 9461, 13, 50813], "temperature": 0, "avg_logprob": -0.13715739202017735, "compression_ratio": 1.4602510460251046, "no_speech_prob": 2.7615509705369856e-12}, {"id": 311, "seek": 147328, "start": 1482.24, "end": 1487.84, "text": " But again, I would not recommend to use perhaps VLLM.", "tokens": [50813, 583, 797, 11, 286, 576, 406, 2748, 281, 764, 4317, 691, 24010, 44, 13, 51093], "temperature": 0, "avg_logprob": -0.13715739202017735, "compression_ratio": 1.4602510460251046, "no_speech_prob": 2.7615509705369856e-12}, {"id": 312, "seek": 147328, "start": 1487.84, "end": 1494.0, "text": " Just stick to Lama CPP and you're going to get probably the best performance here.", "tokens": [51093, 1449, 2897, 281, 441, 2404, 383, 17755, 293, 291, 434, 516, 281, 483, 1391, 264, 1151, 3389, 510, 13, 51401], "temperature": 0, "avg_logprob": -0.13715739202017735, "compression_ratio": 1.4602510460251046, "no_speech_prob": 2.7615509705369856e-12}, {"id": 313, "seek": 147328, "start": 1494.8, "end": 1500.8, "text": " Let's now take a look at the benchmarks, which are arguably one of the most important factor in", "tokens": [51441, 961, 311, 586, 747, 257, 574, 412, 264, 43751, 11, 597, 366, 26771, 472, 295, 264, 881, 1021, 5952, 294, 51741], "temperature": 0, "avg_logprob": -0.13715739202017735, "compression_ratio": 1.4602510460251046, "no_speech_prob": 2.7615509705369856e-12}, {"id": 314, "seek": 150080, "start": 1500.8, "end": 1504.8, "text": " deciding whether or not some of these cards might be good for you.", "tokens": [50365, 17990, 1968, 420, 406, 512, 295, 613, 5632, 1062, 312, 665, 337, 291, 13, 50565], "temperature": 0, "avg_logprob": -0.09721896831805889, "compression_ratio": 1.4166666666666667, "no_speech_prob": 2.5832064352165895e-12}, {"id": 315, "seek": 150080, "start": 1504.8, "end": 1510.32, "text": " Here, I've got the individual repositories for all the cards that I tested.", "tokens": [50565, 1692, 11, 286, 600, 658, 264, 2609, 22283, 2083, 337, 439, 264, 5632, 300, 286, 8246, 13, 50841], "temperature": 0, "avg_logprob": -0.09721896831805889, "compression_ratio": 1.4166666666666667, "no_speech_prob": 2.5832064352165895e-12}, {"id": 316, "seek": 150080, "start": 1510.32, "end": 1514.96, "text": " I should have the MI25 here as well.", "tokens": [50841, 286, 820, 362, 264, 13696, 6074, 510, 382, 731, 13, 51073], "temperature": 0, "avg_logprob": -0.09721896831805889, "compression_ratio": 1.4166666666666667, "no_speech_prob": 2.5832064352165895e-12}, {"id": 317, "seek": 150080, "start": 1514.96, "end": 1520.0, "text": " And I have put all of the benchmark results in the readme.", "tokens": [51073, 400, 286, 362, 829, 439, 295, 264, 18927, 3542, 294, 264, 1401, 1398, 13, 51325], "temperature": 0, "avg_logprob": -0.09721896831805889, "compression_ratio": 1.4166666666666667, "no_speech_prob": 2.5832064352165895e-12}, {"id": 318, "seek": 152000, "start": 1520.0, "end": 1528.16, "text": " So you can probably scroll and find them for Lama CPP and VLLM, token generation and prompt processing.", "tokens": [50365, 407, 291, 393, 1391, 11369, 293, 915, 552, 337, 441, 2404, 383, 17755, 293, 691, 24010, 44, 11, 14862, 5125, 293, 12391, 9007, 13, 50773], "temperature": 0, "avg_logprob": -0.09574057732099368, "compression_ratio": 1.5069124423963134, "no_speech_prob": 2.5238762902529688e-12}, {"id": 319, "seek": 152000, "start": 1528.16, "end": 1533.76, "text": " So this is the V100 and obviously the P100.", "tokens": [50773, 407, 341, 307, 264, 691, 6879, 293, 2745, 264, 430, 6879, 13, 51053], "temperature": 0, "avg_logprob": -0.09574057732099368, "compression_ratio": 1.5069124423963134, "no_speech_prob": 2.5238762902529688e-12}, {"id": 320, "seek": 152000, "start": 1533.76, "end": 1540.56, "text": " But obviously for this video to make things a little bit easier to compare, I just put everything together.", "tokens": [51053, 583, 2745, 337, 341, 960, 281, 652, 721, 257, 707, 857, 3571, 281, 6794, 11, 286, 445, 829, 1203, 1214, 13, 51393], "temperature": 0, "avg_logprob": -0.09574057732099368, "compression_ratio": 1.5069124423963134, "no_speech_prob": 2.5238762902529688e-12}, {"id": 321, "seek": 152000, "start": 1540.56, "end": 1546.72, "text": " You can see the comparison here of the three cards on different models.", "tokens": [51393, 509, 393, 536, 264, 9660, 510, 295, 264, 1045, 5632, 322, 819, 5245, 13, 51701], "temperature": 0, "avg_logprob": -0.09574057732099368, "compression_ratio": 1.5069124423963134, "no_speech_prob": 2.5238762902529688e-12}, {"id": 322, "seek": 154672, "start": 1546.72, "end": 1554.4, "text": " Now, these are modern LLMs that I recommend you run on these cards, mostly for agentic workflows and coding.", "tokens": [50365, 823, 11, 613, 366, 4363, 441, 43, 26386, 300, 286, 2748, 291, 1190, 322, 613, 5632, 11, 5240, 337, 9461, 299, 43461, 293, 17720, 13, 50749], "temperature": 0, "avg_logprob": -0.15984078739466293, "compression_ratio": 1.4675324675324675, "no_speech_prob": 2.2624261405285173e-12}, {"id": 323, "seek": 154672, "start": 1554.4, "end": 1564.24, "text": " The QAN 3.5 and 3.6 families are right now some of the best you can run with 32 to 64 gigabytes of RAM.", "tokens": [50749, 440, 1249, 1770, 805, 13, 20, 293, 805, 13, 21, 4466, 366, 558, 586, 512, 295, 264, 1151, 291, 393, 1190, 365, 8858, 281, 12145, 42741, 295, 14561, 13, 51241], "temperature": 0, "avg_logprob": -0.15984078739466293, "compression_ratio": 1.4675324675324675, "no_speech_prob": 2.2624261405285173e-12}, {"id": 324, "seek": 154672, "start": 1564.24, "end": 1573.92, "text": " And you can see obviously that the V100, the blue one here, it's the best performer at prompt processing and token generation.", "tokens": [51241, 400, 291, 393, 536, 2745, 300, 264, 691, 6879, 11, 264, 3344, 472, 510, 11, 309, 311, 264, 1151, 30248, 412, 12391, 9007, 293, 14862, 5125, 13, 51725], "temperature": 0, "avg_logprob": -0.15984078739466293, "compression_ratio": 1.4675324675324675, "no_speech_prob": 2.2624261405285173e-12}, {"id": 325, "seek": 157392, "start": 1573.92, "end": 1579.04, "text": " Even when the context goes to 32,000 tokens per second.", "tokens": [50365, 2754, 562, 264, 4319, 1709, 281, 8858, 11, 1360, 22667, 680, 1150, 13, 50621], "temperature": 0, "avg_logprob": -0.09677518208821614, "compression_ratio": 1.416184971098266, "no_speech_prob": 1.8465667363243288e-12}, {"id": 326, "seek": 157392, "start": 1579.04, "end": 1589.68, "text": " And that's because this is just a more modern architecture and it's got tensor cores, which essentially allow to do matrix multiplication much better.", "tokens": [50621, 400, 300, 311, 570, 341, 307, 445, 257, 544, 4363, 9482, 293, 309, 311, 658, 40863, 24826, 11, 597, 4476, 2089, 281, 360, 8141, 27290, 709, 1101, 13, 51153], "temperature": 0, "avg_logprob": -0.09677518208821614, "compression_ratio": 1.416184971098266, "no_speech_prob": 1.8465667363243288e-12}, {"id": 327, "seek": 157392, "start": 1589.68, "end": 1593.04, "text": " And that's the core operation in LLMs.", "tokens": [51153, 400, 300, 311, 264, 4965, 6916, 294, 441, 43, 26386, 13, 51321], "temperature": 0, "avg_logprob": -0.09677518208821614, "compression_ratio": 1.416184971098266, "no_speech_prob": 1.8465667363243288e-12}, {"id": 328, "seek": 159304, "start": 1593.04, "end": 1611.44, "text": " Just to give you an idea for a model like 27 billion parameters, you get on the V100, 852 tokens per second in prompt processing with the Q4 quantization, which is a very good quantization.", "tokens": [50365, 1449, 281, 976, 291, 364, 1558, 337, 257, 2316, 411, 7634, 5218, 9834, 11, 291, 483, 322, 264, 691, 6879, 11, 14695, 17, 22667, 680, 1150, 294, 12391, 9007, 365, 264, 1249, 19, 4426, 2144, 11, 597, 307, 257, 588, 665, 4426, 2144, 13, 51285], "temperature": 0, "avg_logprob": -0.13790820042292276, "compression_ratio": 1.2945205479452055, "no_speech_prob": 2.1174021737346838e-12}, {"id": 329, "seek": 161144, "start": 1611.44, "end": 1614.0800000000002, "text": " And that's a very good prompt processing speed.", "tokens": [50365, 400, 300, 311, 257, 588, 665, 12391, 9007, 3073, 13, 50497], "temperature": 0, "avg_logprob": -0.08885503584338773, "compression_ratio": 1.539877300613497, "no_speech_prob": 2.288619597654029e-12}, {"id": 330, "seek": 161144, "start": 1614.0800000000002, "end": 1617.68, "text": " And you get around 34 tokens per second.", "tokens": [50497, 400, 291, 483, 926, 12790, 22667, 680, 1150, 13, 50677], "temperature": 0, "avg_logprob": -0.08885503584338773, "compression_ratio": 1.539877300613497, "no_speech_prob": 2.288619597654029e-12}, {"id": 331, "seek": 161144, "start": 1617.68, "end": 1630.4, "text": " And even when you scale up the context, you see that you're still getting around 622 tokens per second on this 27 billion parameter model, which is a dense model.", "tokens": [50677, 400, 754, 562, 291, 4373, 493, 264, 4319, 11, 291, 536, 300, 291, 434, 920, 1242, 926, 1386, 7490, 22667, 680, 1150, 322, 341, 7634, 5218, 13075, 2316, 11, 597, 307, 257, 18011, 2316, 13, 51313], "temperature": 0, "avg_logprob": -0.08885503584338773, "compression_ratio": 1.539877300613497, "no_speech_prob": 2.288619597654029e-12}, {"id": 332, "seek": 163040, "start": 1630.4, "end": 1633.92, "text": " So these are the hardest model to run.", "tokens": [50365, 407, 613, 366, 264, 13158, 2316, 281, 1190, 13, 50541], "temperature": 0, "avg_logprob": -0.09485077390483782, "compression_ratio": 1.7226890756302522, "no_speech_prob": 2.5740843918181655e-12}, {"id": 333, "seek": 163040, "start": 1633.92, "end": 1636.8000000000002, "text": " And the performance is really good for these.", "tokens": [50541, 400, 264, 3389, 307, 534, 665, 337, 613, 13, 50685], "temperature": 0, "avg_logprob": -0.09485077390483782, "compression_ratio": 1.7226890756302522, "no_speech_prob": 2.5740843918181655e-12}, {"id": 334, "seek": 163040, "start": 1636.8000000000002, "end": 1642.88, "text": " And the tokens per second that you get in token generation is around 28 tokens per second.", "tokens": [50685, 400, 264, 22667, 680, 1150, 300, 291, 483, 294, 14862, 5125, 307, 926, 7562, 22667, 680, 1150, 13, 50989], "temperature": 0, "avg_logprob": -0.09485077390483782, "compression_ratio": 1.7226890756302522, "no_speech_prob": 2.5740843918181655e-12}, {"id": 335, "seek": 163040, "start": 1642.88, "end": 1649.2, "text": " Again, this would be perfectly useful if you were using this model, let's say, in the PI coding agent.", "tokens": [50989, 3764, 11, 341, 576, 312, 6239, 4420, 498, 291, 645, 1228, 341, 2316, 11, 718, 311, 584, 11, 294, 264, 27176, 17720, 9461, 13, 51305], "temperature": 0, "avg_logprob": -0.09485077390483782, "compression_ratio": 1.7226890756302522, "no_speech_prob": 2.5740843918181655e-12}, {"id": 336, "seek": 163040, "start": 1649.2, "end": 1657.2, "text": " Actually, check out the video I've done on coding agents and it's going to give you an idea of the type of performance you can get.", "tokens": [51305, 5135, 11, 1520, 484, 264, 960, 286, 600, 1096, 322, 17720, 12554, 293, 309, 311, 516, 281, 976, 291, 364, 1558, 295, 264, 2010, 295, 3389, 291, 393, 483, 13, 51705], "temperature": 0, "avg_logprob": -0.09485077390483782, "compression_ratio": 1.7226890756302522, "no_speech_prob": 2.5740843918181655e-12}, {"id": 337, "seek": 165720, "start": 1657.2, "end": 1661.52, "text": " So you will also see that the MI25 is missing from some of these.", "tokens": [50365, 407, 291, 486, 611, 536, 300, 264, 13696, 6074, 307, 5361, 490, 512, 295, 613, 13, 50581], "temperature": 0, "avg_logprob": -0.07852014489130142, "compression_ratio": 1.5677966101694916, "no_speech_prob": 3.0443659745915674e-12}, {"id": 338, "seek": 165720, "start": 1661.52, "end": 1664.64, "text": " It's just because it's the first card that I tested.", "tokens": [50581, 467, 311, 445, 570, 309, 311, 264, 700, 2920, 300, 286, 8246, 13, 50737], "temperature": 0, "avg_logprob": -0.07852014489130142, "compression_ratio": 1.5677966101694916, "no_speech_prob": 3.0443659745915674e-12}, {"id": 339, "seek": 165720, "start": 1664.64, "end": 1668.0800000000002, "text": " And back then, I didn't include all the models.", "tokens": [50737, 400, 646, 550, 11, 286, 994, 380, 4090, 439, 264, 5245, 13, 50909], "temperature": 0, "avg_logprob": -0.07852014489130142, "compression_ratio": 1.5677966101694916, "no_speech_prob": 3.0443659745915674e-12}, {"id": 340, "seek": 165720, "start": 1668.0800000000002, "end": 1677.44, "text": " But the ones that I included, you can see that on LLAMA CPP, it does perform close to the P100, but just below it.", "tokens": [50909, 583, 264, 2306, 300, 286, 5556, 11, 291, 393, 536, 300, 322, 441, 43, 38136, 383, 17755, 11, 309, 775, 2042, 1998, 281, 264, 430, 6879, 11, 457, 445, 2507, 309, 13, 51377], "temperature": 0, "avg_logprob": -0.07852014489130142, "compression_ratio": 1.5677966101694916, "no_speech_prob": 3.0443659745915674e-12}, {"id": 341, "seek": 165720, "start": 1677.44, "end": 1679.2, "text": " So be aware of that.", "tokens": [51377, 407, 312, 3650, 295, 300, 13, 51465], "temperature": 0, "avg_logprob": -0.07852014489130142, "compression_ratio": 1.5677966101694916, "no_speech_prob": 3.0443659745915674e-12}, {"id": 342, "seek": 165720, "start": 1679.2, "end": 1683.52, "text": " The MI25 is the one that's going to give you the least performance.", "tokens": [51465, 440, 13696, 6074, 307, 264, 472, 300, 311, 516, 281, 976, 291, 264, 1935, 3389, 13, 51681], "temperature": 0, "avg_logprob": -0.07852014489130142, "compression_ratio": 1.5677966101694916, "no_speech_prob": 3.0443659745915674e-12}, {"id": 343, "seek": 168352, "start": 1683.52, "end": 1687.52, "text": " So you can see that the other ones that you can see in the LLAMA CPP, but this is what you get on LLAMA CPP.", "tokens": [50365, 407, 291, 393, 536, 300, 264, 661, 2306, 300, 291, 393, 536, 294, 264, 441, 43, 38136, 383, 17755, 11, 457, 341, 307, 437, 291, 483, 322, 441, 43, 38136, 383, 17755, 13, 50565], "temperature": 0, "avg_logprob": -0.4354750044802402, "compression_ratio": 1.642512077294686, "no_speech_prob": 1.996579984675506e-12}, {"id": 344, "seek": 168352, "start": 1687.52, "end": 1697.2, "text": " You can also run VLLM, but with some caveats, VLLM is much more sensible to the different GPU architectures.", "tokens": [50565, 509, 393, 611, 1190, 691, 24010, 44, 11, 457, 365, 512, 11730, 1720, 11, 691, 24010, 44, 307, 709, 544, 25380, 281, 264, 819, 18407, 6331, 1303, 13, 51049], "temperature": 0, "avg_logprob": -0.4354750044802402, "compression_ratio": 1.642512077294686, "no_speech_prob": 1.996579984675506e-12}, {"id": 345, "seek": 168352, "start": 1697.2, "end": 1706.56, "text": " You require specific kernels for specific architectures, and some of them are just not available for a lot of these cards.", "tokens": [51049, 509, 3651, 2685, 23434, 1625, 337, 2685, 6331, 1303, 11, 293, 512, 295, 552, 366, 445, 406, 2435, 337, 257, 688, 295, 613, 5632, 13, 51517], "temperature": 0, "avg_logprob": -0.4354750044802402, "compression_ratio": 1.642512077294686, "no_speech_prob": 1.996579984675506e-12}, {"id": 346, "seek": 170656, "start": 1706.56, "end": 1712.3999999999999, "text": " So you are not really able to pick and choose like you do with LLAMA CPP and run anything you want.", "tokens": [50365, 407, 291, 366, 406, 534, 1075, 281, 1888, 293, 2826, 411, 291, 360, 365, 441, 43, 38136, 383, 17755, 293, 1190, 1340, 291, 528, 13, 50657], "temperature": 0, "avg_logprob": -0.08405840598930747, "compression_ratio": 1.325, "no_speech_prob": 1.8111158189490495e-12}, {"id": 347, "seek": 170656, "start": 1712.3999999999999, "end": 1724.56, "text": " But you can see that the MI25, when there are kernels available, actually performs better than the P100 on VLLM.", "tokens": [50657, 583, 291, 393, 536, 300, 264, 13696, 6074, 11, 562, 456, 366, 23434, 1625, 2435, 11, 767, 26213, 1101, 813, 264, 430, 6879, 322, 691, 24010, 44, 13, 51265], "temperature": 0, "avg_logprob": -0.08405840598930747, "compression_ratio": 1.325, "no_speech_prob": 1.8111158189490495e-12}, {"id": 348, "seek": 172456, "start": 1724.56, "end": 1727.28, "text": " So that's something to keep in mind.", "tokens": [50365, 407, 300, 311, 746, 281, 1066, 294, 1575, 13, 50501], "temperature": 0, "avg_logprob": -0.20879085765165442, "compression_ratio": 1.378787878787879, "no_speech_prob": 2.253949197422722e-12}, {"id": 349, "seek": 172456, "start": 1727.28, "end": 1732.1599999999999, "text": " Again, I do not recommend using VLLM with these cards.", "tokens": [50501, 3764, 11, 286, 360, 406, 2748, 1228, 691, 24010, 44, 365, 613, 5632, 13, 50745], "temperature": 0, "avg_logprob": -0.20879085765165442, "compression_ratio": 1.378787878787879, "no_speech_prob": 2.253949197422722e-12}, {"id": 350, "seek": 172456, "start": 1732.1599999999999, "end": 1735.6799999999998, "text": " A lot of models you just cannot easily run.", "tokens": [50745, 316, 688, 295, 5245, 291, 445, 2644, 3612, 1190, 13, 50921], "temperature": 0, "avg_logprob": -0.20879085765165442, "compression_ratio": 1.378787878787879, "no_speech_prob": 2.253949197422722e-12}, {"id": 351, "seek": 172456, "start": 1735.6799999999998, "end": 1744.8, "text": " But on the V100, at least, you can run some quantization of the QAN 3.6 27 billion parameter model.", "tokens": [50921, 583, 322, 264, 691, 6879, 11, 412, 1935, 11, 291, 393, 1190, 512, 4426, 2144, 295, 264, 1249, 1770, 805, 13, 21, 7634, 5218, 13075, 2316, 13, 51377], "temperature": 0, "avg_logprob": -0.20879085765165442, "compression_ratio": 1.378787878787879, "no_speech_prob": 2.253949197422722e-12}, {"id": 352, "seek": 172456, "start": 1745.2, "end": 1748.76, "text": " This is the GPT-Q 4-bit quantization.", "tokens": [51397, 639, 307, 264, 26039, 51, 12, 48, 1017, 12, 5260, 4426, 2144, 13, 51575], "temperature": 0, "avg_logprob": -0.20879085765165442, "compression_ratio": 1.378787878787879, "no_speech_prob": 2.253949197422722e-12}, {"id": 353, "seek": 174876, "start": 1748.76, "end": 1751.72, "text": " And you can see this throughput that you get.", "tokens": [50365, 400, 291, 393, 536, 341, 44629, 300, 291, 483, 13, 50513], "temperature": 0, "avg_logprob": -0.19040824271537163, "compression_ratio": 1.471264367816092, "no_speech_prob": 2.8155719943023794e-12}, {"id": 354, "seek": 174876, "start": 1751.72, "end": 1755.8799999999999, "text": " And you can also run the QAN 3.5 9 billion parameter model.", "tokens": [50513, 400, 291, 393, 611, 1190, 264, 1249, 1770, 805, 13, 20, 1722, 5218, 13075, 2316, 13, 50721], "temperature": 0, "avg_logprob": -0.19040824271537163, "compression_ratio": 1.471264367816092, "no_speech_prob": 2.8155719943023794e-12}, {"id": 355, "seek": 174876, "start": 1755.8799999999999, "end": 1759.56, "text": " But again, I do not recommend running any of these.", "tokens": [50721, 583, 797, 11, 286, 360, 406, 2748, 2614, 604, 295, 613, 13, 50905], "temperature": 0, "avg_logprob": -0.19040824271537163, "compression_ratio": 1.471264367816092, "no_speech_prob": 2.8155719943023794e-12}, {"id": 356, "seek": 174876, "start": 1759.56, "end": 1765.4, "text": " I did try these, and I have all the toolboxes if you want to experiment,", "tokens": [50905, 286, 630, 853, 613, 11, 293, 286, 362, 439, 264, 44593, 279, 498, 291, 528, 281, 5120, 11, 51197], "temperature": 0, "avg_logprob": -0.19040824271537163, "compression_ratio": 1.471264367816092, "no_speech_prob": 2.8155719943023794e-12}, {"id": 357, "seek": 174876, "start": 1765.84, "end": 1770.84, "text": " but probably stick to LLAMA CPP for these older GPUs.", "tokens": [51219, 457, 1391, 2897, 281, 441, 43, 38136, 383, 17755, 337, 613, 4906, 18407, 82, 13, 51469], "temperature": 0, "avg_logprob": -0.19040824271537163, "compression_ratio": 1.471264367816092, "no_speech_prob": 2.8155719943023794e-12}, {"id": 358, "seek": 174876, "start": 1771.76, "end": 1773.28, "text": " So we've seen the benchmarks.", "tokens": [51515, 407, 321, 600, 1612, 264, 43751, 13, 51591], "temperature": 0, "avg_logprob": -0.19040824271537163, "compression_ratio": 1.471264367816092, "no_speech_prob": 2.8155719943023794e-12}, {"id": 359, "seek": 174876, "start": 1773.56, "end": 1778.6, "text": " Now let's discuss the caveats and trade-offs you need to be aware of.", "tokens": [51605, 823, 718, 311, 2248, 264, 11730, 1720, 293, 4923, 12, 19231, 291, 643, 281, 312, 3650, 295, 13, 51857], "temperature": 0, "avg_logprob": -0.19040824271537163, "compression_ratio": 1.471264367816092, "no_speech_prob": 2.8155719943023794e-12}, {"id": 360, "seek": 177876, "start": 1778.76, "end": 1781.64, "text": " First, the GPU architectures.", "tokens": [50365, 2386, 11, 264, 18407, 6331, 1303, 13, 50509], "temperature": 0, "avg_logprob": -0.16674046516418456, "compression_ratio": 1.4460093896713615, "no_speech_prob": 2.3073836678128012e-12}, {"id": 361, "seek": 177876, "start": 1781.64, "end": 1792.12, "text": " The Pascal-based P100 and the Vega-based MI25 do not have hardware paths optimized for modern machine learning.", "tokens": [50509, 440, 41723, 12, 6032, 430, 6879, 293, 264, 48796, 12, 6032, 13696, 6074, 360, 406, 362, 8837, 14518, 26941, 337, 4363, 3479, 2539, 13, 51033], "temperature": 0, "avg_logprob": -0.16674046516418456, "compression_ratio": 1.4460093896713615, "no_speech_prob": 2.3073836678128012e-12}, {"id": 362, "seek": 177876, "start": 1792.12, "end": 1800.76, "text": " The V100 does have Tensor Cores, but it is still a legacy architecture compared to modern GPUs.", "tokens": [51033, 440, 691, 6879, 775, 362, 34306, 383, 2706, 11, 457, 309, 307, 920, 257, 11711, 9482, 5347, 281, 4363, 18407, 82, 13, 51465], "temperature": 0, "avg_logprob": -0.16674046516418456, "compression_ratio": 1.4460093896713615, "no_speech_prob": 2.3073836678128012e-12}, {"id": 363, "seek": 177876, "start": 1801.08, "end": 1805.48, "text": " Because of this, although these cards will be able to run modern LLMs,", "tokens": [51481, 1436, 295, 341, 11, 4878, 613, 5632, 486, 312, 1075, 281, 1190, 4363, 441, 43, 26386, 11, 51701], "temperature": 0, "avg_logprob": -0.16674046516418456, "compression_ratio": 1.4460093896713615, "no_speech_prob": 2.3073836678128012e-12}, {"id": 364, "seek": 180548, "start": 1805.48, "end": 1810.6, "text": " they just cannot match the throughput that a modern GPU gives you.", "tokens": [50365, 436, 445, 2644, 2995, 264, 44629, 300, 257, 4363, 18407, 2709, 291, 13, 50621], "temperature": 0, "avg_logprob": -0.1699386355520665, "compression_ratio": 1.476595744680851, "no_speech_prob": 3.1909325442364134e-12}, {"id": 365, "seek": 180548, "start": 1810.6, "end": 1813.32, "text": " And that's reflected in the price you pay.", "tokens": [50621, 400, 300, 311, 15502, 294, 264, 3218, 291, 1689, 13, 50757], "temperature": 0, "avg_logprob": -0.1699386355520665, "compression_ratio": 1.476595744680851, "no_speech_prob": 3.1909325442364134e-12}, {"id": 366, "seek": 180548, "start": 1813.8, "end": 1816.92, "text": " Second, these setups are noisy.", "tokens": [50781, 5736, 11, 613, 46832, 366, 24518, 13, 50937], "temperature": 0, "avg_logprob": -0.1699386355520665, "compression_ratio": 1.476595744680851, "no_speech_prob": 3.1909325442364134e-12}, {"id": 367, "seek": 180548, "start": 1817.32, "end": 1824.22, "text": " Because these passive cards require high airflow fans to stay cool, the system is loud.", "tokens": [50957, 1436, 613, 14975, 5632, 3651, 1090, 45291, 4499, 281, 1754, 1627, 11, 264, 1185, 307, 6588, 13, 51302], "temperature": 0, "avg_logprob": -0.1699386355520665, "compression_ratio": 1.476595744680851, "no_speech_prob": 3.1909325442364134e-12}, {"id": 368, "seek": 180548, "start": 1824.44, "end": 1827.5, "text": " You will not want this server under your desk.", "tokens": [51313, 509, 486, 406, 528, 341, 7154, 833, 428, 10026, 13, 51466], "temperature": 0, "avg_logprob": -0.1699386355520665, "compression_ratio": 1.476595744680851, "no_speech_prob": 3.1909325442364134e-12}, {"id": 369, "seek": 180548, "start": 1827.88, "end": 1832.58, "text": " This more likely belongs in a garage, a basement, or a dedicated room.", "tokens": [51485, 639, 544, 3700, 12953, 294, 257, 14400, 11, 257, 16893, 11, 420, 257, 8374, 1808, 13, 51720], "temperature": 0, "avg_logprob": -0.1699386355520665, "compression_ratio": 1.476595744680851, "no_speech_prob": 3.1909325442364134e-12}, {"id": 370, "seek": 183258, "start": 1832.58, "end": 1834.9199999999998, "text": " Third, the system bus.", "tokens": [50365, 12548, 11, 264, 1185, 1255, 13, 50482], "temperature": 0, "avg_logprob": -0.20595344664558532, "compression_ratio": 1.38, "no_speech_prob": 2.6244440648470757e-12}, {"id": 371, "seek": 183258, "start": 1834.9199999999998, "end": 1839.26, "text": " These older servers use PCIe Gen 3 slots.", "tokens": [50482, 1981, 4906, 15909, 764, 6465, 40, 68, 3632, 805, 24266, 13, 50699], "temperature": 0, "avg_logprob": -0.20595344664558532, "compression_ratio": 1.38, "no_speech_prob": 2.6244440648470757e-12}, {"id": 372, "seek": 183258, "start": 1839.6399999999999, "end": 1846.28, "text": " A PCIe Gen 3 slot has a theoretical bandwidth of around 16GB per second,", "tokens": [50718, 316, 6465, 40, 68, 3632, 805, 14747, 575, 257, 20864, 23647, 295, 926, 3165, 8769, 680, 1150, 11, 51050], "temperature": 0, "avg_logprob": -0.20595344664558532, "compression_ratio": 1.38, "no_speech_prob": 2.6244440648470757e-12}, {"id": 373, "seek": 183258, "start": 1846.6399999999999, "end": 1854.72, "text": " whereas a modern PCIe Gen 5 slot provides 64GB, four times the speed.", "tokens": [51068, 9735, 257, 4363, 6465, 40, 68, 3632, 1025, 14747, 6417, 12145, 8769, 11, 1451, 1413, 264, 3073, 13, 51472], "temperature": 0, "avg_logprob": -0.20595344664558532, "compression_ratio": 1.38, "no_speech_prob": 2.6244440648470757e-12}, {"id": 374, "seek": 185472, "start": 1854.72, "end": 1859.28, "text": " For single GPU setups, the difference is not that noticeable,", "tokens": [50365, 1171, 2167, 18407, 46832, 11, 264, 2649, 307, 406, 300, 26041, 11, 50593], "temperature": 0, "avg_logprob": -0.14014922595414958, "compression_ratio": 1.4432432432432432, "no_speech_prob": 4.7520221264918394e-12}, {"id": 375, "seek": 185472, "start": 1859.28, "end": 1865.82, "text": " only adding a few seconds when you load the model weights from storage into the GPU memory.", "tokens": [50593, 787, 5127, 257, 1326, 3949, 562, 291, 3677, 264, 2316, 17443, 490, 6725, 666, 264, 18407, 4675, 13, 50920], "temperature": 0, "avg_logprob": -0.14014922595414958, "compression_ratio": 1.4432432432432432, "no_speech_prob": 4.7520221264918394e-12}, {"id": 376, "seek": 185472, "start": 1866.14, "end": 1875.04, "text": " But if you run multi-GPU workloads, the GPUs must constantly exchange information and synchronize their activity.", "tokens": [50936, 583, 498, 291, 1190, 4825, 12, 38, 8115, 32452, 11, 264, 18407, 82, 1633, 6460, 7742, 1589, 293, 19331, 1125, 641, 5191, 13, 51381], "temperature": 0, "avg_logprob": -0.14014922595414958, "compression_ratio": 1.4432432432432432, "no_speech_prob": 4.7520221264918394e-12}, {"id": 377, "seek": 187504, "start": 1875.04, "end": 1885.8999999999999, "text": " So, over a Gen 3 bus, this exchange of information will be slower and it will limit the overall throughput and token generation performance.", "tokens": [50365, 407, 11, 670, 257, 3632, 805, 1255, 11, 341, 7742, 295, 1589, 486, 312, 14009, 293, 309, 486, 4948, 264, 4787, 44629, 293, 14862, 5125, 3389, 13, 50908], "temperature": 0, "avg_logprob": -0.1628827645745076, "compression_ratio": 1.4263959390862944, "no_speech_prob": 5.844971208424088e-12}, {"id": 378, "seek": 187504, "start": 1886.28, "end": 1888.32, "text": " The same is true for the memory.", "tokens": [50927, 440, 912, 307, 2074, 337, 264, 4675, 13, 51029], "temperature": 0, "avg_logprob": -0.1628827645745076, "compression_ratio": 1.4263959390862944, "no_speech_prob": 5.844971208424088e-12}, {"id": 379, "seek": 187504, "start": 1888.7, "end": 1898.24, "text": " The service runs on DDR4-ACC memory, which usually operates at around 2400 to 3200 megatransfer per second.", "tokens": [51048, 440, 2643, 6676, 322, 49272, 19, 12, 4378, 34, 4675, 11, 597, 2673, 22577, 412, 926, 4022, 628, 281, 805, 7629, 10816, 267, 25392, 612, 680, 1150, 13, 51525], "temperature": 0, "avg_logprob": -0.1628827645745076, "compression_ratio": 1.4263959390862944, "no_speech_prob": 5.844971208424088e-12}, {"id": 380, "seek": 189824, "start": 1898.24, "end": 1907.6200000000001, "text": " In comparison, DDR5 starts at 4800 megatransfer and goes much higher, effectively doubling the bandwidth.", "tokens": [50365, 682, 9660, 11, 49272, 20, 3719, 412, 11174, 628, 10816, 267, 25392, 612, 293, 1709, 709, 2946, 11, 8659, 33651, 264, 23647, 13, 50834], "temperature": 0, "avg_logprob": -0.09728164037068684, "compression_ratio": 1.4439024390243902, "no_speech_prob": 5.097007451521085e-12}, {"id": 381, "seek": 189824, "start": 1908.1, "end": 1915.06, "text": " This does not affect you much if you choose models that fully fit inside the GPU memory.", "tokens": [50858, 639, 775, 406, 3345, 291, 709, 498, 291, 2826, 5245, 300, 4498, 3318, 1854, 264, 18407, 4675, 13, 51206], "temperature": 0, "avg_logprob": -0.09728164037068684, "compression_ratio": 1.4439024390243902, "no_speech_prob": 5.097007451521085e-12}, {"id": 382, "seek": 189824, "start": 1915.22, "end": 1922.06, "text": " But if you want to do some offloading to the CPU and to the system memory, this will be a bottleneck.", "tokens": [51214, 583, 498, 291, 528, 281, 360, 512, 766, 2907, 278, 281, 264, 13199, 293, 281, 264, 1185, 4675, 11, 341, 486, 312, 257, 44641, 547, 13, 51556], "temperature": 0, "avg_logprob": -0.09728164037068684, "compression_ratio": 1.4439024390243902, "no_speech_prob": 5.097007451521085e-12}, {"id": 383, "seek": 192206, "start": 1922.06, "end": 1930.2, "text": " As I was editing this video, I realized that I was talking about PCIe 3.0 and DDR4.", "tokens": [50365, 1018, 286, 390, 10000, 341, 960, 11, 286, 5334, 300, 286, 390, 1417, 466, 6465, 40, 68, 805, 13, 15, 293, 49272, 19, 13, 50772], "temperature": 0, "avg_logprob": -0.10822368072251141, "compression_ratio": 1.391304347826087, "no_speech_prob": 4.639720031784922e-12}, {"id": 384, "seek": 192206, "start": 1930.2, "end": 1941.96, "text": " And I chose those for the workstations that we built here because I wanted to show you the cheapest available options to get a viable setup.", "tokens": [50772, 400, 286, 5111, 729, 337, 264, 589, 372, 763, 300, 321, 3094, 510, 570, 286, 1415, 281, 855, 291, 264, 29167, 2435, 3956, 281, 483, 257, 22024, 8657, 13, 51360], "temperature": 0, "avg_logprob": -0.10822368072251141, "compression_ratio": 1.391304347826087, "no_speech_prob": 4.639720031784922e-12}, {"id": 385, "seek": 194196, "start": 1941.96, "end": 1953.8600000000001, "text": " But if you go on bargain hardware, for example, you can find servers and workstations with DDR5 and PCIe 4 and 5.", "tokens": [50365, 583, 498, 291, 352, 322, 34302, 8837, 11, 337, 1365, 11, 291, 393, 915, 15909, 293, 589, 372, 763, 365, 49272, 20, 293, 6465, 40, 68, 1017, 293, 1025, 13, 50960], "temperature": 0, "avg_logprob": -0.13479739516528685, "compression_ratio": 1.3872832369942196, "no_speech_prob": 4.127829588557175e-12}, {"id": 386, "seek": 194196, "start": 1954.1200000000001, "end": 1967.04, "text": " Just to give you an example, these Dell Precision 7960, well, of course, more expensive, but these will come with DDR5 memory.", "tokens": [50973, 1449, 281, 976, 291, 364, 1365, 11, 613, 33319, 6001, 40832, 32803, 4550, 11, 731, 11, 295, 1164, 11, 544, 5124, 11, 457, 613, 486, 808, 365, 49272, 20, 4675, 13, 51619], "temperature": 0, "avg_logprob": -0.13479739516528685, "compression_ratio": 1.3872832369942196, "no_speech_prob": 4.127829588557175e-12}, {"id": 387, "seek": 196704, "start": 1967.04, "end": 1973.28, "text": " And it will also come with PCIe 5.0.", "tokens": [50365, 400, 309, 486, 611, 808, 365, 6465, 40, 68, 1025, 13, 15, 13, 50677], "temperature": 0, "avg_logprob": -0.18616245474134172, "compression_ratio": 1.3142857142857143, "no_speech_prob": 3.848392524791189e-12}, {"id": 388, "seek": 196704, "start": 1973.28, "end": 1979.76, "text": " So these are the main trade-offs you make, which allow you to keep the total cost at this price point.", "tokens": [50677, 407, 613, 366, 264, 2135, 4923, 12, 19231, 291, 652, 11, 597, 2089, 291, 281, 1066, 264, 3217, 2063, 412, 341, 3218, 935, 13, 51001], "temperature": 0, "avg_logprob": -0.18616245474134172, "compression_ratio": 1.3142857142857143, "no_speech_prob": 3.848392524791189e-12}, {"id": 389, "seek": 196704, "start": 1980.1, "end": 1982.6399999999999, "text": " Even with these caveats, the value is clear.", "tokens": [51018, 2754, 365, 613, 11730, 1720, 11, 264, 2158, 307, 1850, 13, 51145], "temperature": 0, "avg_logprob": -0.18616245474134172, "compression_ratio": 1.3142857142857143, "no_speech_prob": 3.848392524791189e-12}, {"id": 390, "seek": 198264, "start": 1982.64, "end": 1991.1000000000001, "text": " If you have a budget of 2,000 to 3,000 pounds and you need 64 gigabytes of VRAM to run certain models,", "tokens": [50365, 759, 291, 362, 257, 4706, 295, 568, 11, 1360, 281, 805, 11, 1360, 8319, 293, 291, 643, 12145, 42741, 295, 13722, 2865, 281, 1190, 1629, 5245, 11, 50788], "temperature": 0, "avg_logprob": -0.06733804151236293, "compression_ratio": 1.447488584474886, "no_speech_prob": 3.0916380566736734e-12}, {"id": 391, "seek": 198264, "start": 1991.5, "end": 1998.0400000000002, "text": " a refurbished enterprise server like the one I showed you is one of the real options available to you today.", "tokens": [50808, 257, 1895, 16659, 4729, 14132, 7154, 411, 264, 472, 286, 4712, 291, 307, 472, 295, 264, 957, 3956, 2435, 281, 291, 965, 13, 51135], "temperature": 0, "avg_logprob": -0.06733804151236293, "compression_ratio": 1.447488584474886, "no_speech_prob": 3.0916380566736734e-12}, {"id": 392, "seek": 198264, "start": 1998.3200000000002, "end": 2008.0800000000002, "text": " You get a pre-built base system that can host four cards for a fraction of the cost of a DIY workstation.", "tokens": [51149, 509, 483, 257, 659, 12, 23018, 3096, 1185, 300, 393, 3975, 1451, 5632, 337, 257, 14135, 295, 264, 2063, 295, 257, 22194, 589, 19159, 13, 51637], "temperature": 0, "avg_logprob": -0.06733804151236293, "compression_ratio": 1.447488584474886, "no_speech_prob": 3.0916380566736734e-12}, {"id": 393, "seek": 200808, "start": 2008.08, "end": 2018.08, "text": " You can even start with cheap cards like the MI25 and the P100 and then upgrade to faster GPUs when you are ready for that switch.", "tokens": [50365, 509, 393, 754, 722, 365, 7084, 5632, 411, 264, 13696, 6074, 293, 264, 430, 6879, 293, 550, 11484, 281, 4663, 18407, 82, 562, 291, 366, 1919, 337, 300, 3679, 13, 50865], "temperature": 0, "avg_logprob": -0.10311933581748706, "compression_ratio": 1.5281385281385282, "no_speech_prob": 3.1904750109196245e-12}, {"id": 394, "seek": 200808, "start": 2018.58, "end": 2027.72, "text": " As I mentioned at the start of the video, bargain hardware is offering a 10% discount on all of the GPUs that they offer at the link in the description.", "tokens": [50890, 1018, 286, 2835, 412, 264, 722, 295, 264, 960, 11, 34302, 8837, 307, 8745, 257, 1266, 4, 11635, 322, 439, 295, 264, 18407, 82, 300, 436, 2626, 412, 264, 2113, 294, 264, 3855, 13, 51347], "temperature": 0, "avg_logprob": -0.10311933581748706, "compression_ratio": 1.5281385281385282, "no_speech_prob": 3.1904750109196245e-12}, {"id": 395, "seek": 200808, "start": 2028.1599999999999, "end": 2033.52, "text": " So if you want to use that discount, just type Donato 10 at checkout.", "tokens": [51369, 407, 498, 291, 528, 281, 764, 300, 11635, 11, 445, 2010, 1468, 2513, 1266, 412, 37153, 13, 51637], "temperature": 0, "avg_logprob": -0.10311933581748706, "compression_ratio": 1.5281385281385282, "no_speech_prob": 3.1904750109196245e-12}, {"id": 396, "seek": 203352, "start": 2033.52, "end": 2038.6399999999999, "text": " Again, this is not an affiliate link and I get zero commission.", "tokens": [50365, 3764, 11, 341, 307, 406, 364, 23975, 2113, 293, 286, 483, 4018, 9221, 13, 50621], "temperature": 0, "avg_logprob": -0.08605766296386719, "compression_ratio": 1.5606694560669456, "no_speech_prob": 3.6016580343134486e-12}, {"id": 397, "seek": 203352, "start": 2038.98, "end": 2047.52, "text": " My goal is just to see what I can do to make hardware a bit more affordable for anyone looking to get started with local inference.", "tokens": [50638, 1222, 3387, 307, 445, 281, 536, 437, 286, 393, 360, 281, 652, 8837, 257, 857, 544, 12028, 337, 2878, 1237, 281, 483, 1409, 365, 2654, 38253, 13, 51065], "temperature": 0, "avg_logprob": -0.08605766296386719, "compression_ratio": 1.5606694560669456, "no_speech_prob": 3.6016580343134486e-12}, {"id": 398, "seek": 203352, "start": 2048.46, "end": 2058.4, "text": " So in the description of this video, you will also find all of the GitHub links to the toolboxes and configurations that I put together for the GPUs that I tested in this video.", "tokens": [51112, 407, 294, 264, 3855, 295, 341, 960, 11, 291, 486, 611, 915, 439, 295, 264, 23331, 6123, 281, 264, 44593, 279, 293, 31493, 300, 286, 829, 1214, 337, 264, 18407, 82, 300, 286, 8246, 294, 341, 960, 13, 51609], "temperature": 0, "avg_logprob": -0.08605766296386719, "compression_ratio": 1.5606694560669456, "no_speech_prob": 3.6016580343134486e-12}, {"id": 399, "seek": 205840, "start": 2058.4, "end": 2066.7200000000003, "text": " So to wrap up, as usual, remember this channel is a hobby project and a significant amount of time goes into making these videos.", "tokens": [50365, 407, 281, 7019, 493, 11, 382, 7713, 11, 1604, 341, 2269, 307, 257, 18240, 1716, 293, 257, 4776, 2372, 295, 565, 1709, 666, 1455, 613, 2145, 13, 50781], "temperature": 0, "avg_logprob": -0.11151992058267399, "compression_ratio": 1.5968992248062015, "no_speech_prob": 3.730486405895128e-12}, {"id": 400, "seek": 205840, "start": 2066.88, "end": 2074.48, "text": " If you find the work useful and want to support my research and me maintaining all of the different toolboxes and containers,", "tokens": [50789, 759, 291, 915, 264, 589, 4420, 293, 528, 281, 1406, 452, 2132, 293, 385, 14916, 439, 295, 264, 819, 44593, 279, 293, 17089, 11, 51169], "temperature": 0, "avg_logprob": -0.11151992058267399, "compression_ratio": 1.5968992248062015, "no_speech_prob": 3.730486405895128e-12}, {"id": 401, "seek": 205840, "start": 2075.0, "end": 2077.84, "text": " the link to support the channel is in the description.", "tokens": [51195, 264, 2113, 281, 1406, 264, 2269, 307, 294, 264, 3855, 13, 51337], "temperature": 0, "avg_logprob": -0.11151992058267399, "compression_ratio": 1.5968992248062015, "no_speech_prob": 3.730486405895128e-12}, {"id": 402, "seek": 205840, "start": 2078.2200000000003, "end": 2081.2000000000003, "text": " You can use Buy Me A Coffee to make a donation.", "tokens": [51356, 509, 393, 764, 19146, 1923, 316, 25481, 281, 652, 257, 19724, 13, 51505], "temperature": 0, "avg_logprob": -0.11151992058267399, "compression_ratio": 1.5968992248062015, "no_speech_prob": 3.730486405895128e-12}, {"id": 403, "seek": 205840, "start": 2081.7400000000002, "end": 2084.44, "text": " Thanks for watching and I'll see you in the next one.", "tokens": [51532, 2561, 337, 1976, 293, 286, 603, 536, 291, 294, 264, 958, 472, 13, 51667], "temperature": 0, "avg_logprob": -0.11151992058267399, "compression_ratio": 1.5968992248062015, "no_speech_prob": 3.730486405895128e-12}, {"id": 404, "seek": 208840, "start": 2088.4, "end": 2118.38, "text": " Thank you.", "tokens": [50365, 1044, 291, 13, 51864], "temperature": 0, "avg_logprob": -0.838227113087972, "compression_ratio": 0.5555555555555556, "no_speech_prob": 1.1097366948986664e-11}], "language": "en"}