WEBVTT

00:00:00.000 --> 00:00:05.740
So DeepSeek version 4 Pro is officially out today. Now you might be confused because this model was

00:00:05.740 --> 00:00:10.680
technically out but it was in general availability but this is the official release meaning that they

00:00:10.680 --> 00:00:15.200
have trained the model a little bit more and produced a stronger version of the model that

00:00:15.200 --> 00:00:19.900
is available today. Now if I were to summarize today's release in one simple sentence it is that

00:00:19.900 --> 00:00:25.000
the price to performance is becoming a very very important thing because not only do we have a new

00:00:25.000 --> 00:00:30.660
model from the DeepSeek team, the SpaceX team also dropped Grok 4.6 and both of these models

00:00:30.660 --> 00:00:36.520
are competing with the best frontier labs at a fraction of their cost. Let's start with Grok 4.6

00:00:36.520 --> 00:00:41.180
and one thing I'm going to say is that I'm genuinely surprised by the SpaceX team. I'm not

00:00:41.180 --> 00:00:45.860
trying to glaze them or Elon Musk or anything. I was just very critical of this lab. In 2025,

00:00:46.240 --> 00:00:51.380
they dropped Grok 4 and other models like that but I wasn't really you know mind blown with their

00:00:51.380 --> 00:00:56.240
performance because they're pretty they're pretty subpar compared to any of the other labs out there

00:00:56.240 --> 00:01:01.680
but in 2026 it looks like things have kind of changed a little because grok 4.5 was quite

00:01:01.680 --> 00:01:08.280
competitive based off what it cost and grok 4.6 is actually not too bad if we take a look at the

00:01:08.280 --> 00:01:13.300
benchmarks for example if we start with the artificial analysis intelligence index this

00:01:13.300 --> 00:01:20.080
model achieves a 61 and fable 5 is at 62 now what's really important to remember is that once again

00:01:20.080 --> 00:01:25.460
this model is quite cheap compared to Fable 5. This is about $2 per million input tokens

00:01:25.460 --> 00:01:31.900
and $6 per million output tokens. And Fable 5 sits at $10 per million input tokens and $50

00:01:31.900 --> 00:01:36.700
per million output tokens. And this is why you start to appreciate this release a little bit

00:01:36.700 --> 00:01:42.240
more. It might not beat the performance of the best models. It's matching them. Even GPT 5.6

00:01:42.240 --> 00:01:47.540
sold, which is a pretty capable model, and this is set at max on the artificial analysis index.

00:01:47.540 --> 00:01:54.660
this model achieves 61 and grok 4.6 61 so it ties it and it's much cheaper and then even on all of

00:01:54.660 --> 00:02:00.420
these other benchmarks for example the gdp valve one it actually beats fable 5 which is at 1741

00:02:00.420 --> 00:02:06.900
and then gpt 5.6 it's at 1728 i'm not sure why they didn't choose opus 5 as well but i guess

00:02:06.900 --> 00:02:11.460
they wanted to choose the quote-unquote strongest model lineup from each lab and they chose fable

00:02:11.460 --> 00:02:16.140
5 for anthropic which is fair and then deep software engineering one which is a critical

00:02:16.140 --> 00:02:23.720
benchmark this model doesn't beat fable 5 or gpt 5.6 soul but it gets close to it it's 65.9 and

00:02:23.720 --> 00:02:29.040
fable 5 sits at 70 but if you're getting results that are pretty close and the model is five times

00:02:29.040 --> 00:02:34.080
cheaper i wouldn't be too disappointed with this result and then same thing with cursor bench 3.2

00:02:34.080 --> 00:02:42.240
the model achieves 69.9 funny number and fable 5 sits at 70.5 so once again closer and grok 4.6

00:02:42.240 --> 00:02:48.200
beats GPT 5.6 soul. The same thing with the Frontier Code, it gets close to Fable 5 beats

00:02:48.200 --> 00:02:53.940
GPT 5.6 soul. So yeah, this model is actually available in Cursor. So the partnership with

00:02:53.940 --> 00:02:58.960
Cursor or I guess the acquisition has really helped SpaceX make some strides in the AI space

00:02:58.960 --> 00:03:04.120
this year. And Grok build is something that, you know, maybe not a lot of us have been using so

00:03:04.120 --> 00:03:09.480
far, but it's probably going to be another platform like Codex or Cloud Code that we start to use.

00:03:09.480 --> 00:03:14.580
but obviously cursor is quite strong as well so you have options available for you to use them in

00:03:14.580 --> 00:03:20.240
both and one thing to note is that they're offering two times usage inside grok build and cursor for

00:03:20.240 --> 00:03:25.320
the first week so if you just want to try it out see what you feel about it then you know it might

00:03:25.320 --> 00:03:29.760
be worth trying it out right now because you get double the usage in the first week now if we were

00:03:29.760 --> 00:03:34.520
to take a look at some of the outputs that people have been generating with grok 4.6 what we're

00:03:34.520 --> 00:03:40.180
looking at right now is a Falcon 9 booster return sequence simulation and this was done in a single

00:03:40.180 --> 00:03:47.620
HTML file and as I said if you expected to get this type of output from Grok in 2025 you would

00:03:47.620 --> 00:03:52.140
be kind of surprised because you wouldn't expect something like this to be generated with Grok but

00:03:52.140 --> 00:03:55.960
now it looks like we have to start taking the Grok team a little bit more serious because

00:03:55.960 --> 00:04:01.160
this output is quite competitive and obviously we're just looking at a simulation and we're just

00:04:01.160 --> 00:04:05.940
basing it off of a visual representation but if the model is able to produce something like this

00:04:05.940 --> 00:04:11.340
consistently then I would expect a lot of people to start adopting Grok because number one it is

00:04:11.340 --> 00:04:15.660
cheaper than the other labs at least at the moment because we don't know if this pricing strategy is

00:04:15.660 --> 00:04:20.360
going to be sustainable for the SpaceX team in the long run but at least for now their models are

00:04:20.360 --> 00:04:25.280
definitely cheaper compared to the others and as I mentioned this model excelled at the artificial

00:04:25.280 --> 00:04:32.580
analysis index. This model jumped to probably number four model. It's tied to GPT 5.6. It's

00:04:32.580 --> 00:04:36.560
pretty much similar. So you could say number three as well. But the models before that are

00:04:36.560 --> 00:04:43.540
Fable 5 and Opus 5, which are only above the model by about a 1% or a 2% difference. So yeah,

00:04:43.580 --> 00:04:48.640
even on this intelligence index, which if you're not familiar with, has nine evaluations. So on

00:04:48.640 --> 00:04:54.220
all of these evaluations, it's kind of matching almost Fable 5 performance, which is crazy to see.

00:04:54.220 --> 00:04:59.760
Before we continue, we just launched the Universe of AI newsletter. If you want to stay on top of

00:04:59.760 --> 00:05:05.060
AI news without having to hunt for it, link is in the description. Don't miss out. And what you see

00:05:05.060 --> 00:05:10.980
on screen right now is a racing game that Grok 4.6 built. And based off of this post, the model took

00:05:10.980 --> 00:05:15.780
about one minute and it was a five word prompt, which was a create a simple racing game in HTML.

00:05:16.280 --> 00:05:21.200
So if you're able to generate something like this easily using Grok 4.6, I think a lot of people

00:05:21.200 --> 00:05:27.160
will be happy. And this is a more detailed analysis of what it costs to run GPT 5.6 on

00:05:27.160 --> 00:05:32.120
the artificial analysis index and what it produced, meaning the output. Both of these models, if you

00:05:32.120 --> 00:05:38.120
remember, scored 61 on the artificial analysis index. Now to run the whole test with GPT 5.6,

00:05:38.480 --> 00:05:45.500
it costs about 2.8k. And then with Grok 4.6, it costs about 1.1k-ish. And this tells you that

00:05:45.500 --> 00:05:50.420
you're getting similar level of performance at half the cost. So yeah, this is a big release for

00:05:50.420 --> 00:05:54.280
the SpaceX team because they just proved once again that they are a lab that you seriously

00:05:54.280 --> 00:05:59.900
start into considering especially in 2026 and I'm going to talk more about DeepSeek version for Pro

00:05:59.900 --> 00:06:05.600
GA but basically what we're seeing today is that both of these releases kind of emphasize the fact

00:06:05.600 --> 00:06:10.520
that performance and all above that is the price at what you're getting for that performance is

00:06:10.520 --> 00:06:15.880
becoming more and more important for all users because we see many labs now focusing on creating

00:06:15.880 --> 00:06:21.380
the best model at the cheapest cost. Last year in 2025, most of the labs were just focused on,

00:06:21.440 --> 00:06:26.680
I would say, creating the strongest model. Yes, cost was important, but I think most of the times

00:06:26.680 --> 00:06:31.620
the frontier labs, meaning OpenAI Anthropic, were kind of more lenient on that fact because they

00:06:31.620 --> 00:06:38.020
didn't have as strong of a competition. Intelligent models that are maybe not always ahead of OpenAI

00:06:38.020 --> 00:06:42.560
Anthropic, but match their performance at a fraction of the cost. So yes, price to performance

00:06:42.560 --> 00:06:49.800
ratio is becoming a critical I would say indicator in 2026. Now this is the updated benchmark chart

00:06:49.800 --> 00:06:54.120
after the release of the new model and one thing you'll see across the board is that it matches

00:06:54.120 --> 00:06:58.360
the top level performance of many of the models. The one thing interesting over here is that they

00:06:58.360 --> 00:07:02.960
haven't put Opus 5 here for some reason. There is Fable 5 here that we can compare this model

00:07:02.960 --> 00:07:08.900
against but one thing you'll notice is that DeepSeek, remember this model costs 43.5 cents

00:07:08.900 --> 00:07:14.940
per million input tokens and 87 cents per million output tokens while the other models all over here

00:07:14.940 --> 00:07:20.140
are way more expensive than that. So the first thing if you look at Terminal Bench 2.1 the model

00:07:20.140 --> 00:07:26.126
scores 87.9. The older version of the model was 72.1 and the flash version was

00:07:26.126 --> 00:07:29.526
of the model was 72.1, and the Flash version, which we got last week, was 82.7. And what's

00:07:29.526 --> 00:07:37.146
crazy is that Fable 5 is 88. Yes, 88. So this model is only 0.1% behind Fable 5. And then on

00:07:37.146 --> 00:07:43.606
the Cyber Gym which is Cybersecurity the model actually beats Fable 5. Fable 5 sits at 83.1

00:07:43.606 --> 00:07:49.546
while DeepSeek version 4 Pro the new one that we got today is at 83.3. Now if this tells you

00:07:49.546 --> 00:07:54.706
something is that this model is once again really geared at Cybersecurity. We saw the Flash model

00:07:54.706 --> 00:07:58.966
also be geared towards Cybersecurity and becoming a model that was quite capable in that area

00:07:58.966 --> 00:08:04.986
and we're seeing the same thing today with the new version 4 Pro which sits at 83.3 outperforming

00:08:04.986 --> 00:08:10.486
Fable 5 on the benchmark. On deep software engineering this model is not beating Fable 5

00:08:10.486 --> 00:08:17.786
because Fable 5 sits at 70 but the model achieves 62.7 which is ahead of GLM 5.2 and a bit behind

00:08:17.786 --> 00:08:24.186
Kimi K3 which sits at 67.5 but one thing again this is a big jump compared to the preview version

00:08:24.186 --> 00:08:29.686
of the model which was at 12.8 so yeah they have really trained this model and have improved it on

00:08:29.686 --> 00:08:34.706
the back end because we can clearly see in this deep software engineering benchmark. And there's

00:08:34.706 --> 00:08:40.266
couple of other benchmarks but one key benchmark that this model excels at is the automation bench

00:08:40.266 --> 00:08:48.346
where the model achieves 31.8 and fable 5 29.1 so once again the model outperforms fable 5 now this

00:08:48.346 --> 00:08:53.166
is a fraction of the cost as i mentioned this is 43 cents and 87 cents for input and output

00:08:53.166 --> 00:08:59.606
respectively versus fable 5 which is at 10.50 so yes if we look at the terminal bench for example

00:08:59.606 --> 00:09:05.106
we are seeing this model achieve a result which is 0.1 behind the best model out there fable 5

00:09:05.106 --> 00:09:12.766
and the pricing is insane to look at 43 cents versus 10 dollars 87 cents versus 50 dollars per

00:09:12.766 --> 00:09:19.446
million input tokens so when we turn that into a fraction this model is 57 times cheaper and earlier

00:09:19.446 --> 00:09:24.586
last week i think i made a video about how deep seek version 4 pro was going to achieve a result

00:09:24.586 --> 00:09:28.866
like this because there was an investor report that leaked and when i saw the pricing where it

00:09:28.866 --> 00:09:34.206
said that it was going to match Fable 5 or even outperform it and be at 57 times cheaper per cost.

00:09:34.366 --> 00:09:38.906
When I was reading their investor report, I didn't really believe it. But today it is proven because

00:09:38.906 --> 00:09:43.786
at least on the benchmark so far, yes, I'm just saying the benchmarks, we are seeing similar level

00:09:43.786 --> 00:09:49.086
performance. Now, if this model actually performs like that in production, it's a little too early

00:09:49.086 --> 00:09:53.506
to tell yet. We still would have to give a couple of weeks and see how it's performing in the long

00:09:53.506 --> 00:09:58.206
run. Because sometimes on the release date, the models perform quite good. But over time,

00:09:58.206 --> 00:10:02.406
they kind of deteriorate in their quality. So we hope DeepSeek doesn't do that. Historically,

00:10:02.486 --> 00:10:06.726
they haven't done that. But let's just see. Just going to put it out there because right now we're

00:10:06.726 --> 00:10:11.886
basing this off of benchmarks. But that's it for today's video. Make sure you guys are subscribed

00:10:11.886 --> 00:10:17.666
to the channel. Follow our new newsletter as well at universeofai.beehive.com. As well,

00:10:17.726 --> 00:10:22.906
subscribe to the main channel, World of AI, and support us on X by following the Universe of AIZ

00:10:22.906 --> 00:10:25.706
as well. Until then, I'll see you guys in the next video.
