WEBVTT

00:00:00.080 --> 00:00:02.639
So, Deepseek version 4 Pro is officially

00:00:02.639 --> 00:00:04.799
out today. Now, you might be confused

00:00:04.799 --> 00:00:06.560
because this model was technically out,

00:00:06.560 --> 00:00:08.800
but it was in general availability, but

00:00:08.800 --> 00:00:10.400
this is the official release, meaning

00:00:10.400 --> 00:00:12.240
that they have trained the model a

00:00:12.240 --> 00:00:14.240
little bit more and produced a stronger

00:00:14.240 --> 00:00:16.000
version of the model that is available

00:00:16.000 --> 00:00:17.520
today. Now, if I were to summarize

00:00:17.520 --> 00:00:19.520
today's release in one simple sentence,

00:00:19.520 --> 00:00:21.520
it is that the price to performance is

00:00:21.520 --> 00:00:23.760
becoming a very, very important thing

00:00:23.760 --> 00:00:25.439
because not only do we have a new model

00:00:25.439 --> 00:00:27.680
from the Deepseek team, the SpaceX team

00:00:27.680 --> 00:00:30.240
also dropped Grock 4.6 six. And both of

00:00:30.240 --> 00:00:32.000
these models are competing with the best

00:00:32.000 --> 00:00:34.079
Frontier Labs at a fraction of their

00:00:34.079 --> 00:00:37.040
cost. Let's start with Gro 4.6. And one

00:00:37.040 --> 00:00:38.800
thing I'm going to say is that I'm

00:00:38.800 --> 00:00:40.879
genuinely surprised by the SpaceX team.

00:00:40.879 --> 00:00:42.480
I'm not trying to glaze them or Elon

00:00:42.480 --> 00:00:44.079
Musk or anything. I was just very

00:00:44.079 --> 00:00:46.480
critical of this lab. In 2025, they

00:00:46.480 --> 00:00:48.320
dropped Grog 4 and other models like

00:00:48.320 --> 00:00:50.399
that, but I wasn't really, you know,

00:00:50.399 --> 00:00:52.239
mind blown with their performance

00:00:52.239 --> 00:00:53.840
because they're pretty they're pretty

00:00:53.840 --> 00:00:56.000
subpar compared to any of the other labs

00:00:56.000 --> 00:00:58.480
out there. But in 2026, it looks like

00:00:58.480 --> 00:01:00.079
things have kind of changed a little

00:01:00.079 --> 00:01:02.399
because Grock 4.5 was quite competitive

00:01:02.399 --> 00:01:05.600
based off what it cost and Grock 4.6 is

00:01:05.600 --> 00:01:07.920
actually not too bad. If we take a look

00:01:07.920 --> 00:01:10.000
at the benchmarks, for example, if we

00:01:10.000 --> 00:01:11.920
start with the artificial analysis

00:01:11.920 --> 00:01:14.240
intelligence index, this model achieves

00:01:14.240 --> 00:01:18.159
a 61 and Fable 5 is at 62. Now, what's

00:01:18.159 --> 00:01:19.680
really important to remember is that

00:01:19.680 --> 00:01:22.159
once again, this model is quite cheap

00:01:22.159 --> 00:01:24.400
compared to Fable 5. This is about $2

00:01:24.400 --> 00:01:26.960
per million input tokens and $6 per

00:01:26.960 --> 00:01:29.439
million output tokens. In Fable 5 sits

00:01:29.439 --> 00:01:32.240
at $10 per million input tokens and $50

00:01:32.240 --> 00:01:34.400
per million output tokens. And this is

00:01:34.400 --> 00:01:36.400
why you start to appreciate this release

00:01:36.400 --> 00:01:38.320
a little bit more. It might not beat the

00:01:38.320 --> 00:01:40.000
performance of the best models is

00:01:40.000 --> 00:01:42.880
matching them. Even GBT 5.6 so which is

00:01:42.880 --> 00:01:44.960
a pretty capable model and this is set

00:01:44.960 --> 00:01:48.000
at max on the artificial analysis index.

00:01:48.000 --> 00:01:51.520
This model achieves 61 and Grock 4.6 61.

00:01:51.520 --> 00:01:53.840
So it ties it and it's much cheaper. And

00:01:53.840 --> 00:01:55.119
then even on all of these other

00:01:55.119 --> 00:01:57.439
benchmarks, for example, the GDP Val

00:01:57.439 --> 00:01:59.600
one, it actually beats Fable 5, which is

00:01:59.600 --> 00:02:04.159
at 1741. And then GPT 5.6, it's at 1728.

00:02:04.159 --> 00:02:06.000
I'm not sure why they didn't choose Opus

00:02:06.000 --> 00:02:07.520
5 as well, but I guess they wanted to

00:02:07.520 --> 00:02:09.759
choose the quote unquote strongest model

00:02:09.759 --> 00:02:11.280
lineup from each lab. And they chose

00:02:11.280 --> 00:02:13.760
Fable 5 for Enthropic, which is fair.

00:02:13.760 --> 00:02:15.440
And then Deep Software Engineering one,

00:02:15.440 --> 00:02:17.520
which is a critical benchmark. This

00:02:17.520 --> 00:02:20.640
model doesn't beat Fable 5 or GPT 5.6

00:02:20.640 --> 00:02:22.480
six soul, but it gets close to it. It's

00:02:22.480 --> 00:02:26.080
65.9 and Fable 5 sits at 70%. But if

00:02:26.080 --> 00:02:27.520
you're getting results that are pretty

00:02:27.520 --> 00:02:29.200
close and the model is five times

00:02:29.200 --> 00:02:30.879
cheaper, I wouldn't be too disappointed

00:02:30.879 --> 00:02:32.560
with this result. And then same thing

00:02:32.560 --> 00:02:34.879
with Cursor Bench 3.2, the model

00:02:34.879 --> 00:02:36.959
achieves 69.9%,

00:02:36.959 --> 00:02:40.080
funny number. And Fable 5 sits at 70.5.

00:02:40.080 --> 00:02:42.400
So once again, closer. And Grock 4.6

00:02:42.400 --> 00:02:45.200
beats GPT 5.6 Soul. Same thing with the

00:02:45.200 --> 00:02:48.000
Frontier Code. It gets close to Fable 5,

00:02:48.000 --> 00:02:51.040
beats GPT 5.6 Soul. So yeah, this model

00:02:51.040 --> 00:02:53.280
is actually available in cursor. So the

00:02:53.280 --> 00:02:55.200
partnership with cursor or I guess the

00:02:55.200 --> 00:02:57.360
acquisition has really helped SpaceX

00:02:57.360 --> 00:02:59.360
make some strides in the AI space this

00:02:59.360 --> 00:03:01.840
year and Grock build is something that

00:03:01.840 --> 00:03:03.680
you know maybe not a lot of us have been

00:03:03.680 --> 00:03:05.599
using so far but it's probably going to

00:03:05.599 --> 00:03:08.159
be another platform like codeex or cloud

00:03:08.159 --> 00:03:10.319
code that we start to use but obviously

00:03:10.319 --> 00:03:12.319
cursor is quite strong as well. So you

00:03:12.319 --> 00:03:14.319
have options available for you to use

00:03:14.319 --> 00:03:16.319
them in both. And one thing to note is

00:03:16.319 --> 00:03:18.560
that they're offering two times usage

00:03:18.560 --> 00:03:20.560
inside Grock build and cursor for the

00:03:20.560 --> 00:03:22.239
first week. So if you just want to try

00:03:22.239 --> 00:03:24.640
it out, see what you feel about it, then

00:03:24.640 --> 00:03:26.400
you know it might be worth trying it out

00:03:26.400 --> 00:03:28.000
right now cuz you get double the usage

00:03:28.000 --> 00:03:30.000
in the first week. Now if we were to

00:03:30.000 --> 00:03:31.519
take a look at some of the outputs that

00:03:31.519 --> 00:03:33.280
people have been generating with Grock

00:03:33.280 --> 00:03:35.519
4.6, what we're looking at right now is

00:03:35.519 --> 00:03:38.080
a Falcon 9 booster return sequence

00:03:38.080 --> 00:03:40.000
simulation. And this was done in a

00:03:40.000 --> 00:03:43.360
single HTML file. And as I said, if you

00:03:43.360 --> 00:03:45.360
expected to get this type of output from

00:03:45.360 --> 00:03:48.560
Grock in 2025, you would be kind of

00:03:48.560 --> 00:03:50.080
surprised because you wouldn't expect

00:03:50.080 --> 00:03:51.680
something like this to be generated with

00:03:51.680 --> 00:03:53.360
Grock. But now it looks like we have to

00:03:53.360 --> 00:03:55.120
start taking the Grock team a little bit

00:03:55.120 --> 00:03:57.360
more serious because this output is

00:03:57.360 --> 00:03:59.360
quite competitive. And obviously, we're

00:03:59.360 --> 00:04:01.120
just looking at a simulation and we're

00:04:01.120 --> 00:04:02.879
just basing it off of a visual

00:04:02.879 --> 00:04:04.959
representation. But if the model is able

00:04:04.959 --> 00:04:06.080
to produce something like this

00:04:06.080 --> 00:04:08.400
consistently, then I would expect a lot

00:04:08.400 --> 00:04:10.159
of people to start adopting Grock

00:04:10.159 --> 00:04:11.840
because number one, it is cheaper than

00:04:11.840 --> 00:04:13.680
the other labs at least at the moment

00:04:13.680 --> 00:04:15.120
because we don't know if this pricing

00:04:15.120 --> 00:04:17.040
strategy is going to be sustainable for

00:04:17.040 --> 00:04:19.040
the SpaceX team in the long run. But at

00:04:19.040 --> 00:04:20.560
least for now, their models are

00:04:20.560 --> 00:04:22.079
definitely cheaper compared to the

00:04:22.079 --> 00:04:24.160
others. And as I mentioned, this model

00:04:24.160 --> 00:04:26.160
excelled at the artificial analysis

00:04:26.160 --> 00:04:28.800
index. This model jumped to probably

00:04:28.800 --> 00:04:32.720
number four model. It's tied to GPT 5.6

00:04:32.720 --> 00:04:34.000
pretty much similar. So you could say

00:04:34.000 --> 00:04:35.759
number three as well, but the models

00:04:35.759 --> 00:04:38.639
before that are Fable 5 and Opus 5,

00:04:38.639 --> 00:04:40.720
which are only above the model by about

00:04:40.720 --> 00:04:43.840
a 1% or a 2% difference. So yeah, even

00:04:43.840 --> 00:04:46.000
on this intelligence index, which if

00:04:46.000 --> 00:04:47.520
you're not familiar with has nine

00:04:47.520 --> 00:04:49.360
evaluations. So on all of these

00:04:49.360 --> 00:04:51.440
evaluations, it's kind of matching

00:04:51.440 --> 00:04:53.600
almost Fable 5 performance, which is

00:04:53.600 --> 00:04:56.240
crazy to see. Before we continue, we

00:04:56.240 --> 00:04:57.840
just launched the Universe of AI

00:04:57.840 --> 00:04:59.600
newsletter. If you want to stay on top

00:04:59.600 --> 00:05:01.600
of AI news without having to hunt for

00:05:01.600 --> 00:05:03.680
it, link is in the description. Don't

00:05:03.680 --> 00:05:05.520
miss out. And what you see on screen

00:05:05.520 --> 00:05:07.680
right now is a racing game that Grock

00:05:07.680 --> 00:05:10.400
4.6 build. And based off of this post,

00:05:10.400 --> 00:05:12.320
the model took about 1 minute and it was

00:05:12.320 --> 00:05:14.400
a fiveword prompt, which was a create a

00:05:14.400 --> 00:05:16.639
simple racing game in HTML. So, if

00:05:16.639 --> 00:05:17.919
you're able to generate something like

00:05:17.919 --> 00:05:20.960
this easily using Grock 4.6, I think a

00:05:20.960 --> 00:05:22.880
lot of people will be happy. And this is

00:05:22.880 --> 00:05:25.199
a more detailed analysis of what it

00:05:25.199 --> 00:05:27.919
costs to run GPT 5.6 6 on the artificial

00:05:27.919 --> 00:05:29.919
analysis index and what it produced

00:05:29.919 --> 00:05:32.000
meaning the output. Both of these models

00:05:32.000 --> 00:05:34.160
if you remember scored 61 on the

00:05:34.160 --> 00:05:36.240
artificial analysis index. Now to run

00:05:36.240 --> 00:05:38.880
the whole test with GPT 5.6 it cost

00:05:38.880 --> 00:05:42.400
about 2.8K and then with Grock 4.6 it

00:05:42.400 --> 00:05:45.360
cost about 1.1Kish. And this tells you

00:05:45.360 --> 00:05:46.960
that you're getting similar level of

00:05:46.960 --> 00:05:49.199
performance at half the cost. So yeah,

00:05:49.199 --> 00:05:51.120
this is a big release for the SpaceX

00:05:51.120 --> 00:05:52.639
team because they just proved once again

00:05:52.639 --> 00:05:54.400
that they are a lab that you seriously

00:05:54.400 --> 00:05:56.320
start need to considering, especially in

00:05:56.320 --> 00:05:58.560
2026. And I'm going to talk more about

00:05:58.560 --> 00:06:01.120
Deep Seek version 4 Pro GA. But

00:06:01.120 --> 00:06:02.880
basically what we're seeing today is

00:06:02.880 --> 00:06:04.639
that both of these releases kind of

00:06:04.639 --> 00:06:07.360
emphasize the fact that performance and

00:06:07.360 --> 00:06:09.199
all above that is the price at what

00:06:09.199 --> 00:06:10.639
you're getting for that performance is

00:06:10.639 --> 00:06:12.479
becoming more and more important for all

00:06:12.479 --> 00:06:14.800
users because we see many labs now

00:06:14.800 --> 00:06:17.039
focusing on creating the best model at

00:06:17.039 --> 00:06:19.600
the cheapest cost. Last year in 2025

00:06:19.600 --> 00:06:21.680
most of the labs were just focused on I

00:06:21.680 --> 00:06:24.400
would say creating the strongest model.

00:06:24.400 --> 00:06:26.080
Yes, cost was important, but I think

00:06:26.080 --> 00:06:27.840
most of the times the frontier labs,

00:06:27.840 --> 00:06:30.080
meaning OpenAI, Enthropic, were kind of

00:06:30.080 --> 00:06:31.759
more lenient on that fact because they

00:06:31.759 --> 00:06:34.000
didn't have as strong of a competition.

00:06:34.000 --> 00:06:36.479
Intelligent models that are maybe not

00:06:36.479 --> 00:06:39.039
always ahead of OpenAI Enthropic, but

00:06:39.039 --> 00:06:40.720
match their performance at a fraction of

00:06:40.720 --> 00:06:42.800
the cost. So yes, price toerformance

00:06:42.800 --> 00:06:45.680
ratio is becoming a critical I would say

00:06:45.680 --> 00:06:48.639
indicator in 2026. Now, this is the

00:06:48.639 --> 00:06:50.479
updated benchmark chart after the

00:06:50.479 --> 00:06:52.240
release of the new model. And one thing

00:06:52.240 --> 00:06:53.919
you'll see across the board is that it

00:06:53.919 --> 00:06:55.919
matches the top level performance of

00:06:55.919 --> 00:06:57.360
many of the models. The one thing

00:06:57.360 --> 00:06:58.479
interesting over here is that they

00:06:58.479 --> 00:07:00.639
haven't put Opus 5 here for some reason.

00:07:00.639 --> 00:07:02.240
There is Fable 5 here that we can

00:07:02.240 --> 00:07:04.000
compare this model against. But one

00:07:04.000 --> 00:07:06.000
thing you'll notice is that Deep Seek,

00:07:06.000 --> 00:07:09.199
remember this model costs 43.5 cents per

00:07:09.199 --> 00:07:12.080
million input tokens and 87 cents per

00:07:12.080 --> 00:07:13.919
million output tokens. While the other

00:07:13.919 --> 00:07:15.680
models all over here are way more

00:07:15.680 --> 00:07:17.440
expensive than that. So the first thing

00:07:17.440 --> 00:07:20.000
if you look at terminal bench 2.1 the

00:07:20.000 --> 00:07:22.319
model scores 87.9.

00:07:22.319 --> 00:07:24.960
The older version of the model was 72.1

00:07:24.960 --> 00:07:27.120
and the flash version which we got last

00:07:27.120 --> 00:07:30.319
week was 82.7. And what's crazy is that

00:07:30.319 --> 00:07:34.000
Fable 5 is 88. Yes, 88. So this model is

00:07:34.000 --> 00:07:37.360
only.1% behind Fable 5. And then on the

00:07:37.360 --> 00:07:39.520
Cyber Gym, which is Cyber Security, the

00:07:39.520 --> 00:07:42.240
model actually beats Fable 5. Fable 5

00:07:42.240 --> 00:07:45.280
sits at 83.1 while Deepseek version 4

00:07:45.280 --> 00:07:47.039
Pro, the new one that we got today, is

00:07:47.039 --> 00:07:49.680
at 83.3. Now, if this tells you

00:07:49.680 --> 00:07:51.680
something is that this model is once

00:07:51.680 --> 00:07:53.759
again really geared at cyber security.

00:07:53.759 --> 00:07:55.599
We saw the Flash model also be geared

00:07:55.599 --> 00:07:57.280
towards cyber security and becoming a

00:07:57.280 --> 00:07:58.800
model that was quite capable in that

00:07:58.800 --> 00:08:00.560
area. And we're seeing the same thing

00:08:00.560 --> 00:08:02.720
today with the new version 4 Pro which

00:08:02.720 --> 00:08:05.840
sits at 83.3 outperforming Fable 5 on

00:08:05.840 --> 00:08:07.680
the benchmark. On deep software

00:08:07.680 --> 00:08:09.919
engineering, this model is not beating

00:08:09.919 --> 00:08:12.639
Fable 5 because Fable 5 sits at 70, but

00:08:12.639 --> 00:08:15.440
the model achieves 62.7 which is ahead

00:08:15.440 --> 00:08:18.639
of GLM 5.2 and a bit behind Kimmy K3

00:08:18.639 --> 00:08:20.879
which sits at 67.5.

00:08:20.879 --> 00:08:22.879
But one thing again, this is a big jump

00:08:22.879 --> 00:08:24.560
compared to the preview version of the

00:08:24.560 --> 00:08:27.360
model which was at 12.8. So yeah, they

00:08:27.360 --> 00:08:28.960
have really trained this model and they

00:08:28.960 --> 00:08:30.560
have improved it on the back end because

00:08:30.560 --> 00:08:32.880
we can clearly see in this deep software

00:08:32.880 --> 00:08:34.880
engineering benchmark. And there's a

00:08:34.880 --> 00:08:36.719
couple of other benchmarks, but one key

00:08:36.719 --> 00:08:39.120
benchmark that this model excels at is

00:08:39.120 --> 00:08:41.279
the automation bench where the model

00:08:41.279 --> 00:08:45.760
achieves 31.8 and Fable 5 29.1. So once

00:08:45.760 --> 00:08:48.000
again, the model outperforms Fable 5.

00:08:48.000 --> 00:08:49.920
Now this is a fraction of the cost as I

00:08:49.920 --> 00:08:52.480
mentioned. This is 43 cents and 87 for

00:08:52.480 --> 00:08:54.640
input and output respectively versus

00:08:54.640 --> 00:08:57.519
Fable 5 which is at $10.50.

00:08:57.519 --> 00:08:59.279
So yes, if we look at the terminal bench

00:08:59.279 --> 00:09:01.279
for example, we are seeing this model

00:09:01.279 --> 00:09:03.600
achieve a result which is 0.1 behind the

00:09:03.600 --> 00:09:06.160
best model out there Fable 5 and the

00:09:06.160 --> 00:09:09.519
pricing is insane to look at $43 versus

00:09:09.519 --> 00:09:13.519
$10 87 versus $50 per million input

00:09:13.519 --> 00:09:15.360
tokens. So when we turn that into a

00:09:15.360 --> 00:09:18.160
fraction, this model is 57 times

00:09:18.160 --> 00:09:20.399
cheaper. And earlier last week, I think

00:09:20.399 --> 00:09:22.160
I made a video about how DeepSeek

00:09:22.160 --> 00:09:24.480
version 4 Pro was going to achieve a

00:09:24.480 --> 00:09:25.839
result like this because there was a

00:09:25.839 --> 00:09:27.839
investor report that leaked. And when I

00:09:27.839 --> 00:09:29.440
saw the pricing where it said that it

00:09:29.440 --> 00:09:30.959
was going to match Fable 5 or even

00:09:30.959 --> 00:09:33.760
outperform it and be at 57 times cheaper

00:09:33.760 --> 00:09:35.440
per cost. When I was reading the

00:09:35.440 --> 00:09:37.120
investor report, I didn't really believe

00:09:37.120 --> 00:09:39.040
it. But today, it is proven because at

00:09:39.040 --> 00:09:41.040
least on the benchmark so far, yes, I'm

00:09:41.040 --> 00:09:42.880
just saying the benchmarks, we are

00:09:42.880 --> 00:09:45.040
seeing similar level performance. Now,

00:09:45.040 --> 00:09:46.959
if this model actually performs like

00:09:46.959 --> 00:09:48.959
that in production, it's a little too

00:09:48.959 --> 00:09:50.720
early to tell yet. We still would have

00:09:50.720 --> 00:09:52.399
to give it a couple of weeks and see how

00:09:52.399 --> 00:09:54.399
it's performing in the long run because

00:09:54.399 --> 00:09:56.399
sometimes on the release day the models

00:09:56.399 --> 00:09:58.480
perform quite good, but over time they

00:09:58.480 --> 00:10:00.240
kind of deteriorate in their quality.

00:10:00.240 --> 00:10:02.000
So, we hope DeepS doesn't do that.

00:10:02.000 --> 00:10:03.680
Historically, they haven't done that,

00:10:03.680 --> 00:10:05.600
but let's just see. Just going to put it

00:10:05.600 --> 00:10:07.120
out there because right now we're basing

00:10:07.120 --> 00:10:09.760
this off of benchmarks. But that's it

00:10:09.760 --> 00:10:11.519
for today's video. Make sure you guys

00:10:11.519 --> 00:10:13.519
are subscribed to the channel. Follow

00:10:13.519 --> 00:10:15.120
our new newsletter as well at

00:10:15.120 --> 00:10:17.350
universeofai.behive.com

00:10:17.350 --> 00:10:17.360
universeofai.behive.com

00:10:17.360 --> 00:10:19.040
as well as subscribe to the main channel

00:10:19.040 --> 00:10:21.519
World of AI and support us on X by

00:10:21.519 --> 00:10:23.760
following the Universe of AIZ as well.

00:10:23.760 --> 00:10:25.360
Until then, I'll see you guys in the

00:10:25.360 --> 00:10:27.519
next
