0:00.250–0:03.749
So , Deepseek version 4 Pro is officially out today .
0:03.759–0:06.460
Now , you might be confused because this model was technically
0:06.470–0:09.810
out , but it was in general availability , but this is the official
0:09.820–0:12.681
release , meaning that they have trained the model a little bit
0:12.691–0:16.111
more and produced a stronger version of the model that is available
0:16.121–0:16.590
today .
0:16.600–0:19.337
Now , if I were to summarize today's release in one simple sentence
0:19.347–0:22.651
, it is that the price to performance is becoming a very ,
0:22.661–0:25.682
very important thing because not only do we have a new model from
0:25.692–0:29.074
the Deepseek team , the SpaceX team also dropped Grock 4 .
0:29.084–0:29.640
6 six .
0:29.650–0:33.063
And both of these models are competing with the best Frontier Labs
0:33.073–0:35.066
at a fraction of their cost .
0:35.076–0:36.431
Let's start with Gro 4 .
0:36.441–0:36.827
6 .
0:36.837–0:39.761
And one thing I'm going to say is that I'm genuinely surprised
0:39.771–0:41.029
by the SpaceX team .
0:41.039–0:43.408
I'm not trying to glaze them or Elon Musk or anything .
0:43.418–0:45.311
I was just very critical of this lab .
0:45.321–0:48.671
In 2025 , they dropped Grog 4 and other models like that ,
0:48.681–0:52.371
but I wasn't really , you know , mind blown with their performance
0:52.381–0:55.400
because they're pretty they're pretty subpar compared to any of
0:55.410–0:56.743
the other labs out there .
0:56.753–1:00.140
But in 2026 , it looks like things have kind of changed a little
1:00.150–1:01.212
because Grock 4 .
1:01.222–1:04.982
5 was quite competitive based off what it cost and Grock 4 .
1:04.992–1:07.268
6 is actually not too bad .
1:07.278–1:09.740
If we take a look at the benchmarks , for example ,
1:09.750–1:13.248
if we start with the artificial analysis intelligence index ,
1:13.258–1:17.633
this model achieves a 61 and Fable 5 is at 62 .
1:17.643–1:20.583
Now , what's really important to remember is that once again ,
1:20.593–1:23.374
this model is quite cheap compared to Fable 5 .
1:23.384–1:27.600
This is about 2 per million input tokens and 6 per million output
1:27.610–1:28.245
tokens .
1:28.255–1:32.844
In Fable 5 sits at 10 per million input tokens and 50 per million
1:32.854–1:33.893
output tokens .
1:33.903–1:36.959
And this is why you start to appreciate this release a little bit
1:36.969–1:37.363
more .
1:37.373–1:40.540
It might not beat the performance of the best models is matching
1:40.550–1:41.037
them .
1:41.047–1:42.390
Even GPT 5 .
1:42.400–1:46.161
6 so which is a pretty capable model and this is set at max on
1:46.171–1:48.126
the artificial analysis index .
1:48.136–1:50.828
This model achieves 61 and Grok 4 .
1:50.838–1:51.632
6 61 .
1:51.642–1:53.804
So it ties it and it's much cheaper .
1:53.814–1:56.216
And then even on all of these other benchmarks ,
1:56.226–1:59.343
for example , the Code Val one , it actually beats Fable 5 ,
1:59.353–2:01.186
which is at 1741 .
2:01.196–2:02.427
And then GPT 5 .
2:02.437–2:04.309
6 , it's at 1728 .
2:04.319–2:06.713
I'm not sure why they didn't choose Opus 5 as well ,
2:06.723–2:09.916
but I guess they wanted to choose the quote unquote strongest model
2:09.926–2:10.959
lineup from each lab .
2:10.969–2:13.903
And they chose Fable 5 for Enthropic , which is fair .
2:13.913–2:16.917
And then Deep Software Engineering one , which is a critical benchmark
2:16.917–2:19.286
.
2:19.286–2:20.585
This model doesn't beat Fable 5 or GPT 5 .
2:20.595–2:22.322
6 six soul , but it gets close to it .
2:22.332–2:23.255
It's 65 .
2:23.265–2:25.839
9 and Fable 5 sits at 70 .
2:25.849–2:28.556
But if you're getting results that are pretty close and the model
2:28.566–2:31.432
is five times cheaper , I wouldn't be too disappointed with this
2:31.442–2:31.991
result .
2:32.001–2:34.068
And then same thing with Cursor Bench 3 .
2:34.078–2:36.221
2 , the model achieves 69 .
2:36.231–2:37.904
9 , funny number .
2:37.914–2:39.617
And Fable 5 sits at 70 .
2:39.627–2:40.141
5 .
2:40.151–2:41.260
So once again , closer .
2:41.270–2:42.199
And Grock 4 .
2:42.209–2:43.693
6 beats GPT 5 .
2:43.703–2:44.860
6 .
2:44.870–2:46.375
Same thing with the Frontier Code .
2:46.385–2:49.287
It gets close to Fable 5 , beats GPT 5 .
2:49.297–2:49.987
6 .
2:49.997–2:52.808
So yeah , this model is actually available in cursor .
2:52.818–2:56.515
So the partnership with cursor or I guess the acquisition has really
2:56.525–3:00.540
helped SpaceX make some strides in the AI space this year and Grock
3:00.550–3:03.740
build is something that you know maybe not a lot of us have been
3:03.750–3:07.341
using so far but it's probably going to be another platform like
3:07.351–3:10.793
codeex or cloud code that we start to use but obviously cursor
3:10.803–3:11.997
is quite strong as well .
3:12.007–3:15.369
So you have options available for you to use them in both .
3:15.379–3:18.660
And one thing to note is that they're offering two times usage
3:18.670–3:21.287
inside Grok Build and Cursor for the first week .
3:21.297–3:23.992
So if you just want to try it out , see what you feel about it
3:24.002–3:27.189
, then you know it might be worth trying it out right now cuz you
3:27.199–3:29.304
get double the usage in the first week .
3:29.314–3:31.907
Now if we were to take a look at some of the outputs that people
3:31.917–3:33.916
have been generating with Grok 4 .
3:33.926–3:37.824
6 , what we're looking at right now is a Falcon 9 booster return
3:37.834–3:38.939
sequence simulation .
3:38.949–3:42.050
And this was done in a single HTML file .
3:42.060–3:45.479
And as I said , if you expected to get this type of output from
3:45.489–3:49.852
Grock in 2025 , you would be kind of surprised because you wouldn't
3:49.862–3:52.240
expect something like this to be generated with Grock .
3:52.250–3:54.958
But now it looks like we have to start taking the Grok team a
3:54.968–3:58.195
little bit more serious because this output is quite competitive
3:58.195–4:00.295
.
4:00.295–4:01.350
And obviously , we're just looking at a simulation and we're just
4:01.360–4:03.987
basing it off of a visual representation .
4:03.997–4:06.624
But if the model is able to produce something like this consistently
4:06.634–4:10.525
, then I would expect a lot of people to start adopting Grock because
4:10.535–4:13.368
number one , it is cheaper than the other labs at least at the
4:13.378–4:15.895
moment because we don't know if this pricing strategy is going
4:15.905–4:18.895
to be sustainable for the SpaceX team in the long run .
4:18.905–4:21.659
But at least for now , their models are definitely cheaper compared
4:21.669–4:22.608
to the others .
4:22.618–4:26.003
And as I mentioned , this model excelled at the artificial analysis
4:26.013–4:26.635
index .
4:26.645–4:30.031
This model jumped to probably number four model .
4:30.041–4:32.180
It's tied to GPT 5 .
4:32.190–4:33.370
6 pretty much similar .
4:33.380–4:36.579
So you could say number three as well , but the models before that
4:36.589–4:40.896
are Fable 5 and Opus 5 , which are only above the model by about
4:40.906–4:43.454
a 1 or a 2 difference .
4:43.464–4:46.492
So yeah , even on this intelligence index , which if you're not
4:46.502–4:48.570
familiar with has nine evaluations .
4:48.580–4:52.646
So on all of these evaluations , it's kind of matching almost Fable
4:52.656–4:55.046
5 performance , which is crazy to see .
4:55.056–4:58.403
Before we continue , we just launched the Universe of AI newsletter
4:58.403–5:00.629
.
5:00.629–5:01.761
If you want to stay on top of AI news without having to hunt for
5:01.771–5:03.595
it , link is in the description .
5:03.605–5:04.710
Don't miss out .
5:04.720–5:07.800
And what you see on screen right now is a racing game that Grok
5:07.810–5:08.320
4 .
5:08.330–5:09.127
6 build .
5:09.137–5:12.168
And based off of this post , the model took about 1 minute and
5:12.178–5:15.191
it was a fiveword prompt , which was a create a simple racing game
5:15.201–5:16.269
in HTML .
5:16.279–5:19.367
So , if you're able to generate something like this easily using
5:19.377–5:20.100
Grock 4 .
5:20.110–5:22.300
6 , I think a lot of people will be happy .
5:22.310–5:26.455
And this is a more detailed analysis of what it costs to run GPT
5:26.465–5:26.785
5 .
5:26.795–5:30.272
6 6 on the artificial analysis index and what it produced meaning
5:30.282–5:31.224
the output .
5:31.234–5:34.713
Both of these models if you remember scored 61 on the artificial
5:34.723–5:35.745
analysis index .
5:35.755–5:38.230
Now to run the whole test with GPT 5 .
5:38.240–5:40.135
6 it cost about 2 .
5:40.145–5:42.008
6 it cost about 2 .
5:42.018–5:43.620
6 it cost about 1 .
5:43.630–5:44.300
1Kish .
5:44.310–5:47.465
And this tells you that you're getting similar level of performance
5:47.475–5:48.519
at half the cost .
5:48.529–5:51.893
So yeah , this is a big release for the SpaceX team because they
5:51.903–5:54.779
just proved once again that they are a lab that you seriously start
5:54.789–5:57.080
need to considering , especially in 2026 .
5:57.090–6:00.824
And I'm going to talk more about Deep Seek version 4 Pro GA .
6:00.834–6:04.350
But basically what we're seeing today is that both of these releases
6:04.360–6:08.119
kind of emphasize the fact that performance and all above that
6:08.129–6:11.005
is the price at what you're getting for that performance is becoming
6:11.015–6:13.201
more and more important for all users
6:13.201–6:14.612
because we see many labs
6:14.622–6:18.379
now focusing on creating the best model at the cheapest cost .
6:18.389–6:21.907
Last year in 2025 most of the labs were just focused on I would
6:21.917–6:24.472
say creating the strongest model .
6:24.482–6:25.571
Yes , cost was important ,
6:25.571–6:27.519
but I think most of the times the frontier
6:27.529–6:29.280
labs , meaning OpenAI , Anthropic ,
6:29.280–6:30.829
were kind of more lenient on
6:30.839–6:34.100
that fact because they didn't have as strong of a competition .
6:34.110–6:38.733
Intelligent models that are maybe not always ahead of OpenAI Enthropic
6:38.743–6:41.447
, but match their performance at a fraction of the cost .
6:41.457–6:45.369
So yes , price to performance ratio is becoming a critical I would
6:45.379–6:47.955
say indicator in 2026 .
6:47.965–6:50.861
Now , this is the updated benchmark chart after the release of
6:50.871–6:51.742
the new model .
6:51.752–6:54.623
And one thing you'll see across the board is that it matches the
6:54.633–6:56.944
top level performance of many of the models .
6:56.954–6:59.345
The one thing interesting over here is that they haven't put Opus
6:59.355–7:00.704
5 here for some reason .
7:00.714–7:03.747
There is Fable 5 here that we can compare this model against .
7:03.757–7:06.068
But one thing you'll notice is that DeepSeek ,
7:06.078–7:08.335
remember this model costs 43 .
7:08.345–7:12.791
5 cents per million input tokens and 87 cents per million output
7:12.801–7:13.511
tokens .
7:13.521–7:16.391
While the other models all over here are way more expensive than
7:16.401–7:16.872
that .
7:16.882–7:19.540
So the first thing if you look at terminal bench 2 .
7:19.550–7:21.629
1 the model scores 87 .
7:21.639–7:22.394
9 .
7:22.404–7:24.656
The older version of the model was 72 .
7:24.666–7:28.557
1 and the flash version which we got last week was 82 .
7:28.567–7:29.277
7 .
7:29.287–7:32.239
And what's crazy is that Fable 5 is 88 .
7:32.249–7:33.279
Yes , 88 .
7:33.289–7:34.887
So this model is only .
7:34.897–7:36.827
1 behind Fable 5 .
7:36.837–7:39.300
And then on the Cyber Gym , which is Cyber Security ,
7:39.310–7:41.629
the model actually beats Fable 5 .
7:41.639–7:43.559
Fable 5 sits at 83 .
7:43.569–7:46.683
1 while Deepseek version 4 Pro , the new one that we got today
7:46.693–7:48.044
3 .
7:48.054–7:48.765
3 .
7:48.775–7:51.966
Now , if this tells you something is that this model is once again
7:51.976–7:53.807
really geared at cyber security .
7:53.817–7:56.850
We saw the Flash model also be geared towards cyber security and
7:56.860–7:59.412
becoming a model that was quite capable in that area .
7:59.422–8:02.534
And we're seeing the same thing today with the new version 4 Pro
8:02.544–8:03.975
which sits at 83 .
8:03.985–8:06.858
3 outperforming Fable 5 on the benchmark .
8:06.868–8:10.379
On deep software engineering , this model is not beating Fable
8:10.389–8:12.491
5 because Fable 5 sits at 70 ,
8:12.491–8:14.497
but the model achieves 62 .
8:14.507–8:16.744
7 which is ahead of GLM 5 .
8:16.754–8:20.209
2 and a bit behind Kimmy K3 which sits at 67 .
8:20.219–8:20.977
5 .
8:20.987–8:24.115
But one thing again , this is a big jump compared to the preview
8:24.125–8:26.286
version of the model which was at 12 .
8:26.296–8:26.769
8 .
8:26.779–8:29.655
So yeah , they have really trained this model and they have improved
8:29.665–8:32.998
it on the back end because we can clearly see in this deep software
8:33.008–8:34.430
engineering benchmark .
8:34.440–8:37.216
And there's a couple of other benchmarks , but one key benchmark
8:37.226–8:41.353
that this model excels at is the automation bench where the model
8:41.363–8:42.680
achieves 31 .
8:42.690–8:44.775
8 and Fable 5 29 .
8:44.785–8:45.332
1 .
8:45.342–8:48.039
So once again , the model outperforms Fable 5 .
8:48.049–8:50.268
Now this is a fraction of the cost as I mentioned .
8:50.278–8:54.378
This is 43 cents and 87 for input and output respectively versus
8:54.388–8:56.300
Fable 5 which is at 10 .
8:56.310–8:57.083
50 .
8:57.093–8:59.787
So yes , if we look at the terminal bench for example ,
8:59.797–9:02.655
we are seeing this model achieve a result which is 0 .
9:02.665–9:07.078
1 behind the best model out there Fable 5 and the pricing is insane
9:07.088–9:14.184
to look at 43 versus 10 87 versus 50 per million input tokens .
9:14.194–9:18.674
So when we turn that into a fraction , this model is 57 times cheaper
9:18.674–9:19.184
.
9:19.184–9:22.285
And earlier last week , I think I made a video about how DeepSeek
9:22.295–9:25.740
version 4 Pro was going to achieve a result like this because there
9:25.750–9:27.495
was a investor report that leaked .
9:27.505–9:30.127
And when I saw the pricing where it said that it was going to match
9:30.137–9:34.470
Fable 5 or even outperform it and be at 57 times cheaper per cost
9:34.470–9:35.751
.
9:35.751–9:37.277
When I was reading the investor report , I didn't really believe
9:37.287–9:37.837
it .
9:37.847–9:40.633
But today , it is proven because at least on the benchmark so far
9:40.643–9:43.829
, yes , I'm just saying the benchmarks , we are seeing similar
9:43.839–9:44.948
level performance .
9:44.958–9:48.303
Now , if this model actually performs like that in production ,
9:48.313–9:50.141
it's a little too early to tell yet .
9:50.151–9:52.763
We still would have to give it a couple of weeks and see how it's
9:52.773–9:55.867
performing in the long run because sometimes on the release day
9:55.877–9:59.340
the models perform quite good , but over time they kind of deteriorate
9:59.350–10:00.380
in their quality .
10:00.390–10:02.005
So , we hope DeepSeek doesn't do that .
10:02.015–10:04.469
Historically , they haven't done that , but let's just see .
10:04.479–10:06.862
Just going to put it out there because right now we're basing this
10:06.872–10:08.529
off of benchmarks .
10:08.539–10:10.051
But that's it for today's video .
10:10.061–10:12.081
Make sure you guys are subscribed to the channel .
10:12.091–10:15.369
Follow our new newsletter as well at universeofai .
10:15.379–10:15.756
behiiv .
10:15.766–10:19.087
com as well as subscribe to the main channel World of AI and support
10:19.097–10:22.063
us on X by following the Universe of AIZ as well .
10:22.073–10:23.114
Until then ,
10:23.114–10:24.618
I'll see you guys
10:24.618–10:25.659
in the next
0:00.250–0:03.749
So , Deepseek version 4 Pro is officially out today .
所以,Deepseek 版本 4 Pro 今天正式推出了。
0:03.759–0:06.460
Now , you might be confused because this model was technically
現在,你可能會感到困惑,因為這個模型在技術上
0:06.470–0:09.810
out , but it was in general availability , but this is the official
已經發布,但處於一般可用性階段,而這是官方
0:09.820–0:12.681
release , meaning that they have trained the model a little bit
發布,意味著他們對模型進行了更多訓練,
0:12.691–0:16.111
more and produced a stronger version of the model that is available
並產出了一個更強大的模型版本,今天即可使用。
0:16.121–0:16.590
today .
今天。
0:16.600–0:19.337
Now , if I were to summarize today's release in one simple sentence
現在,如果我要用一句簡單的話來總結今天的發布,
0:19.347–0:22.651
, it is that the price to performance is becoming a very ,
那就是性價比正變得非常、非常重要的一環,因為
0:22.661–0:25.682
very important thing because not only do we have a new model from
非常重要的一點,因為我們不僅擁有一個來自
0:25.692–0:29.074
the Deepseek team , the SpaceX team also dropped Grock 4 .
Deepseek 團隊的新模型,SpaceX 團隊也推出了 Grock 4。
0:29.084–0:29.640
6 six .
6.6。
0:29.650–0:33.063
And both of these models are competing with the best Frontier Labs
這兩個模型都在以極低的成本與 Frontier Labs 的最佳模型競爭。
0:33.073–0:35.066
at a fraction of their cost .
讓我們從 Gro 4 開始。
0:35.076–0:36.431
Let's start with Gro 4 .
讓我們從 Gro 4 開始。
0:36.441–0:36.827
6 .
我要說的一件事是,我對 SpaceX 團隊真的感到驚訝。
0:36.837–0:39.761
And one thing I'm going to say is that I'm genuinely surprised
我不是在試圖吹捧他們或埃隆·馬斯克或什麼的。
0:39.771–0:41.029
by the SpaceX team .
我只是對這個實驗室非常批評。
0:41.039–0:43.408
I'm not trying to glaze them or Elon Musk or anything .
在 2025 年,他們推出了 Grog 4 和其他類似的模型,
0:43.418–0:45.311
I was just very critical of this lab .
但我並沒有真的,你知道,對他們的表現感到驚艷,
0:45.321–0:48.671
In 2025 , they dropped Grog 4 and other models like that ,
因為與其他實驗室相比,他們的表現相當差。
0:48.681–0:52.371
but I wasn't really , you know , mind blown with their performance
但在 2026 年,情況似乎有點改變,
0:52.381–0:55.400
because they're pretty they're pretty subpar compared to any of
因為跟其他實驗室相比,它們的表現相當遜色
0:55.410–0:56.743
the other labs out there .
5 在成本基礎上相當有競爭力,而 Grock 4。
0:56.753–1:00.140
But in 2026 , it looks like things have kind of changed a little
但在 2026 年,情況似乎有點改變了
1:00.150–1:01.212
because Grock 4 .
如果我們查看基準測試,例如,
1:01.222–1:04.982
5 was quite competitive based off what it cost and Grock 4 .
如果我們從 Artificial Analysis 智能指數開始,
1:04.992–1:07.268
6 is actually not too bad .
這個模型達到了 61 分,而 Fable 5 為 62 分。
1:07.278–1:09.740
If we take a look at the benchmarks , for example ,
現在,真正重要的是要記住,再一次,
1:09.750–1:13.248
if we start with the artificial analysis intelligence index ,
這個模型與 Fable 5 相比非常便宜。
1:13.258–1:17.633
this model achieves a 61 and Fable 5 is at 62 .
這大約是每百萬輸入標記 2 美元,每百萬輸出
1:17.643–1:20.583
Now , what's really important to remember is that once again ,
現在,真正重要的是要記住,再一次,
1:20.593–1:23.374
this model is quite cheap compared to Fable 5 .
在 Fable 5 中,每百萬輸入標記為 10 美元,每百萬
1:23.384–1:27.600
This is about 2 per million input tokens and 6 per million output
這大約是每百萬個輸入 token 2 個,每百萬個輸出
1:27.610–1:28.245
tokens .
這就是為什麼你開始稍微欣賞這次發布的原因。
1:28.255–1:32.844
In Fable 5 sits at 10 per million input tokens and 50 per million
它可能無法擊敗最佳模型的表現,但正在與它們匹敵。
1:32.854–1:33.893
output tokens .
輸出標記
1:33.903–1:36.959
And this is why you start to appreciate this release a little bit
這就是為什麼你會開始稍微欣賞這次發布的原因
1:36.969–1:37.363
more .
更多
1:37.373–1:40.540
It might not beat the performance of the best models is matching
它可能無法擊敗最佳模型的表現,但正在追趕
1:40.550–1:41.037
them .
它們
1:41.047–1:42.390
Even GPT 5 .
即使是 GPT 5。
1:42.400–1:46.161
6 so which is a pretty capable model and this is set at max on
6,這是一個相當強大的模型,並且在人工分析指數上設定為最高。
1:46.171–1:48.126
the artificial analysis index .
人工分析指數
1:48.136–1:50.828
This model achieves 61 and Grok 4 .
該模型達到 61 分,Grok 4。
1:50.838–1:51.632
6 61 .
6,61 分。
1:51.642–1:53.804
So it ties it and it's much cheaper .
所以它並列第一,而且便宜得多。
1:53.814–1:56.216
And then even on all of these other benchmarks ,
然後在這些其他基準測試中,
1:56.226–1:59.343
for example , the Code Val one , it actually beats Fable 5 ,
例如 Code Val 基準,它實際上擊敗了 Fable 5,
1:59.353–2:01.186
which is at 1741 .
後者為 1741 分。
2:01.196–2:02.427
And then GPT 5 .
然後是 GPT 5。
2:02.437–2:04.309
6 , it's at 1728 .
6,它為 1728 分。
2:04.319–2:06.713
I'm not sure why they didn't choose Opus 5 as well ,
我不確定為什麼他們沒有選擇 Opus 5,
2:06.723–2:09.916
but I guess they wanted to choose the quote unquote strongest model
但我猜他們想選擇所謂的「最強」模型
2:09.926–2:10.959
lineup from each lab .
陣容來自每個實驗室。
2:10.969–2:13.903
And they chose Fable 5 for Enthropic , which is fair .
他們為 Enthropic 選擇了 Fable 5,這很公平。
2:13.913–2:16.917
And then Deep Software Engineering one , which is a critical benchmark
然後是 Deep Software Engineering 基準,這是一個關鍵基準
2:16.917–2:19.286
.
。
2:19.286–2:20.585
This model doesn't beat Fable 5 or GPT 5 .
該模型沒有擊敗 Fable 5 或 GPT 5。
2:20.595–2:22.322
6 six soul , but it gets close to it .
6,6 分,但非常接近。
2:22.332–2:23.255
It's 65 .
它是 65。
2:23.265–2:25.839
9 and Fable 5 sits at 70 .
9,而 Fable 5 為 70 分。
2:25.849–2:28.556
But if you're getting results that are pretty close and the model
但如果你得到的結果非常接近,而且模型
2:28.566–2:31.432
is five times cheaper , I wouldn't be too disappointed with this
如果價格便宜五倍,我不會對這個結果太失望
2:31.442–2:31.991
result .
結果感到太失望。
2:32.001–2:34.068
And then same thing with Cursor Bench 3 .
然後是 Cursor Bench 3。
2:34.078–2:36.221
2 , the model achieves 69 .
2,該模型達到 69。
2:36.231–2:37.904
9 , funny number .
9,有趣的數字。
2:37.914–2:39.617
And Fable 5 sits at 70 .
而 Fable 5 為 70。
2:39.627–2:40.141
5 .
[未翻譯]
2:40.151–2:41.260
So once again , closer .
所以再一次,更接近。
2:41.270–2:42.199
And Grock 4 .
而 Grock 4。
2:42.209–2:43.693
6 beats GPT 5 .
6 擊敗了 GPT 5。
2:43.703–2:44.860
6 .
[未翻譯]
2:44.870–2:46.375
Same thing with the Frontier Code .
在 Frontier Code 方面也是如此。
2:46.385–2:49.287
It gets close to Fable 5 , beats GPT 5 .
它非常接近 Fable 5,擊敗了 GPT 5。
2:49.297–2:49.987
6 .
[未翻譯]
2:49.997–2:52.808
So yeah , this model is actually available in cursor .
所以是的,該模型實際上在 cursor 中可用。
2:52.818–2:56.515
So the partnership with cursor or I guess the acquisition has really
所以與 cursor 的合作關係,或者我猜是收購,確實
2:56.525–3:00.540
helped SpaceX make some strides in the AI space this year and Grock
幫助 SpaceX 在今年在 AI 領域取得了一些進展,而 Grock
3:00.550–3:03.740
build is something that you know maybe not a lot of us have been
build 是你知道也許我們中沒有很多人一直在
3:03.750–3:07.341
using so far but it's probably going to be another platform like
到目前為止還沒有人使用,但未來我們可能會開始使用類似 CodeEx 或 Cloud Code 的這類平台,不過顯然 Cursor 也很強大。
3:07.351–3:10.793
codeex or cloud code that we start to use but obviously cursor
因此,你有可用的選項,可以在兩者中選擇使用。
3:10.803–3:11.997
is quite strong as well .
並且有一件事需要注意,Grok Build 和 Cursor 在第一週提供兩倍的用量。
3:12.007–3:15.369
So you have options available for you to use them in both .
所以如果你只是想試試看,看看你的感覺如何,那麼現在嘗試可能值得,因為你在第一週可以獲得雙倍的用量。
3:15.379–3:18.660
And one thing to note is that they're offering two times usage
現在,如果我們來看看人們使用 Grok 4.6 所生成的一些輸出結果。
3:18.670–3:21.287
inside Grok Build and Cursor for the first week .
我們現在看到的是 Falcon 9 助推器返回序列的模擬。
3:21.297–3:23.992
So if you just want to try it out , see what you feel about it
這是在單一 HTML 檔案中完成的。
3:24.002–3:27.189
, then you know it might be worth trying it out right now cuz you
正如我所說,如果你預期在 2025 年從 Grok 獲得這種類型的輸出,你會感到相當驚訝,因為你不會預期 Grok 能生成這樣的內容。
3:27.199–3:29.304
get double the usage in the first week .
但現在看來,我們必須開始更認真地看待 Grok 團隊,因為這項輸出相當具有競爭力。
3:29.314–3:31.907
Now if we were to take a look at some of the outputs that people
顯然,我們只是在看一個模擬,並且僅基於視覺呈現來評估。
3:31.917–3:33.916
have been generating with Grok 4 .
都是使用 Grok 4 生成的。
3:33.926–3:37.824
6 , what we're looking at right now is a Falcon 9 booster return
但至少目前,他們的模型確實比其他模型便宜。
3:37.834–3:38.939
sequence simulation .
正如我所提到的,該模型在人工分析指數中表現出色。
3:38.949–3:42.050
And this was done in a single HTML file .
該模型躍升至可能排名第四的模型。
3:42.060–3:45.479
And as I said , if you expected to get this type of output from
它與 GPT 5.6 並列,表現非常相似。
3:45.489–3:49.852
Grock in 2025 , you would be kind of surprised because you wouldn't
所以你也可以說它是排名第三,但在此之前的是 Falcon 5 和 Opus 5,它們僅比該模型高出約 1 或 2 分的差距。
3:49.862–3:52.240
expect something like this to be generated with Grock .
所以是的,即使是在這個智力指數上,如果你不熟悉,它包含九項評估。
3:52.250–3:54.958
But now it looks like we have to start taking the Grok team a
所以在所有這些評估中,它幾乎與 Falcon 5 的表現相匹配。
3:54.968–3:58.195
little bit more serious because this output is quite competitive
稍微更嚴肅一些,因為這項成果相當具有競爭力
3:58.195–4:00.295
.
。
4:00.295–4:01.350
And obviously , we're just looking at a simulation and we're just
而且顯然,我們只是在看模擬結果,並且只是
4:01.360–4:03.987
basing it off of a visual representation .
基於視覺呈現來進行評估。
4:03.997–4:06.624
But if the model is able to produce something like this consistently
但如果該模型能夠持續產生類似的成果
4:06.634–4:10.525
, then I would expect a lot of people to start adopting Grock because
,那麼我預期會有很多人開始採用 Grock,因為
4:10.535–4:13.368
number one , it is cheaper than the other labs at least at the
第一,它比其他實驗室便宜,至少目前是如此
4:13.378–4:15.895
moment because we don't know if this pricing strategy is going
因為我們不知道這種定價策略
4:15.905–4:18.895
to be sustainable for the SpaceX team in the long run .
對 SpaceX 團隊在長期來看是否具可持續性
4:18.905–4:21.659
But at least for now , their models are definitely cheaper compared
但至少目前,他們的模型確實比其他模型便宜
4:21.669–4:22.608
to the others .
相較於其他模型
4:22.618–4:26.003
And as I mentioned , this model excelled at the artificial analysis
正如我所提到的,該模型在人工分析指數方面表現出色
4:26.013–4:26.635
index .
指數
4:26.645–4:30.031
This model jumped to probably number four model .
該模型躍升至可能排名第四的模型
4:30.041–4:32.180
It's tied to GPT 5 .
它與 GPT 5 並列。
4:32.190–4:33.370
6 pretty much similar .
與 6 幾乎相同。
4:33.380–4:36.579
So you could say number three as well , but the models before that
所以你也可以說它是第三名,但在那之前的模型
4:36.589–4:40.896
are Fable 5 and Opus 5 , which are only above the model by about
是 Fable 5 和 Opus 5,它們僅比該模型高出約
4:40.906–4:43.454
a 1 or a 2 difference .
1 或 2 的差距。
4:43.464–4:46.492
So yeah , even on this intelligence index , which if you're not
所以是的,即使是在這個智力指數上,如果你不
4:46.502–4:48.570
familiar with has nine evaluations .
熟悉的話,它包含九項評估。
4:48.580–4:52.646
So on all of these evaluations , it's kind of matching almost Fable
所以在所有這些評估中,它幾乎與 Fable
4:52.656–4:55.046
5 performance , which is crazy to see .
5的表現,這在視覺上令人瘋狂。
4:55.056–4:58.403
Before we continue , we just launched the Universe of AI newsletter
在我們繼續之前,我們剛剛推出了AI宇宙通訊電子報
4:58.403–5:00.629
.
。
5:00.629–5:01.761
If you want to stay on top of AI news without having to hunt for
如果你想隨時掌握AI新聞,而不必費心去搜尋
5:01.771–5:03.595
it , link is in the description .
它,連結在描述中。
5:03.605–5:04.710
Don't miss out .
別錯過。
5:04.720–5:07.800
And what you see on screen right now is a racing game that Grok
而你現在在螢幕上看到的是一款由Grok
5:07.810–5:08.320
4 .
[未翻譯]
5:08.330–5:09.127
6 build .
.6版本構建的賽車遊戲。
5:09.137–5:12.168
And based off of this post , the model took about 1 minute and
根據這篇帖子,該模型大約花了1分鐘
5:12.178–5:15.191
it was a fiveword prompt , which was a create a simple racing game
,這是一個五個字的提示,內容是創建一個簡單的賽車遊戲
5:15.201–5:16.269
in HTML .
,使用HTML。
5:16.279–5:19.367
So , if you're able to generate something like this easily using
所以,如果你能輕鬆地生成像這樣的東西,使用
5:19.377–5:20.100
Grock 4 .
所以,如果你能輕鬆地使用Grok 4
5:20.110–5:22.300
6 , I think a lot of people will be happy .
.6生成這樣的東西,我認為許多人會很高興。
5:22.310–5:26.455
And this is a more detailed analysis of what it costs to run GPT
這是一份更詳細的分析,說明在人工分析指數上運行GPT
5:26.465–5:26.785
5 .
[未翻譯]
5:26.795–5:30.272
6 6 on the artificial analysis index and what it produced meaning
.6的成本以及它產生的內容,也就是輸出結果。
5:30.282–5:31.224
the output .
如果你記得,這兩個模型在人工分析指數上都獲得了61分。
5:31.234–5:34.713
Both of these models if you remember scored 61 on the artificial
如果你記得,這兩個模型在人工智慧
5:34.723–5:35.745
analysis index .
.6運行整個測試,成本大約是2
5:35.755–5:38.230
Now to run the whole test with GPT 5 .
.6,成本大約是2
5:38.240–5:40.135
6 it cost about 2 .
.6,成本大約是1
5:40.145–5:42.008
6 it cost about 2 .
.1Kish。
5:42.018–5:43.620
6 it cost about 1 .
這告訴你,你以一半的成本獲得了相似水平的表現。
5:43.630–5:44.300
1Kish .
所以,是的,這是SpaceX團隊的一次重大發布,因為他們再次證明
5:44.310–5:47.465
And this tells you that you're getting similar level of performance
他們是一個你認真需要考慮的實驗室,特別是在2026年。
5:47.475–5:48.519
at half the cost .
我將更多地討論Deep Seek版本4 Pro GA。
5:48.529–5:51.893
So yeah , this is a big release for the SpaceX team because they
但基本上,我們今天看到的是,這兩次發布
5:51.903–5:54.779
just proved once again that they are a lab that you seriously start
都強調了一個事實,那就是表現和超越表現的一切
5:54.789–5:57.080
need to considering , especially in 2026 .
,是你為該表現所支付的價格,正變得
5:57.090–6:00.824
And I'm going to talk more about Deep Seek version 4 Pro GA .
越來越重要,對於所有用戶來說
6:00.834–6:04.350
But basically what we're seeing today is that both of these releases
但基本上,我們今天看到的是,這兩項發布
6:04.360–6:08.119
kind of emphasize the fact that performance and all above that
現在專注於以最低的成本創建最佳模型。
6:08.129–6:11.005
is the price at what you're getting for that performance is becoming
去年2025年,大多數實驗室只專注於我會說
6:11.015–6:13.201
more and more important for all users
創建最強大的模型。
6:13.201–6:14.612
because we see many labs
因為我們看到許多實驗室
6:14.622–6:18.379
now focusing on creating the best model at the cheapest cost .
現在專注於以最低的成本打造最佳模型。
6:18.389–6:21.907
Last year in 2025 most of the labs were just focused on I would
去年,2025年,大多數實驗室只專注於我會
6:21.917–6:24.472
say creating the strongest model .
說打造最強大的模型
6:24.482–6:25.571
Yes , cost was important ,
是的,成本很重要,
6:25.571–6:27.519
but I think most of the times the frontier
但我認為大多數時候,前沿
6:27.529–6:29.280
labs , meaning OpenAI , Anthropic ,
實驗室,也就是 OpenAI、Anthropic,
6:29.280–6:30.829
were kind of more lenient on
在這一點上比較寬鬆,
6:30.839–6:34.100
that fact because they didn't have as strong of a competition .
因為它們面臨的競爭沒有那麼激烈。
6:34.110–6:38.733
Intelligent models that are maybe not always ahead of OpenAI Enthropic
這些智能模型可能並不總是領先於 OpenAI 和 Anthropic,
6:38.743–6:41.447
, but match their performance at a fraction of the cost .
但能以更低廉的成本達到相同的性能。
6:41.457–6:45.369
So yes , price to performance ratio is becoming a critical I would
所以是的,性價比會成為我在2026年認為的關鍵指標
6:45.379–6:47.955
say indicator in 2026 .
我認為會成為一個關鍵指標。
6:47.965–6:50.861
Now , this is the updated benchmark chart after the release of
現在,這是新模型發布後更新的基準測試圖表。
6:50.871–6:51.742
the new model .
你會發現一個普遍現象,那就是它達到了
6:51.752–6:54.623
And one thing you'll see across the board is that it matches the
而且你會發現一個普遍現象,那就是它符合了
6:54.633–6:56.944
top level performance of many of the models .
這裡有趣的一點是,出於某種原因,他們沒有將 Opus
6:56.954–6:59.345
The one thing interesting over here is that they haven't put Opus
這裡有趣的一點是他們沒有把 Opus 5 放進來
6:59.355–7:00.704
5 here for some reason .
這裡有 Fable 5,我們可以將這個模型與之進行比較。
7:00.714–7:03.747
There is Fable 5 here that we can compare this model against .
但你會注意到的一點是,DeepSeek,
7:03.757–7:06.068
But one thing you'll notice is that DeepSeek ,
請記住這個模型的成本是每百萬輸入 token 43
7:06.078–7:08.335
remember this model costs 43 .
記得這個模型的費用是 43
7:08.345–7:12.791
5 cents per million input tokens and 87 cents per million output
每百萬個輸入記號5美分,每百萬個輸出記號87美分
7:12.801–7:13.511
tokens .
而其他模型在這裡的成本都要高得多。
7:13.521–7:16.391
While the other models all over here are way more expensive than
所以,如果你查看 Terminal Bench 2.
7:16.401–7:16.872
that .
那。
7:16.882–7:19.540
So the first thing if you look at terminal bench 2 .
所以如果你看 terminal bench 2 的話,第一件事是
7:19.550–7:21.629
1 the model scores 87 .
該模型的舊版本得分為 72.
7:21.639–7:22.394
9 .
九
7:22.404–7:24.656
The older version of the model was 72 .
舊版模型的得分是72
7:24.666–7:28.557
1 and the flash version which we got last week was 82 .
瘋狂的是,Fable 5 的得分為 88.
7:28.567–7:29.277
7 .
[未翻譯]
7:29.287–7:32.239
And what's crazy is that Fable 5 is 88 .
瘋狂的是,Fable 5 高達 88。
7:32.249–7:33.279
Yes , 88 .
所以這個模型僅比
7:33.289–7:34.887
So this model is only .
Fable 5 低了 0.1 分。
7:34.897–7:36.827
1 behind Fable 5 .
然後在 Cyber Gym,也就是網絡安全領域,
7:36.837–7:39.300
And then on the Cyber Gym , which is Cyber Security ,
該模型實際上超越了 Fable 5。
7:39.310–7:41.629
the model actually beats Fable 5 .
Fable 5 的得分為 83.
7:41.639–7:43.559
Fable 5 sits at 83 .
Fable 5 的得分為 83
7:43.569–7:46.683
1 while Deepseek version 4 Pro , the new one that we got today
而 Deepseek 4 Pro 版本,也就是我們今天拿到的新版本
7:46.693–7:48.044
3 .
[未翻譯]
7:48.054–7:48.765
3 .
現在,如果這告訴你什麼的話,那就是這個模型再次
7:48.775–7:51.966
Now , if this tells you something is that this model is once again
現在,如果這告訴你一些事情,那就是這個模型再次
7:51.976–7:53.807
really geared at cyber security .
確實是針對網路安全所設計的
7:53.817–7:56.850
We saw the Flash model also be geared towards cyber security and
我們看到 Flash 模型也針對網路安全領域,
7:56.860–7:59.412
becoming a model that was quite capable in that area .
並成為在該領域相當有能力的模型。
7:59.422–8:02.534
And we're seeing the same thing today with the new version 4 Pro
而今天我們看到新版本 4 Pro 也呈現相同情況,
8:02.544–8:03.975
which sits at 83 .
它在基準測試中得分 83.3,
8:03.985–8:06.858
3 outperforming Fable 5 on the benchmark .
超越了 Fable 5。
8:06.868–8:10.379
On deep software engineering , this model is not beating Fable
在深度軟體工程方面,這個模型並未超越 Fable
8:10.389–8:12.491
5 because Fable 5 sits at 70 ,
5,因為 Fable 5 得分為 70,
8:12.491–8:14.497
but the model achieves 62 .
但該模型取得了 62.7 分,
8:14.507–8:16.744
7 which is ahead of GLM 5 .
領先 GLM 5.2,
8:16.754–8:20.209
2 and a bit behind Kimmy K3 which sits at 67 .
略遜於得分 67.5 的 Kimmy K3。
8:20.219–8:20.977
5 .
。
8:20.987–8:24.115
But one thing again , this is a big jump compared to the preview
但有一點再次強調,與該模型的預覽版相比,這是巨大的進步,
8:24.125–8:26.286
version of the model which was at 12 .
預覽版得分僅為 12.8。
8:26.296–8:26.769
8 .
。
8:26.779–8:29.655
So yeah , they have really trained this model and they have improved
所以是的,他們確實訓練了這個模型,並在後端進行了改進,
8:29.665–8:32.998
it on the back end because we can clearly see in this deep software
因為我們可以清楚地從這個深度軟體
8:33.008–8:34.430
engineering benchmark .
工程基準測試中看到。
8:34.440–8:37.216
And there's a couple of other benchmarks , but one key benchmark
還有其他幾個基準測試,但這個模型擅長的一個關鍵基準測試是自動化基準,
8:37.226–8:41.353
that this model excels at is the automation bench where the model
該模型在自動化基準測試中表現出色,而該模型
8:41.363–8:42.680
achieves 31 .
而 Fable 5 為 29.1。
8:42.690–8:44.775
8 and Fable 5 29 .
。
8:44.785–8:45.332
1 .
。
8:45.342–8:48.039
So once again , the model outperforms Fable 5 .
所以再次強調,該模型優於 Fable 5。
8:48.049–8:50.268
Now this is a fraction of the cost as I mentioned .
現在,正如我提到的,這只是極小的一部分成本。
8:50.278–8:54.378
This is 43 cents and 87 for input and output respectively versus
這分別是輸入和輸出的 43 美分和 87 美分,
8:54.388–8:56.300
Fable 5 which is at 10 .
而 Fable 5 為 10.
8:56.310–8:57.083
50 .
[未翻譯]
8:57.093–8:59.787
So yes , if we look at the terminal bench for example ,
所以是的,如果我們以終端基準測試為例,
8:59.797–9:02.655
we are seeing this model achieve a result which is 0 .
我們看到這個模型取得了與目前最佳模型 Fable 5 僅有 0.1 分之差的結果,
9:02.665–9:07.078
1 behind the best model out there Fable 5 and the pricing is insane
而價格非常驚人,每百萬輸入令牌的價格分別為 43 美分和 10.87 美分對比 50 美分。
9:07.088–9:14.184
to look at 43 versus 10 87 versus 50 per million input tokens .
。
9:14.194–9:18.674
So when we turn that into a fraction , this model is 57 times cheaper
所以當我們將這轉化為比例時,這個模型便宜了 57 倍。
9:18.674–9:19.184
.
。
9:19.184–9:22.285
And earlier last week , I think I made a video about how DeepSeek
而在上週早些時候,我想我做過一個影片,說明 DeepSeek
9:22.295–9:25.740
version 4 Pro was going to achieve a result like this because there
版本 4 Pro 將如何取得這樣的結果,因為有一份投資者報告洩露。
9:25.750–9:27.495
was a investor report that leaked .
。
9:27.505–9:30.127
And when I saw the pricing where it said that it was going to match
當我看到定價時,上面寫說它會匹敵
9:30.137–9:34.470
Fable 5 or even outperform it and be at 57 times cheaper per cost
Fable 5 持平甚至超越它,並且成本每百萬輸入令牌便宜 57 倍。
9:34.470–9:35.751
.
。
9:35.751–9:37.277
When I was reading the investor report , I didn't really believe
當我閱讀那份投資者報告時,我其實不太相信
9:37.287–9:37.837
it .
它。
9:37.847–9:40.633
But today , it is proven because at least on the benchmark so far
但今天,它已經得到證實,因為至少在目前的基準測試中
9:40.643–9:43.829
, yes , I'm just saying the benchmarks , we are seeing similar
,是的,我只是說基準測試,我們看到相似的
9:43.839–9:44.948
level performance .
性能水平。
9:44.958–9:48.303
Now , if this model actually performs like that in production ,
現在,如果這個模型在實際生產中真的表現得如此,
9:48.313–9:50.141
it's a little too early to tell yet .
現在還說得太早。
9:50.151–9:52.763
We still would have to give it a couple of weeks and see how it's
我們仍然需要再給它幾週時間,看看它長期
9:52.773–9:55.867
performing in the long run because sometimes on the release day
長期表現如何還得看後續,因為有時在發布當天
9:55.877–9:59.340
the models perform quite good , but over time they kind of deteriorate
模型表現相當不錯,但隨著時間推移,它們的質量
9:59.350–10:00.380
in their quality .
可能會逐漸下降。
10:00.390–10:02.005
So , we hope DeepSeek doesn't do that .
所以,我們希望 DeepSeek 不要出現這種情況。
10:02.015–10:04.469
Historically , they haven't done that , but let's just see .
從歷史上看,他們沒有這樣做過,但讓我們拭目以待。
10:04.479–10:06.862
Just going to put it out there because right now we're basing this
我只是想把它說出來,因為目前我們是基於基準測試
10:06.872–10:08.529
off of benchmarks .
來得出這個結論。
10:08.539–10:10.051
But that's it for today's video .
但今天的影片就到這裡。
10:10.061–10:12.081
Make sure you guys are subscribed to the channel .
記得訂閱我們的頻道
10:12.091–10:15.369
Follow our new newsletter as well at universeofai .
也請訂閱我們的新聞通訊,網址是 universeofai .
10:15.379–10:15.756
behiiv .
behiv
10:15.766–10:19.087
com as well as subscribe to the main channel World of AI and support
以及訂閱 World of AI 主頻道並支持
10:19.097–10:22.063
us on X by following the Universe of AIZ as well .
我們,關注 Universe of AIZ。
10:22.073–10:23.114
Until then ,
在那之前,
10:23.114–10:24.618
I'll see you guys
我們下次見
10:24.618–10:25.659
in the next
在下一個