實際影片長度:10:26.000。原文、繁中、雙語可點擊句子跳轉影片。
0:00.250–0:03.749
So , Deepseek version 4 Pro is officially out today .
0:03.759–0:06.460
Now , you might be confused because this model was technically
0:06.470–0:09.810
out , but it was in general availability , but this is the official
0:09.820–0:12.681
release , meaning that they have trained the model a little bit
0:12.691–0:16.111
more and produced a stronger version of the model that is available
0:16.121–0:16.590
today .
0:16.600–0:19.337
Now , if I were to summarize today's release in one simple sentence
0:19.347–0:22.651
, it is that the price to performance is becoming a very ,
0:22.661–0:25.682
very important thing because not only do we have a new model from
0:25.692–0:29.074
the Deepseek team , the SpaceX team also dropped Grock 4 .
0:29.084–0:29.640
6 six .
0:29.650–0:33.063
And both of these models are competing with the best Frontier Labs
0:33.073–0:35.066
at a fraction of their cost .
0:35.076–0:36.431
Let's start with Gro 4 .
0:36.441–0:36.827
6 .
0:36.837–0:39.761
And one thing I'm going to say is that I'm genuinely surprised
0:39.771–0:41.029
by the SpaceX team .
0:41.039–0:43.408
I'm not trying to glaze them or Elon Musk or anything .
0:43.418–0:45.311
I was just very critical of this lab .
0:45.321–0:48.671
In 2025 , they dropped Grog 4 and other models like that ,
0:48.681–0:52.371
but I wasn't really , you know , mind blown with their performance
0:52.381–0:55.400
because they're pretty they're pretty subpar compared to any of
0:55.410–0:56.743
the other labs out there .
0:56.753–1:00.140
But in 2026 , it looks like things have kind of changed a little
1:00.150–1:01.212
because Grock 4 .
1:01.222–1:04.982
5 was quite competitive based off what it cost and Grock 4 .
1:04.992–1:07.268
6 is actually not too bad .
1:07.278–1:09.740
If we take a look at the benchmarks , for example ,
1:09.750–1:13.248
if we start with the artificial analysis intelligence index ,
1:13.258–1:17.633
this model achieves a 61 and Fable 5 is at 62 .
1:17.643–1:20.583
Now , what's really important to remember is that once again ,
1:20.593–1:23.374
this model is quite cheap compared to Fable 5 .
1:23.384–1:27.600
This is about 2 per million input tokens and 6 per million output
1:27.610–1:28.245
tokens .
1:28.255–1:32.844
In Fable 5 sits at 10 per million input tokens and 50 per million
1:32.854–1:33.893
output tokens .
1:33.903–1:36.959
And this is why you start to appreciate this release a little bit
1:36.969–1:37.363
more .
1:37.373–1:40.540
It might not beat the performance of the best models is matching
1:40.550–1:41.037
them .
1:41.047–1:42.390
Even GPT 5 .
1:42.400–1:46.161
6 so which is a pretty capable model and this is set at max on
1:46.171–1:48.126
the artificial analysis index .
1:48.136–1:50.828
This model achieves 61 and Grok 4 .
1:50.838–1:51.632
6 61 .
1:51.642–1:53.804
So it ties it and it's much cheaper .
1:53.814–1:56.216
And then even on all of these other benchmarks ,
1:56.226–1:59.343
for example , the Code Val one , it actually beats Fable 5 ,
1:59.353–2:01.186
which is at 1741 .
2:01.196–2:02.427
And then GPT 5 .
2:02.437–2:04.309
6 , it's at 1728 .
2:04.319–2:06.713
I'm not sure why they didn't choose Opus 5 as well ,
2:06.723–2:09.916
but I guess they wanted to choose the quote unquote strongest model
2:09.926–2:10.959
lineup from each lab .
2:10.969–2:13.903
And they chose Fable 5 for Enthropic , which is fair .
2:13.913–2:16.917
And then Deep Software Engineering one , which is a critical benchmark
2:16.917–2:19.286
.
2:19.286–2:20.585
This model doesn't beat Fable 5 or GPT 5 .
2:20.595–2:22.322
6 six soul , but it gets close to it .
2:22.332–2:23.255
It's 65 .
2:23.265–2:25.839
9 and Fable 5 sits at 70 .
2:25.849–2:28.556
But if you're getting results that are pretty close and the model
2:28.566–2:31.432
is five times cheaper , I wouldn't be too disappointed with this
2:31.442–2:31.991
result .
2:32.001–2:34.068
And then same thing with Cursor Bench 3 .
2:34.078–2:36.221
2 , the model achieves 69 .
2:36.231–2:37.904
9 , funny number .
2:37.914–2:39.617
And Fable 5 sits at 70 .
2:39.627–2:40.141
5 .
2:40.151–2:41.260
So once again , closer .
2:41.270–2:42.199
And Grock 4 .
2:42.209–2:43.693
6 beats GPT 5 .
2:43.703–2:44.860
6 .
2:44.870–2:46.375
Same thing with the Frontier Code .
2:46.385–2:49.287
It gets close to Fable 5 , beats GPT 5 .
2:49.297–2:49.987
6 .
2:49.997–2:52.808
So yeah , this model is actually available in cursor .
2:52.818–2:56.515
So the partnership with cursor or I guess the acquisition has really
2:56.525–3:00.540
helped SpaceX make some strides in the AI space this year and Grock
3:00.550–3:03.740
build is something that you know maybe not a lot of us have been
3:03.750–3:07.341
using so far but it's probably going to be another platform like
3:07.351–3:10.793
codeex or cloud code that we start to use but obviously cursor
3:10.803–3:11.997
is quite strong as well .
3:12.007–3:15.369
So you have options available for you to use them in both .
3:15.379–3:18.660
And one thing to note is that they're offering two times usage
3:18.670–3:21.287
inside Grok Build and Cursor for the first week .
3:21.297–3:23.992
So if you just want to try it out , see what you feel about it
3:24.002–3:27.189
, then you know it might be worth trying it out right now cuz you
3:27.199–3:29.304
get double the usage in the first week .
3:29.314–3:31.907
Now if we were to take a look at some of the outputs that people
3:31.917–3:33.916
have been generating with Grok 4 .
3:33.926–3:37.824
6 , what we're looking at right now is a Falcon 9 booster return
3:37.834–3:38.939
sequence simulation .
3:38.949–3:42.050
And this was done in a single HTML file .
3:42.060–3:45.479
And as I said , if you expected to get this type of output from
3:45.489–3:49.852
Grock in 2025 , you would be kind of surprised because you wouldn't
3:49.862–3:52.240
expect something like this to be generated with Grock .
3:52.250–3:54.958
But now it looks like we have to start taking the Grok team a
3:54.968–3:58.195
little bit more serious because this output is quite competitive
3:58.195–4:00.295
.
4:00.295–4:01.350
And obviously , we're just looking at a simulation and we're just
4:01.360–4:03.987
basing it off of a visual representation .
4:03.997–4:06.624
But if the model is able to produce something like this consistently
4:06.634–4:10.525
, then I would expect a lot of people to start adopting Grock because
4:10.535–4:13.368
number one , it is cheaper than the other labs at least at the
4:13.378–4:15.895
moment because we don't know if this pricing strategy is going
4:15.905–4:18.895
to be sustainable for the SpaceX team in the long run .
4:18.905–4:21.659
But at least for now , their models are definitely cheaper compared
4:21.669–4:22.608
to the others .
4:22.618–4:26.003
And as I mentioned , this model excelled at the artificial analysis
4:26.013–4:26.635
index .
4:26.645–4:30.031
This model jumped to probably number four model .
4:30.041–4:32.180
It's tied to GPT 5 .
4:32.190–4:33.370
6 pretty much similar .
4:33.380–4:36.579
So you could say number three as well , but the models before that
4:36.589–4:40.896
are Fable 5 and Opus 5 , which are only above the model by about
4:40.906–4:43.454
a 1 or a 2 difference .
4:43.464–4:46.492
So yeah , even on this intelligence index , which if you're not
4:46.502–4:48.570
familiar with has nine evaluations .
4:48.580–4:52.646
So on all of these evaluations , it's kind of matching almost Fable
4:52.656–4:55.046
5 performance , which is crazy to see .
4:55.056–4:58.403
Before we continue , we just launched the Universe of AI newsletter
4:58.403–5:00.629
.
5:00.629–5:01.761
If you want to stay on top of AI news without having to hunt for
5:01.771–5:03.595
it , link is in the description .
5:03.605–5:04.710
Don't miss out .
5:04.720–5:07.800
And what you see on screen right now is a racing game that Grok
5:07.810–5:08.320
4 .
5:08.330–5:09.127
6 build .
5:09.137–5:12.168
And based off of this post , the model took about 1 minute and
5:12.178–5:15.191
it was a fiveword prompt , which was a create a simple racing game
5:15.201–5:16.269
in HTML .
5:16.279–5:19.367
So , if you're able to generate something like this easily using
5:19.377–5:20.100
Grock 4 .
5:20.110–5:22.300
6 , I think a lot of people will be happy .
5:22.310–5:26.455
And this is a more detailed analysis of what it costs to run GPT
5:26.465–5:26.785
5 .
5:26.795–5:30.272
6 6 on the artificial analysis index and what it produced meaning
5:30.282–5:31.224
the output .
5:31.234–5:34.713
Both of these models if you remember scored 61 on the artificial
5:34.723–5:35.745
analysis index .
5:35.755–5:38.230
Now to run the whole test with GPT 5 .
5:38.240–5:40.135
6 it cost about 2 .
5:40.145–5:42.008
6 it cost about 2 .
5:42.018–5:43.620
6 it cost about 1 .
5:43.630–5:44.300
1Kish .
5:44.310–5:47.465
And this tells you that you're getting similar level of performance
5:47.475–5:48.519
at half the cost .
5:48.529–5:51.893
So yeah , this is a big release for the SpaceX team because they
5:51.903–5:54.779
just proved once again that they are a lab that you seriously start
5:54.789–5:57.080
need to considering , especially in 2026 .
5:57.090–6:00.824
And I'm going to talk more about Deep Seek version 4 Pro GA .
6:00.834–6:04.350
But basically what we're seeing today is that both of these releases
6:04.360–6:08.119
kind of emphasize the fact that performance and all above that
6:08.129–6:11.005
is the price at what you're getting for that performance is becoming
6:11.015–6:13.201
more and more important for all users
6:13.201–6:14.612
because we see many labs
6:14.622–6:18.379
now focusing on creating the best model at the cheapest cost .
6:18.389–6:21.907
Last year in 2025 most of the labs were just focused on I would
6:21.917–6:24.472
say creating the strongest model .
6:24.482–6:25.571
Yes , cost was important ,
6:25.571–6:27.519
but I think most of the times the frontier
6:27.529–6:29.280
labs , meaning OpenAI , Anthropic ,
6:29.280–6:30.829
were kind of more lenient on
6:30.839–6:34.100
that fact because they didn't have as strong of a competition .
6:34.110–6:38.733
Intelligent models that are maybe not always ahead of OpenAI Enthropic
6:38.743–6:41.447
, but match their performance at a fraction of the cost .
6:41.457–6:45.369
So yes , price to performance ratio is becoming a critical I would
6:45.379–6:47.955
say indicator in 2026 .
6:47.965–6:50.861
Now , this is the updated benchmark chart after the release of
6:50.871–6:51.742
the new model .
6:51.752–6:54.623
And one thing you'll see across the board is that it matches the
6:54.633–6:56.944
top level performance of many of the models .
6:56.954–6:59.345
The one thing interesting over here is that they haven't put Opus
6:59.355–7:00.704
5 here for some reason .
7:00.714–7:03.747
There is Fable 5 here that we can compare this model against .
7:03.757–7:06.068
But one thing you'll notice is that DeepSeek ,
7:06.078–7:08.335
remember this model costs 43 .
7:08.345–7:12.791
5 cents per million input tokens and 87 cents per million output
7:12.801–7:13.511
tokens .
7:13.521–7:16.391
While the other models all over here are way more expensive than
7:16.401–7:16.872
that .
7:16.882–7:19.540
So the first thing if you look at terminal bench 2 .
7:19.550–7:21.629
1 the model scores 87 .
7:21.639–7:22.394
9 .
7:22.404–7:24.656
The older version of the model was 72 .
7:24.666–7:28.557
1 and the flash version which we got last week was 82 .
7:28.567–7:29.277
7 .
7:29.287–7:32.239
And what's crazy is that Fable 5 is 88 .
7:32.249–7:33.279
Yes , 88 .
7:33.289–7:34.887
So this model is only .
7:34.897–7:36.827
1 behind Fable 5 .
7:36.837–7:39.300
And then on the Cyber Gym , which is Cyber Security ,
7:39.310–7:41.629
the model actually beats Fable 5 .
7:41.639–7:43.559
Fable 5 sits at 83 .
7:43.569–7:46.683
1 while Deepseek version 4 Pro , the new one that we got today
7:46.693–7:48.044
3 .
7:48.054–7:48.765
3 .
7:48.775–7:51.966
Now , if this tells you something is that this model is once again
7:51.976–7:53.807
really geared at cyber security .
7:53.817–7:56.850
We saw the Flash model also be geared towards cyber security and
7:56.860–7:59.412
becoming a model that was quite capable in that area .
7:59.422–8:02.534
And we're seeing the same thing today with the new version 4 Pro
8:02.544–8:03.975
which sits at 83 .
8:03.985–8:06.858
3 outperforming Fable 5 on the benchmark .
8:06.868–8:10.379
On deep software engineering , this model is not beating Fable
8:10.389–8:12.491
5 because Fable 5 sits at 70 ,
8:12.491–8:14.497
but the model achieves 62 .
8:14.507–8:16.744
7 which is ahead of GLM 5 .
8:16.754–8:20.209
2 and a bit behind Kimmy K3 which sits at 67 .
8:20.219–8:20.977
5 .
8:20.987–8:24.115
But one thing again , this is a big jump compared to the preview
8:24.125–8:26.286
version of the model which was at 12 .
8:26.296–8:26.769
8 .
8:26.779–8:29.655
So yeah , they have really trained this model and they have improved
8:29.665–8:32.998
it on the back end because we can clearly see in this deep software
8:33.008–8:34.430
engineering benchmark .
8:34.440–8:37.216
And there's a couple of other benchmarks , but one key benchmark
8:37.226–8:41.353
that this model excels at is the automation bench where the model
8:41.363–8:42.680
achieves 31 .
8:42.690–8:44.775
8 and Fable 5 29 .
8:44.785–8:45.332
1 .
8:45.342–8:48.039
So once again , the model outperforms Fable 5 .
8:48.049–8:50.268
Now this is a fraction of the cost as I mentioned .
8:50.278–8:54.378
This is 43 cents and 87 for input and output respectively versus
8:54.388–8:56.300
Fable 5 which is at 10 .
8:56.310–8:57.083
50 .
8:57.093–8:59.787
So yes , if we look at the terminal bench for example ,
8:59.797–9:02.655
we are seeing this model achieve a result which is 0 .
9:02.665–9:07.078
1 behind the best model out there Fable 5 and the pricing is insane
9:07.088–9:14.184
to look at 43 versus 10 87 versus 50 per million input tokens .
9:14.194–9:18.674
So when we turn that into a fraction , this model is 57 times cheaper
9:18.674–9:19.184
.
9:19.184–9:22.285
And earlier last week , I think I made a video about how DeepSeek
9:22.295–9:25.740
version 4 Pro was going to achieve a result like this because there
9:25.750–9:27.495
was a investor report that leaked .
9:27.505–9:30.127
And when I saw the pricing where it said that it was going to match
9:30.137–9:34.470
Fable 5 or even outperform it and be at 57 times cheaper per cost
9:34.470–9:35.751
.
9:35.751–9:37.277
When I was reading the investor report , I didn't really believe
9:37.287–9:37.837
it .
9:37.847–9:40.633
But today , it is proven because at least on the benchmark so far
9:40.643–9:43.829
, yes , I'm just saying the benchmarks , we are seeing similar
9:43.839–9:44.948
level performance .
9:44.958–9:48.303
Now , if this model actually performs like that in production ,
9:48.313–9:50.141
it's a little too early to tell yet .
9:50.151–9:52.763
We still would have to give it a couple of weeks and see how it's
9:52.773–9:55.867
performing in the long run because sometimes on the release day
9:55.877–9:59.340
the models perform quite good , but over time they kind of deteriorate
9:59.350–10:00.380
in their quality .
10:00.390–10:02.005
So , we hope DeepSeek doesn't do that .
10:02.015–10:04.469
Historically , they haven't done that , but let's just see .
10:04.479–10:06.862
Just going to put it out there because right now we're basing this
10:06.872–10:08.529
off of benchmarks .
10:08.539–10:10.051
But that's it for today's video .
10:10.061–10:12.081
Make sure you guys are subscribed to the channel .
10:12.091–10:15.369
Follow our new newsletter as well at universeofai .
10:15.379–10:15.756
behiiv .
10:15.766–10:19.087
com as well as subscribe to the main channel World of AI and support
10:19.097–10:22.063
us on X by following the Universe of AIZ as well .
10:22.073–10:23.114
Until then ,
10:23.114–10:24.618
I'll see you guys
10:24.618–10:25.659
in the next
0:00.250–0:03.749
所以,Deepseek 版本 4 Pro 今天正式推出了。
0:03.759–0:06.460
現在,你可能會感到困惑,因為這個模型在技術上
0:06.470–0:09.810
已經發布,但處於一般可用性階段,而這是官方
0:09.820–0:12.681
發布,意味著他們對模型進行了更多訓練,
0:12.691–0:16.111
並產出了一個更強大的模型版本,今天即可使用。
0:16.121–0:16.590
今天。
0:16.600–0:19.337
現在,如果我要用一句簡單的話來總結今天的發布,
0:19.347–0:22.651
那就是性價比正變得非常、非常重要的一環,因為
0:22.661–0:25.682
非常重要的一點,因為我們不僅擁有一個來自
0:25.692–0:29.074
Deepseek 團隊的新模型,SpaceX 團隊也推出了 Grock 4。
0:29.084–0:29.640
6.6。
0:29.650–0:33.063
這兩個模型都在以極低的成本與 Frontier Labs 的最佳模型競爭。
0:33.073–0:35.066
讓我們從 Gro 4 開始。
0:35.076–0:36.431
讓我們從 Gro 4 開始。
0:36.441–0:36.827
我要說的一件事是,我對 SpaceX 團隊真的感到驚訝。
0:36.837–0:39.761
我不是在試圖吹捧他們或埃隆·馬斯克或什麼的。
0:39.771–0:41.029
我只是對這個實驗室非常批評。
0:41.039–0:43.408
在 2025 年,他們推出了 Grog 4 和其他類似的模型,
0:43.418–0:45.311
但我並沒有真的,你知道,對他們的表現感到驚艷,
0:45.321–0:48.671
因為與其他實驗室相比,他們的表現相當差。
0:48.681–0:52.371
但在 2026 年,情況似乎有點改變,
0:52.381–0:55.400
因為跟其他實驗室相比,它們的表現相當遜色
0:55.410–0:56.743
5 在成本基礎上相當有競爭力,而 Grock 4。
0:56.753–1:00.140
但在 2026 年,情況似乎有點改變了
1:00.150–1:01.212
如果我們查看基準測試,例如,
1:01.222–1:04.982
如果我們從 Artificial Analysis 智能指數開始,
1:04.992–1:07.268
這個模型達到了 61 分,而 Fable 5 為 62 分。
1:07.278–1:09.740
現在,真正重要的是要記住,再一次,
1:09.750–1:13.248
這個模型與 Fable 5 相比非常便宜。
1:13.258–1:17.633
這大約是每百萬輸入標記 2 美元,每百萬輸出
1:17.643–1:20.583
現在,真正重要的是要記住,再一次,
1:20.593–1:23.374
在 Fable 5 中,每百萬輸入標記為 10 美元,每百萬
1:23.384–1:27.600
這大約是每百萬個輸入 token 2 個,每百萬個輸出
1:27.610–1:28.245
這就是為什麼你開始稍微欣賞這次發布的原因。
1:28.255–1:32.844
它可能無法擊敗最佳模型的表現,但正在與它們匹敵。
1:32.854–1:33.893
輸出標記
1:33.903–1:36.959
這就是為什麼你會開始稍微欣賞這次發布的原因
1:36.969–1:37.363
更多
1:37.373–1:40.540
它可能無法擊敗最佳模型的表現,但正在追趕
1:40.550–1:41.037
它們
1:41.047–1:42.390
即使是 GPT 5。
1:42.400–1:46.161
6,這是一個相當強大的模型,並且在人工分析指數上設定為最高。
1:46.171–1:48.126
人工分析指數
1:48.136–1:50.828
該模型達到 61 分,Grok 4。
1:50.838–1:51.632
6,61 分。
1:51.642–1:53.804
所以它並列第一,而且便宜得多。
1:53.814–1:56.216
然後在這些其他基準測試中,
1:56.226–1:59.343
例如 Code Val 基準,它實際上擊敗了 Fable 5,
1:59.353–2:01.186
後者為 1741 分。
2:01.196–2:02.427
然後是 GPT 5。
2:02.437–2:04.309
6,它為 1728 分。
2:04.319–2:06.713
我不確定為什麼他們沒有選擇 Opus 5,
2:06.723–2:09.916
但我猜他們想選擇所謂的「最強」模型
2:09.926–2:10.959
陣容來自每個實驗室。
2:10.969–2:13.903
他們為 Enthropic 選擇了 Fable 5,這很公平。
2:13.913–2:16.917
然後是 Deep Software Engineering 基準,這是一個關鍵基準
2:16.917–2:19.286
2:19.286–2:20.585
該模型沒有擊敗 Fable 5 或 GPT 5。
2:20.595–2:22.322
6,6 分,但非常接近。
2:22.332–2:23.255
它是 65。
2:23.265–2:25.839
9,而 Fable 5 為 70 分。
2:25.849–2:28.556
但如果你得到的結果非常接近,而且模型
2:28.566–2:31.432
如果價格便宜五倍,我不會對這個結果太失望
2:31.442–2:31.991
結果感到太失望。
2:32.001–2:34.068
然後是 Cursor Bench 3。
2:34.078–2:36.221
2,該模型達到 69。
2:36.231–2:37.904
9,有趣的數字。
2:37.914–2:39.617
而 Fable 5 為 70。
2:39.627–2:40.141
[未翻譯]
2:40.151–2:41.260
所以再一次,更接近。
2:41.270–2:42.199
而 Grock 4。
2:42.209–2:43.693
6 擊敗了 GPT 5。
2:43.703–2:44.860
[未翻譯]
2:44.870–2:46.375
在 Frontier Code 方面也是如此。
2:46.385–2:49.287
它非常接近 Fable 5,擊敗了 GPT 5。
2:49.297–2:49.987
[未翻譯]
2:49.997–2:52.808
所以是的,該模型實際上在 cursor 中可用。
2:52.818–2:56.515
所以與 cursor 的合作關係,或者我猜是收購,確實
2:56.525–3:00.540
幫助 SpaceX 在今年在 AI 領域取得了一些進展,而 Grock
3:00.550–3:03.740
build 是你知道也許我們中沒有很多人一直在
3:03.750–3:07.341
到目前為止還沒有人使用,但未來我們可能會開始使用類似 CodeEx 或 Cloud Code 的這類平台,不過顯然 Cursor 也很強大。
3:07.351–3:10.793
因此,你有可用的選項,可以在兩者中選擇使用。
3:10.803–3:11.997
並且有一件事需要注意,Grok Build 和 Cursor 在第一週提供兩倍的用量。
3:12.007–3:15.369
所以如果你只是想試試看,看看你的感覺如何,那麼現在嘗試可能值得,因為你在第一週可以獲得雙倍的用量。
3:15.379–3:18.660
現在,如果我們來看看人們使用 Grok 4.6 所生成的一些輸出結果。
3:18.670–3:21.287
我們現在看到的是 Falcon 9 助推器返回序列的模擬。
3:21.297–3:23.992
這是在單一 HTML 檔案中完成的。
3:24.002–3:27.189
正如我所說,如果你預期在 2025 年從 Grok 獲得這種類型的輸出,你會感到相當驚訝,因為你不會預期 Grok 能生成這樣的內容。
3:27.199–3:29.304
但現在看來,我們必須開始更認真地看待 Grok 團隊,因為這項輸出相當具有競爭力。
3:29.314–3:31.907
顯然,我們只是在看一個模擬,並且僅基於視覺呈現來評估。
3:31.917–3:33.916
都是使用 Grok 4 生成的。
3:33.926–3:37.824
但至少目前,他們的模型確實比其他模型便宜。
3:37.834–3:38.939
正如我所提到的,該模型在人工分析指數中表現出色。
3:38.949–3:42.050
該模型躍升至可能排名第四的模型。
3:42.060–3:45.479
它與 GPT 5.6 並列,表現非常相似。
3:45.489–3:49.852
所以你也可以說它是排名第三,但在此之前的是 Falcon 5 和 Opus 5,它們僅比該模型高出約 1 或 2 分的差距。
3:49.862–3:52.240
所以是的,即使是在這個智力指數上,如果你不熟悉,它包含九項評估。
3:52.250–3:54.958
所以在所有這些評估中,它幾乎與 Falcon 5 的表現相匹配。
3:54.968–3:58.195
稍微更嚴肅一些,因為這項成果相當具有競爭力
3:58.195–4:00.295
4:00.295–4:01.350
而且顯然,我們只是在看模擬結果,並且只是
4:01.360–4:03.987
基於視覺呈現來進行評估。
4:03.997–4:06.624
但如果該模型能夠持續產生類似的成果
4:06.634–4:10.525
,那麼我預期會有很多人開始採用 Grock,因為
4:10.535–4:13.368
第一,它比其他實驗室便宜,至少目前是如此
4:13.378–4:15.895
因為我們不知道這種定價策略
4:15.905–4:18.895
對 SpaceX 團隊在長期來看是否具可持續性
4:18.905–4:21.659
但至少目前,他們的模型確實比其他模型便宜
4:21.669–4:22.608
相較於其他模型
4:22.618–4:26.003
正如我所提到的,該模型在人工分析指數方面表現出色
4:26.013–4:26.635
指數
4:26.645–4:30.031
該模型躍升至可能排名第四的模型
4:30.041–4:32.180
它與 GPT 5 並列。
4:32.190–4:33.370
與 6 幾乎相同。
4:33.380–4:36.579
所以你也可以說它是第三名,但在那之前的模型
4:36.589–4:40.896
是 Fable 5 和 Opus 5,它們僅比該模型高出約
4:40.906–4:43.454
1 或 2 的差距。
4:43.464–4:46.492
所以是的,即使是在這個智力指數上,如果你不
4:46.502–4:48.570
熟悉的話,它包含九項評估。
4:48.580–4:52.646
所以在所有這些評估中,它幾乎與 Fable
4:52.656–4:55.046
5的表現,這在視覺上令人瘋狂。
4:55.056–4:58.403
在我們繼續之前,我們剛剛推出了AI宇宙通訊電子報
4:58.403–5:00.629
5:00.629–5:01.761
如果你想隨時掌握AI新聞,而不必費心去搜尋
5:01.771–5:03.595
它,連結在描述中。
5:03.605–5:04.710
別錯過。
5:04.720–5:07.800
而你現在在螢幕上看到的是一款由Grok
5:07.810–5:08.320
[未翻譯]
5:08.330–5:09.127
.6版本構建的賽車遊戲。
5:09.137–5:12.168
根據這篇帖子,該模型大約花了1分鐘
5:12.178–5:15.191
,這是一個五個字的提示,內容是創建一個簡單的賽車遊戲
5:15.201–5:16.269
,使用HTML。
5:16.279–5:19.367
所以,如果你能輕鬆地生成像這樣的東西,使用
5:19.377–5:20.100
所以,如果你能輕鬆地使用Grok 4
5:20.110–5:22.300
.6生成這樣的東西,我認為許多人會很高興。
5:22.310–5:26.455
這是一份更詳細的分析,說明在人工分析指數上運行GPT
5:26.465–5:26.785
[未翻譯]
5:26.795–5:30.272
.6的成本以及它產生的內容,也就是輸出結果。
5:30.282–5:31.224
如果你記得,這兩個模型在人工分析指數上都獲得了61分。
5:31.234–5:34.713
如果你記得,這兩個模型在人工智慧
5:34.723–5:35.745
.6運行整個測試,成本大約是2
5:35.755–5:38.230
.6,成本大約是2
5:38.240–5:40.135
.6,成本大約是1
5:40.145–5:42.008
.1Kish。
5:42.018–5:43.620
這告訴你,你以一半的成本獲得了相似水平的表現。
5:43.630–5:44.300
所以,是的,這是SpaceX團隊的一次重大發布,因為他們再次證明
5:44.310–5:47.465
他們是一個你認真需要考慮的實驗室,特別是在2026年。
5:47.475–5:48.519
我將更多地討論Deep Seek版本4 Pro GA。
5:48.529–5:51.893
但基本上,我們今天看到的是,這兩次發布
5:51.903–5:54.779
都強調了一個事實,那就是表現和超越表現的一切
5:54.789–5:57.080
,是你為該表現所支付的價格,正變得
5:57.090–6:00.824
越來越重要,對於所有用戶來說
6:00.834–6:04.350
但基本上,我們今天看到的是,這兩項發布
6:04.360–6:08.119
現在專注於以最低的成本創建最佳模型。
6:08.129–6:11.005
去年2025年,大多數實驗室只專注於我會說
6:11.015–6:13.201
創建最強大的模型。
6:13.201–6:14.612
因為我們看到許多實驗室
6:14.622–6:18.379
現在專注於以最低的成本打造最佳模型。
6:18.389–6:21.907
去年,2025年,大多數實驗室只專注於我會
6:21.917–6:24.472
說打造最強大的模型
6:24.482–6:25.571
是的,成本很重要,
6:25.571–6:27.519
但我認為大多數時候,前沿
6:27.529–6:29.280
實驗室,也就是 OpenAI、Anthropic,
6:29.280–6:30.829
在這一點上比較寬鬆,
6:30.839–6:34.100
因為它們面臨的競爭沒有那麼激烈。
6:34.110–6:38.733
這些智能模型可能並不總是領先於 OpenAI 和 Anthropic,
6:38.743–6:41.447
但能以更低廉的成本達到相同的性能。
6:41.457–6:45.369
所以是的,性價比會成為我在2026年認為的關鍵指標
6:45.379–6:47.955
我認為會成為一個關鍵指標。
6:47.965–6:50.861
現在,這是新模型發布後更新的基準測試圖表。
6:50.871–6:51.742
你會發現一個普遍現象,那就是它達到了
6:51.752–6:54.623
而且你會發現一個普遍現象,那就是它符合了
6:54.633–6:56.944
這裡有趣的一點是,出於某種原因,他們沒有將 Opus
6:56.954–6:59.345
這裡有趣的一點是他們沒有把 Opus 5 放進來
6:59.355–7:00.704
這裡有 Fable 5,我們可以將這個模型與之進行比較。
7:00.714–7:03.747
但你會注意到的一點是,DeepSeek,
7:03.757–7:06.068
請記住這個模型的成本是每百萬輸入 token 43
7:06.078–7:08.335
記得這個模型的費用是 43
7:08.345–7:12.791
每百萬個輸入記號5美分,每百萬個輸出記號87美分
7:12.801–7:13.511
而其他模型在這裡的成本都要高得多。
7:13.521–7:16.391
所以,如果你查看 Terminal Bench 2.
7:16.401–7:16.872
那。
7:16.882–7:19.540
所以如果你看 terminal bench 2 的話,第一件事是
7:19.550–7:21.629
該模型的舊版本得分為 72.
7:21.639–7:22.394
7:22.404–7:24.656
舊版模型的得分是72
7:24.666–7:28.557
瘋狂的是,Fable 5 的得分為 88.
7:28.567–7:29.277
[未翻譯]
7:29.287–7:32.239
瘋狂的是,Fable 5 高達 88。
7:32.249–7:33.279
所以這個模型僅比
7:33.289–7:34.887
Fable 5 低了 0.1 分。
7:34.897–7:36.827
然後在 Cyber Gym,也就是網絡安全領域,
7:36.837–7:39.300
該模型實際上超越了 Fable 5。
7:39.310–7:41.629
Fable 5 的得分為 83.
7:41.639–7:43.559
Fable 5 的得分為 83
7:43.569–7:46.683
而 Deepseek 4 Pro 版本,也就是我們今天拿到的新版本
7:46.693–7:48.044
[未翻譯]
7:48.054–7:48.765
現在,如果這告訴你什麼的話,那就是這個模型再次
7:48.775–7:51.966
現在,如果這告訴你一些事情,那就是這個模型再次
7:51.976–7:53.807
確實是針對網路安全所設計的
7:53.817–7:56.850
我們看到 Flash 模型也針對網路安全領域,
7:56.860–7:59.412
並成為在該領域相當有能力的模型。
7:59.422–8:02.534
而今天我們看到新版本 4 Pro 也呈現相同情況,
8:02.544–8:03.975
它在基準測試中得分 83.3,
8:03.985–8:06.858
超越了 Fable 5。
8:06.868–8:10.379
在深度軟體工程方面,這個模型並未超越 Fable
8:10.389–8:12.491
5,因為 Fable 5 得分為 70,
8:12.491–8:14.497
但該模型取得了 62.7 分,
8:14.507–8:16.744
領先 GLM 5.2,
8:16.754–8:20.209
略遜於得分 67.5 的 Kimmy K3。
8:20.219–8:20.977
8:20.987–8:24.115
但有一點再次強調,與該模型的預覽版相比,這是巨大的進步,
8:24.125–8:26.286
預覽版得分僅為 12.8。
8:26.296–8:26.769
8:26.779–8:29.655
所以是的,他們確實訓練了這個模型,並在後端進行了改進,
8:29.665–8:32.998
因為我們可以清楚地從這個深度軟體
8:33.008–8:34.430
工程基準測試中看到。
8:34.440–8:37.216
還有其他幾個基準測試,但這個模型擅長的一個關鍵基準測試是自動化基準,
8:37.226–8:41.353
該模型在自動化基準測試中表現出色,而該模型
8:41.363–8:42.680
而 Fable 5 為 29.1。
8:42.690–8:44.775
8:44.785–8:45.332
8:45.342–8:48.039
所以再次強調,該模型優於 Fable 5。
8:48.049–8:50.268
現在,正如我提到的,這只是極小的一部分成本。
8:50.278–8:54.378
這分別是輸入和輸出的 43 美分和 87 美分,
8:54.388–8:56.300
而 Fable 5 為 10.
8:56.310–8:57.083
[未翻譯]
8:57.093–8:59.787
所以是的,如果我們以終端基準測試為例,
8:59.797–9:02.655
我們看到這個模型取得了與目前最佳模型 Fable 5 僅有 0.1 分之差的結果,
9:02.665–9:07.078
而價格非常驚人,每百萬輸入令牌的價格分別為 43 美分和 10.87 美分對比 50 美分。
9:07.088–9:14.184
9:14.194–9:18.674
所以當我們將這轉化為比例時,這個模型便宜了 57 倍。
9:18.674–9:19.184
9:19.184–9:22.285
而在上週早些時候,我想我做過一個影片,說明 DeepSeek
9:22.295–9:25.740
版本 4 Pro 將如何取得這樣的結果,因為有一份投資者報告洩露。
9:25.750–9:27.495
9:27.505–9:30.127
當我看到定價時,上面寫說它會匹敵
9:30.137–9:34.470
Fable 5 持平甚至超越它,並且成本每百萬輸入令牌便宜 57 倍。
9:34.470–9:35.751
9:35.751–9:37.277
當我閱讀那份投資者報告時,我其實不太相信
9:37.287–9:37.837
它。
9:37.847–9:40.633
但今天,它已經得到證實,因為至少在目前的基準測試中
9:40.643–9:43.829
,是的,我只是說基準測試,我們看到相似的
9:43.839–9:44.948
性能水平。
9:44.958–9:48.303
現在,如果這個模型在實際生產中真的表現得如此,
9:48.313–9:50.141
現在還說得太早。
9:50.151–9:52.763
我們仍然需要再給它幾週時間,看看它長期
9:52.773–9:55.867
長期表現如何還得看後續,因為有時在發布當天
9:55.877–9:59.340
模型表現相當不錯,但隨著時間推移,它們的質量
9:59.350–10:00.380
可能會逐漸下降。
10:00.390–10:02.005
所以,我們希望 DeepSeek 不要出現這種情況。
10:02.015–10:04.469
從歷史上看,他們沒有這樣做過,但讓我們拭目以待。
10:04.479–10:06.862
我只是想把它說出來,因為目前我們是基於基準測試
10:06.872–10:08.529
來得出這個結論。
10:08.539–10:10.051
但今天的影片就到這裡。
10:10.061–10:12.081
記得訂閱我們的頻道
10:12.091–10:15.369
也請訂閱我們的新聞通訊,網址是 universeofai .
10:15.379–10:15.756
behiv
10:15.766–10:19.087
以及訂閱 World of AI 主頻道並支持
10:19.097–10:22.063
我們,關注 Universe of AIZ。
10:22.073–10:23.114
在那之前,
10:23.114–10:24.618
我們下次見
10:24.618–10:25.659
在下一個
0:00.250–0:03.749
So , Deepseek version 4 Pro is officially out today .
所以,Deepseek 版本 4 Pro 今天正式推出了。
0:03.759–0:06.460
Now , you might be confused because this model was technically
現在,你可能會感到困惑,因為這個模型在技術上
0:06.470–0:09.810
out , but it was in general availability , but this is the official
已經發布,但處於一般可用性階段,而這是官方
0:09.820–0:12.681
release , meaning that they have trained the model a little bit
發布,意味著他們對模型進行了更多訓練,
0:12.691–0:16.111
more and produced a stronger version of the model that is available
並產出了一個更強大的模型版本,今天即可使用。
0:16.121–0:16.590
today .
今天。
0:16.600–0:19.337
Now , if I were to summarize today's release in one simple sentence
現在,如果我要用一句簡單的話來總結今天的發布,
0:19.347–0:22.651
, it is that the price to performance is becoming a very ,
那就是性價比正變得非常、非常重要的一環,因為
0:22.661–0:25.682
very important thing because not only do we have a new model from
非常重要的一點,因為我們不僅擁有一個來自
0:25.692–0:29.074
the Deepseek team , the SpaceX team also dropped Grock 4 .
Deepseek 團隊的新模型,SpaceX 團隊也推出了 Grock 4。
0:29.084–0:29.640
6 six .
6.6。
0:29.650–0:33.063
And both of these models are competing with the best Frontier Labs
這兩個模型都在以極低的成本與 Frontier Labs 的最佳模型競爭。
0:33.073–0:35.066
at a fraction of their cost .
讓我們從 Gro 4 開始。
0:35.076–0:36.431
Let's start with Gro 4 .
讓我們從 Gro 4 開始。
0:36.441–0:36.827
6 .
我要說的一件事是,我對 SpaceX 團隊真的感到驚訝。
0:36.837–0:39.761
And one thing I'm going to say is that I'm genuinely surprised
我不是在試圖吹捧他們或埃隆·馬斯克或什麼的。
0:39.771–0:41.029
by the SpaceX team .
我只是對這個實驗室非常批評。
0:41.039–0:43.408
I'm not trying to glaze them or Elon Musk or anything .
在 2025 年,他們推出了 Grog 4 和其他類似的模型,
0:43.418–0:45.311
I was just very critical of this lab .
但我並沒有真的,你知道,對他們的表現感到驚艷,
0:45.321–0:48.671
In 2025 , they dropped Grog 4 and other models like that ,
因為與其他實驗室相比,他們的表現相當差。
0:48.681–0:52.371
but I wasn't really , you know , mind blown with their performance
但在 2026 年,情況似乎有點改變,
0:52.381–0:55.400
because they're pretty they're pretty subpar compared to any of
因為跟其他實驗室相比,它們的表現相當遜色
0:55.410–0:56.743
the other labs out there .
5 在成本基礎上相當有競爭力,而 Grock 4。
0:56.753–1:00.140
But in 2026 , it looks like things have kind of changed a little
但在 2026 年,情況似乎有點改變了
1:00.150–1:01.212
because Grock 4 .
如果我們查看基準測試,例如,
1:01.222–1:04.982
5 was quite competitive based off what it cost and Grock 4 .
如果我們從 Artificial Analysis 智能指數開始,
1:04.992–1:07.268
6 is actually not too bad .
這個模型達到了 61 分,而 Fable 5 為 62 分。
1:07.278–1:09.740
If we take a look at the benchmarks , for example ,
現在,真正重要的是要記住,再一次,
1:09.750–1:13.248
if we start with the artificial analysis intelligence index ,
這個模型與 Fable 5 相比非常便宜。
1:13.258–1:17.633
this model achieves a 61 and Fable 5 is at 62 .
這大約是每百萬輸入標記 2 美元,每百萬輸出
1:17.643–1:20.583
Now , what's really important to remember is that once again ,
現在,真正重要的是要記住,再一次,
1:20.593–1:23.374
this model is quite cheap compared to Fable 5 .
在 Fable 5 中,每百萬輸入標記為 10 美元,每百萬
1:23.384–1:27.600
This is about 2 per million input tokens and 6 per million output
這大約是每百萬個輸入 token 2 個,每百萬個輸出
1:27.610–1:28.245
tokens .
這就是為什麼你開始稍微欣賞這次發布的原因。
1:28.255–1:32.844
In Fable 5 sits at 10 per million input tokens and 50 per million
它可能無法擊敗最佳模型的表現,但正在與它們匹敵。
1:32.854–1:33.893
output tokens .
輸出標記
1:33.903–1:36.959
And this is why you start to appreciate this release a little bit
這就是為什麼你會開始稍微欣賞這次發布的原因
1:36.969–1:37.363
more .
更多
1:37.373–1:40.540
It might not beat the performance of the best models is matching
它可能無法擊敗最佳模型的表現,但正在追趕
1:40.550–1:41.037
them .
它們
1:41.047–1:42.390
Even GPT 5 .
即使是 GPT 5。
1:42.400–1:46.161
6 so which is a pretty capable model and this is set at max on
6,這是一個相當強大的模型,並且在人工分析指數上設定為最高。
1:46.171–1:48.126
the artificial analysis index .
人工分析指數
1:48.136–1:50.828
This model achieves 61 and Grok 4 .
該模型達到 61 分,Grok 4。
1:50.838–1:51.632
6 61 .
6,61 分。
1:51.642–1:53.804
So it ties it and it's much cheaper .
所以它並列第一,而且便宜得多。
1:53.814–1:56.216
And then even on all of these other benchmarks ,
然後在這些其他基準測試中,
1:56.226–1:59.343
for example , the Code Val one , it actually beats Fable 5 ,
例如 Code Val 基準,它實際上擊敗了 Fable 5,
1:59.353–2:01.186
which is at 1741 .
後者為 1741 分。
2:01.196–2:02.427
And then GPT 5 .
然後是 GPT 5。
2:02.437–2:04.309
6 , it's at 1728 .
6,它為 1728 分。
2:04.319–2:06.713
I'm not sure why they didn't choose Opus 5 as well ,
我不確定為什麼他們沒有選擇 Opus 5,
2:06.723–2:09.916
but I guess they wanted to choose the quote unquote strongest model
但我猜他們想選擇所謂的「最強」模型
2:09.926–2:10.959
lineup from each lab .
陣容來自每個實驗室。
2:10.969–2:13.903
And they chose Fable 5 for Enthropic , which is fair .
他們為 Enthropic 選擇了 Fable 5,這很公平。
2:13.913–2:16.917
And then Deep Software Engineering one , which is a critical benchmark
然後是 Deep Software Engineering 基準,這是一個關鍵基準
2:16.917–2:19.286
.
2:19.286–2:20.585
This model doesn't beat Fable 5 or GPT 5 .
該模型沒有擊敗 Fable 5 或 GPT 5。
2:20.595–2:22.322
6 six soul , but it gets close to it .
6,6 分,但非常接近。
2:22.332–2:23.255
It's 65 .
它是 65。
2:23.265–2:25.839
9 and Fable 5 sits at 70 .
9,而 Fable 5 為 70 分。
2:25.849–2:28.556
But if you're getting results that are pretty close and the model
但如果你得到的結果非常接近,而且模型
2:28.566–2:31.432
is five times cheaper , I wouldn't be too disappointed with this
如果價格便宜五倍,我不會對這個結果太失望
2:31.442–2:31.991
result .
結果感到太失望。
2:32.001–2:34.068
And then same thing with Cursor Bench 3 .
然後是 Cursor Bench 3。
2:34.078–2:36.221
2 , the model achieves 69 .
2,該模型達到 69。
2:36.231–2:37.904
9 , funny number .
9,有趣的數字。
2:37.914–2:39.617
And Fable 5 sits at 70 .
而 Fable 5 為 70。
2:39.627–2:40.141
5 .
[未翻譯]
2:40.151–2:41.260
So once again , closer .
所以再一次,更接近。
2:41.270–2:42.199
And Grock 4 .
而 Grock 4。
2:42.209–2:43.693
6 beats GPT 5 .
6 擊敗了 GPT 5。
2:43.703–2:44.860
6 .
[未翻譯]
2:44.870–2:46.375
Same thing with the Frontier Code .
在 Frontier Code 方面也是如此。
2:46.385–2:49.287
It gets close to Fable 5 , beats GPT 5 .
它非常接近 Fable 5,擊敗了 GPT 5。
2:49.297–2:49.987
6 .
[未翻譯]
2:49.997–2:52.808
So yeah , this model is actually available in cursor .
所以是的,該模型實際上在 cursor 中可用。
2:52.818–2:56.515
So the partnership with cursor or I guess the acquisition has really
所以與 cursor 的合作關係,或者我猜是收購,確實
2:56.525–3:00.540
helped SpaceX make some strides in the AI space this year and Grock
幫助 SpaceX 在今年在 AI 領域取得了一些進展,而 Grock
3:00.550–3:03.740
build is something that you know maybe not a lot of us have been
build 是你知道也許我們中沒有很多人一直在
3:03.750–3:07.341
using so far but it's probably going to be another platform like
到目前為止還沒有人使用,但未來我們可能會開始使用類似 CodeEx 或 Cloud Code 的這類平台,不過顯然 Cursor 也很強大。
3:07.351–3:10.793
codeex or cloud code that we start to use but obviously cursor
因此,你有可用的選項,可以在兩者中選擇使用。
3:10.803–3:11.997
is quite strong as well .
並且有一件事需要注意,Grok Build 和 Cursor 在第一週提供兩倍的用量。
3:12.007–3:15.369
So you have options available for you to use them in both .
所以如果你只是想試試看,看看你的感覺如何,那麼現在嘗試可能值得,因為你在第一週可以獲得雙倍的用量。
3:15.379–3:18.660
And one thing to note is that they're offering two times usage
現在,如果我們來看看人們使用 Grok 4.6 所生成的一些輸出結果。
3:18.670–3:21.287
inside Grok Build and Cursor for the first week .
我們現在看到的是 Falcon 9 助推器返回序列的模擬。
3:21.297–3:23.992
So if you just want to try it out , see what you feel about it
這是在單一 HTML 檔案中完成的。
3:24.002–3:27.189
, then you know it might be worth trying it out right now cuz you
正如我所說,如果你預期在 2025 年從 Grok 獲得這種類型的輸出,你會感到相當驚訝,因為你不會預期 Grok 能生成這樣的內容。
3:27.199–3:29.304
get double the usage in the first week .
但現在看來,我們必須開始更認真地看待 Grok 團隊,因為這項輸出相當具有競爭力。
3:29.314–3:31.907
Now if we were to take a look at some of the outputs that people
顯然,我們只是在看一個模擬,並且僅基於視覺呈現來評估。
3:31.917–3:33.916
have been generating with Grok 4 .
都是使用 Grok 4 生成的。
3:33.926–3:37.824
6 , what we're looking at right now is a Falcon 9 booster return
但至少目前,他們的模型確實比其他模型便宜。
3:37.834–3:38.939
sequence simulation .
正如我所提到的,該模型在人工分析指數中表現出色。
3:38.949–3:42.050
And this was done in a single HTML file .
該模型躍升至可能排名第四的模型。
3:42.060–3:45.479
And as I said , if you expected to get this type of output from
它與 GPT 5.6 並列,表現非常相似。
3:45.489–3:49.852
Grock in 2025 , you would be kind of surprised because you wouldn't
所以你也可以說它是排名第三,但在此之前的是 Falcon 5 和 Opus 5,它們僅比該模型高出約 1 或 2 分的差距。
3:49.862–3:52.240
expect something like this to be generated with Grock .
所以是的,即使是在這個智力指數上,如果你不熟悉,它包含九項評估。
3:52.250–3:54.958
But now it looks like we have to start taking the Grok team a
所以在所有這些評估中,它幾乎與 Falcon 5 的表現相匹配。
3:54.968–3:58.195
little bit more serious because this output is quite competitive
稍微更嚴肅一些,因為這項成果相當具有競爭力
3:58.195–4:00.295
.
4:00.295–4:01.350
And obviously , we're just looking at a simulation and we're just
而且顯然,我們只是在看模擬結果,並且只是
4:01.360–4:03.987
basing it off of a visual representation .
基於視覺呈現來進行評估。
4:03.997–4:06.624
But if the model is able to produce something like this consistently
但如果該模型能夠持續產生類似的成果
4:06.634–4:10.525
, then I would expect a lot of people to start adopting Grock because
,那麼我預期會有很多人開始採用 Grock,因為
4:10.535–4:13.368
number one , it is cheaper than the other labs at least at the
第一,它比其他實驗室便宜,至少目前是如此
4:13.378–4:15.895
moment because we don't know if this pricing strategy is going
因為我們不知道這種定價策略
4:15.905–4:18.895
to be sustainable for the SpaceX team in the long run .
對 SpaceX 團隊在長期來看是否具可持續性
4:18.905–4:21.659
But at least for now , their models are definitely cheaper compared
但至少目前,他們的模型確實比其他模型便宜
4:21.669–4:22.608
to the others .
相較於其他模型
4:22.618–4:26.003
And as I mentioned , this model excelled at the artificial analysis
正如我所提到的,該模型在人工分析指數方面表現出色
4:26.013–4:26.635
index .
指數
4:26.645–4:30.031
This model jumped to probably number four model .
該模型躍升至可能排名第四的模型
4:30.041–4:32.180
It's tied to GPT 5 .
它與 GPT 5 並列。
4:32.190–4:33.370
6 pretty much similar .
與 6 幾乎相同。
4:33.380–4:36.579
So you could say number three as well , but the models before that
所以你也可以說它是第三名,但在那之前的模型
4:36.589–4:40.896
are Fable 5 and Opus 5 , which are only above the model by about
是 Fable 5 和 Opus 5,它們僅比該模型高出約
4:40.906–4:43.454
a 1 or a 2 difference .
1 或 2 的差距。
4:43.464–4:46.492
So yeah , even on this intelligence index , which if you're not
所以是的,即使是在這個智力指數上,如果你不
4:46.502–4:48.570
familiar with has nine evaluations .
熟悉的話,它包含九項評估。
4:48.580–4:52.646
So on all of these evaluations , it's kind of matching almost Fable
所以在所有這些評估中,它幾乎與 Fable
4:52.656–4:55.046
5 performance , which is crazy to see .
5的表現,這在視覺上令人瘋狂。
4:55.056–4:58.403
Before we continue , we just launched the Universe of AI newsletter
在我們繼續之前,我們剛剛推出了AI宇宙通訊電子報
4:58.403–5:00.629
.
5:00.629–5:01.761
If you want to stay on top of AI news without having to hunt for
如果你想隨時掌握AI新聞,而不必費心去搜尋
5:01.771–5:03.595
it , link is in the description .
它,連結在描述中。
5:03.605–5:04.710
Don't miss out .
別錯過。
5:04.720–5:07.800
And what you see on screen right now is a racing game that Grok
而你現在在螢幕上看到的是一款由Grok
5:07.810–5:08.320
4 .
[未翻譯]
5:08.330–5:09.127
6 build .
.6版本構建的賽車遊戲。
5:09.137–5:12.168
And based off of this post , the model took about 1 minute and
根據這篇帖子,該模型大約花了1分鐘
5:12.178–5:15.191
it was a fiveword prompt , which was a create a simple racing game
,這是一個五個字的提示,內容是創建一個簡單的賽車遊戲
5:15.201–5:16.269
in HTML .
,使用HTML。
5:16.279–5:19.367
So , if you're able to generate something like this easily using
所以,如果你能輕鬆地生成像這樣的東西,使用
5:19.377–5:20.100
Grock 4 .
所以,如果你能輕鬆地使用Grok 4
5:20.110–5:22.300
6 , I think a lot of people will be happy .
.6生成這樣的東西,我認為許多人會很高興。
5:22.310–5:26.455
And this is a more detailed analysis of what it costs to run GPT
這是一份更詳細的分析,說明在人工分析指數上運行GPT
5:26.465–5:26.785
5 .
[未翻譯]
5:26.795–5:30.272
6 6 on the artificial analysis index and what it produced meaning
.6的成本以及它產生的內容,也就是輸出結果。
5:30.282–5:31.224
the output .
如果你記得,這兩個模型在人工分析指數上都獲得了61分。
5:31.234–5:34.713
Both of these models if you remember scored 61 on the artificial
如果你記得,這兩個模型在人工智慧
5:34.723–5:35.745
analysis index .
.6運行整個測試,成本大約是2
5:35.755–5:38.230
Now to run the whole test with GPT 5 .
.6,成本大約是2
5:38.240–5:40.135
6 it cost about 2 .
.6,成本大約是1
5:40.145–5:42.008
6 it cost about 2 .
.1Kish。
5:42.018–5:43.620
6 it cost about 1 .
這告訴你,你以一半的成本獲得了相似水平的表現。
5:43.630–5:44.300
1Kish .
所以,是的,這是SpaceX團隊的一次重大發布,因為他們再次證明
5:44.310–5:47.465
And this tells you that you're getting similar level of performance
他們是一個你認真需要考慮的實驗室,特別是在2026年。
5:47.475–5:48.519
at half the cost .
我將更多地討論Deep Seek版本4 Pro GA。
5:48.529–5:51.893
So yeah , this is a big release for the SpaceX team because they
但基本上,我們今天看到的是,這兩次發布
5:51.903–5:54.779
just proved once again that they are a lab that you seriously start
都強調了一個事實,那就是表現和超越表現的一切
5:54.789–5:57.080
need to considering , especially in 2026 .
,是你為該表現所支付的價格,正變得
5:57.090–6:00.824
And I'm going to talk more about Deep Seek version 4 Pro GA .
越來越重要,對於所有用戶來說
6:00.834–6:04.350
But basically what we're seeing today is that both of these releases
但基本上,我們今天看到的是,這兩項發布
6:04.360–6:08.119
kind of emphasize the fact that performance and all above that
現在專注於以最低的成本創建最佳模型。
6:08.129–6:11.005
is the price at what you're getting for that performance is becoming
去年2025年,大多數實驗室只專注於我會說
6:11.015–6:13.201
more and more important for all users
創建最強大的模型。
6:13.201–6:14.612
because we see many labs
因為我們看到許多實驗室
6:14.622–6:18.379
now focusing on creating the best model at the cheapest cost .
現在專注於以最低的成本打造最佳模型。
6:18.389–6:21.907
Last year in 2025 most of the labs were just focused on I would
去年,2025年,大多數實驗室只專注於我會
6:21.917–6:24.472
say creating the strongest model .
說打造最強大的模型
6:24.482–6:25.571
Yes , cost was important ,
是的,成本很重要,
6:25.571–6:27.519
but I think most of the times the frontier
但我認為大多數時候,前沿
6:27.529–6:29.280
labs , meaning OpenAI , Anthropic ,
實驗室,也就是 OpenAI、Anthropic,
6:29.280–6:30.829
were kind of more lenient on
在這一點上比較寬鬆,
6:30.839–6:34.100
that fact because they didn't have as strong of a competition .
因為它們面臨的競爭沒有那麼激烈。
6:34.110–6:38.733
Intelligent models that are maybe not always ahead of OpenAI Enthropic
這些智能模型可能並不總是領先於 OpenAI 和 Anthropic,
6:38.743–6:41.447
, but match their performance at a fraction of the cost .
但能以更低廉的成本達到相同的性能。
6:41.457–6:45.369
So yes , price to performance ratio is becoming a critical I would
所以是的,性價比會成為我在2026年認為的關鍵指標
6:45.379–6:47.955
say indicator in 2026 .
我認為會成為一個關鍵指標。
6:47.965–6:50.861
Now , this is the updated benchmark chart after the release of
現在,這是新模型發布後更新的基準測試圖表。
6:50.871–6:51.742
the new model .
你會發現一個普遍現象,那就是它達到了
6:51.752–6:54.623
And one thing you'll see across the board is that it matches the
而且你會發現一個普遍現象,那就是它符合了
6:54.633–6:56.944
top level performance of many of the models .
這裡有趣的一點是,出於某種原因,他們沒有將 Opus
6:56.954–6:59.345
The one thing interesting over here is that they haven't put Opus
這裡有趣的一點是他們沒有把 Opus 5 放進來
6:59.355–7:00.704
5 here for some reason .
這裡有 Fable 5,我們可以將這個模型與之進行比較。
7:00.714–7:03.747
There is Fable 5 here that we can compare this model against .
但你會注意到的一點是,DeepSeek,
7:03.757–7:06.068
But one thing you'll notice is that DeepSeek ,
請記住這個模型的成本是每百萬輸入 token 43
7:06.078–7:08.335
remember this model costs 43 .
記得這個模型的費用是 43
7:08.345–7:12.791
5 cents per million input tokens and 87 cents per million output
每百萬個輸入記號5美分,每百萬個輸出記號87美分
7:12.801–7:13.511
tokens .
而其他模型在這裡的成本都要高得多。
7:13.521–7:16.391
While the other models all over here are way more expensive than
所以,如果你查看 Terminal Bench 2.
7:16.401–7:16.872
that .
那。
7:16.882–7:19.540
So the first thing if you look at terminal bench 2 .
所以如果你看 terminal bench 2 的話,第一件事是
7:19.550–7:21.629
1 the model scores 87 .
該模型的舊版本得分為 72.
7:21.639–7:22.394
9 .
7:22.404–7:24.656
The older version of the model was 72 .
舊版模型的得分是72
7:24.666–7:28.557
1 and the flash version which we got last week was 82 .
瘋狂的是,Fable 5 的得分為 88.
7:28.567–7:29.277
7 .
[未翻譯]
7:29.287–7:32.239
And what's crazy is that Fable 5 is 88 .
瘋狂的是,Fable 5 高達 88。
7:32.249–7:33.279
Yes , 88 .
所以這個模型僅比
7:33.289–7:34.887
So this model is only .
Fable 5 低了 0.1 分。
7:34.897–7:36.827
1 behind Fable 5 .
然後在 Cyber Gym,也就是網絡安全領域,
7:36.837–7:39.300
And then on the Cyber Gym , which is Cyber Security ,
該模型實際上超越了 Fable 5。
7:39.310–7:41.629
the model actually beats Fable 5 .
Fable 5 的得分為 83.
7:41.639–7:43.559
Fable 5 sits at 83 .
Fable 5 的得分為 83
7:43.569–7:46.683
1 while Deepseek version 4 Pro , the new one that we got today
而 Deepseek 4 Pro 版本,也就是我們今天拿到的新版本
7:46.693–7:48.044
3 .
[未翻譯]
7:48.054–7:48.765
3 .
現在,如果這告訴你什麼的話,那就是這個模型再次
7:48.775–7:51.966
Now , if this tells you something is that this model is once again
現在,如果這告訴你一些事情,那就是這個模型再次
7:51.976–7:53.807
really geared at cyber security .
確實是針對網路安全所設計的
7:53.817–7:56.850
We saw the Flash model also be geared towards cyber security and
我們看到 Flash 模型也針對網路安全領域,
7:56.860–7:59.412
becoming a model that was quite capable in that area .
並成為在該領域相當有能力的模型。
7:59.422–8:02.534
And we're seeing the same thing today with the new version 4 Pro
而今天我們看到新版本 4 Pro 也呈現相同情況,
8:02.544–8:03.975
which sits at 83 .
它在基準測試中得分 83.3,
8:03.985–8:06.858
3 outperforming Fable 5 on the benchmark .
超越了 Fable 5。
8:06.868–8:10.379
On deep software engineering , this model is not beating Fable
在深度軟體工程方面,這個模型並未超越 Fable
8:10.389–8:12.491
5 because Fable 5 sits at 70 ,
5,因為 Fable 5 得分為 70,
8:12.491–8:14.497
but the model achieves 62 .
但該模型取得了 62.7 分,
8:14.507–8:16.744
7 which is ahead of GLM 5 .
領先 GLM 5.2,
8:16.754–8:20.209
2 and a bit behind Kimmy K3 which sits at 67 .
略遜於得分 67.5 的 Kimmy K3。
8:20.219–8:20.977
5 .
8:20.987–8:24.115
But one thing again , this is a big jump compared to the preview
但有一點再次強調,與該模型的預覽版相比,這是巨大的進步,
8:24.125–8:26.286
version of the model which was at 12 .
預覽版得分僅為 12.8。
8:26.296–8:26.769
8 .
8:26.779–8:29.655
So yeah , they have really trained this model and they have improved
所以是的,他們確實訓練了這個模型,並在後端進行了改進,
8:29.665–8:32.998
it on the back end because we can clearly see in this deep software
因為我們可以清楚地從這個深度軟體
8:33.008–8:34.430
engineering benchmark .
工程基準測試中看到。
8:34.440–8:37.216
And there's a couple of other benchmarks , but one key benchmark
還有其他幾個基準測試,但這個模型擅長的一個關鍵基準測試是自動化基準,
8:37.226–8:41.353
that this model excels at is the automation bench where the model
該模型在自動化基準測試中表現出色,而該模型
8:41.363–8:42.680
achieves 31 .
而 Fable 5 為 29.1。
8:42.690–8:44.775
8 and Fable 5 29 .
8:44.785–8:45.332
1 .
8:45.342–8:48.039
So once again , the model outperforms Fable 5 .
所以再次強調,該模型優於 Fable 5。
8:48.049–8:50.268
Now this is a fraction of the cost as I mentioned .
現在,正如我提到的,這只是極小的一部分成本。
8:50.278–8:54.378
This is 43 cents and 87 for input and output respectively versus
這分別是輸入和輸出的 43 美分和 87 美分,
8:54.388–8:56.300
Fable 5 which is at 10 .
而 Fable 5 為 10.
8:56.310–8:57.083
50 .
[未翻譯]
8:57.093–8:59.787
So yes , if we look at the terminal bench for example ,
所以是的,如果我們以終端基準測試為例,
8:59.797–9:02.655
we are seeing this model achieve a result which is 0 .
我們看到這個模型取得了與目前最佳模型 Fable 5 僅有 0.1 分之差的結果,
9:02.665–9:07.078
1 behind the best model out there Fable 5 and the pricing is insane
而價格非常驚人,每百萬輸入令牌的價格分別為 43 美分和 10.87 美分對比 50 美分。
9:07.088–9:14.184
to look at 43 versus 10 87 versus 50 per million input tokens .
9:14.194–9:18.674
So when we turn that into a fraction , this model is 57 times cheaper
所以當我們將這轉化為比例時,這個模型便宜了 57 倍。
9:18.674–9:19.184
.
9:19.184–9:22.285
And earlier last week , I think I made a video about how DeepSeek
而在上週早些時候,我想我做過一個影片,說明 DeepSeek
9:22.295–9:25.740
version 4 Pro was going to achieve a result like this because there
版本 4 Pro 將如何取得這樣的結果,因為有一份投資者報告洩露。
9:25.750–9:27.495
was a investor report that leaked .
9:27.505–9:30.127
And when I saw the pricing where it said that it was going to match
當我看到定價時,上面寫說它會匹敵
9:30.137–9:34.470
Fable 5 or even outperform it and be at 57 times cheaper per cost
Fable 5 持平甚至超越它,並且成本每百萬輸入令牌便宜 57 倍。
9:34.470–9:35.751
.
9:35.751–9:37.277
When I was reading the investor report , I didn't really believe
當我閱讀那份投資者報告時,我其實不太相信
9:37.287–9:37.837
it .
它。
9:37.847–9:40.633
But today , it is proven because at least on the benchmark so far
但今天,它已經得到證實,因為至少在目前的基準測試中
9:40.643–9:43.829
, yes , I'm just saying the benchmarks , we are seeing similar
,是的,我只是說基準測試,我們看到相似的
9:43.839–9:44.948
level performance .
性能水平。
9:44.958–9:48.303
Now , if this model actually performs like that in production ,
現在,如果這個模型在實際生產中真的表現得如此,
9:48.313–9:50.141
it's a little too early to tell yet .
現在還說得太早。
9:50.151–9:52.763
We still would have to give it a couple of weeks and see how it's
我們仍然需要再給它幾週時間,看看它長期
9:52.773–9:55.867
performing in the long run because sometimes on the release day
長期表現如何還得看後續,因為有時在發布當天
9:55.877–9:59.340
the models perform quite good , but over time they kind of deteriorate
模型表現相當不錯,但隨著時間推移,它們的質量
9:59.350–10:00.380
in their quality .
可能會逐漸下降。
10:00.390–10:02.005
So , we hope DeepSeek doesn't do that .
所以,我們希望 DeepSeek 不要出現這種情況。
10:02.015–10:04.469
Historically , they haven't done that , but let's just see .
從歷史上看,他們沒有這樣做過,但讓我們拭目以待。
10:04.479–10:06.862
Just going to put it out there because right now we're basing this
我只是想把它說出來,因為目前我們是基於基準測試
10:06.872–10:08.529
off of benchmarks .
來得出這個結論。
10:08.539–10:10.051
But that's it for today's video .
但今天的影片就到這裡。
10:10.061–10:12.081
Make sure you guys are subscribed to the channel .
記得訂閱我們的頻道
10:12.091–10:15.369
Follow our new newsletter as well at universeofai .
也請訂閱我們的新聞通訊,網址是 universeofai .
10:15.379–10:15.756
behiiv .
behiv
10:15.766–10:19.087
com as well as subscribe to the main channel World of AI and support
以及訂閱 World of AI 主頻道並支持
10:19.097–10:22.063
us on X by following the Universe of AIZ as well .
我們,關注 Universe of AIZ。
10:22.073–10:23.114
Until then ,
在那之前,
10:23.114–10:24.618
I'll see you guys
我們下次見
10:24.618–10:25.659
in the next
在下一個

影片筆記:DeepSeek V4 Pro Is Out Today And It's A Problem For Every AI Lab!

一句話總結

DeepSeek v4 Pro 與 SpaceX 發布的 Grok 4.6 正式發布,兩者皆以極低的價格提供匹敵頂級實驗室(如 Fable 5、GPT 5.6)的性能,標誌著 2026 年 AI 競爭關鍵從「最強性能」轉向極致的「性價比」。

核心重點

  1. DeepSeek v4 Pro 正式發布 (GA)
  • 經過更多訓練,性能顯著提升,在 Terminal Bench、Cyber Gym、Automation Bench 等多項基準測試中表現優異,部分項目超越 Fable 5。
  • 價格極具競爭力:輸入 43 美分/百萬 token,輸出 87 美分/百萬 token,據稱比 Fable 5 便宜約 57 倍。
  • 驗證了此前根據洩露投資者報告所做的性價比預測。
  1. SpaceX Grok 4.6 發布
  • 被視為 SpaceX 團隊的重大進步,在 Artificial Analysis Intelligence Index、Code Val、Cursor Bench 等測試中與 Fable 5 和 GPT 5.6 打平或略勝。
  • 價格極低:輸入約 2 美元/百萬 token,輸出約 6 美元/百萬 token,性能匹敵頂級模型但成本大幅降低。
  • 已在 Cursor 和 Grok Build 平台可用,首週提供雙倍使用量優惠。
  1. 市場趨勢轉變 (2025 vs 2026)
  • 2025 年:實驗室專注於打造最強模型,對成本較寬鬆。
  • 2026 年:隨著 DeepSeek 和 SpaceX 等競爭對手以極低成本提供匹敵頂級性能的服務,「性價比 (Price to Performance)」成為用戶選擇模型的重要指標。

詳細大綱

A. DeepSeek v4 Pro 正式發布與性價比革命

  • 發布狀態:DeepSeek v4 Pro 今日正式發布 (General Availability, GA)。此前雖已可用但非正式版本,此次發布代表模型經過更多訓練,性能更強。
  • 基準測試表現
  • Terminal Bench 2.1:得分 87.9,僅落後 Fable 5 (88) 0.1 分;舊版為 72.1,Flash 版為 82.7。
  • Cyber Gym (網路安全):得分 83.3,超越 Fable 5 (83.1)。顯示該模型在網路安全領域具備強大能力,延續了 Flash 版的優勢。
  • Deep Software Engineering:得分 62.7,落後 Fable 5 (70),但領先 GLM 5.2 和 Kimmy K3 (67.5)。相比預覽版的 12.8 有巨大提升。
  • Automation Bench:得分 31.8,超越 Fable 5 (29.1)。
  • 價格優勢
  • DeepSeek v4 Pro:輸入 43 美分/百萬 token,輸出 87 美分/百萬 token。
  • 對比 Fable 5:輸入 10.50 美元,輸出 50 美元。
  • 結論:DeepSeek v4 Pro 比 Fable 5 便宜約 57 倍。
  • 預測驗證:此前根據洩露的投資者報告預測其性價比,當時持懷疑態度,但今日基準測試結果證實了預測。
  • 長期表現觀察:目前僅基於基準測試,需觀察數週以確認生產環境中的長期穩定性,避免發布日表現好但隨後品質下降的情況。

B. SpaceX Grok 4.6 發布與評估

  • 團隊轉變:講者對 SpaceX 團隊在 2026 年的表現感到驚訝。2025 年發布的 Grok 4 等模型表現平平,但 2026 年的 Grok 4.5 和 4.6 顯示出競爭力。
  • 基準測試表現
  • Artificial Analysis Intelligence Index:Grok 4.6 得分 61,與 Fable 5 (62) 和 GPT 5.6 (61, Max 設定) 打平或非常接近。
  • Code Val:Grok 4.6 得分超越 Fable 5 (1741) 和 GPT 5.6 (1728)。
  • Deep Software Engineering:得分 65.9,接近 Fable 5 (70) 和 GPT 5.6 (66)。
  • Cursor Bench:得分 69.9,接近 Fable 5 (70.5),並超越 GPT 5.6。
  • Frontier Code:接近 Fable 5,超越 GPT 5.6。
  • 價格對比
  • Grok 4.6:輸入約 2 美元/百萬 token,輸出約 6 美元/百萬 token。
  • Fable 5:輸入 10 美元,輸出 50 美元。
  • 結論:Grok 4.6 性能匹敵頂級模型,但價格極低。
  • 平台與合作
  • Grok 4.6 已在 Cursor 中可用。
  • 提及 SpaceX 與 Cursor 的合作或收購關係有助於其在 AI 領域的進展。
  • 預測 Grok Build 可能成為類似 Codeex 或 Cloud Code 的平台。
  • 優惠:首週在 Grok Build 和 Cursor 中提供雙倍使用量。
  • 輸出範例
  • Falcon 9 助推器返回序列模擬:單個 HTML 文件生成,展現競爭力。
  • 賽車遊戲:使用 5 個字的提示詞 ("create a simple racing game in HTML"),Grok 4.6 在約 1 分鐘內生成。
  • 成本分析
  • 運行 GPT 5.6 在 Artificial Analysis Index 上的完整測試成本約為 1.1K(講者重複提及 2.6 和 1.1Kish,語意稍顯混亂,但強調 GPT 成本較高)。
  • Grok 4.6 提供相似性能但成本僅為一半或更低。

C. 市場總結與展望

  • 競爭格局變化
  • 2025 年:實驗室 (OpenAI, Anthropic 等) 專注於打造最強模型,對成本較寬鬆,因競爭較少。
  • 2026 年:出現智能模型雖未領先但能匹敵頂級性能且成本極低的競爭對手 (如 DeepSeek, SpaceX)。
  • 指標轉變:性價比 (Price to Performance Ratio) 成為 2026 年關鍵指標。
  • 基準圖表更新:新模型發布後,DeepSeek 和 Grok 在多個基準中匹配頂級性能,但 Opus 5 未被列入部分比較中 (講者疑惑為何未包含 Opus 5)。

工具 / 模型 / 名詞整理

  • 模型名稱
  • DeepSeek version 4 Pro (DeepSeek v4 Pro)
  • Grok 4.6 (Grok 4.6)
  • Grok 4.5
  • Grok 4
  • Fable 5 (講者多次提及,疑為某實驗室頂級模型名稱)
  • GPT 5.6 (講者提及,疑為 OpenAI 模型)
  • Opus 5 (講者提及 Anthropic 模型,但未列入部分比較)
  • GLM 5.2
  • Kimmy K3
  • Flash version (DeepSeek Flash)
  • 平台/工具
  • Cursor
  • Grok Build
  • Codeex
  • Cloud Code
  • Artificial Analysis Intelligence Index (Artificial Analysis 智能指數)
  • Terminal Bench 2.1
  • Cyber Gym (Cyber Security 基準測試)
  • Deep Software Engineering (基準測試)
  • Cursor Bench
  • Frontier Code (基準測試)
  • Automation Bench
  • 其他專有名詞
  • SpaceX
  • Deepseek team
  • Frontier Labs (頂級實驗室)
  • OpenAI
  • Anthropic
  • Universe of AI newsletter (Universe of AI 通訊)
  • World of AI (頻道名稱)
  • X (平台)

操作流程整理

  1. 評估 DeepSeek v4 Pro 性能
  • 查看 Terminal Bench 2.1 得分 (87.9)。
  • 查看 Cyber Gym 得分 (83.3)。
  • 查看 Deep Software Engineering 得分 (62.7)。
  • 查看 Automation Bench 得分 (31.8)。
  • 確認價格:輸入 43 美分/百萬 token,輸出 87 美分/百萬 token。
  • 與 Fable 5 進行價格對比 (便宜約 57 倍)。
  1. 評估 Grok 4.6 性能
  • 查看 Artificial Analysis Intelligence Index 得分 (61)。
  • 查看 Code Val 得分 (超越 Fable 5 和 GPT 5.6)。
  • 查看 Deep Software Engineering 得分 (65.9)。
  • 查看 Cursor Bench 得分 (69.9)。
  • 確認價格:輸入約 2 美元/百萬 token,輸出約 6 美元/百萬 token。
  • 確認平台可用性:Cursor 和 Grok Build。
  • 利用首週雙倍使用量優惠進行測試。
  1. 生成範例測試
  • 使用 Grok 4.6 生成 Falcon 9 助推器返回序列模擬 (單個 HTML 文件)。
  • 使用 5 字提示詞 ("create a simple racing game in HTML") 生成賽車遊戲,耗時約 1 分鐘。
  1. 市場趨勢分析
  • 比較 2025 年與 2026 年的競爭重點 (性能 vs. 性價比)。
  • 觀察 DeepSeek 和 Grok 在基準測試中與頂級模型 (Fable 5, GPT 5.6) 的表現對比。
  • 關注生產環境中的長期穩定性。

值得注意的限制或風險

  1. 生產環境穩定性:目前僅基於基準測試,需觀察數週以確認生產環境中的長期穩定性,避免發布日表現好但隨後品質下降的情況。
  2. 基準測試的局限性:基準測試結果可能無法完全反映實際生產環境中的表現,特別是長期運行的穩定性。
  3. 模型名稱與版本的準確性:影片中提及的多個模型名稱 (如 Fable 5, GPT 5.6, Kimmy K3, Opus 5) 和基準測試名稱可能存在辨識錯誤或為非標準命名,需進一步查證。
  4. 價格單位的清晰度:部分價格描述 (如 "1.1Kish", "2.6") 語意不清晰,可能影響對成本的準確評估。
  5. Opus 5 的缺席:講者疑惑為何 Opus 5 未被列入部分比較中,這可能影響對整體競爭格局的完整理解。

逐字稿辨識疑點

  • Fable 5:逐字稿中多次出現 "Fable 5",並與 GPT 5.6、Opus 5 並列為頂級模型。此名稱在常見 AI 模型中較不常見,可能為聽寫錯誤(例如是否為 "Claude 3.5"、"Gemini 1.5 Pro" 或其他模型?),但根據規則僅標記為「需查證」。
  • GPT 5.6:講者提及 "GPT 5.6",目前 OpenAI 主流模型版本號通常不以此格式出現(如 GPT-4o, GPT-4.5 等),此為「需查證」的模型名稱。
  • Kimmy K3:講者提及 "Kimmy K3" 在 Deep Software Engineering 基準中得分 67.5。此名稱疑似為聽寫錯誤(可能指代某特定模型或實驗室產品),標記為「需查證」。
  • GLM 5.2:講者提及 "GLM 5.2"。GLM 系列通常由智譜 AI 開發,版本號格式需查證是否準確。
  • Opus 5:講者提及 "Opus 5" 為 Anthropic 的模型。目前 Anthropic 頂級模型為 Opus (Claude 3 Opus),"Opus 5" 可能為聽寫錯誤或未來預測名稱,標記為「需查證」。
  • 1.1Kish / 2.6:在描述運行 GPT 5.6 的成本時,講者重複提及 "2.6" 和 "1.1Kish",語意不清晰,可能指 1.1K (1100) 或其他單位,標記為「需查證」。
  • Deep Software Engineering one:講者口語中出現 "Deep Software Engineering one",應指 "Deep Software Engineering" 基準測試,"one" 可能為口語贅字或聽誤。
  • Soul:在 "GPT 5.6 six soul" 處,"soul" 可能為 "score"(得分)的聽寫錯誤,或特定術語,標記為「需查證」。
  • Enthropic:講者口語中將 "Anthropic" 聽寫為 "Enthropic",標記為「需查證」。
  • Cursor Bench 3.2:講者提及 "Cursor Bench 3.2",需查證該基準測試版本號是否準確。
  • Terminal Bench 2.1:講者提及 "Terminal Bench 2.1",需查證該基準測試版本號是否準確。
  • Cyber Gym:講者提及 "Cyber Gym" 作為網路安全基準,需查證該基準測試名稱是否準確。
  • Automation Bench:講者提及 "Automation Bench",需查證該基準測試名稱是否準確。
  • Frontier Code:講者提及 "Frontier Code" 作為基準,需查證該基準測試名稱是否準確。
  • Artificial Analysis Intelligence Index:講者提及此指數,需查證其準確名稱是否為 "Artificial Analysis Intelligence Index" 或類似名稱。
  • Universe of AI newsletter:講者提及通訊名稱,需查證是否為 "Universe of AI" 或 "World of AI" 相關通訊。
  • World of AI:講者提及頻道名稱 "World of AI",需查證是否為正確頻道名稱。
  • X:講者提及在 "X" 上關注,指推特平台,標記為「需查證」平台名稱。
  • SpaceX team:講者將 Grok 開發團隊稱為 "SpaceX team",需查證是否準確指代 xAI 團隊。
  • Grok Build:講者提及 "Grok Build" 平台,需查證是否為正確產品名稱。
  • Codeex:講者提及 "Codeex" 平台,需查證是否為正確產品名稱。
  • Cloud Code:講者提及 "Cloud Code" 平台,需查證是否為正確產品名稱。
  • Deepseek version 4 Pro GA:講者提及 "GA" (General Availability),標記為「需查證」縮寫含義。
  • Investor report:講者提及洩露的 "investor report",需查證具體報告名稱。
  • Falcon 9 booster return sequence simulation:講者提及此模擬範例,標記為「需查證」具體內容描述。
  • HTML file:講者提及生成單個 HTML 文件,標記為「需查證」技術細節。
  • 5-word prompt:講者提及 "fiveword prompt",標記為「需查證」提示詞長度描述。
  • Create a simple racing game in HTML:講者提及此提示詞,標記為「需查證」具體提示詞內容。
  • 1 minute:講者提及生成賽車遊戲耗時約 1 分鐘,標記為「需查證」時間描述。
  • 2 per million input tokens:講者提及 Grok 4.6 輸入價格,標記為「需查證」價格單位。
  • 6 per million output tokens:講者提及 Grok 4.6 輸出價格,標記為「需查證」價格單位。
  • 10 per million input tokens:講者提及 Fable 5 輸入價格,標記為「需查證」價格單位。
  • 50 per million output tokens:講者提及 Fable 5 輸出價格,標記為「需查證」價格

生字列表

生字讀音類型中文
general availability/ˈdʒɛnərəl ˌeɪvəˈlɪbɪti/noun phrase一般可用性
price to performance/praɪs tu pərˈfɔrməns/noun phrase性價比
fraction/ˈfrækʃən/noun小部分;分之一
glaze/ɡleɪz/verb吹捧;過度美化
subpar/ˌsʌbˈpɑr/adjective低於標準的;不如預期的
competitive/kəmˈpɛtətɪv/adjective有競爭力的
benchmark/ˈbɛntʃmɑrk/noun基準測試;標準
tie/taɪ/verb並列;打平
lineup/ˈlaɪnʌp/noun陣容;系列
stride/straɪd/noun進展;大步
sustainable/səˈsteɪnəbəl/adjective可持續的
adopt/əˈdɑpt/verb採用;採納
indicator/ˈɪndɪkeɪtər/noun指標;指示器
deteriorate/dɪˈtɪriəreɪt/verb惡化;變壞
outperform/ˌaʊtpərˈfɔrm/verb優於;表現超過
leak/liːk/verb洩露
gear/ɡɪr/verb針對...設計;使適合
preview/ˈprɛvju/noun預覽版;預覽

生字解說

general availability /ˈdʒɛnərəl ˌeɪvəˈlɪbɪti/

noun phrase · C1

意思:一般可用性

解說:指產品或服務已從測試階段正式開放給公眾使用,但可能尚未達到最終的「官方發布」狀態。

影片原句
but it was in general availability , but this is the official release
但處於一般可用性階段,而這是官方發布
延伸例句
The beta version is now in general availability for all users.
這個測試版現在已對所有用戶開放一般可用性。

price to performance /praɪs tu pərˈfɔrməns/

noun phrase · C1

意思:性價比

解說:指產品的性能與其價格之間的比率,常用於評估科技產品的價值。

影片原句
it is that the price to performance is becoming a very , very important thing
那就是性價比正變得非常、非常重要的一環
延伸例句
This laptop offers excellent price to performance for students.
這台筆記型電腦為學生提供了極佳的性價比。

fraction /ˈfrækʃən/

noun · B2

意思:小部分;分之一

解說:在此語境中,指「一小部分」或「極低的比例」,常用來形容成本遠低於對手。

影片原句
at a fraction of their cost
以極低的成本
延伸例句
You can buy this car at a fraction of the price of a luxury brand.
你可以用豪華品牌價格的一小部分買到這輛車。

glaze /ɡleɪz/

verb · C2

意思:吹捧;過度美化

解說:網路俚語,源自「glazing」,意指過度讚美或奉承某人,使其看起來比實際更好。

影片原句
I'm not trying to glaze them or Elon Musk or anything
我不是在試圖吹捧他們或埃隆·馬斯克或什麼的
延伸例句
Stop glazing the product; it has several bugs.
別再吹捧那個產品了;它有好幾個錯誤。

subpar /ˌsʌbˈpɑr/

adjective · C1

意思:低於標準的;不如預期的

解說:形容表現或品質低於預期或行業標準。

影片原句
because they're pretty they're pretty subpar compared to any of the other labs out there
因為與其他實驗室相比,他們的表現相當遜色
延伸例句
The service was subpar during the holiday rush.
在假期高峰期,服務品質低於標準。

competitive /kəmˈpɛtətɪv/

adjective · B2

意思:有競爭力的

解說:指在價格、性能或質量方面足以與對手抗衡。

影片原句
was quite competitive based off what it cost
在成本基礎上相當有競爭力
延伸例句
Their prices are very competitive in the local market.
他們在當地市場的價格非常有競爭力。

benchmark /ˈbɛntʃmɑrk/

noun · C1

意思:基準測試;標準

解說:用於評估系統或模型性能的標準測試或指標。

影片原句
If we take a look at the benchmarks , for example
如果我們查看基準測試,例如
延伸例句
This software sets a new benchmark for speed.
這個軟體為速度設定了新的基準。

tie /taɪ/

verb · B1

意思:並列;打平

解說:指得分或成績相同,沒有勝負之分。

影片原句
So it ties it and it's much cheaper
所以它並列第一,而且便宜得多
延伸例句
The two teams tied at 2-2.
兩隊以 2-2 打平。

lineup /ˈlaɪnʌp/

noun · B2

意思:陣容;系列

解說:指一組人或事物,特別是指產品線或團隊成員。

影片原句
strongest model lineup from each lab
所謂的「最強」模型陣容來自每個實驗室
延伸例句
The company announced a new lineup of smartphones.
公司宣布了一系列新的智慧型手機。

stride /straɪd/

noun · C1

意思:進展;大步

解說:常用於片語「make strides」,意指取得重大進展或進步。

影片原句
helped SpaceX make some strides in the AI space this year
幫助 SpaceX 在今年在 AI 領域取得了一些進展
延伸例句
The team has made great strides in research.
團隊在研究方面取得了巨大進展。

sustainable /səˈsteɪnəbəl/

adjective · C1

意思:可持續的

解說:指能夠長期維持而不耗盡資源或導致負面後果。

影片原句
to be sustainable for the SpaceX team in the long run
對 SpaceX 團隊在長期來看是否具可持續性
延伸例句
Is this business model sustainable in the current economy?
在當前經濟環境下,這個商業模式是可持續的嗎?

adopt /əˈdɑpt/

verb · B2

意思:採用;採納

解說:指開始使用某種技術、方法或習慣。

影片原句
I would expect a lot of people to start adopting Grock
那麼我預期會有很多人開始採用 Grock
延伸例句
Many companies are adopting remote work policies.
許多公司正在採用遠程工作政策。

indicator /ˈɪndɪkeɪtər/

noun · B2

意思:指標;指示器

解說:指顯示某種情況或趨勢的信號或數據。

影片原句
price to performance ratio is becoming a critical I would say indicator in 2026
性價比會成為我在2026年認為的關鍵指標
延伸例句
Unemployment rate is a key economic indicator.
失業率是一個關鍵的經濟指標。

deteriorate /dɪˈtɪriəreɪt/

verb · C1

意思:惡化;變壞

解說:指品質、狀況或健康狀況逐漸變差。

影片原句
but over time they kind of deteriorate in their quality
但隨著時間推移,它們的質量可能會逐漸下降
延伸例句
His health began to deteriorate after the surgery.
手術後,他的健康開始惡化。

outperform /ˌaʊtpərˈfɔrm/

verb · C1

意思:優於;表現超過

解說:指在表現、成績或利潤上超過對手或預期。

影片原句
even outperform it and be at 57 times cheaper per cost
持平甚至超越它,並且成本每百萬輸入令牌便宜 57 倍
延伸例句
The new strategy outperformed the old one significantly.
新策略顯著優於舊策略。

leak /liːk/

verb · B2

意思:洩露

解說:指秘密資訊被非正式地公開或透露出去。

影片原句
because there was a investor report that leaked
因為有一份投資者報告洩露
延伸例句
Details of the merger were leaked to the press.
合併的細節被洩露給媒體。

gear /ɡɪr/

verb · C1

意思:針對...設計;使適合

解說:常用於片語「be geared at/towards」,指針對特定目的或受眾進行設計或調整。

影片原句
this model is once again really geared at cyber security
這個模型再次確實是針對網路安全所設計的
延伸例句
The course is geared towards beginners.
這門課程是針對初學者設計的。

preview /ˈprɛvju/

noun · B1

意思:預覽版;預覽

解說:指在正式發布前展示的早期版本或樣本。

影片原句
compared to the preview version of the model
與該模型的預覽版相比
延伸例句
We watched a preview of the new movie.
我們觀看了一部新電影的預覽版。

句型解說(含實例)

not only do we have..., ... also...

意思:我們不僅有...,...也...

接續:Not only + 助動詞/be + 主語 + 動詞, 主語 + also + 動詞

解說:用於強調兩個同時存在的事實,後半句通常使用倒裝結構以加強語氣。

影片原句
because not only do we have a new model from the Deepseek team , the SpaceX team also dropped Grock 4
因為我們不僅擁有一個來自 Deepseek 團隊的新模型,SpaceX 團隊也推出了 Grock 4
實例
  1. Not only did she finish the project early, but she also exceeded the quality standards.
    她不僅提前完成了項目,還超出了質量標準。
  2. Not only is the price low, but the quality is also high.
    不僅價格低,質量也高。

if I were to summarize... in one sentence

意思:如果我要用一句話來總結...

接續:If + subject + were to + verb, ...

解說:使用虛擬語態表達假設性的總結,語氣較為委婉且正式。

影片原句
if I were to summarize today's release in one simple sentence , it is that the price to performance is becoming a very , very important thing
如果我要用一句簡單的話來總結今天的發布,那就是性價比正變得非常、非常重要的一環
實例
  1. If I were to summarize the book in one sentence, it is about love and loss.
    如果我要用一句話來總結這本書,那就是關於愛與失去。
  2. If I were to describe him in one word, I would say 'brave'.
    如果我要用一個詞來形容他,我會說「勇敢」。

be geared at/towards

意思:針對...設計;以...為目標

接續:Subject + be + geared + at/towards + Object

解說:形容產品、服務或策略是針對特定受眾或目的而設計的。

影片原句
this model is once again really geared at cyber security
這個模型再次確實是針對網路安全所設計的
實例
  1. The marketing campaign is geared towards young adults.
    這次營銷活動是針對年輕成年人設計的。
  2. This software is geared towards professional designers.
    這個軟體是針對專業設計師設計的。

make strides in

意思:在...方面取得進展

接續:Subject + make + strides + in + Noun/Gerund

解說:形容在某個領域取得顯著進步或發展。

影片原句
helped SpaceX make some strides in the AI space this year
幫助 SpaceX 在今年在 AI 領域取得了一些進展
實例
  1. The company has made strides in renewable energy.
    該公司在可再生能源領域取得了進展。
  2. We have made strides in reducing waste.
    我們在減少廢物方面取得了進展。

be tied to

意思:與...並列;與...相同

接續:Subject + be + tied to + Object

解說:指得分、排名或狀態與另一對象相同。

影片原句
It's tied to GPT 5
它與 GPT 5 並列
實例
  1. The winner is tied with the runner-up.
    獲勝者與亞軍並列。
  2. His score is tied to the average of the class.
    他的分數與班級平均分並列。

be based off of

意思:基於...;根據...

接續:Subject + be + based off of + Noun

解說:指某事物是根據另一事物來評估、設計或判斷的。

影片原句
was quite competitive based off what it cost
在成本基礎上相當有競爭力
實例
  1. The decision was based off of the latest data.
    這個決定是基於最新數據做出的。
  2. The design is based off of traditional architecture.
    這個設計是基於傳統建築風格的。

be ahead of / be behind

意思:領先於...;落後於...

接續:Subject + be + ahead of/behind + Object

解說:用於比較得分、進度或地位,指出優劣關係。

影片原句
which is ahead of GLM 5
領先 GLM 5
實例
  1. Our team is ahead of schedule.
    我們的團隊領先於進度表。
  2. He is still behind in his studies.
    他的學習仍然落後。

be early to tell

意思:現在說...還為時過早

接續:It + be + early to tell + Clause

解說:表示目前資訊不足,無法做出最終判斷或結論。

影片原句
it's a little too early to tell yet
現在還說得太早
實例
  1. It's too early to tell if the project will succeed.
    現在說項目是否會成功還為時過早。
  2. It's a bit early to tell how the market will react.
    現在說市場會如何反應還為時過早。