實際影片長度:10:33.000。原文、繁中、雙語可點擊句子跳轉影片。
0:00.000–0:04.410
So , looks like we just got flashed by Google DeepMind once again
0:04.420–0:06.412
because they just dropped Gemini 3.7 Flash today.
0:06.422–0:06.952
Yes, their newest model,
0:06.952–0:07.705
which is once again a Flash model,
0:07.715–0:09.139
Yes , their newest model ,
0:09.139–0:11.162
which is once again a Flash model ,
0:11.172–0:12.046
is here .
0:12.056–0:15.473
And this is only after 3 weeks since Gemini 3.6 Flash.
0:15.483–0:16.544
6 Flash .
0:16.554–0:17.414
And what they are saying is that
0:17.414–0:19.034
Gemini 3.7 Flash is their most intelligent workhorse model.
0:19.044–0:22.649
7 Flash is their most intelligent workhorse model .
0:22.659–0:23.489
And this is coming 3 weeks after 3.6 Flash
0:23.489–0:24.772
because of developer feedback and algorithmic innovations.
0:24.782–0:29.103
6 six flash because of developer feedback and algorithmic innovations
0:29.103–0:35.920
.
0:35.920–0:36.000
And that is why they're saying
0:36.000–0:36.080
the Gemini 3.7 Flash model is,
0:36.080–0:36.160
you know, out with so much improvement cuz
0:36.160–0:36.240
7 Flash model is , you know , out with so much improvement cuz
0:36.240–0:37.105
if we take a look at the benchmarks ,
0:37.105–0:38.350
this model is way better
0:38.360–0:39.552
than the Gemini 3.6 Flash.
0:39.562–0:40.622
6 Flash .
0:40.632–0:42.234
But I think this is also because
0:42.234–0:43.836
we know internally that they're
0:43.846–0:44.926
planning on cancelling Gemini 3.5 Pro completely
0:44.926–0:45.717
and working on Gemini 4 at the moment.
0:45.727–0:49.647
5 Pro completely and working on Gemini 4 at the moment .
0:49.657–0:52.595
So maybe the improvements they made with the Pro model that they're
0:52.605–0:54.432
not going to be dropping anymore .
0:54.442–0:58.360
Maybe they repackaged it as a flash model because this model is
0:58.370–0:59.537
better than 3 .
0:59.547–1:00.449
the 3.5 Pro,
1:00.449–1:03.495
it's quite likely that is the case.
1:03.505–1:04.083
the 3 .
1:04.093–1:04.462
5 Pro ,
1:04.462–1:06.949
it's quite likely that is the case .
1:06.959–1:07.882
We'll just ignore that
1:07.882–1:08.958
and put that into the side .
1:08.968–1:11.450
We don't know when we're getting a new Pro model from the Google
1:11.460–1:12.685
we will see that yes,
1:12.685–1:15.058
it is stronger than the 3.6 Flash model.
1:15.068–1:16.407
will see that yes ,
1:16.407–1:18.319
it is stronger than the 3 .
1:18.329–1:19.628
6 Flash model .
1:19.638–1:20.186
For example ,
1:20.186–1:21.502
when it comes to code quality ,
1:21.502–1:22.654
production code quality
1:22.664–1:23.896
the model is producing 43.6
1:23.896–1:25.128
versus the flash model 34.4.
1:25.138–1:28.110
6 versus the flash model 34 .
1:28.120–1:29.213
4 .
1:29.223–1:30.678
And then Sonet 5 is at 42.7
1:30.678–1:32.352
and the Terra model is at 41.3.
1:32.362–1:35.570
7 and the Terra model is at 41 .
1:35.580–1:36.362
3 .
1:36.372–1:36.514
So ,
1:36.514–1:37.788
one thing to remember ,
1:37.788–1:39.346
since this is a flash model ,
1:39.356–1:41.536
you're not going to see them compare this to Opus 5
1:41.536–1:43.572
or any of the other stronger models .
1:43.582–1:46.775
The reason being cuz this is not their strongest tier .
1:46.785–1:47.650
and the 3 .
1:47.660–1:48.331
1 Pro model
1:48.331–1:51.390
when they finally decide to upgrade it to Gemini 4
1:51.400–1:51.730
Pro ,
1:51.730–1:55.354
then we'll see it being compared to Opus 5
1:55.354–1:56.013
or GPT 5 .
1:56.023–1:56.842
6 Soul .
1:56.852–1:57.427
But anyways ,
1:57.427–1:59.728
we can see that this model is an improvement from
1:59.738–2:00.209
3 .
2:00.219–2:00.757
6 Flash
2:00.757–2:04.699
and I would say that's basically the biggest result or
2:04.709–2:07.182
the biggest update that we saw with the new model
2:07.182–2:07.986
because it does
2:07.996–2:09.982
not really like change up things a lot
2:09.982–2:11.391
is still like not becoming
2:11.401–2:12.565
the number one model
2:12.565–2:14.208
or the number one flash model .
2:14.218–2:15.916
Even on some benchmarks ,
2:15.916–2:18.547
Deepseek version 4 flash is actually
2:18.557–2:19.106
cheaper
2:19.106–2:21.540
and more intelligent than this model .
2:21.550–2:24.249
But let's just take a look at the benchmarks that they have published
2:24.249–2:30.087
.
2:30.087–2:30.167
On Long Horizon software engineering ,
2:30.167–2:30.247
this model is a little bit
2:30.247–2:30.327
behind GPT 5 .
2:30.327–2:30.407
6 Terra ,
2:30.407–2:32.039
which sits at 69 .
2:32.049–2:32.243
6
2:32.243–2:33.991
and then 65 .
2:34.001–2:35.751
3 is the Flash model .
2:35.761–2:37.686
The older Flash model 3 .
2:37.696–2:39.290
6 sits at 48 .
2:39.300–2:40.230
6 .
2:40.240–2:41.987
Sonic 5 sits at 53 .
2:41.997–2:42.941
8 8 .
2:42.951–2:46.380
And the new player that is finally being included on benchmarks
2:46.380–2:46.984
,
2:46.984–2:49.818
which even the Google DeepMind team is considering with their
2:49.828–2:50.388
launch ,
2:50.388–2:51.508
is Muse Spark 1 .
2:51.518–2:51.633
2 ,
2:51.633–2:53.699
which is sitting at 54 .
2:53.709–2:54.544
9 .
2:54.554–2:54.705
So ,
2:54.705–2:57.040
welcome Meta to the benchmark charts
2:57.040–2:58.698
because now we're starting
2:58.708–3:01.254
to see it appear on more and more benchmarks .
3:01.264–3:03.571
And if we also take a look at web development ,
3:03.581–3:07.699
this model is getting a ELO score of 1588 versus their old model
3:07.709–3:08.663
is 1538 .
3:08.673–3:09.773
So , not a crazy difference ,
3:09.773–3:11.345
but compared to everything else out
3:11.355–3:12.726
there , this is number one .
3:12.736–3:14.291
When I say everything else out there ,
3:14.291–3:14.757
once again ,
3:14.767–3:16.671
compared to all the mid tier models ,
3:16.671–3:18.116
this is beating all of them
3:18.126–3:19.891
when it comes to web development .
3:19.901–3:21.486
Now , this model is probably going to be used in enterprises a lot
3:21.486–3:22.979
just because of Google's footprint in the enterprise space
3:22.989–3:25.758
lot just because of Google's footprint in the enterprise space
3:25.758–3:32.640
,
3:32.640–3:32.720
but we are seeing this model achieve on the automation bench
3:32.720–3:32.800
30 .
3:32.800–3:32.880
4 .
3:32.880–3:32.960
And GPT 5 .
3:32.960–3:34.010
6 Terra sits at 23 .
3:34.020–3:34.890
6 .
3:34.900–3:38.560
So , yes , this model is stronger than the other Flash models out
3:38.570–3:41.588
there , but as I said , they haven't included Deepseek version
3:41.598–3:44.142
for Flash because if they do , in some areas ,
3:44.152–3:46.458
that model is actually quite better than the 3 .
3:46.468–3:48.055
7 Flash model .
3:48.065–3:51.409
Before we continue , if you're building AI agents or just messing
3:51.419–3:54.285
around with them , Arcade is worth knowing about .
3:54.295–3:57.318
It's the runtime that lets your agent actually do things instead
3:57.328–4:00.433
of just talking about them because that's the gap right now .
4:00.443–4:02.109
The models are smart enough .
4:02.119–4:05.230
Your agent can figure out exactly what needs to happen in your
4:05.240–4:07.309
email , your Slack , your CRM .
4:07.319–4:09.389
It just can't go in and do it .
4:09.399–4:12.431
And the reason isn't intelligence , it's permissions .
4:12.441–4:15.390
Something has to prove the agent is allowed to act on behalf of
4:15.400–4:18.350
a specific person in a specific account .
4:18.360–4:22.032
That's the messy part everyone runs into , and it's the partit
4:22.042–4:23.393
actually handles for you .
4:23.403–4:26.835
So instead of your agent using one shared login for everybody ,
4:26.845–4:30.516
it acts as whoever is actually signed in with exactly the access
4:30.526–4:31.878
that person has .
4:31.888–4:34.760
If they can't see something , the agent can't either .
4:34.770–4:37.482
And you never have to touch any of that setup yourself .
4:37.492–4:38.843
Then there's the tools .
4:38.853–4:41.725
Arcade has thousands of them already built for Gmail ,
4:41.735–4:45.728
Google Drive , Slack , Notion , Salesforce , most of the apps people
4:45.738–4:49.090
already work in , and they are built specifically for AI to use
4:49.090–4:52.440
.
4:52.440–4:52.520
So the agent gets it right the first time instead of guessing and
4:52.520–4:53.733
failing and retrying .
4:53.743–4:56.855
It also keeps a record of everything , what the agent did ,
4:56.865–5:00.138
for who and where , which matters a lot the moment other people
5:00.148–5:01.739
start using the thing you built .
5:01.749–5:04.140
And it works with whatever you're already using .
5:04.150–5:07.182
Any model , any framework , cloud , cursor , chat ,
5:07.192–5:08.943
GPT , doesn't matter .
5:08.953–5:11.986
So you're not just giving an AI a list of tools and hoping it works
5:11.986–5:16.206
.
5:16.206–5:16.286
You're giving it a place where it can safely take real actions
5:16.286–5:16.469
in real apps .
5:16.479–5:19.030
It's free to start and the link is in the description .
5:19.040–5:22.066
Thank you once again for Arcade for sponsoring today's video .
5:22.076–5:24.378
Now , let's get back into the video .
5:24.388–5:27.178
Now , one thing to note is that the Gemini 3 .
5:27.188–5:31.007
7 Flash model through the end of this year , so end of 2026 ,
5:31.017–5:34.283
they have a cheap pricing model that they're placing on the model
5:34.283–5:38.240
.
5:38.240–5:38.786
75 per 1 million input tokens and 3 .
5:38.796–5:40.822
75 for 1 million output tokens .
5:40.832–5:44.266
So , it's a competitive price , but this is only for the next 6
5:44.276–5:44.899
months .
5:44.909–5:48.305
Because after those six months are done , the model's pricing is
5:48.315–5:50.528
actually , you know , a little bit more expensive .
5:50.538–5:52.922
And now they show it at the bottom over here ,
5:52.932–5:56.990
you can see that after starting January of 2027 ,
5:57.000–5:58.763
it will become 1 .
5:58.773–6:00.728
50 per input and 7 .
6:00.738–6:01.947
50 per output .
6:01.957–6:05.605
So yeah , it's still cheap compared to the other frontier labs
6:05.615–6:08.896
, but it's not as cheap as , for example , MU Spark when the pricing
6:08.906–6:12.352
is updated or even the Deepsee version for Flash .
6:12.362–6:13.955
But across these benchmarks ,
6:13.955–6:15.879
we can see that this model is better
6:15.889–6:18.283
in many areas that they highlighted at the top .
6:18.293–6:21.768
But then they also have some other areas like long video understanding
6:21.778–6:24.452
, which this model excels at 85 .
6:24.462–6:26.501
4 versus 78 .
6:26.511–6:28.298
9 for the Terra model .
6:28.308–6:31.897
And then the old model was also pretty good at that ,
6:31.907–6:32.176
84 .
6:32.186–6:32.465
2 .
6:32.475–6:36.288
Then long context performance the model is at 97
6:36.288–6:37.432
and this model
6:37.442–6:41.273
the GPT 6 Terra one it sits at 93 .
6:41.283–6:42.080
5.
6:42.090–6:44.496
So yeah this model in summary it is better
6:44.496–6:45.840
so it's not all negative
6:45.850–6:48.424
but it's not all like you know that positive
6:48.424–6:49.600
where you are super
6:49.610–6:50.921
excited for Google Deep Mind
6:50.921–6:52.450
because as I said they're probably
6:52.460–6:55.689
still holding off their biggest release for Gemini 4 lineup .
6:55.699–6:58.635
Now , one thing people might have missed in their charts because
6:58.645–7:02.217
these charts sometimes are so messy to read and understand ,
7:02.227–7:03.412
3.
7:03.422–7:05.822
7 flash is worse than GPT 5.
7:05.832–7:06.188
6 Luna ,
7:06.188–7:07.826
which is the model over here ,
7:07.826–7:09.464
which is achieving a higher
7:09.474–7:13.841
score on this benchmark deepware engineering for about three times
7:13.851–7:14.809
the cost .
7:14.819–7:16.992
So , yeah , this is kind of interesting because yeah ,
7:17.002–7:18.340
the cost for Gemini 3.
7:18.350–7:21.276
7 Flash is a little bit more than what it looks like .
7:21.286–7:23.044
Now , some people are a little bit upset
7:23.044–7:23.781
and they're like ,
7:23.791–7:24.994
Oh , disgraceful .
7:25.004–7:28.734
Google left out soul , opus , and fable because Google is incredibly
7:28.744–7:29.447
behind .
7:29.457–7:31.115
But I think one thing we got to remember , guys ,
7:31.125–7:33.336
is that this model is a flash model .
7:33.346–7:35.242
It's not trying to be a pro model
7:35.242–7:36.986
or it's not trying to compete
7:36.996–7:40.458
That is probably going to be Gemini 4 .
7:40.468–7:42.651
That is probably going to be Gemini 4 .
7:42.661–7:44.077
So when Gemini 4 comes out ,
7:44.077–7:46.303
then I think it's okay for us to criticize
7:46.313–7:50.156
them if they don't include Soul , Opus , and Fable in their benchmark
7:50.166–7:51.401
charts because for now ,
7:51.401–7:53.350
I think what they have done is pretty
7:53.360–7:54.069
accurate .
7:54.079–7:55.750
One lab that I would have liked to see
7:55.750–7:56.864
or one model for example
7:56.874–7:59.478
I would have liked to see on that chart would be Deepseek version
7:59.488–7:59.793
4 Flash
7:59.793–8:02.134
because that would kind of spoil their release because
8:02.144–8:05.521
that model is way cheaper compared to the Gemini 3.
8:05.531–8:07.058
7 flash model lineup .
8:07.068–8:10.807
Now this model jumped from number 19 to 8 on the web development
8:10.817–8:11.095
area
8:11.095–8:12.555
and we see it over here now
8:12.555–8:14.224
and couple of models that are
8:14.234–8:17.751
ahead of it are Opus 5 obviously Kim K3 Quinn 3 .
8:17.761–8:20.384
8 Max Cloud Opus 5 Gro 4 .
8:20.394–8:22.409
6 6 , which is a model that came yesterday ,
8:22.409–8:23.353
which was a big win
8:23.363–8:25.900
for SpaceX, Fable 5, and 5.6.
8:25.910–8:26.352
6.
8:26.362–8:27.002
Sol.
8:27.012–8:29.966
So , yeah , this model is trying to compete in the web development
8:29.976–8:33.094
, but it's still behind all of these models , which is ,
8:33.104–8:35.305
you know , expected cuz it's a flash model .
8:35.315–8:37.111
It's not really a pro model .
8:37.121–8:40.296
Today, OpenAI has also launched a weight list for 5.6 so ultra fast mode.
8:40.306–8:42.036
6 so ultra fast mode .
8:42.046–8:45.623
access to now through Cerebras and GPT 5.
8:45.633–8:48.490
6 6o with this chip kind of running it is able to achieve an ultra
8:48.500–8:52.774
6 6o with this chip kind of running it is able to achieve an ultra
8:52.784–8:57.861
fast mode that generates up to 750 output tokens per second which
8:57.871–9:01.494
is about 14 times faster than the standard mode .
9:01.504–9:04.812
So we're getting a really fast version of GPT 5 .
9:04.822–9:09.505
6 so now this is supposed to be used for live or near production
9:09.515–9:13.292
workloads like you know real time voice support commerce what this
9:13.302–9:14.581
6 six sol to do is kind of be really fast in critical situations
9:14.591–9:18.277
6 six soul to do is kind of be really fast in critical situations
9:18.287–9:22.135
when people might be interacting with the AI agent like financial
9:22.145–9:26.395
research security response support I think is going to be a big
9:26.405–9:31.508
area where this new ultrafast mode will be kind of implemented business
9:31.518–9:35.098
and developer agents maybe yeah but I think like support or near
9:35.108–9:38.527
production workloads like real time voice I see this model really
9:38.537–9:42.378
excelling at that and obviously this is still a weightless mode
9:42.388–9:45.323
we don't know how many people are going to get access to this ,
9:45.333–9:47.791
how successful it is or what the pricing is .
9:47.801–9:50.497
I don't know if the pricing has changed because there's no information
9:50.507–9:53.283
in their actual , you know , blog post .
9:53.293–9:56.396
But as I mentioned , couple of areas where they mentioned that
9:56.406–9:58.233
this is going to be really important .
9:58.243–10:01.669
Customer support and voice , commerce , live research and experimentation
10:01.679–10:06.182
, financial research and security , incident response and reliability
10:06.182–10:08.288
.
10:08.288–10:09.514
But yeah , I think this partnership is going to be important for
10:09.524–10:11.044
OpenAI going forward .
10:11.054–10:12.978
But as I said , 14 times the speed .
10:12.988–10:14.589
What does that mean for cost ?
10:14.599–10:16.120
We don't know yet .
10:16.130–10:17.812
But that's it for today's video .
10:17.822–10:20.068
Make sure you guys are subscribed to the channel .
10:20.078–10:21.852
Follow our new newsletter as well at universe-of-ai.beehiiv.com
10:21.852–10:23.626
as well as subscribe to the main channel World of AI and support
10:23.636–10:24.014
behive .
10:24.024–10:26.911
com as well as subscribe to the main channel World of AI and support
10:26.921–10:29.396
us on X by following the Universe of AI as well .
10:29.406–10:32.401
Until then , I'll see you guys in the next
0:00.000–0:04.410
所以,看來我們又被 Google DeepMind 閃了一下
0:04.420–0:06.412
因為他們今天剛發布了 Gemini 3.7 Flash。
0:06.422–0:06.952
是的,他們最新的模型,
0:06.952–0:07.705
這又是一個 Flash 模型,
0:07.715–0:09.139
是的,他們最新的模型,
0:09.139–0:11.162
這又是一個 Flash 模型,
0:11.172–0:12.046
已經問世了。
0:12.056–0:15.473
這距離 Gemini 3.6 Flash 發布才過了 3 週。
0:15.483–0:16.544
[未翻譯]
0:16.554–0:17.414
他們所表示的是
0:17.414–0:19.034
Gemini 3.7 Flash 是他們最智能的主力工作模型。
0:19.044–0:22.649
7 Flash 是他們最智能的主力工作模型。
0:22.659–0:23.489
這是在 3.6 Flash 發布 3 週後推出的
0:23.489–0:24.772
因為開發者的反饋和算法創新。
0:24.782–0:29.103
6 six flash 因為開發者的反饋和算法創新
0:29.103–0:35.920
0:35.920–0:36.000
這就是為什麼他們說
0:36.000–0:36.080
Gemini 3.7 Flash 模型,
0:36.080–0:36.160
你知道,有如此多的改進,因為
0:36.160–0:36.240
7 Flash 模型,你知道,有如此多的改進,因為
0:36.240–0:37.105
如果我們來看看基準測試結果,
0:37.105–0:38.350
這個模型的表現遠比
0:38.360–0:39.552
Gemini 3.6 Flash 好。
0:39.562–0:40.622
6閃電
0:40.632–0:42.234
但我認為這也是因為
0:42.234–0:43.836
我們內部知道他們
0:43.846–0:44.926
正計畫完全取消 Gemini 3.5 Pro
0:44.926–0:45.717
,並目前專注於開發 Gemini 4。
0:45.727–0:49.647
完全取消 3.5 Pro,並目前專注於開發 Gemini 4。
0:49.657–0:52.595
所以也許他們在 Pro 模型上所做的改進,
0:52.605–0:54.432
他們將不再繼續推出。
0:54.442–0:58.360
也許他們將這些改進重新包裝成 Flash 模型,因為這個模型比
0:58.370–0:59.537
3.5
0:59.547–1:00.449
3.5 Pro 更好,
1:00.449–1:03.495
這很有可能就是事實。
1:03.505–1:04.083
3.5
1:04.093–1:04.462
3.5 Pro,
1:04.462–1:06.949
這很有可能就是事實。
1:06.959–1:07.882
我們就暫且忽略這一點
1:07.882–1:08.958
,把它放在一旁吧。
1:08.968–1:11.450
我們不知道何時能從 Google 獲得新的 Pro 型號
1:11.460–1:12.685
我們會看到,是的,
1:12.685–1:15.058
它比 3.6 Flash 模型更強大。
1:15.068–1:16.407
我們會看到,是的,
1:16.407–1:18.319
它比 3 更強大。
1:18.329–1:19.628
6 Flash 模型更強大。
1:19.638–1:20.186
例如,
1:20.186–1:21.502
當涉及到程式碼品質時,
1:21.502–1:22.654
生產環境的程式碼品質
1:22.664–1:23.896
該模型產生了 43.6
1:23.896–1:25.128
而 Flash 模型為 34.4。
1:25.138–1:28.110
6 而 Flash 模型為 34.
1:28.120–1:29.213
[未翻譯]
1:29.223–1:30.678
然後 Sonet 5 為 42.7
1:30.678–1:32.352
而 Terra 模型為 41.3。
1:32.362–1:35.570
7 而 Terra 模型為 41.
1:35.580–1:36.362
[未翻譯]
1:36.372–1:36.514
所以,
1:36.514–1:37.788
有一件事要記住,
1:37.788–1:39.346
因為這是一個 Flash 模型,
1:39.356–1:41.536
你不會看到有人將它與 Opus 5 進行比較
1:41.536–1:43.572
或任何其他更強大的模型。
1:43.582–1:46.775
原因是因為這並非它們最頂級的等級。
1:46.785–1:47.650
以及 3.
1:47.660–1:48.331
1 Pro 模型
1:48.331–1:51.390
當他們最終決定將其升級為 Gemini 4
1:51.400–1:51.730
[未翻譯]
1:51.730–1:55.354
我們才會看到它與 Opus 5
1:55.354–1:56.013
或 GPT 5 進行比較。
1:56.023–1:56.842
第六,靈魂。
1:56.852–1:57.427
不過,
1:57.427–1:59.728
我們可以看到這個模型是
1:59.738–2:00.209
[未翻譯]
2:00.219–2:00.757
6 Flash 的改進版
2:00.757–2:04.699
我認為這基本上就是我們在新模型中看到的最重大成果或
2:04.709–2:07.182
我們在新模型中看到的最重大更新
2:07.182–2:07.986
因為它確實
2:07.996–2:09.982
不太喜歡大幅改變事物
2:09.982–2:11.391
仍然沒有成為
2:11.401–2:12.565
排名第一的模型
2:12.565–2:14.208
或是排名第一的 Flash 模型。
2:14.218–2:15.916
即使在某些基準測試中,
2:15.916–2:18.547
Deepseek 的 v4 Flash 版本實際上
2:18.557–2:19.106
更便宜
2:19.106–2:21.540
且比這個模型更聰明。
2:21.550–2:24.249
但讓我們先來看看他們發布的基準測試結果
2:24.249–2:30.087
2:30.087–2:30.167
在長期軟體工程方面,
2:30.167–2:30.247
這個模型稍微
2:30.247–2:30.327
落後於 GPT 5。
2:30.327–2:30.407
[未翻譯]
2:30.407–2:32.039
得分為 69。
2:32.049–2:32.243
[未翻譯]
2:32.243–2:33.991
然後是 65。
2:34.001–2:35.751
3 是 Flash 模型。
2:35.761–2:37.686
較舊的 Flash 3。
2:37.696–2:39.290
6 得分為 48。
2:39.300–2:40.230
[未翻譯]
2:40.240–2:41.987
5 Sonic 得分為 53。
2:41.997–2:42.941
[未翻譯]
2:42.951–2:46.380
而終於被納入基準測試的新玩家
2:46.380–2:46.984
而終於被納入基準測試的新玩家,連 Google DeepMind 團隊在發布時也在考慮,正是 Muse Spark 1。
2:46.984–2:49.818
就連 Google DeepMind 團隊在他們的
2:49.828–2:50.388
發布
2:50.388–2:51.508
是 Muse Spark 1
2:51.518–2:51.633
[未翻譯]
2:51.633–2:53.699
目前得分為54分
2:53.709–2:54.544
第九名
2:54.554–2:54.705
所以,
2:54.705–2:57.040
歡迎 Meta 進入基準測試排行榜
2:57.040–2:58.698
因為我們現在開始
2:58.708–3:01.254
看到它出現在越來越多的基準測試中。
3:01.264–3:03.571
如果我們也來看看網頁開發,
3:03.581–3:07.699
該模型的 ELO 得分為 1588,而他們舊模型的
3:07.709–3:08.663
得分為 1538。
3:08.673–3:09.773
所以,差異並不算大,
3:09.773–3:11.345
但與其他所有產品相比
3:11.355–3:12.726
這是排名第一的。
3:12.736–3:14.291
當我說與其他所有產品相比時,
3:14.291–3:14.757
再一次,
3:14.767–3:16.671
與所有中階模型相比,
3:16.671–3:18.116
在網頁開發方面,
3:18.126–3:19.891
它擊敗了所有這些模型。
3:19.901–3:21.486
現在,這個模型可能會在企業中被大量使用,
3:21.486–3:22.979
僅僅是因為 Google 在企業領域的影響力,
3:22.989–3:25.758
僅僅是因為 Google 在企業領域的影響力,
3:25.758–3:32.640
3:32.640–3:32.720
但我們看到這個模型在自動化基準測試中
3:32.720–3:32.800
達到了 30。
3:32.800–3:32.880
而 GPT 5 的得分為 30
3:32.880–3:32.960
以及 GPT 5。
3:32.960–3:34.010
6 Terra 位於 23。
3:34.020–3:34.890
[未翻譯]
3:34.900–3:38.560
所以,是的,這個模型比其他現有的 Flash 模型更強大,
3:38.570–3:41.588
但正如我所說,他們沒有將 Deepseek 版本
3:41.598–3:44.142
納入 Flash,因為如果他們這樣做,在某些領域,
3:44.152–3:46.458
該模型實際上比 3 好得多。
3:46.468–3:48.055
7 Flash 模型更好。
3:48.065–3:51.409
在繼續之前,如果你正在構建 AI 代理或只是隨意
3:51.419–3:54.285
嘗試它們,Arcade 值得了解。
3:54.295–3:57.318
它是讓你的 AI 代理真正執行任務,而不只是空談的運行環境
3:57.328–4:00.433
因為這正是目前的缺口所在。
4:00.443–4:02.109
這些模型已經足夠聰明。
4:02.119–4:05.230
你的 AI 代理能精確判斷在你的
4:05.240–4:07.309
電子郵件、Slack 或 CRM 中需要做什麼。
4:07.319–4:09.389
但它無法直接進去執行。
4:09.399–4:12.431
原因不在於智力,而在於權限。
4:12.441–4:15.390
必須有某個機制證明該代理獲授權代表
4:15.400–4:18.350
特定帳戶中的特定人員行事。
4:18.360–4:22.032
這就是每個人都會遇到的棘手部分,而
4:22.042–4:23.393
Arcade 會為你處理這部分。
4:23.403–4:26.835
因此,你的 AI 代理不會使用所有人共用的單一登入資訊,
4:26.845–4:30.516
而是以實際登入者的身份行事,並擁有該人員
4:30.526–4:31.878
所具備的權限。
4:31.888–4:34.760
如果他們看不到某些內容,AI 代理也看不到。
4:34.770–4:37.482
而且你完全不需要自己處理任何設定。
4:37.492–4:38.843
接下來是工具部分。
4:38.853–4:41.725
Arcade 已經內建了數千種工具,適用於 Gmail、
4:41.735–4:45.728
Google Drive、Slack、Notion、Salesforce,以及人們
4:45.738–4:49.090
日常使用的多數應用程式,這些工具都是專門為 AI 使用而設計的
4:49.090–4:52.440
4:52.440–4:52.520
因此,代理程式能一次就正確執行,而不是靠猜測、失敗然後重試。
4:52.520–4:53.733
失敗並重試。
4:53.743–4:56.855
它還會記錄所有內容,包括代理程式做了什麼、
4:56.865–5:00.138
以及對象和地點,這在其他人介入的瞬間顯得至關重要
5:00.148–5:01.739
開始使用你建構的東西時,顯得非常重要。
5:01.749–5:04.140
而且它與你目前正在使用的任何工具都能相容。
5:04.150–5:07.182
任何模型、任何框架、雲端、Cursor、聊天、
5:07.192–5:08.943
GPT,都不重要。
5:08.953–5:11.986
因此,你不僅僅是給 AI 一組工具清單並希望它能運作
5:11.986–5:16.206
5:16.206–5:16.286
你賦予它一個可以在真實應用程式中安全採取實際行動的空間。
5:16.286–5:16.469
在真實應用程式中。
5:16.479–5:19.030
免費開始使用,連結在描述中。
5:19.040–5:22.066
再次感謝 Arcade 贊助今天的影片。
5:22.076–5:24.378
現在,讓我們回到影片中。
5:24.388–5:27.178
現在,有一點需要注意,那就是 Gemini 3.
5:27.188–5:31.007
7 Flash 模型直到今年年底,也就是 2026 年底,
5:31.017–5:34.283
他們為該模型提供了一個廉價的定價方案
5:34.283–5:38.240
5:38.240–5:38.786
每百萬個輸入記號75。
5:38.796–5:40.822
每百萬個輸出記號3.75。
5:40.832–5:44.266
所以,這是一個有競爭力的價格,但僅限於接下來的6
5:44.276–5:44.899
個月。
5:44.909–5:48.305
因為在那六個月結束後,該模型的定價
5:48.315–5:50.528
實際上,你知道,會稍微貴一點。
5:50.538–5:52.922
現在他們在下方顯示了這一點,
5:52.932–5:56.990
你可以看到,從2027年1月開始,
5:57.000–5:58.763
它將變成1.
5:58.773–6:00.728
50的輸入和7.
6:00.738–6:01.947
50的輸出。
6:01.957–6:05.605
所以,是的,與其他前沿實驗室相比,它仍然很便宜,
6:05.615–6:08.896
但不像MU Spark在定價更新時,甚至Flash的Deepsee版本那樣便宜。
6:08.906–6:12.352
或是更新,甚至是 Flash 的 Deepsee 版本。
6:12.362–6:13.955
我們可以看到該模型在他們在上方強調的許多領域中表現更好。
6:13.955–6:15.879
但他們還有一些其他領域,例如長影片理解,
6:15.889–6:18.283
該模型在此方面表現出色,達到85.
6:18.293–6:21.768
但他們也涵蓋其他領域,例如長影片理解
6:21.778–6:24.452
9是Terra模型的得分。
6:24.462–6:26.501
然後,舊模型在該方面也相當不錯,達到84.
6:26.511–6:28.298
9,適用於 Terra 模型。
6:28.308–6:31.897
然後舊模型在該方面也表現得相當不錯,
6:31.907–6:32.176
[未翻譯]
6:32.186–6:32.465
6:32.475–6:36.288
接著是長上下文效能,該模型達到 97
6:36.288–6:37.432
而這個模型
6:37.442–6:41.273
也就是 GPT 6 Terra 版本,它位於 93。
6:41.283–6:42.080
6:42.090–6:44.496
所以,總而言之,這個模型更好
6:44.496–6:45.840
所以並非全是負面評價
6:45.850–6:48.424
但也並非全然如你所知的那麼正面
6:48.424–6:49.600
讓你超級
6:49.610–6:50.921
為 Google DeepMind 感到興奮
6:50.921–6:52.450
因為正如我所說,他們可能
6:52.460–6:55.689
仍在擱置其 Gemini 4 系列的最大規模發布。
6:55.699–6:58.635
現在,人們可能在他們的圖表中錯過了一件事,因為
6:58.645–7:02.217
這些圖表有時難以閱讀和理解,
7:02.227–7:03.412
7:03.422–7:05.822
7 flash 比 GPT 5 差。
7:05.832–7:06.188
6 月神
7:06.188–7:07.826
也就是這裡的模型,
7:07.826–7:09.464
它在這個基準測試中取得了更高的
7:09.474–7:13.841
Deepware 工程分數,成本大約只有三分之一。
7:13.851–7:14.809
成本。
7:14.819–7:16.992
所以,是的,這有點意思,因為確實,
7:17.002–7:18.340
Gemini 3 的成本
7:18.350–7:21.276
7 Flash 的成本比看起來的要多一點。
7:21.286–7:23.044
現在,有些人有點不高興,
7:23.044–7:23.781
他們都說,
7:23.791–7:24.994
「太丟人了。」
7:25.004–7:28.734
Google 沒有納入 Soul、Opus 和 Fable,是因為 Google 遠遠落後。
7:28.744–7:29.447
落後。
7:29.457–7:31.115
但我認為我們必須記住一件事,各位,
7:31.125–7:33.336
這個模型是一款 Flash 模型。
7:33.346–7:35.242
它不是要成為 Pro 模型,
7:35.242–7:36.986
也不是要競爭
7:36.996–7:40.458
那應該是 Gemini 4。
7:40.468–7:42.651
那應該是 Gemini 4。
7:42.661–7:44.077
所以當 Gemini 4 推出時,
7:44.077–7:46.303
我認為我們才有資格批評
7:46.313–7:50.156
如果他們在基準測試中沒有納入 Soul、Opus 和 Fable,我認為我們可以批評他們
7:50.166–7:51.401
因為就目前而言,
7:51.401–7:53.350
我認為他們所做的事情相當
7:53.360–7:54.069
準確。
7:54.079–7:55.750
我希望能看到的一個實驗室,
7:55.750–7:56.864
或者舉例來說一個模型,
7:56.874–7:59.478
我希望能在那張圖表上看到的,會是 Deepseek 第
7:59.488–7:59.793
4 版 Flash
7:59.793–8:02.134
,因為這有點會洩露他們的發布消息,因為
8:02.144–8:05.521
該模型與 Gemini 3.7
8:05.531–8:07.058
Flash 模型陣容相比便宜得多。
8:07.068–8:10.807
現在這個模型在網頁開發領域的排名從第19名躍升至第8名
8:10.817–8:11.095
領域
8:11.095–8:12.555
從第 19 名躍升至第 8 名,
8:12.555–8:14.224
我們現在在這裡看到它,
8:14.234–8:17.751
還有幾個領先的模型,當然是 Opus 5、Kim K3 Quinn 3。
8:17.761–8:20.384
[未翻譯]
8:20.394–8:22.409
6 6,這是一個昨天推出的模型,
8:22.409–8:23.353
這是一大勝利
8:23.363–8:25.900
,對 SpaceX、Fable 5 和 5.6 來說。
8:25.910–8:26.352
[未翻譯]
8:26.362–8:27.002
8:27.012–8:29.966
所以,是的,這個模型試圖在網頁開發領域競爭
8:29.976–8:33.094
,但它仍然落後於這些模型,這
8:33.104–8:35.305
,你知道,是預料之中的,因為它是一個閃電模型。
8:35.315–8:37.111
它並不是真正的專業模型。
8:37.121–8:40.296
今天,OpenAI 也為 5.6 版推出了權重列表,用於超快模式。
8:40.306–8:42.036
6 版的超快模式。
8:42.046–8:45.623
現在可以透過 Cerebras 和 GPT 5 進行存取。
8:45.633–8:48.490
6 6o 使用這款晶片運行,能夠實現超快
8:48.500–8:52.774
6 6o 使用這款晶片運行,能夠實現超快
8:52.784–8:57.861
模式,每秒生成多達 750 個輸出 token,這
8:57.871–9:01.494
大約是標準模式的 14 倍速度。
9:01.504–9:04.812
所以我們得到了一個非常快速的 GPT 5 .
9:04.822–9:09.505
6 版,現在這應該用於即時或接近生產環境
9:09.515–9:13.292
的工作負載,比如你知道的即時語音支援商務,這
9:13.302–9:14.581
6 6 Sol 要做的是在關鍵情況下非常快速
9:14.591–9:18.277
6 6 Sol 要做的是在關鍵情況下非常快速
9:18.287–9:22.135
當人們可能與 AI 代理互動時,例如金融
9:22.145–9:26.395
研究、安全回應支援,我認為這將是一個很大的
9:26.405–9:31.508
這個新的超快模式將在業務領域實施
9:31.518–9:35.098
以及開發者代理,也許是,但我認為像支援或接近
9:35.108–9:38.527
生產環境的工作負載,例如即時語音,我認為這個模型在
9:38.537–9:42.378
這方面表現出色,顯然這仍然是一種無權重模式
9:42.388–9:45.323
我們不知道有多少人會獲得訪問權限,
9:45.333–9:47.791
它的成功程度如何,或者定價是多少。
9:47.801–9:50.497
我不知道定價是否有所改變,因為在他們的實際
9:50.507–9:53.283
部落格文章中沒有相關資訊,你知道的。
9:53.293–9:56.396
但正如我所提到的,他們提到有幾個領域
9:56.406–9:58.233
這部分將非常重要
9:58.243–10:01.669
客戶支援和語音、商務、即時研究和實驗
10:01.679–10:06.182
、金融研究和安全、事件回應和可靠性
10:06.182–10:08.288
10:08.288–10:09.514
不過我認為這項合作對未來很重要
10:09.524–10:11.044
OpenAI 未來的發展將很重要。
10:11.054–10:12.978
但正如我所說,速度快了 14 倍。
10:12.988–10:14.589
這對成本意味著什麼?
10:14.599–10:16.120
我們還不知道。
10:16.130–10:17.812
但今天的影片就到這裡。
10:17.822–10:20.068
記得訂閱我們的頻道
10:20.078–10:21.852
也請訂閱我們的新聞通訊,網址為 universe-of-ai.beehiiv.com
10:21.852–10:23.626
同時訂閱主頻道 World of AI 並支持
10:23.636–10:24.014
behive.com,以及訂閱主頻道 World of AI 並支持
10:24.024–10:26.911
behive.com,同時在 X 上關注 Universe of AI 來支持我們。
10:26.921–10:29.396
behive.com,同時在 X 上關注 Universe of AI 來支持我們。
10:29.406–10:32.401
在那之前,我們下期再見
0:00.000–0:04.410
So , looks like we just got flashed by Google DeepMind once again
所以,看來我們又被 Google DeepMind 閃了一下
0:04.420–0:06.412
because they just dropped Gemini 3.7 Flash today.
因為他們今天剛發布了 Gemini 3.7 Flash。
0:06.422–0:06.952
Yes, their newest model,
是的,他們最新的模型,
0:06.952–0:07.705
which is once again a Flash model,
這又是一個 Flash 模型,
0:07.715–0:09.139
Yes , their newest model ,
是的,他們最新的模型,
0:09.139–0:11.162
which is once again a Flash model ,
這又是一個 Flash 模型,
0:11.172–0:12.046
is here .
已經問世了。
0:12.056–0:15.473
And this is only after 3 weeks since Gemini 3.6 Flash.
這距離 Gemini 3.6 Flash 發布才過了 3 週。
0:15.483–0:16.544
6 Flash .
[未翻譯]
0:16.554–0:17.414
And what they are saying is that
他們所表示的是
0:17.414–0:19.034
Gemini 3.7 Flash is their most intelligent workhorse model.
Gemini 3.7 Flash 是他們最智能的主力工作模型。
0:19.044–0:22.649
7 Flash is their most intelligent workhorse model .
7 Flash 是他們最智能的主力工作模型。
0:22.659–0:23.489
And this is coming 3 weeks after 3.6 Flash
這是在 3.6 Flash 發布 3 週後推出的
0:23.489–0:24.772
because of developer feedback and algorithmic innovations.
因為開發者的反饋和算法創新。
0:24.782–0:29.103
6 six flash because of developer feedback and algorithmic innovations
6 six flash 因為開發者的反饋和算法創新
0:29.103–0:35.920
.
0:35.920–0:36.000
And that is why they're saying
這就是為什麼他們說
0:36.000–0:36.080
the Gemini 3.7 Flash model is,
Gemini 3.7 Flash 模型,
0:36.080–0:36.160
you know, out with so much improvement cuz
你知道,有如此多的改進,因為
0:36.160–0:36.240
7 Flash model is , you know , out with so much improvement cuz
7 Flash 模型,你知道,有如此多的改進,因為
0:36.240–0:37.105
if we take a look at the benchmarks ,
如果我們來看看基準測試結果,
0:37.105–0:38.350
this model is way better
這個模型的表現遠比
0:38.360–0:39.552
than the Gemini 3.6 Flash.
Gemini 3.6 Flash 好。
0:39.562–0:40.622
6 Flash .
6閃電
0:40.632–0:42.234
But I think this is also because
但我認為這也是因為
0:42.234–0:43.836
we know internally that they're
我們內部知道他們
0:43.846–0:44.926
planning on cancelling Gemini 3.5 Pro completely
正計畫完全取消 Gemini 3.5 Pro
0:44.926–0:45.717
and working on Gemini 4 at the moment.
,並目前專注於開發 Gemini 4。
0:45.727–0:49.647
5 Pro completely and working on Gemini 4 at the moment .
完全取消 3.5 Pro,並目前專注於開發 Gemini 4。
0:49.657–0:52.595
So maybe the improvements they made with the Pro model that they're
所以也許他們在 Pro 模型上所做的改進,
0:52.605–0:54.432
not going to be dropping anymore .
他們將不再繼續推出。
0:54.442–0:58.360
Maybe they repackaged it as a flash model because this model is
也許他們將這些改進重新包裝成 Flash 模型,因為這個模型比
0:58.370–0:59.537
better than 3 .
3.5
0:59.547–1:00.449
the 3.5 Pro,
3.5 Pro 更好,
1:00.449–1:03.495
it's quite likely that is the case.
這很有可能就是事實。
1:03.505–1:04.083
the 3 .
3.5
1:04.093–1:04.462
5 Pro ,
3.5 Pro,
1:04.462–1:06.949
it's quite likely that is the case .
這很有可能就是事實。
1:06.959–1:07.882
We'll just ignore that
我們就暫且忽略這一點
1:07.882–1:08.958
and put that into the side .
,把它放在一旁吧。
1:08.968–1:11.450
We don't know when we're getting a new Pro model from the Google
我們不知道何時能從 Google 獲得新的 Pro 型號
1:11.460–1:12.685
we will see that yes,
我們會看到,是的,
1:12.685–1:15.058
it is stronger than the 3.6 Flash model.
它比 3.6 Flash 模型更強大。
1:15.068–1:16.407
will see that yes ,
我們會看到,是的,
1:16.407–1:18.319
it is stronger than the 3 .
它比 3 更強大。
1:18.329–1:19.628
6 Flash model .
6 Flash 模型更強大。
1:19.638–1:20.186
For example ,
例如,
1:20.186–1:21.502
when it comes to code quality ,
當涉及到程式碼品質時,
1:21.502–1:22.654
production code quality
生產環境的程式碼品質
1:22.664–1:23.896
the model is producing 43.6
該模型產生了 43.6
1:23.896–1:25.128
versus the flash model 34.4.
而 Flash 模型為 34.4。
1:25.138–1:28.110
6 versus the flash model 34 .
6 而 Flash 模型為 34.
1:28.120–1:29.213
4 .
[未翻譯]
1:29.223–1:30.678
And then Sonet 5 is at 42.7
然後 Sonet 5 為 42.7
1:30.678–1:32.352
and the Terra model is at 41.3.
而 Terra 模型為 41.3。
1:32.362–1:35.570
7 and the Terra model is at 41 .
7 而 Terra 模型為 41.
1:35.580–1:36.362
3 .
[未翻譯]
1:36.372–1:36.514
So ,
所以,
1:36.514–1:37.788
one thing to remember ,
有一件事要記住,
1:37.788–1:39.346
since this is a flash model ,
因為這是一個 Flash 模型,
1:39.356–1:41.536
you're not going to see them compare this to Opus 5
你不會看到有人將它與 Opus 5 進行比較
1:41.536–1:43.572
or any of the other stronger models .
或任何其他更強大的模型。
1:43.582–1:46.775
The reason being cuz this is not their strongest tier .
原因是因為這並非它們最頂級的等級。
1:46.785–1:47.650
and the 3 .
以及 3.
1:47.660–1:48.331
1 Pro model
1 Pro 模型
1:48.331–1:51.390
when they finally decide to upgrade it to Gemini 4
當他們最終決定將其升級為 Gemini 4
1:51.400–1:51.730
Pro ,
[未翻譯]
1:51.730–1:55.354
then we'll see it being compared to Opus 5
我們才會看到它與 Opus 5
1:55.354–1:56.013
or GPT 5 .
或 GPT 5 進行比較。
1:56.023–1:56.842
6 Soul .
第六,靈魂。
1:56.852–1:57.427
But anyways ,
不過,
1:57.427–1:59.728
we can see that this model is an improvement from
我們可以看到這個模型是
1:59.738–2:00.209
3 .
[未翻譯]
2:00.219–2:00.757
6 Flash
6 Flash 的改進版
2:00.757–2:04.699
and I would say that's basically the biggest result or
我認為這基本上就是我們在新模型中看到的最重大成果或
2:04.709–2:07.182
the biggest update that we saw with the new model
我們在新模型中看到的最重大更新
2:07.182–2:07.986
because it does
因為它確實
2:07.996–2:09.982
not really like change up things a lot
不太喜歡大幅改變事物
2:09.982–2:11.391
is still like not becoming
仍然沒有成為
2:11.401–2:12.565
the number one model
排名第一的模型
2:12.565–2:14.208
or the number one flash model .
或是排名第一的 Flash 模型。
2:14.218–2:15.916
Even on some benchmarks ,
即使在某些基準測試中,
2:15.916–2:18.547
Deepseek version 4 flash is actually
Deepseek 的 v4 Flash 版本實際上
2:18.557–2:19.106
cheaper
更便宜
2:19.106–2:21.540
and more intelligent than this model .
且比這個模型更聰明。
2:21.550–2:24.249
But let's just take a look at the benchmarks that they have published
但讓我們先來看看他們發布的基準測試結果
2:24.249–2:30.087
.
2:30.087–2:30.167
On Long Horizon software engineering ,
在長期軟體工程方面,
2:30.167–2:30.247
this model is a little bit
這個模型稍微
2:30.247–2:30.327
behind GPT 5 .
落後於 GPT 5。
2:30.327–2:30.407
6 Terra ,
[未翻譯]
2:30.407–2:32.039
which sits at 69 .
得分為 69。
2:32.049–2:32.243
6
[未翻譯]
2:32.243–2:33.991
and then 65 .
然後是 65。
2:34.001–2:35.751
3 is the Flash model .
3 是 Flash 模型。
2:35.761–2:37.686
The older Flash model 3 .
較舊的 Flash 3。
2:37.696–2:39.290
6 sits at 48 .
6 得分為 48。
2:39.300–2:40.230
6 .
[未翻譯]
2:40.240–2:41.987
Sonic 5 sits at 53 .
5 Sonic 得分為 53。
2:41.997–2:42.941
8 8 .
[未翻譯]
2:42.951–2:46.380
And the new player that is finally being included on benchmarks
而終於被納入基準測試的新玩家
2:46.380–2:46.984
,
而終於被納入基準測試的新玩家,連 Google DeepMind 團隊在發布時也在考慮,正是 Muse Spark 1。
2:46.984–2:49.818
which even the Google DeepMind team is considering with their
就連 Google DeepMind 團隊在他們的
2:49.828–2:50.388
launch ,
發布
2:50.388–2:51.508
is Muse Spark 1 .
是 Muse Spark 1
2:51.518–2:51.633
2 ,
[未翻譯]
2:51.633–2:53.699
which is sitting at 54 .
目前得分為54分
2:53.709–2:54.544
9 .
第九名
2:54.554–2:54.705
So ,
所以,
2:54.705–2:57.040
welcome Meta to the benchmark charts
歡迎 Meta 進入基準測試排行榜
2:57.040–2:58.698
because now we're starting
因為我們現在開始
2:58.708–3:01.254
to see it appear on more and more benchmarks .
看到它出現在越來越多的基準測試中。
3:01.264–3:03.571
And if we also take a look at web development ,
如果我們也來看看網頁開發,
3:03.581–3:07.699
this model is getting a ELO score of 1588 versus their old model
該模型的 ELO 得分為 1588,而他們舊模型的
3:07.709–3:08.663
is 1538 .
得分為 1538。
3:08.673–3:09.773
So , not a crazy difference ,
所以,差異並不算大,
3:09.773–3:11.345
but compared to everything else out
但與其他所有產品相比
3:11.355–3:12.726
there , this is number one .
這是排名第一的。
3:12.736–3:14.291
When I say everything else out there ,
當我說與其他所有產品相比時,
3:14.291–3:14.757
once again ,
再一次,
3:14.767–3:16.671
compared to all the mid tier models ,
與所有中階模型相比,
3:16.671–3:18.116
this is beating all of them
在網頁開發方面,
3:18.126–3:19.891
when it comes to web development .
它擊敗了所有這些模型。
3:19.901–3:21.486
Now , this model is probably going to be used in enterprises a lot
現在,這個模型可能會在企業中被大量使用,
3:21.486–3:22.979
just because of Google's footprint in the enterprise space
僅僅是因為 Google 在企業領域的影響力,
3:22.989–3:25.758
lot just because of Google's footprint in the enterprise space
僅僅是因為 Google 在企業領域的影響力,
3:25.758–3:32.640
,
3:32.640–3:32.720
but we are seeing this model achieve on the automation bench
但我們看到這個模型在自動化基準測試中
3:32.720–3:32.800
30 .
達到了 30。
3:32.800–3:32.880
4 .
而 GPT 5 的得分為 30
3:32.880–3:32.960
And GPT 5 .
以及 GPT 5。
3:32.960–3:34.010
6 Terra sits at 23 .
6 Terra 位於 23。
3:34.020–3:34.890
6 .
[未翻譯]
3:34.900–3:38.560
So , yes , this model is stronger than the other Flash models out
所以,是的,這個模型比其他現有的 Flash 模型更強大,
3:38.570–3:41.588
there , but as I said , they haven't included Deepseek version
但正如我所說,他們沒有將 Deepseek 版本
3:41.598–3:44.142
for Flash because if they do , in some areas ,
納入 Flash,因為如果他們這樣做,在某些領域,
3:44.152–3:46.458
that model is actually quite better than the 3 .
該模型實際上比 3 好得多。
3:46.468–3:48.055
7 Flash model .
7 Flash 模型更好。
3:48.065–3:51.409
Before we continue , if you're building AI agents or just messing
在繼續之前,如果你正在構建 AI 代理或只是隨意
3:51.419–3:54.285
around with them , Arcade is worth knowing about .
嘗試它們,Arcade 值得了解。
3:54.295–3:57.318
It's the runtime that lets your agent actually do things instead
它是讓你的 AI 代理真正執行任務,而不只是空談的運行環境
3:57.328–4:00.433
of just talking about them because that's the gap right now .
因為這正是目前的缺口所在。
4:00.443–4:02.109
The models are smart enough .
這些模型已經足夠聰明。
4:02.119–4:05.230
Your agent can figure out exactly what needs to happen in your
你的 AI 代理能精確判斷在你的
4:05.240–4:07.309
email , your Slack , your CRM .
電子郵件、Slack 或 CRM 中需要做什麼。
4:07.319–4:09.389
It just can't go in and do it .
但它無法直接進去執行。
4:09.399–4:12.431
And the reason isn't intelligence , it's permissions .
原因不在於智力,而在於權限。
4:12.441–4:15.390
Something has to prove the agent is allowed to act on behalf of
必須有某個機制證明該代理獲授權代表
4:15.400–4:18.350
a specific person in a specific account .
特定帳戶中的特定人員行事。
4:18.360–4:22.032
That's the messy part everyone runs into , and it's the partit
這就是每個人都會遇到的棘手部分,而
4:22.042–4:23.393
actually handles for you .
Arcade 會為你處理這部分。
4:23.403–4:26.835
So instead of your agent using one shared login for everybody ,
因此,你的 AI 代理不會使用所有人共用的單一登入資訊,
4:26.845–4:30.516
it acts as whoever is actually signed in with exactly the access
而是以實際登入者的身份行事,並擁有該人員
4:30.526–4:31.878
that person has .
所具備的權限。
4:31.888–4:34.760
If they can't see something , the agent can't either .
如果他們看不到某些內容,AI 代理也看不到。
4:34.770–4:37.482
And you never have to touch any of that setup yourself .
而且你完全不需要自己處理任何設定。
4:37.492–4:38.843
Then there's the tools .
接下來是工具部分。
4:38.853–4:41.725
Arcade has thousands of them already built for Gmail ,
Arcade 已經內建了數千種工具,適用於 Gmail、
4:41.735–4:45.728
Google Drive , Slack , Notion , Salesforce , most of the apps people
Google Drive、Slack、Notion、Salesforce,以及人們
4:45.738–4:49.090
already work in , and they are built specifically for AI to use
日常使用的多數應用程式,這些工具都是專門為 AI 使用而設計的
4:49.090–4:52.440
.
4:52.440–4:52.520
So the agent gets it right the first time instead of guessing and
因此,代理程式能一次就正確執行,而不是靠猜測、失敗然後重試。
4:52.520–4:53.733
failing and retrying .
失敗並重試。
4:53.743–4:56.855
It also keeps a record of everything , what the agent did ,
它還會記錄所有內容,包括代理程式做了什麼、
4:56.865–5:00.138
for who and where , which matters a lot the moment other people
以及對象和地點,這在其他人介入的瞬間顯得至關重要
5:00.148–5:01.739
start using the thing you built .
開始使用你建構的東西時,顯得非常重要。
5:01.749–5:04.140
And it works with whatever you're already using .
而且它與你目前正在使用的任何工具都能相容。
5:04.150–5:07.182
Any model , any framework , cloud , cursor , chat ,
任何模型、任何框架、雲端、Cursor、聊天、
5:07.192–5:08.943
GPT , doesn't matter .
GPT,都不重要。
5:08.953–5:11.986
So you're not just giving an AI a list of tools and hoping it works
因此,你不僅僅是給 AI 一組工具清單並希望它能運作
5:11.986–5:16.206
.
5:16.206–5:16.286
You're giving it a place where it can safely take real actions
你賦予它一個可以在真實應用程式中安全採取實際行動的空間。
5:16.286–5:16.469
in real apps .
在真實應用程式中。
5:16.479–5:19.030
It's free to start and the link is in the description .
免費開始使用,連結在描述中。
5:19.040–5:22.066
Thank you once again for Arcade for sponsoring today's video .
再次感謝 Arcade 贊助今天的影片。
5:22.076–5:24.378
Now , let's get back into the video .
現在,讓我們回到影片中。
5:24.388–5:27.178
Now , one thing to note is that the Gemini 3 .
現在,有一點需要注意,那就是 Gemini 3.
5:27.188–5:31.007
7 Flash model through the end of this year , so end of 2026 ,
7 Flash 模型直到今年年底,也就是 2026 年底,
5:31.017–5:34.283
they have a cheap pricing model that they're placing on the model
他們為該模型提供了一個廉價的定價方案
5:34.283–5:38.240
.
5:38.240–5:38.786
75 per 1 million input tokens and 3 .
每百萬個輸入記號75。
5:38.796–5:40.822
75 for 1 million output tokens .
每百萬個輸出記號3.75。
5:40.832–5:44.266
So , it's a competitive price , but this is only for the next 6
所以,這是一個有競爭力的價格,但僅限於接下來的6
5:44.276–5:44.899
months .
個月。
5:44.909–5:48.305
Because after those six months are done , the model's pricing is
因為在那六個月結束後,該模型的定價
5:48.315–5:50.528
actually , you know , a little bit more expensive .
實際上,你知道,會稍微貴一點。
5:50.538–5:52.922
And now they show it at the bottom over here ,
現在他們在下方顯示了這一點,
5:52.932–5:56.990
you can see that after starting January of 2027 ,
你可以看到,從2027年1月開始,
5:57.000–5:58.763
it will become 1 .
它將變成1.
5:58.773–6:00.728
50 per input and 7 .
50的輸入和7.
6:00.738–6:01.947
50 per output .
50的輸出。
6:01.957–6:05.605
So yeah , it's still cheap compared to the other frontier labs
所以,是的,與其他前沿實驗室相比,它仍然很便宜,
6:05.615–6:08.896
, but it's not as cheap as , for example , MU Spark when the pricing
但不像MU Spark在定價更新時,甚至Flash的Deepsee版本那樣便宜。
6:08.906–6:12.352
is updated or even the Deepsee version for Flash .
或是更新,甚至是 Flash 的 Deepsee 版本。
6:12.362–6:13.955
But across these benchmarks ,
我們可以看到該模型在他們在上方強調的許多領域中表現更好。
6:13.955–6:15.879
we can see that this model is better
但他們還有一些其他領域,例如長影片理解,
6:15.889–6:18.283
in many areas that they highlighted at the top .
該模型在此方面表現出色,達到85.
6:18.293–6:21.768
But then they also have some other areas like long video understanding
但他們也涵蓋其他領域,例如長影片理解
6:21.778–6:24.452
, which this model excels at 85 .
9是Terra模型的得分。
6:24.462–6:26.501
4 versus 78 .
然後,舊模型在該方面也相當不錯,達到84.
6:26.511–6:28.298
9 for the Terra model .
9,適用於 Terra 模型。
6:28.308–6:31.897
And then the old model was also pretty good at that ,
然後舊模型在該方面也表現得相當不錯,
6:31.907–6:32.176
84 .
[未翻譯]
6:32.186–6:32.465
2 .
6:32.475–6:36.288
Then long context performance the model is at 97
接著是長上下文效能,該模型達到 97
6:36.288–6:37.432
and this model
而這個模型
6:37.442–6:41.273
the GPT 6 Terra one it sits at 93 .
也就是 GPT 6 Terra 版本,它位於 93。
6:41.283–6:42.080
5.
6:42.090–6:44.496
So yeah this model in summary it is better
所以,總而言之,這個模型更好
6:44.496–6:45.840
so it's not all negative
所以並非全是負面評價
6:45.850–6:48.424
but it's not all like you know that positive
但也並非全然如你所知的那麼正面
6:48.424–6:49.600
where you are super
讓你超級
6:49.610–6:50.921
excited for Google Deep Mind
為 Google DeepMind 感到興奮
6:50.921–6:52.450
because as I said they're probably
因為正如我所說,他們可能
6:52.460–6:55.689
still holding off their biggest release for Gemini 4 lineup .
仍在擱置其 Gemini 4 系列的最大規模發布。
6:55.699–6:58.635
Now , one thing people might have missed in their charts because
現在,人們可能在他們的圖表中錯過了一件事,因為
6:58.645–7:02.217
these charts sometimes are so messy to read and understand ,
這些圖表有時難以閱讀和理解,
7:02.227–7:03.412
3.
7:03.422–7:05.822
7 flash is worse than GPT 5.
7 flash 比 GPT 5 差。
7:05.832–7:06.188
6 Luna ,
6 月神
7:06.188–7:07.826
which is the model over here ,
也就是這裡的模型,
7:07.826–7:09.464
which is achieving a higher
它在這個基準測試中取得了更高的
7:09.474–7:13.841
score on this benchmark deepware engineering for about three times
Deepware 工程分數,成本大約只有三分之一。
7:13.851–7:14.809
the cost .
成本。
7:14.819–7:16.992
So , yeah , this is kind of interesting because yeah ,
所以,是的,這有點意思,因為確實,
7:17.002–7:18.340
the cost for Gemini 3.
Gemini 3 的成本
7:18.350–7:21.276
7 Flash is a little bit more than what it looks like .
7 Flash 的成本比看起來的要多一點。
7:21.286–7:23.044
Now , some people are a little bit upset
現在,有些人有點不高興,
7:23.044–7:23.781
and they're like ,
他們都說,
7:23.791–7:24.994
Oh , disgraceful .
「太丟人了。」
7:25.004–7:28.734
Google left out soul , opus , and fable because Google is incredibly
Google 沒有納入 Soul、Opus 和 Fable,是因為 Google 遠遠落後。
7:28.744–7:29.447
behind .
落後。
7:29.457–7:31.115
But I think one thing we got to remember , guys ,
但我認為我們必須記住一件事,各位,
7:31.125–7:33.336
is that this model is a flash model .
這個模型是一款 Flash 模型。
7:33.346–7:35.242
It's not trying to be a pro model
它不是要成為 Pro 模型,
7:35.242–7:36.986
or it's not trying to compete
也不是要競爭
7:36.996–7:40.458
That is probably going to be Gemini 4 .
那應該是 Gemini 4。
7:40.468–7:42.651
That is probably going to be Gemini 4 .
那應該是 Gemini 4。
7:42.661–7:44.077
So when Gemini 4 comes out ,
所以當 Gemini 4 推出時,
7:44.077–7:46.303
then I think it's okay for us to criticize
我認為我們才有資格批評
7:46.313–7:50.156
them if they don't include Soul , Opus , and Fable in their benchmark
如果他們在基準測試中沒有納入 Soul、Opus 和 Fable,我認為我們可以批評他們
7:50.166–7:51.401
charts because for now ,
因為就目前而言,
7:51.401–7:53.350
I think what they have done is pretty
我認為他們所做的事情相當
7:53.360–7:54.069
accurate .
準確。
7:54.079–7:55.750
One lab that I would have liked to see
我希望能看到的一個實驗室,
7:55.750–7:56.864
or one model for example
或者舉例來說一個模型,
7:56.874–7:59.478
I would have liked to see on that chart would be Deepseek version
我希望能在那張圖表上看到的,會是 Deepseek 第
7:59.488–7:59.793
4 Flash
4 版 Flash
7:59.793–8:02.134
because that would kind of spoil their release because
,因為這有點會洩露他們的發布消息,因為
8:02.144–8:05.521
that model is way cheaper compared to the Gemini 3.
該模型與 Gemini 3.7
8:05.531–8:07.058
7 flash model lineup .
Flash 模型陣容相比便宜得多。
8:07.068–8:10.807
Now this model jumped from number 19 to 8 on the web development
現在這個模型在網頁開發領域的排名從第19名躍升至第8名
8:10.817–8:11.095
area
領域
8:11.095–8:12.555
and we see it over here now
從第 19 名躍升至第 8 名,
8:12.555–8:14.224
and couple of models that are
我們現在在這裡看到它,
8:14.234–8:17.751
ahead of it are Opus 5 obviously Kim K3 Quinn 3 .
還有幾個領先的模型,當然是 Opus 5、Kim K3 Quinn 3。
8:17.761–8:20.384
8 Max Cloud Opus 5 Gro 4 .
[未翻譯]
8:20.394–8:22.409
6 6 , which is a model that came yesterday ,
6 6,這是一個昨天推出的模型,
8:22.409–8:23.353
which was a big win
這是一大勝利
8:23.363–8:25.900
for SpaceX, Fable 5, and 5.6.
,對 SpaceX、Fable 5 和 5.6 來說。
8:25.910–8:26.352
6.
[未翻譯]
8:26.362–8:27.002
Sol.
8:27.012–8:29.966
So , yeah , this model is trying to compete in the web development
所以,是的,這個模型試圖在網頁開發領域競爭
8:29.976–8:33.094
, but it's still behind all of these models , which is ,
,但它仍然落後於這些模型,這
8:33.104–8:35.305
you know , expected cuz it's a flash model .
,你知道,是預料之中的,因為它是一個閃電模型。
8:35.315–8:37.111
It's not really a pro model .
它並不是真正的專業模型。
8:37.121–8:40.296
Today, OpenAI has also launched a weight list for 5.6 so ultra fast mode.
今天,OpenAI 也為 5.6 版推出了權重列表,用於超快模式。
8:40.306–8:42.036
6 so ultra fast mode .
6 版的超快模式。
8:42.046–8:45.623
access to now through Cerebras and GPT 5.
現在可以透過 Cerebras 和 GPT 5 進行存取。
8:45.633–8:48.490
6 6o with this chip kind of running it is able to achieve an ultra
6 6o 使用這款晶片運行,能夠實現超快
8:48.500–8:52.774
6 6o with this chip kind of running it is able to achieve an ultra
6 6o 使用這款晶片運行,能夠實現超快
8:52.784–8:57.861
fast mode that generates up to 750 output tokens per second which
模式,每秒生成多達 750 個輸出 token,這
8:57.871–9:01.494
is about 14 times faster than the standard mode .
大約是標準模式的 14 倍速度。
9:01.504–9:04.812
So we're getting a really fast version of GPT 5 .
所以我們得到了一個非常快速的 GPT 5 .
9:04.822–9:09.505
6 so now this is supposed to be used for live or near production
6 版,現在這應該用於即時或接近生產環境
9:09.515–9:13.292
workloads like you know real time voice support commerce what this
的工作負載,比如你知道的即時語音支援商務,這
9:13.302–9:14.581
6 six sol to do is kind of be really fast in critical situations
6 6 Sol 要做的是在關鍵情況下非常快速
9:14.591–9:18.277
6 six soul to do is kind of be really fast in critical situations
6 6 Sol 要做的是在關鍵情況下非常快速
9:18.287–9:22.135
when people might be interacting with the AI agent like financial
當人們可能與 AI 代理互動時,例如金融
9:22.145–9:26.395
research security response support I think is going to be a big
研究、安全回應支援,我認為這將是一個很大的
9:26.405–9:31.508
area where this new ultrafast mode will be kind of implemented business
這個新的超快模式將在業務領域實施
9:31.518–9:35.098
and developer agents maybe yeah but I think like support or near
以及開發者代理,也許是,但我認為像支援或接近
9:35.108–9:38.527
production workloads like real time voice I see this model really
生產環境的工作負載,例如即時語音,我認為這個模型在
9:38.537–9:42.378
excelling at that and obviously this is still a weightless mode
這方面表現出色,顯然這仍然是一種無權重模式
9:42.388–9:45.323
we don't know how many people are going to get access to this ,
我們不知道有多少人會獲得訪問權限,
9:45.333–9:47.791
how successful it is or what the pricing is .
它的成功程度如何,或者定價是多少。
9:47.801–9:50.497
I don't know if the pricing has changed because there's no information
我不知道定價是否有所改變,因為在他們的實際
9:50.507–9:53.283
in their actual , you know , blog post .
部落格文章中沒有相關資訊,你知道的。
9:53.293–9:56.396
But as I mentioned , couple of areas where they mentioned that
但正如我所提到的,他們提到有幾個領域
9:56.406–9:58.233
this is going to be really important .
這部分將非常重要
9:58.243–10:01.669
Customer support and voice , commerce , live research and experimentation
客戶支援和語音、商務、即時研究和實驗
10:01.679–10:06.182
, financial research and security , incident response and reliability
、金融研究和安全、事件回應和可靠性
10:06.182–10:08.288
.
10:08.288–10:09.514
But yeah , I think this partnership is going to be important for
不過我認為這項合作對未來很重要
10:09.524–10:11.044
OpenAI going forward .
OpenAI 未來的發展將很重要。
10:11.054–10:12.978
But as I said , 14 times the speed .
但正如我所說,速度快了 14 倍。
10:12.988–10:14.589
What does that mean for cost ?
這對成本意味著什麼?
10:14.599–10:16.120
We don't know yet .
我們還不知道。
10:16.130–10:17.812
But that's it for today's video .
但今天的影片就到這裡。
10:17.822–10:20.068
Make sure you guys are subscribed to the channel .
記得訂閱我們的頻道
10:20.078–10:21.852
Follow our new newsletter as well at universe-of-ai.beehiiv.com
也請訂閱我們的新聞通訊,網址為 universe-of-ai.beehiiv.com
10:21.852–10:23.626
as well as subscribe to the main channel World of AI and support
同時訂閱主頻道 World of AI 並支持
10:23.636–10:24.014
behive .
behive.com,以及訂閱主頻道 World of AI 並支持
10:24.024–10:26.911
com as well as subscribe to the main channel World of AI and support
behive.com,同時在 X 上關注 Universe of AI 來支持我們。
10:26.921–10:29.396
us on X by following the Universe of AI as well .
behive.com,同時在 X 上關注 Universe of AI 來支持我們。
10:29.406–10:32.401
Until then , I'll see you guys in the next
在那之前,我們下期再見

影片筆記:Google Shipped Gemini 3.7 Flash And OpenAI Made GPT-5.6 14x Faster!

一句話總結

Google DeepMind 發布了定位為「智能工作馬」的 Gemini 3.7 Flash,在代碼與自動化基準測試中表現優異,但定價策略與部分性能仍落後於競爭對手;同時 OpenAI 透過 Cerebras 晶片推出 GPT 5.6 的「超快模式」,實現每秒 750 個 token 的生成速度,主要針對即時語音與商業場景。

核心重點

  1. Gemini 3.7 Flash 發布與定位
  • 在 Gemini 3.6 Flash 發布僅 3 週後推出。
  • Google DeepMind 官方定位為「最智能的工作馬(most intelligent workhorse model)」。
  • 開發動機基於開發者反饋與演算法創新。
  • 講者推測 Google 計劃取消 Gemini 3.5 Pro,專注開發 Gemini 4,而 3.7 Flash 整合了 Pro 級的改進,性能優於 3.5 Pro。
  • 不建議將其與 Opus 5 或 GPT 5 等頂級 Pro 模型直接比較,因其屬於 Flash 等級。
  1. 性能基準測試(Benchmarks)表現
  • 代碼質量:Gemini 3.7 Flash (43.6) 優於 Sonet 5 (42.7) 與 Terra 模型 (41.3)。
  • 長視角軟體工程:略落後於 GPT 5.6 Terra (69.6 vs 65.3)。
  • 網頁開發:ELO 分數 1588,在中階模型中排名第一,排名從第 19 躍升至第 8,但仍落後於 Opus 5、Kim K3 Quinn 3.8 Max Cloud 等頂級模型。
  • 自動化基準測試:表現優於其他 Flash 模型 (30.4 vs GPT 5.6 Terra 23.6)。
  • 長影片理解與長上下文:表現優異,分別為 85.4 和 97。
  • Deepware Engineering 基準測試:落後於 GPT 5.6 Luna,且成本約為後者的三倍。
  1. 定價策略
  • 短期優惠(至 2026 年底):輸入 token $0.75/百萬,輸出 token $3.75/百萬。
  • 長期定價(2027 年 1 月起):輸入 token $1.50/百萬,輸出 token $7.50/百萬。
  • 相比其他前沿實驗室仍屬便宜,但不如 Muse Spark 或 Deepseek Version 4 Flash 便宜。
  1. OpenAI GPT 5.6 超快模式(Ultra Fast Mode)
  • 透過 Cerebras 晶片運行,發布 GPT 5.6 的權重列表。
  • 性能指標:每秒生成高達 750 個輸出 token,速度約為標準模式的 14 倍。
  • 應用場景:即時或近生產負載,包括即時語音支援、商務、金融研究、安全、事件回應與可靠性。
  • 目前尚未公布定價資訊,存取權限也不確定。
  1. 贊助商 Arcade 介紹
  • 定位為 AI Agent 的運行時(Runtime),讓 Agent 能實際執行動作。
  • 核心功能包括權限管理(無需共用單一登入資訊)、內建數千個工具整合(Gmail, Slack, Salesforce 等)、行為記錄追蹤。
  • 支援任何模型、框架、雲端,費用免費開始使用。

詳細大綱

A. Gemini 3.7 Flash 發布與定位

  • 發布時間點:緊接在 Gemini 3.6 Flash 之後,僅相隔 3 週。
  • 官方定位:Google DeepMind 稱其為「最智能的工作馬(most intelligent workhorse model)」。
  • 開發動機:基於開發者反饋與演算法創新。
  • 內部策略推測
  • 講者推測 Google 計劃完全取消 Gemini 3.5 Pro,並專注開發 Gemini 4。
  • 3.7 Flash 可能整合了原本屬於 Pro 模型的改進,因此性能優於 3.5 Pro。
  • 目前不將其與 Opus 5 或 GPT 5 等頂級 Pro 模型直接比較,因為它屬於 Flash 等級。

B. 性能基準測試(Benchmarks)分析

  • 代碼質量(Code Quality)
  • Gemini 3.7 Flash:43.6
  • Sonet 5:42.7
  • Terra 模型:41.3
  • 舊版 Flash 模型:34.4
  • 結論:在生產代碼質量上表現優於其他 Flash 模型。
  • 長視角軟體工程(Long Horizon Software Engineering)
  • GPT 5.6 Terra:69.6
  • Gemini 3.7 Flash:65.3
  • 舊版 Flash 模型(3.6):48.6
  • Sonet 5:53.8
  • Muse Spark 1.2:54.9(首次出現在基準測試中,Meta 的新玩家)
  • 結論:略落後於 GPT 5.6 Terra。
  • 網頁開發(Web Development)
  • ELO 分數:1588(舊版為 1538)。
  • 結論:在中階模型中排名第一,擊敗所有其他中階模型,但仍落後於 Opus 5、Kim K3 Quinn 3.8 Max Cloud、Opus 5 Gro 4.6 6、Fable 5 等頂級模型。排名從第 19 名躍升至第 8 名。
  • 自動化基準測試(Automation Bench)
  • Gemini 3.7 Flash:30.4
  • GPT 5.6 Terra:23.6
  • 結論:強於其他 Flash 模型。
  • 長影片理解(Long Video Understanding)
  • Gemini 3.7 Flash:85.4
  • Terra 模型:78.9
  • 舊版模型:84.2
  • 結論:表現優異。
  • 長上下文性能(Long Context Performance)
  • Gemini 3.7 Flash:97
  • GPT 6 Terra:93.5
  • 結論:表現優異。
  • Deepware Engineering 基準測試
  • Gemini 3.7 Flash 表現落後於 GPT 5.6 Luna,且成本約為後者的三倍。

C. 定價策略

  • 短期優惠(至 2026 年底)
  • 輸入 token:每 100 萬個 $0.75
  • 輸出 token:每 100 萬個 $3.75
  • 註:此價格僅持續 6 個月。
  • 長期定價(2027 年 1 月起)
  • 輸入 token:每 100 萬個 $1.50
  • 輸出 token:每 100 萬個 $7.50
  • 競爭力評估
  • 相比其他前沿實驗室(Frontier Labs)仍屬便宜,但不如 Muse Spark 或 Deepseek Version 4 Flash 便宜。
  • 講者認為 Google 未將 Opus、Soul、Fable 列入基準測試是合理的,因為這是 Flash 模型,應等待 Gemini 4 發布後再進行頂級比較。

D. OpenAI GPT 5.6 超快模式(Ultra Fast Mode)

  • 發布內容:OpenAI 發布 GPT 5.6 的權重列表,透過 Cerebras 晶片運行。
  • 性能指標
  • 每秒生成高達 750 個輸出 token。
  • 速度約為標準模式的 14 倍。
  • 應用場景
  • 即時或近生產負載(Live or near production workloads)。
  • 即時語音支援、商務、金融研究、安全、事件回應與可靠性。
  • 不確定性
  • 尚未公布定價資訊。
  • 尚未確定有多少人能獲得存取權限。

E. 贊助商資訊:Arcade

  • 產品定位:AI Agent 的運行時(Runtime),讓 Agent 能實際執行動作而非僅是討論。
  • 核心功能
  • 權限管理:解決 Agent 代表特定用戶在特定帳戶中行動的權限問題,無需共用單一登入資訊。
  • 工具整合:內建數千個工具,支援 Gmail、Google Drive、Slack、Notion、Salesforce 等應用。
  • 記錄追蹤:記錄 Agent 的行為、對象與地點。
  • 兼容性:支援任何模型、框架、雲端、Cursor、Chat、GPT。
  • 費用:免費開始使用。

工具 / 模型 / 名詞整理

  • Google DeepMind / Google
  • Gemini 3.7 Flash
  • Gemini 3.6 Flash
  • Gemini 3.5 Pro
  • Gemini 4 (預計未來發布)
  • Gemini 4 Pro (預計未來發布)
  • OpenAI
  • GPT 5.6
  • GPT 5.6 Terra
  • GPT 5.6 Luna
  • GPT 5.6 6o (Ultra Fast Mode)
  • GPT 5.6 Soul (提及於基準測試比較中)
  • 其他模型與實驗室
  • Opus 5
  • Sonet 5
  • Terra 模型
  • Deepseek Version 4 Flash
  • Muse Spark 1.2 (Meta)
  • Kim K3 Quinn 3.8 Max Cloud
  • Opus 5 Gro 4.6 6
  • Fable 5
  • Fable 5.6.6 Sol
  • Deepware Engineering (基準測試名稱)
  • 基礎設施與平台
  • Cerebras (晶片供應商)
  • Arcade (贊助商,AI Agent Runtime)
  • Gmail, Google Drive, Slack, Notion, Salesforce (Arcade 支援的應用)
  • Cursor, Chat (開發工具)
  • 其他提及
  • SpaceX (與 Fable 相關)
  • Universe of AI (頻道名稱)
  • World of AI (頻道名稱)
  • Universe-of-ai.beehiiv.com (新聞通訊訂閱連結)
  • X (社群平台)

操作流程整理

  1. 評估 Gemini 3.7 Flash 性能
  • 查看代碼質量基準測試(43.6分)。
  • 查看長視角軟體工程基準測試(65.3分)。
  • 查看網頁開發 ELO 分數(1588分)。
  • 查看自動化基準測試(30.4分)。
  • 查看長影片理解(85.4分)與長上下文性能(97分)。
  • 對比 Deepware Engineering 基準測試結果。
  1. 確認定價策略
  • 確認短期優惠價格(至 2026 年底):輸入 $0.75/百萬,輸出 $3.75/百萬。
  • 確認長期定價(2027 年 1 月起):輸入 $1.50/百萬,輸出 $7.50/百萬。
  1. 了解 OpenAI GPT 5.6 超快模式
  • 確認透過 Cerebras 晶片運行。
  • 確認性能指標:每秒 750 個輸出 token,速度為標準模式 14 倍。
  • 確認應用場景:即時語音支援、商務、金融研究等。
  1. 使用 Arcade 運行 AI Agent
  • 設定權限管理,避免共用單一登入資訊。
  • 整合工具(Gmail, Slack, Salesforce 等)。
  • 記錄 Agent 行為與對象。
  • 確保兼容性(支援任何模型、框架、雲端)。

值得注意的限制或風險

  1. Gemini 3.7 Flash 的定價與成本
  • 長期定價(2027 年起)可能較高,且在某些基準測試(如 Deepware Engineering)中成本約為競爭對手(GPT 5.6 Luna)的三倍。
  1. OpenAI GPT 5.6 超快模式的存取與定價
  • 尚未公布定價資訊。
  • 尚未確定有多少人能獲得存取權限。
  1. 模型性能落後
  • Gemini 3.7 Flash 在長視角軟體工程與 Deepware Engineering 基準測試中落後於 GPT 5.6 Terra 或 Luna。
  1. Arcade 的權限管理
  • 雖然解決了權限問題,但需確保 Agent 在特定帳戶中行動的準確性與安全性。

逐字稿辨識疑點

  • 模型名稱拼寫與版本混淆
  • 逐字稿中多次出現 "Soul""Sol""6 Soul""6.6 Sol",可能指代 OpenAI 的某個模型變體或聽寫錯誤,無法確定具體對應何種官方名稱。
  • "Sonet 5":可能為 "Sonnet 5" (Anthropic 模型) 的聽寫錯誤,但依規則標為疑點。
  • "Kim K3 Quinn 3.8 Max Cloud":此名稱極不尋常,可能是多個模型名稱的混合或聽寫錯誤(如 Kimi, Claude 等),無法查證。
  • "Opus 5 Gro 4.6 6":名稱結構混亂,可能為 "Opus 4" 或 "Grok 4" 等的混合聽寫。
  • "Fable 5""Fable 5.6.6 Sol":目前市場上無知名模型名為 "Fable",可能為聽寫錯誤或極小眾模型。
  • "Deepware Engineering":作為基準測試名稱出現,可能為 "Software Engineering" 的聽寫錯誤,但依規則標為疑點。
  • "GPT 5.6 6o":版本號與模型名稱混合,通常 OpenAI 模型為 GPT-4o 或 GPT-4.5 等,此處 "5.6 6o" 結構不明。
  • "GPT 6 Terra":在長上下文性能部分出現,前文多為 "GPT 5.6 Terra",此處 "GPT 6" 可能為口誤或聽寫錯誤。
  • 價格數值疑點
  • 逐字稿提到 "75 per 1 million input tokens",後文修正為 "0.75",但中間有 "75" 的表述,需確認是否為口誤。
  • 定價部分提到 "end of 2026" 和 "January of 2027",時間點與當前時間關係需查證,但依規則僅記錄原文。
  • 其他名詞
  • "Partit":在 Arcade 介紹中出現 "that's the partit actually handles for you",疑為 "part that" 的連讀或聽寫錯誤。
  • "Weight list":OpenAI 發布 "weight list for 5.6",通常為 "weights" (權重) 或 "waitlist" (等候名單),此處用詞模糊。

可延伸追問

  1. Gemini 3.7 Flash 的「智能工作馬」定位具體意味著哪些開發場景?
  2. OpenAI GPT 5.6 超快模式的定價策略何時會公布?
  3. Arcade 如何確保 AI Agent 在執行動作時的權限安全?
  4. Google 計劃何時發布 Gemini 4?
  5. Muse Spark 1.2 作為 Meta 的新玩家,其技術架構與優勢為何?

生字列表

生字讀音類型中文
flash/flæʃ/閃電(型號);快速
workhorse/ˈwɜːrk.hɔːrs/主力;耐用的工具或人
benchmark/ˈbentʃ.mɑːrk/基準測試;評量標準
footprint/ˈfʊt.prɪnt/足跡;影響力範圍
enterprise/ˈen.tər.praɪz/企業;大公司
automation/ˌɔː.təˈmeɪ.ʃən/自動化
frontier/frʌnˈtɪr/前沿的;領先的
excels/ɪkˈselz/擅長;傑出
context/ˈkɑːn.tekst/上下文;語境
disgraceful/dɪsˈɡreɪs.fəl/可恥的;丟臉的
workload/ˈwɜːrk.loʊd/工作負載
incident/ˈɪn.sɪ.dənt/事件;事故

生字解說

flash /flæʃ/

· B2

意思:閃電(型號);快速

解說:在 AI 模型命名中,「Flash」通常指代速度快、延遲低但可能犧牲部分精確度的模型版本,與「Pro」或「Ultra」相對。

影片原句
because they just dropped Gemini 3.7 Flash today.
因為他們今天剛發布了 Gemini 3.7 Flash。
延伸例句
The new flash model is ideal for real-time applications.
新的閃電型號非常適合即時應用。

workhorse /ˈwɜːrk.hɔːrs/

· C1

意思:主力;耐用的工具或人

解說:原意為「役馬」,引申為在團隊或系統中承擔主要、繁重且可靠工作的模型或人員。

影片原句
Gemini 3.7 Flash is their most intelligent workhorse model.
Gemini 3.7 Flash 是他們最智能的主力工作模型。
延伸例句
This laptop is a reliable workhorse for graphic designers.
這台筆記型電腦是平面設計師可靠的主力工具。

benchmark /ˈbentʃ.mɑːrk/

· C1

意思:基準測試;評量標準

解說:用於評估模型性能、速度或準確性的標準測試集或指標。

影片原句
if we take a look at the benchmarks ,
如果我們來看看基準測試結果,
延伸例句
The model scored highly on the coding benchmark.
該模型在編碼基準測試中獲得了高分。

footprint /ˈfʊt.prɪnt/

· B2

意思:足跡;影響力範圍

解說:在此語境下,指公司在特定市場(如企業領域)的影響力或存在範圍。

影片原句
just because of Google's footprint in the enterprise space
僅僅是因為 Google 在企業領域的影響力,
延伸例句
The company's carbon footprint has decreased significantly.
該公司的碳足跡已顯著減少。

enterprise /ˈen.tər.praɪz/

· B2

意思:企業;大公司

解說:指大型商業組織或機構,通常與「企業級」解決方案相關。

影片原句
Now , this model is probably going to be used in enterprises a lot
現在,這個模型可能會在企業中被大量使用,
延伸例句
Many enterprises are adopting cloud computing solutions.
許多企業正在採用雲端運算解決方案。

automation /ˌɔː.təˈmeɪ.ʃən/

· C1

意思:自動化

解說:使用技術自動執行任務的過程,減少人工干預。

影片原句
but we are seeing this model achieve on the automation bench
但我們看到這個模型在自動化基準測試中
延伸例句
Automation can significantly reduce operational costs.
自動化可以顯著降低運營成本。

frontier /frʌnˈtɪr/

· C1

意思:前沿的;領先的

解說:形容處於技術或知識發展最前線的事物,常指最先進的實驗室或模型。

影片原句
it's still cheap compared to the other frontier labs
所以,是的,與其他前沿實驗室相比,它仍然很便宜,
延伸例句
They are working on the frontier of quantum computing.
他們正在從事量子計算的前沿工作。

excels /ɪkˈselz/

· C1

意思:擅長;傑出

解說:動詞,指在某個領域表現優異或特別出色。

影片原句
which this model excels at 85 .
該模型在此方面表現出色,達到85。
延伸例句
She excels in mathematics and physics.
她在數學和物理方面表現傑出。

context /ˈkɑːn.tekst/

· B2

意思:上下文;語境

解說:在 AI 中,指模型能夠處理和理解的輸入資訊的長度或範圍。

影片原句
Then long context performance the model is at 97
接著是長上下文效能,該模型達到 97
延伸例句
The AI needs more context to understand the joke.
AI 需要更多的上下文才能理解這個笑話。

disgraceful /dɪsˈɡreɪs.fəl/

· C1

意思:可恥的;丟臉的

解說:形容行為或結果極不令人滿意,有損名譽。

影片原句
Oh , disgraceful .
「太丟人了。」
延伸例句
It is disgraceful that they ignored the safety warnings.
他們忽略了安全警告,這真是可恥。

workload /ˈwɜːrk.loʊd/

· C1

意思:工作負載

解說:指系統、人員或模型需要處理的工作量或任務強度。

影片原句
workloads like you know real time voice support commerce
的工作負載,比如你知道的即時語音支援商務,
延伸例句
The server can handle a heavy workload during peak hours.
伺服器在高峰時段可以處理沉重的負載。

incident /ˈɪn.sɪ.dənt/

· B2

意思:事件;事故

解說:指發生的特定事件,特別是在安全或IT領域中的突發狀況。

影片原句
financial research and security , incident response and reliability
、金融研究和安全、事件回應和可靠性
延伸例句
The incident response team acted quickly to contain the breach.
事件回應團隊迅速行動以控制洩露。

句型解說(含實例)

looks like [subject] [verb]

意思:看來……;似乎……

接續:looks like + 子句

解說:用於表達基於現有證據的推測或觀察,語氣較為口語化。

影片原句
So , looks like we just got flashed by Google DeepMind once again
所以,看來我們又被 Google DeepMind 閃了一下
實例
  1. It looks like the meeting has been cancelled.
    看來會議已經取消了。
  2. Looks like it's going to rain later.
    看來晚點會下雨。

once again

意思:再一次;再次

接續:once again + 動詞/形容詞

解說:強調某事重複發生,可能帶有驚訝、無奈或強調的語氣。

影片原句
which is once again a Flash model
這又是一個 Flash 模型,
實例
  1. He once again forgot to lock the door.
    他再一次忘記鎖門了。
  2. Once again, we need to review the budget.
    我們再次需要審查預算。

what they are saying is that [clause]

意思:他們所表示的是……;他們說……

接續:what + 主詞 + 動詞 + be + that + 子句

解說:用於引述或總結他人的觀點、聲明或官方說法,強調內容的重點。

影片原句
And what they are saying is that
他們所表示的是
實例
  1. What they are saying is that the project is delayed.
    他們所表示的是專案延遲了。
  2. What he is saying is that we need more time.
    他所表示的是我們需要更多時間。

the reason being [clause/reason]

意思:原因是……;理由在於……

接續:the reason being + 名詞片語/子句

解說:用於解釋前一句話的原因,語氣較為正式或強調因果關係。

影片原句
The reason being cuz this is not their strongest tier .
原因是因為這並非它們最頂級的等級。
實例
  1. The reason being the lack of funding.
    原因是缺乏資金。
  2. The reason being that he was too busy.
    原因是他太忙了。

when it comes to [noun/gerund]

意思:當涉及到……時;在……方面

接續:when it comes to + 名詞/動名詞

解說:用於將話題轉向特定領域或主題,表示在該特定範圍內的討論。

影片原句
when it comes to code quality ,
當涉及到程式碼品質時,
實例
  1. When it comes to cooking, she is an expert.
    當涉及到烹飪時,她是專家。
  2. When it comes to pricing, we are competitive.
    在定價方面,我們具有競爭力。

worth knowing about

意思:值得了解

接續:worth + 動名詞

解說:表達某事物具有價值或重要性,值得花時間去學習或關注。

影片原句
Arcade is worth knowing about .
Arcade 值得了解。
實例
  1. This book is worth reading.
    這本書值得閱讀。
  2. The new policy is worth discussing.
    新政策值得討論。

instead of [gerund/noun]

意思:而不是……;代替……

接續:instead of + 動名詞/名詞

解說:用於對比兩種選擇,強調採取了後者而非前者。

影片原句
instead of just talking about them
而不只是空談
實例
  1. We should walk instead of driving.
    我們應該走路而不是開車。
  2. He chose tea instead of coffee.
    他選擇了茶而不是咖啡。

not only ... but (also) ...

意思:不僅……而且……

接續:not only + 子句/片語 + but (also) + 子句/片語

解說:用於連接兩個並列的成分,強調兩者都成立,後者往往更重要或更具補充性。

影片原句
So you're not just giving an AI a list of tools and hoping it works . You're giving it a place where it can safely take real actions
因此,你不僅僅是給 AI 一組工具清單並希望它能運作。你賦予它一個可以在真實應用程式中安全採取實際行動的空間。
實例
  1. She is not only smart but also hardworking.
    她不僅聰明而且勤奮。
  2. The app is not only free but also secure.
    該應用程式不僅免費而且安全。