0:00.000–0:04.410
So , looks like we just got flashed by Google DeepMind once again
0:04.420–0:06.412
because they just dropped Gemini 3.7 Flash today.
0:06.422–0:06.952
Yes, their newest model,
0:06.952–0:07.705
which is once again a Flash model,
0:07.715–0:09.139
Yes , their newest model ,
0:09.139–0:11.162
which is once again a Flash model ,
0:11.172–0:12.046
is here .
0:12.056–0:15.473
And this is only after 3 weeks since Gemini 3.6 Flash.
0:15.483–0:16.544
6 Flash .
0:16.554–0:17.414
And what they are saying is that
0:17.414–0:19.034
Gemini 3.7 Flash is their most intelligent workhorse model.
0:19.044–0:22.649
7 Flash is their most intelligent workhorse model .
0:22.659–0:23.489
And this is coming 3 weeks after 3.6 Flash
0:23.489–0:24.772
because of developer feedback and algorithmic innovations.
0:24.782–0:29.103
6 six flash because of developer feedback and algorithmic innovations
0:29.103–0:35.920
.
0:35.920–0:36.000
And that is why they're saying
0:36.000–0:36.080
the Gemini 3.7 Flash model is,
0:36.080–0:36.160
you know, out with so much improvement cuz
0:36.160–0:36.240
7 Flash model is , you know , out with so much improvement cuz
0:36.240–0:37.105
if we take a look at the benchmarks ,
0:37.105–0:38.350
this model is way better
0:38.360–0:39.552
than the Gemini 3.6 Flash.
0:39.562–0:40.622
6 Flash .
0:40.632–0:42.234
But I think this is also because
0:42.234–0:43.836
we know internally that they're
0:43.846–0:44.926
planning on cancelling Gemini 3.5 Pro completely
0:44.926–0:45.717
and working on Gemini 4 at the moment.
0:45.727–0:49.647
5 Pro completely and working on Gemini 4 at the moment .
0:49.657–0:52.595
So maybe the improvements they made with the Pro model that they're
0:52.605–0:54.432
not going to be dropping anymore .
0:54.442–0:58.360
Maybe they repackaged it as a flash model because this model is
0:58.370–0:59.537
better than 3 .
0:59.547–1:00.449
the 3.5 Pro,
1:00.449–1:03.495
it's quite likely that is the case.
1:03.505–1:04.083
the 3 .
1:04.093–1:04.462
5 Pro ,
1:04.462–1:06.949
it's quite likely that is the case .
1:06.959–1:07.882
We'll just ignore that
1:07.882–1:08.958
and put that into the side .
1:08.968–1:11.450
We don't know when we're getting a new Pro model from the Google
1:11.460–1:12.685
we will see that yes,
1:12.685–1:15.058
it is stronger than the 3.6 Flash model.
1:15.068–1:16.407
will see that yes ,
1:16.407–1:18.319
it is stronger than the 3 .
1:18.329–1:19.628
6 Flash model .
1:19.638–1:20.186
For example ,
1:20.186–1:21.502
when it comes to code quality ,
1:21.502–1:22.654
production code quality
1:22.664–1:23.896
the model is producing 43.6
1:23.896–1:25.128
versus the flash model 34.4.
1:25.138–1:28.110
6 versus the flash model 34 .
1:28.120–1:29.213
4 .
1:29.223–1:30.678
And then Sonet 5 is at 42.7
1:30.678–1:32.352
and the Terra model is at 41.3.
1:32.362–1:35.570
7 and the Terra model is at 41 .
1:35.580–1:36.362
3 .
1:36.372–1:36.514
So ,
1:36.514–1:37.788
one thing to remember ,
1:37.788–1:39.346
since this is a flash model ,
1:39.356–1:41.536
you're not going to see them compare this to Opus 5
1:41.536–1:43.572
or any of the other stronger models .
1:43.582–1:46.775
The reason being cuz this is not their strongest tier .
1:46.785–1:47.650
and the 3 .
1:47.660–1:48.331
1 Pro model
1:48.331–1:51.390
when they finally decide to upgrade it to Gemini 4
1:51.400–1:51.730
Pro ,
1:51.730–1:55.354
then we'll see it being compared to Opus 5
1:55.354–1:56.013
or GPT 5 .
1:56.023–1:56.842
6 Soul .
1:56.852–1:57.427
But anyways ,
1:57.427–1:59.728
we can see that this model is an improvement from
1:59.738–2:00.209
3 .
2:00.219–2:00.757
6 Flash
2:00.757–2:04.699
and I would say that's basically the biggest result or
2:04.709–2:07.182
the biggest update that we saw with the new model
2:07.182–2:07.986
because it does
2:07.996–2:09.982
not really like change up things a lot
2:09.982–2:11.391
is still like not becoming
2:11.401–2:12.565
the number one model
2:12.565–2:14.208
or the number one flash model .
2:14.218–2:15.916
Even on some benchmarks ,
2:15.916–2:18.547
Deepseek version 4 flash is actually
2:18.557–2:19.106
cheaper
2:19.106–2:21.540
and more intelligent than this model .
2:21.550–2:24.249
But let's just take a look at the benchmarks that they have published
2:24.249–2:30.087
.
2:30.087–2:30.167
On Long Horizon software engineering ,
2:30.167–2:30.247
this model is a little bit
2:30.247–2:30.327
behind GPT 5 .
2:30.327–2:30.407
6 Terra ,
2:30.407–2:32.039
which sits at 69 .
2:32.049–2:32.243
6
2:32.243–2:33.991
and then 65 .
2:34.001–2:35.751
3 is the Flash model .
2:35.761–2:37.686
The older Flash model 3 .
2:37.696–2:39.290
6 sits at 48 .
2:39.300–2:40.230
6 .
2:40.240–2:41.987
Sonic 5 sits at 53 .
2:41.997–2:42.941
8 8 .
2:42.951–2:46.380
And the new player that is finally being included on benchmarks
2:46.380–2:46.984
,
2:46.984–2:49.818
which even the Google DeepMind team is considering with their
2:49.828–2:50.388
launch ,
2:50.388–2:51.508
is Muse Spark 1 .
2:51.518–2:51.633
2 ,
2:51.633–2:53.699
which is sitting at 54 .
2:53.709–2:54.544
9 .
2:54.554–2:54.705
So ,
2:54.705–2:57.040
welcome Meta to the benchmark charts
2:57.040–2:58.698
because now we're starting
2:58.708–3:01.254
to see it appear on more and more benchmarks .
3:01.264–3:03.571
And if we also take a look at web development ,
3:03.581–3:07.699
this model is getting a ELO score of 1588 versus their old model
3:07.709–3:08.663
is 1538 .
3:08.673–3:09.773
So , not a crazy difference ,
3:09.773–3:11.345
but compared to everything else out
3:11.355–3:12.726
there , this is number one .
3:12.736–3:14.291
When I say everything else out there ,
3:14.291–3:14.757
once again ,
3:14.767–3:16.671
compared to all the mid tier models ,
3:16.671–3:18.116
this is beating all of them
3:18.126–3:19.891
when it comes to web development .
3:19.901–3:21.486
Now , this model is probably going to be used in enterprises a lot
3:21.486–3:22.979
just because of Google's footprint in the enterprise space
3:22.989–3:25.758
lot just because of Google's footprint in the enterprise space
3:25.758–3:32.640
,
3:32.640–3:32.720
but we are seeing this model achieve on the automation bench
3:32.720–3:32.800
30 .
3:32.800–3:32.880
4 .
3:32.880–3:32.960
And GPT 5 .
3:32.960–3:34.010
6 Terra sits at 23 .
3:34.020–3:34.890
6 .
3:34.900–3:38.560
So , yes , this model is stronger than the other Flash models out
3:38.570–3:41.588
there , but as I said , they haven't included Deepseek version
3:41.598–3:44.142
for Flash because if they do , in some areas ,
3:44.152–3:46.458
that model is actually quite better than the 3 .
3:46.468–3:48.055
7 Flash model .
3:48.065–3:51.409
Before we continue , if you're building AI agents or just messing
3:51.419–3:54.285
around with them , Arcade is worth knowing about .
3:54.295–3:57.318
It's the runtime that lets your agent actually do things instead
3:57.328–4:00.433
of just talking about them because that's the gap right now .
4:00.443–4:02.109
The models are smart enough .
4:02.119–4:05.230
Your agent can figure out exactly what needs to happen in your
4:05.240–4:07.309
email , your Slack , your CRM .
4:07.319–4:09.389
It just can't go in and do it .
4:09.399–4:12.431
And the reason isn't intelligence , it's permissions .
4:12.441–4:15.390
Something has to prove the agent is allowed to act on behalf of
4:15.400–4:18.350
a specific person in a specific account .
4:18.360–4:22.032
That's the messy part everyone runs into , and it's the partit
4:22.042–4:23.393
actually handles for you .
4:23.403–4:26.835
So instead of your agent using one shared login for everybody ,
4:26.845–4:30.516
it acts as whoever is actually signed in with exactly the access
4:30.526–4:31.878
that person has .
4:31.888–4:34.760
If they can't see something , the agent can't either .
4:34.770–4:37.482
And you never have to touch any of that setup yourself .
4:37.492–4:38.843
Then there's the tools .
4:38.853–4:41.725
Arcade has thousands of them already built for Gmail ,
4:41.735–4:45.728
Google Drive , Slack , Notion , Salesforce , most of the apps people
4:45.738–4:49.090
already work in , and they are built specifically for AI to use
4:49.090–4:52.440
.
4:52.440–4:52.520
So the agent gets it right the first time instead of guessing and
4:52.520–4:53.733
failing and retrying .
4:53.743–4:56.855
It also keeps a record of everything , what the agent did ,
4:56.865–5:00.138
for who and where , which matters a lot the moment other people
5:00.148–5:01.739
start using the thing you built .
5:01.749–5:04.140
And it works with whatever you're already using .
5:04.150–5:07.182
Any model , any framework , cloud , cursor , chat ,
5:07.192–5:08.943
GPT , doesn't matter .
5:08.953–5:11.986
So you're not just giving an AI a list of tools and hoping it works
5:11.986–5:16.206
.
5:16.206–5:16.286
You're giving it a place where it can safely take real actions
5:16.286–5:16.469
in real apps .
5:16.479–5:19.030
It's free to start and the link is in the description .
5:19.040–5:22.066
Thank you once again for Arcade for sponsoring today's video .
5:22.076–5:24.378
Now , let's get back into the video .
5:24.388–5:27.178
Now , one thing to note is that the Gemini 3 .
5:27.188–5:31.007
7 Flash model through the end of this year , so end of 2026 ,
5:31.017–5:34.283
they have a cheap pricing model that they're placing on the model
5:34.283–5:38.240
.
5:38.240–5:38.786
75 per 1 million input tokens and 3 .
5:38.796–5:40.822
75 for 1 million output tokens .
5:40.832–5:44.266
So , it's a competitive price , but this is only for the next 6
5:44.276–5:44.899
months .
5:44.909–5:48.305
Because after those six months are done , the model's pricing is
5:48.315–5:50.528
actually , you know , a little bit more expensive .
5:50.538–5:52.922
And now they show it at the bottom over here ,
5:52.932–5:56.990
you can see that after starting January of 2027 ,
5:57.000–5:58.763
it will become 1 .
5:58.773–6:00.728
50 per input and 7 .
6:00.738–6:01.947
50 per output .
6:01.957–6:05.605
So yeah , it's still cheap compared to the other frontier labs
6:05.615–6:08.896
, but it's not as cheap as , for example , MU Spark when the pricing
6:08.906–6:12.352
is updated or even the Deepsee version for Flash .
6:12.362–6:13.955
But across these benchmarks ,
6:13.955–6:15.879
we can see that this model is better
6:15.889–6:18.283
in many areas that they highlighted at the top .
6:18.293–6:21.768
But then they also have some other areas like long video understanding
6:21.778–6:24.452
, which this model excels at 85 .
6:24.462–6:26.501
4 versus 78 .
6:26.511–6:28.298
9 for the Terra model .
6:28.308–6:31.897
And then the old model was also pretty good at that ,
6:31.907–6:32.176
84 .
6:32.186–6:32.465
2 .
6:32.475–6:36.288
Then long context performance the model is at 97
6:36.288–6:37.432
and this model
6:37.442–6:41.273
the GPT 6 Terra one it sits at 93 .
6:41.283–6:42.080
5.
6:42.090–6:44.496
So yeah this model in summary it is better
6:44.496–6:45.840
so it's not all negative
6:45.850–6:48.424
but it's not all like you know that positive
6:48.424–6:49.600
where you are super
6:49.610–6:50.921
excited for Google Deep Mind
6:50.921–6:52.450
because as I said they're probably
6:52.460–6:55.689
still holding off their biggest release for Gemini 4 lineup .
6:55.699–6:58.635
Now , one thing people might have missed in their charts because
6:58.645–7:02.217
these charts sometimes are so messy to read and understand ,
7:02.227–7:03.412
3.
7:03.422–7:05.822
7 flash is worse than GPT 5.
7:05.832–7:06.188
6 Luna ,
7:06.188–7:07.826
which is the model over here ,
7:07.826–7:09.464
which is achieving a higher
7:09.474–7:13.841
score on this benchmark deepware engineering for about three times
7:13.851–7:14.809
the cost .
7:14.819–7:16.992
So , yeah , this is kind of interesting because yeah ,
7:17.002–7:18.340
the cost for Gemini 3.
7:18.350–7:21.276
7 Flash is a little bit more than what it looks like .
7:21.286–7:23.044
Now , some people are a little bit upset
7:23.044–7:23.781
and they're like ,
7:23.791–7:24.994
Oh , disgraceful .
7:25.004–7:28.734
Google left out soul , opus , and fable because Google is incredibly
7:28.744–7:29.447
behind .
7:29.457–7:31.115
But I think one thing we got to remember , guys ,
7:31.125–7:33.336
is that this model is a flash model .
7:33.346–7:35.242
It's not trying to be a pro model
7:35.242–7:36.986
or it's not trying to compete
7:36.996–7:40.458
That is probably going to be Gemini 4 .
7:40.468–7:42.651
That is probably going to be Gemini 4 .
7:42.661–7:44.077
So when Gemini 4 comes out ,
7:44.077–7:46.303
then I think it's okay for us to criticize
7:46.313–7:50.156
them if they don't include Soul , Opus , and Fable in their benchmark
7:50.166–7:51.401
charts because for now ,
7:51.401–7:53.350
I think what they have done is pretty
7:53.360–7:54.069
accurate .
7:54.079–7:55.750
One lab that I would have liked to see
7:55.750–7:56.864
or one model for example
7:56.874–7:59.478
I would have liked to see on that chart would be Deepseek version
7:59.488–7:59.793
4 Flash
7:59.793–8:02.134
because that would kind of spoil their release because
8:02.144–8:05.521
that model is way cheaper compared to the Gemini 3.
8:05.531–8:07.058
7 flash model lineup .
8:07.068–8:10.807
Now this model jumped from number 19 to 8 on the web development
8:10.817–8:11.095
area
8:11.095–8:12.555
and we see it over here now
8:12.555–8:14.224
and couple of models that are
8:14.234–8:17.751
ahead of it are Opus 5 obviously Kim K3 Quinn 3 .
8:17.761–8:20.384
8 Max Cloud Opus 5 Gro 4 .
8:20.394–8:22.409
6 6 , which is a model that came yesterday ,
8:22.409–8:23.353
which was a big win
8:23.363–8:25.900
for SpaceX, Fable 5, and 5.6.
8:25.910–8:26.352
6.
8:26.362–8:27.002
Sol.
8:27.012–8:29.966
So , yeah , this model is trying to compete in the web development
8:29.976–8:33.094
, but it's still behind all of these models , which is ,
8:33.104–8:35.305
you know , expected cuz it's a flash model .
8:35.315–8:37.111
It's not really a pro model .
8:37.121–8:40.296
Today, OpenAI has also launched a weight list for 5.6 so ultra fast mode.
8:40.306–8:42.036
6 so ultra fast mode .
8:42.046–8:45.623
access to now through Cerebras and GPT 5.
8:45.633–8:48.490
6 6o with this chip kind of running it is able to achieve an ultra
8:48.500–8:52.774
6 6o with this chip kind of running it is able to achieve an ultra
8:52.784–8:57.861
fast mode that generates up to 750 output tokens per second which
8:57.871–9:01.494
is about 14 times faster than the standard mode .
9:01.504–9:04.812
So we're getting a really fast version of GPT 5 .
9:04.822–9:09.505
6 so now this is supposed to be used for live or near production
9:09.515–9:13.292
workloads like you know real time voice support commerce what this
9:13.302–9:14.581
6 six sol to do is kind of be really fast in critical situations
9:14.591–9:18.277
6 six soul to do is kind of be really fast in critical situations
9:18.287–9:22.135
when people might be interacting with the AI agent like financial
9:22.145–9:26.395
research security response support I think is going to be a big
9:26.405–9:31.508
area where this new ultrafast mode will be kind of implemented business
9:31.518–9:35.098
and developer agents maybe yeah but I think like support or near
9:35.108–9:38.527
production workloads like real time voice I see this model really
9:38.537–9:42.378
excelling at that and obviously this is still a weightless mode
9:42.388–9:45.323
we don't know how many people are going to get access to this ,
9:45.333–9:47.791
how successful it is or what the pricing is .
9:47.801–9:50.497
I don't know if the pricing has changed because there's no information
9:50.507–9:53.283
in their actual , you know , blog post .
9:53.293–9:56.396
But as I mentioned , couple of areas where they mentioned that
9:56.406–9:58.233
this is going to be really important .
9:58.243–10:01.669
Customer support and voice , commerce , live research and experimentation
10:01.679–10:06.182
, financial research and security , incident response and reliability
10:06.182–10:08.288
.
10:08.288–10:09.514
But yeah , I think this partnership is going to be important for
10:09.524–10:11.044
OpenAI going forward .
10:11.054–10:12.978
But as I said , 14 times the speed .
10:12.988–10:14.589
What does that mean for cost ?
10:14.599–10:16.120
We don't know yet .
10:16.130–10:17.812
But that's it for today's video .
10:17.822–10:20.068
Make sure you guys are subscribed to the channel .
10:20.078–10:21.852
Follow our new newsletter as well at universe-of-ai.beehiiv.com
10:21.852–10:23.626
as well as subscribe to the main channel World of AI and support
10:23.636–10:24.014
behive .
10:24.024–10:26.911
com as well as subscribe to the main channel World of AI and support
10:26.921–10:29.396
us on X by following the Universe of AI as well .
10:29.406–10:32.401
Until then , I'll see you guys in the next
0:00.000–0:04.410
So , looks like we just got flashed by Google DeepMind once again
所以,看來我們又被 Google DeepMind 閃了一下
0:04.420–0:06.412
because they just dropped Gemini 3.7 Flash today.
因為他們今天剛發布了 Gemini 3.7 Flash。
0:06.422–0:06.952
Yes, their newest model,
是的,他們最新的模型,
0:06.952–0:07.705
which is once again a Flash model,
這又是一個 Flash 模型,
0:07.715–0:09.139
Yes , their newest model ,
是的,他們最新的模型,
0:09.139–0:11.162
which is once again a Flash model ,
這又是一個 Flash 模型,
0:11.172–0:12.046
is here .
已經問世了。
0:12.056–0:15.473
And this is only after 3 weeks since Gemini 3.6 Flash.
這距離 Gemini 3.6 Flash 發布才過了 3 週。
0:15.483–0:16.544
6 Flash .
[未翻譯]
0:16.554–0:17.414
And what they are saying is that
他們所表示的是
0:17.414–0:19.034
Gemini 3.7 Flash is their most intelligent workhorse model.
Gemini 3.7 Flash 是他們最智能的主力工作模型。
0:19.044–0:22.649
7 Flash is their most intelligent workhorse model .
7 Flash 是他們最智能的主力工作模型。
0:22.659–0:23.489
And this is coming 3 weeks after 3.6 Flash
這是在 3.6 Flash 發布 3 週後推出的
0:23.489–0:24.772
because of developer feedback and algorithmic innovations.
因為開發者的反饋和算法創新。
0:24.782–0:29.103
6 six flash because of developer feedback and algorithmic innovations
6 six flash 因為開發者的反饋和算法創新
0:29.103–0:35.920
.
。
0:35.920–0:36.000
And that is why they're saying
這就是為什麼他們說
0:36.000–0:36.080
the Gemini 3.7 Flash model is,
Gemini 3.7 Flash 模型,
0:36.080–0:36.160
you know, out with so much improvement cuz
你知道,有如此多的改進,因為
0:36.160–0:36.240
7 Flash model is , you know , out with so much improvement cuz
7 Flash 模型,你知道,有如此多的改進,因為
0:36.240–0:37.105
if we take a look at the benchmarks ,
如果我們來看看基準測試結果,
0:37.105–0:38.350
this model is way better
這個模型的表現遠比
0:38.360–0:39.552
than the Gemini 3.6 Flash.
Gemini 3.6 Flash 好。
0:39.562–0:40.622
6 Flash .
6閃電
0:40.632–0:42.234
But I think this is also because
但我認為這也是因為
0:42.234–0:43.836
we know internally that they're
我們內部知道他們
0:43.846–0:44.926
planning on cancelling Gemini 3.5 Pro completely
正計畫完全取消 Gemini 3.5 Pro
0:44.926–0:45.717
and working on Gemini 4 at the moment.
,並目前專注於開發 Gemini 4。
0:45.727–0:49.647
5 Pro completely and working on Gemini 4 at the moment .
完全取消 3.5 Pro,並目前專注於開發 Gemini 4。
0:49.657–0:52.595
So maybe the improvements they made with the Pro model that they're
所以也許他們在 Pro 模型上所做的改進,
0:52.605–0:54.432
not going to be dropping anymore .
他們將不再繼續推出。
0:54.442–0:58.360
Maybe they repackaged it as a flash model because this model is
也許他們將這些改進重新包裝成 Flash 模型,因為這個模型比
0:58.370–0:59.537
better than 3 .
3.5
0:59.547–1:00.449
the 3.5 Pro,
3.5 Pro 更好,
1:00.449–1:03.495
it's quite likely that is the case.
這很有可能就是事實。
1:03.505–1:04.083
the 3 .
3.5
1:04.093–1:04.462
5 Pro ,
3.5 Pro,
1:04.462–1:06.949
it's quite likely that is the case .
這很有可能就是事實。
1:06.959–1:07.882
We'll just ignore that
我們就暫且忽略這一點
1:07.882–1:08.958
and put that into the side .
,把它放在一旁吧。
1:08.968–1:11.450
We don't know when we're getting a new Pro model from the Google
我們不知道何時能從 Google 獲得新的 Pro 型號
1:11.460–1:12.685
we will see that yes,
我們會看到,是的,
1:12.685–1:15.058
it is stronger than the 3.6 Flash model.
它比 3.6 Flash 模型更強大。
1:15.068–1:16.407
will see that yes ,
我們會看到,是的,
1:16.407–1:18.319
it is stronger than the 3 .
它比 3 更強大。
1:18.329–1:19.628
6 Flash model .
6 Flash 模型更強大。
1:19.638–1:20.186
For example ,
例如,
1:20.186–1:21.502
when it comes to code quality ,
當涉及到程式碼品質時,
1:21.502–1:22.654
production code quality
生產環境的程式碼品質
1:22.664–1:23.896
the model is producing 43.6
該模型產生了 43.6
1:23.896–1:25.128
versus the flash model 34.4.
而 Flash 模型為 34.4。
1:25.138–1:28.110
6 versus the flash model 34 .
6 而 Flash 模型為 34.
1:28.120–1:29.213
4 .
[未翻譯]
1:29.223–1:30.678
And then Sonet 5 is at 42.7
然後 Sonet 5 為 42.7
1:30.678–1:32.352
and the Terra model is at 41.3.
而 Terra 模型為 41.3。
1:32.362–1:35.570
7 and the Terra model is at 41 .
7 而 Terra 模型為 41.
1:35.580–1:36.362
3 .
[未翻譯]
1:36.372–1:36.514
So ,
所以,
1:36.514–1:37.788
one thing to remember ,
有一件事要記住,
1:37.788–1:39.346
since this is a flash model ,
因為這是一個 Flash 模型,
1:39.356–1:41.536
you're not going to see them compare this to Opus 5
你不會看到有人將它與 Opus 5 進行比較
1:41.536–1:43.572
or any of the other stronger models .
或任何其他更強大的模型。
1:43.582–1:46.775
The reason being cuz this is not their strongest tier .
原因是因為這並非它們最頂級的等級。
1:46.785–1:47.650
and the 3 .
以及 3.
1:47.660–1:48.331
1 Pro model
1 Pro 模型
1:48.331–1:51.390
when they finally decide to upgrade it to Gemini 4
當他們最終決定將其升級為 Gemini 4
1:51.400–1:51.730
Pro ,
[未翻譯]
1:51.730–1:55.354
then we'll see it being compared to Opus 5
我們才會看到它與 Opus 5
1:55.354–1:56.013
or GPT 5 .
或 GPT 5 進行比較。
1:56.023–1:56.842
6 Soul .
第六,靈魂。
1:56.852–1:57.427
But anyways ,
不過,
1:57.427–1:59.728
we can see that this model is an improvement from
我們可以看到這個模型是
1:59.738–2:00.209
3 .
[未翻譯]
2:00.219–2:00.757
6 Flash
6 Flash 的改進版
2:00.757–2:04.699
and I would say that's basically the biggest result or
我認為這基本上就是我們在新模型中看到的最重大成果或
2:04.709–2:07.182
the biggest update that we saw with the new model
我們在新模型中看到的最重大更新
2:07.182–2:07.986
because it does
因為它確實
2:07.996–2:09.982
not really like change up things a lot
不太喜歡大幅改變事物
2:09.982–2:11.391
is still like not becoming
仍然沒有成為
2:11.401–2:12.565
the number one model
排名第一的模型
2:12.565–2:14.208
or the number one flash model .
或是排名第一的 Flash 模型。
2:14.218–2:15.916
Even on some benchmarks ,
即使在某些基準測試中,
2:15.916–2:18.547
Deepseek version 4 flash is actually
Deepseek 的 v4 Flash 版本實際上
2:18.557–2:19.106
cheaper
更便宜
2:19.106–2:21.540
and more intelligent than this model .
且比這個模型更聰明。
2:21.550–2:24.249
But let's just take a look at the benchmarks that they have published
但讓我們先來看看他們發布的基準測試結果
2:24.249–2:30.087
.
。
2:30.087–2:30.167
On Long Horizon software engineering ,
在長期軟體工程方面,
2:30.167–2:30.247
this model is a little bit
這個模型稍微
2:30.247–2:30.327
behind GPT 5 .
落後於 GPT 5。
2:30.327–2:30.407
6 Terra ,
[未翻譯]
2:30.407–2:32.039
which sits at 69 .
得分為 69。
2:32.049–2:32.243
6
[未翻譯]
2:32.243–2:33.991
and then 65 .
然後是 65。
2:34.001–2:35.751
3 is the Flash model .
3 是 Flash 模型。
2:35.761–2:37.686
The older Flash model 3 .
較舊的 Flash 3。
2:37.696–2:39.290
6 sits at 48 .
6 得分為 48。
2:39.300–2:40.230
6 .
[未翻譯]
2:40.240–2:41.987
Sonic 5 sits at 53 .
5 Sonic 得分為 53。
2:41.997–2:42.941
8 8 .
[未翻譯]
2:42.951–2:46.380
And the new player that is finally being included on benchmarks
而終於被納入基準測試的新玩家
2:46.380–2:46.984
,
而終於被納入基準測試的新玩家,連 Google DeepMind 團隊在發布時也在考慮,正是 Muse Spark 1。
2:46.984–2:49.818
which even the Google DeepMind team is considering with their
就連 Google DeepMind 團隊在他們的
2:49.828–2:50.388
launch ,
發布
2:50.388–2:51.508
is Muse Spark 1 .
是 Muse Spark 1
2:51.518–2:51.633
2 ,
[未翻譯]
2:51.633–2:53.699
which is sitting at 54 .
目前得分為54分
2:53.709–2:54.544
9 .
第九名
2:54.554–2:54.705
So ,
所以,
2:54.705–2:57.040
welcome Meta to the benchmark charts
歡迎 Meta 進入基準測試排行榜
2:57.040–2:58.698
because now we're starting
因為我們現在開始
2:58.708–3:01.254
to see it appear on more and more benchmarks .
看到它出現在越來越多的基準測試中。
3:01.264–3:03.571
And if we also take a look at web development ,
如果我們也來看看網頁開發,
3:03.581–3:07.699
this model is getting a ELO score of 1588 versus their old model
該模型的 ELO 得分為 1588,而他們舊模型的
3:07.709–3:08.663
is 1538 .
得分為 1538。
3:08.673–3:09.773
So , not a crazy difference ,
所以,差異並不算大,
3:09.773–3:11.345
but compared to everything else out
但與其他所有產品相比
3:11.355–3:12.726
there , this is number one .
這是排名第一的。
3:12.736–3:14.291
When I say everything else out there ,
當我說與其他所有產品相比時,
3:14.291–3:14.757
once again ,
再一次,
3:14.767–3:16.671
compared to all the mid tier models ,
與所有中階模型相比,
3:16.671–3:18.116
this is beating all of them
在網頁開發方面,
3:18.126–3:19.891
when it comes to web development .
它擊敗了所有這些模型。
3:19.901–3:21.486
Now , this model is probably going to be used in enterprises a lot
現在,這個模型可能會在企業中被大量使用,
3:21.486–3:22.979
just because of Google's footprint in the enterprise space
僅僅是因為 Google 在企業領域的影響力,
3:22.989–3:25.758
lot just because of Google's footprint in the enterprise space
僅僅是因為 Google 在企業領域的影響力,
3:25.758–3:32.640
,
,
3:32.640–3:32.720
but we are seeing this model achieve on the automation bench
但我們看到這個模型在自動化基準測試中
3:32.720–3:32.800
30 .
達到了 30。
3:32.800–3:32.880
4 .
而 GPT 5 的得分為 30
3:32.880–3:32.960
And GPT 5 .
以及 GPT 5。
3:32.960–3:34.010
6 Terra sits at 23 .
6 Terra 位於 23。
3:34.020–3:34.890
6 .
[未翻譯]
3:34.900–3:38.560
So , yes , this model is stronger than the other Flash models out
所以,是的,這個模型比其他現有的 Flash 模型更強大,
3:38.570–3:41.588
there , but as I said , they haven't included Deepseek version
但正如我所說,他們沒有將 Deepseek 版本
3:41.598–3:44.142
for Flash because if they do , in some areas ,
納入 Flash,因為如果他們這樣做,在某些領域,
3:44.152–3:46.458
that model is actually quite better than the 3 .
該模型實際上比 3 好得多。
3:46.468–3:48.055
7 Flash model .
7 Flash 模型更好。
3:48.065–3:51.409
Before we continue , if you're building AI agents or just messing
在繼續之前,如果你正在構建 AI 代理或只是隨意
3:51.419–3:54.285
around with them , Arcade is worth knowing about .
嘗試它們,Arcade 值得了解。
3:54.295–3:57.318
It's the runtime that lets your agent actually do things instead
它是讓你的 AI 代理真正執行任務,而不只是空談的運行環境
3:57.328–4:00.433
of just talking about them because that's the gap right now .
因為這正是目前的缺口所在。
4:00.443–4:02.109
The models are smart enough .
這些模型已經足夠聰明。
4:02.119–4:05.230
Your agent can figure out exactly what needs to happen in your
你的 AI 代理能精確判斷在你的
4:05.240–4:07.309
email , your Slack , your CRM .
電子郵件、Slack 或 CRM 中需要做什麼。
4:07.319–4:09.389
It just can't go in and do it .
但它無法直接進去執行。
4:09.399–4:12.431
And the reason isn't intelligence , it's permissions .
原因不在於智力,而在於權限。
4:12.441–4:15.390
Something has to prove the agent is allowed to act on behalf of
必須有某個機制證明該代理獲授權代表
4:15.400–4:18.350
a specific person in a specific account .
特定帳戶中的特定人員行事。
4:18.360–4:22.032
That's the messy part everyone runs into , and it's the partit
這就是每個人都會遇到的棘手部分,而
4:22.042–4:23.393
actually handles for you .
Arcade 會為你處理這部分。
4:23.403–4:26.835
So instead of your agent using one shared login for everybody ,
因此,你的 AI 代理不會使用所有人共用的單一登入資訊,
4:26.845–4:30.516
it acts as whoever is actually signed in with exactly the access
而是以實際登入者的身份行事,並擁有該人員
4:30.526–4:31.878
that person has .
所具備的權限。
4:31.888–4:34.760
If they can't see something , the agent can't either .
如果他們看不到某些內容,AI 代理也看不到。
4:34.770–4:37.482
And you never have to touch any of that setup yourself .
而且你完全不需要自己處理任何設定。
4:37.492–4:38.843
Then there's the tools .
接下來是工具部分。
4:38.853–4:41.725
Arcade has thousands of them already built for Gmail ,
Arcade 已經內建了數千種工具,適用於 Gmail、
4:41.735–4:45.728
Google Drive , Slack , Notion , Salesforce , most of the apps people
Google Drive、Slack、Notion、Salesforce,以及人們
4:45.738–4:49.090
already work in , and they are built specifically for AI to use
日常使用的多數應用程式,這些工具都是專門為 AI 使用而設計的
4:49.090–4:52.440
.
。
4:52.440–4:52.520
So the agent gets it right the first time instead of guessing and
因此,代理程式能一次就正確執行,而不是靠猜測、失敗然後重試。
4:52.520–4:53.733
failing and retrying .
失敗並重試。
4:53.743–4:56.855
It also keeps a record of everything , what the agent did ,
它還會記錄所有內容,包括代理程式做了什麼、
4:56.865–5:00.138
for who and where , which matters a lot the moment other people
以及對象和地點,這在其他人介入的瞬間顯得至關重要
5:00.148–5:01.739
start using the thing you built .
開始使用你建構的東西時,顯得非常重要。
5:01.749–5:04.140
And it works with whatever you're already using .
而且它與你目前正在使用的任何工具都能相容。
5:04.150–5:07.182
Any model , any framework , cloud , cursor , chat ,
任何模型、任何框架、雲端、Cursor、聊天、
5:07.192–5:08.943
GPT , doesn't matter .
GPT,都不重要。
5:08.953–5:11.986
So you're not just giving an AI a list of tools and hoping it works
因此,你不僅僅是給 AI 一組工具清單並希望它能運作
5:11.986–5:16.206
.
。
5:16.206–5:16.286
You're giving it a place where it can safely take real actions
你賦予它一個可以在真實應用程式中安全採取實際行動的空間。
5:16.286–5:16.469
in real apps .
在真實應用程式中。
5:16.479–5:19.030
It's free to start and the link is in the description .
免費開始使用,連結在描述中。
5:19.040–5:22.066
Thank you once again for Arcade for sponsoring today's video .
再次感謝 Arcade 贊助今天的影片。
5:22.076–5:24.378
Now , let's get back into the video .
現在,讓我們回到影片中。
5:24.388–5:27.178
Now , one thing to note is that the Gemini 3 .
現在,有一點需要注意,那就是 Gemini 3.
5:27.188–5:31.007
7 Flash model through the end of this year , so end of 2026 ,
7 Flash 模型直到今年年底,也就是 2026 年底,
5:31.017–5:34.283
they have a cheap pricing model that they're placing on the model
他們為該模型提供了一個廉價的定價方案
5:34.283–5:38.240
.
。
5:38.240–5:38.786
75 per 1 million input tokens and 3 .
每百萬個輸入記號75。
5:38.796–5:40.822
75 for 1 million output tokens .
每百萬個輸出記號3.75。
5:40.832–5:44.266
So , it's a competitive price , but this is only for the next 6
所以,這是一個有競爭力的價格,但僅限於接下來的6
5:44.276–5:44.899
months .
個月。
5:44.909–5:48.305
Because after those six months are done , the model's pricing is
因為在那六個月結束後,該模型的定價
5:48.315–5:50.528
actually , you know , a little bit more expensive .
實際上,你知道,會稍微貴一點。
5:50.538–5:52.922
And now they show it at the bottom over here ,
現在他們在下方顯示了這一點,
5:52.932–5:56.990
you can see that after starting January of 2027 ,
你可以看到,從2027年1月開始,
5:57.000–5:58.763
it will become 1 .
它將變成1.
5:58.773–6:00.728
50 per input and 7 .
50的輸入和7.
6:00.738–6:01.947
50 per output .
50的輸出。
6:01.957–6:05.605
So yeah , it's still cheap compared to the other frontier labs
所以,是的,與其他前沿實驗室相比,它仍然很便宜,
6:05.615–6:08.896
, but it's not as cheap as , for example , MU Spark when the pricing
但不像MU Spark在定價更新時,甚至Flash的Deepsee版本那樣便宜。
6:08.906–6:12.352
is updated or even the Deepsee version for Flash .
或是更新,甚至是 Flash 的 Deepsee 版本。
6:12.362–6:13.955
But across these benchmarks ,
我們可以看到該模型在他們在上方強調的許多領域中表現更好。
6:13.955–6:15.879
we can see that this model is better
但他們還有一些其他領域,例如長影片理解,
6:15.889–6:18.283
in many areas that they highlighted at the top .
該模型在此方面表現出色,達到85.
6:18.293–6:21.768
But then they also have some other areas like long video understanding
但他們也涵蓋其他領域,例如長影片理解
6:21.778–6:24.452
, which this model excels at 85 .
9是Terra模型的得分。
6:24.462–6:26.501
4 versus 78 .
然後,舊模型在該方面也相當不錯,達到84.
6:26.511–6:28.298
9 for the Terra model .
9,適用於 Terra 模型。
6:28.308–6:31.897
And then the old model was also pretty good at that ,
然後舊模型在該方面也表現得相當不錯,
6:31.907–6:32.176
84 .
[未翻譯]
6:32.186–6:32.465
2 .
二
6:32.475–6:36.288
Then long context performance the model is at 97
接著是長上下文效能,該模型達到 97
6:36.288–6:37.432
and this model
而這個模型
6:37.442–6:41.273
the GPT 6 Terra one it sits at 93 .
也就是 GPT 6 Terra 版本,它位於 93。
6:41.283–6:42.080
5.
五
6:42.090–6:44.496
So yeah this model in summary it is better
所以,總而言之,這個模型更好
6:44.496–6:45.840
so it's not all negative
所以並非全是負面評價
6:45.850–6:48.424
but it's not all like you know that positive
但也並非全然如你所知的那麼正面
6:48.424–6:49.600
where you are super
讓你超級
6:49.610–6:50.921
excited for Google Deep Mind
為 Google DeepMind 感到興奮
6:50.921–6:52.450
because as I said they're probably
因為正如我所說,他們可能
6:52.460–6:55.689
still holding off their biggest release for Gemini 4 lineup .
仍在擱置其 Gemini 4 系列的最大規模發布。
6:55.699–6:58.635
Now , one thing people might have missed in their charts because
現在,人們可能在他們的圖表中錯過了一件事,因為
6:58.645–7:02.217
these charts sometimes are so messy to read and understand ,
這些圖表有時難以閱讀和理解,
7:02.227–7:03.412
3.
三
7:03.422–7:05.822
7 flash is worse than GPT 5.
7 flash 比 GPT 5 差。
7:05.832–7:06.188
6 Luna ,
6 月神
7:06.188–7:07.826
which is the model over here ,
也就是這裡的模型,
7:07.826–7:09.464
which is achieving a higher
它在這個基準測試中取得了更高的
7:09.474–7:13.841
score on this benchmark deepware engineering for about three times
Deepware 工程分數,成本大約只有三分之一。
7:13.851–7:14.809
the cost .
成本。
7:14.819–7:16.992
So , yeah , this is kind of interesting because yeah ,
所以,是的,這有點意思,因為確實,
7:17.002–7:18.340
the cost for Gemini 3.
Gemini 3 的成本
7:18.350–7:21.276
7 Flash is a little bit more than what it looks like .
7 Flash 的成本比看起來的要多一點。
7:21.286–7:23.044
Now , some people are a little bit upset
現在,有些人有點不高興,
7:23.044–7:23.781
and they're like ,
他們都說,
7:23.791–7:24.994
Oh , disgraceful .
「太丟人了。」
7:25.004–7:28.734
Google left out soul , opus , and fable because Google is incredibly
Google 沒有納入 Soul、Opus 和 Fable,是因為 Google 遠遠落後。
7:28.744–7:29.447
behind .
落後。
7:29.457–7:31.115
But I think one thing we got to remember , guys ,
但我認為我們必須記住一件事,各位,
7:31.125–7:33.336
is that this model is a flash model .
這個模型是一款 Flash 模型。
7:33.346–7:35.242
It's not trying to be a pro model
它不是要成為 Pro 模型,
7:35.242–7:36.986
or it's not trying to compete
也不是要競爭
7:36.996–7:40.458
That is probably going to be Gemini 4 .
那應該是 Gemini 4。
7:40.468–7:42.651
That is probably going to be Gemini 4 .
那應該是 Gemini 4。
7:42.661–7:44.077
So when Gemini 4 comes out ,
所以當 Gemini 4 推出時,
7:44.077–7:46.303
then I think it's okay for us to criticize
我認為我們才有資格批評
7:46.313–7:50.156
them if they don't include Soul , Opus , and Fable in their benchmark
如果他們在基準測試中沒有納入 Soul、Opus 和 Fable,我認為我們可以批評他們
7:50.166–7:51.401
charts because for now ,
因為就目前而言,
7:51.401–7:53.350
I think what they have done is pretty
我認為他們所做的事情相當
7:53.360–7:54.069
accurate .
準確。
7:54.079–7:55.750
One lab that I would have liked to see
我希望能看到的一個實驗室,
7:55.750–7:56.864
or one model for example
或者舉例來說一個模型,
7:56.874–7:59.478
I would have liked to see on that chart would be Deepseek version
我希望能在那張圖表上看到的,會是 Deepseek 第
7:59.488–7:59.793
4 Flash
4 版 Flash
7:59.793–8:02.134
because that would kind of spoil their release because
,因為這有點會洩露他們的發布消息,因為
8:02.144–8:05.521
that model is way cheaper compared to the Gemini 3.
該模型與 Gemini 3.7
8:05.531–8:07.058
7 flash model lineup .
Flash 模型陣容相比便宜得多。
8:07.068–8:10.807
Now this model jumped from number 19 to 8 on the web development
現在這個模型在網頁開發領域的排名從第19名躍升至第8名
8:10.817–8:11.095
area
領域
8:11.095–8:12.555
and we see it over here now
從第 19 名躍升至第 8 名,
8:12.555–8:14.224
and couple of models that are
我們現在在這裡看到它,
8:14.234–8:17.751
ahead of it are Opus 5 obviously Kim K3 Quinn 3 .
還有幾個領先的模型,當然是 Opus 5、Kim K3 Quinn 3。
8:17.761–8:20.384
8 Max Cloud Opus 5 Gro 4 .
[未翻譯]
8:20.394–8:22.409
6 6 , which is a model that came yesterday ,
6 6,這是一個昨天推出的模型,
8:22.409–8:23.353
which was a big win
這是一大勝利
8:23.363–8:25.900
for SpaceX, Fable 5, and 5.6.
,對 SpaceX、Fable 5 和 5.6 來說。
8:25.910–8:26.352
6.
[未翻譯]
8:26.362–8:27.002
Sol.
解
8:27.012–8:29.966
So , yeah , this model is trying to compete in the web development
所以,是的,這個模型試圖在網頁開發領域競爭
8:29.976–8:33.094
, but it's still behind all of these models , which is ,
,但它仍然落後於這些模型,這
8:33.104–8:35.305
you know , expected cuz it's a flash model .
,你知道,是預料之中的,因為它是一個閃電模型。
8:35.315–8:37.111
It's not really a pro model .
它並不是真正的專業模型。
8:37.121–8:40.296
Today, OpenAI has also launched a weight list for 5.6 so ultra fast mode.
今天,OpenAI 也為 5.6 版推出了權重列表,用於超快模式。
8:40.306–8:42.036
6 so ultra fast mode .
6 版的超快模式。
8:42.046–8:45.623
access to now through Cerebras and GPT 5.
現在可以透過 Cerebras 和 GPT 5 進行存取。
8:45.633–8:48.490
6 6o with this chip kind of running it is able to achieve an ultra
6 6o 使用這款晶片運行,能夠實現超快
8:48.500–8:52.774
6 6o with this chip kind of running it is able to achieve an ultra
6 6o 使用這款晶片運行,能夠實現超快
8:52.784–8:57.861
fast mode that generates up to 750 output tokens per second which
模式,每秒生成多達 750 個輸出 token,這
8:57.871–9:01.494
is about 14 times faster than the standard mode .
大約是標準模式的 14 倍速度。
9:01.504–9:04.812
So we're getting a really fast version of GPT 5 .
所以我們得到了一個非常快速的 GPT 5 .
9:04.822–9:09.505
6 so now this is supposed to be used for live or near production
6 版,現在這應該用於即時或接近生產環境
9:09.515–9:13.292
workloads like you know real time voice support commerce what this
的工作負載,比如你知道的即時語音支援商務,這
9:13.302–9:14.581
6 six sol to do is kind of be really fast in critical situations
6 6 Sol 要做的是在關鍵情況下非常快速
9:14.591–9:18.277
6 six soul to do is kind of be really fast in critical situations
6 6 Sol 要做的是在關鍵情況下非常快速
9:18.287–9:22.135
when people might be interacting with the AI agent like financial
當人們可能與 AI 代理互動時,例如金融
9:22.145–9:26.395
research security response support I think is going to be a big
研究、安全回應支援,我認為這將是一個很大的
9:26.405–9:31.508
area where this new ultrafast mode will be kind of implemented business
這個新的超快模式將在業務領域實施
9:31.518–9:35.098
and developer agents maybe yeah but I think like support or near
以及開發者代理,也許是,但我認為像支援或接近
9:35.108–9:38.527
production workloads like real time voice I see this model really
生產環境的工作負載,例如即時語音,我認為這個模型在
9:38.537–9:42.378
excelling at that and obviously this is still a weightless mode
這方面表現出色,顯然這仍然是一種無權重模式
9:42.388–9:45.323
we don't know how many people are going to get access to this ,
我們不知道有多少人會獲得訪問權限,
9:45.333–9:47.791
how successful it is or what the pricing is .
它的成功程度如何,或者定價是多少。
9:47.801–9:50.497
I don't know if the pricing has changed because there's no information
我不知道定價是否有所改變,因為在他們的實際
9:50.507–9:53.283
in their actual , you know , blog post .
部落格文章中沒有相關資訊,你知道的。
9:53.293–9:56.396
But as I mentioned , couple of areas where they mentioned that
但正如我所提到的,他們提到有幾個領域
9:56.406–9:58.233
this is going to be really important .
這部分將非常重要
9:58.243–10:01.669
Customer support and voice , commerce , live research and experimentation
客戶支援和語音、商務、即時研究和實驗
10:01.679–10:06.182
, financial research and security , incident response and reliability
、金融研究和安全、事件回應和可靠性
10:06.182–10:08.288
.
。
10:08.288–10:09.514
But yeah , I think this partnership is going to be important for
不過我認為這項合作對未來很重要
10:09.524–10:11.044
OpenAI going forward .
OpenAI 未來的發展將很重要。
10:11.054–10:12.978
But as I said , 14 times the speed .
但正如我所說,速度快了 14 倍。
10:12.988–10:14.589
What does that mean for cost ?
這對成本意味著什麼?
10:14.599–10:16.120
We don't know yet .
我們還不知道。
10:16.130–10:17.812
But that's it for today's video .
但今天的影片就到這裡。
10:17.822–10:20.068
Make sure you guys are subscribed to the channel .
記得訂閱我們的頻道
10:20.078–10:21.852
Follow our new newsletter as well at universe-of-ai.beehiiv.com
也請訂閱我們的新聞通訊,網址為 universe-of-ai.beehiiv.com
10:21.852–10:23.626
as well as subscribe to the main channel World of AI and support
同時訂閱主頻道 World of AI 並支持
10:23.636–10:24.014
behive .
behive.com,以及訂閱主頻道 World of AI 並支持
10:24.024–10:26.911
com as well as subscribe to the main channel World of AI and support
behive.com,同時在 X 上關注 Universe of AI 來支持我們。
10:26.921–10:29.396
us on X by following the Universe of AI as well .
behive.com,同時在 X 上關注 Universe of AI 來支持我們。
10:29.406–10:32.401
Until then , I'll see you guys in the next
在那之前,我們下期再見