WEBVTT

00:00:00.240 --> 00:00:02.480
So, looks like we just got flashed by

00:00:02.480 --> 00:00:04.640
Google DeepMind [music] once again

00:00:04.640 --> 00:00:06.879
because they just dropped Gemini 3.7

00:00:06.879 --> 00:00:09.440
Flash today. Yes, their newest model,

00:00:09.440 --> 00:00:11.679
which is once again a Flash model, is

00:00:11.679 --> 00:00:14.160
here. And this is only after 3 weeks

00:00:14.160 --> 00:00:17.279
since Gemini 3.6 Flash. And what they

00:00:17.279 --> 00:00:20.000
are saying is that Gemini 3.7 Flash is

00:00:20.000 --> 00:00:22.800
their most intelligent workhorse model.

00:00:22.800 --> 00:00:25.359
And this is coming 3 weeks after 3.6 six

00:00:25.359 --> 00:00:27.840
flash because of developer feedback and

00:00:27.840 --> 00:00:30.480
algorithmic innovations. And that is why

00:00:30.480 --> 00:00:32.480
they're saying the Gemini 3.7 Flash

00:00:32.480 --> 00:00:34.320
model is, you know, out with so much

00:00:34.320 --> 00:00:36.160
improvement cuz if we take a look at the

00:00:36.160 --> 00:00:38.399
benchmarks, this model is way better

00:00:38.399 --> 00:00:41.360
than the Gemini 3.6 Flash. But I think

00:00:41.360 --> 00:00:43.440
this is also because we know internally

00:00:43.440 --> 00:00:44.879
that they're planning on cancelelling

00:00:44.879 --> 00:00:48.160
Gemini 3.5 Pro completely and working on

00:00:48.160 --> 00:00:50.480
Gemini 4 at the moment. So maybe the

00:00:50.480 --> 00:00:51.920
improvements they made with the Pro

00:00:51.920 --> 00:00:53.520
model that they're not going to be

00:00:53.520 --> 00:00:55.840
dropping anymore. Maybe they repackaged

00:00:55.840 --> 00:00:58.320
it as a flash model because this model

00:00:58.320 --> 00:01:01.199
is better than 3.6 and if the timeline

00:01:01.199 --> 00:01:03.120
is three weeks plus they're planning on

00:01:03.120 --> 00:01:05.519
cancelling the 3.5 Pro, it's quite

00:01:05.519 --> 00:01:07.600
likely that is the case. We'll just

00:01:07.600 --> 00:01:09.119
ignore that and put that into the side.

00:01:09.119 --> 00:01:10.479
We don't know when we're getting a new

00:01:10.479 --> 00:01:12.080
Pro model from the Google Deep Mind

00:01:12.080 --> 00:01:14.400
team, but this model across many of the

00:01:14.400 --> 00:01:17.119
benchmarks we will see that yes, it is

00:01:17.119 --> 00:01:19.920
stronger than the 3.6 Flash model. For

00:01:19.920 --> 00:01:21.840
example, when it comes to code quality,

00:01:21.840 --> 00:01:24.159
production code quality, the model is

00:01:24.159 --> 00:01:26.159
producing 43.6%

00:01:26.159 --> 00:01:29.680
versus the flash model 34.4%.

00:01:29.680 --> 00:01:33.520
And then Sonet 5 is at 42.7%

00:01:33.520 --> 00:01:36.799
and the Terra model is at 41.3%.

00:01:36.799 --> 00:01:38.799
So, one thing to remember, since this is

00:01:38.799 --> 00:01:40.240
a flash model, you're not going to see

00:01:40.240 --> 00:01:42.159
them compare this to Opus 5 or any of

00:01:42.159 --> 00:01:44.240
the other stronger models. The reason

00:01:44.240 --> 00:01:46.240
being cuz this is not their strongest

00:01:46.240 --> 00:01:49.119
tier. and the 3.1 Pro model when they

00:01:49.119 --> 00:01:51.600
finally decide to upgrade it to Gemini 4

00:01:51.600 --> 00:01:54.479
Pro, then we'll see it being compared to

00:01:54.479 --> 00:01:57.920
Opus 5 or GPT 5.6 Soul. But anyways, we

00:01:57.920 --> 00:01:59.280
can see that this model is an

00:01:59.280 --> 00:02:01.520
improvement from 3.6 Flash and I would

00:02:01.520 --> 00:02:04.560
say that's basically the biggest result

00:02:04.560 --> 00:02:06.479
or the biggest update that we saw with

00:02:06.479 --> 00:02:08.560
the new model because it does not really

00:02:08.560 --> 00:02:10.720
like change up things a lot is still

00:02:10.720 --> 00:02:12.560
like not becoming the number one model

00:02:12.560 --> 00:02:15.040
or the number one flash model. Even on

00:02:15.040 --> 00:02:17.200
some benchmarks, Deepseek version 4

00:02:17.200 --> 00:02:19.920
flash is actually cheaper and more

00:02:19.920 --> 00:02:22.160
intelligent than this model. But let's

00:02:22.160 --> 00:02:23.680
just take a look at the benchmarks that

00:02:23.680 --> 00:02:25.840
they have published. On Long Horizon

00:02:25.840 --> 00:02:28.080
software engineering, this model is a

00:02:28.080 --> 00:02:30.959
little bit behind GPT 5.6 Terra, which

00:02:30.959 --> 00:02:32.959
sits at 69.6%

00:02:32.959 --> 00:02:36.400
and then 65.3% is the Flash model. The

00:02:36.400 --> 00:02:40.800
older Flash model 3.6 sits at 48.6%.

00:02:40.800 --> 00:02:44.000
Sonic 5 sits at 53.8% 8%. And the new

00:02:44.000 --> 00:02:46.160
player that is finally being included on

00:02:46.160 --> 00:02:48.000
benchmarks, which even the Google

00:02:48.000 --> 00:02:49.920
DeepMind team is considering with their

00:02:49.920 --> 00:02:52.480
launch, is Muse Spark 1.2, which is

00:02:52.480 --> 00:02:54.640
sitting at 54.9%.

00:02:54.640 --> 00:02:57.519
So, welcome Meta to the benchmark charts

00:02:57.519 --> 00:02:59.360
because now we're starting to see it

00:02:59.360 --> 00:03:01.519
appear on more and more benchmarks. And

00:03:01.519 --> 00:03:02.800
if we also take a look at web

00:03:02.800 --> 00:03:05.360
development, this model is getting a ELO

00:03:05.360 --> 00:03:08.159
score of 1588 versus their old model is

00:03:08.159 --> 00:03:10.319
1538. So, not a crazy difference, but

00:03:10.319 --> 00:03:11.840
compared to everything else out there,

00:03:11.840 --> 00:03:13.360
this is number one. When I say

00:03:13.360 --> 00:03:14.879
everything else out there, once again,

00:03:14.879 --> 00:03:16.720
compared to all the mid-tier models,

00:03:16.720 --> 00:03:18.640
this is beating all of them when it

00:03:18.640 --> 00:03:20.640
comes to web development. Now, this

00:03:20.640 --> 00:03:22.480
model is probably going to be used in

00:03:22.480 --> 00:03:24.159
enterprises a lot just because of

00:03:24.159 --> 00:03:26.000
Google's footprint in the enterprise

00:03:26.000 --> 00:03:28.159
space, but we are seeing this model

00:03:28.159 --> 00:03:31.599
achieve on the automation bench 30.4%.

00:03:31.599 --> 00:03:35.440
And GPT 5.6 Terra sits at 23.6%.

00:03:35.440 --> 00:03:37.599
So, yes, this model is stronger than the

00:03:37.599 --> 00:03:39.760
other Flash models out there, but as I

00:03:39.760 --> 00:03:41.440
said, they haven't included Deepseek

00:03:41.440 --> 00:03:43.599
version for Flash because if they do, in

00:03:43.599 --> 00:03:45.519
some areas, that model is actually quite

00:03:45.519 --> 00:03:48.480
better than the 3.7 Flash model. Before

00:03:48.480 --> 00:03:50.319
we continue, if you're building AI

00:03:50.319 --> 00:03:52.480
agents or just messing around with them,

00:03:52.480 --> 00:03:54.799
Arcade is worth knowing about. It's the

00:03:54.799 --> 00:03:56.799
runtime that lets your agent actually do

00:03:56.799 --> 00:03:58.319
things instead of just talking about

00:03:58.319 --> 00:04:00.560
them because that's the gap right now.

00:04:00.560 --> 00:04:02.799
The models are smart enough. Your agent

00:04:02.799 --> 00:04:04.720
can figure out exactly what needs to

00:04:04.720 --> 00:04:06.560
happen in your email, your Slack, your

00:04:06.560 --> 00:04:09.680
CRM. It just can't go in and do it. And

00:04:09.680 --> 00:04:11.519
the reason isn't intelligence, it's

00:04:11.519 --> 00:04:13.680
permissions. Something has to prove the

00:04:13.680 --> 00:04:15.840
agent is allowed to act on behalf of a

00:04:15.840 --> 00:04:18.479
specific person in a specific account.

00:04:18.479 --> 00:04:20.239
That's the messy part everyone runs

00:04:20.239 --> 00:04:22.479
into, and it's the partit actually

00:04:22.479 --> 00:04:24.320
handles for you. So instead of your

00:04:24.320 --> 00:04:26.240
agent using one shared login for

00:04:26.240 --> 00:04:28.160
everybody, it acts as whoever is

00:04:28.160 --> 00:04:30.320
actually signed in with exactly the

00:04:30.320 --> 00:04:32.560
access that person has. If they can't

00:04:32.560 --> 00:04:34.880
see something, the agent can't either.

00:04:34.880 --> 00:04:36.479
And you never have to touch any of that

00:04:36.479 --> 00:04:38.960
setup yourself. Then there's the tools.

00:04:38.960 --> 00:04:40.800
Arcade has thousands of them already

00:04:40.800 --> 00:04:43.120
built for Gmail, Google Drive, Slack,

00:04:43.120 --> 00:04:45.520
Notion, Salesforce, most of the apps

00:04:45.520 --> 00:04:47.360
people already work in, and they are

00:04:47.360 --> 00:04:49.759
built specifically for AI to use. So the

00:04:49.759 --> 00:04:51.440
agent gets it right the first time

00:04:51.440 --> 00:04:53.040
instead of guessing and failing and

00:04:53.040 --> 00:04:55.120
retrying. It also keeps a record of

00:04:55.120 --> 00:04:57.520
everything, what the agent did, for who

00:04:57.520 --> 00:04:59.440
and where, which matters a lot the

00:04:59.440 --> 00:05:00.880
moment other people start using the

00:05:00.880 --> 00:05:02.639
thing you built. And it works with

00:05:02.639 --> 00:05:04.560
whatever you're already using. Any

00:05:04.560 --> 00:05:06.960
model, any framework, cloud, cursor,

00:05:06.960 --> 00:05:09.520
chat, GPT, doesn't matter. So you're not

00:05:09.520 --> 00:05:11.139
just giving an AI a list of tools and

00:05:11.139 --> 00:05:12.720
[music] hoping it works. You're giving

00:05:12.720 --> 00:05:14.800
it a place where it can safely take real

00:05:14.800 --> 00:05:17.360
actions in real apps. It's free to start

00:05:17.360 --> 00:05:19.120
and the link is in the description.

00:05:19.120 --> 00:05:20.880
Thank you once again for Arcade for

00:05:20.880 --> 00:05:22.960
sponsoring today's video. Now, let's get

00:05:22.960 --> 00:05:25.440
back into the video. Now, one thing to

00:05:25.440 --> 00:05:28.560
note is that the Gemini 3.7 Flash model

00:05:28.560 --> 00:05:30.320
through the end of this year, so end of

00:05:30.320 --> 00:05:33.360
2026, they have a cheap pricing model

00:05:33.360 --> 00:05:36.160
that they're placing on the model. 75

00:05:36.160 --> 00:05:39.840
per 1 million input tokens and $3.75 for

00:05:39.840 --> 00:05:41.680
1 million output tokens. So, it's a

00:05:41.680 --> 00:05:43.759
competitive price, but this is only for

00:05:43.759 --> 00:05:45.919
the next 6 months. Because after those

00:05:45.919 --> 00:05:48.240
six months are done, the model's pricing

00:05:48.240 --> 00:05:50.160
is actually, you know, a little bit more

00:05:50.160 --> 00:05:51.919
expensive. And now they show it at the

00:05:51.919 --> 00:05:54.720
bottom over here, you can see that after

00:05:54.720 --> 00:05:58.160
starting January of 2027, it will become

00:05:58.160 --> 00:06:02.479
$1.50 per input and $7.50 per output. So

00:06:02.479 --> 00:06:04.720
yeah, it's still cheap compared to the

00:06:04.720 --> 00:06:06.560
other frontier labs, but it's not as

00:06:06.560 --> 00:06:08.639
cheap as, for example, MU Spark when the

00:06:08.639 --> 00:06:11.199
pricing is updated or even the Deepsee

00:06:11.199 --> 00:06:13.280
version for Flash. But across these

00:06:13.280 --> 00:06:15.440
benchmarks, we can see that this model

00:06:15.440 --> 00:06:17.360
is better in many areas that they

00:06:17.360 --> 00:06:18.880
highlighted at the top. But then they

00:06:18.880 --> 00:06:21.199
also have some other areas like long

00:06:21.199 --> 00:06:22.960
video understanding, which this model

00:06:22.960 --> 00:06:27.199
excels at 85.4% versus 78.9%

00:06:27.199 --> 00:06:29.120
for the Terra model. And then the old

00:06:29.120 --> 00:06:30.880
model was also pretty good at that,

00:06:30.880 --> 00:06:32.550
84.2%.

00:06:32.550 --> 00:06:32.560
84.2%.

00:06:32.560 --> 00:06:34.800
Then long context performance the model

00:06:34.800 --> 00:06:39.440
is at 97% and this model the GPT 6 Terra

00:06:39.440 --> 00:06:42.160
one it sits at 93.5%.

00:06:42.160 --> 00:06:44.160
So yeah this model in summary it is

00:06:44.160 --> 00:06:46.479
better so it's not all negative but it's

00:06:46.479 --> 00:06:48.639
not all like you know that positive

00:06:48.639 --> 00:06:50.479
where you are super excited for Google

00:06:50.479 --> 00:06:52.240
Deep Mind because as I said they're

00:06:52.240 --> 00:06:53.840
probably still holding off their biggest

00:06:53.840 --> 00:06:56.319
release for Gemini 4 lineup. Now, one

00:06:56.319 --> 00:06:58.160
thing people might have missed in their

00:06:58.160 --> 00:06:59.759
charts because these charts sometimes

00:06:59.759 --> 00:07:02.800
are so messy to read and understand, but

00:07:02.800 --> 00:07:06.160
the 3.7 flash is worse than GPT 5.6

00:07:06.160 --> 00:07:08.160
Luna, which is the model over here,

00:07:08.160 --> 00:07:10.560
which is achieving a higher score on

00:07:10.560 --> 00:07:13.199
this benchmark deepware engineering for

00:07:13.199 --> 00:07:15.520
about three times the cost. So, yeah,

00:07:15.520 --> 00:07:17.039
this is kind of interesting because

00:07:17.039 --> 00:07:19.599
yeah, the cost for Gemini 3.7 Flash is a

00:07:19.599 --> 00:07:21.440
little bit more than what it looks like.

00:07:21.440 --> 00:07:23.360
Now, some people are a little bit upset

00:07:23.360 --> 00:07:25.120
and they're like, "Oh, disgraceful.

00:07:25.120 --> 00:07:27.520
Google left out soul, opus, and fable

00:07:27.520 --> 00:07:29.680
because Google is incredibly behind. But

00:07:29.680 --> 00:07:31.039
I think one thing we got to remember,

00:07:31.039 --> 00:07:33.120
guys, is that this model is a flash

00:07:33.120 --> 00:07:35.919
model. It's not trying to be a pro model

00:07:35.919 --> 00:07:38.000
or it's not trying to compete with Opus

00:07:38.000 --> 00:07:40.639
or Soul or Fable 5 level models yet.

00:07:40.639 --> 00:07:42.800
That is probably going to be Gemini 4.

00:07:42.800 --> 00:07:44.880
So when Gemini 4 comes out, then I think

00:07:44.880 --> 00:07:47.199
it's okay for us to criticize them if

00:07:47.199 --> 00:07:49.440
they don't include Soul, Opus, and Fable

00:07:49.440 --> 00:07:51.199
in their benchmark charts because for

00:07:51.199 --> 00:07:53.199
now, I think what they have done is

00:07:53.199 --> 00:07:55.120
pretty accurate. One lab that I would

00:07:55.120 --> 00:07:56.639
have liked to see or one model for

00:07:56.639 --> 00:07:58.000
example I would have liked to see on

00:07:58.000 --> 00:07:59.840
that chart would be Deepseek version 4

00:07:59.840 --> 00:08:01.520
Flash because that would kind of spoil

00:08:01.520 --> 00:08:03.520
their release because that model is way

00:08:03.520 --> 00:08:06.319
cheaper compared to the Gemini 3.7 flash

00:08:06.319 --> 00:08:08.800
model lineup. Now this model jumped from

00:08:08.800 --> 00:08:11.199
number 19 to 8 on the web development

00:08:11.199 --> 00:08:13.599
area and we see it over here now and

00:08:13.599 --> 00:08:15.120
couple of models that are ahead of it

00:08:15.120 --> 00:08:18.160
are Opus 5 obviously Kim K3 Quinn 3.8

00:08:18.160 --> 00:08:21.680
Max Cloud Opus 5 Gro 4.6 6, which is a

00:08:21.680 --> 00:08:23.360
model that came yesterday, which was a

00:08:23.360 --> 00:08:26.720
big win for SpaceX, Fable 5, and 5.6.

00:08:26.720 --> 00:08:28.960
Soul. So, yeah, this model is trying to

00:08:28.960 --> 00:08:31.120
compete in the web development, but it's

00:08:31.120 --> 00:08:33.120
still behind all of these models, which

00:08:33.120 --> 00:08:35.120
is, you know, expected cuz it's a flash

00:08:35.120 --> 00:08:37.279
model. It's not really a pro model.

00:08:37.279 --> 00:08:39.599
Today, OpenAI has also launched a weight

00:08:39.599 --> 00:08:42.479
list for 5.6 so ultra fast mode. Now,

00:08:42.479 --> 00:08:44.959
this is possible because of the fast

00:08:44.959 --> 00:08:46.640
chips that they have access to now

00:08:46.640 --> 00:08:49.839
through Cabus and GPT 5.6 6o with this

00:08:49.839 --> 00:08:52.160
chip kind of running it is able to

00:08:52.160 --> 00:08:54.000
achieve an ultra fast mode that

00:08:54.000 --> 00:08:57.120
generates up to 750 output tokens per

00:08:57.120 --> 00:09:00.399
second which is about 14 times faster

00:09:00.399 --> 00:09:02.399
than the standard mode. So we're getting

00:09:02.399 --> 00:09:06.320
a really fast version of GPT 5.6 so now

00:09:06.320 --> 00:09:08.880
this is supposed to be used for live or

00:09:08.880 --> 00:09:10.720
near production workloads like you know

00:09:10.720 --> 00:09:13.279
real-time voice support commerce what

00:09:13.279 --> 00:09:15.839
this allows GPT 5.6 six soul to do is

00:09:15.839 --> 00:09:17.680
kind of be really fast in critical

00:09:17.680 --> 00:09:19.519
situations when people might be

00:09:19.519 --> 00:09:21.680
interacting with the AI agent like

00:09:21.680 --> 00:09:24.640
financial research security response

00:09:24.640 --> 00:09:26.480
support I think is going to be a big

00:09:26.480 --> 00:09:29.360
area where this new ultraast mode will

00:09:29.360 --> 00:09:32.000
be kind of implemented business and

00:09:32.000 --> 00:09:33.920
developer agents maybe yeah but I think

00:09:33.920 --> 00:09:35.600
like support or near production

00:09:35.600 --> 00:09:37.760
workloads like real-time voice I see

00:09:37.760 --> 00:09:40.160
this model really excelling at that and

00:09:40.160 --> 00:09:42.160
obviously this is still a weightless

00:09:42.160 --> 00:09:44.160
mode we don't know how many people are

00:09:44.160 --> 00:09:45.760
going to get access to this, how

00:09:45.760 --> 00:09:47.920
successful it is or what the pricing is.

00:09:47.920 --> 00:09:49.440
I don't know if the pricing has changed

00:09:49.440 --> 00:09:51.200
because there's no information in their

00:09:51.200 --> 00:09:54.000
actual, you know, blog post. But as I

00:09:54.000 --> 00:09:55.920
mentioned, couple of areas where they

00:09:55.920 --> 00:09:57.360
mentioned that this is going to be

00:09:57.360 --> 00:09:59.360
really important. Customer support and

00:09:59.360 --> 00:10:01.200
voice, commerce, live research and

00:10:01.200 --> 00:10:03.760
experimentation, financial research and

00:10:03.760 --> 00:10:05.839
security, incident response and

00:10:05.839 --> 00:10:07.760
reliability. But yeah, I think this

00:10:07.760 --> 00:10:09.680
partnership is going to be important for

00:10:09.680 --> 00:10:12.240
OpenAI going forward. But as I said, 14

00:10:12.240 --> 00:10:14.240
times the speed. What does that mean for

00:10:14.240 --> 00:10:16.800
cost? We don't know yet. But that's it

00:10:16.800 --> 00:10:18.560
for today's video. Make sure you guys

00:10:18.560 --> 00:10:20.480
are subscribed to the channel. Follow

00:10:20.480 --> 00:10:22.160
our new newsletter as well at

00:10:22.160 --> 00:10:24.389
universeai.behive.com

00:10:24.389 --> 00:10:24.399
universeai.behive.com

00:10:24.399 --> 00:10:26.079
as well as subscribe to the main channel

00:10:26.079 --> 00:10:28.560
World of AI and support us on X by

00:10:28.560 --> 00:10:30.800
following the Universe of AIZ as well.

00:10:30.800 --> 00:10:32.399
Until then, I'll see you guys in the

00:10:32.399 --> 00:10:34.560
next
