WEBVTT

00:00:00.000 --> 00:00:05.580
So looks like we just got flashed by Google DeepMind once again because they just dropped

00:00:05.580 --> 00:00:11.820
Gemini 3.7 flash today. Yes, their newest model, which is once again a flash model, is here.

00:00:12.200 --> 00:00:18.520
And this is only after three weeks since Gemini 3.6 flash. And what they are saying is that Gemini

00:00:18.520 --> 00:00:25.160
3.7 flash is their most intelligent workhorse model. And this is coming three weeks after 3.6

00:00:25.160 --> 00:00:30.940
flash because of developer feedback and algorithmic innovations and that is why they're saying that

00:00:30.940 --> 00:00:35.960
gemini 3.7 flash model is you know out with so much improvement because if you take a look at

00:00:35.960 --> 00:00:42.280
the benchmarks this model is way better than a gemini 3.6 flash but i think this is also because

00:00:42.280 --> 00:00:47.920
we know internally that they're planning on canceling gemini 3.5 pro completely and working

00:00:47.920 --> 00:00:52.880
on gemini 4 at the moment so maybe the improvements they made with the pro model that they're not

00:00:52.880 --> 00:00:58.660
going to be dropping anymore maybe they repackaged it as a flash model because this model is better

00:00:58.660 --> 00:01:05.220
than 3.6 and if the timeline is three weeks plus they're planning on canceling the 3.5 pro it's

00:01:05.220 --> 00:01:09.520
quite likely that is the case we'll just ignore that and put that into the side we don't know

00:01:09.520 --> 00:01:14.300
when we're getting a new pro model from the google deep mine team but this model across many of the

00:01:14.300 --> 00:01:20.700
benchmarks we will see that yes it is stronger than the 3.6 flash model for example when it comes

00:01:20.700 --> 00:01:27.740
to code quality production code quality the model is producing 43.6 percent versus the flash model

00:01:27.740 --> 00:01:37.320
34.4 percent and then sonnet 5 is at 42.7 percent and the terra model is at 41.3 percent so one thing

00:01:37.320 --> 00:01:42.080
to remember since this is a flash model you're not going to see them compare this to opus 5 or any of

00:01:42.080 --> 00:01:48.200
the other stronger models the reason being because this is not their strongest tier and the 3.1 pro

00:01:48.200 --> 00:01:53.780
model when they finally decide to upgrade it to Gemini 4 Pro probably then we'll see it being

00:01:53.780 --> 00:01:59.760
compared to Opus 5 or GPT 5.6 Soul but anyways we can see that this model is an improvement from

00:01:59.760 --> 00:02:06.200
3.6 Flash and I would say that's basically the biggest result or the biggest update that we saw

00:02:06.200 --> 00:02:11.080
with the new model because it does not really like change up things a lot it's still like not

00:02:11.080 --> 00:02:16.360
becoming the number one model or the number one Flash model even on some benchmarks DeepSeek

00:02:16.360 --> 00:02:22.520
version for flash it's actually cheaper and more intelligent than this model but let's just take a

00:02:22.520 --> 00:02:27.500
look at the benchmarks that they have published on long horizon software engineering this model

00:02:27.500 --> 00:02:35.320
is a little bit behind gpd 5.6 tera which sits at 69.6 percent and then 65.3 percent is the flash

00:02:35.320 --> 00:02:43.640
model the older flash model 3.6 sits at 48.6 percent sonnet 5 sits at 53.8 percent and the

00:02:43.640 --> 00:02:48.580
new player that is finally being included on benchmarks which even the google deep mine team

00:02:48.580 --> 00:02:55.920
is considering with their launch is musepark 1.2 which is sitting at 54.9 so welcome meta to the

00:02:55.920 --> 00:03:01.660
benchmark charts because now we're starting to see it appear on more and more benchmarks and if we

00:03:01.660 --> 00:03:07.540
also take a look at web development this model is getting an elo score of 1588 versus their old

00:03:07.540 --> 00:03:12.020
model is 1538 so not a crazy difference but compared to everything else out there this is

00:03:12.020 --> 00:03:16.460
number one. When I say everything else out there, once again, compared to all the mid-tier models,

00:03:16.580 --> 00:03:21.500
this is beating all of them when it comes to web development. Now, this model is probably going to

00:03:21.500 --> 00:03:26.180
be used in enterprises a lot just because of Google's footprint in the enterprise space.

00:03:26.480 --> 00:03:33.500
But we are seeing this model achieve on the automation bench 30.4%. And GPT 5.6 Terra sits

00:03:33.500 --> 00:03:39.840
at 23.6%. So yes, this model is stronger than the other Flash models out there. But as I said,

00:03:39.840 --> 00:03:44.760
they haven't included DeepSeq version for Flash because if they do, in some areas that model is

00:03:44.760 --> 00:03:50.640
actually quite better than the 3.7 Flash model. Before we continue, if you're building AI agents

00:03:50.640 --> 00:03:55.600
or just messing around with them, Arcade is worth knowing about. It's the runtime that lets your

00:03:55.600 --> 00:04:00.240
agent actually do things instead of just talking about them. Because that's the gap right now.

00:04:00.540 --> 00:04:05.540
The models are smart enough. Your agent can figure out exactly what needs to happen in your email,

00:04:05.540 --> 00:04:11.300
your Slack, your CRM. It just can't go in and do it. And the reason isn't intelligence, it's

00:04:11.300 --> 00:04:16.740
permissions. Something has to prove the agent is allowed to act on behalf of a specific person

00:04:16.740 --> 00:04:21.940
in a specific account. That's the messy part everyone runs into, and it's the part ArcGate

00:04:21.940 --> 00:04:26.740
actually handles for you. So instead of your agent using one shared login for everybody,

00:04:26.740 --> 00:04:31.860
it acts as whoever is actually signed in with exactly the access that person has.

00:04:31.860 --> 00:04:37.140
If they can't see something, the agent can't either, and you never have to touch any of that setup yourself.

00:04:37.540 --> 00:04:38.540
Then there's the tools.

00:04:38.960 --> 00:04:44.000
Arcade has thousands of them already built for Gmail, Google Drive, Slack, Notion, Salesforce,

00:04:44.580 --> 00:04:49.120
most of the apps people already work in, and they are built specifically for AI to use,

00:04:49.120 --> 00:04:53.420
so the agent gets it right the first time instead of guessing and failing and retrying.

00:04:53.880 --> 00:04:57.960
It also keeps a record of everything, what the agent did for who and where,

00:04:58.280 --> 00:05:01.400
which matters a lot the moment other people start using the thing you built.

00:05:01.400 --> 00:05:07.640
And it works with whatever you're already using. Any model, any framework, cloud, cursor, chat GPT,

00:05:07.780 --> 00:05:11.960
doesn't matter. So you're not just giving an AI a list of tools and hoping it works,

00:05:12.260 --> 00:05:17.520
you're giving it a place where it can safely take real actions in real apps. It's free to start and

00:05:17.520 --> 00:05:21.780
the link is in the description. Thank you once again for Arcade for sponsoring today's video.

00:05:22.100 --> 00:05:28.700
Now let's get back into the video. Now one thing to note is that the Gemini 3.7 flash model through

00:05:28.700 --> 00:05:34.420
the end of this year so end of 2026 they have a cheap pricing model that they're placing on the

00:05:34.420 --> 00:05:40.880
model 75 cents per 1 million input tokens and 3 dollars and 75 cents for 1 million output tokens

00:05:40.880 --> 00:05:46.360
so it's a competitive price but this is only for the next six months because after those six months

00:05:46.360 --> 00:05:51.520
are done the model's pricing is actually you know a little bit more expensive and now they show it

00:05:51.520 --> 00:05:58.320
at the bottom over here you can see that after starting January of 2027 it will become a dollar

00:05:58.320 --> 00:06:04.700
and 50 per input and seven dollars and 50 per output so yeah it's still cheap compared to the

00:06:04.700 --> 00:06:09.460
other frontier labs but it's not as cheap as for example muse spark when the pricing is updated

00:06:09.460 --> 00:06:15.040
or even deep seek version for flash but across these benchmarks we can see that this model

00:06:15.040 --> 00:06:20.000
is better in many areas that they highlighted at the top but then they also have some other areas

00:06:20.000 --> 00:06:28.260
like long video understanding which this model excels at 85.4% versus 78.9% for the Terra model

00:06:28.260 --> 00:06:34.440
and then the old model was also pretty good at that 84.2%. Then long context performance the

00:06:34.440 --> 00:06:43.420
model is at 97% and this model the GPT-6 Terra one it sits at 93.5%. So yeah this model in summary

00:06:43.420 --> 00:06:49.240
it is better so it's not all negative but it's not all like you know that positive where you are

00:06:49.240 --> 00:06:53.440
super excited for Google DeepMind because as I said they're probably still holding off their

00:06:53.440 --> 00:06:58.600
biggest release for Gemini 4 lineup. Now one thing people might have missed in their charts because

00:06:58.600 --> 00:07:06.080
these charts sometimes are so messy to read and understand but the 3.7 flash is worse than GPT 5.6

00:07:06.080 --> 00:07:11.240
Luna which is the model over here which is achieving a higher score on this benchmark

00:07:11.240 --> 00:07:16.480
deep software engineering for about three times the cost. So yeah this is kind of interesting

00:07:16.480 --> 00:07:20.960
because yeah the cost for gemini 3.7 flash is a little bit more than what it looks like

00:07:20.960 --> 00:07:25.860
now some people are a little bit upset and they're like oh disgraceful google left out

00:07:25.860 --> 00:07:30.860
so opus and fable because google is incredibly behind but i think one thing we got to remember

00:07:30.860 --> 00:07:33.162
guys is that this model is a flash model

00:07:34.162 --> 00:07:40.222
to be a pro model or it's not trying to compete with opus or soul or fable five level models yet

00:07:40.222 --> 00:07:45.522
that is probably going to be gemini 4 so when gemini 4 comes out then i think it's okay for

00:07:45.522 --> 00:07:50.762
us to criticize them if they don't include Sol, Opus, and Fable in their benchmark charts because

00:07:50.762 --> 00:07:55.642
for now I think what they have done is pretty accurate. One lab that I would have liked to see

00:07:55.642 --> 00:08:00.062
or one model for example I would have liked to see on that chart would be DeepSeq version 4 Flash

00:08:00.062 --> 00:08:04.642
because that would kind of spoil their release because that model is way cheaper compared to

00:08:04.642 --> 00:08:11.022
the Gemini 3.7 Flash model lineup. Now this model jumped from number 19 to 8 on the web development

00:08:11.022 --> 00:08:16.242
area and we see it over here now and a couple of models that are ahead of it are Opus 5 obviously,

00:08:16.422 --> 00:08:23.662
Kimi K3, Quen 3.8 Max, Cloud Opus 5, Grok 4.6 which is a model that came yesterday which was a big win

00:08:23.662 --> 00:08:30.222
for SpaceX, Fable 5 and 5.6 Sol. So yeah this model is trying to compete in the web development

00:08:30.222 --> 00:08:35.282
but it's still behind all of these models which is you know expected because it's a flash model

00:08:35.282 --> 00:08:42.042
it's not really a pro model. Today OpenAI has also launched a waitlist for 5.6 SOL ultra fast mode.

00:08:42.202 --> 00:08:47.382
Now this is possible because of the fast chips that they have access to now through Cerebris

00:08:47.382 --> 00:08:53.922
and GPT 5.6 SOL with this chip kind of running it is able to achieve an ultra fast mode that

00:08:53.922 --> 00:09:01.342
generates up to 750 output tokens per second which is about 14 times faster than the standard mode.

00:09:01.342 --> 00:09:07.822
So we're getting a really fast version of GPT 5.6 Soul. And now this is supposed to be used for

00:09:07.822 --> 00:09:12.522
live or near production workloads like, you know, real-time voice, support, commerce.

00:09:12.762 --> 00:09:19.002
What this allows GPT 5.6 Soul to do is kind of be really fast in critical situations when people

00:09:19.002 --> 00:09:24.982
might be interacting with the AI agent, like financial research, security response, support,

00:09:25.082 --> 00:09:30.382
I think is going to be a big area where this new ultra fast mode will be kind of implemented.

00:09:30.382 --> 00:09:35.962
business and developer agents maybe yeah but i think like support or near production workloads

00:09:35.962 --> 00:09:41.322
like real-time voice i see this model really excelling at that and obviously this is still

00:09:41.322 --> 00:09:46.562
a waitlist mode we don't know how many people are going to get access to this how successful it is

00:09:46.562 --> 00:09:50.982
or what the pricing is i don't know if the pricing has changed because there's no information in their

00:09:50.982 --> 00:09:56.622
actual you know blog post but as i mentioned a couple of areas where they mentioned that this

00:09:56.622 --> 00:10:01.822
is going to be really important customer support and voice commerce live research and experimentation

00:10:01.822 --> 00:10:08.062
financial research and security incident response and reliability but yeah i think this partnership

00:10:08.062 --> 00:10:13.642
is going to be important for open ai going forward but as i said 14 times the speed what does that

00:10:13.642 --> 00:10:18.902
mean for cost we don't know yet but that's it for today's video make sure you guys are subscribed

00:10:18.902 --> 00:10:25.282
to the channel follow our new newsletter as well at universeofai.beehive.com as well subscribe to

00:10:25.282 --> 00:10:30.402
the main channel, World of AI, and support us on X by following the universe of AIZ as well.

00:10:30.762 --> 00:10:32.722
Until then, I'll see you guys in the next video.
