1
00:00:00,000 --> 00:00:05,580
So looks like we just got flashed by Google DeepMind once again because they just dropped

2
00:00:05,580 --> 00:00:11,820
Gemini 3.7 flash today. Yes, their newest model, which is once again a flash model, is here.

3
00:00:12,200 --> 00:00:18,520
And this is only after three weeks since Gemini 3.6 flash. And what they are saying is that Gemini

4
00:00:18,520 --> 00:00:25,160
3.7 flash is their most intelligent workhorse model. And this is coming three weeks after 3.6

5
00:00:25,160 --> 00:00:30,940
flash because of developer feedback and algorithmic innovations and that is why they're saying that

6
00:00:30,940 --> 00:00:35,960
gemini 3.7 flash model is you know out with so much improvement because if you take a look at

7
00:00:35,960 --> 00:00:42,280
the benchmarks this model is way better than a gemini 3.6 flash but i think this is also because

8
00:00:42,280 --> 00:00:47,920
we know internally that they're planning on canceling gemini 3.5 pro completely and working

9
00:00:47,920 --> 00:00:52,880
on gemini 4 at the moment so maybe the improvements they made with the pro model that they're not

10
00:00:52,880 --> 00:00:58,660
going to be dropping anymore maybe they repackaged it as a flash model because this model is better

11
00:00:58,660 --> 00:01:05,220
than 3.6 and if the timeline is three weeks plus they're planning on canceling the 3.5 pro it's

12
00:01:05,220 --> 00:01:09,520
quite likely that is the case we'll just ignore that and put that into the side we don't know

13
00:01:09,520 --> 00:01:14,300
when we're getting a new pro model from the google deep mine team but this model across many of the

14
00:01:14,300 --> 00:01:20,700
benchmarks we will see that yes it is stronger than the 3.6 flash model for example when it comes

15
00:01:20,700 --> 00:01:27,740
to code quality production code quality the model is producing 43.6 percent versus the flash model

16
00:01:27,740 --> 00:01:37,320
34.4 percent and then sonnet 5 is at 42.7 percent and the terra model is at 41.3 percent so one thing

17
00:01:37,320 --> 00:01:42,080
to remember since this is a flash model you're not going to see them compare this to opus 5 or any of

18
00:01:42,080 --> 00:01:48,200
the other stronger models the reason being because this is not their strongest tier and the 3.1 pro

19
00:01:48,200 --> 00:01:53,780
model when they finally decide to upgrade it to Gemini 4 Pro probably then we'll see it being

20
00:01:53,780 --> 00:01:59,760
compared to Opus 5 or GPT 5.6 Soul but anyways we can see that this model is an improvement from

21
00:01:59,760 --> 00:02:06,200
3.6 Flash and I would say that's basically the biggest result or the biggest update that we saw

22
00:02:06,200 --> 00:02:11,080
with the new model because it does not really like change up things a lot it's still like not

23
00:02:11,080 --> 00:02:16,360
becoming the number one model or the number one Flash model even on some benchmarks DeepSeek

24
00:02:16,360 --> 00:02:22,520
version for flash it's actually cheaper and more intelligent than this model but let's just take a

25
00:02:22,520 --> 00:02:27,500
look at the benchmarks that they have published on long horizon software engineering this model

26
00:02:27,500 --> 00:02:35,320
is a little bit behind gpd 5.6 tera which sits at 69.6 percent and then 65.3 percent is the flash

27
00:02:35,320 --> 00:02:43,640
model the older flash model 3.6 sits at 48.6 percent sonnet 5 sits at 53.8 percent and the

28
00:02:43,640 --> 00:02:48,580
new player that is finally being included on benchmarks which even the google deep mine team

29
00:02:48,580 --> 00:02:55,920
is considering with their launch is musepark 1.2 which is sitting at 54.9 so welcome meta to the

30
00:02:55,920 --> 00:03:01,660
benchmark charts because now we're starting to see it appear on more and more benchmarks and if we

31
00:03:01,660 --> 00:03:07,540
also take a look at web development this model is getting an elo score of 1588 versus their old

32
00:03:07,540 --> 00:03:12,020
model is 1538 so not a crazy difference but compared to everything else out there this is

33
00:03:12,020 --> 00:03:16,460
number one. When I say everything else out there, once again, compared to all the mid-tier models,

34
00:03:16,580 --> 00:03:21,500
this is beating all of them when it comes to web development. Now, this model is probably going to

35
00:03:21,500 --> 00:03:26,180
be used in enterprises a lot just because of Google's footprint in the enterprise space.

36
00:03:26,480 --> 00:03:33,500
But we are seeing this model achieve on the automation bench 30.4%. And GPT 5.6 Terra sits

37
00:03:33,500 --> 00:03:39,840
at 23.6%. So yes, this model is stronger than the other Flash models out there. But as I said,

38
00:03:39,840 --> 00:03:44,760
they haven't included DeepSeq version for Flash because if they do, in some areas that model is

39
00:03:44,760 --> 00:03:50,640
actually quite better than the 3.7 Flash model. Before we continue, if you're building AI agents

40
00:03:50,640 --> 00:03:55,600
or just messing around with them, Arcade is worth knowing about. It's the runtime that lets your

41
00:03:55,600 --> 00:04:00,240
agent actually do things instead of just talking about them. Because that's the gap right now.

42
00:04:00,540 --> 00:04:05,540
The models are smart enough. Your agent can figure out exactly what needs to happen in your email,

43
00:04:05,540 --> 00:04:11,300
your Slack, your CRM. It just can't go in and do it. And the reason isn't intelligence, it's

44
00:04:11,300 --> 00:04:16,740
permissions. Something has to prove the agent is allowed to act on behalf of a specific person

45
00:04:16,740 --> 00:04:21,940
in a specific account. That's the messy part everyone runs into, and it's the part ArcGate

46
00:04:21,940 --> 00:04:26,740
actually handles for you. So instead of your agent using one shared login for everybody,

47
00:04:26,740 --> 00:04:31,860
it acts as whoever is actually signed in with exactly the access that person has.

48
00:04:31,860 --> 00:04:37,140
If they can't see something, the agent can't either, and you never have to touch any of that setup yourself.

49
00:04:37,540 --> 00:04:38,540
Then there's the tools.

50
00:04:38,960 --> 00:04:44,000
Arcade has thousands of them already built for Gmail, Google Drive, Slack, Notion, Salesforce,

51
00:04:44,580 --> 00:04:49,120
most of the apps people already work in, and they are built specifically for AI to use,

52
00:04:49,120 --> 00:04:53,420
so the agent gets it right the first time instead of guessing and failing and retrying.

53
00:04:53,880 --> 00:04:57,960
It also keeps a record of everything, what the agent did for who and where,

54
00:04:58,280 --> 00:05:01,400
which matters a lot the moment other people start using the thing you built.

55
00:05:01,400 --> 00:05:07,640
And it works with whatever you're already using. Any model, any framework, cloud, cursor, chat GPT,

56
00:05:07,780 --> 00:05:11,960
doesn't matter. So you're not just giving an AI a list of tools and hoping it works,

57
00:05:12,260 --> 00:05:17,520
you're giving it a place where it can safely take real actions in real apps. It's free to start and

58
00:05:17,520 --> 00:05:21,780
the link is in the description. Thank you once again for Arcade for sponsoring today's video.

59
00:05:22,100 --> 00:05:28,700
Now let's get back into the video. Now one thing to note is that the Gemini 3.7 flash model through

60
00:05:28,700 --> 00:05:34,420
the end of this year so end of 2026 they have a cheap pricing model that they're placing on the

61
00:05:34,420 --> 00:05:40,880
model 75 cents per 1 million input tokens and 3 dollars and 75 cents for 1 million output tokens

62
00:05:40,880 --> 00:05:46,360
so it's a competitive price but this is only for the next six months because after those six months

63
00:05:46,360 --> 00:05:51,520
are done the model's pricing is actually you know a little bit more expensive and now they show it

64
00:05:51,520 --> 00:05:58,320
at the bottom over here you can see that after starting January of 2027 it will become a dollar

65
00:05:58,320 --> 00:06:04,700
and 50 per input and seven dollars and 50 per output so yeah it's still cheap compared to the

66
00:06:04,700 --> 00:06:09,460
other frontier labs but it's not as cheap as for example muse spark when the pricing is updated

67
00:06:09,460 --> 00:06:15,040
or even deep seek version for flash but across these benchmarks we can see that this model

68
00:06:15,040 --> 00:06:20,000
is better in many areas that they highlighted at the top but then they also have some other areas

69
00:06:20,000 --> 00:06:28,260
like long video understanding which this model excels at 85.4% versus 78.9% for the Terra model

70
00:06:28,260 --> 00:06:34,440
and then the old model was also pretty good at that 84.2%. Then long context performance the

71
00:06:34,440 --> 00:06:43,420
model is at 97% and this model the GPT-6 Terra one it sits at 93.5%. So yeah this model in summary

72
00:06:43,420 --> 00:06:49,240
it is better so it's not all negative but it's not all like you know that positive where you are

73
00:06:49,240 --> 00:06:53,440
super excited for Google DeepMind because as I said they're probably still holding off their

74
00:06:53,440 --> 00:06:58,600
biggest release for Gemini 4 lineup. Now one thing people might have missed in their charts because

75
00:06:58,600 --> 00:07:06,080
these charts sometimes are so messy to read and understand but the 3.7 flash is worse than GPT 5.6

76
00:07:06,080 --> 00:07:11,240
Luna which is the model over here which is achieving a higher score on this benchmark

77
00:07:11,240 --> 00:07:16,480
deep software engineering for about three times the cost. So yeah this is kind of interesting

78
00:07:16,480 --> 00:07:20,960
because yeah the cost for gemini 3.7 flash is a little bit more than what it looks like

79
00:07:20,960 --> 00:07:25,860
now some people are a little bit upset and they're like oh disgraceful google left out

80
00:07:25,860 --> 00:07:30,860
so opus and fable because google is incredibly behind but i think one thing we got to remember

81
00:07:30,860 --> 00:07:33,162
guys is that this model is a flash model

82
00:07:34,162 --> 00:07:40,222
to be a pro model or it's not trying to compete with opus or soul or fable five level models yet

83
00:07:40,222 --> 00:07:45,522
that is probably going to be gemini 4 so when gemini 4 comes out then i think it's okay for

84
00:07:45,522 --> 00:07:50,762
us to criticize them if they don't include Sol, Opus, and Fable in their benchmark charts because

85
00:07:50,762 --> 00:07:55,642
for now I think what they have done is pretty accurate. One lab that I would have liked to see

86
00:07:55,642 --> 00:08:00,062
or one model for example I would have liked to see on that chart would be DeepSeq version 4 Flash

87
00:08:00,062 --> 00:08:04,642
because that would kind of spoil their release because that model is way cheaper compared to

88
00:08:04,642 --> 00:08:11,022
the Gemini 3.7 Flash model lineup. Now this model jumped from number 19 to 8 on the web development

89
00:08:11,022 --> 00:08:16,242
area and we see it over here now and a couple of models that are ahead of it are Opus 5 obviously,

90
00:08:16,422 --> 00:08:23,662
Kimi K3, Quen 3.8 Max, Cloud Opus 5, Grok 4.6 which is a model that came yesterday which was a big win

91
00:08:23,662 --> 00:08:30,222
for SpaceX, Fable 5 and 5.6 Sol. So yeah this model is trying to compete in the web development

92
00:08:30,222 --> 00:08:35,282
but it's still behind all of these models which is you know expected because it's a flash model

93
00:08:35,282 --> 00:08:42,042
it's not really a pro model. Today OpenAI has also launched a waitlist for 5.6 SOL ultra fast mode.

94
00:08:42,202 --> 00:08:47,382
Now this is possible because of the fast chips that they have access to now through Cerebris

95
00:08:47,382 --> 00:08:53,922
and GPT 5.6 SOL with this chip kind of running it is able to achieve an ultra fast mode that

96
00:08:53,922 --> 00:09:01,342
generates up to 750 output tokens per second which is about 14 times faster than the standard mode.

97
00:09:01,342 --> 00:09:07,822
So we're getting a really fast version of GPT 5.6 Soul. And now this is supposed to be used for

98
00:09:07,822 --> 00:09:12,522
live or near production workloads like, you know, real-time voice, support, commerce.

99
00:09:12,762 --> 00:09:19,002
What this allows GPT 5.6 Soul to do is kind of be really fast in critical situations when people

100
00:09:19,002 --> 00:09:24,982
might be interacting with the AI agent, like financial research, security response, support,

101
00:09:25,082 --> 00:09:30,382
I think is going to be a big area where this new ultra fast mode will be kind of implemented.

102
00:09:30,382 --> 00:09:35,962
business and developer agents maybe yeah but i think like support or near production workloads

103
00:09:35,962 --> 00:09:41,322
like real-time voice i see this model really excelling at that and obviously this is still

104
00:09:41,322 --> 00:09:46,562
a waitlist mode we don't know how many people are going to get access to this how successful it is

105
00:09:46,562 --> 00:09:50,982
or what the pricing is i don't know if the pricing has changed because there's no information in their

106
00:09:50,982 --> 00:09:56,622
actual you know blog post but as i mentioned a couple of areas where they mentioned that this

107
00:09:56,622 --> 00:10:01,822
is going to be really important customer support and voice commerce live research and experimentation

108
00:10:01,822 --> 00:10:08,062
financial research and security incident response and reliability but yeah i think this partnership

109
00:10:08,062 --> 00:10:13,642
is going to be important for open ai going forward but as i said 14 times the speed what does that

110
00:10:13,642 --> 00:10:18,902
mean for cost we don't know yet but that's it for today's video make sure you guys are subscribed

111
00:10:18,902 --> 00:10:25,282
to the channel follow our new newsletter as well at universeofai.beehive.com as well subscribe to

112
00:10:25,282 --> 00:10:30,402
the main channel, World of AI, and support us on X by following the universe of AIZ as well.

113
00:10:30,762 --> 00:10:32,722
Until then, I'll see you guys in the next video.
