start	end	text
240	2480	So, looks like we just got flashed by
2480	4640	Google DeepMind [music] once again
4640	6879	because they just dropped Gemini 3.7
6879	9440	Flash today. Yes, their newest model,
9440	11679	which is once again a Flash model, is
11679	14160	here. And this is only after 3 weeks
14160	17279	since Gemini 3.6 Flash. And what they
17279	20000	are saying is that Gemini 3.7 Flash is
20000	22800	their most intelligent workhorse model.
22800	25359	And this is coming 3 weeks after 3.6 six
25359	27840	flash because of developer feedback and
27840	30480	algorithmic innovations. And that is why
30480	32480	they're saying the Gemini 3.7 Flash
32480	34320	model is, you know, out with so much
34320	36160	improvement cuz if we take a look at the
36160	38399	benchmarks, this model is way better
38399	41360	than the Gemini 3.6 Flash. But I think
41360	43440	this is also because we know internally
43440	44879	that they're planning on cancelelling
44879	48160	Gemini 3.5 Pro completely and working on
48160	50480	Gemini 4 at the moment. So maybe the
50480	51920	improvements they made with the Pro
51920	53520	model that they're not going to be
53520	55840	dropping anymore. Maybe they repackaged
55840	58320	it as a flash model because this model
58320	61199	is better than 3.6 and if the timeline
61199	63120	is three weeks plus they're planning on
63120	65519	cancelling the 3.5 Pro, it's quite
65519	67600	likely that is the case. We'll just
67600	69119	ignore that and put that into the side.
69119	70479	We don't know when we're getting a new
70479	72080	Pro model from the Google Deep Mind
72080	74400	team, but this model across many of the
74400	77119	benchmarks we will see that yes, it is
77119	79920	stronger than the 3.6 Flash model. For
79920	81840	example, when it comes to code quality,
81840	84159	production code quality, the model is
84159	86159	producing 43.6%
86159	89680	versus the flash model 34.4%.
89680	93520	And then Sonet 5 is at 42.7%
93520	96799	and the Terra model is at 41.3%.
96799	98799	So, one thing to remember, since this is
98799	100240	a flash model, you're not going to see
100240	102159	them compare this to Opus 5 or any of
102159	104240	the other stronger models. The reason
104240	106240	being cuz this is not their strongest
106240	109119	tier. and the 3.1 Pro model when they
109119	111600	finally decide to upgrade it to Gemini 4
111600	114479	Pro, then we'll see it being compared to
114479	117920	Opus 5 or GPT 5.6 Soul. But anyways, we
117920	119280	can see that this model is an
119280	121520	improvement from 3.6 Flash and I would
121520	124560	say that's basically the biggest result
124560	126479	or the biggest update that we saw with
126479	128560	the new model because it does not really
128560	130720	like change up things a lot is still
130720	132560	like not becoming the number one model
132560	135040	or the number one flash model. Even on
135040	137200	some benchmarks, Deepseek version 4
137200	139920	flash is actually cheaper and more
139920	142160	intelligent than this model. But let's
142160	143680	just take a look at the benchmarks that
143680	145840	they have published. On Long Horizon
145840	148080	software engineering, this model is a
148080	150959	little bit behind GPT 5.6 Terra, which
150959	152959	sits at 69.6%
152959	156400	and then 65.3% is the Flash model. The
156400	160800	older Flash model 3.6 sits at 48.6%.
160800	164000	Sonic 5 sits at 53.8% 8%. And the new
164000	166160	player that is finally being included on
166160	168000	benchmarks, which even the Google
168000	169920	DeepMind team is considering with their
169920	172480	launch, is Muse Spark 1.2, which is
172480	174640	sitting at 54.9%.
174640	177519	So, welcome Meta to the benchmark charts
177519	179360	because now we're starting to see it
179360	181519	appear on more and more benchmarks. And
181519	182800	if we also take a look at web
182800	185360	development, this model is getting a ELO
185360	188159	score of 1588 versus their old model is
188159	190319	1538. So, not a crazy difference, but
190319	191840	compared to everything else out there,
191840	193360	this is number one. When I say
193360	194879	everything else out there, once again,
194879	196720	compared to all the mid-tier models,
196720	198640	this is beating all of them when it
198640	200640	comes to web development. Now, this
200640	202480	model is probably going to be used in
202480	204159	enterprises a lot just because of
204159	206000	Google's footprint in the enterprise
206000	208159	space, but we are seeing this model
208159	211599	achieve on the automation bench 30.4%.
211599	215440	And GPT 5.6 Terra sits at 23.6%.
215440	217599	So, yes, this model is stronger than the
217599	219760	other Flash models out there, but as I
219760	221440	said, they haven't included Deepseek
221440	223599	version for Flash because if they do, in
223599	225519	some areas, that model is actually quite
225519	228480	better than the 3.7 Flash model. Before
228480	230319	we continue, if you're building AI
230319	232480	agents or just messing around with them,
232480	234799	Arcade is worth knowing about. It's the
234799	236799	runtime that lets your agent actually do
236799	238319	things instead of just talking about
238319	240560	them because that's the gap right now.
240560	242799	The models are smart enough. Your agent
242799	244720	can figure out exactly what needs to
244720	246560	happen in your email, your Slack, your
246560	249680	CRM. It just can't go in and do it. And
249680	251519	the reason isn't intelligence, it's
251519	253680	permissions. Something has to prove the
253680	255840	agent is allowed to act on behalf of a
255840	258479	specific person in a specific account.
258479	260239	That's the messy part everyone runs
260239	262479	into, and it's the partit actually
262479	264320	handles for you. So instead of your
264320	266240	agent using one shared login for
266240	268160	everybody, it acts as whoever is
268160	270320	actually signed in with exactly the
270320	272560	access that person has. If they can't
272560	274880	see something, the agent can't either.
274880	276479	And you never have to touch any of that
276479	278960	setup yourself. Then there's the tools.
278960	280800	Arcade has thousands of them already
280800	283120	built for Gmail, Google Drive, Slack,
283120	285520	Notion, Salesforce, most of the apps
285520	287360	people already work in, and they are
287360	289759	built specifically for AI to use. So the
289759	291440	agent gets it right the first time
291440	293040	instead of guessing and failing and
293040	295120	retrying. It also keeps a record of
295120	297520	everything, what the agent did, for who
297520	299440	and where, which matters a lot the
299440	300880	moment other people start using the
300880	302639	thing you built. And it works with
302639	304560	whatever you're already using. Any
304560	306960	model, any framework, cloud, cursor,
306960	309520	chat, GPT, doesn't matter. So you're not
309520	311139	just giving an AI a list of tools and
311139	312720	[music] hoping it works. You're giving
312720	314800	it a place where it can safely take real
314800	317360	actions in real apps. It's free to start
317360	319120	and the link is in the description.
319120	320880	Thank you once again for Arcade for
320880	322960	sponsoring today's video. Now, let's get
322960	325440	back into the video. Now, one thing to
325440	328560	note is that the Gemini 3.7 Flash model
328560	330320	through the end of this year, so end of
330320	333360	2026, they have a cheap pricing model
333360	336160	that they're placing on the model. 75
336160	339840	per 1 million input tokens and $3.75 for
339840	341680	1 million output tokens. So, it's a
341680	343759	competitive price, but this is only for
343759	345919	the next 6 months. Because after those
345919	348240	six months are done, the model's pricing
348240	350160	is actually, you know, a little bit more
350160	351919	expensive. And now they show it at the
351919	354720	bottom over here, you can see that after
354720	358160	starting January of 2027, it will become
358160	362479	$1.50 per input and $7.50 per output. So
362479	364720	yeah, it's still cheap compared to the
364720	366560	other frontier labs, but it's not as
366560	368639	cheap as, for example, MU Spark when the
368639	371199	pricing is updated or even the Deepsee
371199	373280	version for Flash. But across these
373280	375440	benchmarks, we can see that this model
375440	377360	is better in many areas that they
377360	378880	highlighted at the top. But then they
378880	381199	also have some other areas like long
381199	382960	video understanding, which this model
382960	387199	excels at 85.4% versus 78.9%
387199	389120	for the Terra model. And then the old
389120	390880	model was also pretty good at that,
390880	392550	84.2%.
392550	392560	84.2%.
392560	394800	Then long context performance the model
394800	399440	is at 97% and this model the GPT 6 Terra
399440	402160	one it sits at 93.5%.
402160	404160	So yeah this model in summary it is
404160	406479	better so it's not all negative but it's
406479	408639	not all like you know that positive
408639	410479	where you are super excited for Google
410479	412240	Deep Mind because as I said they're
412240	413840	probably still holding off their biggest
413840	416319	release for Gemini 4 lineup. Now, one
416319	418160	thing people might have missed in their
418160	419759	charts because these charts sometimes
419759	422800	are so messy to read and understand, but
422800	426160	the 3.7 flash is worse than GPT 5.6
426160	428160	Luna, which is the model over here,
428160	430560	which is achieving a higher score on
430560	433199	this benchmark deepware engineering for
433199	435520	about three times the cost. So, yeah,
435520	437039	this is kind of interesting because
437039	439599	yeah, the cost for Gemini 3.7 Flash is a
439599	441440	little bit more than what it looks like.
441440	443360	Now, some people are a little bit upset
443360	445120	and they're like, "Oh, disgraceful.
445120	447520	Google left out soul, opus, and fable
447520	449680	because Google is incredibly behind. But
449680	451039	I think one thing we got to remember,
451039	453120	guys, is that this model is a flash
453120	455919	model. It's not trying to be a pro model
455919	458000	or it's not trying to compete with Opus
458000	460639	or Soul or Fable 5 level models yet.
460639	462800	That is probably going to be Gemini 4.
462800	464880	So when Gemini 4 comes out, then I think
464880	467199	it's okay for us to criticize them if
467199	469440	they don't include Soul, Opus, and Fable
469440	471199	in their benchmark charts because for
471199	473199	now, I think what they have done is
473199	475120	pretty accurate. One lab that I would
475120	476639	have liked to see or one model for
476639	478000	example I would have liked to see on
478000	479840	that chart would be Deepseek version 4
479840	481520	Flash because that would kind of spoil
481520	483520	their release because that model is way
483520	486319	cheaper compared to the Gemini 3.7 flash
486319	488800	model lineup. Now this model jumped from
488800	491199	number 19 to 8 on the web development
491199	493599	area and we see it over here now and
493599	495120	couple of models that are ahead of it
495120	498160	are Opus 5 obviously Kim K3 Quinn 3.8
498160	501680	Max Cloud Opus 5 Gro 4.6 6, which is a
501680	503360	model that came yesterday, which was a
503360	506720	big win for SpaceX, Fable 5, and 5.6.
506720	508960	Soul. So, yeah, this model is trying to
508960	511120	compete in the web development, but it's
511120	513120	still behind all of these models, which
513120	515120	is, you know, expected cuz it's a flash
515120	517279	model. It's not really a pro model.
517279	519599	Today, OpenAI has also launched a weight
519599	522479	list for 5.6 so ultra fast mode. Now,
522479	524959	this is possible because of the fast
524959	526640	chips that they have access to now
526640	529839	through Cabus and GPT 5.6 6o with this
529839	532160	chip kind of running it is able to
532160	534000	achieve an ultra fast mode that
534000	537120	generates up to 750 output tokens per
537120	540399	second which is about 14 times faster
540399	542399	than the standard mode. So we're getting
542399	546320	a really fast version of GPT 5.6 so now
546320	548880	this is supposed to be used for live or
548880	550720	near production workloads like you know
550720	553279	real-time voice support commerce what
553279	555839	this allows GPT 5.6 six soul to do is
555839	557680	kind of be really fast in critical
557680	559519	situations when people might be
559519	561680	interacting with the AI agent like
561680	564640	financial research security response
564640	566480	support I think is going to be a big
566480	569360	area where this new ultraast mode will
569360	572000	be kind of implemented business and
572000	573920	developer agents maybe yeah but I think
573920	575600	like support or near production
575600	577760	workloads like real-time voice I see
577760	580160	this model really excelling at that and
580160	582160	obviously this is still a weightless
582160	584160	mode we don't know how many people are
584160	585760	going to get access to this, how
585760	587920	successful it is or what the pricing is.
587920	589440	I don't know if the pricing has changed
589440	591200	because there's no information in their
591200	594000	actual, you know, blog post. But as I
594000	595920	mentioned, couple of areas where they
595920	597360	mentioned that this is going to be
597360	599360	really important. Customer support and
599360	601200	voice, commerce, live research and
601200	603760	experimentation, financial research and
603760	605839	security, incident response and
605839	607760	reliability. But yeah, I think this
607760	609680	partnership is going to be important for
609680	612240	OpenAI going forward. But as I said, 14
612240	614240	times the speed. What does that mean for
614240	616800	cost? We don't know yet. But that's it
616800	618560	for today's video. Make sure you guys
618560	620480	are subscribed to the channel. Follow
620480	622160	our new newsletter as well at
622160	624389	universeai.behive.com
624389	624399	universeai.behive.com
624399	626079	as well as subscribe to the main channel
626079	628560	World of AI and support us on X by
628560	630800	following the Universe of AIZ as well.
630800	632399	Until then, I'll see you guys in the
632399	634560	next
