start	end	text
0	5580	So looks like we just got flashed by Google DeepMind once again because they just dropped
5580	11820	Gemini 3.7 flash today. Yes, their newest model, which is once again a flash model, is here.
12200	18520	And this is only after three weeks since Gemini 3.6 flash. And what they are saying is that Gemini
18520	25160	3.7 flash is their most intelligent workhorse model. And this is coming three weeks after 3.6
25160	30940	flash because of developer feedback and algorithmic innovations and that is why they're saying that
30940	35960	gemini 3.7 flash model is you know out with so much improvement because if you take a look at
35960	42280	the benchmarks this model is way better than a gemini 3.6 flash but i think this is also because
42280	47920	we know internally that they're planning on canceling gemini 3.5 pro completely and working
47920	52880	on gemini 4 at the moment so maybe the improvements they made with the pro model that they're not
52880	58660	going to be dropping anymore maybe they repackaged it as a flash model because this model is better
58660	65220	than 3.6 and if the timeline is three weeks plus they're planning on canceling the 3.5 pro it's
65220	69520	quite likely that is the case we'll just ignore that and put that into the side we don't know
69520	74300	when we're getting a new pro model from the google deep mine team but this model across many of the
74300	80700	benchmarks we will see that yes it is stronger than the 3.6 flash model for example when it comes
80700	87740	to code quality production code quality the model is producing 43.6 percent versus the flash model
87740	97320	34.4 percent and then sonnet 5 is at 42.7 percent and the terra model is at 41.3 percent so one thing
97320	102080	to remember since this is a flash model you're not going to see them compare this to opus 5 or any of
102080	108200	the other stronger models the reason being because this is not their strongest tier and the 3.1 pro
108200	113780	model when they finally decide to upgrade it to Gemini 4 Pro probably then we'll see it being
113780	119760	compared to Opus 5 or GPT 5.6 Soul but anyways we can see that this model is an improvement from
119760	126200	3.6 Flash and I would say that's basically the biggest result or the biggest update that we saw
126200	131080	with the new model because it does not really like change up things a lot it's still like not
131080	136360	becoming the number one model or the number one Flash model even on some benchmarks DeepSeek
136360	142520	version for flash it's actually cheaper and more intelligent than this model but let's just take a
142520	147500	look at the benchmarks that they have published on long horizon software engineering this model
147500	155320	is a little bit behind gpd 5.6 tera which sits at 69.6 percent and then 65.3 percent is the flash
155320	163640	model the older flash model 3.6 sits at 48.6 percent sonnet 5 sits at 53.8 percent and the
163640	168580	new player that is finally being included on benchmarks which even the google deep mine team
168580	175920	is considering with their launch is musepark 1.2 which is sitting at 54.9 so welcome meta to the
175920	181660	benchmark charts because now we're starting to see it appear on more and more benchmarks and if we
181660	187540	also take a look at web development this model is getting an elo score of 1588 versus their old
187540	192020	model is 1538 so not a crazy difference but compared to everything else out there this is
192020	196460	number one. When I say everything else out there, once again, compared to all the mid-tier models,
196580	201500	this is beating all of them when it comes to web development. Now, this model is probably going to
201500	206180	be used in enterprises a lot just because of Google's footprint in the enterprise space.
206480	213500	But we are seeing this model achieve on the automation bench 30.4%. And GPT 5.6 Terra sits
213500	219840	at 23.6%. So yes, this model is stronger than the other Flash models out there. But as I said,
219840	224760	they haven't included DeepSeq version for Flash because if they do, in some areas that model is
224760	230640	actually quite better than the 3.7 Flash model. Before we continue, if you're building AI agents
230640	235600	or just messing around with them, Arcade is worth knowing about. It's the runtime that lets your
235600	240240	agent actually do things instead of just talking about them. Because that's the gap right now.
240540	245540	The models are smart enough. Your agent can figure out exactly what needs to happen in your email,
245540	251300	your Slack, your CRM. It just can't go in and do it. And the reason isn't intelligence, it's
251300	256740	permissions. Something has to prove the agent is allowed to act on behalf of a specific person
256740	261940	in a specific account. That's the messy part everyone runs into, and it's the part ArcGate
261940	266740	actually handles for you. So instead of your agent using one shared login for everybody,
266740	271860	it acts as whoever is actually signed in with exactly the access that person has.
271860	277140	If they can't see something, the agent can't either, and you never have to touch any of that setup yourself.
277540	278540	Then there's the tools.
278960	284000	Arcade has thousands of them already built for Gmail, Google Drive, Slack, Notion, Salesforce,
284580	289120	most of the apps people already work in, and they are built specifically for AI to use,
289120	293420	so the agent gets it right the first time instead of guessing and failing and retrying.
293880	297960	It also keeps a record of everything, what the agent did for who and where,
298280	301400	which matters a lot the moment other people start using the thing you built.
301400	307640	And it works with whatever you're already using. Any model, any framework, cloud, cursor, chat GPT,
307780	311960	doesn't matter. So you're not just giving an AI a list of tools and hoping it works,
312260	317520	you're giving it a place where it can safely take real actions in real apps. It's free to start and
317520	321780	the link is in the description. Thank you once again for Arcade for sponsoring today's video.
322100	328700	Now let's get back into the video. Now one thing to note is that the Gemini 3.7 flash model through
328700	334420	the end of this year so end of 2026 they have a cheap pricing model that they're placing on the
334420	340880	model 75 cents per 1 million input tokens and 3 dollars and 75 cents for 1 million output tokens
340880	346360	so it's a competitive price but this is only for the next six months because after those six months
346360	351520	are done the model's pricing is actually you know a little bit more expensive and now they show it
351520	358320	at the bottom over here you can see that after starting January of 2027 it will become a dollar
358320	364700	and 50 per input and seven dollars and 50 per output so yeah it's still cheap compared to the
364700	369460	other frontier labs but it's not as cheap as for example muse spark when the pricing is updated
369460	375040	or even deep seek version for flash but across these benchmarks we can see that this model
375040	380000	is better in many areas that they highlighted at the top but then they also have some other areas
380000	388260	like long video understanding which this model excels at 85.4% versus 78.9% for the Terra model
388260	394440	and then the old model was also pretty good at that 84.2%. Then long context performance the
394440	403420	model is at 97% and this model the GPT-6 Terra one it sits at 93.5%. So yeah this model in summary
403420	409240	it is better so it's not all negative but it's not all like you know that positive where you are
409240	413440	super excited for Google DeepMind because as I said they're probably still holding off their
413440	418600	biggest release for Gemini 4 lineup. Now one thing people might have missed in their charts because
418600	426080	these charts sometimes are so messy to read and understand but the 3.7 flash is worse than GPT 5.6
426080	431240	Luna which is the model over here which is achieving a higher score on this benchmark
431240	436480	deep software engineering for about three times the cost. So yeah this is kind of interesting
436480	440960	because yeah the cost for gemini 3.7 flash is a little bit more than what it looks like
440960	445860	now some people are a little bit upset and they're like oh disgraceful google left out
445860	450860	so opus and fable because google is incredibly behind but i think one thing we got to remember
450860	453162	guys is that this model is a flash model
454162	460222	to be a pro model or it's not trying to compete with opus or soul or fable five level models yet
460222	465522	that is probably going to be gemini 4 so when gemini 4 comes out then i think it's okay for
465522	470762	us to criticize them if they don't include Sol, Opus, and Fable in their benchmark charts because
470762	475642	for now I think what they have done is pretty accurate. One lab that I would have liked to see
475642	480062	or one model for example I would have liked to see on that chart would be DeepSeq version 4 Flash
480062	484642	because that would kind of spoil their release because that model is way cheaper compared to
484642	491022	the Gemini 3.7 Flash model lineup. Now this model jumped from number 19 to 8 on the web development
491022	496242	area and we see it over here now and a couple of models that are ahead of it are Opus 5 obviously,
496422	503662	Kimi K3, Quen 3.8 Max, Cloud Opus 5, Grok 4.6 which is a model that came yesterday which was a big win
503662	510222	for SpaceX, Fable 5 and 5.6 Sol. So yeah this model is trying to compete in the web development
510222	515282	but it's still behind all of these models which is you know expected because it's a flash model
515282	522042	it's not really a pro model. Today OpenAI has also launched a waitlist for 5.6 SOL ultra fast mode.
522202	527382	Now this is possible because of the fast chips that they have access to now through Cerebris
527382	533922	and GPT 5.6 SOL with this chip kind of running it is able to achieve an ultra fast mode that
533922	541342	generates up to 750 output tokens per second which is about 14 times faster than the standard mode.
541342	547822	So we're getting a really fast version of GPT 5.6 Soul. And now this is supposed to be used for
547822	552522	live or near production workloads like, you know, real-time voice, support, commerce.
552762	559002	What this allows GPT 5.6 Soul to do is kind of be really fast in critical situations when people
559002	564982	might be interacting with the AI agent, like financial research, security response, support,
565082	570382	I think is going to be a big area where this new ultra fast mode will be kind of implemented.
570382	575962	business and developer agents maybe yeah but i think like support or near production workloads
575962	581322	like real-time voice i see this model really excelling at that and obviously this is still
581322	586562	a waitlist mode we don't know how many people are going to get access to this how successful it is
586562	590982	or what the pricing is i don't know if the pricing has changed because there's no information in their
590982	596622	actual you know blog post but as i mentioned a couple of areas where they mentioned that this
596622	601822	is going to be really important customer support and voice commerce live research and experimentation
601822	608062	financial research and security incident response and reliability but yeah i think this partnership
608062	613642	is going to be important for open ai going forward but as i said 14 times the speed what does that
613642	618902	mean for cost we don't know yet but that's it for today's video make sure you guys are subscribed
618902	625282	to the channel follow our new newsletter as well at universeofai.beehive.com as well subscribe to
625282	630402	the main channel, World of AI, and support us on X by following the universe of AIZ as well.
630762	632722	Until then, I'll see you guys in the next video.
