So, looks like we just got flashed by Google DeepMind [music] once again because they just dropped Gemini 3.7 Flash today. Yes, their newest model, which is once again a Flash model, is here. And this is only after 3 weeks since Gemini 3.6 Flash. And what they are saying is that Gemini 3.7 Flash is their most intelligent workhorse model. And this is coming 3 weeks after 3.6 six flash because of developer feedback and algorithmic innovations. And that is why they're saying the Gemini 3.7 Flash model is, you know, out with so much improvement cuz if we take a look at the benchmarks, this model is way better than the Gemini 3.6 Flash. But I think this is also because we know internally that they're planning on cancelelling Gemini 3.5 Pro completely and working on Gemini 4 at the moment. So maybe the improvements they made with the Pro model that they're not going to be dropping anymore. Maybe they repackaged it as a flash model because this model is better than 3.6 and if the timeline is three weeks plus they're planning on cancelling the 3.5 Pro, it's quite likely that is the case. We'll just ignore that and put that into the side. We don't know when we're getting a new Pro model from the Google Deep Mind team, but this model across many of the benchmarks we will see that yes, it is stronger than the 3.6 Flash model. For example, when it comes to code quality, production code quality, the model is producing 43.6% versus the flash model 34.4%. And then Sonet 5 is at 42.7% and the Terra model is at 41.3%. So, one thing to remember, since this is a flash model, you're not going to see them compare this to Opus 5 or any of the other stronger models. The reason being cuz this is not their strongest tier. and the 3.1 Pro model when they finally decide to upgrade it to Gemini 4 Pro, then we'll see it being compared to Opus 5 or GPT 5.6 Soul. But anyways, we can see that this model is an improvement from 3.6 Flash and I would say that's basically the biggest result or the biggest update that we saw with the new model because it does not really like change up things a lot is still like not becoming the number one model or the number one flash model. Even on some benchmarks, Deepseek version 4 flash is actually cheaper and more intelligent than this model. But let's just take a look at the benchmarks that they have published. On Long Horizon software engineering, this model is a little bit behind GPT 5.6 Terra, which sits at 69.6% and then 65.3% is the Flash model. The older Flash model 3.6 sits at 48.6%. Sonic 5 sits at 53.8% 8%. And the new player that is finally being included on benchmarks, which even the Google DeepMind team is considering with their launch, is Muse Spark 1.2, which is sitting at 54.9%. So, welcome Meta to the benchmark charts because now we're starting to see it appear on more and more benchmarks. And if we also take a look at web development, this model is getting a ELO score of 1588 versus their old model is 1538. So, not a crazy difference, but compared to everything else out there, this is number one. When I say everything else out there, once again, compared to all the mid-tier models, this is beating all of them when it comes to web development. Now, this model is probably going to be used in enterprises a lot just because of Google's footprint in the enterprise space, but we are seeing this model achieve on the automation bench 30.4%. And GPT 5.6 Terra sits at 23.6%. So, yes, this model is stronger than the other Flash models out there, but as I said, they haven't included Deepseek version for Flash because if they do, in some areas, that model is actually quite better than the 3.7 Flash model. Before we continue, if you're building AI agents or just messing around with them, Arcade is worth knowing about. It's the runtime that lets your agent actually do things instead of just talking about them because that's the gap right now. The models are smart enough. Your agent can figure out exactly what needs to happen in your email, your Slack, your CRM. It just can't go in and do it. And the reason isn't intelligence, it's permissions. Something has to prove the agent is allowed to act on behalf of a specific person in a specific account. That's the messy part everyone runs into, and it's the partit actually handles for you. So instead of your agent using one shared login for everybody, it acts as whoever is actually signed in with exactly the access that person has. If they can't see something, the agent can't either. And you never have to touch any of that setup yourself. Then there's the tools. Arcade has thousands of them already built for Gmail, Google Drive, Slack, Notion, Salesforce, most of the apps people already work in, and they are built specifically for AI to use. So the agent gets it right the first time instead of guessing and failing and retrying. It also keeps a record of everything, what the agent did, for who and where, which matters a lot the moment other people start using the thing you built. And it works with whatever you're already using. Any model, any framework, cloud, cursor, chat, GPT, doesn't matter. So you're not just giving an AI a list of tools and [music] hoping it works. You're giving it a place where it can safely take real actions in real apps. It's free to start and the link is in the description. Thank you once again for Arcade for sponsoring today's video. Now, let's get back into the video. Now, one thing to note is that the Gemini 3.7 Flash model through the end of this year, so end of 2026, they have a cheap pricing model that they're placing on the model. 75 per 1 million input tokens and $3.75 for 1 million output tokens. So, it's a competitive price, but this is only for the next 6 months. Because after those six months are done, the model's pricing is actually, you know, a little bit more expensive. And now they show it at the bottom over here, you can see that after starting January of 2027, it will become $1.50 per input and $7.50 per output. So yeah, it's still cheap compared to the other frontier labs, but it's not as cheap as, for example, MU Spark when the pricing is updated or even the Deepsee version for Flash. But across these benchmarks, we can see that this model is better in many areas that they highlighted at the top. But then they also have some other areas like long video understanding, which this model excels at 85.4% versus 78.9% for the Terra model. And then the old model was also pretty good at that, 84.2%. 84.2%. Then long context performance the model is at 97% and this model the GPT 6 Terra one it sits at 93.5%. So yeah this model in summary it is better so it's not all negative but it's not all like you know that positive where you are super excited for Google Deep Mind because as I said they're probably still holding off their biggest release for Gemini 4 lineup. Now, one thing people might have missed in their charts because these charts sometimes are so messy to read and understand, but the 3.7 flash is worse than GPT 5.6 Luna, which is the model over here, which is achieving a higher score on this benchmark deepware engineering for about three times the cost. So, yeah, this is kind of interesting because yeah, the cost for Gemini 3.7 Flash is a little bit more than what it looks like. Now, some people are a little bit upset and they're like, "Oh, disgraceful. Google left out soul, opus, and fable because Google is incredibly behind. But I think one thing we got to remember, guys, is that this model is a flash model. It's not trying to be a pro model or it's not trying to compete with Opus or Soul or Fable 5 level models yet. That is probably going to be Gemini 4. So when Gemini 4 comes out, then I think it's okay for us to criticize them if they don't include Soul, Opus, and Fable in their benchmark charts because for now, I think what they have done is pretty accurate. One lab that I would have liked to see or one model for example I would have liked to see on that chart would be Deepseek version 4 Flash because that would kind of spoil their release because that model is way cheaper compared to the Gemini 3.7 flash model lineup. Now this model jumped from number 19 to 8 on the web development area and we see it over here now and couple of models that are ahead of it are Opus 5 obviously Kim K3 Quinn 3.8 Max Cloud Opus 5 Gro 4.6 6, which is a model that came yesterday, which was a big win for SpaceX, Fable 5, and 5.6. Soul. So, yeah, this model is trying to compete in the web development, but it's still behind all of these models, which is, you know, expected cuz it's a flash model. It's not really a pro model. Today, OpenAI has also launched a weight list for 5.6 so ultra fast mode. Now, this is possible because of the fast chips that they have access to now through Cabus and GPT 5.6 6o with this chip kind of running it is able to achieve an ultra fast mode that generates up to 750 output tokens per second which is about 14 times faster than the standard mode. So we're getting a really fast version of GPT 5.6 so now this is supposed to be used for live or near production workloads like you know real-time voice support commerce what this allows GPT 5.6 six soul to do is kind of be really fast in critical situations when people might be interacting with the AI agent like financial research security response support I think is going to be a big area where this new ultraast mode will be kind of implemented business and developer agents maybe yeah but I think like support or near production workloads like real-time voice I see this model really excelling at that and obviously this is still a weightless mode we don't know how many people are going to get access to this, how successful it is or what the pricing is. I don't know if the pricing has changed because there's no information in their actual, you know, blog post. But as I mentioned, couple of areas where they mentioned that this is going to be really important. Customer support and voice, commerce, live research and experimentation, financial research and security, incident response and reliability. But yeah, I think this partnership is going to be important for OpenAI going forward. But as I said, 14 times the speed. What does that mean for cost? We don't know yet. But that's it for today's video. Make sure you guys are subscribed to the channel. Follow our new newsletter as well at universeai.behive.com universeai.behive.com as well as subscribe to the main channel World of AI and support us on X by following the Universe of AIZ as well. Until then, I'll see you guys in the next