{
  "text": "So looks like we just got flashed by Google DeepMind once again because they just dropped\nGemini 3.7 flash today. Yes, their newest model, which is once again a flash model, is here.\nAnd this is only after three weeks since Gemini 3.6 flash. And what they are saying is that Gemini\n3.7 flash is their most intelligent workhorse model. And this is coming three weeks after 3.6\nflash because of developer feedback and algorithmic innovations and that is why they're saying that\ngemini 3.7 flash model is you know out with so much improvement because if you take a look at\nthe benchmarks this model is way better than a gemini 3.6 flash but i think this is also because\nwe know internally that they're planning on canceling gemini 3.5 pro completely and working\non gemini 4 at the moment so maybe the improvements they made with the pro model that they're not\ngoing to be dropping anymore maybe they repackaged it as a flash model because this model is better\nthan 3.6 and if the timeline is three weeks plus they're planning on canceling the 3.5 pro it's\nquite likely that is the case we'll just ignore that and put that into the side we don't know\nwhen we're getting a new pro model from the google deep mine team but this model across many of the\nbenchmarks we will see that yes it is stronger than the 3.6 flash model for example when it comes\nto code quality production code quality the model is producing 43.6 percent versus the flash model\n34.4 percent and then sonnet 5 is at 42.7 percent and the terra model is at 41.3 percent so one thing\nto remember since this is a flash model you're not going to see them compare this to opus 5 or any of\nthe other stronger models the reason being because this is not their strongest tier and the 3.1 pro\nmodel when they finally decide to upgrade it to Gemini 4 Pro probably then we'll see it being\ncompared to Opus 5 or GPT 5.6 Soul but anyways we can see that this model is an improvement from\n3.6 Flash and I would say that's basically the biggest result or the biggest update that we saw\nwith the new model because it does not really like change up things a lot it's still like not\nbecoming the number one model or the number one Flash model even on some benchmarks DeepSeek\nversion for flash it's actually cheaper and more intelligent than this model but let's just take a\nlook at the benchmarks that they have published on long horizon software engineering this model\nis a little bit behind gpd 5.6 tera which sits at 69.6 percent and then 65.3 percent is the flash\nmodel the older flash model 3.6 sits at 48.6 percent sonnet 5 sits at 53.8 percent and the\nnew player that is finally being included on benchmarks which even the google deep mine team\nis considering with their launch is musepark 1.2 which is sitting at 54.9 so welcome meta to the\nbenchmark charts because now we're starting to see it appear on more and more benchmarks and if we\nalso take a look at web development this model is getting an elo score of 1588 versus their old\nmodel is 1538 so not a crazy difference but compared to everything else out there this is\nnumber one. When I say everything else out there, once again, compared to all the mid-tier models,\nthis is beating all of them when it comes to web development. Now, this model is probably going to\nbe used in enterprises a lot just because of Google's footprint in the enterprise space.\nBut we are seeing this model achieve on the automation bench 30.4%. And GPT 5.6 Terra sits\nat 23.6%. So yes, this model is stronger than the other Flash models out there. But as I said,\nthey haven't included DeepSeq version for Flash because if they do, in some areas that model is\nactually quite better than the 3.7 Flash model. Before we continue, if you're building AI agents\nor just messing around with them, Arcade is worth knowing about. It's the runtime that lets your\nagent actually do things instead of just talking about them. Because that's the gap right now.\nThe models are smart enough. Your agent can figure out exactly what needs to happen in your email,\nyour Slack, your CRM. It just can't go in and do it. And the reason isn't intelligence, it's\npermissions. Something has to prove the agent is allowed to act on behalf of a specific person\nin a specific account. That's the messy part everyone runs into, and it's the part ArcGate\nactually handles for you. So instead of your agent using one shared login for everybody,\nit acts as whoever is actually signed in with exactly the access that person has.\nIf they can't see something, the agent can't either, and you never have to touch any of that setup yourself.\nThen there's the tools.\nArcade has thousands of them already built for Gmail, Google Drive, Slack, Notion, Salesforce,\nmost of the apps people already work in, and they are built specifically for AI to use,\nso the agent gets it right the first time instead of guessing and failing and retrying.\nIt also keeps a record of everything, what the agent did for who and where,\nwhich matters a lot the moment other people start using the thing you built.\nAnd it works with whatever you're already using. Any model, any framework, cloud, cursor, chat GPT,\ndoesn't matter. So you're not just giving an AI a list of tools and hoping it works,\nyou're giving it a place where it can safely take real actions in real apps. It's free to start and\nthe link is in the description. Thank you once again for Arcade for sponsoring today's video.\nNow let's get back into the video. Now one thing to note is that the Gemini 3.7 flash model through\nthe end of this year so end of 2026 they have a cheap pricing model that they're placing on the\nmodel 75 cents per 1 million input tokens and 3 dollars and 75 cents for 1 million output tokens\nso it's a competitive price but this is only for the next six months because after those six months\nare done the model's pricing is actually you know a little bit more expensive and now they show it\nat the bottom over here you can see that after starting January of 2027 it will become a dollar\nand 50 per input and seven dollars and 50 per output so yeah it's still cheap compared to the\nother frontier labs but it's not as cheap as for example muse spark when the pricing is updated\nor even deep seek version for flash but across these benchmarks we can see that this model\nis better in many areas that they highlighted at the top but then they also have some other areas\nlike long video understanding which this model excels at 85.4% versus 78.9% for the Terra model\nand then the old model was also pretty good at that 84.2%. Then long context performance the\nmodel is at 97% and this model the GPT-6 Terra one it sits at 93.5%. So yeah this model in summary\nit is better so it's not all negative but it's not all like you know that positive where you are\nsuper excited for Google DeepMind because as I said they're probably still holding off their\nbiggest release for Gemini 4 lineup. Now one thing people might have missed in their charts because\nthese charts sometimes are so messy to read and understand but the 3.7 flash is worse than GPT 5.6\nLuna which is the model over here which is achieving a higher score on this benchmark\ndeep software engineering for about three times the cost. So yeah this is kind of interesting\nbecause yeah the cost for gemini 3.7 flash is a little bit more than what it looks like\nnow some people are a little bit upset and they're like oh disgraceful google left out\nso opus and fable because google is incredibly behind but i think one thing we got to remember\nguys is that this model is a flash model\nto be a pro model or it's not trying to compete with opus or soul or fable five level models yet\nthat is probably going to be gemini 4 so when gemini 4 comes out then i think it's okay for\nus to criticize them if they don't include Sol, Opus, and Fable in their benchmark charts because\nfor now I think what they have done is pretty accurate. One lab that I would have liked to see\nor one model for example I would have liked to see on that chart would be DeepSeq version 4 Flash\nbecause that would kind of spoil their release because that model is way cheaper compared to\nthe Gemini 3.7 Flash model lineup. Now this model jumped from number 19 to 8 on the web development\narea and we see it over here now and a couple of models that are ahead of it are Opus 5 obviously,\nKimi K3, Quen 3.8 Max, Cloud Opus 5, Grok 4.6 which is a model that came yesterday which was a big win\nfor SpaceX, Fable 5 and 5.6 Sol. So yeah this model is trying to compete in the web development\nbut it's still behind all of these models which is you know expected because it's a flash model\nit's not really a pro model. Today OpenAI has also launched a waitlist for 5.6 SOL ultra fast mode.\nNow this is possible because of the fast chips that they have access to now through Cerebris\nand GPT 5.6 SOL with this chip kind of running it is able to achieve an ultra fast mode that\ngenerates up to 750 output tokens per second which is about 14 times faster than the standard mode.\nSo we're getting a really fast version of GPT 5.6 Soul. And now this is supposed to be used for\nlive or near production workloads like, you know, real-time voice, support, commerce.\nWhat this allows GPT 5.6 Soul to do is kind of be really fast in critical situations when people\nmight be interacting with the AI agent, like financial research, security response, support,\nI think is going to be a big area where this new ultra fast mode will be kind of implemented.\nbusiness and developer agents maybe yeah but i think like support or near production workloads\nlike real-time voice i see this model really excelling at that and obviously this is still\na waitlist mode we don't know how many people are going to get access to this how successful it is\nor what the pricing is i don't know if the pricing has changed because there's no information in their\nactual you know blog post but as i mentioned a couple of areas where they mentioned that this\nis going to be really important customer support and voice commerce live research and experimentation\nfinancial research and security incident response and reliability but yeah i think this partnership\nis going to be important for open ai going forward but as i said 14 times the speed what does that\nmean for cost we don't know yet but that's it for today's video make sure you guys are subscribed\nto the channel follow our new newsletter as well at universeofai.beehive.com as well subscribe to\nthe main channel, World of AI, and support us on X by following the universe of AIZ as well.\nUntil then, I'll see you guys in the next video.",
  "language": "en",
  "duration": 633.16175,
  "segments": [
    {
      "start": 0.0,
      "end": 5.58,
      "text": "So looks like we just got flashed by Google DeepMind once again because they just dropped",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07356654167175293,
      "no_speech_prob": 1.8395497798640026e-12,
      "compression_ratio": 1.728110599078341
    },
    {
      "start": 5.58,
      "end": 11.82,
      "text": "Gemini 3.7 flash today. Yes, their newest model, which is once again a flash model, is here.",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07356654167175293,
      "no_speech_prob": 1.8395497798640026e-12,
      "compression_ratio": 1.728110599078341
    },
    {
      "start": 12.2,
      "end": 18.52,
      "text": "And this is only after three weeks since Gemini 3.6 flash. And what they are saying is that Gemini",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07356654167175293,
      "no_speech_prob": 1.8395497798640026e-12,
      "compression_ratio": 1.728110599078341
    },
    {
      "start": 18.52,
      "end": 25.16,
      "text": "3.7 flash is their most intelligent workhorse model. And this is coming three weeks after 3.6",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07356654167175293,
      "no_speech_prob": 1.8395497798640026e-12,
      "compression_ratio": 1.728110599078341
    },
    {
      "start": 25.16,
      "end": 30.94,
      "text": "flash because of developer feedback and algorithmic innovations and that is why they're saying that",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03237514158265781,
      "no_speech_prob": 1.201742264728134e-12,
      "compression_ratio": 1.867704280155642
    },
    {
      "start": 30.94,
      "end": 35.96,
      "text": "gemini 3.7 flash model is you know out with so much improvement because if you take a look at",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03237514158265781,
      "no_speech_prob": 1.201742264728134e-12,
      "compression_ratio": 1.867704280155642
    },
    {
      "start": 35.96,
      "end": 42.28,
      "text": "the benchmarks this model is way better than a gemini 3.6 flash but i think this is also because",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03237514158265781,
      "no_speech_prob": 1.201742264728134e-12,
      "compression_ratio": 1.867704280155642
    },
    {
      "start": 42.28,
      "end": 47.92,
      "text": "we know internally that they're planning on canceling gemini 3.5 pro completely and working",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03237514158265781,
      "no_speech_prob": 1.201742264728134e-12,
      "compression_ratio": 1.867704280155642
    },
    {
      "start": 47.92,
      "end": 52.88,
      "text": "on gemini 4 at the moment so maybe the improvements they made with the pro model that they're not",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03237514158265781,
      "no_speech_prob": 1.201742264728134e-12,
      "compression_ratio": 1.867704280155642
    },
    {
      "start": 52.88,
      "end": 58.66,
      "text": "going to be dropping anymore maybe they repackaged it as a flash model because this model is better",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03933087984720866,
      "no_speech_prob": 1.1289260542016177e-12,
      "compression_ratio": 1.7455197132616487
    },
    {
      "start": 58.66,
      "end": 65.22,
      "text": "than 3.6 and if the timeline is three weeks plus they're planning on canceling the 3.5 pro it's",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03933087984720866,
      "no_speech_prob": 1.1289260542016177e-12,
      "compression_ratio": 1.7455197132616487
    },
    {
      "start": 65.22,
      "end": 69.52,
      "text": "quite likely that is the case we'll just ignore that and put that into the side we don't know",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03933087984720866,
      "no_speech_prob": 1.1289260542016177e-12,
      "compression_ratio": 1.7455197132616487
    },
    {
      "start": 69.52,
      "end": 74.3,
      "text": "when we're getting a new pro model from the google deep mine team but this model across many of the",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03933087984720866,
      "no_speech_prob": 1.1289260542016177e-12,
      "compression_ratio": 1.7455197132616487
    },
    {
      "start": 74.3,
      "end": 80.7,
      "text": "benchmarks we will see that yes it is stronger than the 3.6 flash model for example when it comes",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.03933087984720866,
      "no_speech_prob": 1.1289260542016177e-12,
      "compression_ratio": 1.7455197132616487
    },
    {
      "start": 80.7,
      "end": 87.74,
      "text": "to code quality production code quality the model is producing 43.6 percent versus the flash model",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05110936164855957,
      "no_speech_prob": 8.758711529492647e-13,
      "compression_ratio": 1.8785046728971964
    },
    {
      "start": 87.74,
      "end": 97.32,
      "text": "34.4 percent and then sonnet 5 is at 42.7 percent and the terra model is at 41.3 percent so one thing",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05110936164855957,
      "no_speech_prob": 8.758711529492647e-13,
      "compression_ratio": 1.8785046728971964
    },
    {
      "start": 97.32,
      "end": 102.08,
      "text": "to remember since this is a flash model you're not going to see them compare this to opus 5 or any of",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05110936164855957,
      "no_speech_prob": 8.758711529492647e-13,
      "compression_ratio": 1.8785046728971964
    },
    {
      "start": 102.08,
      "end": 108.2,
      "text": "the other stronger models the reason being because this is not their strongest tier and the 3.1 pro",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05110936164855957,
      "no_speech_prob": 8.758711529492647e-13,
      "compression_ratio": 1.8785046728971964
    },
    {
      "start": 108.2,
      "end": 113.78,
      "text": "model when they finally decide to upgrade it to Gemini 4 Pro probably then we'll see it being",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05864520694898522,
      "no_speech_prob": 1.421644052479465e-12,
      "compression_ratio": 1.6596491228070176
    },
    {
      "start": 113.78,
      "end": 119.76,
      "text": "compared to Opus 5 or GPT 5.6 Soul but anyways we can see that this model is an improvement from",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05864520694898522,
      "no_speech_prob": 1.421644052479465e-12,
      "compression_ratio": 1.6596491228070176
    },
    {
      "start": 119.76,
      "end": 126.2,
      "text": "3.6 Flash and I would say that's basically the biggest result or the biggest update that we saw",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05864520694898522,
      "no_speech_prob": 1.421644052479465e-12,
      "compression_ratio": 1.6596491228070176
    },
    {
      "start": 126.2,
      "end": 131.08,
      "text": "with the new model because it does not really like change up things a lot it's still like not",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05864520694898522,
      "no_speech_prob": 1.421644052479465e-12,
      "compression_ratio": 1.6596491228070176
    },
    {
      "start": 131.08,
      "end": 136.36,
      "text": "becoming the number one model or the number one Flash model even on some benchmarks DeepSeek",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05864520694898522,
      "no_speech_prob": 1.421644052479465e-12,
      "compression_ratio": 1.6596491228070176
    },
    {
      "start": 136.36,
      "end": 142.52,
      "text": "version for flash it's actually cheaper and more intelligent than this model but let's just take a",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.06749648463969328,
      "no_speech_prob": 9.21529671148169e-13,
      "compression_ratio": 1.7098214285714286
    },
    {
      "start": 142.52,
      "end": 147.5,
      "text": "look at the benchmarks that they have published on long horizon software engineering this model",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.06749648463969328,
      "no_speech_prob": 9.21529671148169e-13,
      "compression_ratio": 1.7098214285714286
    },
    {
      "start": 147.5,
      "end": 155.32,
      "text": "is a little bit behind gpd 5.6 tera which sits at 69.6 percent and then 65.3 percent is the flash",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.06749648463969328,
      "no_speech_prob": 9.21529671148169e-13,
      "compression_ratio": 1.7098214285714286
    },
    {
      "start": 155.32,
      "end": 163.64,
      "text": "model the older flash model 3.6 sits at 48.6 percent sonnet 5 sits at 53.8 percent and the",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.06749648463969328,
      "no_speech_prob": 9.21529671148169e-13,
      "compression_ratio": 1.7098214285714286
    },
    {
      "start": 163.64,
      "end": 168.58,
      "text": "new player that is finally being included on benchmarks which even the google deep mine team",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.06445804210977817,
      "no_speech_prob": 1.1028875295665541e-12,
      "compression_ratio": 1.6573426573426573
    },
    {
      "start": 168.58,
      "end": 175.92,
      "text": "is considering with their launch is musepark 1.2 which is sitting at 54.9 so welcome meta to the",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.06445804210977817,
      "no_speech_prob": 1.1028875295665541e-12,
      "compression_ratio": 1.6573426573426573
    },
    {
      "start": 175.92,
      "end": 181.66,
      "text": "benchmark charts because now we're starting to see it appear on more and more benchmarks and if we",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.06445804210977817,
      "no_speech_prob": 1.1028875295665541e-12,
      "compression_ratio": 1.6573426573426573
    },
    {
      "start": 181.66,
      "end": 187.54,
      "text": "also take a look at web development this model is getting an elo score of 1588 versus their old",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.06445804210977817,
      "no_speech_prob": 1.1028875295665541e-12,
      "compression_ratio": 1.6573426573426573
    },
    {
      "start": 187.54,
      "end": 192.02,
      "text": "model is 1538 so not a crazy difference but compared to everything else out there this is",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.06445804210977817,
      "no_speech_prob": 1.1028875295665541e-12,
      "compression_ratio": 1.6573426573426573
    },
    {
      "start": 192.02,
      "end": 196.46,
      "text": "number one. When I say everything else out there, once again, compared to all the mid-tier models,",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07064480426882909,
      "no_speech_prob": 1.5796107911275614e-12,
      "compression_ratio": 1.6054421768707483
    },
    {
      "start": 196.58,
      "end": 201.5,
      "text": "this is beating all of them when it comes to web development. Now, this model is probably going to",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07064480426882909,
      "no_speech_prob": 1.5796107911275614e-12,
      "compression_ratio": 1.6054421768707483
    },
    {
      "start": 201.5,
      "end": 206.18,
      "text": "be used in enterprises a lot just because of Google's footprint in the enterprise space.",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07064480426882909,
      "no_speech_prob": 1.5796107911275614e-12,
      "compression_ratio": 1.6054421768707483
    },
    {
      "start": 206.48,
      "end": 213.5,
      "text": "But we are seeing this model achieve on the automation bench 30.4%. And GPT 5.6 Terra sits",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07064480426882909,
      "no_speech_prob": 1.5796107911275614e-12,
      "compression_ratio": 1.6054421768707483
    },
    {
      "start": 213.5,
      "end": 219.84,
      "text": "at 23.6%. So yes, this model is stronger than the other Flash models out there. But as I said,",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07064480426882909,
      "no_speech_prob": 1.5796107911275614e-12,
      "compression_ratio": 1.6054421768707483
    },
    {
      "start": 219.84,
      "end": 224.76,
      "text": "they haven't included DeepSeq version for Flash because if they do, in some areas that model is",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07140385477166426,
      "no_speech_prob": 9.61982013145124e-13,
      "compression_ratio": 1.6317567567567568
    },
    {
      "start": 224.76,
      "end": 230.64,
      "text": "actually quite better than the 3.7 Flash model. Before we continue, if you're building AI agents",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07140385477166426,
      "no_speech_prob": 9.61982013145124e-13,
      "compression_ratio": 1.6317567567567568
    },
    {
      "start": 230.64,
      "end": 235.6,
      "text": "or just messing around with them, Arcade is worth knowing about. It's the runtime that lets your",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07140385477166426,
      "no_speech_prob": 9.61982013145124e-13,
      "compression_ratio": 1.6317567567567568
    },
    {
      "start": 235.6,
      "end": 240.24,
      "text": "agent actually do things instead of just talking about them. Because that's the gap right now.",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07140385477166426,
      "no_speech_prob": 9.61982013145124e-13,
      "compression_ratio": 1.6317567567567568
    },
    {
      "start": 240.54,
      "end": 245.54,
      "text": "The models are smart enough. Your agent can figure out exactly what needs to happen in your email,",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07140385477166426,
      "no_speech_prob": 9.61982013145124e-13,
      "compression_ratio": 1.6317567567567568
    },
    {
      "start": 245.54,
      "end": 251.3,
      "text": "your Slack, your CRM. It just can't go in and do it. And the reason isn't intelligence, it's",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.0749070794732721,
      "no_speech_prob": 1.216037361952138e-12,
      "compression_ratio": 1.6568265682656826
    },
    {
      "start": 251.3,
      "end": 256.74,
      "text": "permissions. Something has to prove the agent is allowed to act on behalf of a specific person",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.0749070794732721,
      "no_speech_prob": 1.216037361952138e-12,
      "compression_ratio": 1.6568265682656826
    },
    {
      "start": 256.74,
      "end": 261.94,
      "text": "in a specific account. That's the messy part everyone runs into, and it's the part ArcGate",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.0749070794732721,
      "no_speech_prob": 1.216037361952138e-12,
      "compression_ratio": 1.6568265682656826
    },
    {
      "start": 261.94,
      "end": 266.74,
      "text": "actually handles for you. So instead of your agent using one shared login for everybody,",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.0749070794732721,
      "no_speech_prob": 1.216037361952138e-12,
      "compression_ratio": 1.6568265682656826
    },
    {
      "start": 266.74,
      "end": 271.86,
      "text": "it acts as whoever is actually signed in with exactly the access that person has.",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.0749070794732721,
      "no_speech_prob": 1.216037361952138e-12,
      "compression_ratio": 1.6568265682656826
    },
    {
      "start": 271.86,
      "end": 277.14,
      "text": "If they can't see something, the agent can't either, and you never have to touch any of that setup yourself.",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.16463276939670535,
      "no_speech_prob": 1.0523920060054315e-12,
      "compression_ratio": 1.7375
    },
    {
      "start": 277.54,
      "end": 278.54,
      "text": "Then there's the tools.",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.16463276939670535,
      "no_speech_prob": 1.0523920060054315e-12,
      "compression_ratio": 1.7375
    },
    {
      "start": 278.96,
      "end": 284.0,
      "text": "Arcade has thousands of them already built for Gmail, Google Drive, Slack, Notion, Salesforce,",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.16463276939670535,
      "no_speech_prob": 1.0523920060054315e-12,
      "compression_ratio": 1.7375
    },
    {
      "start": 284.58,
      "end": 289.12,
      "text": "most of the apps people already work in, and they are built specifically for AI to use,",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.16463276939670535,
      "no_speech_prob": 1.0523920060054315e-12,
      "compression_ratio": 1.7375
    },
    {
      "start": 289.12,
      "end": 293.42,
      "text": "so the agent gets it right the first time instead of guessing and failing and retrying.",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.16463276939670535,
      "no_speech_prob": 1.0523920060054315e-12,
      "compression_ratio": 1.7375
    },
    {
      "start": 293.88,
      "end": 297.96,
      "text": "It also keeps a record of everything, what the agent did for who and where,",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.16463276939670535,
      "no_speech_prob": 1.0523920060054315e-12,
      "compression_ratio": 1.7375
    },
    {
      "start": 298.28,
      "end": 301.4,
      "text": "which matters a lot the moment other people start using the thing you built.",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.16463276939670535,
      "no_speech_prob": 1.0523920060054315e-12,
      "compression_ratio": 1.7375
    },
    {
      "start": 301.4,
      "end": 307.64,
      "text": "And it works with whatever you're already using. Any model, any framework, cloud, cursor, chat GPT,",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.10897723388671875,
      "no_speech_prob": 1.8981762776176803e-12,
      "compression_ratio": 1.614864864864865
    },
    {
      "start": 307.78,
      "end": 311.96,
      "text": "doesn't matter. So you're not just giving an AI a list of tools and hoping it works,",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.10897723388671875,
      "no_speech_prob": 1.8981762776176803e-12,
      "compression_ratio": 1.614864864864865
    },
    {
      "start": 312.26,
      "end": 317.52,
      "text": "you're giving it a place where it can safely take real actions in real apps. It's free to start and",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.10897723388671875,
      "no_speech_prob": 1.8981762776176803e-12,
      "compression_ratio": 1.614864864864865
    },
    {
      "start": 317.52,
      "end": 321.78,
      "text": "the link is in the description. Thank you once again for Arcade for sponsoring today's video.",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.10897723388671875,
      "no_speech_prob": 1.8981762776176803e-12,
      "compression_ratio": 1.614864864864865
    },
    {
      "start": 322.1,
      "end": 328.7,
      "text": "Now let's get back into the video. Now one thing to note is that the Gemini 3.7 flash model through",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.10897723388671875,
      "no_speech_prob": 1.8981762776176803e-12,
      "compression_ratio": 1.614864864864865
    },
    {
      "start": 328.7,
      "end": 334.42,
      "text": "the end of this year so end of 2026 they have a cheap pricing model that they're placing on the",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.047414999741774336,
      "no_speech_prob": 1.102784096679299e-12,
      "compression_ratio": 1.7455357142857142
    },
    {
      "start": 334.42,
      "end": 340.88,
      "text": "model 75 cents per 1 million input tokens and 3 dollars and 75 cents for 1 million output tokens",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.047414999741774336,
      "no_speech_prob": 1.102784096679299e-12,
      "compression_ratio": 1.7455357142857142
    },
    {
      "start": 340.88,
      "end": 346.36,
      "text": "so it's a competitive price but this is only for the next six months because after those six months",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.047414999741774336,
      "no_speech_prob": 1.102784096679299e-12,
      "compression_ratio": 1.7455357142857142
    },
    {
      "start": 346.36,
      "end": 351.52,
      "text": "are done the model's pricing is actually you know a little bit more expensive and now they show it",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.047414999741774336,
      "no_speech_prob": 1.102784096679299e-12,
      "compression_ratio": 1.7455357142857142
    },
    {
      "start": 351.52,
      "end": 358.32,
      "text": "at the bottom over here you can see that after starting January of 2027 it will become a dollar",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07676051487432461,
      "no_speech_prob": 1.7829340434594165e-12,
      "compression_ratio": 1.72992700729927
    },
    {
      "start": 358.32,
      "end": 364.7,
      "text": "and 50 per input and seven dollars and 50 per output so yeah it's still cheap compared to the",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07676051487432461,
      "no_speech_prob": 1.7829340434594165e-12,
      "compression_ratio": 1.72992700729927
    },
    {
      "start": 364.7,
      "end": 369.46,
      "text": "other frontier labs but it's not as cheap as for example muse spark when the pricing is updated",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07676051487432461,
      "no_speech_prob": 1.7829340434594165e-12,
      "compression_ratio": 1.72992700729927
    },
    {
      "start": 369.46,
      "end": 375.04,
      "text": "or even deep seek version for flash but across these benchmarks we can see that this model",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07676051487432461,
      "no_speech_prob": 1.7829340434594165e-12,
      "compression_ratio": 1.72992700729927
    },
    {
      "start": 375.04,
      "end": 380.0,
      "text": "is better in many areas that they highlighted at the top but then they also have some other areas",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07676051487432461,
      "no_speech_prob": 1.7829340434594165e-12,
      "compression_ratio": 1.72992700729927
    },
    {
      "start": 380.0,
      "end": 388.26,
      "text": "like long video understanding which this model excels at 85.4% versus 78.9% for the Terra model",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.047229664996989724,
      "no_speech_prob": 9.072133777716929e-13,
      "compression_ratio": 1.620253164556962
    },
    {
      "start": 388.26,
      "end": 394.44,
      "text": "and then the old model was also pretty good at that 84.2%. Then long context performance the",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.047229664996989724,
      "no_speech_prob": 9.072133777716929e-13,
      "compression_ratio": 1.620253164556962
    },
    {
      "start": 394.44,
      "end": 403.42,
      "text": "model is at 97% and this model the GPT-6 Terra one it sits at 93.5%. So yeah this model in summary",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.047229664996989724,
      "no_speech_prob": 9.072133777716929e-13,
      "compression_ratio": 1.620253164556962
    },
    {
      "start": 403.42,
      "end": 409.24,
      "text": "it is better so it's not all negative but it's not all like you know that positive where you are",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.047229664996989724,
      "no_speech_prob": 9.072133777716929e-13,
      "compression_ratio": 1.620253164556962
    },
    {
      "start": 409.24,
      "end": 413.44,
      "text": "super excited for Google DeepMind because as I said they're probably still holding off their",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05204198945243404,
      "no_speech_prob": 1.4440298999954249e-12,
      "compression_ratio": 1.580536912751678
    },
    {
      "start": 413.44,
      "end": 418.6,
      "text": "biggest release for Gemini 4 lineup. Now one thing people might have missed in their charts because",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05204198945243404,
      "no_speech_prob": 1.4440298999954249e-12,
      "compression_ratio": 1.580536912751678
    },
    {
      "start": 418.6,
      "end": 426.08,
      "text": "these charts sometimes are so messy to read and understand but the 3.7 flash is worse than GPT 5.6",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05204198945243404,
      "no_speech_prob": 1.4440298999954249e-12,
      "compression_ratio": 1.580536912751678
    },
    {
      "start": 426.08,
      "end": 431.24,
      "text": "Luna which is the model over here which is achieving a higher score on this benchmark",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05204198945243404,
      "no_speech_prob": 1.4440298999954249e-12,
      "compression_ratio": 1.580536912751678
    },
    {
      "start": 431.24,
      "end": 436.48,
      "text": "deep software engineering for about three times the cost. So yeah this is kind of interesting",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.05204198945243404,
      "no_speech_prob": 1.4440298999954249e-12,
      "compression_ratio": 1.580536912751678
    },
    {
      "start": 436.48,
      "end": 440.96,
      "text": "because yeah the cost for gemini 3.7 flash is a little bit more than what it looks like",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07750610252479453,
      "no_speech_prob": 1.1202418116404433e-12,
      "compression_ratio": 1.6062176165803108
    },
    {
      "start": 440.96,
      "end": 445.86,
      "text": "now some people are a little bit upset and they're like oh disgraceful google left out",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07750610252479453,
      "no_speech_prob": 1.1202418116404433e-12,
      "compression_ratio": 1.6062176165803108
    },
    {
      "start": 445.86,
      "end": 450.86,
      "text": "so opus and fable because google is incredibly behind but i think one thing we got to remember",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07750610252479453,
      "no_speech_prob": 1.1202418116404433e-12,
      "compression_ratio": 1.6062176165803108
    },
    {
      "start": 450.86,
      "end": 453.162,
      "text": "guys is that this model is a flash model",
      "chunk": 0,
      "language": "en",
      "avg_logprob": -0.07750610252479453,
      "no_speech_prob": 1.1202418116404433e-12,
      "compression_ratio": 1.6062176165803108
    },
    {
      "start": 454.162,
      "end": 460.222,
      "text": "to be a pro model or it's not trying to compete with opus or soul or fable five level models yet",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.03837726529964731,
      "no_speech_prob": 1.748642463467176e-12,
      "compression_ratio": 1.8604651162790697
    },
    {
      "start": 460.222,
      "end": 465.522,
      "text": "that is probably going to be gemini 4 so when gemini 4 comes out then i think it's okay for",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.03837726529964731,
      "no_speech_prob": 1.748642463467176e-12,
      "compression_ratio": 1.8604651162790697
    },
    {
      "start": 465.522,
      "end": 470.762,
      "text": "us to criticize them if they don't include Sol, Opus, and Fable in their benchmark charts because",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.054536696138053106,
      "no_speech_prob": 1.0984870782090872e-12,
      "compression_ratio": 1.694736842105263
    },
    {
      "start": 470.762,
      "end": 475.642,
      "text": "for now I think what they have done is pretty accurate. One lab that I would have liked to see",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.054536696138053106,
      "no_speech_prob": 1.0984870782090872e-12,
      "compression_ratio": 1.694736842105263
    },
    {
      "start": 475.642,
      "end": 480.062,
      "text": "or one model for example I would have liked to see on that chart would be DeepSeq version 4 Flash",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.054536696138053106,
      "no_speech_prob": 1.0984870782090872e-12,
      "compression_ratio": 1.694736842105263
    },
    {
      "start": 480.062,
      "end": 484.642,
      "text": "because that would kind of spoil their release because that model is way cheaper compared to",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.054536696138053106,
      "no_speech_prob": 1.0984870782090872e-12,
      "compression_ratio": 1.694736842105263
    },
    {
      "start": 484.642,
      "end": 491.022,
      "text": "the Gemini 3.7 Flash model lineup. Now this model jumped from number 19 to 8 on the web development",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.054536696138053106,
      "no_speech_prob": 1.0984870782090872e-12,
      "compression_ratio": 1.694736842105263
    },
    {
      "start": 491.022,
      "end": 496.242,
      "text": "area and we see it over here now and a couple of models that are ahead of it are Opus 5 obviously,",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.11288458108901978,
      "no_speech_prob": 1.3045795997299048e-12,
      "compression_ratio": 1.5291828793774318
    },
    {
      "start": 496.422,
      "end": 503.662,
      "text": "Kimi K3, Quen 3.8 Max, Cloud Opus 5, Grok 4.6 which is a model that came yesterday which was a big win",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.11288458108901978,
      "no_speech_prob": 1.3045795997299048e-12,
      "compression_ratio": 1.5291828793774318
    },
    {
      "start": 503.662,
      "end": 510.222,
      "text": "for SpaceX, Fable 5 and 5.6 Sol. So yeah this model is trying to compete in the web development",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.11288458108901978,
      "no_speech_prob": 1.3045795997299048e-12,
      "compression_ratio": 1.5291828793774318
    },
    {
      "start": 510.222,
      "end": 515.282,
      "text": "but it's still behind all of these models which is you know expected because it's a flash model",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.11288458108901978,
      "no_speech_prob": 1.3045795997299048e-12,
      "compression_ratio": 1.5291828793774318
    },
    {
      "start": 515.282,
      "end": 522.042,
      "text": "it's not really a pro model. Today OpenAI has also launched a waitlist for 5.6 SOL ultra fast mode.",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.07899854621108697,
      "no_speech_prob": 1.2546096973820031e-12,
      "compression_ratio": 1.5461847389558232
    },
    {
      "start": 522.202,
      "end": 527.382,
      "text": "Now this is possible because of the fast chips that they have access to now through Cerebris",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.07899854621108697,
      "no_speech_prob": 1.2546096973820031e-12,
      "compression_ratio": 1.5461847389558232
    },
    {
      "start": 527.382,
      "end": 533.922,
      "text": "and GPT 5.6 SOL with this chip kind of running it is able to achieve an ultra fast mode that",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.07899854621108697,
      "no_speech_prob": 1.2546096973820031e-12,
      "compression_ratio": 1.5461847389558232
    },
    {
      "start": 533.922,
      "end": 541.342,
      "text": "generates up to 750 output tokens per second which is about 14 times faster than the standard mode.",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.07899854621108697,
      "no_speech_prob": 1.2546096973820031e-12,
      "compression_ratio": 1.5461847389558232
    },
    {
      "start": 541.342,
      "end": 547.822,
      "text": "So we're getting a really fast version of GPT 5.6 Soul. And now this is supposed to be used for",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.12016314473645441,
      "no_speech_prob": 1.1925345693580836e-12,
      "compression_ratio": 1.6258741258741258
    },
    {
      "start": 547.822,
      "end": 552.522,
      "text": "live or near production workloads like, you know, real-time voice, support, commerce.",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.12016314473645441,
      "no_speech_prob": 1.1925345693580836e-12,
      "compression_ratio": 1.6258741258741258
    },
    {
      "start": 552.762,
      "end": 559.002,
      "text": "What this allows GPT 5.6 Soul to do is kind of be really fast in critical situations when people",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.12016314473645441,
      "no_speech_prob": 1.1925345693580836e-12,
      "compression_ratio": 1.6258741258741258
    },
    {
      "start": 559.002,
      "end": 564.982,
      "text": "might be interacting with the AI agent, like financial research, security response, support,",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.12016314473645441,
      "no_speech_prob": 1.1925345693580836e-12,
      "compression_ratio": 1.6258741258741258
    },
    {
      "start": 565.082,
      "end": 570.382,
      "text": "I think is going to be a big area where this new ultra fast mode will be kind of implemented.",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.12016314473645441,
      "no_speech_prob": 1.1925345693580836e-12,
      "compression_ratio": 1.6258741258741258
    },
    {
      "start": 570.382,
      "end": 575.962,
      "text": "business and developer agents maybe yeah but i think like support or near production workloads",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.030362447102864582,
      "no_speech_prob": 1.6427473530783443e-12,
      "compression_ratio": 1.749090909090909
    },
    {
      "start": 575.962,
      "end": 581.322,
      "text": "like real-time voice i see this model really excelling at that and obviously this is still",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.030362447102864582,
      "no_speech_prob": 1.6427473530783443e-12,
      "compression_ratio": 1.749090909090909
    },
    {
      "start": 581.322,
      "end": 586.562,
      "text": "a waitlist mode we don't know how many people are going to get access to this how successful it is",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.030362447102864582,
      "no_speech_prob": 1.6427473530783443e-12,
      "compression_ratio": 1.749090909090909
    },
    {
      "start": 586.562,
      "end": 590.982,
      "text": "or what the pricing is i don't know if the pricing has changed because there's no information in their",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.030362447102864582,
      "no_speech_prob": 1.6427473530783443e-12,
      "compression_ratio": 1.749090909090909
    },
    {
      "start": 590.982,
      "end": 596.622,
      "text": "actual you know blog post but as i mentioned a couple of areas where they mentioned that this",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.030362447102864582,
      "no_speech_prob": 1.6427473530783443e-12,
      "compression_ratio": 1.749090909090909
    },
    {
      "start": 596.622,
      "end": 601.822,
      "text": "is going to be really important customer support and voice commerce live research and experimentation",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.07032611882575204,
      "no_speech_prob": 2.289275106287514e-12,
      "compression_ratio": 1.7464788732394365
    },
    {
      "start": 601.822,
      "end": 608.062,
      "text": "financial research and security incident response and reliability but yeah i think this partnership",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.07032611882575204,
      "no_speech_prob": 2.289275106287514e-12,
      "compression_ratio": 1.7464788732394365
    },
    {
      "start": 608.062,
      "end": 613.642,
      "text": "is going to be important for open ai going forward but as i said 14 times the speed what does that",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.07032611882575204,
      "no_speech_prob": 2.289275106287514e-12,
      "compression_ratio": 1.7464788732394365
    },
    {
      "start": 613.642,
      "end": 618.902,
      "text": "mean for cost we don't know yet but that's it for today's video make sure you guys are subscribed",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.07032611882575204,
      "no_speech_prob": 2.289275106287514e-12,
      "compression_ratio": 1.7464788732394365
    },
    {
      "start": 618.902,
      "end": 625.282,
      "text": "to the channel follow our new newsletter as well at universeofai.beehive.com as well subscribe to",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.07032611882575204,
      "no_speech_prob": 2.289275106287514e-12,
      "compression_ratio": 1.7464788732394365
    },
    {
      "start": 625.282,
      "end": 630.402,
      "text": "the main channel, World of AI, and support us on X by following the universe of AIZ as well.",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.12289900896025867,
      "no_speech_prob": 4.632463575064694e-13,
      "compression_ratio": 1.194915254237288
    },
    {
      "start": 630.762,
      "end": 632.722,
      "text": "Until then, I'll see you guys in the next video.",
      "chunk": 1,
      "language": "en",
      "avg_logprob": -0.12289900896025867,
      "no_speech_prob": 4.632463575064694e-13,
      "compression_ratio": 1.194915254237288
    }
  ],
  "segmented_transcription": {
    "status": "SEGMENTED_TRANSCRIPT_OK",
    "created_at": "2026-08-14T13:39:22",
    "manifest": "/Users/sagawa/AI_WORK/video_notes/url/20260814_133848__Google_Shipped_Gemini_3.7_Flash_And_OpenAI_Made_GPT-5.6_14x_Faster/segmented_work/chunk_manifest.json",
    "audio": "/Users/sagawa/AI_WORK/video_notes/url/20260814_133848__Google_Shipped_Gemini_3.7_Flash_And_OpenAI_Made_GPT-5.6_14x_Faster/audio_16k_mono.wav",
    "audio_sha256": "218a01a1ea2ce64f9e1cffdb1efe75d13d3f8b298d90cca4bd896b82760b12d3",
    "audio_duration": 633.16175,
    "model": "mlx-community/whisper-large-v3-turbo",
    "language_hint": "en",
    "condition_on_previous_text": false,
    "chunk_count": 2,
    "resumed_chunks": 0,
    "merged_segment_count": 113,
    "final_segment_end": 632.722,
    "global_repeat_run": 1,
    "global_repeat_text": "So looks like we just got flashed by Google DeepMind once again because they just dropped",
    "tail_repeat_run": 1,
    "tail_repeat_text": "us to criticize them if they don't include Sol, Opus, and Fable in their benchmark charts because",
    "hallucination_region_count": 0,
    "hallucination_regions": [],
    "qc_ok": true,
    "qc_errors": []
  }
}