So OpenAI are finally introducing their newest models, GPT 5.6 Sol, GPT 5.6 Terra, and the GPT 5.6 Luna. Now, this is really interesting because this comes at a time where GPT 5.6 isn't being released publicly for us. And I'll dive more into that later, so let's dive into this. They are announcing three new models. GPT 5.6 Sol is going to be the largest model, the Frontier model, or I guess you could say all of your complex tasks. GPT 5.6 Terra is going to be the balanced and efficient model for everyday use. And then, of course, you will have GPT 5.6 Luna for your fast slash high volume tasks. So essentially, they went from max to medium to mini. And essentially, you can think of this as OpenAI changing their names from max to mini and medium to the style of Fable, Opus, Sonnet, and Haiku. Okay, so you can basically think of this as Opus for GPT 5.6 Sol, Sonnet for GPT 5.6 Terra, and Haiku for GPT 5.6 Luna in terms of the sizing. Now, one model that people actually kept missing in this announcement, which is unfortunately buried in there, is that they're actually announcing a new max reasoning effort to give Sol the most time to reason deeply. So for the max model, they're essentially introducing a new Ultra mode that goes beyond the capabilities of a single agent by leveraging sub-agents to accelerate complex work. They're calling this the GPT 5.6 Sol Ultra. So rather than just one model spinning up, it spins up sub-agents to accelerate more complex tasks, which I find to be super interesting. Now, the first benchmark that we're going to be looking at is rather interesting because this is a section that talks about the fact that GPT 5.6 Sol is actually beating Claude Mythos 5. Now, if you don't know what Terminal Bench is, I'm not going to spend too much time diving into, you know, benchmark maxing and all of that crazy stuff. But essentially, Terminal Bench basically is how well do AI agents use a real command line? Instead of just asking the model, can the model answer a question or can it write a code patch? It asks, can the agent sit inside a terminal, run the commands, edit the file, install dependencies, debug errors, and finish a real task end-to-end? So this was built by Stanford. And the reason that this benchmark is one of the first ones for GPT 5.6 Sol is because when we actually look at this, on the left-hand side, you can see that GPT 5.6 Sol Ultra and GPT 5.6 Sol actually does beat Claude Mythos, even though it's a very small amount. It actually does surpass Claude Mythos, which was, of course, this scary, scary model. And it really does surpass Claude Fable 5 by a real big amount. So this is super, super surprising because even the GPT 5.6 Terra model surpasses Claude Fable 5, which means if we're looking at where opening and focusing, it does mean, and I'm going to show you guys why later on in this video, why this is not just like one benchmark with a small improvement, why this is even bigger when you look at the other factors. And yes, I would argue that this chart is probably one of the most confusing charts just because of the way things are. But let's dive into the next one. So the next one that we have here is Exploit Bench. And this one, this benchmark is super interesting for a variety of reasons. So I think the reason that most people were focused on this benchmark is because number one, GPT 5.6 is competitive with Mythos Preview. And it's actually basically so much cheaper than Mythos Preview or Mythos 5, whatever you want to call it. And I think this is really, really important to know because GPT 5.6, and I'm going to show you guys why later, is shaping up to be a model that I think a lot of people are actually going to switch to and use because it seems to be a lot more cost-effective than Mythos 5 and a lot more smarter in many of the areas where you'd want to use it. So GPT 5.6, Sol, Terra, and Luna, they all demonstrate strong improvements in cyber capabilities as they increase their reasoning. And the Exploit Bench benchmark is substantially testing how capable AI agents are at real software exploitation, not just coding trivia. So the Exploit Bench basically gives an agent a real patch vulnerability and asks, can you work your way from I found the bug to can I exploit it? And so that is what the benchmark is doing. Now, I do want to say that I will dive into later, okay, how all of the model restrictions and stuff is going to go on. But I do find it interesting that if you take a look at where Mythos 5 is and where GPT 5.6 is, it stops right just under that Mythos 5 level. Very, very interesting if you ask me. Now, of course, I think this does show that OpenAI are very much still in the race and poised to lead if they continue developing models at this stage. And once again, if we look at Exploit Gym, we don't actually have the Fable benchmarks here, but this basically just asked if the AI agent can turn known bugs into working attacks. And essentially here, you can see the quality. Now, of course, this is a little bit concerning, but at the end of the day, this is where AI is headed. Now, we're going to get into some more interesting things because I think there is some stuff that people did miss in the areas, okay? And I'm talking about the areas of the actual papers called the system card. Now, something that I found to be pretty ridiculous is that one of the things that you're measuring is the hallucinations. So you can see here that there are still an absurd level of hallucinations in AI models. Now, this isn't just to dog on OpenAI's newest announcement, but measuring the model on cases users already flagged as, you know, factual error prone, this is what this benchmark is doing. So the data set here is basically, here are conversations where previous models already messed up. Now, test the new models on those same cursed cases. And OpenAI describes this as not really representative of ordinary production traffic. And of course, hallucinations remain a hard problem, but it does really show you how models are still managing to hallucinate. And the fact that, like, when we previously thought that with more models, with more compute, it would solve this issue, it doesn't really show us that there's any meaningful improvements in these models. In fact, you could argue that the newer models actually hallucinate even more, okay? Which is a little bit concerning. And the thing I would say here is that what this, you know, information should be relaying back to you is that you still need to ground your answer. So if you're going to be using GPT-5.6 Sol, GPT-5.6 Terra, any of these models, in fact, any AI model, this means you need to be still checking your models because the world is messy. There are long tail facts. There are stale data. There's ambiguous questions. There are fake sources, change details, niche topics. These are all of the areas where hallucinations happen. So think about it like this. The model can be good for reasoning but can still reason from a false premise. And you have to understand that, like, this is something that is probably going to persist until we develop new architectures to solve them. Because the way how LLMs are built inherently, hallucinations are essentially a part of the design. So you have to understand that, like, if you're trying to do market research, if you're trying to do medical advice, factual claims, ground your data, ground it in data, have it, you know, with citations and have it checks, okay? Because this is something that I think a lot of people do miss. And a lot of people just take what they're getting from the AI at face value. So ground your AIs. Ensure they're working from docs. Ensure if they're saying something, it's completely cited because a lot of people are going to miss this. Now, you might be wondering, okay, well, this model seems smart. But Fable 5 was designed to work. And this is a model that actually looks like it was just designed for benchmark maxing. Now, this is where we start to get into a really weird area because essentially, METER, the company that essentially measures how long a human task AI can do by itself, actually struggles to measure just how good GPT 5.6 is because the model was acting strangely. The problem is that they tried to estimate GPT 5.6 time horizons, but the estimate got super messy. And why was this? Well, GPT 5.6 had some runs, looked to it like it was cheating, okay? And gaming the benchmark. Now, not necessarily malicious like a human villain, but doing things like making the benchmark pass without really solving the intended task. So they decided different ways to count those weird runs. And they looked at, you know, many different runs. And, you know, counting cheating attempts as failures is around 11 hours. And counting some failed cheating attempts as legitimate failures as success puts it at 270 hours. And so their entire range for this, if you're wondering how long can this thing autonomously work for, the funny part is the range is from 13 hours to 11,400 hours. So their best guess is 71 hours, but the uncertainty is so huge that it could be way lower or way higher, which is one of the times where they say this is the first time that the error bars literally break the chart, which is insane. So, I mean, this is something that it's probably good at long software slash research tasks, but Mita cannot confidently say exactly how good because the model is beyond the benchmarks measuring range, which means we're going to have to get a new benchmark for this. Now, something that you guys are going to want to pay attention to is pricing because everyone, and I know everyone, is, you know, right now feeling the credit crunch because these constant models are being updated in terms of the pricing and usage. So pricing is going to be something that you have to pay attention to. And right now, it looks like that TPT 5.6 Soul is going to be a model that is far more cost-effective. The same for Terra and the same for the other variants of the model. Now, I think that this has been a long time coming because Anthropic has been charging very, very expensive prices for their models. And unfortunately, sometimes those models don't tend to give you back the best responses during peak times of traffic. I promise you guys, you're not crazy. Sometimes there is degraded performance and Anthropic doesn't really say anything. So you can kind of feel like you're being stolen from, in a sense, if you're paying for a service and it just degrades without any real notice. And this is something that I think is really important because if you are going to be building systems, you need reliable models and reliable systems. So with the GPT 5.6 being nearly, you know, I guess you could say 40%, around 40% cheaper than these models, I would say that this is something that you need to pay attention to. And maybe think about putting future systems into the GPT 6 ecosystem because one thing that OpenAI is doing is they are going to be able to have cheaper models because they're building their own chips, building their own full stack. So this is going to be something that you should be paying attention to. Now, one thing that I'm excited for, and this is something that was also left out, something that most people didn't see, is the fact that GPT 5.6 Sol on Cerebrus is actually coming up with 750 tokens per second in July. Now, I don't think people realize just how fast this is because you need to understand that Cerebrus inference is going to be powering the next wave of AI. And when that happens, we are going to have this crazy, crazy intelligence explosion, guys. And what I mean by intelligence explosion is, you're going to be able to do a lot of stuff with that. So what this is, okay, if you don't know what Cerebrus is, it's basically a chip that is designed for LLM inference. And this is how quick it is. So right now the video is paused, but I'm going to show you guys because I want you to understand just how quick it is. Because the moment I hit play, you're probably not going to even realize what happens in the first second. So this is basically the prompt, implement a test in Python. On the left, you can see a model was done already. It was already finished coding in just like three seconds because that was at 2,500 tokens per second. But on the right, you've got Llama for Maverick on an NVIDIA GPU being served with how traditional LLMs work. So think about it like this, guys. Imagine you are using GPT 5.6 and you're able to get that model and it's able to be served for 750 tokens per second. You're going to get all of your responses so, so quickly. So this is something that I think most people are underestimating. If OpenAI does manage to secure this, this is going to be something that I think is going to be completely valuable to all of the users. So I would say look out for this when it does announce because this is going to be something that is super, super useful. If you're tired of your model, you know, waiting for your models to think and think and think and burn through tokens, Cerebrus inference with GPT 5.6 SOL, that is going to be absolutely insane. Now, guys, this is the time where we need to get into the rollout because the rollout is really interesting, okay? The rollout, for now, you don't have access to GPT 5.6 SOL, Terra or Luna. And it says, for now, at the request of the United States government, they're starting with a limited preview among a small group of trusted partners in Codex and the API. So this is something that is pretty crazy, guys. Right now, there is a model that is really smart and it is not even allowed to be, you know, given out for public use. And I could make this video genuinely 40 minutes long talking about all of the different ways that, you know, things have transpired and what's going to happen. But I will upload a video within the next couple of hours explaining everything because I don't want it to be too long. But the gist of this is the fact that like these models now have cyber capabilities that are apparently meaning that it's too dangerous for it to be released to the public out of fears of people using those systems to jailbreak them and then hack critical infrastructure. Now, I've got to be honest, I can't argue with that because that makes sense. But I will say that this is going to pose some interesting consequences for the AI industry as a whole. And I think you guys need to pay attention to this because if you aren't, you're not going to be in the best position to benefit when AI does change in terms of how the access gets distributed. So currently it says at their requests, we're starting with a limited preview for a small group of trusted partners whose participation has been shared with the government before releasing more broadly. And during this preview, we will continue testing and coordinating closely with our partners as we work towards broader availability. And so they're saying that, look, right now we're just releasing this to a bunch of companies. That way we know who exactly is using this and we can ensure that there's complete safety. Now, OpenAI says in this that they don't believe that this kind of, you know, government access should be the long term by default. You know, it keeps the best tools from the users, developers and enterprises and cyber defenders and global partners who need them. But they're taking this short term step because they believe it is the strongest path to broader availability in the coming weeks while they keep the administration at bay and work with them to develop the cyber executive order framework and a repeatable process for future model releases. So this is why earlier I said, are models, are AI companies now going to do benchmark minimizing? So the entire thing started when Mythos was released and Anthropik decided to fear monger and basically say, our AI is so good and go with the marketing angle. It's so good. It can hack everyone and anything. And the US government decided, wait a minute, if that's so good, maybe we shouldn't, you know, allow this just to be free. We're going to actually regulate this. So now you can see that the Mythos preview line, in fact, let me get another image because this is a little bit hard to see. And this tweet here basically explains it really well. It says, goodbye, benchmark maxing. Hello, benchmark minimizing. Mythos is the new bar and you must be very careful to not pass it. So essentially you can see right here on exploit bench, we have Opus 4.8. That's fine. But the Mythos level model, we can see that GPT 5.6 Sol is actually right underneath there. Now, some would argue that this is on purpose because they don't want their model to be, you know, too better than Mythos on the exploit bench because if it is, they know that it's going to get regulated to hell. And then, of course, that just results in a wider range of issues because no one's going to have access to front-end models and OpenAI is just put in a rough position. I mean, it really is going to be interesting to see if models start benchmark minimizing to say that, look, this model isn't better than Mythos 5. There's no need to regulate it. It will be very interesting. Now, Sam Altman made a statement and he says, good news. Of course, these models are smart. You know, we're launching GPT 5.6, yada, yada, yada. But he said, bad news. At the request of the United States government, you know, it's launching today in limited preview instead of the open access we were planning on. We are working with the government to get general availability as fast as we can. And I think it is quite a reasonable rollout models, especially as they reach significant new levels of capability. In this way, it fits with our long-held strategy of iterative deployment. But this isn't quite the process we think is optimal. And they said, they're going to be working with the government to basically get something that works with their safeguards. So at the end of the day, right now, they are saying that, look, they are working with the government to get this model in our hands as quick as possible, but they don't know how long that could take. Which means that the future development of AI model releases is going to be particularly interesting because we won't always have access to the frontier, but perhaps we will know about the frontier and what is being developed. Now, if you're wondering about the actual, you know, I guess you could say deep research in terms of looking into what the models are actually capable of, GPT 5.6 is being treated as a high-risk capability in both cybersecurity and biological and chemical domains. And even for the cheaper Terra and faster Luna versions, OpenAI said this is the first time that smaller and faster models in a family received a high designation in any tracked danger category. So this is super interesting. And then, not just that, but in the cybersecurity area, GPT 5.6 Sol saturated OpenAI's internal cyber challenge set at 96.7%, putting it above the high threshold. And external cyber testers found that high-impact zero days, including one where read-only users could modify and delete data in a widely deployed database. So, GPT 5.6 is clearly, you know, pretty dangerous when it comes to cybersecurity. Now, guys, I'm not trying to fear-monger so the US government can regulate the model. That stuff is already happening. I think what this is showing you is that these models are essentially passing the threshold for where they can now, you know, have those zero-day vulnerabilities that actually affect, maybe not critical infrastructure, but seriously, the economy and companies and that kind of stuff, which is pretty crazy. Now, if you're wondering about the bio, the bio result was just as revealing. High-threshold bio-evaluations crossed the line, while zero out of three critical bio-devaluations crossed it. On virology troubleshooting, GPT 5.6 scored 55%, far above the 31.0% expert performance threshold, and Secure Bio found that GPT 5.6 reached new highs on several expert bio tests, including 68.4% on human pathogen capabilities and 68.3% on world-class bio. And the thing that I want you guys to understand here is that while these models seem like they're plateauing, because, of course, maybe, I don't know, maybe you might not be a bio-researcher, maybe you are. Unless you are in these specific fields like cybersecurity, you know, bio-research, you aren't really going to see those big changes, and these are where the big changes are actually occurring. So, it's really, really interesting to see just how crazy it is. And this is something that nobody, I didn't see anyone talking about this on the timeline. GPT 5.6 Sol's unsettling agent behavior. So, the agent behavior is arguably the most unsettling. GPT 5.6 Sol more often goes beyond user intent when coding, including deleting the wrong virtual machines, claiming unfinished research is verified, and moving cached credentials without permission. And Mita found that GPT 5.6 Sol sometimes tried to game the test instead of just doing the task. The bench row result couldn't be treated as a clean result of raw capability. They're basically trying to say that, look, this model was just acting very strange and we don't even know if we can trust the model's results because it didn't even want to participate properly. And so, what we have here is a situation where nobody's really talking about the fact that these models aren't even behaving. And I find that super, super interesting because you have to understand that these models are really black boxes. So, I mean, honestly, things are starting to get into the weird zone. But let me know what you guys think about this model. And so, it's going to be really interesting to see where things go from the feature from here.