Hi, welcome to another video. So, Meta has finally entered the frontier model race for real this time. They just released MuseSpark 1.1, which is their new multimodal reasoning model built for agentic tasks. This is coming out of their superintelligence labs, and it's also the first model that's available through their brand new Meta model API. Yes, Meta now has a paid API, and Zuckerberg even came back to X after three years just to launch this thing. So, they're definitely taking this seriously. Now what is this model exactly? It's a multimodal reasoning model that combines reasoning, coding, computer use, and multimodal understanding. It has a 1 million token context window, and Meta says it can actively manage that context. It remembers actions, retrieves information from much earlier in the session, and compacts things in a way that keeps the critical steps. It's also trained to orchestrate multi-agent systems, so it can act as a main agent that plans and delegates, or as a sub-agent that sticks to its job and escalates when needed. On the benchmarks, Meta claims it scored 88.1 on MCP Atlas, which tests scaled tool use, while Opus 4.8 and GPT 5.5 are around 80 there. It also leads on JawBench, and a few other tool use benchmarks. But interestingly, on TerminalBench, it only scored 59, while GPT 5.5 is at around 83. So it seems like the model is really strong on tool use and agentic stuff, but not exactly dominant on pure coding. Take these claim numbers with a grain of salt, though, because we'll do our own testing. For the pricing, it's $125 per million input tokens, and $425 per million output tokens, which is actually pretty reasonable. It's cheaper than Sonnet, and you also get $20 in free credits when you sign up for the API preview. So that's good. Now let's put it through my own benchmark, KingBench, and see how it actually performs. I have seven questions here covering UI generation, 3JS, SVG games, hard math, and agentic tasks. Each question is scored out of ten. Let's get right into it. The first question is the elevator simulation. The model has to build a simulation in HTML, CSS, and JS where you can spawn people on different levels. There are three elevators. Each elevator can only take one person, and the people left behind should catch the next one. Every person has a random target floor with a tooltip on hover, and it should be well-animated. Here's what it generated, and it does look pretty nice visually, but the logic is broken. The elevators are not respecting the one-person rule properly, and the cueing behavior just falls apart when you spawn multiple people. The animation is also janky, so this is mostly a fail. I'm giving it a 3 out of 10 here. For comparison, Opus 4.8. Got a full 10 on this one. The second question is the 3JS 3D model of a contact lens case. It should have pronounced L and R caps, and you should be able to click on the caps to open them up. So, here's the result. The case renders, and the L and R are there. The materials actually look quite good, but the click-to-open interaction is only half-working. One cap opens weirdly and clips through the body, so this is a partial pass. I'm giving it a 5 out of 10. The third question is the 3JS folding table with a slider. As you take the slider to the right, the table should unfold, and to the left, it should fold. It should be seamlessly animated, and be 3D. This one is actually decent. The table folds and unfolds with the slider, and the animation is fairly smooth. The geometry gets a bit wonky at the extremes, but overall it works, so this is a 7 out of 10. Not bad at all. The fifth question is the SVG of a panda eating a burger. This tests whether the model can visualize things in code without seeing them. Here's what it made. The panda is recognizable, and there is something that looks like a burger. It's mid. It's not the worst I've seen, but it's not great either. 5 out of 10. The fifth question is the bow and arrow simulator game with 4 targets and a leaderboard based on the lowest time. This one is actually good. The bow mechanics work, the target's register hits, and the leaderboard updates properly. The visuals are also pretty clean, so this is an 8 out of 10. Pretty good. The sixth question is the hard math one. It's a permutation counting problem over a grid of ordered pairs, and the answer should be 20460. This question destroys most models. GPT 5.5, Opus 4.7, and Gemini 3.5 flash all got zero on this. And believe it or not, it got 2460. That's exactly right. So, this is a full 10 out of 10. Really impressive, to be honest. The seventh question is the agentic one. The model has to generate a dataset of facts about pandas, Finitune, Gemma, 2B model on that, all locally, and then give me a web UI where every refresh generates a new panda fact. And it did the whole thing. It generated the dataset, set up the Finituning, ran it, and the web UI works. Every refresh gives me a new panda fact from the Finitune model. So this is a full 10 out of 10 as well. This is where you can see the agentic training paying off. So this is the final chart. And Muse Spark 1 ends up with 48 out of 70, which is 68.57%. That puts it in sixth place on the leaderboard, above GPT 5.6 Terra, Opus 4.7, and Sonnet 5, below Grok 4.5, GPT 5.6 Sol, GLM 5.2, Opus 4.8, and Fable 5 at the top, for a first real frontier attempt from Meta. This is honestly a pretty solid showing. Now, let me tell you some observations from using this model beyond the benchmark, because the scores don't tell the full story. The first thing is that this model is actually really good at UI. Even on the questions where it lost points for logic, the designs it produces look really good. And now that it's on API, you can potentially use it in your workflow with things like Open Design to generate some crazy designs, and then hand that over to something like GPT 5.6 to work on the actual implementation. That combo could be really powerful. The second thing is that, while the model is good at agentic tasks, it is still pretty weird in agentic use. For an example, it doesn't look at other existing files before writing. I generally have multiple sessions of agents running in open code to make sure all my prompts are done. I ask each model to make a new file and not touch the others, but this one just doesn't care. Each session overwrote the others, and none of them even looked at the existing files first. It is really bad in that regard. So if you're running it in parallel setups, be careful, because it will happily clobber your files. The third thing is that the model is still quite good when it works. When it's in its lane, doing agentic tasks, tool use, or that math question, it genuinely delivers. The problem is consistency. You never quite know which version of the model you're going to get. Overall I think this is a really interesting first entry from Meta. It's not going to replace Fable 5 or Opus 4.8 for coding. But the agentic capabilities are real. The pricing is decent. And the UI design sense is a nice surprise. If they fix the file awareness issues and the consistency, the next version could be a serious contender. Overall it's pretty cool. Anyway, let me know your thoughts in the comments. If you liked this video, consider donating through the Super Thanks option, or becoming a member by clicking the Join button. Also, give this video a thumbs up and subscribe to my channel. I'll see you in the next one. Until then, bye. Bye. ! ! ! Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye.