start	end	text
80	2639	So, Deepseek version 4 Pro is officially
2639	4799	out today. Now, you might be confused
4799	6560	because this model was technically out,
6560	8800	but it was in general availability, but
8800	10400	this is the official release, meaning
10400	12240	that they have trained the model a
12240	14240	little bit more and produced a stronger
14240	16000	version of the model that is available
16000	17520	today. Now, if I were to summarize
17520	19520	today's release in one simple sentence,
19520	21520	it is that the price to performance is
21520	23760	becoming a very, very important thing
23760	25439	because not only do we have a new model
25439	27680	from the Deepseek team, the SpaceX team
27680	30240	also dropped Grock 4.6 six. And both of
30240	32000	these models are competing with the best
32000	34079	Frontier Labs at a fraction of their
34079	37040	cost. Let's start with Gro 4.6. And one
37040	38800	thing I'm going to say is that I'm
38800	40879	genuinely surprised by the SpaceX team.
40879	42480	I'm not trying to glaze them or Elon
42480	44079	Musk or anything. I was just very
44079	46480	critical of this lab. In 2025, they
46480	48320	dropped Grog 4 and other models like
48320	50399	that, but I wasn't really, you know,
50399	52239	mind blown with their performance
52239	53840	because they're pretty they're pretty
53840	56000	subpar compared to any of the other labs
56000	58480	out there. But in 2026, it looks like
58480	60079	things have kind of changed a little
60079	62399	because Grock 4.5 was quite competitive
62399	65600	based off what it cost and Grock 4.6 is
65600	67920	actually not too bad. If we take a look
67920	70000	at the benchmarks, for example, if we
70000	71920	start with the artificial analysis
71920	74240	intelligence index, this model achieves
74240	78159	a 61 and Fable 5 is at 62. Now, what's
78159	79680	really important to remember is that
79680	82159	once again, this model is quite cheap
82159	84400	compared to Fable 5. This is about $2
84400	86960	per million input tokens and $6 per
86960	89439	million output tokens. In Fable 5 sits
89439	92240	at $10 per million input tokens and $50
92240	94400	per million output tokens. And this is
94400	96400	why you start to appreciate this release
96400	98320	a little bit more. It might not beat the
98320	100000	performance of the best models is
100000	102880	matching them. Even GBT 5.6 so which is
102880	104960	a pretty capable model and this is set
104960	108000	at max on the artificial analysis index.
108000	111520	This model achieves 61 and Grock 4.6 61.
111520	113840	So it ties it and it's much cheaper. And
113840	115119	then even on all of these other
115119	117439	benchmarks, for example, the GDP Val
117439	119600	one, it actually beats Fable 5, which is
119600	124159	at 1741. And then GPT 5.6, it's at 1728.
124159	126000	I'm not sure why they didn't choose Opus
126000	127520	5 as well, but I guess they wanted to
127520	129759	choose the quote unquote strongest model
129759	131280	lineup from each lab. And they chose
131280	133760	Fable 5 for Enthropic, which is fair.
133760	135440	And then Deep Software Engineering one,
135440	137520	which is a critical benchmark. This
137520	140640	model doesn't beat Fable 5 or GPT 5.6
140640	142480	six soul, but it gets close to it. It's
142480	146080	65.9 and Fable 5 sits at 70%. But if
146080	147520	you're getting results that are pretty
147520	149200	close and the model is five times
149200	150879	cheaper, I wouldn't be too disappointed
150879	152560	with this result. And then same thing
152560	154879	with Cursor Bench 3.2, the model
154879	156959	achieves 69.9%,
156959	160080	funny number. And Fable 5 sits at 70.5.
160080	162400	So once again, closer. And Grock 4.6
162400	165200	beats GPT 5.6 Soul. Same thing with the
165200	168000	Frontier Code. It gets close to Fable 5,
168000	171040	beats GPT 5.6 Soul. So yeah, this model
171040	173280	is actually available in cursor. So the
173280	175200	partnership with cursor or I guess the
175200	177360	acquisition has really helped SpaceX
177360	179360	make some strides in the AI space this
179360	181840	year and Grock build is something that
181840	183680	you know maybe not a lot of us have been
183680	185599	using so far but it's probably going to
185599	188159	be another platform like codeex or cloud
188159	190319	code that we start to use but obviously
190319	192319	cursor is quite strong as well. So you
192319	194319	have options available for you to use
194319	196319	them in both. And one thing to note is
196319	198560	that they're offering two times usage
198560	200560	inside Grock build and cursor for the
200560	202239	first week. So if you just want to try
202239	204640	it out, see what you feel about it, then
204640	206400	you know it might be worth trying it out
206400	208000	right now cuz you get double the usage
208000	210000	in the first week. Now if we were to
210000	211519	take a look at some of the outputs that
211519	213280	people have been generating with Grock
213280	215519	4.6, what we're looking at right now is
215519	218080	a Falcon 9 booster return sequence
218080	220000	simulation. And this was done in a
220000	223360	single HTML file. And as I said, if you
223360	225360	expected to get this type of output from
225360	228560	Grock in 2025, you would be kind of
228560	230080	surprised because you wouldn't expect
230080	231680	something like this to be generated with
231680	233360	Grock. But now it looks like we have to
233360	235120	start taking the Grock team a little bit
235120	237360	more serious because this output is
237360	239360	quite competitive. And obviously, we're
239360	241120	just looking at a simulation and we're
241120	242879	just basing it off of a visual
242879	244959	representation. But if the model is able
244959	246080	to produce something like this
246080	248400	consistently, then I would expect a lot
248400	250159	of people to start adopting Grock
250159	251840	because number one, it is cheaper than
251840	253680	the other labs at least at the moment
253680	255120	because we don't know if this pricing
255120	257040	strategy is going to be sustainable for
257040	259040	the SpaceX team in the long run. But at
259040	260560	least for now, their models are
260560	262079	definitely cheaper compared to the
262079	264160	others. And as I mentioned, this model
264160	266160	excelled at the artificial analysis
266160	268800	index. This model jumped to probably
268800	272720	number four model. It's tied to GPT 5.6
272720	274000	pretty much similar. So you could say
274000	275759	number three as well, but the models
275759	278639	before that are Fable 5 and Opus 5,
278639	280720	which are only above the model by about
280720	283840	a 1% or a 2% difference. So yeah, even
283840	286000	on this intelligence index, which if
286000	287520	you're not familiar with has nine
287520	289360	evaluations. So on all of these
289360	291440	evaluations, it's kind of matching
291440	293600	almost Fable 5 performance, which is
293600	296240	crazy to see. Before we continue, we
296240	297840	just launched the Universe of AI
297840	299600	newsletter. If you want to stay on top
299600	301600	of AI news without having to hunt for
301600	303680	it, link is in the description. Don't
303680	305520	miss out. And what you see on screen
305520	307680	right now is a racing game that Grock
307680	310400	4.6 build. And based off of this post,
310400	312320	the model took about 1 minute and it was
312320	314400	a fiveword prompt, which was a create a
314400	316639	simple racing game in HTML. So, if
316639	317919	you're able to generate something like
317919	320960	this easily using Grock 4.6, I think a
320960	322880	lot of people will be happy. And this is
322880	325199	a more detailed analysis of what it
325199	327919	costs to run GPT 5.6 6 on the artificial
327919	329919	analysis index and what it produced
329919	332000	meaning the output. Both of these models
332000	334160	if you remember scored 61 on the
334160	336240	artificial analysis index. Now to run
336240	338880	the whole test with GPT 5.6 it cost
338880	342400	about 2.8K and then with Grock 4.6 it
342400	345360	cost about 1.1Kish. And this tells you
345360	346960	that you're getting similar level of
346960	349199	performance at half the cost. So yeah,
349199	351120	this is a big release for the SpaceX
351120	352639	team because they just proved once again
352639	354400	that they are a lab that you seriously
354400	356320	start need to considering, especially in
356320	358560	2026. And I'm going to talk more about
358560	361120	Deep Seek version 4 Pro GA. But
361120	362880	basically what we're seeing today is
362880	364639	that both of these releases kind of
364639	367360	emphasize the fact that performance and
367360	369199	all above that is the price at what
369199	370639	you're getting for that performance is
370639	372479	becoming more and more important for all
372479	374800	users because we see many labs now
374800	377039	focusing on creating the best model at
377039	379600	the cheapest cost. Last year in 2025
379600	381680	most of the labs were just focused on I
381680	384400	would say creating the strongest model.
384400	386080	Yes, cost was important, but I think
386080	387840	most of the times the frontier labs,
387840	390080	meaning OpenAI, Enthropic, were kind of
390080	391759	more lenient on that fact because they
391759	394000	didn't have as strong of a competition.
394000	396479	Intelligent models that are maybe not
396479	399039	always ahead of OpenAI Enthropic, but
399039	400720	match their performance at a fraction of
400720	402800	the cost. So yes, price toerformance
402800	405680	ratio is becoming a critical I would say
405680	408639	indicator in 2026. Now, this is the
408639	410479	updated benchmark chart after the
410479	412240	release of the new model. And one thing
412240	413919	you'll see across the board is that it
413919	415919	matches the top level performance of
415919	417360	many of the models. The one thing
417360	418479	interesting over here is that they
418479	420639	haven't put Opus 5 here for some reason.
420639	422240	There is Fable 5 here that we can
422240	424000	compare this model against. But one
424000	426000	thing you'll notice is that Deep Seek,
426000	429199	remember this model costs 43.5 cents per
429199	432080	million input tokens and 87 cents per
432080	433919	million output tokens. While the other
433919	435680	models all over here are way more
435680	437440	expensive than that. So the first thing
437440	440000	if you look at terminal bench 2.1 the
440000	442319	model scores 87.9.
442319	444960	The older version of the model was 72.1
444960	447120	and the flash version which we got last
447120	450319	week was 82.7. And what's crazy is that
450319	454000	Fable 5 is 88. Yes, 88. So this model is
454000	457360	only.1% behind Fable 5. And then on the
457360	459520	Cyber Gym, which is Cyber Security, the
459520	462240	model actually beats Fable 5. Fable 5
462240	465280	sits at 83.1 while Deepseek version 4
465280	467039	Pro, the new one that we got today, is
467039	469680	at 83.3. Now, if this tells you
469680	471680	something is that this model is once
471680	473759	again really geared at cyber security.
473759	475599	We saw the Flash model also be geared
475599	477280	towards cyber security and becoming a
477280	478800	model that was quite capable in that
478800	480560	area. And we're seeing the same thing
480560	482720	today with the new version 4 Pro which
482720	485840	sits at 83.3 outperforming Fable 5 on
485840	487680	the benchmark. On deep software
487680	489919	engineering, this model is not beating
489919	492639	Fable 5 because Fable 5 sits at 70, but
492639	495440	the model achieves 62.7 which is ahead
495440	498639	of GLM 5.2 and a bit behind Kimmy K3
498639	500879	which sits at 67.5.
500879	502879	But one thing again, this is a big jump
502879	504560	compared to the preview version of the
504560	507360	model which was at 12.8. So yeah, they
507360	508960	have really trained this model and they
508960	510560	have improved it on the back end because
510560	512880	we can clearly see in this deep software
512880	514880	engineering benchmark. And there's a
514880	516719	couple of other benchmarks, but one key
516719	519120	benchmark that this model excels at is
519120	521279	the automation bench where the model
521279	525760	achieves 31.8 and Fable 5 29.1. So once
525760	528000	again, the model outperforms Fable 5.
528000	529920	Now this is a fraction of the cost as I
529920	532480	mentioned. This is 43 cents and 87 for
532480	534640	input and output respectively versus
534640	537519	Fable 5 which is at $10.50.
537519	539279	So yes, if we look at the terminal bench
539279	541279	for example, we are seeing this model
541279	543600	achieve a result which is 0.1 behind the
543600	546160	best model out there Fable 5 and the
546160	549519	pricing is insane to look at $43 versus
549519	553519	$10 87 versus $50 per million input
553519	555360	tokens. So when we turn that into a
555360	558160	fraction, this model is 57 times
558160	560399	cheaper. And earlier last week, I think
560399	562160	I made a video about how DeepSeek
562160	564480	version 4 Pro was going to achieve a
564480	565839	result like this because there was a
565839	567839	investor report that leaked. And when I
567839	569440	saw the pricing where it said that it
569440	570959	was going to match Fable 5 or even
570959	573760	outperform it and be at 57 times cheaper
573760	575440	per cost. When I was reading the
575440	577120	investor report, I didn't really believe
577120	579040	it. But today, it is proven because at
579040	581040	least on the benchmark so far, yes, I'm
581040	582880	just saying the benchmarks, we are
582880	585040	seeing similar level performance. Now,
585040	586959	if this model actually performs like
586959	588959	that in production, it's a little too
588959	590720	early to tell yet. We still would have
590720	592399	to give it a couple of weeks and see how
592399	594399	it's performing in the long run because
594399	596399	sometimes on the release day the models
596399	598480	perform quite good, but over time they
598480	600240	kind of deteriorate in their quality.
600240	602000	So, we hope DeepS doesn't do that.
602000	603680	Historically, they haven't done that,
603680	605600	but let's just see. Just going to put it
605600	607120	out there because right now we're basing
607120	609760	this off of benchmarks. But that's it
609760	611519	for today's video. Make sure you guys
611519	613519	are subscribed to the channel. Follow
613519	615120	our new newsletter as well at
615120	617350	universeofai.behive.com
617350	617360	universeofai.behive.com
617360	619040	as well as subscribe to the main channel
619040	621519	World of AI and support us on X by
621519	623760	following the Universe of AIZ as well.
623760	625360	Until then, I'll see you guys in the
625360	627519	next
