1
00:00:00,000 --> 00:00:05,740
So DeepSeek version 4 Pro is officially out today. Now you might be confused because this model was

2
00:00:05,740 --> 00:00:10,680
technically out but it was in general availability but this is the official release meaning that they

3
00:00:10,680 --> 00:00:15,200
have trained the model a little bit more and produced a stronger version of the model that

4
00:00:15,200 --> 00:00:19,900
is available today. Now if I were to summarize today's release in one simple sentence it is that

5
00:00:19,900 --> 00:00:25,000
the price to performance is becoming a very very important thing because not only do we have a new

6
00:00:25,000 --> 00:00:30,660
model from the DeepSeek team, the SpaceX team also dropped Grok 4.6 and both of these models

7
00:00:30,660 --> 00:00:36,520
are competing with the best frontier labs at a fraction of their cost. Let's start with Grok 4.6

8
00:00:36,520 --> 00:00:41,180
and one thing I'm going to say is that I'm genuinely surprised by the SpaceX team. I'm not

9
00:00:41,180 --> 00:00:45,860
trying to glaze them or Elon Musk or anything. I was just very critical of this lab. In 2025,

10
00:00:46,240 --> 00:00:51,380
they dropped Grok 4 and other models like that but I wasn't really you know mind blown with their

11
00:00:51,380 --> 00:00:56,240
performance because they're pretty they're pretty subpar compared to any of the other labs out there

12
00:00:56,240 --> 00:01:01,680
but in 2026 it looks like things have kind of changed a little because grok 4.5 was quite

13
00:01:01,680 --> 00:01:08,280
competitive based off what it cost and grok 4.6 is actually not too bad if we take a look at the

14
00:01:08,280 --> 00:01:13,300
benchmarks for example if we start with the artificial analysis intelligence index this

15
00:01:13,300 --> 00:01:20,080
model achieves a 61 and fable 5 is at 62 now what's really important to remember is that once again

16
00:01:20,080 --> 00:01:25,460
this model is quite cheap compared to Fable 5. This is about $2 per million input tokens

17
00:01:25,460 --> 00:01:31,900
and $6 per million output tokens. And Fable 5 sits at $10 per million input tokens and $50

18
00:01:31,900 --> 00:01:36,700
per million output tokens. And this is why you start to appreciate this release a little bit

19
00:01:36,700 --> 00:01:42,240
more. It might not beat the performance of the best models. It's matching them. Even GPT 5.6

20
00:01:42,240 --> 00:01:47,540
sold, which is a pretty capable model, and this is set at max on the artificial analysis index.

21
00:01:47,540 --> 00:01:54,660
this model achieves 61 and grok 4.6 61 so it ties it and it's much cheaper and then even on all of

22
00:01:54,660 --> 00:02:00,420
these other benchmarks for example the gdp valve one it actually beats fable 5 which is at 1741

23
00:02:00,420 --> 00:02:06,900
and then gpt 5.6 it's at 1728 i'm not sure why they didn't choose opus 5 as well but i guess

24
00:02:06,900 --> 00:02:11,460
they wanted to choose the quote-unquote strongest model lineup from each lab and they chose fable

25
00:02:11,460 --> 00:02:16,140
5 for anthropic which is fair and then deep software engineering one which is a critical

26
00:02:16,140 --> 00:02:23,720
benchmark this model doesn't beat fable 5 or gpt 5.6 soul but it gets close to it it's 65.9 and

27
00:02:23,720 --> 00:02:29,040
fable 5 sits at 70 but if you're getting results that are pretty close and the model is five times

28
00:02:29,040 --> 00:02:34,080
cheaper i wouldn't be too disappointed with this result and then same thing with cursor bench 3.2

29
00:02:34,080 --> 00:02:42,240
the model achieves 69.9 funny number and fable 5 sits at 70.5 so once again closer and grok 4.6

30
00:02:42,240 --> 00:02:48,200
beats GPT 5.6 soul. The same thing with the Frontier Code, it gets close to Fable 5 beats

31
00:02:48,200 --> 00:02:53,940
GPT 5.6 soul. So yeah, this model is actually available in Cursor. So the partnership with

32
00:02:53,940 --> 00:02:58,960
Cursor or I guess the acquisition has really helped SpaceX make some strides in the AI space

33
00:02:58,960 --> 00:03:04,120
this year. And Grok build is something that, you know, maybe not a lot of us have been using so

34
00:03:04,120 --> 00:03:09,480
far, but it's probably going to be another platform like Codex or Cloud Code that we start to use.

35
00:03:09,480 --> 00:03:14,580
but obviously cursor is quite strong as well so you have options available for you to use them in

36
00:03:14,580 --> 00:03:20,240
both and one thing to note is that they're offering two times usage inside grok build and cursor for

37
00:03:20,240 --> 00:03:25,320
the first week so if you just want to try it out see what you feel about it then you know it might

38
00:03:25,320 --> 00:03:29,760
be worth trying it out right now because you get double the usage in the first week now if we were

39
00:03:29,760 --> 00:03:34,520
to take a look at some of the outputs that people have been generating with grok 4.6 what we're

40
00:03:34,520 --> 00:03:40,180
looking at right now is a Falcon 9 booster return sequence simulation and this was done in a single

41
00:03:40,180 --> 00:03:47,620
HTML file and as I said if you expected to get this type of output from Grok in 2025 you would

42
00:03:47,620 --> 00:03:52,140
be kind of surprised because you wouldn't expect something like this to be generated with Grok but

43
00:03:52,140 --> 00:03:55,960
now it looks like we have to start taking the Grok team a little bit more serious because

44
00:03:55,960 --> 00:04:01,160
this output is quite competitive and obviously we're just looking at a simulation and we're just

45
00:04:01,160 --> 00:04:05,940
basing it off of a visual representation but if the model is able to produce something like this

46
00:04:05,940 --> 00:04:11,340
consistently then I would expect a lot of people to start adopting Grok because number one it is

47
00:04:11,340 --> 00:04:15,660
cheaper than the other labs at least at the moment because we don't know if this pricing strategy is

48
00:04:15,660 --> 00:04:20,360
going to be sustainable for the SpaceX team in the long run but at least for now their models are

49
00:04:20,360 --> 00:04:25,280
definitely cheaper compared to the others and as I mentioned this model excelled at the artificial

50
00:04:25,280 --> 00:04:32,580
analysis index. This model jumped to probably number four model. It's tied to GPT 5.6. It's

51
00:04:32,580 --> 00:04:36,560
pretty much similar. So you could say number three as well. But the models before that are

52
00:04:36,560 --> 00:04:43,540
Fable 5 and Opus 5, which are only above the model by about a 1% or a 2% difference. So yeah,

53
00:04:43,580 --> 00:04:48,640
even on this intelligence index, which if you're not familiar with, has nine evaluations. So on

54
00:04:48,640 --> 00:04:54,220
all of these evaluations, it's kind of matching almost Fable 5 performance, which is crazy to see.

55
00:04:54,220 --> 00:04:59,760
Before we continue, we just launched the Universe of AI newsletter. If you want to stay on top of

56
00:04:59,760 --> 00:05:05,060
AI news without having to hunt for it, link is in the description. Don't miss out. And what you see

57
00:05:05,060 --> 00:05:10,980
on screen right now is a racing game that Grok 4.6 built. And based off of this post, the model took

58
00:05:10,980 --> 00:05:15,780
about one minute and it was a five word prompt, which was a create a simple racing game in HTML.

59
00:05:16,280 --> 00:05:21,200
So if you're able to generate something like this easily using Grok 4.6, I think a lot of people

60
00:05:21,200 --> 00:05:27,160
will be happy. And this is a more detailed analysis of what it costs to run GPT 5.6 on

61
00:05:27,160 --> 00:05:32,120
the artificial analysis index and what it produced, meaning the output. Both of these models, if you

62
00:05:32,120 --> 00:05:38,120
remember, scored 61 on the artificial analysis index. Now to run the whole test with GPT 5.6,

63
00:05:38,480 --> 00:05:45,500
it costs about 2.8k. And then with Grok 4.6, it costs about 1.1k-ish. And this tells you that

64
00:05:45,500 --> 00:05:50,420
you're getting similar level of performance at half the cost. So yeah, this is a big release for

65
00:05:50,420 --> 00:05:54,280
the SpaceX team because they just proved once again that they are a lab that you seriously

66
00:05:54,280 --> 00:05:59,900
start into considering especially in 2026 and I'm going to talk more about DeepSeek version for Pro

67
00:05:59,900 --> 00:06:05,600
GA but basically what we're seeing today is that both of these releases kind of emphasize the fact

68
00:06:05,600 --> 00:06:10,520
that performance and all above that is the price at what you're getting for that performance is

69
00:06:10,520 --> 00:06:15,880
becoming more and more important for all users because we see many labs now focusing on creating

70
00:06:15,880 --> 00:06:21,380
the best model at the cheapest cost. Last year in 2025, most of the labs were just focused on,

71
00:06:21,440 --> 00:06:26,680
I would say, creating the strongest model. Yes, cost was important, but I think most of the times

72
00:06:26,680 --> 00:06:31,620
the frontier labs, meaning OpenAI Anthropic, were kind of more lenient on that fact because they

73
00:06:31,620 --> 00:06:38,020
didn't have as strong of a competition. Intelligent models that are maybe not always ahead of OpenAI

74
00:06:38,020 --> 00:06:42,560
Anthropic, but match their performance at a fraction of the cost. So yes, price to performance

75
00:06:42,560 --> 00:06:49,800
ratio is becoming a critical I would say indicator in 2026. Now this is the updated benchmark chart

76
00:06:49,800 --> 00:06:54,120
after the release of the new model and one thing you'll see across the board is that it matches

77
00:06:54,120 --> 00:06:58,360
the top level performance of many of the models. The one thing interesting over here is that they

78
00:06:58,360 --> 00:07:02,960
haven't put Opus 5 here for some reason. There is Fable 5 here that we can compare this model

79
00:07:02,960 --> 00:07:08,900
against but one thing you'll notice is that DeepSeek, remember this model costs 43.5 cents

80
00:07:08,900 --> 00:07:14,940
per million input tokens and 87 cents per million output tokens while the other models all over here

81
00:07:14,940 --> 00:07:20,140
are way more expensive than that. So the first thing if you look at Terminal Bench 2.1 the model

82
00:07:20,140 --> 00:07:26,126
scores 87.9. The older version of the model was 72.1 and the flash version was

83
00:07:26,126 --> 00:07:29,526
of the model was 72.1, and the Flash version, which we got last week, was 82.7. And what's

84
00:07:29,526 --> 00:07:37,146
crazy is that Fable 5 is 88. Yes, 88. So this model is only 0.1% behind Fable 5. And then on

85
00:07:37,146 --> 00:07:43,606
the Cyber Gym which is Cybersecurity the model actually beats Fable 5. Fable 5 sits at 83.1

86
00:07:43,606 --> 00:07:49,546
while DeepSeek version 4 Pro the new one that we got today is at 83.3. Now if this tells you

87
00:07:49,546 --> 00:07:54,706
something is that this model is once again really geared at Cybersecurity. We saw the Flash model

88
00:07:54,706 --> 00:07:58,966
also be geared towards Cybersecurity and becoming a model that was quite capable in that area

89
00:07:58,966 --> 00:08:04,986
and we're seeing the same thing today with the new version 4 Pro which sits at 83.3 outperforming

90
00:08:04,986 --> 00:08:10,486
Fable 5 on the benchmark. On deep software engineering this model is not beating Fable 5

91
00:08:10,486 --> 00:08:17,786
because Fable 5 sits at 70 but the model achieves 62.7 which is ahead of GLM 5.2 and a bit behind

92
00:08:17,786 --> 00:08:24,186
Kimi K3 which sits at 67.5 but one thing again this is a big jump compared to the preview version

93
00:08:24,186 --> 00:08:29,686
of the model which was at 12.8 so yeah they have really trained this model and have improved it on

94
00:08:29,686 --> 00:08:34,706
the back end because we can clearly see in this deep software engineering benchmark. And there's

95
00:08:34,706 --> 00:08:40,266
couple of other benchmarks but one key benchmark that this model excels at is the automation bench

96
00:08:40,266 --> 00:08:48,346
where the model achieves 31.8 and fable 5 29.1 so once again the model outperforms fable 5 now this

97
00:08:48,346 --> 00:08:53,166
is a fraction of the cost as i mentioned this is 43 cents and 87 cents for input and output

98
00:08:53,166 --> 00:08:59,606
respectively versus fable 5 which is at 10.50 so yes if we look at the terminal bench for example

99
00:08:59,606 --> 00:09:05,106
we are seeing this model achieve a result which is 0.1 behind the best model out there fable 5

100
00:09:05,106 --> 00:09:12,766
and the pricing is insane to look at 43 cents versus 10 dollars 87 cents versus 50 dollars per

101
00:09:12,766 --> 00:09:19,446
million input tokens so when we turn that into a fraction this model is 57 times cheaper and earlier

102
00:09:19,446 --> 00:09:24,586
last week i think i made a video about how deep seek version 4 pro was going to achieve a result

103
00:09:24,586 --> 00:09:28,866
like this because there was an investor report that leaked and when i saw the pricing where it

104
00:09:28,866 --> 00:09:34,206
said that it was going to match Fable 5 or even outperform it and be at 57 times cheaper per cost.

105
00:09:34,366 --> 00:09:38,906
When I was reading their investor report, I didn't really believe it. But today it is proven because

106
00:09:38,906 --> 00:09:43,786
at least on the benchmark so far, yes, I'm just saying the benchmarks, we are seeing similar level

107
00:09:43,786 --> 00:09:49,086
performance. Now, if this model actually performs like that in production, it's a little too early

108
00:09:49,086 --> 00:09:53,506
to tell yet. We still would have to give a couple of weeks and see how it's performing in the long

109
00:09:53,506 --> 00:09:58,206
run. Because sometimes on the release date, the models perform quite good. But over time,

110
00:09:58,206 --> 00:10:02,406
they kind of deteriorate in their quality. So we hope DeepSeek doesn't do that. Historically,

111
00:10:02,486 --> 00:10:06,726
they haven't done that. But let's just see. Just going to put it out there because right now we're

112
00:10:06,726 --> 00:10:11,886
basing this off of benchmarks. But that's it for today's video. Make sure you guys are subscribed

113
00:10:11,886 --> 00:10:17,666
to the channel. Follow our new newsletter as well at universeofai.beehive.com. As well,

114
00:10:17,726 --> 00:10:22,906
subscribe to the main channel, World of AI, and support us on X by following the Universe of AIZ

115
00:10:22,906 --> 00:10:25,706
as well. Until then, I'll see you guys in the next video.
