實際影片長度:8:00.000。原文、繁中、雙語可點擊句子跳轉影片。
0:00.000–0:03.414
Today, we are unpacking how a small independent lab,
0:03.414–0:05.820
Emperor AI, just dropped an absolute
0:05.820–0:09.240
powerhouse of a model that's going toe-to-toe with big tech.
0:09.240–0:11.010
Look, if you follow the AI space,
0:11.010–0:13.240
you know there is a lot of noise out there.
0:13.240–0:16.028
We get bombarded with hype cycles, PR stunts, and
0:16.028–0:18.000
vaporware that you can't even use.
0:18.000–0:20.600
But this update, it's genuinely worth your time.
0:20.600–0:24.209
We're talking about a highly capable model that runs completely locally,
0:24.209–0:25.140
for free, natively
0:25.140–0:26.436
understands images, and
0:26.436–0:29.460
can hold an immense amount of data in its memory all at once,
0:29.460–0:30.460
without breaking a sweat.
0:30.460–0:33.387
It is a massive leap forward for the open source community,
0:33.387–0:34.880
and it fundamentally changes
0:34.880–0:37.740
what independent developers have access to.
0:37.740–0:39.020
So here's our agenda for today.
0:39.020–0:41.107
We'll meet Emperos QFOS 27B,
0:41.107–0:45.280
dive into its massive million-word memory, check out its
0:45.280–0:49.274
vision and speed, explore some real-world business use cases,
0:49.274–0:51.460
and finally, go over the fine print
0:51.460–0:52.620
and next steps.
0:52.620–0:54.840
Alright, section 1.
0:54.840–0:57.400
Meet Emperos QFOS 27B.
0:57.400–1:00.115
To really get why this release is so impressive,
1:00.115–1:02.220
you've got to know the team behind it.
1:02.220–1:04.122
Emperos AI is a laser-focused,
1:04.122–1:07.780
independent lab making some serious ways with open-weight
1:07.780–1:08.780
models.
1:08.780–1:10.751
QFOS 27B is basically the bigger,
1:10.751–1:14.340
incredibly capable sibling to their hit 9 billion parameter
1:14.340–1:15.340
model.
1:15.340–1:18.175
For this release, they started with the QFOS 3.
1:18.175–1:20.720
527B architecture as their base, and then
1:20.720–1:23.400
ran it through a really rigorous training sequence.
1:23.400–1:28.220
They pushed it through full-parameter supervised fine-tuning to nail down structure and formatting,
1:28.220–1:30.219
then direct preference optimization so
1:30.219–1:33.040
it actually knows what a high-quality answer looks like,
1:33.040–1:35.040
and capped it off with ESFT.
1:35.040–1:41.320
Now, doing a full-parameter fine-tune on a 27 billion-parameter model is brutally resource-intensive.
1:41.320–1:45.540
Most small labs would just take shortcuts to save on compute costs, not Emperos.
1:45.540–1:46.901
They refused to cut corners,
1:46.901–1:48.319
meaning every single capabil
1:48.319–1:50.360
ity from that base model was kept perfectly
1:50.360–1:51.360
intact.
1:51.360–1:53.360
Which brings us to section 2.
1:53.360–1:55.360
A million-word memory.
1:55.360–1:58.180
Over 1 million tokens.
1:58.180–1:59.797
For an open-weight release,
1:59.797–2:03.300
hitting a context window this massive is just a wild technical
2:03.300–2:04.300
achievement.
2:04.300–2:07.740
They pulled this off using a technique called yarn scaling.
2:07.740–2:09.631
Without getting totally bogged down in the math,
2:09.631–2:11.200
yarn scaling essentially stretches the
2:11.200–2:12.920
model's positional embeddings.
2:12.920–2:17.140
It basically tricks the neural network into handling incredibly long sequences of text without
2:17.140–2:19.940
losing its mind or degrading the output quality.
2:19.940–2:20.940
The result?
2:20.940–2:22.993
It's a model that stays perfectly coherent even when
2:22.993–2:24.300
you just flood it with an absolute
2:24.300–2:25.860
mountain of data.
2:25.860–2:28.082
To put that million tokens into perspective,
2:28.082–2:30.480
it's practically like holding an entire library
2:30.480–2:32.420
in its head all at once.
2:32.420–2:33.420
We've all been there, right?
2:33.420–2:34.388
You're using an AI and
2:34.388–2:37.560
it completely forgets the instructions you gave it just three prompts
2:37.560–2:38.560
ago.
2:38.560–2:41.460
It's incredibly frustrating and completely derails your workflow.
2:41.460–2:45.300
But with a context window this big, that problem vanishes.
2:45.300–2:47.526
You can hand it hundreds of pages of meeting notes,
2:47.526–2:49.540
your company's whole codebase, or literally
2:49.540–2:52.180
every customer conversation you've ever had.
2:52.180–2:54.380
And it looks at all of it simultaneously.
2:54.380–2:56.780
It remembers all those microscopic details and
2:56.780–2:58.940
gives you a flawless, fully context-aware
2:58.940–2:59.940
answer.
2:59.940–3:03.720
Let's move on to Section 3, Vision and Speed.
3:03.720–3:07.900
Because Emperor kept the full vision tower totally intact during that intense training
3:07.900–3:11.960
pipeline, this model doesn't just read giant walls of text.
3:11.960–3:14.852
It can instantly pull meaning out of screenshots,
3:14.852–3:17.400
complex charts, or even messy handwriting.
3:17.400–3:20.040
Think about how much friction there is in manual data entry.
3:20.040–3:24.540
Normally, you're stuck transcribing a chaotic whiteboard after a long meeting or typing out
3:24.540–3:27.820
data points from a financial chart just to analyze them.
3:27.820–3:30.675
Now you just take a screenshot or snap a photo of that whiteboard,
3:30.675–3:31.680
hand it directly to the
3:31.680–3:35.000
model, and it spots the patterns and pulls the meaning instantly.
3:35.000–3:38.260
You literally save hours of tedious manual work.
3:38.260–3:40.500
And it is incredibly fast.
3:40.500–3:43.720
Standard models can act a bit like a turtle when generating text.
3:43.720–3:45.876
They suffer from this bottleneck because
3:45.876–3:48.340
they predict one single word, pause, think, and
3:48.340–3:49.740
then predict the next one.
3:49.740–3:52.540
It works, but it can be agonizingly slow.
3:52.540–3:55.185
Kwaithos 27b leaps ahead like lightning because
3:55.185–3:57.960
it utilizes a native multi-token prediction head,
3:57.960–3:59.380
or MTP.
3:59.380–4:01.139
Instead of calculating one word at a time,
4:01.139–4:03.200
it confidently predicts chunks of several words
4:03.200–4:04.200
all at once.
4:04.200–4:07.520
It's basically a highly accurate deep learning autocomplete.
4:07.520–4:10.519
This speculative decoding saves a massive amount of inference time,
4:10.519–4:11.760
giving you incredible speed
4:11.760–4:14.580
without sacrificing any response quality.
4:14.580–4:17.620
This right here is what we call the rare trifecta.
4:17.620–4:21.780
An MTP head, a vision tower, and a full context window.
4:21.780–4:24.423
Usually when independent devs fine-tune a model this big,
4:24.423–4:25.880
they have to make really painful
4:25.880–4:28.340
compromises just to save on compute and memory.
4:28.340–4:31.066
They'll strip out the vision tower, drop the MTP head,
4:31.066–4:32.580
or slash the context window to
4:32.580–4:34.460
a fraction of its size.
4:34.460–4:37.960
Putting all three to work in perfect harmony without breaking the model is incredibly rare
4:37.960–4:39.620
in the open-wake community.
4:39.620–4:42.547
And Para didn't compromise, giving us a tool that's smart,
4:42.547–4:44.280
fast, and wonderfully versatile.
4:44.280–4:47.960
Okay, section 4, business use cases.
4:47.960–4:50.700
Let's see how you can actually apply this.
4:50.700–4:53.080
Let's put this into a real-world scenario.
4:53.080–4:55.120
Say you are a social media manager.
4:55.120–4:56.282
It's Monday morning and
4:56.282–4:58.780
it's time for the dreaded weekly analytics review.
4:58.780–5:00.044
Instead of endless scrolling and
5:00.044–5:02.300
manually copying data into a spreadsheet, you just feed the
5:02.300–5:04.430
model 20 of your top posts all at once,
5:04.430–5:07.660
including the attached images and engagement metrics.
5:07.660–5:09.871
Because of that massive context window and
5:09.871–5:12.020
intact vision tower, it digests a month's
5:12.020–5:14.240
worth of content planning in seconds.
5:14.240–5:17.206
It tells you exactly what topics resonated most with your audience and
5:17.206–5:18.060
helps you instantly
5:18.060–5:21.960
draft clear, perfectly tailored replies to common customer questions.
5:21.960–5:24.300
It's a complete game-changer for a daily workflow.
5:24.300–5:29.120
Finally, section 5, the fine print and next steps.
5:29.120–5:30.660
First up, the license.
5:30.660–5:34.520
Kwethos 27B runs on the open Apache 2.0 license.
5:34.520–5:37.432
If you are tired of the restrictive licenses big tech often pushes,
5:37.432–5:38.420
where things are gated
5:38.420–5:42.429
behind non-commercial use-only clauses or massive enterprise fees,
5:42.429–5:44.060
this is a breath of fresh air.
5:44.060–5:46.460
It means anyone can download this model.
5:46.460–5:48.300
Use it in a private business.
5:48.300–5:49.960
Integrate it into a commercial product.
5:49.960–5:52.580
And even make money from it completely freely.
5:52.580–5:55.200
There are no surprise usage caps hiding in the background.
5:55.200–5:58.260
It is absolute freedom to build exactly what you want.
5:58.260–6:00.520
Now, you do need a quick heads up on safety.
6:00.520–6:04.000
Ampra explicitly calls this an uncensored model.
6:04.000–6:07.520
Unlike commercial models from major corporations that heavily filter their outputs right out
6:07.520–6:10.820
of the box, this gives you raw unfiltered power.
6:10.820–6:13.100
It has far fewer default guardrails.
6:13.100–6:16.373
What this means in practice is that if you're building a customer-facing product,
6:16.373–6:17.180
the responsibility
6:17.180–6:19.820
is entirely on you to build in the safety nets.
6:19.820–6:22.177
You have to add your own application-level safety controls,
6:22.177–6:23.240
like secondary moderation
6:23.240–6:24.760
endpoints or strict system prompts,
6:24.760–6:26.820
to make sure the model behaves exactly how you want
6:26.820–6:27.760
it to in public.
6:27.760–6:32.180
Also, keep in mind this current release is a version 1 pre-RL checkpoint.
6:32.180–6:34.380
This is completely normal for open-weight releases.
6:34.380–6:38.120
It just means the model hasn't gone through its final reinforcement learning polish yet.
6:38.120–6:41.716
That final RLHF step is what usually makes a model a bit more chatty,
6:41.716–6:43.000
agreeable, and precise
6:43.000–6:45.440
with highly specific formatting instructions.
6:45.440–6:47.582
So while this version 1 is brilliantly capable and
6:47.582–6:49.520
ready to deploy right now, you can absolutely
6:49.520–6:53.280
expect an even sharper, more refined version 2 in the near future.
6:53.280–6:56.200
And looking at Empero's roadmap, they are moving fast.
6:56.200–7:01.160
They already have a specialized terminal coding agent called Abacus in the works.
7:01.160–7:02.311
Even more exciting,
7:02.311–7:03.394
they are training a
7:03.394–7:06.440
small in-house mixture of experts model named Clair.
7:06.440–7:08.678
MOE models are insanely efficient because
7:08.678–7:11.600
they route queries to specialized subnetworks, saving
7:11.600–7:14.940
a ton of compute while radically boosting intelligence.
7:14.940–7:19.261
Seeing an independent lab successfully building an MOE model is super ambitious,
7:19.261–7:20.200
and it definitely
7:20.200–7:24.640
makes Empero AI a vital team to watch in the open-source space.
7:24.640–7:26.700
Which brings us to our final thought.
7:26.700–7:30.419
Looking at everything Coithos 27B brings to the table,
7:30.419–7:32.440
it forces us to ask, if a small,
7:32.440–7:35.085
independent lab can deliver a 1 million token memory,
7:35.085–7:37.260
native vision, and multi-token prediction
7:37.260–7:41.355
entirely for free, how soon until the gap between open-source and
7:41.355–7:43.440
big tech completely disappears?
7:43.440–7:47.110
We are seeing definitive proof of how incredibly fast this community is moving,
7:47.110–7:48.260
empowering individuals
7:48.260–7:52.360
with tools that would have required a massive corporate budget just a year ago.
7:52.360–7:54.420
Thank you so much for joining me for this explainer,
7:54.420–7:56.240
and I hope you leave today feeling inspired by
7:56.240–7:59.480
the incredible capabilities now freely available at your fingertips.
0:00.000–0:03.414
今天,我們來拆解一個小型獨立實驗室
0:03.414–0:05.820
Emperor AI 剛剛發布了一款
0:05.820–0:09.240
絕對強悍的模型,它將與大型科技公司正面競爭。
0:09.240–0:11.010
聽著,如果你關注 AI 領域,
0:11.010–0:13.240
你就知道外面充滿了許多噪音。
0:13.240–0:16.028
我們不斷被炒作週期、公關噱頭,
0:16.028–0:18.000
以及那些你甚至無法使用的虛構產品所轟炸。
0:18.000–0:20.600
但這次的更新,確實值得你花時間關注。
0:20.600–0:24.209
我們談論的是一個高度強大的模型,它可以完全在本地運行,
0:24.209–0:25.140
免費,原生
0:25.140–0:26.436
理解圖像,並且
0:26.436–0:29.460
能夠一次性在記憶體中容納海量數據,
0:29.460–0:30.460
毫不費力。
0:30.460–0:33.387
這對開源社群來說是一大進步,
0:33.387–0:34.880
它根本性地改變了
0:34.880–0:37.740
獨立開發者可獲得的資源。
0:37.740–0:39.020
所以,這是我們今天的議程。
0:39.020–0:41.107
我們將認識 Emperos QFOS 27B,
0:41.107–0:45.280
深入探討其龐大的百萬字詞記憶體,查看其
0:45.280–0:49.274
視覺能力和速度,探索一些實際的商業用例,
0:49.274–0:51.460
最後,仔細檢視細則
0:51.460–0:52.620
以及後續步驟。
0:52.620–0:54.840
好了,第一部分。
0:54.840–0:57.400
認識 Emperos QFOS 27B。
0:57.400–1:00.115
要真正理解這次發布為何如此令人印象深刻,
1:00.115–1:02.220
你必須了解背後的團隊。
1:02.220–1:04.122
Emperos AI 是一個專注度極高的
1:04.122–1:07.780
獨立實驗室,正在開權重
1:07.780–1:08.780
模型領域取得重大進展。
1:08.780–1:10.751
QFOS 27B 基本上是其熱門的 90 億參數
1:10.751–1:14.340
模型更大、
1:14.340–1:15.340
能力極強的兄弟版本。
1:15.340–1:18.175
對於這次發布,他們從 QFOS 3.
1:18.175–1:20.720
5 27B 架構作為基礎,然後
1:20.720–1:23.400
對其進行了非常嚴格的訓練序列。
1:23.400–1:28.220
他們通過全參數監督微調來確定結構和格式,
1:28.220–1:30.219
然後進行直接偏好優化,以便
1:30.219–1:33.040
它真正知道高質量答案的樣貌,
1:33.040–1:35.040
最後以 ESFT 收尾。
1:35.040–1:41.320
現在,對一個 270 億參數的模型進行全參數微調是非常耗費資源的。
1:41.320–1:45.540
大多數小型實驗室為了節省運算成本,通常會走捷徑,但Emperos沒有。
1:45.540–1:46.901
他們拒絕偷工減料,
1:46.901–1:48.319
意味著基礎模型中的每項能力
1:48.319–1:50.360
都完美地
1:50.360–1:51.360
保持完整。
1:51.360–1:53.360
這就帶我們進入第二部分。
1:53.360–1:55.360
百萬字記憶。
1:55.360–1:58.180
超過100萬個token。
1:58.180–1:59.797
對於開放權重的發布來說,
1:59.797–2:03.300
達到如此龐大的上下文視窗,簡直是一個瘋狂的技術
2:03.300–2:04.300
成就。
2:04.300–2:07.740
他們使用了一種稱為Yarn Scaling的技術來實現這一點。
2:07.740–2:09.631
不必完全陷入數學細節,
2:09.631–2:11.200
Yarn Scaling本質上擴展了
2:11.200–2:12.920
模型的位置嵌入。
2:12.920–2:17.140
它基本上欺騙神經網絡,使其能夠處理極其長的文字序列,而不會
2:17.140–2:19.940
失去理智或降低輸出品質。
2:19.940–2:20.940
結果如何?
2:20.940–2:22.993
這是一個即使在
2:22.993–2:24.300
你只是傾倒大量
2:24.300–2:25.860
數據山時,也能保持完美連貫性的模型。
2:25.860–2:28.082
為了讓這100萬個token更具體,
2:28.082–2:30.480
這 practically 就像同時將整個圖書館
2:30.480–2:32.420
裝進它的腦袋裡。
2:32.420–2:33.420
我們都有過這種經歷,對吧?
2:33.420–2:34.388
當你使用AI時,
2:34.388–2:37.560
它完全忘記了你在三個提示前給它的指令。
2:37.560–2:38.560
這令人非常沮喪,並完全打亂了你的工作流程。
2:38.560–2:41.460
但擁有如此大的上下文視窗,這個問題就消失了。
2:41.460–2:45.300
你可以交給它數百頁的會議記錄,
2:45.300–2:47.526
你們公司的整個程式碼庫,或者 literally
2:47.526–2:49.540
你曾經有過的每一位客戶對話。
2:49.540–2:52.180
它會同時檢視所有這些內容。
2:52.180–2:54.380
它記住所有那些微小的細節,並
2:54.380–2:56.780
給出完美、完全具備情境意識的
2:56.780–2:58.940
回答。
2:58.940–2:59.940
讓我們繼續看第三部分:視覺與速度。
2:59.940–3:03.720
因為Emperor在該密集訓練
3:03.720–3:07.900
管線中完全保留了完整的視覺塔,這個模型不僅能閱讀巨大的文字牆。
3:07.900–3:11.960
管線,這個模型不只是讀取龐大的文字牆。
3:11.960–3:14.852
它能從截圖中瞬間提取意義,
3:14.852–3:17.400
複雜的圖表,甚至是凌亂的手寫字跡。
3:17.400–3:20.040
想像一下手動輸入資料有多麼繁瑣。
3:20.040–3:24.540
通常,你只能在漫長的會議後,被迫抄錄混亂的白板內容,
3:24.540–3:27.820
或是從財務圖表中輸入數據點,以便進行分析。
3:27.820–3:30.675
現在,你只需截圖或拍下白板的照片,
3:30.675–3:31.680
直接交給
3:31.680–3:35.000
模型,它就能立即識別模式並提取意義。
3:35.000–3:38.260
你實際上節省了數小時繁瑣的手動工作時間。
3:38.260–3:40.500
而且它的速度極快。
3:40.500–3:43.720
標準模型在生成文字時,有時會像烏龜一樣緩慢。
3:43.720–3:45.876
它們面臨這種瓶頸,因為
3:45.876–3:48.340
它們一次只預測一個單詞,暫停,思考,然後
3:48.340–3:49.740
再預測下一個。
3:49.740–3:52.540
這確實有效,但速度可能慢得令人痛苦。
3:52.540–3:55.185
Kwaithos 27b 像閃電一樣躍進,因為
3:55.185–3:57.960
它利用了原生的多標記預測頭,
3:57.960–3:59.380
即 MTP。
3:59.380–4:01.139
與其一次計算一個單詞,
4:01.139–4:03.200
它自信地一次預測幾個單詞的區塊
4:03.200–4:04.200
4:04.200–4:07.520
這基本上是一個高度準確的深度學習自動完成功能。
4:07.520–4:10.519
這種推測解碼節省了大量的推理時間,
4:10.519–4:11.760
為你帶來驚人的速度,
4:11.760–4:14.580
同時不犧牲任何回應品質。
4:14.580–4:17.620
這裡正是我們所稱的稀有三重奏。
4:17.620–4:21.780
一個 MTP 頭、一個視覺塔,以及完整的上下文視窗。
4:21.780–4:24.423
通常,當獨立開發人員微調這麼大的模型時,
4:24.423–4:25.880
他們必須做出非常痛苦的
4:25.880–4:28.340
妥協,以節省運算和記憶體。
4:28.340–4:31.066
他們會移除視覺塔,丟棄 MTP 頭,
4:31.066–4:32.580
或將上下文視窗縮減至
4:32.580–4:34.460
其大小的一小部分。
4:34.460–4:37.960
在開放原始碼社群中,將這三者完美協同運作而不破壞模型,極為罕見
4:37.960–4:39.620
4:39.620–4:42.547
而 Para 沒有妥協,為我們提供了一個聰明、
4:42.547–4:44.280
快速且極具多功能性的工具。
4:44.280–4:47.960
好的,第 4 節,商業應用案例。
4:47.960–4:50.700
讓我們看看你如何實際應用這一點。
4:50.700–4:53.080
讓我們將此置於真實世界的場景中。
4:53.080–4:55.120
假設你是一位社群媒體經理。
4:55.120–4:56.282
週一早上,
4:56.282–4:58.780
又到了令人畏懼的每週數據分析檢討時間。
4:58.780–5:00.044
與其無止盡地滑動螢幕,
5:00.044–5:02.300
並手動將資料複製到試算表中,你只需將
5:02.300–5:04.430
你的前 20 篇熱門貼文一次餵給
5:04.430–5:07.660
模型,包含附上的圖片和互動指標。
5:07.660–5:09.871
由於龐大的上下文視窗和
5:09.871–5:12.020
完整的視覺塔結構,它能在幾秒內消化
5:12.020–5:14.240
一個月的內容規劃。
5:14.240–5:17.206
它會明確告訴你哪些主題最能引起觀眾共鳴,並
5:17.206–5:18.060
幫助你立即
5:18.060–5:21.960
草擬清晰且完美客製化的回覆,以應對常見的客户問題。
5:21.960–5:24.300
這對日常工作流程來說,是完全的遊戲規則改變者。
5:24.300–5:29.120
最後,第 5 部分,細則和後續步驟。
5:29.120–5:30.660
首先來看授權條款。
5:30.660–5:34.520
Kwethos 27B 採用開放的 Apache 2.0 授權。
5:34.520–5:37.432
如果你厭倦了大型科技公司經常推銷的限制性授權,
5:37.432–5:38.420
那些授權往往
5:38.420–5:42.429
被非商業用途條款或龐大的企業費用所限制,
5:42.429–5:44.060
這簡直是一股清新之風。
5:44.060–5:46.460
這意味著任何人都可以下載這個模型。
5:46.460–5:48.300
在私人企業中使用。
5:48.300–5:49.960
將其整合到商業產品中。
5:49.960–5:52.580
甚至完全自由地從中獲利。
5:52.580–5:55.200
背景中不會隱藏令人驚訝的使用量上限。
5:55.200–5:58.260
這賦予了你絕對的自由,可以構建你想要的任何東西。
5:58.260–6:00.520
現在,你需要快速了解安全方面的注意事項。
6:00.520–6:04.000
Ampra 明確指出這是一個未經審查的模型。
6:04.000–6:07.520
與大型公司提供的商業模型不同,後者在出廠時就會對輸出進行嚴格過濾,
6:07.520–6:10.820
這個模型則賦予你原始且未經過濾的力量。
6:10.820–6:13.100
它預設的安全防護欄要少得多。
6:13.100–6:16.373
這在實務上的意思是,如果你正在構建面向客戶的產品,
6:16.373–6:17.180
責任
6:17.180–6:19.820
完全在於你必須建立安全網。
6:19.820–6:22.177
你必須添加自己應用層級的安全控制,
6:22.177–6:23.240
例如次要的審核
6:23.240–6:24.760
端點或嚴格的系統提示,
6:24.760–6:26.820
以確保模型在公開場合的行為完全符合你的期望
6:26.820–6:27.760
6:27.760–6:32.180
此外,請記住這個當前版本是一個預強化學習(pre-RL)的 v1 檢查點。
6:32.180–6:34.380
對於開放權重的發布來說,這完全正常。
6:34.380–6:38.120
這只意味著該模型尚未經過最終的強化學習打磨。
6:38.120–6:41.716
最終的 RLHF 步驟通常會讓模型變得更加健談、
6:41.716–6:43.000
順從,
6:43.000–6:45.440
並且在處理高度特定的格式指令時更加精確。
6:45.440–6:47.582
因此,雖然這個 v1 版本功能強大且
6:47.582–6:49.520
現在就可以部署,但你絕對可以
6:49.520–6:53.280
期待在不久的將來會有更銳利、更完善的 v2 版本。
6:53.280–6:56.200
從 Empero 的路線圖來看,他們的進展迅速。
6:56.200–7:01.160
他們已經在開發一個名為 Abacus 的專用終端編碼代理。
7:01.160–7:02.311
更令人興奮的是,
7:02.311–7:03.394
他們正在訓練一個
7:03.394–7:06.440
名為 Clair 的小型內部混合專家(MoE)模型。
7:06.440–7:08.678
MoE 模型之所以效率極高,是因為
7:08.678–7:11.600
它們會將查詢路由到專門的子網絡,從而節省
7:11.600–7:14.940
大量運算資源,同時大幅提升智能。
7:14.940–7:19.261
看到獨立實驗室成功構建 MoE 模型是非常雄心勃勃的,
7:19.261–7:20.200
這也確實
7:20.200–7:24.640
使 Empero AI 成為開源領域中一個值得關注的重要團隊。
7:24.640–7:26.700
這引出了我們最後的想法。
7:26.700–7:30.419
回顧 Coithos 27B 所帶來的一切,
7:30.419–7:32.440
它迫使我們思考,如果一個小型、
7:32.440–7:35.085
獨立的實驗室能夠免費提供 100 萬 token 的記憶、
7:35.085–7:37.260
原生視覺和多 token 預測功能,
7:37.260–7:41.355
那麼開源與
7:41.355–7:43.440
大型科技公司之間的差距完全消失還需要多久?
7:43.440–7:47.110
我們正在看到確鑿的證據,證明這個社區的進展速度有多快,
7:47.110–7:48.260
賦予個人
7:48.260–7:52.360
強大的工具,而在一年前,這些工具需要龐大的企業預算才能獲得。
7:52.360–7:54.420
非常感謝你參與這次解說,
7:54.420–7:56.240
我希望你今天離開時,能感受到
7:56.240–7:59.480
現在唾手可得的強大功能所帶來的啟發。
0:00.000–0:03.414
Today, we are unpacking how a small independent lab,
今天,我們來拆解一個小型獨立實驗室
0:03.414–0:05.820
Emperor AI, just dropped an absolute
Emperor AI 剛剛發布了一款
0:05.820–0:09.240
powerhouse of a model that's going toe-to-toe with big tech.
絕對強悍的模型,它將與大型科技公司正面競爭。
0:09.240–0:11.010
Look, if you follow the AI space,
聽著,如果你關注 AI 領域,
0:11.010–0:13.240
you know there is a lot of noise out there.
你就知道外面充滿了許多噪音。
0:13.240–0:16.028
We get bombarded with hype cycles, PR stunts, and
我們不斷被炒作週期、公關噱頭,
0:16.028–0:18.000
vaporware that you can't even use.
以及那些你甚至無法使用的虛構產品所轟炸。
0:18.000–0:20.600
But this update, it's genuinely worth your time.
但這次的更新,確實值得你花時間關注。
0:20.600–0:24.209
We're talking about a highly capable model that runs completely locally,
我們談論的是一個高度強大的模型,它可以完全在本地運行,
0:24.209–0:25.140
for free, natively
免費,原生
0:25.140–0:26.436
understands images, and
理解圖像,並且
0:26.436–0:29.460
can hold an immense amount of data in its memory all at once,
能夠一次性在記憶體中容納海量數據,
0:29.460–0:30.460
without breaking a sweat.
毫不費力。
0:30.460–0:33.387
It is a massive leap forward for the open source community,
這對開源社群來說是一大進步,
0:33.387–0:34.880
and it fundamentally changes
它根本性地改變了
0:34.880–0:37.740
what independent developers have access to.
獨立開發者可獲得的資源。
0:37.740–0:39.020
So here's our agenda for today.
所以,這是我們今天的議程。
0:39.020–0:41.107
We'll meet Emperos QFOS 27B,
我們將認識 Emperos QFOS 27B,
0:41.107–0:45.280
dive into its massive million-word memory, check out its
深入探討其龐大的百萬字詞記憶體,查看其
0:45.280–0:49.274
vision and speed, explore some real-world business use cases,
視覺能力和速度,探索一些實際的商業用例,
0:49.274–0:51.460
and finally, go over the fine print
最後,仔細檢視細則
0:51.460–0:52.620
and next steps.
以及後續步驟。
0:52.620–0:54.840
Alright, section 1.
好了,第一部分。
0:54.840–0:57.400
Meet Emperos QFOS 27B.
認識 Emperos QFOS 27B。
0:57.400–1:00.115
To really get why this release is so impressive,
要真正理解這次發布為何如此令人印象深刻,
1:00.115–1:02.220
you've got to know the team behind it.
你必須了解背後的團隊。
1:02.220–1:04.122
Emperos AI is a laser-focused,
Emperos AI 是一個專注度極高的
1:04.122–1:07.780
independent lab making some serious ways with open-weight
獨立實驗室,正在開權重
1:07.780–1:08.780
models.
模型領域取得重大進展。
1:08.780–1:10.751
QFOS 27B is basically the bigger,
QFOS 27B 基本上是其熱門的 90 億參數
1:10.751–1:14.340
incredibly capable sibling to their hit 9 billion parameter
模型更大、
1:14.340–1:15.340
model.
能力極強的兄弟版本。
1:15.340–1:18.175
For this release, they started with the QFOS 3.
對於這次發布,他們從 QFOS 3.
1:18.175–1:20.720
527B architecture as their base, and then
5 27B 架構作為基礎,然後
1:20.720–1:23.400
ran it through a really rigorous training sequence.
對其進行了非常嚴格的訓練序列。
1:23.400–1:28.220
They pushed it through full-parameter supervised fine-tuning to nail down structure and formatting,
他們通過全參數監督微調來確定結構和格式,
1:28.220–1:30.219
then direct preference optimization so
然後進行直接偏好優化,以便
1:30.219–1:33.040
it actually knows what a high-quality answer looks like,
它真正知道高質量答案的樣貌,
1:33.040–1:35.040
and capped it off with ESFT.
最後以 ESFT 收尾。
1:35.040–1:41.320
Now, doing a full-parameter fine-tune on a 27 billion-parameter model is brutally resource-intensive.
現在,對一個 270 億參數的模型進行全參數微調是非常耗費資源的。
1:41.320–1:45.540
Most small labs would just take shortcuts to save on compute costs, not Emperos.
大多數小型實驗室為了節省運算成本,通常會走捷徑,但Emperos沒有。
1:45.540–1:46.901
They refused to cut corners,
他們拒絕偷工減料,
1:46.901–1:48.319
meaning every single capabil
意味著基礎模型中的每項能力
1:48.319–1:50.360
ity from that base model was kept perfectly
都完美地
1:50.360–1:51.360
intact.
保持完整。
1:51.360–1:53.360
Which brings us to section 2.
這就帶我們進入第二部分。
1:53.360–1:55.360
A million-word memory.
百萬字記憶。
1:55.360–1:58.180
Over 1 million tokens.
超過100萬個token。
1:58.180–1:59.797
For an open-weight release,
對於開放權重的發布來說,
1:59.797–2:03.300
hitting a context window this massive is just a wild technical
達到如此龐大的上下文視窗,簡直是一個瘋狂的技術
2:03.300–2:04.300
achievement.
成就。
2:04.300–2:07.740
They pulled this off using a technique called yarn scaling.
他們使用了一種稱為Yarn Scaling的技術來實現這一點。
2:07.740–2:09.631
Without getting totally bogged down in the math,
不必完全陷入數學細節,
2:09.631–2:11.200
yarn scaling essentially stretches the
Yarn Scaling本質上擴展了
2:11.200–2:12.920
model's positional embeddings.
模型的位置嵌入。
2:12.920–2:17.140
It basically tricks the neural network into handling incredibly long sequences of text without
它基本上欺騙神經網絡,使其能夠處理極其長的文字序列,而不會
2:17.140–2:19.940
losing its mind or degrading the output quality.
失去理智或降低輸出品質。
2:19.940–2:20.940
The result?
結果如何?
2:20.940–2:22.993
It's a model that stays perfectly coherent even when
這是一個即使在
2:22.993–2:24.300
you just flood it with an absolute
你只是傾倒大量
2:24.300–2:25.860
mountain of data.
數據山時,也能保持完美連貫性的模型。
2:25.860–2:28.082
To put that million tokens into perspective,
為了讓這100萬個token更具體,
2:28.082–2:30.480
it's practically like holding an entire library
這 practically 就像同時將整個圖書館
2:30.480–2:32.420
in its head all at once.
裝進它的腦袋裡。
2:32.420–2:33.420
We've all been there, right?
我們都有過這種經歷,對吧?
2:33.420–2:34.388
You're using an AI and
當你使用AI時,
2:34.388–2:37.560
it completely forgets the instructions you gave it just three prompts
它完全忘記了你在三個提示前給它的指令。
2:37.560–2:38.560
ago.
這令人非常沮喪,並完全打亂了你的工作流程。
2:38.560–2:41.460
It's incredibly frustrating and completely derails your workflow.
但擁有如此大的上下文視窗,這個問題就消失了。
2:41.460–2:45.300
But with a context window this big, that problem vanishes.
你可以交給它數百頁的會議記錄,
2:45.300–2:47.526
You can hand it hundreds of pages of meeting notes,
你們公司的整個程式碼庫,或者 literally
2:47.526–2:49.540
your company's whole codebase, or literally
你曾經有過的每一位客戶對話。
2:49.540–2:52.180
every customer conversation you've ever had.
它會同時檢視所有這些內容。
2:52.180–2:54.380
And it looks at all of it simultaneously.
它記住所有那些微小的細節,並
2:54.380–2:56.780
It remembers all those microscopic details and
給出完美、完全具備情境意識的
2:56.780–2:58.940
gives you a flawless, fully context-aware
回答。
2:58.940–2:59.940
answer.
讓我們繼續看第三部分:視覺與速度。
2:59.940–3:03.720
Let's move on to Section 3, Vision and Speed.
因為Emperor在該密集訓練
3:03.720–3:07.900
Because Emperor kept the full vision tower totally intact during that intense training
管線中完全保留了完整的視覺塔,這個模型不僅能閱讀巨大的文字牆。
3:07.900–3:11.960
pipeline, this model doesn't just read giant walls of text.
管線,這個模型不只是讀取龐大的文字牆。
3:11.960–3:14.852
It can instantly pull meaning out of screenshots,
它能從截圖中瞬間提取意義,
3:14.852–3:17.400
complex charts, or even messy handwriting.
複雜的圖表,甚至是凌亂的手寫字跡。
3:17.400–3:20.040
Think about how much friction there is in manual data entry.
想像一下手動輸入資料有多麼繁瑣。
3:20.040–3:24.540
Normally, you're stuck transcribing a chaotic whiteboard after a long meeting or typing out
通常,你只能在漫長的會議後,被迫抄錄混亂的白板內容,
3:24.540–3:27.820
data points from a financial chart just to analyze them.
或是從財務圖表中輸入數據點,以便進行分析。
3:27.820–3:30.675
Now you just take a screenshot or snap a photo of that whiteboard,
現在,你只需截圖或拍下白板的照片,
3:30.675–3:31.680
hand it directly to the
直接交給
3:31.680–3:35.000
model, and it spots the patterns and pulls the meaning instantly.
模型,它就能立即識別模式並提取意義。
3:35.000–3:38.260
You literally save hours of tedious manual work.
你實際上節省了數小時繁瑣的手動工作時間。
3:38.260–3:40.500
And it is incredibly fast.
而且它的速度極快。
3:40.500–3:43.720
Standard models can act a bit like a turtle when generating text.
標準模型在生成文字時,有時會像烏龜一樣緩慢。
3:43.720–3:45.876
They suffer from this bottleneck because
它們面臨這種瓶頸,因為
3:45.876–3:48.340
they predict one single word, pause, think, and
它們一次只預測一個單詞,暫停,思考,然後
3:48.340–3:49.740
then predict the next one.
再預測下一個。
3:49.740–3:52.540
It works, but it can be agonizingly slow.
這確實有效,但速度可能慢得令人痛苦。
3:52.540–3:55.185
Kwaithos 27b leaps ahead like lightning because
Kwaithos 27b 像閃電一樣躍進,因為
3:55.185–3:57.960
it utilizes a native multi-token prediction head,
它利用了原生的多標記預測頭,
3:57.960–3:59.380
or MTP.
即 MTP。
3:59.380–4:01.139
Instead of calculating one word at a time,
與其一次計算一個單詞,
4:01.139–4:03.200
it confidently predicts chunks of several words
它自信地一次預測幾個單詞的區塊
4:03.200–4:04.200
all at once.
4:04.200–4:07.520
It's basically a highly accurate deep learning autocomplete.
這基本上是一個高度準確的深度學習自動完成功能。
4:07.520–4:10.519
This speculative decoding saves a massive amount of inference time,
這種推測解碼節省了大量的推理時間,
4:10.519–4:11.760
giving you incredible speed
為你帶來驚人的速度,
4:11.760–4:14.580
without sacrificing any response quality.
同時不犧牲任何回應品質。
4:14.580–4:17.620
This right here is what we call the rare trifecta.
這裡正是我們所稱的稀有三重奏。
4:17.620–4:21.780
An MTP head, a vision tower, and a full context window.
一個 MTP 頭、一個視覺塔,以及完整的上下文視窗。
4:21.780–4:24.423
Usually when independent devs fine-tune a model this big,
通常,當獨立開發人員微調這麼大的模型時,
4:24.423–4:25.880
they have to make really painful
他們必須做出非常痛苦的
4:25.880–4:28.340
compromises just to save on compute and memory.
妥協,以節省運算和記憶體。
4:28.340–4:31.066
They'll strip out the vision tower, drop the MTP head,
他們會移除視覺塔,丟棄 MTP 頭,
4:31.066–4:32.580
or slash the context window to
或將上下文視窗縮減至
4:32.580–4:34.460
a fraction of its size.
其大小的一小部分。
4:34.460–4:37.960
Putting all three to work in perfect harmony without breaking the model is incredibly rare
在開放原始碼社群中,將這三者完美協同運作而不破壞模型,極為罕見
4:37.960–4:39.620
in the open-wake community.
4:39.620–4:42.547
And Para didn't compromise, giving us a tool that's smart,
而 Para 沒有妥協,為我們提供了一個聰明、
4:42.547–4:44.280
fast, and wonderfully versatile.
快速且極具多功能性的工具。
4:44.280–4:47.960
Okay, section 4, business use cases.
好的,第 4 節,商業應用案例。
4:47.960–4:50.700
Let's see how you can actually apply this.
讓我們看看你如何實際應用這一點。
4:50.700–4:53.080
Let's put this into a real-world scenario.
讓我們將此置於真實世界的場景中。
4:53.080–4:55.120
Say you are a social media manager.
假設你是一位社群媒體經理。
4:55.120–4:56.282
It's Monday morning and
週一早上,
4:56.282–4:58.780
it's time for the dreaded weekly analytics review.
又到了令人畏懼的每週數據分析檢討時間。
4:58.780–5:00.044
Instead of endless scrolling and
與其無止盡地滑動螢幕,
5:00.044–5:02.300
manually copying data into a spreadsheet, you just feed the
並手動將資料複製到試算表中,你只需將
5:02.300–5:04.430
model 20 of your top posts all at once,
你的前 20 篇熱門貼文一次餵給
5:04.430–5:07.660
including the attached images and engagement metrics.
模型,包含附上的圖片和互動指標。
5:07.660–5:09.871
Because of that massive context window and
由於龐大的上下文視窗和
5:09.871–5:12.020
intact vision tower, it digests a month's
完整的視覺塔結構,它能在幾秒內消化
5:12.020–5:14.240
worth of content planning in seconds.
一個月的內容規劃。
5:14.240–5:17.206
It tells you exactly what topics resonated most with your audience and
它會明確告訴你哪些主題最能引起觀眾共鳴,並
5:17.206–5:18.060
helps you instantly
幫助你立即
5:18.060–5:21.960
draft clear, perfectly tailored replies to common customer questions.
草擬清晰且完美客製化的回覆,以應對常見的客户問題。
5:21.960–5:24.300
It's a complete game-changer for a daily workflow.
這對日常工作流程來說,是完全的遊戲規則改變者。
5:24.300–5:29.120
Finally, section 5, the fine print and next steps.
最後,第 5 部分,細則和後續步驟。
5:29.120–5:30.660
First up, the license.
首先來看授權條款。
5:30.660–5:34.520
Kwethos 27B runs on the open Apache 2.0 license.
Kwethos 27B 採用開放的 Apache 2.0 授權。
5:34.520–5:37.432
If you are tired of the restrictive licenses big tech often pushes,
如果你厭倦了大型科技公司經常推銷的限制性授權,
5:37.432–5:38.420
where things are gated
那些授權往往
5:38.420–5:42.429
behind non-commercial use-only clauses or massive enterprise fees,
被非商業用途條款或龐大的企業費用所限制,
5:42.429–5:44.060
this is a breath of fresh air.
這簡直是一股清新之風。
5:44.060–5:46.460
It means anyone can download this model.
這意味著任何人都可以下載這個模型。
5:46.460–5:48.300
Use it in a private business.
在私人企業中使用。
5:48.300–5:49.960
Integrate it into a commercial product.
將其整合到商業產品中。
5:49.960–5:52.580
And even make money from it completely freely.
甚至完全自由地從中獲利。
5:52.580–5:55.200
There are no surprise usage caps hiding in the background.
背景中不會隱藏令人驚訝的使用量上限。
5:55.200–5:58.260
It is absolute freedom to build exactly what you want.
這賦予了你絕對的自由,可以構建你想要的任何東西。
5:58.260–6:00.520
Now, you do need a quick heads up on safety.
現在,你需要快速了解安全方面的注意事項。
6:00.520–6:04.000
Ampra explicitly calls this an uncensored model.
Ampra 明確指出這是一個未經審查的模型。
6:04.000–6:07.520
Unlike commercial models from major corporations that heavily filter their outputs right out
與大型公司提供的商業模型不同,後者在出廠時就會對輸出進行嚴格過濾,
6:07.520–6:10.820
of the box, this gives you raw unfiltered power.
這個模型則賦予你原始且未經過濾的力量。
6:10.820–6:13.100
It has far fewer default guardrails.
它預設的安全防護欄要少得多。
6:13.100–6:16.373
What this means in practice is that if you're building a customer-facing product,
這在實務上的意思是,如果你正在構建面向客戶的產品,
6:16.373–6:17.180
the responsibility
責任
6:17.180–6:19.820
is entirely on you to build in the safety nets.
完全在於你必須建立安全網。
6:19.820–6:22.177
You have to add your own application-level safety controls,
你必須添加自己應用層級的安全控制,
6:22.177–6:23.240
like secondary moderation
例如次要的審核
6:23.240–6:24.760
endpoints or strict system prompts,
端點或嚴格的系統提示,
6:24.760–6:26.820
to make sure the model behaves exactly how you want
以確保模型在公開場合的行為完全符合你的期望
6:26.820–6:27.760
it to in public.
6:27.760–6:32.180
Also, keep in mind this current release is a version 1 pre-RL checkpoint.
此外,請記住這個當前版本是一個預強化學習(pre-RL)的 v1 檢查點。
6:32.180–6:34.380
This is completely normal for open-weight releases.
對於開放權重的發布來說,這完全正常。
6:34.380–6:38.120
It just means the model hasn't gone through its final reinforcement learning polish yet.
這只意味著該模型尚未經過最終的強化學習打磨。
6:38.120–6:41.716
That final RLHF step is what usually makes a model a bit more chatty,
最終的 RLHF 步驟通常會讓模型變得更加健談、
6:41.716–6:43.000
agreeable, and precise
順從,
6:43.000–6:45.440
with highly specific formatting instructions.
並且在處理高度特定的格式指令時更加精確。
6:45.440–6:47.582
So while this version 1 is brilliantly capable and
因此,雖然這個 v1 版本功能強大且
6:47.582–6:49.520
ready to deploy right now, you can absolutely
現在就可以部署,但你絕對可以
6:49.520–6:53.280
expect an even sharper, more refined version 2 in the near future.
期待在不久的將來會有更銳利、更完善的 v2 版本。
6:53.280–6:56.200
And looking at Empero's roadmap, they are moving fast.
從 Empero 的路線圖來看,他們的進展迅速。
6:56.200–7:01.160
They already have a specialized terminal coding agent called Abacus in the works.
他們已經在開發一個名為 Abacus 的專用終端編碼代理。
7:01.160–7:02.311
Even more exciting,
更令人興奮的是,
7:02.311–7:03.394
they are training a
他們正在訓練一個
7:03.394–7:06.440
small in-house mixture of experts model named Clair.
名為 Clair 的小型內部混合專家(MoE)模型。
7:06.440–7:08.678
MOE models are insanely efficient because
MoE 模型之所以效率極高,是因為
7:08.678–7:11.600
they route queries to specialized subnetworks, saving
它們會將查詢路由到專門的子網絡,從而節省
7:11.600–7:14.940
a ton of compute while radically boosting intelligence.
大量運算資源,同時大幅提升智能。
7:14.940–7:19.261
Seeing an independent lab successfully building an MOE model is super ambitious,
看到獨立實驗室成功構建 MoE 模型是非常雄心勃勃的,
7:19.261–7:20.200
and it definitely
這也確實
7:20.200–7:24.640
makes Empero AI a vital team to watch in the open-source space.
使 Empero AI 成為開源領域中一個值得關注的重要團隊。
7:24.640–7:26.700
Which brings us to our final thought.
這引出了我們最後的想法。
7:26.700–7:30.419
Looking at everything Coithos 27B brings to the table,
回顧 Coithos 27B 所帶來的一切,
7:30.419–7:32.440
it forces us to ask, if a small,
它迫使我們思考,如果一個小型、
7:32.440–7:35.085
independent lab can deliver a 1 million token memory,
獨立的實驗室能夠免費提供 100 萬 token 的記憶、
7:35.085–7:37.260
native vision, and multi-token prediction
原生視覺和多 token 預測功能,
7:37.260–7:41.355
entirely for free, how soon until the gap between open-source and
那麼開源與
7:41.355–7:43.440
big tech completely disappears?
大型科技公司之間的差距完全消失還需要多久?
7:43.440–7:47.110
We are seeing definitive proof of how incredibly fast this community is moving,
我們正在看到確鑿的證據,證明這個社區的進展速度有多快,
7:47.110–7:48.260
empowering individuals
賦予個人
7:48.260–7:52.360
with tools that would have required a massive corporate budget just a year ago.
強大的工具,而在一年前,這些工具需要龐大的企業預算才能獲得。
7:52.360–7:54.420
Thank you so much for joining me for this explainer,
非常感謝你參與這次解說,
7:54.420–7:56.240
and I hope you leave today feeling inspired by
我希望你今天離開時,能感受到
7:56.240–7:59.480
the incredible capabilities now freely available at your fingertips.
現在唾手可得的強大功能所帶來的啟發。

影片筆記:NEW Qwythos-27B-v1 is INSANE!

一句話總結

獨立實驗室 Emperor AI 發布了開源模型 QFOS 27B(逐字稿中亦出現 Kwaithos、Kwethos、Coithos 等變體),該模型具備罕見的「三合一」能力:超過 100 萬 token 的上下文窗口、原生視覺處理能力,以及透過多 Token 預測頭(MTP)實現的高速推理。模型採用 Apache 2.0 許可證,但為無審查(Uncensored)版本,目前為強化學習前的檢查點。

核心重點

  1. 模型發布與開發者
  • 開發者為獨立實驗室 Emperor AI(逐字稿中亦出現 Emperor、Emperos、Ampra、Para 等稱呼)。
  • 發布的模型主要稱為 QFOS 27B,但逐字稿中多次出現名稱不一致的情況,包括 Kwaithos 27bKwethos 27BCoithos 27B 以及 Emperos QFOS 27B
  • 該模型是基於前代熱門模型 QFOS 9 billion parameter model 及架構 QFOS 3.5 27B 進行開發。
  1. 罕見的「三合一」技術架構
  • 極大上下文窗口:支援超過 100 萬 token 的記憶容量。
  • 原生視覺能力:保留完整的 Vision tower,可直接處理截圖、圖表及手寫內容。
  • 高速推理:採用原生多 Token 預測頭(MTP, Multi-Token Prediction),實現類似自動補全的快速生成速度。
  • 影片強調獨立開發者通常需犧牲其中一項以節省運算或記憶體,但 Emperor AI 未做妥協。
  1. 訓練與技術細節
  • 訓練架構基於 QFOS 3.5 27B。
  • 執行全參數監督微調(full-parameter supervised fine-tuning)以規範結構。
  • 執行直接偏好優化(direct preference optimization)以提升回答品質。
  • 最後階段使用 ESFT(需查證具體含義)。
  • 使用 Yarn scaling 技術拉伸位置嵌入(positional embeddings),以實現長序列處理。
  1. 許可證與安全性
  • 採用 Apache 2.0 開源許可證,完全免費且可商用,無企業費用或隱藏上限。
  • 明確標示為 Uncensored(無審查) 模型,預設防護欄較少。
  • 使用者需自行負責安全控制,特別是用於面向客戶的產品時。
  1. 版本狀態與未來路線圖
  • 目前版本為 Version 1 Pre-RL checkpoint(強化學習前的檢查點)。
  • 尚未經過最終的 RLHF(人類反饋強化學習)打磨,因此在聊天性、順從性及特定格式指令的精確度上可能不如最終版。
  • Emperor AI 展示了後續路線圖,包括專項編碼代理 Abacus 及混合專家模型(MOE)Clair

詳細大綱

Section 1: 認識 Emperos QFOS 27B

  • 開發團隊背景:Emperor AI 是一個專注於開源權重(open-weight)模型的獨立實驗室。
  • 模型定位:QFOS 27B 是其熱門的 90 億參數模型的更大、更強大版本。
  • 訓練流程
  • 基於 QFOS 3.5 27B 架構。
  • 執行全參數監督微調(full-parameter supervised fine-tuning)。
  • 執行直接偏好優化(direct preference optimization)。
  • 最後階段使用 ESFT。
  • 資源投入:拒絕為節省運算成本而妥協,保留了基礎模型的所有能力。

Section 2: 百萬字詞記憶(Context Window)

  • 技術實現:使用 Yarn scaling 技術,拉伸模型的位置嵌入(positional embeddings),使其能處理極長序列而不降低品質。
  • 容量規模:超過 100 萬 token(約等於一整個圖書館的資料)。
  • 應用優勢
  • 解決傳統 AI 忘記早期指令的問題。
  • 可同時處理數百頁會議記錄、公司程式碼庫或所有客戶對話。
  • 提供具備完整上下文感知的精準回答。

Section 3: 視覺能力與速度

  • 視覺能力(Vision)
  • 訓練過程中保留完整的 Vision tower。
  • 可識別截圖、複雜圖表、雜亂手寫。
  • 應用:減少手動數據輸入,直接從白板或財務圖表中提取意義。
  • 推理速度(Speed)
  • 傳統模型逐字預測,速度慢。
  • QFOS 27B 使用 原生多 token 預測頭(MTP, Multi-Token Prediction)
  • 機制:像高度準確的深度學習自動補全,一次預測多個詞塊。
  • 效果:透過推測解碼(speculative decoding)大幅節省推論時間,實現「閃電般」的速度。
  • 罕見的三合一(Rare Trifecta)
  • 同時具備 MTP 頭、Vision tower 和完整上下文窗口。
  • 獨立開發者通常需犧牲其中一項以節省運算/記憶體,Emperor AI 未做妥協。

Section 4: 商業應用案例

  • 場景:社群媒體經理的每週分析審查。
  • 操作流程
  1. 一次性輸入 20 篇最佳貼文(包含圖片和互動指標)。
  2. 利用大上下文窗口和視覺能力,瞬間消化一個月內容規劃。
  3. 識別受眾共鳴的話題。
  4. 即時草擬針對常見客戶問題的客製化回覆。
  • 價值:將耗時的數據複製與分析轉化為秒級處理,改變日常工作流程。

Section 5: 細則與後續步驟(Fine Print & Next Steps)

  • 許可證(License)
  • 採用 Apache 2.0 開源許可證。
  • 無商業使用限制,無企業費用,無隱藏使用上限。
  • 允許私人商業使用、商業產品整合及獲利。
  • 安全警示(Safety)
  • 明確標示為 Uncensored(無審查) 模型。
  • 預設防護欄(guardrails)較少,提供原始未過濾輸出。
  • 使用者責任:若用於面向客戶的產品,需自行開發應用層級的安全控制(如次要審核端點、嚴格系統提示詞)。
  • 版本狀態
  • 目前為 Version 1 Pre-RL checkpoint(強化學習前的檢查點)。
  • 尚未經過最終的 RLHF(人類反饋強化學習)打磨,因此在聊天性、順從性及特定格式指令的精確度上可能不如最終版。
  • 預期未來會有更精煉的 Version 2。
  • Emperor AI 路線圖
  • Abacus:專項終端編碼代理(terminal coding agent)。
  • Clair:內部訓練的小型混合專家模型(Mixture of Experts, MOE)。
  • MOE 優勢:將查詢路由至專用子網路,節省運算並提升智能。
  • 意義:獨立實驗室成功建構 MOE 模型顯示其雄心與潛力。

工具 / 模型 / 名詞整理

  • Emperor AI:開發該模型的獨立實驗室(逐字稿中亦出現 Emperor、Emperos、Ampra、Para 等稱呼)。
  • QFOS 27B:主要討論的模型名稱。
  • Emperos QFOS 27B:逐字稿中出現的模型變體名稱。
  • Kwaithos 27b:逐字稿中出現的模型變體名稱。
  • Kwethos 27B:逐字稿中出現的模型變體名稱。
  • Coithos 27B:逐字稿中出現的模型變體名稱。
  • QFOS 3.5 27B:作為基礎架構的版本。
  • QFOS 9 billion parameter model:該實驗室之前的熱門模型。
  • Yarn scaling:用於擴展上下文窗口的技術。
  • MTP (Multi-Token Prediction):原生多 token 預測頭。
  • ESFT:訓練流程的最後階段技術(需查證具體含義)。
  • Apache 2.0:模型使用的許可證。
  • RLHF:人類反饋強化學習(文中提及為未來版本可能包含的步驟)。
  • Abacus:專項終端編碼代理(terminal coding agent)。
  • Clair:小型混合專家模型(MOE)。
  • MOE (Mixture of Experts):混合專家模型架構。
  • full-parameter supervised fine-tuning:全參數監督微調。
  • direct preference optimization:直接偏好優化。
  • speculative decoding:推測解碼。

操作流程整理

商業應用案例:社群媒體經理的每週分析審查

  1. 輸入數據:一次性輸入 20 篇最佳貼文(包含圖片和互動指標)。
  2. 處理內容:利用大上下文窗口和視覺能力,瞬間消化一個月內容規劃。
  3. 分析洞察:識別受眾共鳴的話題。
  4. 生成回覆:即時草擬針對常見客戶問題的客製化回覆。
  • 結果:將耗時的數據複製與分析轉化為秒級處理,改變日常工作流程。

值得注意的限制或風險

  1. 無審查(Uncensored)
  • 模型明確標示為無審查,預設防護欄較少,提供原始未過濾輸出。
  • 使用者需自行負責安全控制,特別是用於面向客戶的產品時,需開發應用層級的安全控制(如次要審核端點、嚴格系統提示詞)。
  1. 版本狀態限制
  • 目前為 Version 1 Pre-RL checkpoint(強化學習前的檢查點)。
  • 尚未經過最終的 RLHF(人類反饋強化學習)打磨。
  • 在聊天性、順從性及特定格式指令的精確度上可能不如最終版。
  1. 名稱辨識不確定性
  • 模型名稱與實驗室名稱在逐字稿中存在大量不一致,可能影響搜尋與確認官方資源。

逐字稿辨識疑點

  • 模型名稱不一致
  • 逐字稿中模型名稱多次變化,包括 Emperos QFOS 27BQFOS 27BKwaithos 27bKwethos 27BCoithos 27B。需查證正確官方名稱。
  • 實驗室名稱拼寫
  • 文中出現 Emperor AIEmperos AIEmperoAmpraPara。需查證實驗室正確名稱。
  • 技術術語拼寫
  • ESFT:需查證此縮寫在該上下文中的具體含義(通常為 SFT 或 DPO 之後的步驟,常見縮寫可能為 E-SFT 或其他變體)。
  • yarn scaling:需查證此技術的標準名稱(通常指 YaRN 或相關位置嵌入擴展技術)。
  • full-parameter supervised fine-tuning:通常簡稱為 SFT,文中描述為全參數微調。
  • direct preference optimization:通常簡稱為 DPO。
  • 產品名稱
  • Abacus:專項編碼代理名稱。
  • Clair:MOE 模型名稱。
  • 語音辨識疑似錯誤
  • Emperos vs Emperor
  • Kwaithos / Kwethos / CoithosQFOS 之間的差異。
  • AmpraEmperor 的混用。
  • ParaEmperor 的混用。

可延伸追問

  1. Emperor AI 實驗室發布的模型正確官方名稱為何?(QFOS、Kwaithos、Kwethos 或 Coithos?)
  2. ESFT 技術在該模型訓練中的具體定義與作用為何?
  3. Yarn scaling 技術的具體實現方式及其對模型性能的實際影響數據為何?
  4. 預計何時發布經過 RLHF 打磨的 Version 2?
  5. Abacus 編碼代理與 Clair MOE 模型的具體發布時間表與功能細節為何?

生字列表

生字讀音類型中文
unpacking/ʌnˈpækɪŋ/拆解、深入分析
powerhouse/ˈpaʊərhaʊs/強悍的實體、 powerhouse(在此指能力極強的事物)
toe-to-toe/tuː tuː tuː/正面競爭、不相上下
bombarded/bɒmˈbɑːrdɪd/不斷被轟炸、受到大量衝擊
vaporware/ˈveɪpəwɛr/虛構產品、空頭承諾的產品
natively/ˈneɪtɪvli/原生地、本質上
breaking a sweat/breɪkɪŋ ə swɛt/毫不費力、輕鬆應對
laser-focused/ˈleɪzər ˈfoʊkəst/專注度極高的、目標明確的
cut corners/kʌt ˈkɔːrnərz/偷工減料、走捷徑
bogged down/ˈbɒɡɪd daʊn/陷入困境、被拖累
derails/dɪˈreɪlz/使脫軌、打亂(計畫或流程)
friction/ˈfrɪkʃən/摩擦、阻力、繁瑣
agonizingly/ˈæɡənaɪzɪŋli/令人痛苦地、極度緩慢地
trifecta/traɪˈfɛktə/三重奏、三項全能
compromises/ˈkɒmprəmaɪzɪz/妥協、折衷
gated/ɡeɪtɪd/受限的、有門檻的
heads up/hɛdz ʌp/預警、提醒
roadmap/ˈroʊdmæp/路線圖、發展藍圖

生字解說

unpacking /ʌnˈpækɪŋ/

· C1

意思:拆解、深入分析

解說:原意為「打開包裹」,在此比喻將複雜的概念或技術細節逐一拆解並解釋清楚。

影片原句
Today, we are unpacking how a small independent lab,
今天,我們來拆解一個小型獨立實驗室
延伸例句
Let's unpack the key features of this new software.
讓我們來拆解這個新軟體的關鍵功能。

powerhouse /ˈpaʊərhaʊs/

· B2

意思:強悍的實體、 powerhouse(在此指能力極強的事物)

解說:原意為發電廠,引申為具有強大能量或能力的人或事物。

影片原句
Emperor AI, just dropped an absolute powerhouse of a model that's going toe-to-toe with big tech.
Emperor AI 剛剛發布了一款絕對強悍的模型,它將與大型科技公司正面競爭。
延伸例句
She is a powerhouse in the field of quantum computing.
她是量子計算領域的一位強悍人物。

toe-to-toe /tuː tuː tuː/

· C1

意思:正面競爭、不相上下

解說:形容雙方勢均力敵,進行激烈的競爭或對抗。

影片原句
Emperor AI, just dropped an absolute powerhouse of a model that's going toe-to-toe with big tech.
Emperor AI 剛剛發布了一款絕對強悍的模型,它將與大型科技公司正面競爭。
延伸例句
The two candidates are going toe-to-toe in the final debate.
兩位候選人在最後一場辯論中正面競爭。

bombarded /bɒmˈbɑːrdɪd/

· B2

意思:不斷被轟炸、受到大量衝擊

解說:原意為炮轟,在此比喻被大量的資訊、廣告或訊息不斷衝擊。

影片原句
We get bombarded with hype cycles, PR stunts, and vaporware that you can't even use.
我們不斷被炒作週期、公關噱頭,以及那些你甚至無法使用的虛構產品所轟炸。
延伸例句
Customers are bombarded with promotional emails every day.
顧客每天都會收到大量的促銷電子郵件。

vaporware /ˈveɪpəwɛr/

· C1

意思:虛構產品、空頭承諾的產品

解說:指已經宣傳但尚未發布,或實際上無法運行的軟體或硬體產品。

影片原句
We get bombarded with hype cycles, PR stunts, and vaporware that you can't even use.
我們不斷被炒作週期、公關噱頭,以及那些你甚至無法使用的虛構產品所轟炸。
延伸例句
The company announced a revolutionary phone, but it turned out to be vaporware.
該公司宣布了一款革命性的手機,但結果證明是虛構產品。

natively /ˈneɪtɪvli/

· C1

意思:原生地、本質上

解說:指系統或軟體內建支援某種功能,無需額外安裝外掛或轉換層。

影片原句
We're talking about a highly capable model that runs completely locally, for free, natively understands images,
我們談論的是一個高度強大的模型,它可以完全在本地運行,免費,原生理解圖像,並且
延伸例句
This application natively supports dark mode.
此應用程式原生支援深色模式。

breaking a sweat /breɪkɪŋ ə swɛt/

· B2

意思:毫不費力、輕鬆應對

解說:字面意為「流汗」,比喻處理高難度任務時輕鬆自如,不需付出額外努力。

影片原句
can hold an immense amount of data in its memory all at once, without breaking a sweat.
能夠一次性在記憶體中容納海量數據,毫不費力。
延伸例句
He solved the complex equation without breaking a sweat.
他毫不費力地解開了這個複雜的方程式。

laser-focused /ˈleɪzər ˈfoʊkəst/

· C1

意思:專注度極高的、目標明確的

解說:形容人或團隊將全部精力集中在單一目標上,排除其他干擾。

影片原句
Emperos AI is a laser-focused, independent lab making some serious ways with open-weight models.
Emperos AI 是一個專注度極高的獨立實驗室,正在開權重模型領域取得重大進展。
延伸例句
The team is laser-focused on delivering the project by Friday.
團隊專注於在週五前完成專案交付。

cut corners /kʌt ˈkɔːrnərz/

· B2

意思:偷工減料、走捷徑

解說:指為了節省時間、金錢或精力而省略必要的步驟或降低標準。

影片原句
They refused to cut corners, meaning every single capability from that base model was kept perfectly intact.
他們拒絕偷工減料,意味著基礎模型中的每項能力都完美地保持完整。
延伸例句
Don't cut corners on safety checks.
不要在安全檢查上偷工減料。

bogged down /ˈbɒɡɪd daʊn/

· C1

意思:陷入困境、被拖累

解說:指因過於複雜或繁瑣的細節而無法前進或理解。

影片原句
Without getting totally bogged down in the math, yarn scaling essentially stretches the model's positional embeddings.
不必完全陷入數學細節,Yarn Scaling本質上擴展了模型的位置嵌入。
延伸例句
The project was bogged down by bureaucratic red tape.
專案因官僚主義的繁文縟節而陷入困境。

derails /dɪˈreɪlz/

· B2

意思:使脫軌、打亂(計畫或流程)

解說:原意為火車脫軌,比喻使計畫、對話或工作流程中斷或失敗。

影片原句
It's incredibly frustrating and completely derails your workflow.
但擁有如此大的上下文視窗,這個問題就消失了。
延伸例句
Unexpected technical issues derailed our launch schedule.
意外的技術問題打亂了我們的發布計畫。

friction /ˈfrɪkʃən/

· B2

意思:摩擦、阻力、繁瑣

解說:在此比喻工作流程中因不便或複雜步驟而產生的阻礙或麻煩。

影片原句
Think about how much friction there is in manual data entry.
想像一下手動輸入資料有多麼繁瑣。
延伸例句
Removing friction from the user experience is crucial for adoption.
消除使用者體驗中的阻力對於採用至關重要。

agonizingly /ˈæɡənaɪzɪŋli/

· C1

意思:令人痛苦地、極度緩慢地

解說:形容某種狀態(如等待或速度)讓人感到焦慮或痛苦。

影片原句
It works, but it can be agonizingly slow.
這確實有效,但速度可能慢得令人痛苦。
延伸例句
The process was agonizingly slow, taking days to complete.
這個過程慢得令人痛苦,花了幾天才完成。

trifecta /traɪˈfɛktə/

· C1

意思:三重奏、三項全能

解說:原指賭馬中連續贏得前三名的彩券,在此比喻三個優秀特徵的完美結合。

影片原句
This right here is what we call the rare trifecta. An MTP head, a vision tower, and a full context window.
這裡正是我們所稱的稀有三重奏。一個 MTP 頭、一個視覺塔,以及完整的上下文視窗。
延伸例句
The car offers speed, luxury, and efficiency—a true trifecta.
這輛車提供了速度、豪華和效率——真正的三重奏。

compromises /ˈkɒmprəmaɪzɪz/

· B2

意思:妥協、折衷

解說:指為了達成目標或節省成本而做出的讓步或犧牲。

影片原句
Usually when independent devs fine-tune a model this big, they have to make really painful compromises just to save on compute and memory.
通常,當獨立開發人員微調這麼大的模型時,他們必須做出非常痛苦的妥協,以節省運算和記憶體。
延伸例句
We had to make some compromises on the design to meet the deadline.
為了趕上截止日期,我們不得不在設計上做一些妥協。

gated /ɡeɪtɪd/

· C1

意思:受限的、有門檻的

解說:指資源或功能被限制在特定群體或條件下,無法自由獲取。

影片原句
If you are tired of the restrictive licenses big tech often pushes, where things are gated behind non-commercial use-only clauses or massive enterprise fees,
如果你厭倦了大型科技公司經常推銷的限制性授權,那些授權往往被非商業用途條款或龐大的企業費用所限制,
延伸例句
The premium features are gated behind a subscription paywall.
高級功能被訂閱付費牆所限制。

heads up /hɛdz ʌp/

· B2

意思:預警、提醒

解說:指提前告知某人關於潛在問題或重要資訊。

影片原句
Now, you do need a quick heads up on safety.
現在,你需要快速了解安全方面的注意事項。
延伸例句
Thanks for the heads up about the meeting change.
謝謝你提醒我會議時間變更。

roadmap /ˈroʊdmæp/

· B2

意思:路線圖、發展藍圖

解說:指產品或技術未來的發展計畫和目標時間表。

影片原句
And looking at Empero's roadmap, they are moving fast.
從 Empero 的路線圖來看,他們的進展迅速。
延伸例句
The company released a detailed roadmap for the next three years.
公司發布了未來三年的詳細發展路線圖。

句型解說(含實例)

just dropped a [noun]

意思:剛剛發布了一款... / 剛剛推出...

接續:Subject + just dropped + a/an + adjective + noun

解說:用於描述科技公司或開發者剛剛發布新產品或更新,強調其新穎性和重要性。

影片原句
Emperor AI, just dropped an absolute powerhouse of a model that's going toe-to-toe with big tech.
Emperor AI 剛剛發布了一款絕對強悍的模型,它將與大型科技公司正面競爭。
實例
  1. Apple just dropped a new iPhone model with advanced AI features.
    蘋果公司剛剛發布了一款具有先進 AI 功能的新 iPhone 型號。
  2. The startup just dropped a revolutionary app for remote work.
    這家初創公司剛剛發布了一款用於遠端工作的革命性應用程式。

runs completely locally, for free, natively [verb]

意思:完全在本地運行,免費,原生地...

接續:runs + adverb (completely) + adverb (locally), + prepositional phrase (for free), + adverb (natively) + verb

解說:用於強調軟體或模型的幾個關鍵優勢:隱私性(本地運行)、成本(免費)和技術架構(原生支援)。

影片原句
We're talking about a highly capable model that runs completely locally, for free, natively understands images, and can hold an immense amount of data in its memory all at once, without breaking a sweat.
我們談論的是一個高度強大的模型,它可以完全在本地運行,免費,原生理解圖像,並且能夠一次性在記憶體中容納海量數據,毫不費力。
實例
  1. This software runs completely locally, for free, natively supports offline mode.
    此軟體完全在本地運行,免費,原生支援離線模式。
  2. The new processor runs completely locally, for free, natively handles encryption.
    新處理器完全在本地運行,免費,原生處理加密。

To put that [noun] into perspective, it's practically like [gerund phrase]

意思:為了讓...更具體,這 practically 就像...

接續:To put that + noun + into perspective, + it's practically like + gerund phrase

解說:用於將抽象的技術指標轉化為具體、易懂的生活比喻,幫助聽眾理解其規模或影響。

影片原句
To put that million tokens into perspective, it's practically like holding an entire library in its head all at once.
為了讓這100萬個token更具體,這 practically 就像同時將整個圖書館裝進它的腦袋裡。
實例
  1. To put that budget into perspective, it's practically like buying a small island.
    為了讓那個預算更具體,這 practically 就像買下一座小島。
  2. To put that speed into perspective, it's practically like traveling across the country in an hour.
    為了讓那個速度更具體,這 practically 就像在一小時內穿越全國。

Instead of [gerund phrase], you just [verb phrase]

意思:與其...,你只需...

接續:Instead of + gerund phrase, + you just + verb phrase

解說:用於對比傳統繁瑣的方法與新技術帶來的簡便操作,強調效率提升。

影片原句
Instead of endless scrolling and manually copying data into a spreadsheet, you just feed the model 20 of your top posts all at once,
與其無止盡地滑動螢幕,並手動將資料複製到試算表中,你只需將你的前 20 篇熱門貼文一次餵給模型,
實例
  1. Instead of writing code from scratch, you just use the AI assistant.
    與其從頭開始編寫程式碼,你只需使用 AI 助手。
  2. Instead of waiting in line, you just order online.
    與其排隊,你只需線上訂購。

It tells you exactly what [clause] and helps you [verb phrase]

意思:它明確告訴你...,並幫助你...

接續:It tells you exactly + what-clause + and + helps you + verb phrase

解說:用於描述 AI 模型的具體功能:提供精確的分析結果,並輔助用戶進行後續操作。

影片原句
It tells you exactly what topics resonated most with your audience and helps you instantly draft clear, perfectly tailored replies to common customer questions.
它會明確告訴你哪些主題最能引起觀眾共鳴,並幫助你立即草擬清晰且完美客製化的回覆,以應對常見的客户問題。
實例
  1. The dashboard tells you exactly what users are clicking and helps you optimize the layout.
    儀表板明確告訴你用戶點擊了什麼,並幫助你優化佈局。
  2. The report tells you exactly where the budget is spent and helps you cut unnecessary costs.
    報告明確告訴你預算花在哪裡,並幫助你削減不必要的成本。

Unlike [noun phrase] that [relative clause], this [noun] [verb phrase]

意思:與...不同,這個......

接續:Unlike + noun phrase + that + relative clause, + this + noun + verb phrase

解說:用於對比傳統商業產品與新開源產品的差異,強調新產品的優勢(如未經過濾、自由度高)。

影片原句
Unlike commercial models from major corporations that heavily filter their outputs right out of the box, this gives you raw unfiltered power.
與大型公司提供的商業模型不同,後者在出廠時就會對輸出進行嚴格過濾,這個模型則賦予你原始且未經過濾的力量。
實例
  1. Unlike traditional cars that require frequent maintenance, this electric vehicle needs almost no upkeep.
    與需要頻繁維護的傳統汽車不同,這款電動車幾乎不需要保養。
  2. Unlike open-source software that is free, proprietary software often has high licensing fees.
    與免費的開源軟體不同,專有軟體通常有高額的授權費。

if a [noun phrase] can [verb phrase], how soon until [clause]?

意思:如果...能夠...,那麼...還需要多久?

接續:if + a + noun phrase + can + verb phrase, + how soon until + clause?

解說:用於提出反思性問題,強調技術進步的速度和對現有格局的衝擊。

影片原句
Looking at everything Coithos 27B brings to the table, it forces us to ask, if a small, independent lab can deliver a 1 million token memory, native vision, and multi-token prediction entirely for free, how soon until the gap between open-source and big tech completely disappears?
回顧 Coithos 27B 所帶來的一切,它迫使我們思考,如果一個小型、獨立的實驗室能夠免費提供 100 萬 token 的記憶、原生視覺和多 token 預測功能,那麼開源與大型科技公司之間的差距完全消失還需要多久?
實例
  1. If a small team can build a global platform, how soon until the monopoly breaks?
    如果一個小團隊能夠建立全球平台,那麼壟斷何時會打破?
  2. If renewable energy can be cheaper than fossil fuels, how soon until the transition is complete?
    如果可再生能源比化石燃料更便宜,那麼過渡何時能完成?