Experts: Kimi K3 Development Was Not Based Solely on Distillling Fable
Analysis

Experts: Kimi K3 Development Was Not Based Solely on Distillling Fable

White House science advisor claims Moonshot copied the US model, but AI researchers express doubt over distillation feasibility

4 min read
Based on original reporting byTechCrunchTranslated, summarized and given business context by our systemHow we work

Executive summary

Key Takeaways

  • White House science advisor Michael Kratsios accused Moonshot of copying Anthropic’s Fable model using chips not approved for export to China.

  • Researcher Braden Hancock notes that Fable was released only on July 1, making the distillation and training of Kimi K3 within two weeks timeline-wise impossible.

  • Advanced reinforcement learning required for high-capability distillation could require tens of millions of agents, making frontier lab API usage extremely expensive and slow.

  • Anthropic accused Moonshot, DeepSeek, and MiniMax earlier this year of executing millions of capability-extraction queries against its models.

  • Despite export bans, Nvidia’s advanced Grace Blackwell 300 (GB300) chips are entering China through a black market and servers located in Thailand.

Experts: Kimi K3 Development Was Not Based Solely on Distillling Fable

  • White House science advisor Michael Kratsios accused Moonshot of copying Anthropic’s Fable model using chips...
  • Researcher Braden Hancock notes that Fable was released only on July 1, making the distillation...
  • Advanced reinforcement learning required for high-capability distillation could require tens of millions of agents, making...
  • Anthropic accused Moonshot, DeepSeek, and MiniMax earlier this year of executing millions of capability-extraction queries...
  • Despite export bans, Nvidia’s advanced Grace Blackwell 300 (GB300) chips are entering China through a...

According to a report in TechCrunch, artificial intelligence experts are casting doubt on claims that Chinese company Moonshot's large open-weight language model, Kimi K3, achieved its advanced capabilities solely through the distillation of the American company Anthropic's Fable model. The debate arises following allegations by U.S. officials regarding systematic copying and the use of banned chips.

U.S. Administration Allegations and Watermarks

Michael Kratsios, White House science advisor, stated that Moonshot, the company behind the development of the Kimi K3 model—currently the largest available open-weight LLM—built its model by copying Anthropic’s Fable model. According to Kratsios, the company did this while using chips that are not approved for export to China. "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable," Kratsios wrote amid reported discussions regarding a ban on Chinese open-weight models, which are stirring controversy in the AI sector. Moonshot did not respond to questions regarding its training process, and Kratsios did not share further details regarding the sources of his allegations.

Kratsios's comments echoed remarks by Treasury Secretary Scott Bessent, who noted that "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that’s unacceptable." However, it is not clear what these watermarks consist of, and the U.S. Treasury Department did not respond to an inquiry on the matter.

Technological Doubts and Timelines Too Short

Despite the official allegations, AI experts are expressing significant skepticism that distillation—the process of querying a large language model to understand its inner workings and copy its capabilities—is the explanation for the advanced capabilities displayed by Kimi K3.

Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, explained to TechCrunch that timelines simply do not allow for it. "I don’t think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Hancock said. "There’s just not even frankly time, right? Fable's only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks."

Nathan Lambert, an AI researcher at the Allen Institute for AI, expressed a similar view in a recently released podcast. Lambert argued that the impact of distillation is decreasing as Chinese models get closer to the frontier of technology and the training regime shifts to reinforcement learning. According to him, "if it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won’t see this, from supervised fine-tuning (SFT) alone."

Limitations of Fine-Tuning and the Need for Reinforcement Learning

Performing distillation requires a research lab to systematically query the target model to generate data that can be used in post-training stages. Sometimes this explicitly involves asking the model to explain its chain-of-thought to understand how it solves problems. In other cases, the prompts and responses of the model are used to train a new model in a process known as supervised fine-tuning (SFT).

This fine-tuning process is the reason why a model ostensibly built by a third party might claim during a conversation that it is Anthropic's Claude. According to Lambert, this is the stage where the model "picks up its manners." However, Lambert believes that the benefits of SFT are becoming less important as models become more complex.

To distill capabilities similar to those of Fable, reinforcement learning techniques would likely be required. In many cases, this means using an agent of the larger model to grade the smaller model's responses, and adjusting based on the grade given. The more advanced techniques also require highly significant infrastructure; large reinforcement learning runs can require tens of millions of agents. Using a frontier lab's API for this purpose would be insanely expensive and potentially a time bottleneck, because these models are pretty slow and, frankly, might not even provide a performance uplift.

History of Claims and Common Industry Practice

Nevertheless, it appears that previous frontier models might have contributed to Moonshot's developments. Earlier this year, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models. Anthropic claimed it identified millions of exchanges between its models and users identified by IP addresses and metadata of these companies. These queries were described as distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use. Anthropic did not respond to queries from TechCrunch regarding Fable distillation.

At the same time, distillation is considered common among many AI companies, and not just in China. Elon Musk testified earlier this year that his company, SpaceXAI, distilled OpenAI models to develop Grok, adding that the practice was common in the industry. The line between distillation and developing synthetic datasets can be fairly blurry.

Hancock added that, in general, Americans are understating the technical expertise of these Chinese teams: "One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. If American models ground to a halt, I think China's progress would slow, but would still continue. They’re not just riding coattails here."

Chip Smuggling Routes and Data Center Oversight

It is hard to separate distillation claims from the second part of Kratsios's allegations—the claim that Moonshot obtained advanced Nvidia Grace Blackwell 300 chips (known as GB300), and also accessed servers equipped with GB300 chips in Thailand. The export of these chips to China is banned, but according to Sam Bresnick, a research fellow at Georgetown University's Center for Security and Emerging Technology (CSET), a black market exists for them.

In May, the founder of Supermicro, an American server builder, was indicted for smuggling advanced chips into China. Bresnick noted that he is a proponent of "know your customer" (KYC) laws for data centers across the world: "If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing."

President Joe Biden's Department of Commerce proposed federal "know your customer" rules for data centers in 2024, but no further progress appears to have been made under Donald Trump's administration. However, exporters shipping advanced chips abroad are required to ensure they are only used for approved purposes.

Questions & Answers

FAQ

This article was produced by our AI-assisted system: translation, summarization and business context based on original reporting by TechCrunch. Read about our editorial process. Link to the original source.

Enjoyed the article?

Subscribe to our newsletter for the latest AI updates straight to your inbox

More from TechCrunch

All articles from TechCrunch
כיצד מגבלות הבטיחות של הבינה המלאכותית מקשות על חוקרי סייבר התקפי
חדשות
5 דקות
מ־TechCrunch

כיצד מגבלות הבטיחות של הבינה המלאכותית מקשות על חוקרי סייבר התקפי

דיווח חדש מאתר TechCrunch חושף כי מגבלות הבטיחות (guardrails) המחמירות שהטילו ענקיות הבינה המלאכותית כדי למנוע שימוש לרעה במודלים שלהן, פוגעות כעת דווקא בחוקרי אבטחת מידע לגיטימיים ובמגני רשתות. חוקרים בתחום הסייבר ההתקפי, המשתמשים בכלים אלו לאיתור חולשות, מדווחים על חסימות תכופות וחוסר עקביות במענה של המערכות, אפילו במסגרת תוכניות סינון מיוחדות של OpenAI ו-Anthropic. בעקבות המגבלות והחשש מדליפת מידע לענן, חוקרים רבים נאלצים לנטוש את המערכות האמריקאיות המפוקחות ולעבור לשימוש במודלים מקומיים של קוד פתוח, לעיתים של חברות זרות כגון מודל GLM הסיני. מומחים מזהירים כי המגבלות הנוכחיות עלולות לפגוע ביכולת ההגנה של המערב אל מול גל מתקפות סייבר עתידי.

קרא עוד
AMD קוראת תיגר על Nvidia עם מערכת המסד Helios למחשוב AI
מוצר חדש
3 דקות
מ־TechCrunch

AMD קוראת תיגר על Nvidia עם מערכת המסד Helios למחשוב AI

לפי דיווח של TechCrunch, חברת השבבים AMD משיקה את מערכת ה-Helios, מערכת שרתים שלמה במסד (rack-scale system) המיועדת להתחרות ישירות במובילת השוק Nvidia. המערכת החדשה, שתוכננה עבור מעבדות ה-AI הגדולות בעולם כדי לאמן ולהריץ את מודלי הקצה התובעניים ביותר, כבר גייסה רשימה של לקוחות ענק הכוללת את מיקרוסופט, OpenAI, Meta, Oracle ו-Anthropic. מיקרוסופט מתכננת להרחיב את תשתית הענן Azure באמצעותה, ואילו Anthropic הכריזה על שיתוף פעולה לפריסת עד שני גיגוואט של מעבדים גרפיים. בנוסף, AMD הציגה את מעבד ה-Venice-X המיועד למרכזי נתונים שיושק ב-2027. מנכ"לית AMD, ד"ר ליסה סו, מעריכה כי העלייה בביקוש למחשוב המונעת על ידי בינה מלאכותית סוכנתית (agentic AI) תוביל את שוק מאיצי ה-AI לשווי של כ-1.4 טריליון דולר עד שנת 2030.

קרא עוד
Runway משיקה מנתב מודלים של בינה מלאכותית למדיה גנרטיבית
מוצר חדש
4 דקות
מ־TechCrunch

Runway משיקה מנתב מודלים של בינה מלאכותית למדיה גנרטיבית

חברת הסטארט-אפ Runway משיקה את Runway Media Router, כלי חדש המאפשר למפתחים לנתב באופן אוטומטי בקשות יצירת תמונה, וידאו ואודיו למודל המתאים ביותר. הכלי, שהושק באמצעות פלטפורמת המפתחים Runway Dev, מנתב את הבקשות על פי העדפות המפתח לגבי איכות, מהירות או עלות. המהלך מסמן את שאיפתה של Runway להפוך לשכבת תשתית ואורקסטרציה עבור תעשיית המדיה הגנרטיבית, בתקופה שבה השוק נעשה צפוף ותחרותי במיוחד והחברה מתמודדת עם תחרות עזה מצד ענקיות טכנולוגיה כמו גוגל, בייטדאנס ועליבאבא.

קרא עוד
אנבידיה שולחת מעבדים גרפיים לירח: שבבי Jetson של החברה בחלל
חדשות
3 דקות
מ־TechCrunch

אנבידיה שולחת מעבדים גרפיים לירח: שבבי Jetson של החברה בחלל

לפי דיווח באתר TechCrunch, חברת הסטארט-אפ Lunar Outpost תשתמש בשבבי Nvidia Jetson כדי לשלוט במערכת הלידאר של רכב השטח הירחי הבא שלה, מה שצפוי להפוך אותו למעבד הגרפי (GPU) הראשון על פני השטח של הירח. השימוש בפלטפורמת השבבים הזו נועד לסייע לרכבי השטח לעבד נתונים מחיישנים באופן מקומי ולקבל החלטות מהירות בסביבות הקיצוניות של הירח. בנוסף, אנבידיה משתפת פעולה עם Firefly Aerospace להפעלת לוויין עיבוד תמונות במסלול סביב הירח. המשימות מתוכננות לשיגור על גבי טילי Falcon 9 של SpaceX לפני סוף השנה הנוכחית, כחלק מהמאמצים הרחבים של נאס"א וחברות פרטיות לבסס נוכחות מתמשכת על הירח לקראת חזרת אסטרונאוטים המתוכננת לשנת 2028.

קרא עוד

More articles you might like

All articles
בינה מלאכותית ועלייתן של אפליקציות הבידור האוניברסליות
ניתוח
4 דקות
מ־TechCrunch

בינה מלאכותית ועלייתן של אפליקציות הבידור האוניברסליות

המאבק בעולם אפליקציות הבידור משתנה: פלטפורמות כמו נטפליקס, ספוטיפיי, יוטיוב וטיקטוק אינן מסתפקות עוד בפורמט תוכן יחיד. הן שואפות להפוך לאפליקציות בידור אוניברסליות המרכזות מוזיקה, וידאו, פודקאסטים, משחקים וקניות תחת קורת גג אחת, במטרה להשתלט על הזמן הפנוי של המשתמשים ולמנוע מעבר לפלטפורמות מתחרות. הבינה המלאכותית משחקת תפקיד מרכזי במהפכה זו, החל משיפור המלצות והתאמה אישית של תכנים במגוון פורמטים, דרך האצת תהליכי פיתוח קוד, ועד להפקת תוכן יוצר וייעול כלי פרסום.

קרא עוד
וורטו מציגה את אלפאפולד: האם סוכן AI מצדיק תג מחיר של 6,880 דולר?
ניתוח
6 דקות
מ־TechCrunch

וורטו מציגה את אלפאפולד: האם סוכן AI מצדיק תג מחיר של 6,880 דולר?

סקירה מקיפה של מכשיר ה-Vertu Alphafold החדש, שמחירו מתחיל ב-6,880 דולר ומיועד למנהלים בכירים. הבדיקה שנערכה על ידי TechCrunch בחנה את תפקודו של סוכן הבינה המלאכותית המובנה, Hermes Agent, במשימות ניהוליות יומיומיות כמו שליחת הודעות בזמן אמת, תכנון נסיעות וניתוח מסמכים פיננסיים, בהשוואה לסייען ה-Gemini ב-Samsung Galaxy Z Fold 7. הסקירה חושפת כי ה-Hermes מציג נטייה לפעול באופן עצמאי אך סובל מחוסר עקביות וטעויות דיוק, בעוד החומרה של המכשיר מבוססת במידה רבה על מכשיר ה-ZTE Nubia Fold הזול משמעותית. למרות מעטפת היוקרה ושבב האבטחה הייעודי A5, קשה להצדיק את תוספת המחיר הגבוהה מול החלופות הבשלות בשוק.

קרא עוד
משלוח האופניים שאבד והמאבק המתיש בצ'אטבוטים של שירות הלקוחות
ניתוח
5 דקות
מ־Wired

משלוח האופניים שאבד והמאבק המתיש בצ'אטבוטים של שירות הלקוחות

כתבה במגזין WIRED מתארת את חוויותיו המתישות של העיתונאי דילון תומפסון, אשר ניסה לאתר משלוח של אופניים חשמליים בשווי 2,000 דולר שנעלמו, ומצא את עצמו לכוד במשך חודשים ב"גיהנום של צ'אטבוטים". הכתבה מפרטת כיצד חברות משתמשות בבינה מלאכותית ובטקטיקות של "בוץ" (sludge) דיגיטלי המייצרות חיכוך מכוון כדי למנוע גישה לנציגים אנושיים, במקביל לצמצום כוח האדם שבו מדווחים 31% ממנהלי השירות. מומחים מסבירים כי לחצים מצד משקיעים מובילים חברות להשקיע ב-AI מתוך "כשל השקעה שקועה", גם כשהדבר פוגע קשות בחוויית הלקוח ומשטח את רמת השירות הניתנת לצרכנים.

קרא עוד
באילו מקרים כדאי להשתמש ב-Claude Code ובאילו ב-n8n?
ניתוח
5 דקות
מ־n8n

באילו מקרים כדאי להשתמש ב-Claude Code ובאילו ב-n8n?

בפוסט שפורסם בבלוג של n8n על ידי אופיר פרוסאק, נבחנת הדילמה בין שימוש ב-Claude Code לבין n8n לבניית אוטומציות. פרוסאק, המשתמש בשני הכלים מדי יום, מסביר כי לא מדובר בבחירה בלעדית אלא בכלים משלימים. המענה לשאלה תלוי בחמישה משתנים: אופי התהליך, הגורם שמקבל החלטות (חוקים דטרמיניסטיים או AI), בעלי התפקידים המעורבים בתחזוקה, דרישות ההרצה והאמינות (במיוחד בקנה מידה רחב), וההשלכות של כשלים (מהירות מול סובלנות לסיכונים). במקרים מורכבים ובעלי סיכון, מומלץ לשלב ביניהם על ידי בניית ה-workflow ב-n8n ושימוש ב-Claude Code עם שרת ה-MCP של n8n כדי להאיץ את תהליך הפיתוח.

קרא עוד