How AI Guardrails Are Impeding Offensive Cybersecurity Researchers
News

How AI Guardrails Are Impeding Offensive Cybersecurity Researchers

Security researchers warn that tech giants' restrictions harm defenders and push them to foreign open-source models

5 min read
Based on original reporting byTechCrunchTranslated, summarized and given business context by our systemHow we work

Executive summary

Key Takeaways

  • In June, the U.S. government imposed export control restrictions on two Anthropic AI models (Mythos and Fable), with the Fable 5 model returning to general access on July 1.

  • Researchers are criticizing the effectiveness of two key vetting programs: OpenAI’s Trusted Access for Cyber and Anthropic's Cyber Verification Program.

  • Security researcher Mark Dowd, who has spent decades finding and selling vulnerabilities to governments, criticizes the arbitrary decisions made by AI companies.

  • Strict guardrails are pushing researchers to switch to locally run open-source models, including China's GLM model, which is free to download.

How AI Guardrails Are Impeding Offensive Cybersecurity Researchers

  • In June, the U.S. government imposed export control restrictions on two Anthropic AI models (Mythos...
  • Researchers are criticizing the effectiveness of two key vetting programs: OpenAI’s Trusted Access for Cyber...
  • Security researcher Mark Dowd, who has spent decades finding and selling vulnerabilities to governments, criticizes...
  • Strict guardrails are pushing researchers to switch to locally run open-source models, including China's GLM...

According to a report by TechCrunch by Lorenzo Franceschi-Bicchierai, the guardrails and strict vetting programs established by AI giants to prevent malicious hackers from exploiting their models are beginning to complicate the work of legitimate network defenders and offensive cybersecurity researchers. These researchers, whose role is to identify vulnerabilities in systems and develop ways to exploit them before criminal actors do, are running into blocks from vetting mechanisms or facing inconsistency in system responses. These restrictions often lead to the opposite result, forcing researchers to abandon regulated American systems and transition to using open-source models, sometimes from foreign companies, which do not include any vetting.

Export Controls and Vetting Programs of Tech Giants

Over the past several months, AI giants have developed special vetted programs and strict guardrails to limit the use of their models by malicious hackers. In June, the U.S. government imposed export control restrictions on Anthropic's much-hyped AI models, known as Mythos and Fable. This step was prompted, at least in part, by a report claiming that it was possible to bypass the guardrails of these models, which were designed to prevent users from using them to build and execute malicious cyberattacks.

Regardless of whether the incident was indeed motivated by fears of a jailbreak, the fact is that Anthropic has repeatedly marketed the Mythos model as some kind of "doomsday cybermachine" that can only be given to carefully vetted users, and even then with strict guardrails implemented in practice. (The export controls on Fable 5 and Mythos 5 have since been lifted; Fable 5 returned to general access on July 1, while Mythos 5 has been reintroduced only to vetted U.S. organizations as part of the government's review process.)

This type of gatekeeping is not unique to the Mythos model. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers special programs they can apply to in order to get vetted and—if approved—receive access to models with fewer cybersecurity restrictions. OpenAI's program is called "Trusted Access for Cyber", while Anthropic's program is named the "Cyber Verification Program".

Criticism from Offensive Security Researchers

These restrictions have drawn widespread criticism, particularly from researchers whose job is to find unknown security vulnerabilities in systems and find ways to exploit them before cybercriminals do. During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said: "it's not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what's not." Dowd has spent decades finding and selling "zero days"—previously unknown software flaws and the exploits that take advantage of them—to Western governments, rather than reporting them to software makers so they get patched. Governments pay premium prices for these vulnerabilities precisely because they stay open, which aids intelligence operations. Dowd admitted that the nature of his work might make him biased, but he is not the only one holding this view.

"Like a Hammer": A Tool That Is Both Offensive and Defensive

Offensive cybersecurity researchers—those who proactively probe systems to locate weaknesses—described to TechCrunch how they use AI tools and deal with their guardrails. Chris Anley, the chief scientist at security consulting giant NCC Group, noted that asking an AI model to try to exploit a bug is a key step in confirming that a real vulnerability worth fixing actually exists. But if a guardrail causes the model to refuse to answer the question outright, the restriction hurts defenders.

"This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base," Anley explained. "So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can't really be unpicked." Anley compared the tool to a hammer: "You can't build a house without a hammer. It's definitely a tool but it's also irreducibly a weapon as well." When he and his colleagues run into such a roadblock, they sometimes turn to open-source AI models that come with no guardrails at all.

Concerns Over Data Leaks and Moving to Local Models

Paolo Stagno, the chief technology officer (CTO) of CrowdFense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying that AI companies "essentially treat customers like children who need babysitting" with their vetted programs and guardrails. Stagno noted that he and his colleagues use frontier models—but solely for reverse engineering. They avoid using AI to help find vulnerabilities or build exploits because feeding these materials into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future model training runs. For this stage, he said, they use open-source models run locally, as they do not rely on sharing data outside of the model.

Another security researcher, Giuseppe Cali, who finds zero-days and develops exploits, noted that guardrails do not impede his work. This is because he does not use AI for offensive work; instead, he uses it for initial reverse engineering, to understand the code he is analyzing, and to build supporting tools. For these tasks, he explains, AI tools can speed up the process and allow him to focus on discovering vulnerabilities. “I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow," Cali said. "I am jealous of my bugs, and I like this game too much to let models play it for me.”

System Inconsistency and the Push Toward Chinese Models

One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity because he is not authorized to talk to the press, said his employer is not part of Anthropic's CVP program, and as a result, its tools are barely useful for finding vulnerabilities because the guardrails are too strict. "If it catches wind we're doing anything security related, it just stops and isn't usable," the person said.

Chris Thompson, chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con—an offensive security and AI-focused event—shared that in his experience using frontier AI models, the guardrails can be inconsistent and work differently every day. This is true even inside the looser boundaries of Anthropic and OpenAI’s vetted programs.

"I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program," Thompson said. "Instead of analyzing a vulnerability and reasoning through the exploitability, you're trying to find why you're getting inconsistent results or why are models over-sanitizing the output."

As a consequence, researchers rely on or are pushed toward Chinese open-source models like GLM—freely downloadable models that can be run locally with no vetting or usage restrictions, Thompson said. "You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems," he noted. "I think it's more harmful than good to have these guardrails in place."

Rather than tightening restrictions further, Thompson called for the leading AI labs to open up their programs, provide responsible access, and also hold those who abuse their tools accountable. Otherwise, he argues, defenders will lose the AI race. "There's this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before," Thompson said. "But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now."

Questions & Answers

FAQ

This article was produced by our AI-assisted system: translation, summarization and business context based on original reporting by TechCrunch. Read about our editorial process. Link to the original source.

Enjoyed the article?

Subscribe to our newsletter for the latest AI updates straight to your inbox

More from TechCrunch

All articles from TechCrunch
AMD קוראת תיגר על Nvidia עם מערכת המסד Helios למחשוב AI
מוצר חדש
3 דקות
מ־TechCrunch

AMD קוראת תיגר על Nvidia עם מערכת המסד Helios למחשוב AI

לפי דיווח של TechCrunch, חברת השבבים AMD משיקה את מערכת ה-Helios, מערכת שרתים שלמה במסד (rack-scale system) המיועדת להתחרות ישירות במובילת השוק Nvidia. המערכת החדשה, שתוכננה עבור מעבדות ה-AI הגדולות בעולם כדי לאמן ולהריץ את מודלי הקצה התובעניים ביותר, כבר גייסה רשימה של לקוחות ענק הכוללת את מיקרוסופט, OpenAI, Meta, Oracle ו-Anthropic. מיקרוסופט מתכננת להרחיב את תשתית הענן Azure באמצעותה, ואילו Anthropic הכריזה על שיתוף פעולה לפריסת עד שני גיגוואט של מעבדים גרפיים. בנוסף, AMD הציגה את מעבד ה-Venice-X המיועד למרכזי נתונים שיושק ב-2027. מנכ"לית AMD, ד"ר ליסה סו, מעריכה כי העלייה בביקוש למחשוב המונעת על ידי בינה מלאכותית סוכנתית (agentic AI) תוביל את שוק מאיצי ה-AI לשווי של כ-1.4 טריליון דולר עד שנת 2030.

קרא עוד
Runway משיקה מנתב מודלים של בינה מלאכותית למדיה גנרטיבית
מוצר חדש
4 דקות
מ־TechCrunch

Runway משיקה מנתב מודלים של בינה מלאכותית למדיה גנרטיבית

חברת הסטארט-אפ Runway משיקה את Runway Media Router, כלי חדש המאפשר למפתחים לנתב באופן אוטומטי בקשות יצירת תמונה, וידאו ואודיו למודל המתאים ביותר. הכלי, שהושק באמצעות פלטפורמת המפתחים Runway Dev, מנתב את הבקשות על פי העדפות המפתח לגבי איכות, מהירות או עלות. המהלך מסמן את שאיפתה של Runway להפוך לשכבת תשתית ואורקסטרציה עבור תעשיית המדיה הגנרטיבית, בתקופה שבה השוק נעשה צפוף ותחרותי במיוחד והחברה מתמודדת עם תחרות עזה מצד ענקיות טכנולוגיה כמו גוגל, בייטדאנס ועליבאבא.

קרא עוד
אנבידיה שולחת מעבדים גרפיים לירח: שבבי Jetson של החברה בחלל
חדשות
3 דקות
מ־TechCrunch

אנבידיה שולחת מעבדים גרפיים לירח: שבבי Jetson של החברה בחלל

לפי דיווח באתר TechCrunch, חברת הסטארט-אפ Lunar Outpost תשתמש בשבבי Nvidia Jetson כדי לשלוט במערכת הלידאר של רכב השטח הירחי הבא שלה, מה שצפוי להפוך אותו למעבד הגרפי (GPU) הראשון על פני השטח של הירח. השימוש בפלטפורמת השבבים הזו נועד לסייע לרכבי השטח לעבד נתונים מחיישנים באופן מקומי ולקבל החלטות מהירות בסביבות הקיצוניות של הירח. בנוסף, אנבידיה משתפת פעולה עם Firefly Aerospace להפעלת לוויין עיבוד תמונות במסלול סביב הירח. המשימות מתוכננות לשיגור על גבי טילי Falcon 9 של SpaceX לפני סוף השנה הנוכחית, כחלק מהמאמצים הרחבים של נאס"א וחברות פרטיות לבסס נוכחות מתמשכת על הירח לקראת חזרת אסטרונאוטים המתוכננת לשנת 2028.

קרא עוד
סטארטאפ שבבי ה-AI של Etched גייס לפי שווי של 10.3 מיליארד דולר
חדשות
4 דקות
מ־TechCrunch

סטארטאפ שבבי ה-AI של Etched גייס לפי שווי של 10.3 מיליארד דולר

סטארטאפ שבבי הבינה המלאכותית Etched, שהוקם בשנת 2022 על ידי שלושה נושרי הרווארד, השלים סבב גיוס הון של 300 מיליון דולר לפי שווי חברה מרשים של 10.3 מיליארד דולר. סבב הגיוס הובל על ידי קרן Sequoia המפורסמת ובהשתתפות משקיעים בולטים כמו Andreessen Horowitz, SK Hynix, פיטר תיל ואנדריי קארפאטי. בכך הכפילה החברה את שוויה המוערך בתוך שבעה חודשים בלבד, לאחר שהוערכה ב-5 מיליארד דולר בדצמבר האחרון. החברה מפתחת מערכות שבבים ייעודיות לבינה מלאכותית, הכוללות פתרונות חומרה מתקדמים לייעול שלבי ההסקה (prefill ו-decode) במהירות גבוהה ובעלויות נמוכות.

קרא עוד

More articles you might like

All articles
הפרסומת החדשה של מטא ל-AI משתמשת בשיר על סוף העולם
חדשות
4 דקות
מ־Wired

הפרסומת החדשה של מטא ל-AI משתמשת בשיר על סוף העולם

לפי דיווח במגזין WIRED, ענקית הטכנולוגיה מטא פרסמה לאחרונה פרסומת חדשה לקידום טכנולוגיית הבינה המלאכותית שלה באינסטגרם, תחת המסר שהטכנולוגיה לא תשאיר אותנו מאחור. אלא שהמוזיקה המלווה את הסרטון האופטימי היא השיר "Five Years" של דיוויד בואי משנת 1972, העוסק באפוקליפסה ובחורבן כדור הארץ בתוך חמש שנים. בעוד דובר מטא טוען כי בואי ראה בשיר ביטוי לאופטימיות לעתיד טכנולוגי, מומחי מוזיקה והיסטוריונים מצביעים על כך שהשיר נכתב מתוך עצב עמוק ומתאר סיוט דיסטופי מובהק. המקרה מצטרף לשורה של פרסומות מצד ענקיות טכנולוגיה כמו גוגל ואנתרופיק, המנסות להפיג את חששות הציבור מהשפעות הטכנולוגיה.

קרא עוד
מודלי ה-AI הפתוחים מסין מאתגרים את האסטרטגיה של סיליקון ואלי
חדשות
5 דקות
מ־Wired

מודלי ה-AI הפתוחים מסין מאתגרים את האסטרטגיה של סיליקון ואלי

מאמר זה מבוסס על דיווח של מגזין WIRED העוסק באתגר הגובר שמציבות מעבדות בינה מלאכותית סיניות בפני חברות הטכנולוגיה הגדולות בארה"ב. עם השקתם של דגמים מתקדמים בקוד פתוח, כמו Kimi K3 של חברת Moonshot AI ו-GLM 5.2 של Z.ai, חברות בסין מציעות חלופות נגישות ויציבות לדגמים החסומים של OpenAI ו-Anthropic. בעוד שארצות הברית מטילה הגבלות ייצוא ומעכבת השקות בשל חששות בטיחות, דגמי הקוד הפתוח הסיניים מאומצים על ידי חוקרים וסטארט-אפים במערב למשימות מורכבות כמו פיתוח אתרים וניתוח אירועי סייבר, ומציבים סימן שאלה סביב הצורך במימון אינסופי לפיתוח דגמים סגורים.

קרא עוד
גוגל משקיעה 40 מיליון דולר במשימת ג'נסיס הלאומית של ארה"ב
חדשות
3 דקות
מ־DeepMind

גוגל משקיעה 40 מיליון דולר במשימת ג'נסיס הלאומית של ארה"ב

בבלוג הרשמי של גוגל קלאוד הוכרז על הרחבת התמיכה של החברה במשימת ג'נסיס (Genesis Mission) הלאומית של ארצות הברית, באמצעות הקצאה של 40 מיליון דולר באסימוני בינה מלאכותית ובקרדיטים לענן. גוגל תעניק למעבדות הלאומיות של משרד האנרגיה האמריקאי (DOE) גישה לכלי בינה מלאכותית מתקדמים מבית Google DeepMind, בהם AlphaEvolve, AlphaFold 3 ו-AlphaGenome, לצד רישיונות שימוש ב-Gemini for Government למשך שנה עבור עשרות אלפי עובדים. הכלים כבר משולבים במעבדות כגון PNNL ו-NLR, ומסייעים בקיצור זמני כיול חומרה ובמיפוי מערכות מתמטיות מורכבות.

קרא עוד
סטארט-אפ אבטחת הסייבר Glow נחשף עם שווי של 1.2 מיליארד דולר
חדשות
4 דקות
מ־TechCrunch

סטארט-אפ אבטחת הסייבר Glow נחשף עם שווי של 1.2 מיליארד דולר

חברת אבטחת הסייבר Glow יצאה מפעילות בחשאיות (Stealth) כחד-קרן והכריזה על גיוס של 180 מיליון דולר בסבב Series A לפי שווי של 1.2 מיליארד דולר. הסטארט-אפ, שהוקם על ידי יוצאי החברות Meta ו-Snowflake וממוקם בפאלו אלטו, מפתח פלטפורמת אבטחה ייחודית המיועדת להתמודד עם האתגרים החדשים שמציבה הבינה המלאכותית לנקודות קצה בארגונים — החל ממחשבי עובדים ועד שרתים. בעזרת סוכני AI מתמחים, הפלטפורמה מנטרת, מנהלת ואוכפת מדיניות אבטחה בזמן אמת, ומסייעת במניעת חדירה של תוכנות מסוכנות, סוכני AI זדוניים וכלי פיתוח מזיקים לסביבה הארגונית עוד לפני כניסתם. החברה מעסיקה כיום קרוב ל-100 עובדים, רובם הגדול בישראל, וכבר משרתת לקוחות משלמים במגוון תעשיות.

קרא עוד