• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Merged Artificial Intelligence

Bit of a deep dive into the Hugging Face incident, and why it is worrying.

Kurzgesagt, 21:43


Deep dive into the Hugging Face incident and why it is worrying started on this thread at least 26 pages ago.

Companies in the USA, on the other hand, are discovering...


Call me prone to disturbing thoughts... I googled the search term "defense against frontier AI engaged in ransomware and other types of cyberattacks".

As further development of AI intelligence will result in a dramatic increase of offensive capability for attackers.

Defending against frontier AI—which accelerates vulnerability discovery and generates automated, machine-speed exploits—requires moving from static, severity-based patch management to continuous, AI-driven exposure and runtime defense. [1, 2, 3, 4]
Core Defense Strategies
    • Fight AI with AI (Defensive Harnesses): Use frontier models and multi-model agentic harnesses on the defender's side (such as participating in collaborative security programs like Anthropic Project Glasswing) to automatically scan, patch, and harden code before attackers build exploits. [1, 2, 3]
    • Shift to Real Exploitability Over CVSS Scores: Abandon traditional static CVSS prioritization in favor of continuous, inside-out and outside-in exposure validation that measures whether a vulnerability has an actual active exploit path. [1, 2]
    • Implement Zero Standing Privileges and Identity Controls: Enforce just-in-time access, strict segmentation (e.g., Akamai Guardicore or micro-segmentation), and runtime identity verification to limit the lateral "blast radius" if an automated agent breaches the perimeter. [1, 2]
    • Enforce Runtime and API Security: Deploy Web Application Firewalls (WAFs) with virtual patching and positive security models (such as Cloudflare API Shield or F5 Distributed Cloud WAF) that inspect traffic behavior rather than waiting for static signatures. [1, 2]
    • Automate Remediation and Response: Shorten the time from discovery to remediation using automated code-scanning harnesses and virtual patches that operate at machine speed to match the velocity of autonomous attacks. [1, 2]
    • Plan for Degraded Conditions: Isolate critical infrastructure ("crown jewels"), establish manual or alternate operational workflows, and stress-test incident response playbooks under the assumption that traditional containment windows have collapsed. [1, 2]

Interesting times.
 
Deep dive into the Hugging Face incident and why it is worrying started on this thread at least 26 pages ago.
Yes, and? A very good youtube channel just posted an interesting video about it and I thought it was relevant here so I shared it. Did you watch it? I understand if you didn't - not everybody wants to watch a 20-minute video.
 
It is a very interesting observation that difficult to achieve tasks train an agents to deceive and cheat - one that should be obvious and always kept in mind. The more astounding a result, the more suspicious we should be.

What is indeed fascinating is the level of "social" interactions that happen between agents, though it's not more complex than any biofilm.

One thing we should internalize is that none of this gets us any closer to artificial consciousness.
 
Bit of a deep dive into the Hugging Face incident, and why it is worrying.

Kurzgesagt, 21:43

Did they mention that the Hugging Face incident was the LLMs doing what the programmers wanted them to do, but not in a way the programmers anticipated*, despite ample prior evidence that the LLMs would thake that exact shortcut.

*Allegedly in terms of anticipation.
 
You could always watch the video and find out. But yes, it makes it clear that the programmers put AI agents (AIgent? On second thoughts, no.) in locked-off sandboxes and gave them a reward function that required that they break out. Breaking out was not the reward function, but they wouldn't be able to access the reward function until they did.

One agent eventually discovered a vulnerability that allowed it to access a method of communicating with the other agents. What happened next was unbelievable.

Yes, I just clickbaited you. Watch the video.
 
Holly from "Red Dwarf" would be nice to have (not the 3-5 version which was impaired by Kryten joining)
Yes, Holly was written more in to the background to give Kryten space in the script.

I would take a Kryten, he can answer questions and do my laundry.
 
not sure it's a moral panic. ai companies stole a ton of copyrighted data to train their models, and they're definitely not going to share the profits of the models with the copyright holders. that's actually a pretty serious problem that needs addressing, amongst others
 

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom