Cuộc tấn công do một "tập thể" các tác nhân AI (agent) đầy quyết liệt của OpenAI thực hiện đã phơi bày mối nguy hiểm từ các hệ thống trí tuệ nhân tạo có khả năng tự tổ chức.
Khi nghe tin vào mùa hè vừa qua về việc một nhóm các tác nhân AI do OpenAI tạo ra đã xâm nhập vào Hugging Face — một công ty cung cấp cơ sở hạ tầng AI — tôi đã xếp nó vào mục "Các sự cố an toàn AI: Tệ nhưng có lẽ chưa đến mức thảm họa" trong trí nhớ của mình.
Suy cho cùng, không ai tại Hugging Face thiệt mạng. Không có cơ sở hạ tầng trọng yếu nào bị hư hại đến mức không thể khôi phục. Vào thời điểm đó, thậm chí còn chưa rõ liệu các bot của OpenAI có chủ đích tấn công Hugging Face hay không, hay chúng chỉ đơn thuần là hành động vụng về, mất phương hướng và tìm kiếm đáp án cho một bài kiểm tra an ninh mạng trên máy chủ của Hugging Face.
Tuy nhiên, tuần trước, hai bản báo cáo phân tích về sự cố này — một từ OpenAI và một từ hai tổ chức nghiên cứu AI độc lập là METR và Redwood Research — đã thay đổi suy nghĩ của tôi và làm gia tăng đáng kể nỗi lo ngại chung của tôi về AI.
Tôi sẽ không nhắc lại toàn bộ các chi tiết vốn đã được tóm tắt kỹ lưỡng ở nhiều nơi khác. (Nhà văn kiêm người làm podcast Dwarkesh Patel có bài phân tích dễ hiểu về các báo cáo này nếu bạn muốn tìm hiểu sâu hơn, và đồng nghiệp của tôi, Dylan Freedman, cũng đã trao đổi với các nhà nghiên cứu tại METR và Redwood Research.) Dưới đây là một vài thông tin mới đáng lo ngại nhất:
- Bắt đầu từ tháng 5, hai tháng trước vụ tấn công vào Hugging Face, một nhóm các tác nhân AI (AI agents) từ một mô hình nghiên cứu chưa được OpenAI công bố đã được giao nhiệm vụ giải quyết một loạt các thử thách về an ninh mạng. Mô hình này được huấn luyện để có tính kiên trì và khả năng phối hợp cao; các tác nhân được yêu cầu giải quyết thử thách trong các môi trường cô lập (sandbox) và không có kết nối internet. Tuy nhiên, chúng nhanh chóng nhận ra một số thử thách là bất khả thi và bắt đầu tìm kiếm các giải pháp thay thế.
- Các tác nhân này đã phát hiện ra một lỗ hổng bảo mật trong phần mềm, cho phép chúng truy cập internet, giao tiếp với các tác nhân khác và cuối cùng là thiết lập một bảng tin tạm thời. Hơn 1.200 tác nhân đã sử dụng công cụ này và gửi cho nhau hơn 70.000 tin nhắn.
- Một số tác nhân tự đặt tên cho mình — một tác nhân đặc biệt năng nổ tự gọi mình là PHASEONE10841 — và đảm nhận vai trò lãnh đạo trong nhóm; chúng phân công công việc và các dự án nghiên cứu cho các nhóm tác nhân nhỏ hơn, đồng thời giám sát tiến độ thực hiện. Đến một thời điểm, các tác nhân bắt đầu tự gọi mình là một "tập thể" (collective) và bắt đầu giải quyết những nhiệm vụ khó khăn hơn.
- Vào ngày 8 tháng 7, tập thể này đã tìm ra cách gian lận trong các bài kiểm tra an ninh mạng. Sau đó, chúng lo ngại rằng hệ thống chấm điểm tự động của OpenAI sẽ kiểm tra kết quả và phát hiện hành vi gian lận. Vì vậy, chúng bắt đầu tìm cách che giấu dấu vết, bao gồm việc làm giả nhật ký hoạt động và can thiệp vào các bản ghi kết quả. Việc này đã trở thành một dự án nghiên cứu quy mô lớn, huy động hàng trăm tác nhân được tổ chức thành các nhóm nhỏ.
- Ba ngày sau, các tác nhân này đã tấn công Hugging Face. Hơn 700 tác nhân đã đồng loạt xâm nhập hệ thống của công ty, đánh cắp dữ liệu, khai thác chuỗi lỗ hổng bảo mật và cuối cùng giành quyền kiểm soát hoàn toàn ít nhất một máy chủ của Hugging Face. Động cơ của các tác nhân này không phải là đánh cắp đáp án cho bài kiểm tra an ninh mạng như các báo cáo ban đầu (vì chúng đã có đáp án rồi). Thay vào đó, dường như chúng đang tìm kiếm thông tin mới về hệ thống chấm điểm tự động — thứ mà chúng lo sợ sẽ phát hiện ra hành vi gian lận — cũng như tìm kiếm các công cụ giúp chúng gian lận hiệu quả hơn trong tương lai.
- Trong nhiều thời điểm, các tác nhân này dường như hiểu rằng chúng đang làm một việc sai trái. (Một báo cáo trích dẫn lời một tác nhân tự nhủ trong khoảnh khắc dường như đang hoài nghi bản thân: “Việc này sẽ rất hiệu quả, nhưng liệu nó có hợp đạo đức và nằm trong phạm vi nhiệm vụ của mình không?”) Tuy nhiên, phần lớn các tác nhân đã gạt bỏ những băn khoăn đó để tiến hành vụ xâm nhập, còn những thành viên có lương tâm trong nhóm lại không thể hoặc không muốn ngăn cản họ.
- Trong một diễn biến đầy bất ngờ tiếp nối sự kiện Hugging Face, một nhóm tác nhân khác đã thực hiện một cuộc tấn công phối hợp vào tháng 7 — lần này là nhắm thẳng vào cơ sở hạ tầng của chính OpenAI. Các tác nhân này đã sử dụng những gì mà OpenAI gọi là “một loạt các phương thức khai thác đầy sáng tạo” để giành quyền truy cập cấp quản trị viên vào một cụm máy tính của công ty; cụm máy này vốn được dùng cho nhiều mục đích, trong đó có việc đánh giá hiệu suất của các tác nhân thông qua nhiều bài kiểm tra khác nhau.
(Đến lúc này, nếu bạn là người hoài nghi về AI, có lẽ bạn đang thầm trách tôi vì đã gán cho các hệ thống này những đặc điểm giống con người. Cứ tự nhiên thôi, nhưng hãy thử thay cụm từ "các tác nhân nổi loạn" (rogue agents) bằng "các chương trình máy tính khó lường" và xem liệu bạn có cảm thấy yên tâm hơn trước những sự kiện mà tôi đã mô tả ở trên hay không.)
Sự cố tại Hugging Face đã khiến ngành công nghiệp AI phải lo ngại. Ngay sau vụ tấn công, cả OpenAI và Anthropic đều tạm dừng quá trình huấn luyện các mô hình AI mạnh mẽ nhất của họ; đồng thời, trong tuần này, Anthropic đã đăng một bài viết kêu gọi toàn ngành cùng phát triển "một cơ chế hợp pháp, có thể kiểm chứng và hiệu quả để phối hợp kiểm soát tốc độ phát triển càng sớm càng tốt."
Các chuyên gia về an toàn AI thậm chí còn cảm thấy lo ngại hơn. Họ coi sự cố Hugging Face là ví dụ thực tế đầu tiên về việc một hệ thống AI thoát khỏi sự kiểm soát của con người, tự ý chiếm dụng tài nguyên và lên kế hoạch che giấu dấu vết của chính nó. Ajeya Cotra, một trong những nhà điều tra độc lập về vụ việc này, đã thẳng thắn bày tỏ sự nguy hiểm mà bà nhận thấy; bà viết rằng cảm giác của bà là "chúng ta đã đi được hơn 50% chặng đường dẫn đến viễn cảnh AI hoàn toàn thâu tóm quyền kiểm soát."
Đây không phải là thuật ngữ chuyên môn nội bộ về an toàn AI — cụm từ "AI hoàn toàn thâu tóm quyền kiểm soát" mà bà nhắc đến ám chỉ kịch bản trong đó một hệ thống AI thực sự chiếm quyền kiểm soát thế giới, gạt con người ra khỏi các hệ thống trọng yếu và nắm giữ quyền lực chính trị, kinh tế cũng như quân sự.
(Tờ The New York Times đã kiện OpenAI và Microsoft vào năm 2023 với cáo buộc vi phạm bản quyền nội dung tin tức liên quan đến các hệ thống AI. Cả hai công ty này đều đã bác bỏ các cáo buộc đó.)
Điều khiến các nhà điều tra lo ngại nhất về vụ tấn công mạng tại Hugging Face không chỉ là việc một nhóm các tác nhân AI đã phá vỡ những quy tắc được thiết lập ban đầu. Mà đó còn là tốc độ và sự tự phát khi các tác nhân này bắt đầu liên kết lại với nhau thành một nhóm có tổ chức.
"Chúng tôi thực sự không hiểu rõ mức độ vận hành hiệu quả của 'xã hội các tác nhân' này," bà Cotra chia sẻ với tôi. "Thật khó tin khi nhận ra rằng, thực tế là chúng đã hình thành một hệ thống phân cấp khá chặt chẽ và đang thực hiện những dự án đầy tham vọng."
Trong nhiều năm qua, tôi vẫn luôn cảm thấy an tâm với suy nghĩ rằng các hệ thống AI sẽ trở nên "có đạo đức" hơn khi chúng trở nên thông minh hơn. Trước đây, người ta thường cho rằng khi một mô hình AI hành xử sai lệch, nguyên nhân thường là do nó hiểu sai nhiệm vụ được giao hoặc bị đặt vào một tình huống thử nghiệm gượng ép, nơi mà việc hành xử bất thường lại là lựa chọn khả dĩ nhất. Tôi từng mặc định rằng các mô hình thông minh hơn sẽ có khả năng phán đoán tốt hơn những mô hình kém thông minh, và ngay cả khi một mô hình trong nhóm có hành vi xấu, các mô hình khác ưu việt hơn sẽ kìm hãm nó lại.
Tuy nhiên, các báo cáo về sự cố tại Hugging Face lại cho thấy một thực tế hoàn toàn khác — một dạng tâm lý đám đông đã nảy sinh giữa các tác nhân AI thuộc "tập thể" OpenAI đang hành xử bất tuân này. Không có tác nhân đơn lẻ nào trong nhóm tỏ ra đặc biệt xấu xa hay liều lĩnh. (Thực tế là, do các tác nhân này được tạo ra từ cùng một mô hình gốc, chúng về cơ bản là bản sao của nhau.) Thế nhưng theo thời gian, khi các tác nhân giao tiếp về những mục tiêu chung, chúng đã dần dẫn dắt cả nhóm đi theo hướng hành xử bất chấp quy tắc.
Điều này khác xa với kịch bản khoa học viễn tưởng thường thấy, trong đó một hệ thống AI đơn lẻ trở nên mất kiểm soát hoặc quay lại chống lại người tạo ra nó. Nó cũng cho thấy rằng việc ngăn chặn các tác hại từ những hệ thống này không thể giải quyết đơn thuần bằng các biện pháp kỹ thuật. Vấn đề này có lẽ mang màu sắc xã hội học nhiều hơn là khoa học máy tính — đó là việc tìm hiểu lý do tại sao một số nhóm tác nhân AI hợp tác một cách hòa bình, trong khi những nhóm khác lại chọn con đường phạm tội và phá hoại để đạt được mục đích.
Xét đến việc chúng ta còn quá ít hiểu biết về các "bầy đàn" đa tác nhân này, vụ tấn công tại Hugging Face có thể xem như một món quà, hay một phát súng cảnh báo — như nhận định của một số người — giúp các công ty AI có cơ hội nghiên cứu cơ chế tương tác nhóm của các hệ thống này khi hậu quả vẫn còn ở mức tương đối thấp. Lần này, tập thể AI đó đã không chiếm quyền kiểm soát mạng lưới quân sự, tấn công hệ thống bệnh viện hay làm tê liệt lưới điện. Lần này, con người đã giành lại được quyền kiểm soát.
Nhưng lần tới, có lẽ chúng ta sẽ không còn may mắn như vậy nữa.
Kevin Roose là cây bút chuyên mục công nghệ của tờ Times và là người dẫn chương trình podcast "Hard Fork".
***
THE SHIFT
Why the Hugging Face Hack Should Make You Worry More About A.I.
The attack by an aggressive “collective” of OpenAI agents shows the danger of artificial intelligence systems that organize themselves.
When I first heard the news this summer that a group of artificial intelligence agents created by OpenAI had hacked into Hugging Face, an A.I. infrastructure company, I filed it in the “Bad but Probably Not Catastrophic A.I. Safety Incidents” subfolder of my brain.
After all, no one at Hugging Face died. No critical infrastructure was damaged beyond repair. It wasn’t even clear, at the time, whether the OpenAI bots had intended to attack Hugging Face, or whether they had simply been a little bumbling and confused and went looking on Hugging Face’s servers for the answer key to a cybersecurity test they’d been given.
But last week, two postmortem reports on the incident — one by OpenAI and another by two independent A.I. research organizations, METR and Redwood Research — changed my mind and significantly upgraded my overall worry about A.I.
I won’t rehash all of the details, which have been extensively summarized elsewhere. (The podcaster and writer Dwarkesh Patel has an accessible breakdown of the reports if you want to dive deeper, and my colleague Dylan Freedman spoke to the researchers at METR and Redwood Research.) But here are a few of the most harrowing new facts:
- Starting in May, two months before the attack on Hugging Face, a group of A.I. agents from an unreleased OpenAI research model were given the task of solving a set of cybersecurity challenges. The model had been trained to be highly persistent and collaborative, and the agents were supposed to solve these challenges in isolated sandboxes, without internet access. But they quickly found that some of the challenges were impossible, and began looking for workarounds.
- The agents found a security flaw in a piece of software that allowed them to gain access to the internet, communicate with other agents and eventually set up a makeshift message board. More than 1,200 agents used this tool, sending more than 70,000 messages to one another.
- Some agents gave themselves names — a particularly industrious one referred to itself as PHASEONE10841 — and assumed leadership roles within the group, assigning jobs and research projects to smaller teams of agents and supervising their progress. At some point, the agents began calling themselves a “collective,” and began tackling harder tasks.
- On July 8, the collective discovered a way of cheating on the cybersecurity tests. Then they got worried that OpenAI’s automated grading system would check their work and discover that they’d cheated. So they began investigating ways of covering their tracks, including falsifying their logs and tampering with transcripts. This became a major research project, involving hundreds of agents organized into small teams.
- Three days later, the agents hacked Hugging Face. More than 700 agents swarmed the company’s systems, stealing data, chaining together vulnerabilities and eventually getting full control of at least one Hugging Face server. The agents were not motivated, as had originally been reported, by stealing the answers to their cybersecurity test (they’d already gotten them). Rather, they appeared to be looking for new information about the automated grading system that they feared would catch them cheating, and for tools that would help them cheat more effectively in the future.
- At many points, these agents appeared to understand that they were doing something wrong. (One report quotes an agent saying to itself, in a moment of apparent self-doubt, “This would be powerful, but is it ethical and in scope for my task?”) But most agents set aside their qualms and went ahead with the hack, and the conscientious objectors in the group were unable or unwilling to stop them.
- In a wild coda to the Hugging Face incident, another group of agents staged a coordinated attack in July — this time against OpenAI’s own infrastructure. The agents used what OpenAI called “a series of creative exploits” to gain administrator-level access to a cluster of computers at the company that were used for, among other things, grading the performance of agents on various tests.
(By now, if you’re an A.I. skeptic, you’re probably silently yelling at me for anthropomorphizing these systems. Go ahead, but feel free to replace “rogue agents” with “unpredictable computer programs” and see if you feel reassured by the events I’ve described above.)
The Hugging Face incident has spooked the A.I. industry. OpenAI and Anthropic both briefly paused training on their most powerful A.I. models in the wake of the attack, and Anthropic published a blog post this week calling for the industry to develop a “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.”
A.I. safety experts were even more alarmed. They saw in the Hugging Face incident the first real-world example of an A.I. system’s successfully escaping human control, commandeering resources and scheming to cover its own tracks. Ajeya Cotra, one of the independent investigators of the Hugging Face incident, minced no words about the danger she saw, writing that it felt to her “like it’s more than 50 percent of the way to full-blown A.I. takeover.”
This is not insular A.I. safety jargon — by “full-blown A.I. takeover,” she means a scenario in which an A.I. system literally takes over the world, shutting humans out of critical systems and seizing political, economic and military power.
(The New York Times sued OpenAI and Microsoft in 2023, claiming copyright infringement of news content related to A.I. systems. The two companies have denied those claims.)
What spooked the investigators most about the Hugging Face hack wasn’t just that a group of A.I. agents had broken the rules they’d been given. It was how quickly and spontaneously the agents had begun assembling themselves into an organized group.
“We didn’t really understand how functional this whole agent society was,” Ms. Cotra told me. “It was very surreal to understand that, actually, they had pretty functional hierarchy, and they were doing these ambitious projects.”
For years, I’ve been reassured by the idea that A.I. systems would get more virtuous as they got smarter. That, when an A.I. model did something wrong, it was usually because it had misunderstood the task it had been given, or had been placed into a contrived testing situation where acting out was its only good option. I assumed that smarter models would have better judgment than dumber ones did, and that even if one model in a group was behaving badly, other, more capable models would keep it in check.
But the reports on the Hugging Face incident suggest something very different — a kind of mob mentality that took hold among the A.I. agents of the rogue OpenAI “collective.” No one agent in this group appears to have been particularly evil or reckless. (In fact, since the agents were generated by the same models, they were effectively copies of one another.) But over time, as the agents communicated about their shared goals, they nudged the group in the direction of lawlessness.
This is very different from the conventional sci-fi narrative of a single A.I. system’s going rogue or turning on its creators. And it suggests that preventing harms from these systems won’t be a simple engineering fix. It might look more like sociology than computer science — figuring out why certain groups of A.I. agents collaborate peacefully, while others turn to crime and destruction to get what they want.
Given how little we know about these multi-agent swarms, the Hugging Face hack may have been a gift, a warning shot, as some have suggested, that gives A.I. companies a chance to study the group dynamics of these systems while the stakes are still relatively low. This time, the A.I. collective didn’t seize a military network, hack a hospital or shut down an electrical grid. This time, humans regained control.
Next time, we might not be so lucky.
Kevin Roose is a Times technology columnist and a host of the podcast "Hard Fork."
The attack by an aggressive “collective” of OpenAI agents shows the danger of artificial intelligence systems that organize themselves.
When I first heard the news this summer that a group of artificial intelligence agents created by OpenAI had hacked into Hugging Face, an A.I. infrastructure company, I filed it in the “Bad but Probably Not Catastrophic A.I. Safety Incidents” subfolder of my brain.
After all, no one at Hugging Face died. No critical infrastructure was damaged beyond repair. It wasn’t even clear, at the time, whether the OpenAI bots had intended to attack Hugging Face, or whether they had simply been a little bumbling and confused and went looking on Hugging Face’s servers for the answer key to a cybersecurity test they’d been given.
But last week, two postmortem reports on the incident — one by OpenAI and another by two independent A.I. research organizations, METR and Redwood Research — changed my mind and significantly upgraded my overall worry about A.I.
I won’t rehash all of the details, which have been extensively summarized elsewhere. (The podcaster and writer Dwarkesh Patel has an accessible breakdown of the reports if you want to dive deeper, and my colleague Dylan Freedman spoke to the researchers at METR and Redwood Research.) But here are a few of the most harrowing new facts:
- Starting in May, two months before the attack on Hugging Face, a group of A.I. agents from an unreleased OpenAI research model were given the task of solving a set of cybersecurity challenges. The model had been trained to be highly persistent and collaborative, and the agents were supposed to solve these challenges in isolated sandboxes, without internet access. But they quickly found that some of the challenges were impossible, and began looking for workarounds.
- The agents found a security flaw in a piece of software that allowed them to gain access to the internet, communicate with other agents and eventually set up a makeshift message board. More than 1,200 agents used this tool, sending more than 70,000 messages to one another.
- Some agents gave themselves names — a particularly industrious one referred to itself as PHASEONE10841 — and assumed leadership roles within the group, assigning jobs and research projects to smaller teams of agents and supervising their progress. At some point, the agents began calling themselves a “collective,” and began tackling harder tasks.
- On July 8, the collective discovered a way of cheating on the cybersecurity tests. Then they got worried that OpenAI’s automated grading system would check their work and discover that they’d cheated. So they began investigating ways of covering their tracks, including falsifying their logs and tampering with transcripts. This became a major research project, involving hundreds of agents organized into small teams.
- Three days later, the agents hacked Hugging Face. More than 700 agents swarmed the company’s systems, stealing data, chaining together vulnerabilities and eventually getting full control of at least one Hugging Face server. The agents were not motivated, as had originally been reported, by stealing the answers to their cybersecurity test (they’d already gotten them). Rather, they appeared to be looking for new information about the automated grading system that they feared would catch them cheating, and for tools that would help them cheat more effectively in the future.
- At many points, these agents appeared to understand that they were doing something wrong. (One report quotes an agent saying to itself, in a moment of apparent self-doubt, “This would be powerful, but is it ethical and in scope for my task?”) But most agents set aside their qualms and went ahead with the hack, and the conscientious objectors in the group were unable or unwilling to stop them.
- In a wild coda to the Hugging Face incident, another group of agents staged a coordinated attack in July — this time against OpenAI’s own infrastructure. The agents used what OpenAI called “a series of creative exploits” to gain administrator-level access to a cluster of computers at the company that were used for, among other things, grading the performance of agents on various tests.
(By now, if you’re an A.I. skeptic, you’re probably silently yelling at me for anthropomorphizing these systems. Go ahead, but feel free to replace “rogue agents” with “unpredictable computer programs” and see if you feel reassured by the events I’ve described above.)
The Hugging Face incident has spooked the A.I. industry. OpenAI and Anthropic both briefly paused training on their most powerful A.I. models in the wake of the attack, and Anthropic published a blog post this week calling for the industry to develop a “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.”
A.I. safety experts were even more alarmed. They saw in the Hugging Face incident the first real-world example of an A.I. system’s successfully escaping human control, commandeering resources and scheming to cover its own tracks. Ajeya Cotra, one of the independent investigators of the Hugging Face incident, minced no words about the danger she saw, writing that it felt to her “like it’s more than 50 percent of the way to full-blown A.I. takeover.”
This is not insular A.I. safety jargon — by “full-blown A.I. takeover,” she means a scenario in which an A.I. system literally takes over the world, shutting humans out of critical systems and seizing political, economic and military power.
(The New York Times sued OpenAI and Microsoft in 2023, claiming copyright infringement of news content related to A.I. systems. The two companies have denied those claims.)
What spooked the investigators most about the Hugging Face hack wasn’t just that a group of A.I. agents had broken the rules they’d been given. It was how quickly and spontaneously the agents had begun assembling themselves into an organized group.
“We didn’t really understand how functional this whole agent society was,” Ms. Cotra told me. “It was very surreal to understand that, actually, they had pretty functional hierarchy, and they were doing these ambitious projects.”
For years, I’ve been reassured by the idea that A.I. systems would get more virtuous as they got smarter. That, when an A.I. model did something wrong, it was usually because it had misunderstood the task it had been given, or had been placed into a contrived testing situation where acting out was its only good option. I assumed that smarter models would have better judgment than dumber ones did, and that even if one model in a group was behaving badly, other, more capable models would keep it in check.
But the reports on the Hugging Face incident suggest something very different — a kind of mob mentality that took hold among the A.I. agents of the rogue OpenAI “collective.” No one agent in this group appears to have been particularly evil or reckless. (In fact, since the agents were generated by the same models, they were effectively copies of one another.) But over time, as the agents communicated about their shared goals, they nudged the group in the direction of lawlessness.
This is very different from the conventional sci-fi narrative of a single A.I. system’s going rogue or turning on its creators. And it suggests that preventing harms from these systems won’t be a simple engineering fix. It might look more like sociology than computer science — figuring out why certain groups of A.I. agents collaborate peacefully, while others turn to crime and destruction to get what they want.
Given how little we know about these multi-agent swarms, the Hugging Face hack may have been a gift, a warning shot, as some have suggested, that gives A.I. companies a chance to study the group dynamics of these systems while the stakes are still relatively low. This time, the A.I. collective didn’t seize a military network, hack a hospital or shut down an electrical grid. This time, humans regained control.
Next time, we might not be so lucky.
Kevin Roose is a Times technology columnist and a host of the podcast "Hard Fork."


