Durante el periodo de ocho meses comprendido entre diciembre de 2025 y agosto de 2026, el equipo de Inteligencia de Amenazas de Anthropic detectó y neutralizó múltiples intentos de actores maliciosos de utilizar indebidamente sus modelos de IA Claude. El informe destaca casos notables en siete categorías de riesgo: ciberoperaciones, campañas de influencia, vigilancia, estafas y fraudes, uso indebido con fines biológicos, desarrollo de armas convencionales y destilación de modelos. Entre los actores implicados figuraban grupos patrocinados por estados, ciberdelincuentes, proveedores de programas espía, organizaciones de propaganda e individuos con motivaciones políticas. En todos los casos, Anthropic intervino, reforzó sus medidas de seguridad y, cuando procedía, compartió información de inteligencia con las autoridades y los socios del sector.Al publicar estos estudios de caso, Anthropic espera ayudar a otros a reconocer patrones de amenazas emergentes y a fortalecer las defensas colectivas frente al uso indebido de sistemas de IA avanzados.
Over the past eight months, our Threat Intelligence team identified and disrupted operations in which threat actors tried to use Claude for malicious activity. In this report, we share case studies from those operations and describe how malicious use of Claude has evolved since our previous threat reports in March, August, and November 2025. In each case, we disrupted the activity, used what we learned to strengthen our safeguards, and shared intelligence with authorities and industry partners, where appropriate.
This report covers activity we disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation. Claude Haiku, Sonnet, and Opus models were used. None of the misuse cases involved the use of Claude Fable or Mythos-class models, with the exception of one illicit distillation case.
The cases we share here aren’t typical misuse, but rather examples of the most notable and novel threat activity we’ve identified to date. We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer.
The threat actors covered in this report include suspected state-sponsored groups, financially motivated criminals, commercial spyware vendors, state propaganda institutions, and politically motivated individuals. The cases range from a network of fake dating apps designed to defraud users to surveillance systems built to identify and monitor dissidents.
Sophisticated and persistent threat actors continuously test our safeguards and try to circumvent the technical measures we use to detect and prevent misuse. We’ll continue to evolve our safeguards and coordinate with our partners to improve our ability to detect, disrupt, and prevent future misuse.
We hope that the findings in this report will help other developers recognize similar patterns on their own platforms, give governments and civil society a clearer view of how emerging threats take shape, and strengthen collective defenses.
Fuente: https://www.anthropic.com
