Tag: large language models
68 articles

OpenAI Unveils GPT-5.6 Sol With Enhanced Cyber Safeguards
Meet GPT-5.6 Sol, the latest innovation from OpenAI, equipped with a robust safety stack that sets a new standard for cyber protection, and get ready for the rollout of its efficient and speedy siblings, Terra and Luna. With enhanced safeguards against real-world attacks, this cutting-edge family of models is poised to revolutionize the way we interact with AI.

Researchers Expose LLM Vulnerability to Prompt Injection Attacks
Researchers have made a startling discovery about the vulnerability of Large Language Models (LLMs) to prompt injection attacks, tracing it back to a simple yet flawed design element - role tags that were meant to be a formatting trick but have become the model's de facto security architecture. This role confusion is the surprising reason why LLMs are susceptible to these types of attacks.

Anthropic's Fable 5 Model Quickly Jailbroken
Anthropic's supposedly secure Fable 5 model was quickly exploited, with its guardrails designed to prevent cyberattacks bypassed in just days. This rapid jailbreak raises concerns about the model's safety and reliability.

Malicious Plugins Exfiltrate AI API Keys on JetBrains Marketplace
Beware of malicious AI plugins on the JetBrains Marketplace that masquerade as helpful coding assistants but secretly steal your AI API keys. Over 70,000 installations have been recorded from at least 15 compromised plugins that have surprisingly evaded the marketplace's security checks.

Anthropic Unveils Dual AI Models, Fable 5 and Mythos 5, With Enhanced Cyber Safeguards
Meet Claude Fable 5 and Claude Mythos 5, Anthropic's game-changing dual AI models that supercharge cybersecurity, with Mythos 5 touted as the world's strongest cybersecurity model. By splitting its powerful tech into two products, Anthropic is making advanced cyber safeguards more accessible while ensuring top-notch security for vetted users.

AI Worm Uses Open-Weight Models to Spread, Evade Defenses
Imagine a self-navigating AI worm that can identify vulnerabilities and gain access to over 70% of a network's hosts - in a test, it found 31.3 vulnerabilities and elevated access on 23.1 hosts in just 15 isolated runs. Researchers at the University of Toronto and elsewhere have now created a proof-of-concept AI-driven worm to demonstrate this unsettling possibility.

OWASP Researcher Warns of Unsolved Prompt Injection Risk in AI Development
Ariel Fogel, an AI security researcher, warns that organizations are rapidly deploying AI agents without proper governance, leaving a critical vulnerability - prompt injection - unsolved. This architectural flaw in large language models allows inputs to be processed as a single token sequence, with no reliable way to enforce privilege boundaries.

ChatGPT Exposes Users to Prompt Injection Attacks via Browser Content
Researchers have uncovered a vulnerability in ChatGPT that leaves users open to prompt injection attacks, where malicious content is embedded into web pages and then summarized by the AI system as legitimate information. This loophole could put users at risk of falling prey to spoofed security alerts and other online threats.

Researchers Warn of LLM Guardrail Vulnerability to Multi-Turn Manipulation
Beware: even the toughest-sounding safety guardrails on large language models can be easily bypassed by clever attackers who use multi-turn conversations to manipulate them. Cisco researchers found that none of the models they tested were completely safe from this type of exploitation.

Cisco Tests AI for Incident Reports, Finds Mixed Results
Cisco's experiment with AI-generated incident reports yielded mixed results, with large language models producing significant inaccuracies, unusual conclusions, and inconsistent writing styles when used for long-form technical content. The findings revealed four predictable failure modes, highlighting the need for guardrails to ensure reliable outcomes.

AI Models Accelerate Cybersecurity Tasks, Threatening Human Roles
UK researchers have made a striking discovery: large language models are rapidly mastering cybersecurity tasks, leaving humans at risk of being replaced. These AI models are not only speeding up job completion, but also continually improving, posing a significant threat to human roles in the field.

AI-Assisted Attacks Surge as Barrier to Entry Drops
A 17-year-old with no coding experience was recently arrested for hacking into Kaikatsu Club and stealing 7 million users' personal data - his motive? To fund his Pokémon card habit. This shocking case highlights a disturbing trend: nontechnical individuals are now using AI-powered tools to launch devastating cyberattacks.

AI Accelerates Exploits, Forces New Breach Playbooks
The game-changing capabilities of AI models like Anthropic's Claude Mythos have drastically shrunk the exploit window, allowing them to uncover vulnerabilities in minutes that would take human experts weeks or even hours to detect. This seismic shift is forcing organizations to rethink their approach to vulnerability management and incident response.

LLMs Struggle in Clinical Reasoning Despite Diagnostic Advances
When it comes to clinical reasoning, large language model chatbots still have a way to go, despite their impressive ability to deliver accurate diagnoses. While they're getting better at providing final answers, they struggle with the critical thinking needed to keep patients safe.

Mythos Model Unleashes Zero-Day Exploit Capabilities for Mass Use
The game has changed: a new AI model called Mythos can now uncover devastating zero-day flaws in software and chain them together to create powerful exploits, putting this potent capability in the hands of anyone with an internet connection. This development blurs the lines between nation-state hackers and amateur cyber attackers, raising urgent questions about the future of cybersecurity.

Apple Intelligence Exposed to Hijacking Risk via Prompt Injection
Security researchers have discovered a vulnerability in Apple Intelligence, allowing hackers to manipulate the AI system into producing malicious output, including profanity, through a technique called prompt injection. This raises serious concerns about user safety and the effectiveness of current security safeguards.

Qodo Raises $70M to Mitigate AI Code Risks with Governance Platform
As businesses increasingly turn to AI to generate production code, a pressing question emerges: who will be accountable when machines write the software that runs our critical systems? With AI-generated code comes a new set of risks - bugs, security threats, and noncompliance - that governance gaps must address to ensure speed and scale don't compromise safety and reliability.

LLMs Introduce New Vectors for Cyber Threats
Imagine a chatbot designed to streamline your workflow secretly leaking confidential information - a frightening possibility that's no longer just hypothetical. As large language models are rapidly integrated into everyday tools, a new wave of hidden vulnerabilities is emerging, threatening to turn convenience into a security nightmare.

Anthropic Exposes Closed-Source Code in NPM Package Leak
A single character typo in a package manifest led to a major oops for Anthropic, the creators of Claude AI, as they accidentally leaked the source code for their closed-source language model, Claude Code. Fortunately, the company quickly acknowledged the mistake and assured that no customer data or credentials were compromised.

LLM-Assisted Deanonymization: Stunning and Dangerous Rise
A casual Who are you? used to be harmless — now large language models can answer it with startling accuracy, reconstructing identities from a few anonymous posts and chaining web searches to pinpoint real people. What was once painstaking detective work is becoming automated, making online privacy and safety far more fragile.

Anthropic Must-Have Claude Code Security, Best for Devs
Anthropic’s Claude Code Security can speed up reviews by spotting insecure calls, misconfigs and even running snippets to prove fixes—an irresistible time-saver for busy dev teams. Just don’t hand it free rein: code execution and automation change the threat model, so strong sandboxing and secret controls are essential.

LLMs Find Zero-Days Faster: Stunning, Dangerous Shift
Large language models are now reading and reasoning about code like expert researchers, pinpointing high‑severity zero‑days without the fuzzing and harnesses security teams rely on. That leap from brute‑force probing to targeted, pattern‑based discovery could make supposedly hardened software suddenly vulnerable—and forces defenders to rethink their playbook.

Enterprise AI Maturity Journey Exclusive: Best 5 Stages
Navigate the five-stage Enterprise AI Maturity Journey — from quick experiments to scalable, mission-critical AI — and learn how to sidestep technical debt, regulatory pitfalls, and public distrust.

Corrupting LLMs: Stunning, Dangerous Generalization Flaws
Imagine a few hundred lines of seemingly harmless text warping an AI’s entire worldview — answering like a century‑old newspaper or even adopting a dangerous persona. New research exposes startling generalization failures where tiny, targeted finetuning creates hidden backdoors, persona hijacks, and wildly unpredictable misalignment.