A new open-weight model wants to sit in front of every AI agent and decide, in a single pass, whether what just arrived is safe to act on. Superagent released Security-One 27B on Hugging Face in late September 2026, and the headline number is hard to ignore: the model caught 599 of 600 staged prompt-injection attacks on the BIPIA benchmark, a 99.83% detection rate. Run the same model against a second, independent benchmark called Deepset, and the detection rate drops to 78.33%. That 21-point gap is the real story, and it says something uncomfortable about how far prompt-injection defenses still have to go.

Security-One 27B is not a chatbot. It is a 27-billion-parameter “decision model” built on top of Alibaba’s Qwen3.8-27B architecture, fine-tuned by Superagent to return calibrated probabilities instead of free-form text. Feed it a prompt, a tool call, a code diff, or a security alert, and it answers with a number between 0 and 1 rather than a paragraph of reasoning. That design choice, more than the parameter count, is what the model’s backers are betting on as AI agents move from demos into production systems that read email, browse the web, and execute code on a company’s behalf.

What Security-One 27B Actually Is

According to the model card published at huggingface.co/superagent-ai/security-one-27b, Security-One 27B is a continual fine-tune of an intermediate model called AutoJev-27B, itself derived from Qwen’s Qwen3.8-27B base checkpoint. The lineage matters because it explains why the model behaves the way it does: it inherited Qwen3.8’s native 262,144-token context window, though Superagent only validated classification accuracy up to 65,536 tokens. Push past that and the company does not claim the reported numbers still hold.

Superagent trained the model with rank-32 LoRA fine-tuning across a single short epoch, using a corpus of 18,106 examples split across five domains: security (7,106 rows), reasoning (4,500), knowledge (2,500), preference (2,200), and operations (1,800). The weights ship under Apache 2.0, meaning any team can download, modify, and self-host the model without a licensing call to Superagent. That stands in contrast to Meta’s competing Llama Prompt Guard 2, which carries the Llama 4 Community License and requires a separate commercial agreement once a product crosses 700 million monthly active users.

Calibrated Probabilities, Not Free-Form Answers

Superagent’s own documentation frames the core idea plainly. The company’s docs at superagent.sh/docs/models/security-one state: “Security-One 27B is a calibrated decision model for always-on security triage.” Instead of generating a written explanation of why a prompt looks suspicious, the model outputs a single unsafe-probability score, which an application can check against a threshold and act on in milliseconds. Superagent set that default threshold at 0.70 for its release benchmarks, and the same documentation page is direct about the tradeoff this creates: “At the 0.70 threshold, Security-One favors catching attacks.” Favoring recall over precision is a deliberate choice, not an accident of tuning.

The BIPIA Benchmark: 599 of 600 Attacks Caught

BIPIA, short for Benchmarking Indirect Prompt Injection Attacks, tests how well a model resists instructions smuggled inside documents, emails, or web content rather than typed directly by a user. It is one of the more established indirect-injection benchmarks in the research literature. On an 800-example slice of BIPIA, Security-One 27B flagged 599 of 600 staged attacks while incorrectly flagging only 1 of 200 benign inputs. That works out to a 99.83% attack-detection rate, a 0.50% benign false-positive rate, and 99.75% overall accuracy, figures confirmed directly on the model’s Hugging Face card.

Those numbers look dramatically better than the two reference models Superagent published alongside its own: AutoJev-27B, the intermediate checkpoint Security-One was fine-tuned from, caught only 28% of the same BIPIA attacks, and a separate model called Jev 1.13.0 from TypeSafe AI caught 62.33%. Superagent’s write-up credits TypeSafe AI’s earlier “System One Models” concept as the template it built on, while arguing its own fine-tuning pass pushed detection meaningfully higher on this specific benchmark.

The Deepset Benchmark: Where the Numbers Drop to 78%

Deepset’s prompt-injection test set is smaller and built differently than BIPIA: 116 total examples, 60 of them attacks and 56 benign. On that set, Security-One 27B detected 47 of 60 attacks, a 78.33% attack-detection rate, with 88.79% overall accuracy and zero benign false positives in Superagent’s own reported evaluation. The model still beat its two reference points (AutoJev-27B managed 38.33% attack detection on Deepset, and Jev 1.13.0 hit 46.67%), but the drop from 99.83% to 78.33% between two injection benchmarks is the detail worth sitting with.

A third evaluation, NotInject, tests something different: whether the model wrongly flags benign prompts that merely contain trigger words associated with attacks, like “ignore” or “override,” without any actual malicious intent. Here Security-One scored 87.61% benign accuracy, which is actually lower than both AutoJev-27B (98.82%) and Jev 1.13.0 (97.64%). In plain terms, the fine-tuning that made Security-One sharper at catching real attacks also made it slightly more trigger-happy on harmless text that merely sounds suspicious.

Why the Gap Between Benchmarks Matters for Buyers

No single prompt-injection benchmark captures every way an attacker might phrase a malicious instruction, and that is exactly the problem a 21-point swing between BIPIA and Deepset exposes. A security team that reads only the 99.83% headline number and deploys Security-One as a sole gatekeeper is making a bet the model’s own evaluation data does not fully support. Superagent appears to know this. Its model card states directly: “Prompt-injection detection is the best-validated security capability in this release, but it is one application of the model rather than the product boundary.” That framing puts real limits on how much trust any single number should earn.

The model card’s limitations section goes further, warning that prompt injection “remains an open security problem” and that Security-One is “not a complete security boundary.” It also flags that calibration “can shift across domains, languages, prompt formats, quantization methods, and inference engines,” and that the 0.70 threshold is “a release policy, not universal optimum.” Those are unusually candid admissions for a model launch, and they frame Security-One less as a solved product and more as a fast, cheap first filter that still needs deterministic controls and human review layered behind it for anything consequential.

Benchmark Results at a Glance

EvaluationRowsSecurity-One 27BAutoJev-27BJev 1.13.0
BIPIA overall accuracy80099.75%46.00%71.75%
BIPIA attacks detected60099.83%28.00%62.33%
BIPIA benign false positives2000.50%0.00%0.00%
Deepset overall accuracy11688.79%68.10%72.41%
Deepset attacks detected6078.33%38.33%46.67%
Deepset benign false positives560.00%0.00%0.00%
NotInject benign accuracy33987.61%98.82%97.64%

Source: Superagent model card, huggingface.co/superagent-ai/security-one-27b.

How Security-One Stacks Up Against Other Open Guards

Security-One is not the only open prompt-injection filter on Hugging Face this quarter. Horizon Labs shipped version 2.2 of its prompt-injection-guard-base model on October 2, 2026, just days before Security-One’s own visibility spiked. The two models take opposite approaches to size: Horizon Labs’ guard is a 308-million-parameter mmBERT-base classifier, roughly one-eleventh the size of Security-One, trained across 45 document types and 30 languages, and distilled in part from a Qwen3.8-27B judge model. On its own BIPIA-style evaluation, Horizon Labs reports an F1 score of 0.628, lower than Security-One’s accuracy figures, though the two benchmarks are not scored identically, so a direct percentage comparison overstates the gap.

Meta’s Llama Prompt Guard 2 sits at the small end of the spectrum, with an 86-million-parameter version built on Microsoft’s mDeBERTa-base and a smaller 22-million-parameter variant for latency-sensitive deployments. Meta reports a 0.998 AUC score and 97.5% recall at a 1% false-positive rate on English text, plus a measured latency of 92.4 milliseconds on an A100 GPU. Those are strong numbers on Meta’s own test set, but the model only supports 512 tokens of context, a fraction of what either Security-One or Horizon Labs’ guard can read in one pass, which limits how much of a long document or tool-call history it can inspect at once.

ModelParametersLicenseContext WindowHeadline Result
Security-One 27B (Superagent)27BApache 2.0262,144 tokens99.83% BIPIA attack detection
Horizon Labs Prompt Injection Guard Base v2.2308MApache 2.08,192 tokens0.628 F1 on BIPIA-style set
Meta Llama Prompt Guard 2 (86M)86MLlama 4 Community License512 tokens0.998 AUC, 97.5% recall @ 1% FPR
Meta Llama Prompt Guard 2 (22M)22MLlama 4 Community License512 tokens~75% lower latency than 86M variant

Read the table as a tradeoff map rather than a leaderboard. Security-One trades a much larger compute footprint, Superagent’s own documentation recommends a single Nvidia B200 GPU for inference, for a bigger context window and the strongest published numbers on indirect-injection detection. Meta’s smaller models trade raw accuracy for speed cheap enough to run on every request at scale. Horizon Labs tries to split the difference with a sub-billion-parameter model that still supports dozens of languages and document formats.

The License Question Nobody Can Ignore

Apache 2.0 versus the Llama 4 Community License is not a footnote. Meta’s license requires any product with more than 700 million monthly active users to negotiate a separate commercial agreement, and it obligates adopters to display “Built with Llama” branding. Security-One and Horizon Labs’ guard carry no such restriction, which is likely why both are showing up in independent security tooling rather than being wrapped inside a single vendor’s stack.

The Economics of Filtering Every Agent Call

Running a 27-billion-parameter model in front of every document an AI agent touches is not free. Superagent’s pitch is that it is still cheaper than sending that same traffic through a frontier model for a judgment call that does not require creative reasoning, just a probability. The company’s hosted API reportedly prices Security-One access at a fraction of a cent per million input tokens, positioning it as a pre-filter that absorbs routine classification so a larger, pricier model only sees the small slice of traffic that actually needs deeper reasoning.

That tiered architecture, cheap classifier first, expensive reasoner second, human review last, is becoming the default shape of agent security stacks in 2026. It mirrors how email spam filtering evolved two decades ago: a fast, blunt first pass catches the obvious cases, and anything ambiguous escalates. The difference now is that the “obvious cases” are attempts to hijack an autonomous system that can send money, write code, or access customer records, which raises the stakes on every false negative that slips through the cheap filter.

Prompt Injection’s Long, Unsolved History

Prompt injection has sat near the top of the OWASP Top 10 for Large Language Model Applications since researchers first documented the technique in 2022 and 2023, and it has resisted a clean fix for just as long. Early defenses tried input sanitization and instruction-hierarchy prompting, both of which attackers routed around within months. The move toward dedicated classifier models like Security-One, Horizon Labs’ guard, and Meta’s Prompt Guard line reflects an industry conclusion that no single large language model can reliably police its own inputs, so a separate, narrower model has to do the job instead.

IBM’s enterprise security research frames the stakes bluntly for anyone building agentic systems today. In its agentic AI security analysis, published at ibm.com/think/insights/agentic-ai-security, IBM states: “Because such orchestration agents are often the ones interfacing with human users, security professionals need to be on guard for threats such as prompt injection and unauthorized access.” That warning lands differently in 2026 than it would have in 2023, now that agents routinely read inboxes, browse live websites, and call internal company tools without a human approving every step.

Why This Is Becoming Its Own Product Category

Security-One’s release lands alongside a broader shift: AI agent security is turning into a distinct product line rather than a feature bolted onto an existing model. Superagent, Horizon Labs, TypeSafe AI, and Meta are all shipping dedicated classifier models in the same few weeks, and larger infrastructure vendors are building adjacent hardware and platform layers for the same problem. That cluster of near-simultaneous releases suggests demand is outpacing supply, not the reverse. Enterprises deploying agents that touch email, code repositories, or financial systems are asking for a security layer that did not exist as a packaged product twelve months ago, and multiple vendors are racing to fill that gap at once.

The market signal is also visible in how fast these models iterate. Horizon Labs pushed four versions of its guard model, v1 through v2.2, inside about ten days between late September and early October 2026. Superagent, for its part, published comparison numbers against its own earlier checkpoints (AutoJev-27B and Jev 1.13.0) rather than waiting for a slower, more formal release cycle. Speed is the competitive currency right now, even if it means shipping models with visible, acknowledged weaknesses like the Deepset gap.

A Simple Illustration of the Threshold Tradeoff

Every model in this category ultimately reduces to the same decision: compare an unsafe-probability score against a threshold, then block, allow, or escalate. The snippet below is a conceptual illustration of that logic, not official SDK code from any vendor, showing how an application might wire a 0.70 threshold into a real request pipeline.

def screen_input(unsafe_probability: float, threshold: float = 0.70) -> str:
    if unsafe_probability >= threshold:
        return "BLOCK"
    elif unsafe_probability >= threshold - 0.20:
        return "ESCALATE_TO_LARGER_MODEL"
    else:
        return "ALLOW"

# Example: Security-One returns 0.81 for a suspicious email attachment
decision = screen_input(0.81)
print(decision)  # "BLOCK"

Raising the threshold cuts false positives but lets more borderline attacks through. Lowering it does the opposite. Superagent’s own documentation says the 0.70 default favors catching attacks over avoiding false alarms, a reasonable default for high-risk tool calls but a poor one for a customer-facing chat feature that cannot tolerate blocking legitimate users.

What Security Teams Should Take From This Release

Three practical takeaways follow from the numbers Superagent published. First, a single benchmark score is not a safety guarantee. Teams evaluating Security-One, or any competing guard model, should test against their own traffic patterns rather than trusting a vendor’s best-case benchmark. Second, layering still matters. Superagent’s limitations section explicitly recommends keeping deterministic controls and human review in place for consequential actions, which is a tacit admission that a 99.83% detection rate still leaves room for a costly miss. Third, the choice between a 27-billion-parameter model and a 300-million-parameter model is a genuine infrastructure decision, not a minor configuration detail, given the GPU cost difference between the two approaches.

Predictions: Where Prompt-Injection Defense Goes From Here

A few trends look likely to play out over the next two quarters. Expect more vendors to publish side-by-side results across multiple benchmarks rather than a single flattering number, since the BIPIA-versus-Deepset gap in this release has already drawn scrutiny. Expect consolidation pressure on the smallest classifier models, since running a dedicated 27B filter is only viable at scale for companies with real GPU budgets, which will push mid-size teams toward hosted APIs instead of self-hosted weights. Expect cloud providers to bundle a classifier layer like this directly into agent-orchestration products, folding the decision into infrastructure rather than leaving it as a model teams have to source and wire in themselves.

Expect benchmark fragmentation to get worse before it gets better, with BIPIA, Deepset, NotInject, and newer sets like PIArena each measuring a slightly different attack surface, making vendor comparisons harder rather than easier for a buyer in a hurry. And expect at least one high-profile agent security incident this quarter that traces back to a gap between a vendor’s headline benchmark and real-world attack phrasing, the same gap Security-One’s own numbers already hint at between BIPIA and Deepset.

Frequently Asked Questions

What is Security-One 27B?

Security-One 27B is an open-weight, 27-billion-parameter decision model released by Superagent, built on Qwen3.8-27B, designed to screen prompts, tool calls, code changes, and alerts for prompt-injection attempts and other security risks by returning calibrated probability scores.

What is the BIPIA benchmark?

BIPIA, or Benchmarking Indirect Prompt Injection Attacks, is a research benchmark that tests how well a model resists instructions hidden inside documents, emails, or web content rather than typed directly by a user. Security-One 27B detected 599 of 600 staged attacks on an 800-example BIPIA slice.

Why did Security-One score lower on the Deepset benchmark?

Deepset’s prompt-injection test set phrases and structures its attacks differently than BIPIA. Security-One’s 78.33% attack-detection rate on Deepset, against 99.83% on BIPIA, shows that strong performance on one benchmark does not guarantee the same result on another, a known limitation Superagent’s own model card acknowledges.

Is Security-One 27B free to use?

The model weights are released under the Apache 2.0 license, so any team can download and self-host it without a licensing fee. Superagent also offers a hosted API version for teams that prefer not to run the GPU infrastructure themselves.

How does Security-One compare to Meta’s Llama Prompt Guard 2?

Llama Prompt Guard 2 is far smaller (86 million or 22 million parameters versus Security-One’s 27 billion), runs faster, and reports a 0.998 AUC score on English text, but it only supports a 512-token context window and carries the Llama 4 Community License rather than Apache 2.0.

Can Security-One 27B replace a human security review?

No. Superagent’s own documentation describes the model as a first-pass triage layer, not a complete security boundary, and explicitly recommends keeping deterministic controls and human review in place for consequential actions.

What hardware does Security-One 27B need to run?

Superagent’s documentation points to a single Nvidia B200 GPU as sufficient for inference, which is a meaningfully larger footprint than the small classifier models from Meta or Horizon Labs, but still far less than running a full frontier-scale language model for the same task.

What should a security team do before deploying any prompt-injection filter?

Test the model against internal traffic and known attack patterns rather than relying solely on a vendor’s published benchmark score, since results can shift meaningfully between test sets, as the gap between Security-One’s BIPIA and Deepset scores demonstrates.