Nace.AI has open-sourced Drex 1.5, a decision model built to score options instead of writing paragraphs. The company published the weights on September 28, 2026, and the release has since been picked up across AI trade press as a notable entry in the growing field of small, task-specific models built for enterprise pipelines rather than chat windows.

The headline number is 58.08, the score Drex 1.5 posted on Decision Index 0.3.1, a benchmark built around choosing, ranking, and judging rather than generating prose. Nace.AI says that result makes Drex 1.5 the top-ranked model under 10 billion parameters on that board. An earlier version of the same benchmark, Decision Index 0.3.1, put Drex 1.5 in first place under 10B parameters, according to AI Daily Post’s coverage of the release. For readers tracking the open-weight AI hardware story, Drex 1.5 is a compact example of a bigger shift: models built to run cheaply on modest GPUs while still beating larger systems at a narrow job.

What Nace.AI Shipped on September 28

Drex 1.5 carries the official model identifier drex-v1.5. Nace.AI lists it at roughly 9 billion parameters, reported more precisely as 8.95B, shipped in bf16 precision at approximately 18 GB on disk. That puts it squarely in the same weight class as popular general-purpose open models, but the design goal is different. Rather than compete on writing quality, Drex 1.5 is meant to replace the layer of custom classifiers, rule engines, and reranking scripts that production AI systems often bolt on around a chat model.

The weights and model card are hosted on Hugging Face under the nace-ai/drex-v1.5 repository, which lists the architecture, context limits, and benchmark subscores Nace.AI reported at launch. Nace.AI describes Drex 1.5 as live, with a 128K context window, sub-second latency, and pricing of $0.04 per 1 million input tokens, plus a 250 million token free allowance for developers who want to test it before committing to production traffic.

How Drex 1.5 Decides Without Writing a Word

Most language models answer a question by generating a sequence of tokens one at a time, sampling from a probability distribution until they reach a stop condition. Drex 1.5 skips that step entirely. According to Nace.AI, “it does not write.” Instead, a caller feeds the model a state (plain text or structured JSON) plus a set of typed questions, and the model returns a probability for every option supplied, all in a single forward pass.

Nace.AI explained the mechanic directly: “You give it a state and a set of typed questions, and one forward pass returns a probability for every option you supplied.” Drex 1.5 supports three question formats: choice (pick among several labeled options), noul (a yes/no judgment), and score (an ordinal rating). That typed-question design is what separates Drex 1.5 from a general chatbot asked to “pick the best answer” in free text, where the output has to be parsed back into a usable format and often drifts from the expected schema.

In practice, this kind of model slots into places where engineers currently use embeddings plus a reranker, a fine-tuned classifier, or a prompt that asks a general LLM to output “A, B, or C” and hope it follows instructions. Nace.AI frames the result as a calibrated probability on every option, in one forward pass, a phrasing that signals the company is competing on decision quality and consistency rather than fluency.

Inside the Architecture: A Qwen-Based Backbone Plus a Pointer Head

Drex 1.5’s model card, published in the nace-ai/drex-decision-models repository on GitHub, identifies the base as MiMo-V2.6-Distill-Qwen-9B with an added Kev pointer head. Nace.AI’s own description is blunt about the setup: “A decision model built on MiMo-V2.6-Distill-Qwen-9B with a Kev pointer head.”

Starting from a distilled Qwen-family base gives Nace.AI a model that already understands language, reasoning chains, and long documents. The pointer head is the part that turns that general competence into structured output. Instead of letting the model choose the next token from its full vocabulary, the pointer head constrains its attention to the finite set of options the caller supplied, then converts the model’s internal representation into a probability over just those options. That is a meaningfully different inference path than asking a chat model to output JSON and parsing the response, since there is no token generation, no retry logic for malformed output, and no risk of the model inventing an option that was not on the list.

Drex 1.5 by the Numbers

The table below collects the specifications Nace.AI has published for Drex 1.5 since the September 28 release.

SpecDetail
DeveloperNace.AI
Model name / IDDrex 1.5 (drex-v1.5)
Parameters~9B (reported as 8.95B)
Precision / disk sizebf16, approximately 18 GB
Base modelMiMo-V2.6-Distill-Qwen-9B
Added componentKev pointer head
Context window131,072 tokens (128K)
Supported question typeschoice, noul (yes/no), score
Weights availableSeptember 28, 2026
API price$0.04 per 1M input tokens
Free allowance250M tokens
Latency claimSub-second (per Nace.AI)
Decision Index 0.3.1 score58.08, #1 under 10B parameters

The Decision Index Score: 58.08, and What It Actually Measures

Decision Index is the benchmark Nace.AI uses to rank Drex 1.5 against other compact models, and the company has run Drex through two versions of it. On Decision Index 0.3.1, Drex 1.5 scored 58.08, which Nace.AI lists as the highest result among models under 10 billion parameters. On an earlier pass, Decision Index 0.2.1, Nace.AI reported Drex 1.5 at 58.08 on Decision Index 0.3.1, placing first under 10B parameters.

Those two numbers are not directly comparable to each other since they come from different benchmark versions, and that distinction matters for anyone reading the scores at face value. A model’s rank on 0.2.1 does not automatically translate to the same rank on 0.3.1, even when the raw score lands in a similar range. What both results share is the same basic claim: among open and hosted models small enough to run on a single consumer or workstation GPU, Drex 1.5 is positioned by its maker as the strongest option for structured decision tasks specifically, not for open-ended writing, coding, or general knowledge recall.

It is also worth separating benchmark marketing from independent verification. The Decision Index numbers currently in circulation are self-reported by Nace.AI, a pattern that is common across the open-weight model market and not unique to this release. Readers evaluating Drex 1.5 for a production workload should run their own evaluation against their own classification or ranking task before trusting a vendor leaderboard position.

Pricing: $0.04 Per Million Tokens, With a 250M-Token Trial

Nace.AI set Drex 1.5’s hosted API price at $0.04 per 1 million input tokens, with a free allowance of 250 million tokens for developers evaluating the model. That pricing structure targets high-volume, low-margin workloads: content moderation queues, lead-scoring pipelines, support-ticket routing, and other jobs that run millions of small classification calls a day rather than a handful of long conversations.

Drex 1.5 is also listed as a hosted model on OpenRouter’s model directory, which gives developers a second, provider-neutral route to call the model through a single API key alongside other open-weight options. For teams that want to self-host instead, Nace.AI’s developer documentation covers authentication for the hosted endpoint at console.nace.ai’s Drex authentication guide, while the open weights on Hugging Face let anyone download and run the model without paying per-token fees at all, at the cost of managing their own GPU capacity.

Context Window and Latency: 128K Tokens, Claimed Sub-Second Response

Drex 1.5 accepts up to 131,072 tokens of context, roughly 128,000 tokens, which covers long contracts, multi-document support tickets, or full codebases when the task is to classify or score rather than summarize. Nace.AI describes the resulting latency as sub-second, a claim that, if it holds at the high end of that context range in production, would make Drex 1.5 fast enough to sit inline in request paths that currently can’t afford the delay of a generative model’s token-by-token output.

That speed advantage is structural rather than incidental. Because Drex 1.5 returns a probability distribution in one forward pass instead of streaming tokens, its latency profile looks more like a traditional classifier’s than a chatbot’s, even while it carries a 9B-parameter language backbone capable of reading long, messy real-world documents. That combination, large-model comprehension with small-model response time, is the core pitch Nace.AI is making to engineering teams evaluating whether to replace brittle rule-based routing logic.

What It Takes to Run Drex 1.5 on Your Own Hardware

Because the weights are open, Drex 1.5’s roughly 9 billion parameters can be run on local hardware rather than through Nace.AI’s hosted endpoint, and the math follows the same rules as any other model its size. At full bf16 precision, Drex 1.5’s reported 18 GB footprint needs headroom for the model plus the KV cache that long contexts demand, which in practice points toward a GPU with at least 24 GB of VRAM for comfortable single-GPU inference. Quantized down to 8-bit, the weight footprint drops to roughly 9 GB, often workable on 12 to 16 GB cards. Pushed to 4-bit, weights shrink to around 4.5 GB, within reach of many 8 to 12 GB consumer GPUs, though accuracy and long-context stability typically soften at that level of compression.

PrecisionApprox. weight sizePractical minimum GPU VRAM
bf16 / FP16 (full)~18 GB24 GB or more
8-bit quantized~9 GB12-16 GB
4-bit quantized~4.5 GB8-12 GB
CPU / system RAM onlyWeights plus runtime overheadNo GPU required, slower

That 128K context window complicates the picture, since the figures above cover model weights only. A long-context KV cache adds real memory on top of the base footprint, so running Drex 1.5 near its full 128K limit will generally need more VRAM than the quantization table suggests, or a serving stack that offloads part of the cache to system RAM. For most local deployments, keeping working context in the 8K-32K range is the more realistic target on consumer-grade hardware, which still covers the bulk of document review, support-ticket, and moderation use cases most teams reach for a decision model to solve.

Competitive Landscape: Decision Models vs General-Purpose LLMs

Drex 1.5 is not the first model built specifically to score rather than generate. Shattered.io reported last month on Microsoft’s Decision-1, another 9B-parameter model that Microsoft said processes decision workloads dramatically faster than conventional generation-based approaches. The fact that two separate companies shipped similarly sized decision-specific models within weeks of each other suggests the category is forming around a 9B parameter sweet spot: large enough to understand complex, multi-part business logic, small enough to run cheaply at the volumes classification workloads require.

That contrasts sharply with the trajectory of general-purpose open models, which have been racing upward in parameter count. Mistral’s recently previewed Large 4 model reportedly carries over a trillion total parameters with tens of billions active at inference, and Nvidia’s Nemotron 3.5 line has pushed from 30B toward a stated roadmap goal of a trillion parameters. Those systems are optimized for broad reasoning, coding, and conversation. Drex 1.5 is optimized for the opposite problem: doing one narrow job, cheaply, at scale, without the overhead of a model built to do everything.

Put differently, buyers evaluating AI infrastructure in late 2026 increasingly have two separate shopping lists. One is for flagship reasoning and generation models, where parameter count and benchmark leaderboards like those driving Mistral’s and Nvidia’s releases still dominate the conversation. The other is for workhorse models like Drex 1.5 that never touch a chat interface at all, instead running quietly behind an API gateway, scoring millions of routine decisions a day. Nace.AI is explicitly building for the second list.

Historical Context: From Rule Engines to Decision-Native Models

The lineage behind Drex 1.5 runs through a decade of machine learning infrastructure that predates the current generative AI boom. Before transformer-based language models existed, companies routing support tickets or scoring fraud risk relied on hand-tuned logistic regression models, gradient-boosted trees, or rigid rule engines. Those systems were fast and cheap to run but brittle, requiring constant retraining whenever the underlying data distribution shifted even slightly.

When transformer-based embeddings and encoder models like BERT arrived, they replaced much of that hand-tuned infrastructure with learned representations that generalized far better, though still as narrow classifiers bolted onto whatever the business needed scored. The generative AI wave that followed flipped the pattern again, with engineering teams increasingly prompting a general chatbot to “classify this as A, B, or C” because it was available and good enough, even though it was neither the fastest nor the cheapest way to get a reliable label.

Drex 1.5 represents something closer to a return trip, applying a modern, Qwen-derived language backbone to the narrow classification problem that encoder models used to own, but with the broader world knowledge and long-context comprehension that only became possible once large language models existed. That combination, old problem, new backbone, is the throughline connecting Drex 1.5 to the broader open-source AI story Shattered.io has tracked through releases like Reflection AI’s Beam and the wider push toward efficient, openly licensed alternatives to closed frontier models.

Market Impact: Why an Open-Sourced Decision Model Matters Now

Open-sourcing Drex 1.5 rather than keeping it behind a closed API is a deliberate distribution strategy. By putting the weights on Hugging Face at the same time it opened a paid hosted endpoint, Nace.AI is betting that developer trust and adoption built through free experimentation will convert into hosted API revenue once teams move from prototyping to production scale. That is the same playbook several other open-weight AI companies have run in 2026, part of a broader shift Shattered.io covered when the open-source share of newly released frontier-adjacent models climbed sharply this year as companies like Mistral and Reflection AI shipped openly licensed systems alongside, not instead of, paid hosting.

For enterprise buyers, the practical impact is a lower floor on the cost of adding structured decision-making to an AI pipeline. At $0.04 per million input tokens, with 250 million tokens free, a team can prototype a moderation or routing system for a few dollars rather than committing to a generative model’s higher per-call cost for a job that never needed free-text generation in the first place. That price point also puts competitive pressure on both the encoder-classifier vendors that have dominated this niche and on general-purpose model providers that have been pricing classification calls at generative-model rates.

There is a hardware angle too. Because Drex 1.5 fits on a single mid-range GPU even before quantization, it lowers the barrier for companies that want to run AI decision infrastructure on-premises for compliance or latency reasons, rather than routing sensitive data through a third-party API at all. That is a meaningfully different infrastructure footprint than the multi-GPU clusters required for trillion-parameter models like Nvidia’s Nemotron roadmap or Mistral’s newly previewed Large 4, and it is likely to appeal specifically to regulated industries, healthcare, finance, and government contractors among them, that have been slower to adopt generative AI broadly.

Analysis: Strengths and Open Questions

Drex 1.5’s core strength is architectural clarity. By constraining the model’s output to a calibrated probability over a finite, caller-supplied set of options, Nace.AI sidesteps an entire category of production headaches that come with asking a generative model to behave like a classifier, malformed JSON, hallucinated options outside the supplied list, and inconsistent formatting across calls. That alone is a real engineering win for anyone currently gluing together prompt templates and regex parsers to extract a decision from a chat model’s free-text reply.

The open questions sit mostly around independent verification and durability of the benchmark lead. The Decision Index scores driving Drex 1.5’s marketing come from Nace.AI itself, and the gap between Drex 1.5 and whatever sits in second place under 10B parameters has not been independently reproduced in a published third-party evaluation as of this writing. Benchmark leadership in the fast-moving small-model category has also proven short-lived before. A narrow lead claimed in September can evaporate within weeks once a competitor ships an update, so treating 58.08 as a durable crown rather than a snapshot would be premature.

There is also the practical matter of ecosystem maturity. Drex 1.5’s typed-question interface, choice, noul, and score, is a clean design, but it is also a new convention that developer tooling, orchestration frameworks, and monitoring dashboards have not yet built native support for. Teams adopting it early will likely need to write their own glue code around the API shape until the surrounding tooling catches up, a cost that is easy to underestimate when a model’s benchmark numbers look this competitive on paper.

Predictions: Where Decision-Native Models Go From Here

  • More vendors will ship their own decision-specific models in the 7B-13B range over the next two quarters, following the pattern set by both Drex 1.5 and Microsoft’s Decision-1, as the category proves it can win dedicated benchmark attention separate from general LLM leaderboards.
  • Independent benchmark groups will likely publish third-party Decision Index evaluations or a competing standardized benchmark within the next few months, since vendor-reported scores alone will not satisfy enterprise procurement teams evaluating multiple options.
  • Expect hosted API pricing for decision/classification-only models to keep falling toward the sub-$0.05-per-million-token range Nace.AI has already set, pressuring both legacy classifier vendors and general LLM providers that currently charge generation-model rates for simple scoring calls.
  • Orchestration and agent frameworks will likely add native support for typed-question decision APIs like Drex 1.5’s choice/noul/score format within the next year, formalizing what is currently a bespoke integration pattern.
  • On-premises and regulated-industry adoption of small open-weight decision models will grow faster than cloud-hosted adoption, driven by the same data-residency and latency concerns that have slowed generative AI rollout in healthcare, finance, and government settings.

Frequently Asked Questions

What is Drex 1.5?
Drex 1.5 is an open-weight, roughly 9-billion-parameter decision model from Nace.AI, released on September 28, 2026. Instead of generating text, it takes a state plus typed questions and returns a probability for each supplied option in a single forward pass.

How is Drex 1.5 different from a chatbot like ChatGPT or Claude?
Chatbots generate free-form text token by token and are typically asked to format a decision as part of their written response. Drex 1.5 skips text generation entirely and returns structured probabilities over a fixed set of options the caller defines in the request.

How much does Drex 1.5 cost to use?
Nace.AI’s hosted API prices Drex 1.5 at $0.04 per 1 million input tokens, with a 250 million token free allowance for developers testing the model before committing to production use.

What hardware do I need to run Drex 1.5 locally?
At full bf16 precision, Drex 1.5’s roughly 18 GB footprint is comfortable on a GPU with 24 GB or more of VRAM. Quantized to 8-bit it fits in roughly 9 GB, and 4-bit quantization brings it down to around 4.5 GB, workable on many consumer GPUs with 8 to 12 GB of VRAM, though long-context inference near the full 128K window will need additional memory for the KV cache.

What is the Decision Index benchmark?
Decision Index is the benchmark Nace.AI uses to evaluate decision-making models rather than text generation. Drex 1.5 scored 58.08 on Decision Index 0.3.1 and 58.28 on the earlier Decision Index 0.2.1, both reported by Nace.AI as the top result among models under 10 billion parameters.

Is Drex 1.5 really open source?
Yes. Nace.AI published the model weights on Hugging Face alongside the model card detailing its architecture and benchmark results, in addition to offering a separate paid hosted API for teams that prefer not to self-host.

How does Drex 1.5 compare to Microsoft’s Decision-1?
Both are roughly 9-billion-parameter models built specifically for decision and classification tasks rather than open-ended generation, released within weeks of each other in 2026. That overlap suggests multiple companies are converging on a similar parameter range for this category, though the two models use different architectures and have not been directly benchmarked against each other in a published head-to-head test.

What’s next for decision-native AI models?
Expect more vendors to release similarly sized decision-specific models, independent benchmarks to emerge to verify vendor-reported scores, and hosted pricing for this category to keep falling as competition increases through the rest of 2026 and into 2027.