OrcaRouter Shrinks 27B Security AI to Fit One GPU
OrcaRouter has released OrcaSAQ-2 Cyber 27B, a highly compressed security model that allows researchers to run advanced vulnerability analysis locally on a single consumer GPU.

OrcaRouter has launched OrcaSAQ-2 Cyber 27B Uncensored, a compressed version of the Qwen3.8-27B Uncensored model designed specifically for local cybersecurity operations. By applying quantization, the developers reduced the model's footprint by 71.3 percent, shrinking it from 54.7 GB in BF16 format down to just 15.7 GB. This dramatic reduction allows the weights to run efficiently on a single consumer GPU with 16 GB or 24 GB of VRAM, removing the need for expensive enterprise hardware.
Despite the heavy compression, the model maintains high fidelity to its unquantized counterpart. Evaluated on 16,376 predicted WikiText-2 tokens, the GGUF checkpoint demonstrated a mere 0.80 percent increase in perplexity, a 94.4 percent Top-1 agreement rate, and a mean Kullback-Leibler divergence of 0.020 compared to the BF16 baseline. It preserves a massive 262K context window using a hybrid Gated DeltaNet and full-attention architecture. Additionally, the model achieves a single-stream inference speed of 27.6 tokens per second on a 24 GB GPU by utilizing DFlash2 lossless speculative decoding.
For security practitioners, this release enables highly sensitive tasks to be performed entirely offline. Teams can analyze proprietary source code, live system logs, suspicious binaries, and unpatched zero-day vulnerabilities without exposing confidential data to third-party cloud providers. Because the model is an abliterated checkpoint with its learned refusal behaviors removed, researchers can conduct authorized defensive red teaming and exploit-adjacent vulnerability research without encountering artificial safety blocks.
The model is distributed under the Apache-2.0 license in the GGUF format, making it highly accessible for integration into existing workflows. It is fully compatible with popular local LLM runtimes and tools, including llama.cpp, Ollama, LM Studio, vLLM, and Docker Model Runner, allowing developers to deploy it immediately in their local environments.
This is our own summary of reporting by AlphaSignal



