Skip to main content
← Back to Blog

Qwen3.8-27B: The Benchmarks, the Uncensored Forks, and the Real Local-AI Story

• Dataxad Team

Qwen3.8-27B coding benchmarks, local memory needs, independent tests, and the evidence behind viral third-party uncensored derivatives for benign work.

Qwen3.8-27B is exactly the kind of model that makes local AI interesting: small enough to quantize for a well-equipped personal workstation, yet ambitious enough to target repository-level coding, terminal agents, computer use, document understanding, and long-running professional tasks.

It is also a useful test of how quickly four separate stories become tangled online. There is the official Qwen model. There are Qwen’s own benchmark charts. There are independent local tests. And there are third-party “uncensored” derivatives whose release posts went viral on X.

Those stories overlap, but they are not interchangeable.

The evidence supports Qwen3.8-27B as an unusually capable and locally deployable 27B dense model. It does not support calling the official checkpoint uncensored, transferring its BF16 benchmark scores to every quantized fork, or assuming that refusal removal makes a model smarter.

An evidence map separating the official Qwen3.8-27B model, benchmarks, local quantizations, and third-party uncensored forks

First: the exact model name matters

The canonical name is Qwen3.8-27B. Qwen’s official repository dates the 27B weight release to August 14, 2026. The spelling matters because “Qwen 3.8 27B” can be mistaken for Qwen3-8B, an 8-billion-parameter model from an older generation.

The official model card describes a post-trained, dense 27B vision-language model under the Apache-2.0 license. “Open weights under Apache-2.0” is more precise than “fully open source”: users can download and modify the weights under a permissive license, but Qwen has not published the complete training dataset and pipeline.

Architecturally, this is not a conventional all-attention transformer. Its 64 language-model layers are arranged as 16 repeated groups of three Gated DeltaNet linear-attention layers followed by one full-attention layer. It has a 5,120 hidden dimension, grouped-query attention in the full-attention blocks, a vision encoder, and a multi-token-prediction head.

The native context length is 262,144 tokens and Qwen documents extension to one million tokens with YaRN. That is a capability boundary, not a memory guarantee. Qwen itself warns that current static YaRN configurations can reduce short-context quality, so RoPE scaling should be enabled only when it is actually needed.

Thinking is on by default. Qwen exposes xhigh, medium, and low reasoning effort and can preserve reasoning state across turns. Its documentation makes an unusually practical point: reducing reasoning effort can shorten an individual turn while increasing failed attempts and retries, so total task time can still rise.

What the official benchmark table says

Qwen reports broad gains over Qwen3.6-27B, particularly in coding agents, long-horizon work, instruction following, and visual computer use.

BenchmarkQwen3.8-27BQwen3.6-27BOpus 4.6 Max
Terminal Bench 2.1, Terminus73.063.478.2
SWE-bench Pro61.753.553.4
NL2Repo-Bench42.336.247.6
DeepSWE 1.142.213.3—
QwenSWEBench79.049.363.8
CoWorkBench70.761.068.2
JobBench33.421.8—
IFBench79.569.162.5
GPQA Diamond89.287.891.3
Humanity’s Last Exam30.824.040.0
LiveCodeBench v690.383.988.8
OSWorld-Verified84.363.972.7
WebArena-Verified64.848.8—
AndroidWorld81.970.362.0

These are Qwen-reported results, not one independent tournament with identical conditions in every column. The methodology notes are part of the result:

  • Most SWE-bench Pro models were rerun through a Claude Code harness at temperature 1.0, top-p 0.95, and 256K context. The Opus figure is its separately reported official score.
  • Qwen says problematic tasks were corrected before its baseline reruns.
  • QwenSWEBench and CoWorkBench are in-house benchmarks.
  • QwenSWEBench reports average-at-three with an eight-hour timeout and up to 32,768 output tokens.
  • Humanity’s Last Exam uses GPT-4o judging, while Vision2Web uses another model as judge.
  • Several visual benchmark annotations were manually corrected.

The safe reading is “Qwen reports a large generational gain and several notable wins.” The table does not prove “Qwen3.8-27B beats Opus overall.” It loses on Terminal Bench, NL2Repo, GPQA Diamond, and Humanity’s Last Exam in the same published table.

Independent testing narrows the claim

A reproducible community study on an RTX 5090 provides a valuable counterweight. On that hardware, quantization, runtime, and agent scaffold, Qwen3.8-27B scored 56.2% on a full k=1 Terminal Bench 2.1 run—not Qwen’s reported 73.0. It still beat Qwen3.6-27B on the same setup, 56.2% to 48.3%.

The same study found Qwen3.8 at 66.2% on SWE-bench Verified versus Qwen3.6 at 69.4%, a difference the author treats as close enough to require caution. It documented timeouts, parser behavior, harness failures, sampling choices, and the model’s tendency to submit earlier.

That is not a reason to dismiss the launch benchmarks. It is the reason to test the exact deployment. A model score is the product of weights, quantization, runtime, prompt, tools, scaffold, timeout policy, parser, and sometimes a judge model. Change enough of those and the number changes with them.

Artificial Analysis offers another view of the hosted model. Its August 30 snapshot placed Qwen3.8-27B at an Intelligence Index of 52 with xhigh reasoning, falling to 44 at medium, 43 at low, and 35 without reasoning. The measured Alibaba API output rate was roughly 46–56 tokens per second across modes. Those are hosted API results, not local GGUF or Apple Silicon benchmarks, and both scores and rankings can change.

How local is “local” for a 27B dense model?

The official BF16 weights occupy about 55.6 GB. The official FP8 checkpoint is about 30.9 GB. Community GGUF conversions from Unsloth range from roughly 29 GB at Q8_0 to 16.5 GB at Q4_K_M, 12 GB at IQ3_S, and 8.4 GB at IQ2_S. A vision projector adds roughly another 0.9 GB when used.

ArtifactApprox. file sizePractical planning boundary
Official BF1655.6 GB64 GB can be tight; 96 GB gives safer runtime and context headroom
Official FP830.9 GBPlan around 40 GB or more for weights, cache, and runtime
GGUF Q8_029 GB32 GB is usually too tight for a full local workflow
GGUF Q6_K22 GB32 GB is a plausible workstation target
GGUF Q5_K_M19.8 GB24 GB may be tight once context and vision are added
GGUF Q4_K_M16.5 GB24 GB is the defensible minimum; 32 GB is more comfortable
GGUF IQ3_S12 GBLower memory, with a larger quality tradeoff to validate

File size is not runtime memory. KV cache, recurrent state, inference-engine allocations, vision components, and the operating system all compete for capacity. A model that loads is not necessarily a model that can sustain the context length or concurrency you need.

Our updated Local AI Capacity Planner now includes Qwen3.8-27B’s observed Q4 size, architecture-derived full-attention KV estimate, and the new Apple M6 and M5 Ultra configurations. Use it to shortlist hardware, then run the exact artifact before purchasing or promising a deployment.

Qwen officially documents Transformers, vLLM, SGLang, and TokenSpeed support; its repository also points to llama.cpp/GGUF and MLX routes. Framework support evolves quickly, so “supported” should still be verified against the features you need: vision, YaRN, thinking preservation, and multi-token prediction may arrive on different schedules.

The uncensored model is not an official Qwen release

The viral “uncensored Qwen” story refers to third-party modified checkpoints. It helps to keep four terms separate:

  • Open weights means the parameters are downloadable under a license.
  • Uncensored or abliterated usually means a third party has modified refusal behavior.
  • Fine-tuned means additional training changed the model’s behavior.
  • Quantized means numerical precision was reduced to change size and performance.

A model can be several of these at once, but none is a synonym for another.

The most visible derivative came from OrcaRouter in BF16, FP8, GGUF, and NVFP4 forms. Its current model card describes directional ablation or orthogonalization: a refusal direction was identified in the residual stream and removed from matrices that write into it. That is a new behavior-modified checkpoint, not a hidden setting in Qwen’s official weights.

What actually trended on X

The highest-engagement post we directly observed about the derivative was OrcaRouter’s August 15 FP8 release post. On August 30 its public page showed approximately 7.6 million views, 7,000 likes, 657 reposts, 133 replies, and 6,500 bookmarks. The post framed the weights for AI red teaming and security research.

That virality is real attention. It is not a benchmark.

Direct X postAug. 30 snapshotWhat it establishes
FP8 uncensored release7.6M views; 7K likes; 657 reposts; 6.5K bookmarksA third-party derivative became highly visible
NVFP4 derivative release126K views; 1.4K likes; 64 reposts; 1.3K bookmarksStrong interest in a smaller Blackwell-oriented artifact
Hugging Face ranking claim149.6K views; 222 likes; 25 reposts; 175 bookmarksA point-in-time popularity claim, not a quality ranking
Hosted API announcement42K views; 143 likes; 10 repostsAvailability through the uploader’s service
Acceptable-use clarification15K views; 18 likesEven the hosted “uncensored” route retains an abuse policy and guardrails

Engagement counts are dynamic. X search required a login during research, so this is the highest-engagement set we directly observed, not a claim that we exhaustively ranked every post on the platform.

The distinction between downloadable weights and a hosted API also matters. A provider may apply acceptable-use rules, filters, logging, or account controls around a derivative even when the downloadable checkpoint itself has weaker refusal behavior.

Does refusal removal help ordinary work?

The plausible benign benefit is narrower—and more useful—than the hype: fewer false refusals.

A coding assistant can encounter legitimate requests that superficially resemble risky material. Examples include auditing an authorized application, reviewing exploit mitigations, analyzing suspicious code in a sandbox, discussing abuse patterns in a safety report, writing defensive filters, or handling frank medical, legal, political, and fictional text. An overcautious model may stop even when the task is permitted and well-scoped.

OrcaRouter’s own limited evaluation reports benign over-refusal on XSTest-safe falling from 5.6% for stock Qwen to 0.4% for its derivative in non-thinking mode. Its small capability samples changed by +0.4 on MMLU, −0.8 on MMLU-Pro, −1.3 on GSM8K, and −0.6 on CMMLU.

Those results are uploader-reported, small-sample, text-only, and based partly on a rule-based refusal classifier. The card explicitly says they are not publication-grade. They support “less likely to refuse this sample,” not “more intelligent.”

Evidence from other derivatives reinforces the boundary. OBLITERATUS reports its modified version dropping about 2.1 MMLU points, with a larger STEM decline; stock and modified versions each completed seven of eight selected real-world tasks. A Jonathan Coletti derivative reports far fewer refusals with an average capability decline near half a point in limited zero-shot tests, while explicitly noting that coding, math, multilingual, generative, and vision performance were not evaluated.

The reasonable conclusion is that refusal modification can reduce friction while preserving much—but not necessarily all—general capability. It has not been shown to improve coding correctness.

Why fewer refusals can create more engineering work

Refusal behavior is only one safety layer, and often a blunt one. Removing it does not add factuality, authorization, least privilege, or tool isolation. In an agent, a compliant wrong action can be more damaging than a refusal.

The independent RTX 5090 study did not publish a prompt-injection evaluation for Qwen3.8. That missing evidence is itself an important boundary: strong coding and agent benchmarks do not establish resistance to hostile tool output. Local control makes the surrounding harness more important, not less.

For benign production use, pair any low-refusal derivative with:

  • a sandboxed runtime and allowlisted tools;
  • read-only defaults for repositories, browsers, and databases;
  • explicit confirmation before sending messages, spending money, changing production, or deleting data;
  • provenance and logs for retrieved content and tool output;
  • task-level evaluations on both success and unsafe side effects;
  • a separate policy layer when users or the public can submit prompts.

This is not about restoring a single universal filter. It is about moving control from an opaque refusal inside the model to explicit, auditable boundaries around the system.

Where Qwen3.8-27B is genuinely compelling

The official model’s strongest local use cases do not depend on an uncensored fork:

Private repository work

A 27B model can be kept near proprietary code and documents. Qwen’s coding and agent benchmark gains make it a serious candidate for issue analysis, test generation, constrained refactors, code review, and repository search—provided the team evaluates its own stack.

Multimodal technical work

Native image and video input means one model can reason about screenshots, diagrams, PDFs, UI state, and code. That reduces the integration burden of routing every visual task through a separate model.

Long-running office and research tasks

CoWorkBench, JobBench, and instruction-following gains point toward document synthesis, spreadsheet-adjacent analysis, research organization, and tool-driven workflows. Again, the published scores are evidence to test, not a service-level agreement.

A controllable reasoning budget

Reasoning effort and preserved thinking offer a practical knob for different tasks. Low effort may suit formatting or extraction; complex coding may justify xhigh. Measure complete task time and retry rate, not only tokens per turn.

Offline continuity

Once weights and runtime are installed, a local checkpoint can keep working through an API outage or policy change. The trade is operational ownership: model storage, updates, evaluation, security, and hardware are now yours.

Frequently asked questions

Is Qwen3.8-27B officially uncensored?

No. Qwen released an Apache-2.0 open-weight model. “Uncensored” checkpoints are third-party modifications from OrcaRouter and other community publishers.

Does the uncensored fork code better?

There is no credible evidence for that conclusion. Limited derivative evaluations show fewer refusals with roughly preserved or somewhat reduced general scores, and several cards explicitly did not test coding.

Can Qwen3.8-27B run in 16 GB?

A heavily quantized file can be smaller than 16 GB, but the widely used Q4_K_M GGUF is about 16.5 GB before a vision projector, context cache, runtime allocations, and the operating system. Plan around 24 GB as a practical minimum for Q4 and validate the exact workload.

Does the model beat Opus 4.6 Max?

It wins some rows in Qwen’s table and loses others. Different harnesses and reporting methods also limit direct comparison. “Competitive on several published tasks” is defensible; “better overall” is not.

Is 262K context free once the weights fit?

No. Context consumes memory and compute. Very long context can materially change capacity and latency, and Qwen warns that static long-context scaling can hurt shorter inputs.

Bottom line

Qwen3.8-27B matters because it compresses a broad agentic and multimodal capability set into a dense model that local hardware can plausibly run. Qwen’s launch results are strong enough to justify evaluation; independent results are different enough to reject automatic conclusions.

The uncensored derivatives matter for another reason. They expose real demand for models that do not falsely refuse legitimate, sensitive-looking work. The viral X posts quantify that demand. They do not prove a capability gain, and they do not remove the need for sandboxing, authorization, and measured task-level evaluation.

Use the official checkpoint when its behavior fits. Test a derivative when false refusals are the actual bottleneck. In either case, evaluate the exact quantization, runtime, tools, and hardware that will do the work.

Sources and methodology

Model files, cards, benchmark pages, X engagement, and availability were checked on August 30, 2026. Dynamic counts are labeled as snapshots. Vendor and uploader results are labeled by provenance; the independent local study is identified separately.

Need help implementing this?

Book a Consultation