Qwen 3.8-27B is a strong local reasoning model, but its default `xhigh` mode can spend a large number of tokens before answering. ThinkingCap and Swift are post-trained derivatives designed to reduce that reasoning overhead while preserving most of the base model's accuracy. The useful question is not simply which model has the highest benchmark score. It is which trade-off—baseline fidelity, shorter traces, licensing, or consistency—fits your workload.
Quick answer
ThinkingCap and Swift both cut Qwen 3.8-27B reasoning substantially while staying close to the base model on published evaluations. In one independent Aider comparison, ThinkingCap matched the base retry-pass rate with 41% fewer median tokens, while Swift was slightly faster and shorter. Choose only after testing your own prompts, hardware, quantization, and license requirements.
What people searching this comparison usually need to know
The current search intent around these three models is practical rather than academic. Users want to know whether a reasoning-efficient fine-tune is a safe drop-in replacement for Qwen 3.8-27B, how much latency it can save, whether accuracy changes, which model is better for coding, and whether the weights can be used commercially.
Those questions cannot be answered by one headline percentage. The published results come from different harnesses, output caps, seeds, quantizations, and hardware. A vendor's self-evaluation can show whether its own derivative behaves as intended under controlled conditions. An independent test can show whether the claimed direction survives a different workload. Neither alone proves universal superiority.
The strongest conclusion supported by the available evidence is narrower: both derivatives reduce reasoning length materially; both remain close enough to the base model on several tests to merit workload-specific evaluation; and the base remains the clean reference when maximum fidelity and Apache-2.0 licensing simplicity matter more than efficiency.
The three models in one table
| Model | What it is | Published efficiency claim | Main strength | Main caveat |
|---|---|---|---|---|
| Qwen 3.8-27B | Original 27B dense vision-language model | Reference baseline | Official model, broad capability, Apache-2.0 | Can overthink and produce long reasoning traces |
| ThinkingCap-Qwen3.8-27B | BottleCap AI fine-tune of Qwen 3.8-27B | 37.2% lower mean thinking tokens across its 12-benchmark suite | Close macro accuracy, detailed evaluation, several deployment formats | PolyForm Small Business license plus personal-use grant; longer tail than Swift in one independent test |
| Swift-Qwen3.8-27B | UkisAI reasoning-efficient derivative | 41.0% lower mean thinking at `xhigh`; 58.3% lower median on GPQA-Diamond | Very short traces and consistent efficiency | Swift Open License has a commercial revenue threshold; some accuracy losses vary by benchmark |
All three share the same basic 27B Qwen lineage, but they are not interchangeable from a governance perspective. The model card, license, quantization, serving stack, and version should be recorded as part of any deployment decision.
Qwen 3.8-27B: the reference baseline
The official Qwen model card describes Qwen 3.8-27B as a 27-billion-parameter dense causal language model with a vision encoder. It supports image and video understanding, tool use, agent execution, and adjustable reasoning effort. Its native context length is 262,144 tokens and Qwen documents extension up to one million tokens with scaling techniques.
Qwen publishes strong official results, including 73.0 on Terminal-Bench 2.1, 61.7 on SWE-bench Pro, 79.5 on IFBench, and 89.2 on GPQA Diamond. These numbers establish why the base model is attractive, but they should not be mixed directly with derivative results from other harnesses. Benchmark implementations, agents, timeouts, output limits, and sample counts can change the result.
The base model's advantage is clarity. It is the upstream artifact, carries Apache-2.0 licensing, and is the model against which both fine-tunes define their gains. If your system already meets latency and throughput targets, remaining on the base can be the lowest-risk choice.
Its disadvantage is the reason these derivatives exist. Qwen enables thinking by default and allows `reasoning_effort` to be adjusted. At high effort, difficult prompts can generate very long traces. That may improve difficult-task performance, but it can also increase wall-clock latency, consume output budgets, reduce concurrency, and make a local interface feel slow.
Before replacing the model, test whether simply lowering `reasoning_effort` solves the problem. Swift's model card provides a useful example: on GPQA-Diamond, the base at `medium` used fewer tokens than Swift at `xhigh`, but its score was about four percentage points lower. That illustrates the real decision: a fine-tune may preserve high-effort accuracy more effectively than lowering the base model's effort, but the outcome depends on the task.
ThinkingCap: broad token savings with close macro accuracy
BottleCap AI reports ThinkingCap against the BF16 Qwen base across twelve benchmarks at `reasoning_effort=xhigh`, using the same sampling settings and serving stack. Its macro-average accuracy was 85.8 versus 86.6 for the base, while the equal-weight average reduction in thinking tokens was 37.2%.
The per-benchmark results matter more than the macro headline. ThinkingCap reduced mean thinking tokens by 43.1% on GPQA-Diamond, 57.3% on MMLU-Pro, 65.5% on MMMLU, 46.4% on IFBench, and 38.6% on the long-context AA-LCR evaluation. Accuracy changes were usually small, but not uniformly zero. AIME 2026 fell from 98.13% to 94.27%, while AA-LCR rose from 81.75% to 84.00% and LiveCodeBench v6 was effectively flat.
This profile suggests a conservative interpretation. ThinkingCap is not proven to be a smarter general model. Its stated objective is to preserve the answer quality and style of Qwen while reducing unnecessary reasoning. The evidence supports meaningful efficiency gains with a small average accuracy cost, plus variation by domain.
The evaluation is unusually well documented. BottleCap lists hardware, vLLM version, speculative decoding settings, output caps, task counts, seeds, confidence intervals, truncation, and looping behavior. That detail improves auditability, but it remains an evaluation produced by the model's creator.
Deployment options include BF16, FP8, GGUF, NVFP4, and an NVFP4 W4A4 build. The model card describes use with Transformers, vLLM, SGLang, llama.cpp-compatible tools, LM Studio, and Ollama-compatible GGUF workflows. Quantization can change both quality and performance, so the BF16 claims should not be copied automatically onto a particular 4-bit build.
Swift: aggressive efficiency with a different trade-off profile
UkisAI describes Swift as a reasoning-efficient derivative trained by identifying tokens associated with overthinking, penalizing them during fine-tuning, and restoring accuracy through further post-training. The model card says Swift also uses a transfer component derived from BottleCap AI's earlier ThinkingCap-Qwen3.6-27B work.
In UkisAI's BF16 comparison, Swift reduced mean thinking tokens by 41.0% on GPQA-Diamond while moving from 88.38% to 88.28%. It reported reductions of 46.2% on MMLU-Pro, 42.2% on IFBench, 26.7% on AIME 2026, 26.5% on Terminal-Bench 2.1, and 24.3% on LiveCodeBench v6. Median-token reductions were often larger than mean reductions.
Accuracy again varied. MMLU-Pro moved from 85.47% to 84.95%; AIME 2026 from 98.67% to 94.00%; Terminal-Bench 2.1 from 66.74% to 65.84%; and LiveCodeBench v6 rose from 76.76% to 81.55%. UkisAI warns that the apparent LiveCodeBench improvement is influenced by truncation behavior and should not be read as a simple capability gain.
Swift's main practical appeal is consistency of shortening. In the independent Aider evaluation discussed below, its median completion tokens and seconds per case were slightly lower than ThinkingCap's, and the tester reported a shorter mean tail. That does not prove Swift is universally faster, but it gives a plausible reason to test it when long outlier traces are the operational problem.
Swift also exposes a research API at the time of review, but an externally hosted free endpoint should be treated as an evaluation convenience rather than production infrastructure. Availability, rate limits, privacy, and terms can change. Local deployment or a controlled provider remains the safer basis for a durable system.
What the independent Aider comparison adds
A community comparison tested Q8_0 versions of all three models with llama.cpp 0.5.0, two runs per model, and an Aider-focused coding suite. This is valuable because it uses the same harness and quantization across the three candidates.
| Model at `xhigh` | First-try pass | Retry pass | Well-formed diff | Median tokens | Seconds per case | Tokens per solve |
|---|---|---|---|---|---|---|
| ThinkingCap | 27.1% | 77.6% | 100.0% | 7,436 | 777 | 12.8K |
| Qwen base | 27.1% | 77.6% | 99.1% | 12,547 | 1,481 | 19.3K |
| Swift | 30.8% | 75.7% | 98.1% | 7,301 | 750 | 12.1K |
The comparison supports three useful points.
First, both derivatives cut median tokens by roughly forty percent relative to the base in this workload. Second, ThinkingCap matched the base model's retry-pass result exactly in these runs, while Swift was within the tester's estimated noise range. Third, the efficiency gain also appeared as lower seconds per case, not merely shorter visible traces.
The limits are equally important. Two runs do not establish a universal ranking. The suite focuses on coding and diff generation. Language-specific results varied: the author reported meaningful differences across C++, JavaScript, and Python, while other languages were statistically similar. The test used Q8_0, so it cannot validate every GGUF, AWQ, NVFP4, or MLX build.
This is why “Swift wins” or “ThinkingCap wins” is too strong. The independent data says both are credible efficiency candidates and that their differences are small enough to require workload-specific testing.
Which model should you choose?
Choose Qwen 3.8-27B when baseline fidelity is the priority
Use the official base when you want the upstream reference, the simplest Apache-2.0 licensing position, and the least post-training uncertainty. It is also the right control model for any evaluation. If latency is acceptable, the operational benefit of changing may not justify a new model, license, and validation cycle.
Choose ThinkingCap when auditability and balanced preservation matter
ThinkingCap is attractive when you want substantial token savings backed by a detailed multi-benchmark methodology and close macro accuracy. Its independent coding result also matched the base retry-pass rate in the reported runs. Check the license carefully before commercial deployment, especially organizational revenue and use-case eligibility.
Choose Swift when reducing long reasoning traces is the dominant goal
Swift deserves evaluation when median and tail latency are the main problem. Its published GPQA median reduction is large, and the independent Aider comparison showed slightly fewer median tokens and seconds per case than ThinkingCap. Again, the Swift Open License is not identical to Apache-2.0 and includes a revenue threshold for free commercial use.
Licensing can decide the result before benchmarks do
Qwen 3.8-27B is listed under Apache-2.0 on Hugging Face.
ThinkingCap is listed under PolyForm Small Business 1.0.0 plus a BottleCap personal-use grant. That is a different permission model from the base.
Swift states that the underlying Qwen model remains Apache-2.0, while UkisAI's fine-tuned contribution uses the Swift Open License v1.0. The model card says personal, research, educational, evaluation, and commercial use are free for individuals and organizations with gross annual revenue, including affiliates, up to US$1 million; above that threshold, a separate enterprise license is required.
This article is not legal advice. Read the current license files yourself and record the exact model revision used. A benchmark advantage is irrelevant if the license does not fit the intended deployment.
A fair evaluation plan for your own workload
Use a paired test rather than asking which model “feels better.”
- Select 50 to 200 representative prompts from real work. Remove secrets and personal data.
- Use the same quantization class, serving engine, context, sampling settings, reasoning effort, and output cap.
- Run multiple seeds for stochastic tasks.
- Score task success before measuring speed. A shorter wrong answer is not an optimization.
- Measure input tokens, thinking tokens, answer tokens, time to first token, total latency, tokens per second, truncations, loops, tool calls, and peak memory.
- Separate easy prompts from hard prompts. Efficiency changes can behave differently by difficulty.
- Review long-tail failures, not only medians. A model that is usually short but occasionally loops may still hurt interactive use.
- Repeat the test after changing quantization or runtime.
For coding agents, include repository edits, test execution, malformed diffs, retry behavior, and success after feedback. For RAG, include citation correctness and retrieval-grounded answers. For multimodal work, test your real image and document types rather than assuming text benchmarks transfer.
Hardware and quantization considerations
A 27B model can run locally in several quantized formats, but “runs” and “runs well for an agent” are different standards. The base model card documents a long native context and large potential outputs. KV cache, vision inputs, tool transcripts, and long reasoning traces add memory pressure beyond the weight file.
Shorter reasoning can improve throughput and user-perceived latency even when decode speed in tokens per second is unchanged. It can also reduce the chance of hitting an output cap. However, model-loading time, prompt prefill, speculative decoding, flash-attention kernels, and the specific quantizer may dominate some workloads.
Compare complete serving configurations, not just model names. Record engine and version, quantization, GPU or Apple Silicon model, memory, context length, batch size, speculative settings, and whether vision is enabled.
Limitations and scope
The vendor tables are not directly comparable because the teams used different benchmark settings and, in some cases, different base scores. The independent Aider comparison improves comparability across the three candidates but covers one coding-oriented harness with two runs and one quantization.
Published accuracy is not a guarantee for your prompts. Fine-tunes can improve one behavior while subtly changing refusal, tool use, multilingual performance, formatting, or calibration. Free hosted research endpoints are not guaranteed services. Model cards and licenses can also change.
The safe conclusion is therefore conditional: ThinkingCap and Swift are credible ways to trade a small, task-dependent amount of accuracy for materially shorter reasoning. Qwen remains the reference and the simplest licensed baseline.
Final verdict
There is no universal winner.
For the cleanest baseline and Apache-2.0 licensing, start with Qwen 3.8-27B and test lower reasoning effort.
For a well-documented efficiency fine-tune with close macro accuracy, test ThinkingCap.
For aggressive shortening and slightly better median efficiency in the independent coding comparison, test Swift.
Do not select from the vendor headline alone. Run all three under the same serving stack and accept a derivative only if it preserves the success rate that matters to your application.
Sources and evidence
This comparison uses the official Qwen 3.8-27B model card; the ThinkingCap model card and BottleCap AI methodology post; the Swift model card; and an independent community Aider comparison mirrored by Prismix and linked to its original LocalLLaMA post. Simon Willison's hands-on review provides additional context on why overthinking became a practical concern. Figures are attributed to their original evaluation and are not recomputed as if they belonged to one shared leaderboard.
For related local-AI resources, see Free AI APIs, free AI tiers, free GPU credits, and the FreeAI Tokens editorial hub.