
Swift 1.5 Cuts Qwen3.8-27B Thinking, With Trade-offs
Published by AINave Editorial
UkisAI’s Swift-1.5-Qwen3.8-27B-GGUF is a local, llama.cpp-compatible version of a Qwen3.8-27B derivative built to use fewer reasoning tokens. The trade-off in its reported tests is concrete: it scores higher on coding and agent benchmarks, but lower on several math and reasoning benchmarks.
Fewer thinking tokens, not a universal quality win
On GPQA-Diamond, Swift 1.5 averaged 8,717 thinking tokens against 15,014 for the base model, while scoring 88.59% versus 88.28%. The reported evaluation also found fewer tokens on coding and agent tasks. UkisAI describes its training approach as penalizing tokens associated with pathological overthinking, then using reinforcement learning and outcome-based preference optimization to recover accuracy. The release reports a 9.18× speed-up on “several tasks,” but provides no general throughput figure or hardware details, so that number should not be treated as an expected local-inference speed-up. The reported evaluation gives the token counts and scores.
The gains are clearest in the tests aimed at coding and tool use. Swift 1.5 scored 81.71% on LiveCodeBench v6, compared with 76.76% for Qwen3.8-27B, using 8,448 rather than 11,184 mean reasoning tokens. On Terminal-Bench 2.1, it scored 72.13% versus 69.21%; mean tokens across agent calls fell from 52,265 to 43,733. The main evaluation used five repeats, and Terminal-Bench used five trials per task. Those conditions make the comparison more useful than an isolated score, but neither benchmark settles performance on a particular codebase or tool setup. The benchmark results and evaluation conditions are task-specific evidence, not a general guarantee.
The benchmark gains come with a trade-off
Swift 1.5 scored below the base model on IFBench, ERQA, AIME 2026 and HMMT November 2025. For example, its AIME 2026 score was 96.00%, compared with 98.67% for Qwen3.8-27B. So lower token use does not mean the same level of performance across every reasoning task. The reported results show both the coding and agent gains and these regressions.
That pattern makes Swift a workload-specific alternative rather than a drop-in upgrade. If an application is dominated by coding or agent tasks, the reported results are encouraging; if it depends on the weaker benchmarks, the base model’s scores matter more. The evidence supports that distinction, not a broad claim that Swift is better or worse overall.
GGUF size is not a memory requirement
The available GGUF files range from 29.0 GB for Q80 to 8.9 GB for IQ2XXS. These are file sizes, not RAM or VRAM requirements. The release’s quantization measurements compare each file’s output distribution with the Swift 1.5 BF16 source; they are not direct task-accuracy scores. At 32k, reported KLD rises from 0.0006 for Q80 to 0.2769 for IQ2XXS, illustrating how much the smallest tier diverges on that measure. The GGUF sizes and divergence figures do not establish the quality of every quantized model on a real workload.
The release recommends a current llama.cpp-compatible runtime such as llama-server, but the supplied material gives no loading command or memory guidance. It also does not state a maximum context length for this GGUF package; context sizes reported for evaluations are settings used in those tests, not a package limit. For local deployment, the practical choice is therefore two-part: whether Swift’s benchmark profile fits the task, and whether a particular quantization works within the machine’s memory and quality constraints. The runtime recommendation and evaluation details leave those deployment requirements open.





















