Qwen3.8 Flash Next GSQ RCO Abliterated GGUF
Qwen3.8 Flash Next GSQ RCO Abliterated GGUF by SC117, a image-text-to-text model with multimodal capabilities. Understand and compare multimodal features, benchmarks, and capabilities.
Comparison
| Feature | Qwen3.8 Flash Next GSQ RCO Abliterated GGUF | Interfaze |
|---|---|---|
| Input Modalities | text, image, video | image, text, audio, video, document |
| Native OCR | No | Yes |
| Long Document Processing | No | Yes |
| Language Support | unknown | 162+ |
| Native Speech-to-Text | No | Yes |
| Native Object Detection | No | Yes |
| Guardrail Controls | Yes | Yes |
| Context Input Size | 262.1K | 1M |
| Tool Calling | Yes | Tool calling supported + built in browser, code execution and web search |
Scaling
| Feature | Qwen3.8 Flash Next GSQ RCO Abliterated GGUF | Interfaze |
|---|---|---|
| Scaling | Self-hosted/Provider-hosted with quantization | Unlimited |
View model card on Hugging Face
llama-mtmd-cli -m IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00001-of-00002.gguf
--mmproj mmproj-Qwen3.8-Flash-Next-BF16.gguf -lm mmap --lazy-mode on
--image photo.jpg -p "Describe this image."Speculative decoding with the MTP draft headllama-cli -m IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00001-of-00002.gguf
-md mtp-Qwen3.8-Flash-Next-Q8_0.gguf --draft-max 4 -lm mmap --lazy-mode on -ngl 99-lm mmap --lazy-mode on keeps the n-gram table memory-mapped on disk: it is 28.8 GB and one row is read per token. Shard 1 wants to be resident (VRAM or RAM); shard 2 is fine on an SSD. The vision projector is a standard clip GGUF shared with the upstream tiers, 0.91 GB. Native context is 262,144 tokens and the KV cache grows with it, so start at -c 32768 and work up.