Moved → Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF

The APEX quants of Kwaipilot/KAT-Coder-V2.5-Dev now live in a single repo:

Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF

file size
KAT-Coder-V2.5-Dev-MTP-APEX-i-quality.gguf 20.72 GB recommended — includes a working MTP head for speculative decoding
KAT-Coder-V2.5-Dev-APEX-dynamic.gguf 11.86 GiB sized for a 16 GB card
kat-coder.imatrix 192 MB the importance matrix, reusable for your own tiers
model-00014-of-mtp.safetensors 1.69 GB the bf16 MTP head, if you want to redo the transplant

KAT-Coder ships mtp_num_hidden_layers: 0 — no MTP head at all, so no speculative decoding is possible out of the box, and that is true of the vendor release and of every other quant of this model I am aware of. The build above transplants Qwen3.6-35B-A3B's own trained head onto it: 2.03× on a hard agentic-coding suite with correctness unchanged (100% on both suites either way — the head only drafts, the main model verifies).

This repo previously held only the imatrix, while its README described quants that were never uploaded here. Everything is now in the repo linked above.

Downloads last month
412
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Myric/KAT-Coder-V2.5-Dev-APEX-GGUF

Quantized
(62)
this model
MiniMax H3 Video Generator 20 free credits · Text & image to video Try Free →