Instructions to use bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4
- SGLang
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4 with Docker Model Runner:
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4
Request access to the BottleCap AI model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Please fill out the form below to request access. Name, Company Name, and Company Email are optional — just enter NA if you'd prefer not to share. Your email may be used to send you information about BottleCap AI model updates and early access before public release.
Log in or Sign Up to review the conditions and access this model content.
bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4
FP4 (NVFP4) quantization of bottlecapai/ThinkingCap-Qwen3.6-27B — capability of Qwen3.6-27B with 50% less thinking tokens on average, achieved by finetuning Qwen3.6-27B (Qwen Team, 2026) while preserving the original answer quality and style.
➡️ Full model description, evaluation results (multi-seed, statistically tested), recommended sampling params, and citation: see the main model card at bottlecapai/ThinkingCap-Qwen3.6-27B.
About this quantization
Weights are stored at 4 bits in NVIDIA's NVFP4 format (E2M1 values with per-group-of-16 FP8 scales) while activations stay in bf16 — weight-only NVFP4A16 (W4A16), produced data-free with llm-compressor in the compressed-tensors nvfp4-pack-quantized format that vLLM loads natively. ≈19 GB instead of ≈55 GB bf16 — ~2.9× smaller. Being weight-only (W4A16), it serves through vLLM's MarlinNvFp4LinearKernel on both Hopper (H100/H200) and Blackwell: the 4-bit weights unpack to bf16 for compute, so the decode win is memory bandwidth (≈2× bf16 tok/s), not FP4 tensor cores (native cutlass FP4 compute would require a W4A4 / FP4-activation checkpoint). The model's MTP (multi-token-prediction) head is preserved, so self-speculative decoding works (≈3.2 accepted tokens/step).
Kept in bf16: lm_head, the MTP head, the vision tower, and the two small Gated-DeltaNet input gates linear_attn.in_proj_a / in_proj_b (48-wide output, not a multiple of the 4-bit kernel's tile size and fused to width 96 at serve time, so quantizing them breaks loading on non-Blackwell GPUs; keeping them bf16 costs ~60 MB). The remaining group-quantizable Linear layers — MLP, self-attention, and the other linear-attention projections (in_proj_qkv / in_proj_z / out_proj) — carry the 4-bit weights.
Usage
vllm serve bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4
Speculative decoding (MTP)
The MTP (multi-token-prediction / NextN) head is kept in bf16, so vLLM can run self-speculative decoding for a decode speed-up — no separate draft model needed. Add --speculative-config when serving (requires vLLM ≥ 0.24.0):
vllm serve bottlecapai/ThinkingCap-Qwen3.6-27B-NVFP4 --speculative-config '{"method":"mtp","num_speculative_tokens":3}'
Speculative decoding is lossless — the output is identical to standard decoding. It accepts ≈3.2 drafts per verify step here, for a further ≈1.5× on top of the finetune's token savings (see below).
Served on vLLM; SGLang cannot currently load weight-only NVFP4 for this architecture. For local llama.cpp / Ollama / LM Studio use, see the GGUF quantizations at bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF; for near-lossless FP8 at half the memory (loads in vLLM and SGLang), see bottlecapai/ThinkingCap-Qwen3.6-27B-FP8.
Expected performance
Measured on our internal serving harness on 8× RTX PRO 6000 (Blackwell) with vLLM 0.25.0, one configuration per GPU. Each cell is N=200 questions/dataset × 3 generation seeds over a fixed problem subset, batch size 16, sampled decoding (temperature 1.0, top_p 0.95, top_k 20); acc is the mean ± 95% CI across the seeds. Generation budgets are 24,576 tokens (MMLU-Pro) and 16,384 tokens (RealWorldQA).
FP4 matches the bf16 finetune's accuracy (all 95% CIs overlap) while decoding ≈2.3× faster (52.7 vs 22.6 tok/s, MMLU-Pro standard — the memory-bandwidth win of 4-bit weights via the Marlin kernel), and MTP self-speculative decoding (≈3.1–3.2 accepted tokens per verify step) adds a further ≈1.5× — stacking with the finetune's token savings to ≈4.7–8.4× faster per task than the unquantized base (MMLU-Pro 11.6 s vs 97.3 s; RealWorldQA 7.3 s vs 34.2 s).
All timings are measured per request during the eval. median tokens = median completion length (the finetune's lever); tok/s = per-request steady-state decode rate (prefill excluded via time-to-first-token); task s = measured end-to-end wall-clock per request under the batch-size-16 concurrency; speedup = task s of the base model in standard decoding ÷ task s of the row. unsloth NVFP4 and NVIDIA NVFP4 are community NVFP4 quants of the base model, shown for comparison; ThinkingCap-FP8 is our FP8 sibling (bottlecapai/ThinkingCap-Qwen3.6-27B-FP8).
NVFP4 kernel (important for reading tok/s). On this hardware the vLLM NVFP4 kernel is chosen by the checkpoint's FP4 scheme, not a flag, so the NVFP4 rows are not on identical footing: this model and NVIDIA NVFP4 are weight-only W4A16 → MarlinNvFp4LinearKernel (4-bit weights unpacked for compute; a memory-bandwidth decode win), whereas unsloth NVFP4 is W4A4 (weights and activations FP4, mixed-precision) → FlashInferCutlassNvFp4LinearKernel, the cutlass/flashinfer FP4 GEMM. vLLM has no cutlass path for weight-only W4A16 on sm_120, so those fall back to Marlin — a checkpoint-format property, bit-exact-ish and not a source of accuracy difference. On Hopper (H100/H200) there is no native FP4 compute at all, so every NVFP4 serves via Marlin weight-unpack; the cutlass FP4 tensor-core advantage (W4A4) shows mostly on compute-bound prefill / high batch and on datacenter Blackwell.
MMLU-Pro (reasoning)
| config | acc | median tokens | tok/s | task s | speedup | accept_len |
|---|---|---|---|---|---|---|
| Qwen3.6-27B base bf16 · standard | 0.902 ± 0.019 | 2186 | 22.6 | 97.3 | 1.00× | — |
| Qwen3.6-27B base bf16 · MTP | 0.892 ± 0.031 | 2117 | 47.7 | 45.6 | 2.13× | 3.23 |
| unsloth NVFP4 (base) · standard | 0.878 ± 0.038 | 2186 | 44.5 | 49.1 | 1.98× | — |
| unsloth NVFP4 (base) · MTP | 0.888 ± 0.047 | 2123 | 87.4 | 25.4 | 3.83× | 3.21 |
| NVIDIA NVFP4 (base) · standard | 0.892 ± 0.026 | 2032 | 51.0 | 40.0 | 2.43× | — |
| NVIDIA NVFP4 (base) · MTP | 0.907 ± 0.036 | 1996 | 93.7 | 21.6 | 4.50× | 3.23 |
| ThinkingCap bf16 · standard | 0.890 ± 0.022 | 963 | 23.4 | 41.6 | 2.34× | — |
| ThinkingCap bf16 · MTP | 0.895 ± 0.012 | 910 | 51.1 | 18.8 | 5.18× | 3.26 |
| ThinkingCap-NVFP4 · standard | 0.882 ± 0.031 | 926 | 52.7 | 17.7 | 5.50× | — |
| ThinkingCap-NVFP4 · MTP | 0.885 ± 0.033 | 878 | 81.3 | 11.6 | 8.39× | 3.24 |
RealWorldQA (vision)
| config | acc | median tokens | tok/s | task s | speedup | accept_len |
|---|---|---|---|---|---|---|
| Qwen3.6-27B base bf16 · standard | 0.802 ± 0.019 | 718 | 21.7 | 34.2 | 1.00× | — |
| Qwen3.6-27B base bf16 · MTP | 0.798 ± 0.064 | 752 | 42.2 | 18.7 | 1.83× | 3.09 |
| unsloth NVFP4 (base) · standard | 0.788 ± 0.019 | 777 | 43.8 | 18.7 | 1.83× | — |
| unsloth NVFP4 (base) · MTP | 0.783 ± 0.040 | 731 | 79.2 | 10.9 | 3.14× | 3.08 |
| NVIDIA NVFP4 (base) · standard | 0.785 ± 0.012 | 732 | 46.6 | 17.3 | 1.98× | — |
| NVIDIA NVFP4 (base) · MTP | 0.790 ± 0.050 | 700 | 79.8 | 10.5 | 3.26× | 3.09 |
| ThinkingCap bf16 · standard | 0.818 ± 0.047 | 344 | 21.5 | 17.4 | 1.97× | — |
| ThinkingCap bf16 · MTP | 0.805 ± 0.045 | 372 | 42.5 | 9.9 | 3.45× | 3.09 |
| ThinkingCap-NVFP4 · standard | 0.802 ± 0.007 | 317 | 40.9 | 9.2 | 3.72× | — |
| ThinkingCap-NVFP4 · MTP | 0.800 ± 0.012 | 320 | 55.5 | 7.3 | 4.68× | 3.09 |
- Downloads last month
- 66,117
