Qwen3.8-2.4T-A95B-int4

int4 quantization of Qwen/Qwen3.8-2.4T-A95B, produced with compressed-tensors by streaming the checkpoint tensor-by-tensor (the model is never fully instantiated).

4-bit integer weights in groups of 128 with a BF16 scale. Smallest of the set and the highest error.

All quantizations of this model

Variant Format Size vs BF16 Mean rel. error Linears quantized Left BF16
Qwen3.8-2.4T-A95B-FP8 float-quantized 2453.05 GB 50% 0.0264 143569 0
Qwen3.8-2.4T-A95B-NVFP4 nvfp4-pack-quantized 1382.45 GB 28% 0.0952 143569 0
Qwen3.8-2.4T-A95B-int4 ← this one pack-quantized 1268.00 GB 26% 0.1118 143569 0

Mean relative error is ||dequant(W) - W|| / ||W||, averaged over a sample of quantized Linear layers, measured against the original BF16 weights. Lower is better.

This variant

Format pack-quantized
Weight bits 4
Group size 128
Strategy group
Linears quantized 143569
Left in BF16 0
Shards 304
On disk 1268.00 GB
Mean relative error 0.1118
Shape/dtype conformance failures 0

Use with vLLM

vllm serve dudeman2512/Qwen3.8-2.4T-A95B-int4

How this was made

Every produced tensor is checked for shape/dtype conformance against what the server expects, then reconstruction error is measured against the source BF16 weights, before anything is published. The numbers in the table above are those measurements — not estimates.

Downloads last month
51
Safetensors
Model size
2.4T params
Tensor type
BF16
·
I32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dudeman2512/Qwen3.8-2.4T-A95B-int4

Quantized
(30)
this model