NLLB-200 Distilled 600M — FP16 Half Precision

🌐 Live Web Demo: Try 100% offline in your browser at huggingface.co/spaces/rudrakshrakeshzodage/nllb-200-onnx-browser-demo

This model is part of a suite of optimized/quantized versions of the base model. Other variants in this direction:

Base model: facebook/nllb-200-distilled-600M
Precision: FP16 (16-bit Float Half Precision)
Languages: 200 (all NLLB-200 supported languages)


🏗️ Detailed Tokenizer & Model Architecture Pipeline

💾 Model Size Reduction Comparison (FP32 vs FP16 vs INT8 vs NF4)

Model Size Comparison

📈 200+ Language Performance Benchmark Chart

200+ Language Performance Benchmark


📜 Citation & Credits

Quantization research, engineering, and benchmarking developed by Rudraksh Rakesh Zodage:


Downloads last month
79
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rudrakshrakeshzodage/nllb-200-distilled-600M-fp16

Quantized
(28)
this model

Space using rudrakshrakeshzodage/nllb-200-distilled-600M-fp16 1

MiniMax H3 Video Generator 20 free credits · Text & image to video Try Free →