Llama-2-13b-chat-hf - bnb 4bit

Description

This model is 4bit quantized version of Llama-2-13b-chat-hf using bitsandbytes. It's designed for fine-tuning! The PAD token is set as UNK.

Safetensors

Model size

6.87B params

Tensor type

F32

FP16

Inference Providers NEW

This model is not currently available via any of the supported Inference Providers.

The model cannot be deployed to the HF Inference API: The model authors have turned it off explicitly.

Base model

Quantized

(19)

this model