[Bug?] google/gemma-3n-E4B-it fails with AttributeError: 'Gemma3nConfig' object has no attribute 'vocab_size'

#15
by EsatKara - opened

Environment

  • vLLM version: 0.9.1
  • Transformers version: 4.41.2
  • CUDA / GPU: RTX 4080 (driver 550.xx)

Full Traceback

Notes

  • The model works in transformers directly.
  • I suspect the config.json of the model lacks vocab_size, which vLLM expects.
  • Should we patch the model config or wait for vLLM to adapt?
Google org

Hi @EsatKara , Gemma-3n-E4Bmodels requires Transformers version 4.53.0. Please update your Transformers library to at least version 4.53.0 to use Gemma-3n-E4B and let us know if you still face the same issue. Thank you

Hi @EsatKara , Gemma-3n-E4Bmodels requires Transformers version 4.53.0. Please update your Transformers library to at least version 4.53.0 to use Gemma-3n-E4B and let us know if you still face the same issue. Thank you

I upgraded Transformers to 4.53.0 with vllm of 0.9.1, but encountered the same issue

Hi @EsatKara I’m using the same environment as you. I suspect the root cause is that the config.json does not explicitly define top-level keys like vocab_size, num_hidden_layers, num_attention_heads, etc. Instead, these values are nested under the "audio_config" field, like so:

json
"audio_config": {
"vocab_size": 128,
"hidden_size": 1536,
"conf_num_attention_heads": 8,
"conf_num_hidden_layers": 12,
...
}
In my case, I manually copied relevant values (like vocab_size, num_attention_heads, num_hidden_layers, etc.) from audio_config and added them as top-level fields in config.json. After that, VLLM was able to launch successfully.

However, I’m not sure if this workaround would affect model performance or correctness — since this is not an officially supported or standard method. I also suspect this issue could be related to VLLM not fully supporting models with nested configs, or possibly a version mismatch.

@EsatKara vllm 0.9.1 does not support gemma 3n yet, see https://github.com/vllm-project/vllm/pull/20134 the PR only got merged 4 days ago. You need to wait for new release or build from source.

This comment has been hidden (marked as Off-Topic)

HI, @Lqqs
could you please share the config file.the json that you edited

HI @kloblok
Here is my JSON file. I hope it will be helpful to you.
I am currently using transformers version 4.53.0 and vllm version 0.7.4. With the following config.json, I am now able to run vLLM successfully.

{
"architectures": [
"Gemma3nForConditionalGeneration"
],
"hidden_activation": "gelu_pytorch_tanh",
"hidden_size": 2048,
"hidden_size_per_layer_input": 256,
"initializer_range": 0.02,
"intermediate_size": 16384,
"laurel_rank": 64,
"vocab_size": 262400,
"head_dim": 256,
"base_model_tp_plan": {},
"max_position_embeddings": 32768,
"model_type": "gemma3n_text",
"num_attention_heads": 8,
"num_hidden_layers": 35,
"num_key_value_heads": 2,
"num_kv_shared_layers": 15,
"query_pre_attn_scalar": 256,
"rms_norm_eps": 1e-06,
"rope_local_base_freq": 10000.0,
"rope_scaling": null,
"rope_theta": 1000000.0,
"sliding_window": 512,
"torch_dtype": "bfloat16",
"use_cache": true,
"audio_config": {
"conf_attention_chunk_size": 12,
"conf_attention_context_left": 13,
"conf_attention_context_right": 0,
"conf_attention_logit_cap": 50.0,
"conf_conv_kernel_size": 5,
"conf_num_attention_heads": 8,
"conf_num_hidden_layers": 12,
"conf_positional_bias_size": 256,
"conf_reduction_factor": 4,
"conf_residual_weight": 0.5,
"gradient_clipping": 10000000000.0,
"hidden_size": 1536,
"input_feat_size": 128,
"model_type": "gemma3n_audio",
"rms_norm_eps": 1e-06,
"sscp_conv_channel_size": [
128,
32
],
"sscp_conv_eps": 0.001,
"sscp_conv_kernel_size": [
[
3,
3
],
[
3,
3
]
],
"sscp_conv_stride_size": [
[
2,
2
],
[
2,
2
]
],
"torch_dtype": "bfloat16",
"vocab_offset": 262273,
"vocab_size":128
},
"audio_soft_tokens_per_image": 188,
"audio_token_id": 262273,
"boa_token_id": 256000,
"boi_token_id": 255999,
"eoa_token_id": 262272,
"eoi_token_id": 262144,
"eos_token_id": [
1,
106
],
"image_token_id": 262145,
"initializer_range": 0.02,
"model_type": "gemma3n",
"text_config": {
"activation_sparsity_pattern": [
0.95,
0.95,
0.95,
0.95,
0.95,
0.95,
0.95,
0.95,
0.95,
0.95,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0,
0.0
],
"altup_active_idx": 0,
"altup_coef_clip": 120.0,
"altup_correct_scale": true,
"altup_lr_multiplier": 1.0,
"altup_num_inputs": 4,
"attention_bias": false,
"attention_dropout": 0.0,
"final_logit_softcapping": 30.0,
"head_dim": 256,
"hidden_activation": "gelu_pytorch_tanh",
"hidden_size": 2048,
"hidden_size_per_layer_input": 256,
"initializer_range": 0.02,
"intermediate_size": 16384,
"laurel_rank": 64,
"layer_types": [
"sliding_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"full_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"full_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"full_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"full_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"full_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"full_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"sliding_attention",
"full_attention"
],
"max_position_embeddings": 32768,
"model_type": "gemma3n_text",
"num_attention_heads": 8,
"num_hidden_layers": 35,
"num_key_value_heads": 2,
"num_kv_shared_layers": 15,
"query_pre_attn_scalar": 256,
"rms_norm_eps": 1e-06,
"rope_local_base_freq": 10000.0,
"rope_scaling": null,
"rope_theta": 1000000.0,
"sliding_window": 512,
"torch_dtype": "bfloat16",
"use_cache": true,
"vocab_size": 262400,
"vocab_size_per_layer_input": 262144
},
"torch_dtype": "bfloat16",
"transformers_version": "4.53.0.dev0",
"vision_config": {
"architecture": "mobilenetv5_300m_enc",
"do_pooling": true,
"hidden_size": 2048,
"initializer_range": 0.02,
"label_names": [
"LABEL_0",
"LABEL_1"
],
"model_type": "gemma3n_vision",
"num_classes": 2,
"rms_norm_eps": 1e-06,
"torch_dtype": "bfloat16",
"vocab_offset": 262144,
"vocab_size": 128
},
"vision_soft_tokens_per_image": 256
}

Additionally, using fp32 precision during vLLM deployment may lead to runtime errors.DEBUG 07-01 10:59:16 client.py:171] Heartbeat successful. Sampling probabilities contain NaN or Inf values. So far, the model runs more stably with bfloat16 (bf16) precision.

Sign up or log in to comment

MiniMax H3 Video Generator 20 free credits · Text & image to video Try Free →