ProCreations/grug-27b | INT4 (W4A16)
#42
by INC4AI - opened
Pipeline Failure Report
Model: ProCreations/grug-27b
Quantization Scheme: INT4 (W4A16)
Failed Phase: quantize
Run ID: grug-27b-AutoRound-W4A16-RTN
Error Category: out_of_memory
Full Error Log
10:45:06 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect"
10:45:06 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/ProCreations/grug-27b/03637fc89a53196be770ae6b751ba52aaac46621/tokenizer_config.json "HTTP/1.1 200 OK"
10:45:06 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/ProCreations/grug-27b/03637fc89a53196be770ae6b751ba52aaac46621/tokenizer_config.json "HTTP/1.1 200 OK"
10:45:06 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect"
10:45:06 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/ProCreations/grug-27b/03637fc89a53196be770ae6b751ba52aaac46621/tokenizer_config.json "HTTP/1.1 200 OK"
10:45:06 [INFO] HTTP Request: GET https://huggingface.co/api/models/ProCreations/grug-27b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found"
10:45:07 [INFO] HTTP Request: GET https://huggingface.co/api/models/ProCreations/grug-27b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK"
10:45:07 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/vocab.json "HTTP/1.1 404 Not Found"
10:45:07 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/merges.txt "HTTP/1.1 404 Not Found"
10:45:07 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/tokenizer.json "HTTP/1.1 302 Found"
10:45:08 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/added_tokens.json "HTTP/1.1 404 Not Found"
10:45:09 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/special_tokens_map.json "HTTP/1.1 404 Not Found"
10:45:09 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/chat_template.jinja "HTTP/1.1 307 Temporary Redirect"
10:45:09 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/ProCreations/grug-27b/045d620eea94c9b0993f4df1444653ad30b608d0/chat_template.jinja "HTTP/1.1 200 OK"
10:45:09 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/ProCreations/grug-27b/045d620eea94c9b0993f4df1444653ad30b608d0/chat_template.jinja "HTTP/1.1 200 OK"
10:45:11 [INFO] HTTP Request: GET https://huggingface.co/api/models/ProCreations/grug-27b "HTTP/1.1 200 OK"
10:45:11 [INFO] Loading model...
10:45:11 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
10:45:12 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/ProCreations/grug-27b/045d620eea94c9b0993f4df1444653ad30b608d0/config.json "HTTP/1.1 200 OK"
10:45:12 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/adapter_config.json "HTTP/1.1 404 Not Found"
10:45:12 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
10:45:12 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/ProCreations/grug-27b/045d620eea94c9b0993f4df1444653ad30b608d0/config.json "HTTP/1.1 200 OK"
10:45:12 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/model.safetensors "HTTP/1.1 404 Not Found"
10:45:12 [INFO] HTTP Request: HEAD https://huggingface.co/ProCreations/grug-27b/resolve/main/model.safetensors.index.json "HTTP/1.1 307 Temporary Redirect"
10:45:13 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/ProCreations/grug-27b/045d620eea94c9b0993f4df1444653ad30b608d0/model.safetensors.index.json "HTTP/1.1 200 OK"
10:45:13 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/ProCreations/grug-27b/045d620eea94c9b0993f4df1444653ad30b608d0/model.safetensors.index.json "HTTP/1.1 200 OK"
10:45:13 [INFO] HTTP Request: GET https://huggingface.co/api/models/ProCreations/grug-27b/revision/main "HTTP/1.1 200 OK"
10:45:13 [INFO] HTTP Request: GET https://huggingface.co/api/models/ProCreations/grug-27b/tree/045d620eea94c9b0993f4df1444653ad30b608d0?recursive=true&expand=false "HTTP/1.1 200 OK"
[33;1m2026-07-23 10:52:07 WARNING modeling_qwen3_5.py L427: The fast path is not available because one of the required library is not installed. Falling back to torch implementation. To install follow https://github.com/fla-org/flash-linear-attention#installation and https://github.com/Dao-AILab/causal-conv1d[0m
10:52:11 [ERROR] Quantization failed: CUDA out of memory. Tried to allocate 170.00 MiB. GPU 0 has a total capacity of 31.37 GiB of which 72.25 MiB is free. Including non-PyTorch memory, this process has 31.29 GiB memory in use. Of the allocated memory 30.80 GiB is allocated by PyTorch, and 2.80 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf)
Traceback (most recent call last):
File "/root/_work/1/s/auto_quant/phases/quantize.py", line 479, in <module>
quantize(args)
File "/root/_work/1/s/auto_quant/phases/quantize.py", line 293, in quantize
model = AutoModelForCausalLM.from_pretrained(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/utils/common.py", line 140, in patched
return underlying_func(klass, *args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/models/auto/auto_factory.py", line 402, in from_pretrained
return model_class.from_pretrained(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/modeling_utils.py", line 4456, in from_pretrained
loading_info, disk_offload_index = cls._load_pretrained_model(model, state_dict, checkpoint_files, load_config)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/modeling_utils.py", line 4590, in _load_pretrained_model
loading_info, disk_offload_index = convert_and_load_state_dict_in_model(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/core_model_loading.py", line 1695, in convert_and_load_state_dict_in_model
realized_value = mapping.convert(
^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/core_model_loading.py", line 990, in convert
collected_tensors = self.materialize_tensors()
^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/core_model_loading.py", line 952, in materialize_tensors
tensors = [future.result() for future in tensors if future.result() is not None]
^^^^^^^^^^^^^^^
File "/root/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 456, in result
return self.__get_result()
^^^^^^^^^^^^^^^^^^^
File "/root/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result
raise self._exception
File "/root/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/concurrent/futures/thread.py", line 59, in run
result = self.fn(*self.args, **self.kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/core_model_loading.py", line 1239, in _job
return _materialize_copy(tensor, device, dtype)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/core_model_loading.py", line 1217, in _materialize_copy
tensor = tensor.to(device=device, dtype=dtype)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 170.00 MiB. GPU 0 has a total capacity of 31.37 GiB of which 72.25 MiB is free. Including non-PyTorch memory, this process has 31.29 GiB memory in use. Of the allocated memory 30.80 GiB is allocated by PyTorch, and 2.80 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf)
Auto-generated by error_analysis pipeline. cc @lvkaokao