llama.cpp supports?
llama.cpp + GGUF supports?
Commenting so I get things in my notifications since huggingface doesn't have way to do that without commenting
+1
+1 For llama.cpp. Thank you for gifting us the model.
Did you ever manage to port support into llama.cpp mainline before? For Ling-2.6-flash took me some trying to get it running
https://huggingface.co/ljupco/Ling-2.6-flash-GGUF
On a llama.cpp branch here
https://github.com/ljubomirj/llama.cpp/tree/LJ-Ling-2.6-flash-r2
Being 'AI-assisted' it's against the llama.cpp policy, so never attempted a PR. Maybe you are knowledgeable enough to do a real PR for Ling-3-Flash for llama.cpp this time around? Excellent sizing, exactly what I needed in between the 30B-A3B class of models and the 280B-A13B like deepsseek, mimo etc. Post training extending context to 1M would be a blast too. But we can do with 256K context at A5M. Thank you.
+1
I'll give it a go soonish and report back if it runs or my computer gets infected with the curse of ra
I tried it but I seem to be having problems in llama.cpp,
EDIT: I didn't read the thing saying it needs the turboquant build lol, I just tried a PR as written below. I'll try it with turboquant maybe. Anyway, here's what I did before I read that:
I tried it with the pr at https://github.com/ggml-org/llama.cpp/pull/26608 and had to have qwen make a bazillion changes to get it to run with it lol.
it worked but it throws special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect and it randomly has spaces in its output
but it ran pretty fast actually on my 2x3090s + 48GB RAM, about 285-327pp, 50tg, 120k ctx, probably could push it more idk.
Anyway, I tried using the chat template file at ling's repo but the error is still there.
EDIT:
Omg that took forever but I finally managed to get it to work in turboquant. I will test and let everyone know :)
EDIT 2:
OK so it's still kind of having problems tokenizing unfortunately, but I can't tell if that's my low quant (IQ3_XXS) or if it's something else I'm doing wrong. It's not a bad model, but it's hard to use because it keeps messing up tool calls due to the token issues - but it knows how to do it! If the tool calls were fixed, it would be at least on par with 27B, but I can't tell if it's worse or better just yet.
EDIT 3: There's a solution! Check https://huggingface.co/AtomicChat/Ling-3.0-flash-GGUF/discussions/1#6a74ff181717f81592cdecce and https://github.com/ggml-org/llama.cpp/commit/0266ebca66bd95b7a85d37b8ca08ccf9812b85cc
I just replaced the files in my llama cpp with those in the turboquant repo and it works without any tool call issues. It's similar to qwen 27B so far, but I need to do more testing. I suspect it has much better world knowledge due to having way more params, but I don't know for sure. I need some good tests haha.
I've also had issues with tool calling that involved multiple arguments - fixed it by modifying jinja2 template to use JSON when calling tools.
Modified Jinja2 Template:
{#- Bailing V3 chat template -#}
{#- Supports: thinking option, tool calling -#}
{#- ==================== thinking option normalization ==================== -#}
{%- if enable_thinking is defined %}
{%- if enable_thinking %}
{%- set thinking_option = 'on' %}
{%- else %}
{%- set thinking_option = 'off' %}
{%- endif %}
{%- elif thinking_option is not defined %}
{%- set thinking_option = 'on' %}
{%- endif %}
{#- ==================== preserved thinking ==================== -#}
{% set preserved_thinking = true %}
{#- ==================== system message ==================== -#}
{{- 'SYSTEM' }}
{%- if tools %}
{%- if messages[0].role == 'system' %}
{{- messages[0].content + '\n' }}
{%- endif %}
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within XML tags:\n" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n\n\nIf none of the functions can be used, point it out. If the given question lacks the parameters required by the function, also point it out.\nIf you need to use a function, for each function call, output the function name and arguments within the following JSON format:\n{function-name}\n{\n "arg1": "value1",\n "arg2": "value2"\n}\n\nDO NOT use or tags. Those are invalid and will cause errors.\n\nIMPORTANT: and are structural markers ONLY. They must NEVER appear as regular text in your response. Use them exclusively as an opening/closing pair around your internal reasoning, then output your actual response after .\n" }}
{%- if messages[0].role == 'system' and messages[0].content is string and ('detailed thinking on' in messages[0].content or 'detailed thinking off' in messages[0].content) %}
{{- '<|role_end|>' }}
{%- else %}
{{- 'detailed thinking ' + thinking_option + '<|role_end|>' }}
{%- endif %}
{%- else %}
{%- if messages[0].role == 'system' %}
{%- if 'detailed thinking on' in messages[0].content or 'detailed thinking off' in messages[0].content %}
{{- messages[0].content + '<|role_end|>' }}
{%- else %}
{{- messages[0].content + '\n' }}
{{- 'detailed thinking ' + thinking_option + '<|role_end|>' }}
{%- endif %}
{% else %}
{{- 'detailed thinking ' + thinking_option + '<|role_end|>' }}
{%- endif %}
{%- endif %}
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
{%- for message in messages[::-1] %}
{%- set index = (messages|length - 1) - loop.index0 %}
{%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('') and message.content.endswith('')) %}
{%- set ns.multi_step_tool = false %}
{%- set ns.last_query_index = index %}
{%- endif %}
{%- endfor %}
{%- for message in messages %}
{%- if message.content is string %}
{%- set content = message.content %}
{%- else %}
{%- set content = '' %}
{%- endif %}
{%- if message.role == "user" %}
{{- 'HUMAN' + message.content + '<|role_end|>' }}
{%- elif message.role == "system" and not loop.first %}
{{- 'SYSTEM' + message.content + '<|role_end|>' }}
{%- elif message.role == "assistant" %}
{%- set reasoning_content = '' %}
{%- if message.reasoning_content is string and message.reasoning_content != '' %}
{%- set reasoning_content = message.reasoning_content %}
{%- else %}
{%- if '' in content %}
{%- set _parts = content.split('') %}
{%- set reasoning_content = _parts[0].rstrip('\n').split('')[-1].lstrip('\n') %}
{# Rejoin everything after the first as content (handles leaked tags gracefully) #}
{%- set content = '\n'.join(_parts[1:]) %}
{%- endif %}
{%- endif %}
{%- if preserved_thinking or loop.index0 > ns.last_query_index %}
{%- if reasoning_content != '' %}
{{- 'ASSISTANT' + '\n' + reasoning_content.strip('\n') + '' + content.lstrip('\n') }}
{%- else %}
{{- 'ASSISTANT\n' + content }}
{%- endif %}
{%- else %}
{{- 'ASSISTANT\n' + content }}
{%- endif %}
{%- if message.tool_calls %}
{%- for tool_call in message.tool_calls %}
{%- if (loop.first and content) or (not loop.first) %}
{{- '\n' }}
{%- endif %}
{%- set tc = tool_call %}
{%- if tool_call.function %}
{%- set tc = tool_call.function %}
{%- endif %}
{{- '' + tc.name + '\n{\n' }}
{%- for k, v in tc.arguments.items() %}
{{- ' "' + k + '": ' }}
{{- v | tojson(ensure_ascii=False) }}
{%- if not loop.last %}
{{- ',' }}
{%- endif %}
{{- '\n' }}
{%- endfor %}
{{- '}\n' }}
{%- endfor %}
{%- endif %}
{{- '<|role_end|>' }}
{%- elif message.role == "tool" %}
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
{{- 'OBSERVATION' }}
{%- endif %}
{{- '\n\n' }}
{{- content }}
{{- '\n' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|role_end|>' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{#- ==================== generation prompt ==================== -#}
{%- if add_generation_prompt %}
{{- 'ASSISTANT' }}
{%- if thinking_option == 'on' %}
{{- '\n' }}
{%- elif thinking_option == 'off' %}
{{- '\n' }}
{%- endif %}
{%- endif %}
I've also had issues with tool calling that involved multiple arguments - fixed it by modifying jinja2 template to use JSON when calling tools.
Modified Jinja2 Template:{#- Bailing V3 chat template -#}
{#- Supports: thinking option, tool calling -#}{#- ==================== thinking option normalization ==================== -#}
{%- if enable_thinking is defined %}
{%- if enable_thinking %}
{%- set thinking_option = 'on' %}
{%- else %}
{%- set thinking_option = 'off' %}
{%- endif %}
{%- elif thinking_option is not defined %}
{%- set thinking_option = 'on' %}
{%- endif %}{#- ==================== preserved thinking ==================== -#}
{% set preserved_thinking = true %}{#- ==================== system message ==================== -#}
{{- 'SYSTEM' }}
{%- if tools %}
{%- if messages[0].role == 'system' %}
{{- messages[0].content + '\n' }}
{%- endif %}
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within XML tags:\n" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n\n\nIf none of the functions can be used, point it out. If the given question lacks the parameters required by the function, also point it out.\nIf you need to use a function, for each function call, output the function name and arguments within the following JSON format:\n{function-name}\n{\n "arg1": "value1",\n "arg2": "value2"\n}\n\nDO NOT use or tags. Those are invalid and will cause errors.\n\nIMPORTANT: and are structural markers ONLY. They must NEVER appear as regular text in your response. Use them exclusively as an opening/closing pair around your internal reasoning, then output your actual response after .\n" }}
{%- if messages[0].role == 'system' and messages[0].content is string and ('detailed thinking on' in messages[0].content or 'detailed thinking off' in messages[0].content) %}
{{- '<|role_end|>' }}
{%- else %}
{{- 'detailed thinking ' + thinking_option + '<|role_end|>' }}
{%- endif %}
{%- else %}
{%- if messages[0].role == 'system' %}
{%- if 'detailed thinking on' in messages[0].content or 'detailed thinking off' in messages[0].content %}
{{- messages[0].content + '<|role_end|>' }}
{%- else %}
{{- messages[0].content + '\n' }}
{{- 'detailed thinking ' + thinking_option + '<|role_end|>' }}
{%- endif %}
{% else %}
{{- 'detailed thinking ' + thinking_option + '<|role_end|>' }}
{%- endif %}
{%- endif %}
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
{%- for message in messages[::-1] %}
{%- set index = (messages|length - 1) - loop.index0 %}
{%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('') and message.content.endswith('')) %}
{%- set ns.multi_step_tool = false %}
{%- set ns.last_query_index = index %}
{%- endif %}
{%- endfor %}
{%- for message in messages %}
{%- if message.content is string %}
{%- set content = message.content %}
{%- else %}
{%- set content = '' %}
{%- endif %}
{%- if message.role == "user" %}
{{- 'HUMAN' + message.content + '<|role_end|>' }}
{%- elif message.role == "system" and not loop.first %}
{{- 'SYSTEM' + message.content + '<|role_end|>' }}
{%- elif message.role == "assistant" %}
{%- set reasoning_content = '' %}
{%- if message.reasoning_content is string and message.reasoning_content != '' %}
{%- set reasoning_content = message.reasoning_content %}
{%- else %}
{%- if '' in content %}
{%- set _parts = content.split('') %}
{%- set reasoning_content = _parts[0].rstrip('\n').split('')[-1].lstrip('\n') %}
{# Rejoin everything after the first as content (handles leaked tags gracefully) #}
{%- set content = '\n'.join(_parts[1:]) %}
{%- endif %}
{%- endif %}
{%- if preserved_thinking or loop.index0 > ns.last_query_index %}
{%- if reasoning_content != '' %}
{{- 'ASSISTANT' + '\n' + reasoning_content.strip('\n') + '' + content.lstrip('\n') }}
{%- else %}
{{- 'ASSISTANT\n' + content }}
{%- endif %}
{%- else %}
{{- 'ASSISTANT\n' + content }}
{%- endif %}
{%- if message.tool_calls %}
{%- for tool_call in message.tool_calls %}
{%- if (loop.first and content) or (not loop.first) %}
{{- '\n' }}
{%- endif %}
{%- set tc = tool_call %}
{%- if tool_call.function %}
{%- set tc = tool_call.function %}
{%- endif %}
{{- '' + tc.name + '\n{\n' }}
{%- for k, v in tc.arguments.items() %}
{{- ' "' + k + '": ' }}
{{- v | tojson(ensure_ascii=False) }}
{%- if not loop.last %}
{{- ',' }}
{%- endif %}
{{- '\n' }}
{%- endfor %}
{{- '}\n' }}
{%- endfor %}
{%- endif %}
{{- '<|role_end|>' }}
{%- elif message.role == "tool" %}
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
{{- 'OBSERVATION' }}
{%- endif %}
{{- '\n\n' }}
{{- content }}
{{- '\n' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|role_end|>' }}
{%- endif %}
{%- endif %}
{%- endfor %}{#- ==================== generation prompt ==================== -#}
{%- if add_generation_prompt %}
{{- 'ASSISTANT' }}
{%- if thinking_option == 'on' %}
{{- '\n' }}
{%- elif thinking_option == 'off' %}
{{- '\n' }}
{%- endif %}
{%- endif %}
Thank you so much for sharing