Merlin Research

non-profit
Activity Feed

AI & ML interests

Independent AI safety lab. Stockholm, Sweden. We test deployed LLM agents under adversarial conditions and measure behavioral alignment in production β€” not in controlled benchmarks.

Recent Activity

squ11z1Β  updated a model 17 days ago
Merlin-Research/Pluto
squ11z1Β  updated a model 17 days ago
Merlin-Research/HybridIntelligence-0.5B
View all activity

DedeProGamesΒ 
posted an update about 16 hours ago
view post
Post
76
πŸš€ Introducing the GRM-3.2 Family

The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.

GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.

GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.

GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.

All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many stepsβ€”whether on a server, a local workstation, or an edge device.

Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf

Organization:
OrionLLM

DedeProGamesΒ 
posted an update 16 days ago
view post
Post
218
who want GRM-2.6 with 9b?
DedeProGamesΒ 
posted an update about 1 month ago
DedeProGamesΒ 
posted an update 3 months ago
view post
Post
376
πŸš€ Introducing the GRM-2.6 Family

The GRM-2.6 family is a new generation of reasoning-focused models from Orion LLM Labs, built for difficult tasks, coding, STEM, terminal agents, and advanced local AI workflows.

GRM-2.6-Plus is the main high-capability model in the family: a 27B-class reasoning model based on Qwen3.6, designed for strong structured reasoning, coding, agentic use, and practical local deployment.

GRM-2.6-Opus builds on GRM-2.6-Plus as a merge with an Opus-style reasoning distilled model, improving structured reasoning behavior, terminal-agent workflows, coding ability, and complex problem solving.

Both models are designed for users who want powerful reasoning models that remain practical for research, local inference, coding, and agent experiments.

Models:
GRM-2.6-Plus: OrionLLM/GRM-2.6-Plus
GRM-2.6-Opus: OrionLLM/GRM-2.6-Opus

Organization:
OrionLLM
DedeProGamesΒ 
posted an update 3 months ago
view post
Post
5186
πŸš€ Introducing the GRM-2.6 Family

The GRM-2.6 family is a new generation of reasoning-focused models from Orion LLM Labs, built for difficult tasks, coding, STEM, terminal agents, and advanced local AI workflows.

GRM-2.6-Plus is the main high-capability model in the family: a 27B-class reasoning model based on Qwen3.6, designed for strong structured reasoning, coding, agentic use, and practical local deployment.

GRM-2.6-Opus builds on GRM-2.6-Plus as a merge with an Opus-style reasoning distilled model, improving structured reasoning behavior, terminal-agent workflows, coding ability, and complex problem solving.

Both models are designed for users who want powerful reasoning models that remain practical for research, local inference, coding, and agent experiments.

Models:
GRM-2.6-Plus: OrionLLM/GRM-2.6-Plus
GRM-2.6-Opus: OrionLLM/GRM-2.6-Opus

Organization:
OrionLLM
  • 2 replies
Β·
DedeProGamesΒ 
posted an update 3 months ago
view post
Post
8592
GRaPE 2 Pro is now available.

SL-AI/GRaPE-2-Pro

This is the flagship model of the GRaPE 2 family and the largest model I have trained to date, sitting at 27B parameters. It is built on Qwen3.5-27B and trained on a closed-source proprietary dataset, with roughly half of post-training focused on code and the rest split between STEAM subjects and structured logical reasoning. It punches seriously above its weight class.

GRaPE 2 Pro supports multimodal input (image + text) and features 6 thinking modes via the <thinking_mode> tag. This gives you real control over how hard the model thinks, from skipping the reasoning phase entirely with minimal, all the way up to xtra-Hi for deep, extended thought on hard problems. For most agentic use, auto or low is the move to keep things snappy.

It also runs on consumer hardware. You can get it going with as low as 12GB of VRAM on a quantized build.

If you want to try it out and give feedback, that would be really appreciated. Email us at contact@skinnertopia.com
  • 1 reply
Β·
DedeProGamesΒ 
posted an update 4 months ago
view post
Post
3487
πŸ”₯ GRM-2.5 - The most POWERFUL model for local inference

The GRM-2.5 is the newest model from Orion LLM Labs. It has consistent RAW reasoning and is capable of generating very precise responses, similar to large models, while maintaining a parameter size of 4b.

The GRM-2.5 family consists of these models:
OrionLLM/GRM-2.5 (4b)
OrionLLM/GRM-2.5-Air (0.8b)

Furthermore, the GRM-2.5 is the best option for local agentic environments, being very good in code, terminal agent, etc. It is capable of generating 1000 lines of consistent code and programming like large models.
The GRM-2.5 is the best base for FineTune to date and has vision, which means it can interpret images and videos.
  • 1 reply
Β·
DedeProGamesΒ 
posted an update 4 months ago
view post
Post
3062
πŸ”₯ GRM2 - The small one that surpasses the big ones.
What if a 3-parameter model can beat a 32-parameter model in every benchmark? We prove that it can.
GRM2 is a 3b params model based on the llama architecture, trained for long reasoning and high performance in complex tasks - the first 3b params model to outperform qwen3-32b in ALL benchmarks, and outperform o3-mini in almost all benchmarks.
πŸ€— Model: OrionLLM/GRM2-3b
The first 3b params model to generate over 1000 lines of code and achieve a score of 39.0 in xBench-DeepSearch-2510.

πŸš€ Chat with GRM:
https://huggingface.co/spaces/DedeProGames/GRM2-Chat

πŸ† Download official GGUFs: OrionLLM/GRM2-3b-GGUF
DedeProGamesΒ 
posted an update 4 months ago
view post
Post
1792
Introducing GRM2, a powerful 3 billion parameter model designed for long-term reasoning and high performance in complex tasks.

Even with only 3 billion parameters, it outperforms qwen3-32b in several benchmarks and complex reasoning tasks.

With just 3 billion parameters, it can also generate extensive and complex code with over 1000 lines, utilize tools comparable to larger models, and is perfect for agentic tasks.

GRM2 is licensed under Apache 2.0, making it ideal as a base for FineTune in other tasks.

GRM2 Model Page: OrionLLM/GRM2-3b
Official GRM2 GGUFs Quantizations: OrionLLM/GRM2-3b-GGUF
DedeProGamesΒ 
posted an update 4 months ago
view post
Post
193
Introducing GRM-Coder, a 14b params code model based on Qwen3-14b .

On LiveCodeBench v6 (01/08/2024 - 01/05/2025), we achieved a Pass@1 accuracy ofΒ 67.87%, upΒ 7.08%Β from the baseline Pass@1 accuracy ofΒ 60.79%Β of Qwen3-14B.
OrionLLM/GRM-Coder-14b

DedeProGamesΒ 
posted an update 4 months ago
view post
Post
5279
Introducing GRM2, a powerful 3 billion parameter model designed for long-term reasoning and high performance in complex tasks.

Even with only 3 billion parameters, it outperforms qwen3-32b in several benchmarks and complex reasoning tasks.

With just 3 billion parameters, it can also generate extensive and complex code with over 1000 lines, utilize tools comparable to larger models, and is perfect for agentic tasks.

GRM2 is licensed under Apache 2.0, making it ideal as a base for FineTune in other tasks.
You can see more here: OrionLLM/GRM2-3b