Instructions to use 45th/comfyui-model-pack with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use 45th/comfyui-model-pack with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("45th/comfyui-model-pack", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 45th/comfyui-model-pack with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 45th/comfyui-model-pack # Run inference directly in the terminal: llama cli -hf 45th/comfyui-model-pack
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 45th/comfyui-model-pack # Run inference directly in the terminal: llama cli -hf 45th/comfyui-model-pack
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 45th/comfyui-model-pack # Run inference directly in the terminal: ./llama-cli -hf 45th/comfyui-model-pack
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 45th/comfyui-model-pack # Run inference directly in the terminal: ./build/bin/llama-cli -hf 45th/comfyui-model-pack
Use Docker
docker model run hf.co/45th/comfyui-model-pack
- LM Studio
- Jan
- Ollama
How to use 45th/comfyui-model-pack with Ollama:
ollama run hf.co/45th/comfyui-model-pack
- Unsloth Desktop
- Pi
How to use 45th/comfyui-model-pack with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 45th/comfyui-model-pack
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "45th/comfyui-model-pack" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use 45th/comfyui-model-pack with Docker Model Runner:
docker model run hf.co/45th/comfyui-model-pack
- Lemonade
How to use 45th/comfyui-model-pack with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 45th/comfyui-model-pack
Run and chat with the model
lemonade run user.comfyui-model-pack-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use 45th/comfyui-model-pack with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 45th/comfyui-model-pack
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default 45th/comfyui-model-pack
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use 45th/comfyui-model-pack with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 45th/comfyui-model-pack
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "45th/comfyui-model-pack" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download custom_nodes/comfyui-technodes/quant_nodes.py from 45th/comfyui-model-pack: direct link, hf CLI and curl.
- Browser
- Download file 4.28 kB
-
https://huggingface.co/45th/comfyui-model-pack/resolve/main/custom_nodes/comfyui-technodes/quant_nodes.py
- Command line
-
hf download hf://45th/comfyui-model-pack/custom_nodes/comfyui-technodes/quant_nodes.py
-
curl -L -o quant_nodes.py https://huggingface.co/45th/comfyui-model-pack/resolve/main/custom_nodes/comfyui-technodes/quant_nodes.py
4.28 kB
| import torch | |
| import copy | |
| import folder_paths | |
| import comfy_extras.nodes_model_merging | |
| def quantize_tensor(tensor, num_bits=8, dtype=torch.float16, dequant=True): | |
| """ | |
| Quantizes a tensor to a specified number of bits. | |
| Args: | |
| tensor (torch.Tensor): The input tensor to be quantized. | |
| num_bits (int): The number of bits to use for quantization (default: 8). | |
| dtype(torch.dtype): The datatype to use for the output (default: torch.float16). | |
| dequant (bool): Whether to dequantize or not (default: true). | |
| Returns: | |
| torch.Tensor: The quantized tensor. | |
| """ | |
| # Determine the minimum and maximum values of the tensor | |
| min_val = tensor.min() | |
| max_val = tensor.max() | |
| # Calculate the scale factor and zero point | |
| qmin = 0 | |
| qmax = 2 ** num_bits - 1 | |
| scale = (max_val - min_val) / (qmax - qmin) | |
| zero_point = qmin - torch.round(min_val / scale) | |
| # Quantize the tensor | |
| quantized_tensor = torch.round(tensor / scale + zero_point) | |
| quantized_tensor = torch.clamp(quantized_tensor, qmin, qmax) | |
| # Convert the quantized tensor to the datatype | |
| dequantized_tensor = quantized_tensor.to(dtype) | |
| if dequant: | |
| # De-quantize the tensor | |
| dequantized_tensor = (dequantized_tensor - zero_point) * scale | |
| return dequantized_tensor | |
| def quantize_model(model, in_bits, mid_bits, out_bits, dtype=torch.float16, dequant=True): | |
| # Clone the base model to create a new one | |
| quantized_model = model.clone() | |
| # Get the key patches from the model with the prefix "diffusion_model." | |
| key_patches = quantized_model.get_key_patches("diffusion_model.") | |
| # Iterate over each key patch in the patches | |
| for key in key_patches: | |
| if ".input_" in key: | |
| num_bits = in_bits | |
| elif ".middle_" in key: | |
| num_bits = mid_bits | |
| elif ".output_" in key: | |
| num_bits = out_bits | |
| else: | |
| num_bits = 8 | |
| quantized_tensor = quantize_tensor(key_patches[key][0], num_bits, dtype, dequant) | |
| quantized_model.add_patches({key: (quantized_tensor,)}, 1, 0) | |
| # Return the quantized model | |
| return quantized_model | |
| def quantize_clip(clip, bits, dtype=torch.float16, dequant=True): | |
| # Clone the base model to create a new one | |
| quantized_clip = clip.clone() | |
| # Get the key patches from the model with the prefix "diffusion_model." | |
| key_patches = quantized_clip.get_key_patches() | |
| # Iterate over each key patch in the patches | |
| for key in key_patches: | |
| quantized_tensor = quantize_tensor(key_patches[key][0], bits, dtype, dequant) | |
| quantized_clip.add_patches({key: (quantized_tensor,)}, 1, 0) | |
| # Return the quantized model | |
| return quantized_clip | |
| def quantize_vae(vae, bits, dtype=torch.float16, dequant=True): | |
| # Create a clone of the VAE model | |
| quantized_vae = copy.deepcopy(vae) | |
| # Get the state dictionary from the clone | |
| state_dict = quantized_vae.first_stage_model.state_dict() | |
| # Iterate over each key-value pair in the state dictionary | |
| for key, value in state_dict.items(): | |
| state_dict[key] = quantize_tensor(value, bits, dtype, dequant) | |
| # Load the quantized state dictionary back into the clone | |
| quantized_vae.first_stage_model.load_state_dict(state_dict) | |
| # Return the quantized clone | |
| return quantized_vae | |
| class ModelQuant: | |
| def INPUT_TYPES(cls): | |
| return { | |
| "required": { | |
| "model": ["MODEL"], | |
| "in_bits": ("INT", {"default": 8, "min": 1, "max": 8}), | |
| "mid_bits": ("INT", {"default": 8, "min": 1, "max": 8}), | |
| "out_bits": ("INT", {"default": 8, "min": 1, "max": 8}), | |
| } | |
| } | |
| RETURN_TYPES = ["MODEL"] | |
| FUNCTION = "quant_model" | |
| CATEGORY = "TechNodes/quantization" | |
| def quant_model(self, model, in_bits, mid_bits, out_bits): | |
| return [quantize_model(model, in_bits, mid_bits, out_bits)] | |
| class ClipQuant: | |
| def INPUT_TYPES(cls): | |
| return { | |
| "required": { | |
| "clip": ["CLIP"], | |
| "bits": ("INT", {"default": 8, "min": 1, "max": 8}), | |
| } | |
| } | |
| RETURN_TYPES = ["CLIP"] | |
| FUNCTION = "quant_clip" | |
| CATEGORY = "TechNodes/quantization" | |
| def quant_clip(self, clip, bits): | |
| return [quantize_clip(clip, bits)] | |
| class VAEQuant: | |
| def INPUT_TYPES(cls): | |
| return { | |
| "required": { | |
| "vae": ["VAE"], | |
| "bits": ("INT", {"default": 8, "min": 1, "max": 8}), | |
| } | |
| } | |
| RETURN_TYPES = ["VAE"] | |
| FUNCTION = "quant_vae" | |
| CATEGORY = "TechNodes/quantization" | |
| def quant_vae(self, vae, bits): | |
| return [quantize_vae(vae, bits)] | |