Instructions to use vtava/Laya-Integrated-Memory-V22 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vtava/Laya-Integrated-Memory-V22 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vtava/Laya-Integrated-Memory-V22")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vtava/Laya-Integrated-Memory-V22", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vtava/Laya-Integrated-Memory-V22 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vtava/Laya-Integrated-Memory-V22" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Laya-Integrated-Memory-V22", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/vtava/Laya-Integrated-Memory-V22
- SGLang
How to use vtava/Laya-Integrated-Memory-V22 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vtava/Laya-Integrated-Memory-V22" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Laya-Integrated-Memory-V22", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vtava/Laya-Integrated-Memory-V22" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Laya-Integrated-Memory-V22", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use vtava/Laya-Integrated-Memory-V22 with Docker Model Runner:
docker model run hf.co/vtava/Laya-Integrated-Memory-V22
Download loader.py from vtava/Laya-Integrated-Memory-V22: direct link, hf CLI and curl.
- Browser
- Download file 1.03 kB
-
https://huggingface.co/vtava/Laya-Integrated-Memory-V22/resolve/main/loader.py
- Command line
-
hf download hf://vtava/Laya-Integrated-Memory-V22/loader.py
-
curl -L -o loader.py https://huggingface.co/vtava/Laya-Integrated-Memory-V22/resolve/main/loader.py
1.03 kB
| from pathlib import Path | |
| import torch | |
| from huggingface_hub import snapshot_download | |
| import laya | |
| from tinycenn_lm.laya_lab.core import LayaLabConfig | |
| from tinycenn_lm.laya_lab.factory import make_replacement | |
| def load_model(repo_id, device=None, token=None): | |
| root = Path(snapshot_download( | |
| repo_id, | |
| repo_type="model", | |
| token=token, | |
| allow_patterns=["adapter.pt", "report.json", "model_meta.json"], | |
| )) | |
| payload = torch.load(root / "adapter.pt", map_location="cpu", weights_only=False) | |
| cfg = LayaLabConfig(**payload["lab_config"]) | |
| agent = laya.load(cfg.model_id, device=device, token=token) | |
| for layer_id, spec in payload["adapters"].items(): | |
| idx = int(layer_id) | |
| layer = agent.model.encoder.layers[idx] | |
| replacement = make_replacement(layer.attn, cfg) | |
| replacement.load_state_dict(spec["state_dict"], strict=True) | |
| replacement.to(agent.device).eval().requires_grad_(False) | |
| layer.attn = replacement | |
| agent.model.eval() | |
| return agent | |