Text Generation
Transformers
ONNX
GGUF
PyTorch
English
qtensorformer
tensor-networks
model-compression
adaptive-computation
kv-cache-compression
hardware-aware
energy-aware
green-ai
custom-code
multimodal
ollama
webgpu
triton
low-rank-adaptation
mixture-of-depths
edge-ai
custom_code
Instructions to use Premchan369/Q-TensorFormer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Premchan369/Q-TensorFormer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Premchan369/Q-TensorFormer", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Premchan369/Q-TensorFormer", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Premchan369/Q-TensorFormer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Premchan369/Q-TensorFormer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Premchan369/Q-TensorFormer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Premchan369/Q-TensorFormer
- SGLang
How to use Premchan369/Q-TensorFormer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Premchan369/Q-TensorFormer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Premchan369/Q-TensorFormer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Premchan369/Q-TensorFormer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Premchan369/Q-TensorFormer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Premchan369/Q-TensorFormer with Docker Model Runner:
docker model run hf.co/Premchan369/Q-TensorFormer
Premchandyadav369
feat(frontier): achieve 10/10 with Triton kernels, GGUF/Ollama exporter, WebGPU runtime, multimodal vision, technical report, and 81 tests
4e689f6 Download web/index.html from Premchan369/Q-TensorFormer: direct link, hf CLI and curl.
- Browser
- Download file 2.69 kB
-
https://huggingface.co/Premchan369/Q-TensorFormer/resolve/main/web/index.html
- Command line
-
hf download hf://Premchan369/Q-TensorFormer/web/index.html
-
curl -L -o index.html https://huggingface.co/Premchan369/Q-TensorFormer/resolve/main/web/index.html
2.69 kB
| <html lang="en"> | |
| <head> | |
| <meta charset="UTF-8"> | |
| <title>Q-TensorFormer WebGPU Browser Inference</title> | |
| <script src="https://cdn.jsdelivr.net/npm/onnxruntime-web/dist/ort.webgpu.min.js"></script> | |
| <style> | |
| body { font-family: -apple-system, BlinkMacSystemFont, "Segoe UI", Roboto, sans-serif; max-width: 800px; margin: 40px auto; padding: 20px; line-height: 1.6; } | |
| h1 { color: #2563eb; } | |
| .box { background: #f8fafc; border: 1px solid #e2e8f0; border-radius: 8px; padding: 20px; margin-bottom: 20px; } | |
| button { background: #2563eb; color: white; border: none; padding: 10px 20px; border-radius: 6px; cursor: pointer; font-size: 16px; } | |
| button:hover { background: #1d4ed8; } | |
| #output { background: #1e293b; color: #38bdf8; font-family: monospace; padding: 15px; border-radius: 6px; white-space: pre-wrap; min-height: 100px; } | |
| </style> | |
| </head> | |
| <body> | |
| <h1>⚛️ Q-TensorFormer: Zero-Server WebGPU Inference</h1> | |
| <div class="box"> | |
| <p>Run closed-loop Tensor-Train transformer generation locally inside your browser using WebGPU hardware acceleration.</p> | |
| <button id="runBtn" onclick="runInference()">Initialize WebGPU & Generate</button> | |
| </div> | |
| <div id="output">Click the button above to load and run the model via WebGPU...</div> | |
| <script> | |
| async function runInference() { | |
| const out = document.getElementById('output'); | |
| out.innerText = "Loading ONNX model via WebGPU..."; | |
| try { | |
| // Initialize ONNX Runtime Web session | |
| const session = await ort.InferenceSession.create('./qtensorformer.onnx', { | |
| executionProviders: ['webgpu', 'wasm'] | |
| }); | |
| out.innerText = "Model loaded successfully! Running forward pass...\n"; | |
| // Create dummy input | |
| const inputData = new BigInt64Array([101n, 2054n, 2003n, 1037n]); | |
| const tensor = new ort.Tensor('int64', inputData, [1, 4]); | |
| const start = performance.now(); | |
| const feeds = { input_ids: tensor }; | |
| const results = await session.run(feeds); | |
| const elapsed = (performance.now() - start).toFixed(2); | |
| out.innerText += `Generation completed in ${elapsed} ms!\nLogits shape: ${results.logits.dims}\nWebGPU acceleration active!`; | |
| } catch (err) { | |
| out.innerText = "WebGPU execution status: " + err.message + "\n(Ensure qtensorformer.onnx is served in the same directory)"; | |
| } | |
| } | |
| </script> | |
| </body> | |
| </html> | |