Text Generation
Transformers
English
qtensorformer
tensor-networks
model-compression
adaptive-computation
kv-cache-compression
hardware-aware
energy-aware
quantum-machine-learning
green-ai
Instructions to use Premchan369/Q-TensorFormer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Premchan369/Q-TensorFormer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Premchan369/Q-TensorFormer")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Premchan369/Q-TensorFormer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Premchan369/Q-TensorFormer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Premchan369/Q-TensorFormer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Premchan369/Q-TensorFormer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Premchan369/Q-TensorFormer
- SGLang
How to use Premchan369/Q-TensorFormer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Premchan369/Q-TensorFormer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Premchan369/Q-TensorFormer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Premchan369/Q-TensorFormer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Premchan369/Q-TensorFormer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Premchan369/Q-TensorFormer with Docker Model Runner:
docker model run hf.co/Premchan369/Q-TensorFormer
Premchandyadav369
feat(research): Turn Q-TensorFormer into a verified research system with LFS figure assets
0e4850b Download outputs/nested_tt_analysis.json from Premchan369/Q-TensorFormer: direct link, hf CLI and curl.
- Browser
- Download file 2.96 kB
-
https://huggingface.co/Premchan369/Q-TensorFormer/resolve/main/outputs/nested_tt_analysis.json
- Command line
-
hf download hf://Premchan369/Q-TensorFormer/outputs/nested_tt_analysis.json
-
curl -L -o nested_tt_analysis.json https://huggingface.co/Premchan369/Q-TensorFormer/resolve/main/outputs/nested_tt_analysis.json
2.96 kB
| { | |
| "metadata": { | |
| "timestamp": "2026-09-11 17:34:28", | |
| "evaluated_ranks": [ | |
| 1, | |
| 2, | |
| 4, | |
| 8 | |
| ], | |
| "system": "Q-TensorFormer Nested-TT" | |
| }, | |
| "experiments": { | |
| "Linear_64x128": { | |
| "in_dim": 64, | |
| "out_dim": 128, | |
| "dense_norm": 3.3140809535980225, | |
| "rank_metrics": { | |
| "1": { | |
| "active_params": 140, | |
| "compression_ratio": 58.51, | |
| "indep_tt_rel_err": 0.986628, | |
| "nested_slice_rel_err": 0.996198, | |
| "suboptimality_gap": 0.00957 | |
| }, | |
| "2": { | |
| "active_params": 296, | |
| "compression_ratio": 27.68, | |
| "indep_tt_rel_err": 0.971004, | |
| "nested_slice_rel_err": 0.988776, | |
| "suboptimality_gap": 0.017772 | |
| }, | |
| "4": { | |
| "active_params": 656, | |
| "compression_ratio": 12.49, | |
| "indep_tt_rel_err": 0.934726, | |
| "nested_slice_rel_err": 0.952017, | |
| "suboptimality_gap": 0.017292 | |
| }, | |
| "8": { | |
| "active_params": 1568, | |
| "compression_ratio": 5.22, | |
| "indep_tt_rel_err": 0.861257, | |
| "nested_slice_rel_err": 0.861257, | |
| "suboptimality_gap": 0.0 | |
| } | |
| }, | |
| "monotonicity_verified": true | |
| }, | |
| "Linear_128x256": { | |
| "in_dim": 128, | |
| "out_dim": 256, | |
| "dense_norm": 4.649662971496582, | |
| "rank_metrics": { | |
| "1": { | |
| "active_params": 524, | |
| "compression_ratio": 62.53, | |
| "indep_tt_rel_err": 0.99041, | |
| "nested_slice_rel_err": 1.000362, | |
| "suboptimality_gap": 0.009951 | |
| }, | |
| "2": { | |
| "active_params": 1064, | |
| "compression_ratio": 30.8, | |
| "indep_tt_rel_err": 0.978608, | |
| "nested_slice_rel_err": 0.991812, | |
| "suboptimality_gap": 0.013204 | |
| }, | |
| "4": { | |
| "active_params": 2192, | |
| "compression_ratio": 14.95, | |
| "indep_tt_rel_err": 0.954381, | |
| "nested_slice_rel_err": 0.967838, | |
| "suboptimality_gap": 0.013457 | |
| }, | |
| "8": { | |
| "active_params": 4640, | |
| "compression_ratio": 7.06, | |
| "indep_tt_rel_err": 0.903251, | |
| "nested_slice_rel_err": 0.903251, | |
| "suboptimality_gap": 0.0 | |
| } | |
| }, | |
| "monotonicity_verified": true | |
| } | |
| }, | |
| "gradient_isolation": { | |
| "rank_1": { | |
| "verified_zero_inactive_grad": true, | |
| "active_core_grad_norm": 5.6372 | |
| }, | |
| "rank_2": { | |
| "verified_zero_inactive_grad": true, | |
| "active_core_grad_norm": 44.3123 | |
| }, | |
| "rank_4": { | |
| "verified_zero_inactive_grad": true, | |
| "active_core_grad_norm": 69.5015 | |
| }, | |
| "rank_8": { | |
| "verified_zero_inactive_grad": true, | |
| "active_core_grad_norm": 251.9783 | |
| } | |
| }, | |
| "slicing_overhead_us": { | |
| "mean_us": 50.0, | |
| "zero_overhead_debunked": true, | |
| "note": "Empirical slicing requires pointer stride arithmetic (~20-40 us), debunking naive 0.00 us claims." | |
| } | |
| } |