Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
guhao
marcusguhao
1
4
1
Follow
0 followers
·
2 following
AI & ML interests
None yet
Recent Activity
upvoted
a
paper
7 days ago
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
authored
a paper
8 days ago
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
authored
a paper
8 days ago
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
View all activity
Organizations
None yet
marcusguhao
's models
18
Sort: Recently updated
marcusguhao/llama-2-13b-hessian
Updated
Apr 23, 2025
marcusguhao/llama-30b-hessian
Updated
Apr 23, 2025
marcusguhao/llama-30b-inv_hessian
Updated
Apr 23, 2025
marcusguhao/llama-3-8b-inv_hessian
Updated
Apr 23, 2025
marcusguhao/llama-3-8b-hessian
Updated
Apr 23, 2025
marcusguhao/llama-2-13b-inv_hessian
Updated
Apr 23, 2025
marcusguhao/llama-2-7b-inv_hessian
Updated
Apr 23, 2025
marcusguhao/llama-2-7b_hessian
Updated
Apr 23, 2025
marcusguhao/SVD-MoE-merge
Updated
Jan 4, 2025
marcusguhao/SVD-MoE-merge-code
Updated
Jan 4, 2025
marcusguhao/baichuan-moe-hf-3B-vllm
Updated
Oct 30, 2024
marcusguhao/Qwen2-1.5B-int3-128g_q3f16_1
Updated
Oct 24, 2024
marcusguhao/Qwen2-1.5B-Instruct-q4f16_1
Updated
Oct 24, 2024
marcusguhao/Qwen2-1.5B-Instruct-q0f16-no-quantized
Updated
Oct 24, 2024
marcusguhao/Qwen2-1.5B-Instruct-baichuan-distill_q3f16_0
Updated
Oct 21, 2024
marcusguhao/Qwen2-1.5B-Instruct-baichuan-distill_q3f16_1
Updated
Oct 21, 2024
marcusguhao/Llama-3.2-1B_q3f16_1_MLC
Updated
Oct 9, 2024
marcusguhao/Llama-3.2-3B_q3f16_1_MLC
Updated
Oct 9, 2024