Models runnable on my workstation (VRAM/RAM friendly). Practical shortlist for local experiments and ICE integration.
🔄 In a Training Loop
Francesco Maiomascio
francescomaiomascio
AI & ML interests
Public reference hub for ML baselines, evaluation resources, and demos.
Publishing is restricted to reproducible artifacts with clear scope, limitations, and documentation.
Recent Activity
liked a model 5 days ago
deepseek-ai/DeepSeek-V4-Flash liked a model 7 days ago
poolside/Laguna-S-2.1 liked a model 4 months ago
bartowski/Qwen_Qwen3-8B-GGUFOrganizations
None yet
ICE • Code & Tool-Use
Code-oriented LLMs and tool-use models for agent execution workflows.
Used to evaluate planning, refactoring, and structured action outputs.
-
Qwen/Qwen2.5-Coder-32B-Instruct
Text Generation • 33B • Updated • 1.2M • • 2.09k -
Qwen/Qwen2.5-Coder-7B-Instruct
Text Generation • 8B • Updated • 2.01M • • 765 -
deepseek-ai/deepseek-coder-6.7b-instruct
Text Generation • 7B • Updated • 480k • 506 -
deepseek-ai/deepseek-coder-33b-instruct
Text Generation • 33B • Updated • 3.16k • 579
ICE • Multimodal (Vision + Speech)
Vision-language and speech models for multimodal IO and perception tasks.
Reference set for captioning, OCR-ish flows, ASR, and VLM reasoning.
-
Qwen/Qwen2.5-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 9.2M • • 1.66k -
llava-hf/llava-v1.6-mistral-7b-hf
Image-Text-to-Text • 8B • Updated • 717k • 313 -
Salesforce/blip2-opt-2.7b
Image-Text-to-Text • 4B • Updated • 520k • 446 -
google/siglip-so400m-patch14-384
Zero-Shot Image Classification • 0.9B • Updated • 2.26M • 681
Research • Archive
Long-term archive of papers, models, datasets, and tools worth revisiting.
Curated for reference, replication, and future deep dives.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 9 -
Agent-as-a-Judge: Evaluate Agents with Agents
Paper • 2410.10934 • Published • 23 -
Agent-SafetyBench: Evaluating the Safety of LLM Agents
Paper • 2412.14470 • Published • 12 -
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
Paper • 2501.11067 • Published • 13
ICE • Core LLM Baselines
Baseline LLMs for ICE experimentation and regression checks.
Reference set for capability, cost, and stability comparisons.
-
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.95M • • 6.44k -
meta-llama/Llama-3.1-70B-Instruct
Text Generation • 71B • Updated • 822k • • 937 -
meta-llama/Llama-3.1-8B
Text Generation • 8B • Updated • 1.53M • • 2.35k -
meta-llama/Llama-3.1-70B
Text Generation • 71B • Updated • 44.2k • • 429
ICE • Retrieval (Embeddings + Rerankers)
Embedding models and rerankers for RAG, search, and semantic indexing.
Reference set for retrieval quality, latency, and multilingual coverage.
ICE • Safety / Guardrails
Safety classifiers and guard models for prompt/content filtering and policy checks.
Reference set for runtime gating, moderation, and risk control lay
Infra • Serving & Optimization
Inference engines, quantization, serving stacks, and perf tooling. Reference list for deployment and latency/cost work.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 9 -
Seesaw: High-throughput LLM Inference via Model Re-sharding
Paper • 2503.06433 • Published -
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters
Paper • 2504.08791 • Published • 141 - Running342
Evaluation Guidebook
📝342Explore LLM benchmark scores over time
Local • Workstation-Ready (≤14B)
Models runnable on my workstation (VRAM/RAM friendly). Practical shortlist for local experiments and ICE integration.
ICE • Core LLM Baselines
Baseline LLMs for ICE experimentation and regression checks.
Reference set for capability, cost, and stability comparisons.
-
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.95M • • 6.44k -
meta-llama/Llama-3.1-70B-Instruct
Text Generation • 71B • Updated • 822k • • 937 -
meta-llama/Llama-3.1-8B
Text Generation • 8B • Updated • 1.53M • • 2.35k -
meta-llama/Llama-3.1-70B
Text Generation • 71B • Updated • 44.2k • • 429
ICE • Code & Tool-Use
Code-oriented LLMs and tool-use models for agent execution workflows.
Used to evaluate planning, refactoring, and structured action outputs.
-
Qwen/Qwen2.5-Coder-32B-Instruct
Text Generation • 33B • Updated • 1.2M • • 2.09k -
Qwen/Qwen2.5-Coder-7B-Instruct
Text Generation • 8B • Updated • 2.01M • • 765 -
deepseek-ai/deepseek-coder-6.7b-instruct
Text Generation • 7B • Updated • 480k • 506 -
deepseek-ai/deepseek-coder-33b-instruct
Text Generation • 33B • Updated • 3.16k • 579
ICE • Retrieval (Embeddings + Rerankers)
Embedding models and rerankers for RAG, search, and semantic indexing.
Reference set for retrieval quality, latency, and multilingual coverage.
ICE • Multimodal (Vision + Speech)
Vision-language and speech models for multimodal IO and perception tasks.
Reference set for captioning, OCR-ish flows, ASR, and VLM reasoning.
-
Qwen/Qwen2.5-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 9.2M • • 1.66k -
llava-hf/llava-v1.6-mistral-7b-hf
Image-Text-to-Text • 8B • Updated • 717k • 313 -
Salesforce/blip2-opt-2.7b
Image-Text-to-Text • 4B • Updated • 520k • 446 -
google/siglip-so400m-patch14-384
Zero-Shot Image Classification • 0.9B • Updated • 2.26M • 681
ICE • Safety / Guardrails
Safety classifiers and guard models for prompt/content filtering and policy checks.
Reference set for runtime gating, moderation, and risk control lay
Research • Archive
Long-term archive of papers, models, datasets, and tools worth revisiting.
Curated for reference, replication, and future deep dives.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 9 -
Agent-as-a-Judge: Evaluate Agents with Agents
Paper • 2410.10934 • Published • 23 -
Agent-SafetyBench: Evaluating the Safety of LLM Agents
Paper • 2412.14470 • Published • 12 -
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
Paper • 2501.11067 • Published • 13
Infra • Serving & Optimization
Inference engines, quantization, serving stacks, and perf tooling. Reference list for deployment and latency/cost work.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 9 -
Seesaw: High-throughput LLM Inference via Model Re-sharding
Paper • 2503.06433 • Published -
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters
Paper • 2504.08791 • Published • 141 - Running342
Evaluation Guidebook
📝342Explore LLM benchmark scores over time