AI & ML interests

Open science and open source

spillai 
posted an update 24 days ago
view post
Post
3202
We're excited to introduce VLM Run Gateway - a new unified OpenAI-compatible API for running open-weight VLMs, OCR VLMs and ViT-based vision models.

https://vlm.run/gateway
Full model catalog: https://vlm.run/gateway/models
Blog post announcement: https://www.vlm.run/blog/introducing-gateway

Try different models on the gateway simply by updating the model name. Free to use and no sign-up required for now (in alpha).

$ uvx vlmrun gw models
$ uvx vlmrun gw chat <doc>.pdf -m glm-ocr
$ uvx vlmrun gw chat <doc>.pdf -m deepseek-ocr-2
$ uvx vlmrun gw chat <doc>.pdf -m pp-ocrv6
$ uvx vlmrun gw chat <img>.jpg -m qwen/qwen3.5-0.8b -p "describe the image"
$ uvx vlmrun gw chat <vid>.mp4 -m qwen/qwen3.5-0.8b -p "describe the video"

jeffboudier 
posted an update 3 months ago
alvarobartt 
posted an update 4 months ago
view post
Post
1426
Open agents on AWS SageMaker AI with open models from the Hugging Face Hub!

> Deploy an open model from the Hugging Face Hub on SageMaker AI
> Connect the deployed model to Strands Agents
> Add built-in and custom tools for tool calling
> Expose external capabilities through MCP integration
> Bonus: talk to your agent and visualize traces with Gradio

https://alvarobartt.com/agents-on-aws-sagemaker
alvarobartt 
posted an update 4 months ago
view post
Post
4087
Latest hf-mem release added a breakdown of Mixture-of-Experts (MoE) memory usage!

TL; DR MoEs can be misleading to reason about from active parameters alone, since each token only activates a subset of experts, while the serving setup still needs to account for the full resident memory footprint.

🧠 hf-mem now splits MoE memory into base model weights, routed experts, and KV cache
🏗️ Dense models usually load and use most weights every forward pass, while MoEs load many experts but only route each token to a few of them
⚡ Active params isn't the same as memory footprint, especially for sparse architectures
📦 Runtime memory is about what is used per request/token, while loading memory also includes the expert weights that need to be resident
📚 KV cache can still dominate depending on context length, batch size, and concurrency
🔀 Expert Parallelism (EP) helps shard experts across accelerators when expert weights dominate
🚀 Data Parallelism (DP) + EP is often a good fit for throughput-oriented MoE serving

Check the repository at https://github.com/alvarobartt/hf-mem
spillai 
posted an update 5 months ago
view post
Post
8798
mm-ctx – fast, multimodal context for agents.

LLM-based agents handle text incredibly well, but images, videos, or PDFs with visual content are hard to interpret. mm-ctx gives your CLI agent multi-modal skills.

Try it interactively in Spaces: vlm-run/mm-ctx

Readme: https://vlm-run.github.io/mm/
PyPI: https://pypi.org/project/mm-ctx
SKILL.md: https://github.com/vlm-run/skills/blob/main/skills/mm-cli-skill/SKILL.md

mm-ctx is meant to feel familiar: the UNIX tools we already love (find/cat/grep/wc), rebuilt for file types LLMs can't read natively and designed to work with agents via the CLI.
- mm grep "invoice #1234" ~/Downloads searches across PDFs and returns line-numbered matches
- mm cat <document>.pdf returns a metadata description of the file
- mm cat <photo>.jpg returns a caption of the photo
- mm cat <video>.mp4 returns a caption of the video

A few things we obsessed over:
⚡ Speed: Rust core for the hot paths
🏠 Local-first, BYO model: Uses any OpenAI-compatible endpoint: Ollama, vLLM/SGLang, LMStudio with any multimodal LLM (Gemma4, Qwen3.5, GLM-4.6V).
🔗 Composable: stdin + structured outputs
🤖 Drops into any agent via mm-cli-skills: Claude Code, Codex, Gemini CLI, OpenClaw.

We’d love to hear your feedback! Especially on the CLI and what file types and workflows you would like to see next.
  • 2 replies
·
oncody 
posted an update 6 months ago
view post
Post
230
Are Large Language Models actually becoming more intelligent, or just better at seeming intelligent?

There is a noticeable shift happening in the LLM space.

Models today can:

Generate cleaner and more structured code.
Explain complex topics in simpler ways.
Maintain longer and more coherent conversations.

Yet at the same time, they still:

Produce confident hallucinations.
Fail in multi-step reasoning tasks.
Break under slightly unfamiliar or challenging inputs.

This raises a critical question.

Are we advancing intelligence, or optimizing presentation?

Most improvements so far seem driven by:

Larger datasets.
Increased scale.
Alignment techniques like RLHF.

But these do not necessarily lead to genuine reasoning ability.

What still appears fundamentally missing:

Persistent memory across interactions.
True reasoning rather than pattern completion.
Grounded understanding connected to real-world context.

Reliable self-correction and verification mechanisms.

If current scaling trends start to plateau, the next breakthrough will not come from doing more of the same.

So the real question for the community is:

If you were designing the next generation of AI systems, where would you focus?

A. Larger models and compute
B. Higher-quality and structured data
C. Agent-based systems with tool use and memory
D. New architectures beyond transformers

This is not just a technical discussion. It defines where AI is actually heading over the next few years.

I am interested to hear how others are thinking about this.
alvarobartt 
posted an update 7 months ago
view post
Post
4217
Learn how to deploy Microsoft Research VibeVoice ASR on Microsoft Azure Foundry with Hugging Face to generate rich audio transcriptions with Who, When, and What! 💥

> 🕒 60-minute single-pass processing, no chunking or stitching
> 👤 Customized hotwords to guide recognition on domain-specific content
> 📝 Rich transcription: joint ASR + diarization + timestamping in one pass
> 🌍 50+ languages with automatic detection and code-switching support
> 🤗 Deployed on Microsoft Foundry via an OpenAI-compatible Chat Completions API

https://huggingface.co/docs/microsoft-azure/foundry/examples/deploy-vibevoice-asr
jorgemunozl 
posted an update 8 months ago
view post
Post
516
just published a short article about something that bit me hard while porting PI05’s subtask prediction to PyTorch: left vs right alignment in transformer padding.
turns out JAX (what Physical Intelligence used) and Hugging Face use opposite padding conventions — and if you don’t catch it, your model silently produces nonsense instead of crashing. no NaN, no error, just garbled subtasks 🤡
i walk through the full tensor pipeline — images → embeddings → pad masks → attention masks → position IDs — and show exactly where the mismatch corrupts everything. also included the implementation file with the fix.
if you’ve ever ported a model between frameworks or messed with custom attention patterns, i think you will enjoy it
  • 1 reply
·
alvarobartt 
posted an update 8 months ago
view post
Post
3601
💥 hf-mem v0.4.1 now also estimates KV cache memory requirements for any context length and batch size with the --experimental flag!

uvx hf-mem --model-id ... --experimental will automatically pull the required information from the Hugging Face Hub to include the KV cache estimation, when applicable.

💡 Alternatively, you can also set the --max-model-len, --batch-size and --kv-cache-dtype arguments (à la vLLM) manually if preferred.
  • 1 reply
·
multimodalart 
posted an update 12 months ago
view post
Post
38809
Want to iterate on a Hugging Face Space with an LLM?

Now you can easily convert any HF entire repo (Model, Dataset or Space) to a text file and feed it to a language model!

multimodalart/repo2txt
  • 3 replies
·
jeffboudier 
posted an update about 1 year ago
view post
Post
3465
Quick 30s demo of the new Hub > Azure AI integration to deploy HF models in your own Azure account. Now with Py and CLI!

GG @alvarobartt @kramp @pagezyhf
BrigitteTousi 
posted an update about 1 year ago
BrigitteTousi 
posted an update about 1 year ago
view post
Post
711
New interactive viz from AI World showing OpenAI's new open model gpt-oss-120b breaking into the top 50 most liked models of all time on the Hub in under a day! ☄️☄️☄️
BrigitteTousi 
posted an update about 1 year ago
view post
Post
715
This is what Hugging Face is all about. We want everyone, hobbyists, researchers and industry alike, to be able to contribute to AI because everyone is affected by it. Kudos to HF's @irenesolaiman for spreading the word!🔥🤗
jeffboudier 
posted an update over 1 year ago
view post
Post
631
AMD summer hackathons are here!
A chance to get hands-on with MI300X GPUs and accelerate models.
🇫🇷 Paris - Station F - July 5-6
🇮🇳 Mumbai - July 12-13
🇮🇳 Bengaluru - July 19-20

Hugging Face and GPU Mode will be on site and on July 6 in Paris @ror will share lessons learned while building new kernels to accelerate Llama 3.1 405B on ROCm

Register to Paris event: https://lu.ma/fmvdjmur?tk=KeAbiP
All dates: https://lu.ma/calendar/cal-3sxhD5FdxWsMDIz
multimodalart 
posted an update over 1 year ago
view post
Post
18611
Self-Forcing - a real-time video distilled model from Wan 2.1 by @adobe is out, and they open sourced it 🐐

I've built a live real time demo on Spaces 📹💨

multimodalart/self-forcing
  • 6 replies
·
jeffboudier 
posted an update over 1 year ago
view post
Post
1792
Today we launched Training Cluster as a Service, to make the new DGX Cloud Lepton supercloud easily accessible to AI researchers.

Hugging Face will collaborate with NVIDIA to provision and set up GPU training clusters to make them available for the duration of training runs.

Hugging Face organizations can sign up here: https://huggingface.co/training-cluster