Instructions to use HuiLin0220/Medcat with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use HuiLin0220/Medcat with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="HuiLin0220/Medcat")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("HuiLin0220/Medcat", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use HuiLin0220/Medcat with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "HuiLin0220/Medcat" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "HuiLin0220/Medcat", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/HuiLin0220/Medcat
- SGLang
How to use HuiLin0220/Medcat with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "HuiLin0220/Medcat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "HuiLin0220/Medcat", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "HuiLin0220/Medcat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "HuiLin0220/Medcat", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use HuiLin0220/Medcat with Docker Model Runner:
docker model run hf.co/HuiLin0220/Medcat
Download NOTICE from HuiLin0220/Medcat: direct link, hf CLI and curl.
- Browser
- Download file 679 Bytes
-
https://huggingface.co/HuiLin0220/Medcat/resolve/main/NOTICE
- Command line
-
hf download hf://HuiLin0220/Medcat/NOTICE
-
curl -L -o NOTICE https://huggingface.co/HuiLin0220/Medcat/resolve/main/NOTICE
679 Bytes
| Medcat V10 | |
| Copyright 2026 Hui Lin | |
| Medcat inference code and adapters derive in part from ME-VLIP, distributed | |
| under the Apache License, Version 2.0: | |
| https://github.com/BioMedIA-MBZUAI/ME-VLIP | |
| The bundled InternVL3-8B-hf base model derives from InternVL by OpenGVLab. | |
| InternVL is distributed under the MIT License; see LICENSE-INTERNVL. | |
| Qwen is licensed under the Qwen LICENSE AGREEMENT, Copyright (c) Alibaba | |
| Cloud. All Rights Reserved. See LICENSE-QWEN. | |
| The bundled base-model files are redistributed without numerical | |
| modification from OpenGVLab/InternVL3-8B-hf. Medcat task/source adapters and | |
| auxiliary heads are separate fine-tuned components loaded at inference time. | |