See It from My Perspective: Diagnosing the Western Cultural Bias of Large Vision-Language Models in Image Understanding
Paper • 2406.11665 • Published • 1
How to use amitha/mllava-llama2-en with Transformers:
# Use a pipeline as a high-level helper
# Warning: Pipeline type "visual-question-answering" is no longer supported in transformers v5.
# You must load the model directly (see below) or downgrade to v4.x with:
# pip install "transformers<5.0.0"
from transformers import pipeline
pipe = pipeline("visual-question-answering", model="amitha/mllava-llama2-en", trust_remote_code=True) # pip install -U transformers accelerate
# Load model directly
from transformers import AutoModelForVisualQuestionAnswering
model = AutoModelForVisualQuestionAnswering.from_pretrained("amitha/mllava-llama2-en", trust_remote_code=True, device_map="auto")The English Llama2-7B-Chat VLM trained via LORA for https://arxiv.org/abs/2406.11665.