Instructions to use Caraaaaa/text_image_captioning with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Caraaaaa/text_image_captioning with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("image-to-text", model="Caraaaaa/text_image_captioning")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Caraaaaa/text_image_captioning") model = AutoModelForMultimodalLM.from_pretrained("Caraaaaa/text_image_captioning", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| datasets: | |
| - Caraaaaa/non_text_image_captioning | |
| pipeline_tag: image-to-text | |
| This is a [GenerativeImage2Text](https://huggingface.co/microsoft/git-base) model finetuned on [non-text images](https://huggingface.co/datasets/Caraaaaa/non_text_image_captioning) extracted from documents (i.e.PDF). It is used to analyze the content of the image and produce a descriptive caption. | |
| It is part of a [project]((https://github.com/caraaaaa/doc_accessibility?tab=readme-ov-file)) to build a software solution capable of processing offline documents (PDFs, Word, PowerPoint, PPT, etc.) to detect WCAG accessibility issues. | |
| Example document with non-text images: | |
|  | |
| Extracted Image: | |
|  | |
| Generated caption: | |
| "Indication of correct signature" |