TEM-LLM / README.md
LabSmart's picture
Update README.md
5f8ef08 verified
|
Raw
History Blame Contribute Delete
2.63 kB
---
license: llama2
language:
- en
base_model:
- liuhaotian/llava-v1.5-7b
pipeline_tag: visual-question-answering
tags:
- tem
- materials-science
- multimodal
---
# TEM-LLM
TEM-LLM is a domain-specific multimodal large language model for
Transmission Electron Microscopy (TEM) image analysis, fine-tuned
from LLaVA-v1.5-7B using a difficulty-aware curriculum learning strategy.
This model is associated with the following work (currently under review):
> *AI-Assisted Materials Characterization: A Curriculum-Guided
> Multimodal Framework for Transmission Electron Microscopy*
---
## Model Details
| | |
|---|---|
| **Base Model** | LLaVA-v1.5-7B (Vicuna-v1.5-7B) |
| **Training Strategy** | Difficulty-Aware Curriculum Learning (4 stages) |
| **Training Data** | ~216K QA pairs across 40K TEM images |
| **Domain** | Transmission Electron Microscopy (TEM) |
| **Modalities** | CTEM, HR-TEM, STEM, Diffraction |
| **Fine-tuning Method** | LoRA (Low-Rank Adaptation) |
| **License** | LLaMA 2 |
---
## Intended Use
TEM-LLM is designed for:
- **Image Classification**: Identifying TEM imaging modalities
(CTEM, HRTEM, STEM, diffraction)
- **Microstructural Feature Detection**: Localizing and counting
structural elements such as nanoparticles and grain boundaries
- **Feature Analysis & Interpretation**: Providing scientific
explanations of observed microstructural phenomena
- **Comprehensive Image Description**: Integrating multiple aspects
of visual understanding into structured descriptions
---
## Important: Inference Requirements
TEM-LLM is built on LLaVA and **cannot be loaded directly via
`transformers`**. Inference requires the LLaVA repository.
**Step 1 — Clone LLaVA:**
```bash
git clone https://github.com/haotian-liu/LLaVA.git
cd LLaVA
pip install -e .
```
**Step 2 — Download weights:**
```python
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="LabSmart/TEM-LLM",
local_dir="./TEM-LLM"
)
```
**Step 3 — Run inference:**
```bash
python -m llava.serve.cli \
--model-path "./TEM-LLM" \
--image-file "path/to/your/tem_image.jpg" \
--load-4bit
```
For the complete pipeline including data crawling, preprocessing,
knowledge distillation, training, and evaluation scripts, please
refer to the [GitHub repository](https://github.com/SemiMIRTLab/TEM-LLM).
---
## Limitations
- TEM-LLM is optimized for TEM image analysis and may produce
unreliable outputs when applied to out-of-scope questions or
out-of-domain images.
- The model provides approximate qualitative descriptions rather
than exact quantitative measurements.