You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

HAR-Agent: Multilingual Multimodal Human Activity Recognition via Knowledge-Distilled LLM Reasoning Read Paper

Overview

HAR-Agent is a multimodal, multilingual human activity recognition (HAR) system that combines vision-language-audio perception with knowledge-distilled LLM reasoning. It achieves sophisticated activity understanding on consumer-grade hardware โ€” the best student model requires under 1 GB VRAM.

This repository contains all trained LoRA adapter weights for 12 fine-tuned models (2 teachers + 10 students) across two training paradigms, plus cached teacher logits and comprehensive evaluation results.

Key Results

Model Method Params Accuracy F1 (Macro) VRAM
IT-T-72B IT 72B 60.3% 0.621 ~40 GB
IT-S-1.5B IT 1.5B 41.4% 0.405 0.8 GB
IT-S-3B IT 3B 39.8% 0.388 1.7 GB
SFT-T-72B SFT 72B 42.6% 0.401 ~40 GB
Audio (multilingual) IT-S-1.5B + Whisper โ€” 89.2% 0.890 ~6 GB

Inverse scaling finding: The smallest IT student (1.5B) outperforms all larger students (3Bโ€“32B), making the most accessible model also the most capable.

Architecture

HAR-Agent uses a three-pathway multimodal architecture:

  1. Visual pathway: LLaVA-NeXT 7B generates natural language scene descriptions from video frames (17 uniformly sampled frames per clip), with Short-Term Memory (STM) and semantic deduplication (cosine similarity threshold 0.85โ€“0.92).
  2. Audio pathway: Whisper-medium transcribes spoken activity descriptions in any language, enabling multilingual HAR across 5+ languages.
  3. Text pathway: Direct text input for integration with existing systems.

All pathways converge to a unified text representation, which a single LLM reasoning module classifies into one of 14 activity classes.

Knowledge Distillation Pipeline

  • Teacher: Qwen2.5-72B-Instruct fine-tuned with QLoRA (4-bit NormalFloat)
  • Students: Qwen2.5-{32B, 14B, 7B, 3B, 1.5B}-Instruct distilled from the teacher
  • IT (Instruction Tuning): Preserves the generative CausalLM paradigm; distillation via full-vocabulary KL divergence in shared token space
  • SFT (Supervised Fine-Tuning): Adds a classification head; distillation via logit matching

IT outperforms SFT by 17โ€“28 percentage points at every model scale (all p < 0.0001).

Quick Start

Loading the Best Student (IT-S-1.5B)

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch

# Load base model
base_model = "Qwen/Qwen2.5-1.5B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(base_model)
model = AutoModelForCausalLM.from_pretrained(
    base_model,
    torch_dtype=torch.float16,
    device_map="auto",
)

# Load LoRA adapter
model = PeftModel.from_pretrained(model, "khashayargh/HAR-Agent/student_outputs_it_1_5b")

# Classify an activity from a scene description
prompt = """Based on the following scene description, classify the human activity.

Scene: A person is seen bending their knees and lowering their body onto a chair,
with their hands resting on the armrests for support.

Activity classes: Bending, CarryingObject, Cleaning, ClosingCan, Drinking,
LiftingObject, OpeningCan, PuttingDownObjects, Reaching, SittingDown,
StairsClimbingDown, StairsClimbingUp, StandingUp, Walking

The activity is:"""

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=10, temperature=0.3, top_p=0.95)
prediction = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(prediction.strip())  # โ†’ SittingDown

Loading the Teacher (IT-T-72B)

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel

# 4-bit quantisation for ~40 GB VRAM
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype="float16",
)

base_model = "Qwen/Qwen2.5-72B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
    base_model,
    quantization_config=bnb_config,
    device_map="auto",
)
model = PeftModel.from_pretrained(model, "khashayargh/HAR-Agent/teacher_outputs_it_72b")

Activity Classes (14)

Class Description
Bending Upper body forward flexion
CarryingObject Transporting an object while walking
Cleaning Wiping, sweeping, or tidying surfaces
ClosingCan Sealing a container with a lid
Drinking Raising a vessel to the mouth
LiftingObject Picking up an object from a surface
OpeningCan Removing a lid from a container
PuttingDownObjects Placing held objects onto a surface
Reaching Extending arm(s) toward a target
SittingDown Transitioning from standing to seated
StairsClimbingDown Descending a staircase
StairsClimbingUp Ascending a staircase
StandingUp Transitioning from seated to standing
Walking Forward locomotion at normal pace

Datasets

  • RHM-HAR (Herts HAR RobotView): 6,701 video clips (5,354 train / 1,347 val), 14 activities, stratified 80/20 split by class label
  • Toyota Smarthome: 186 samples, 8 mapped activity classes โ€” used for cross-domain generalisation evaluation
  • Multilingual Audio: 250 TTS samples (5 activities ร— 5 languages: English US, English UK, Chinese, Spanish, Farsi)

Training Details

All models use QLoRA (4-bit NormalFloat quantisation) with LoRA rank r=64, ฮฑ=128, dropout=0.05.

Model LR Epochs Batch KD Temp KD ฮฑ GPUs
IT teachers 5e-5 3 4 โ€” โ€” 4ร—A100 80GB
IT students 5e-5 3 4โ€“8 2.0 0.5 1โ€“4ร—A100
SFT teachers 2e-5 5 4 โ€” โ€” 4ร—A100 80GB
SFT students 2e-5 5 4โ€“8 2.0 0.5 1โ€“4ร—A100

Total training time for all 12 models: ~1 month on the UHHPC cluster (NVIDIA A100 80 GB).

Deployment Requirements

Configuration VRAM Latency Hardware
IT-T-72B + LLaVA 48 GB ~14s 4ร—A100 (research only)
IT-S-1.5B + LLaVA 9 GB ~12s RTX 3060 12GB
IT-S-1.5B + Whisper 6 GB ~0.6s RTX 3060 (multilingual audio)
IT-S-1.5B only (text) 0.8 GB ~0.2s Any GPU

Knowledge Distillation Metrics

Student TSA Cohen's ฮบ Compression CES
IT-S-1.5B 0.457 0.417 48ร— 0.95
IT-S-3B 0.454 0.412 24ร— 1.89
IT-S-14B 0.434 0.377 5.1ร— 8.43
SFT-S-7B (best SFT) 0.078 0.042 10.3ร— 0.76

Cross-Domain Generalisation (Toyota Smarthome)

Model RHM-HAR F1 Toyota F1 Drop
IT-T-72B 0.621 0.308 50.5%
IT-S-1.5B 0.405 0.221 45.4%
IT-S-14B 0.336 0.261 22.3%

Citation

@article{ghamati2026har,
  title={HAR-Agent: Multilingual Multimodal Activity Recognition via Knowledge-Distilled LLM Reasoning},
  author={Ghamati, Khashayar and Alashti, Mohammad Reza Shahabian and Fallahirahmatabadi, Ali and Zaraki, Abolfazl},
  year={2026},
  publisher={Authorea}
}

License

Apache 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support