RoBERTa-large span extractor for Tweet Sentiment Extraction
Status: trained, evaluated, and released for reproducibility.
This repository is reserved for the EE5438 Group 36 baseline. Given a tweet and its supplied sentiment label, the planned model will score start and end positions independently and return a contiguous span copied from the tweet by character offset.
The experiment uses one fixed split of the Kaggle training set: 21,984 train, 2,748 validation, and 2,748 local-test examples after excluding one record with missing text. The local test will remain untouched until model selection is frozen.
The three-seed validation scores are 0.69897, 0.70521, and 0.70338 (mean 0.70252, sample standard deviation 0.00321). The selected seed-42 reference scores 0.70283 on the held-out local test. These are local split measurements, not Kaggle leaderboard claims.
The released artifact is the selected seed-42 checkpoint. The parent checkpoint is FacebookAI/roberta-large; preprocessing uses the fixed split and character-offset contract documented in the project repository. Reported scores are local split measurements, not Kaggle leaderboard claims. The upstream parent license and terms apply.
Project code and plan: https://github.com/ELVISIO119/ee5438-group36-tweet-sentiment-extraction
- Downloads last month
- 14