Comprehensive Sentiment Analysis on Text Data
Overview
This project implements a comprehensive sentiment analysis system for textual data, focusing on movie reviews. The system uses deep learning techniques, specifically Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units, to classify the sentiment of text as positive or negative.
The project includes data preprocessing, model architecture design, training, evaluation, and deployment considerations for sentiment classification.
Dataset
The dataset consists of movie reviews with sentiment labels:
Data Files
movie_train.jsonl: Training set with movie reviews and sentiment labelsmovie_dev.jsonl: Development/validation setmovie_test.jsonl: Test set for final evaluationmovie.jsonl: Complete dataset
Data Structure
Each entry in JSONL format contains:
text: The movie review textlabel: Sentiment label (0 for negative, 1 for positive)id: Unique identifier for the review
Sample Data
{"text": "This movie was excellent! The acting was superb and the plot kept me engaged throughout.", "label": 1, "id": "12345"}
{"text": "Terrible film. Waste of time and money. The story made no sense.", "label": 0, "id": "67890"}
Dependencies
- Python 3.x
- PyTorch: Deep learning framework
- torchtext: Text processing utilities for PyTorch
- pandas: Data manipulation
- numpy: Numerical operations
- scikit-learn: Evaluation metrics
- matplotlib: Visualization
- tensorboard: Experiment tracking
Installation
- Install PyTorch (adjust for your system):
pip install torch torchvision torchaudio
- Install other dependencies:
pip install torchtext pandas numpy scikit-learn matplotlib tensorboard
Usage
Data Preprocessing
- Load and preprocess the data:
import pandas as pd
import torch
from torchtext.data.utils import get_tokenizer
from torchtext.vocab import build_vocab_from_iterator
# Load data
train_df = pd.read_json('data/movie_train.jsonl', lines=True)
dev_df = pd.read_json('data/movie_dev.jsonl', lines=True)
test_df = pd.read_json('data/movie_test.jsonl', lines=True)
# Tokenization
tokenizer = get_tokenizer('basic_english')
def yield_tokens(data_iter):
for text in data_iter:
yield tokenizer(text)
# Build vocabulary
vocab = build_vocab_from_iterator(yield_tokens(train_df['text']), specials=["<unk>"])
vocab.set_default_index(vocab["<unk>"])
- Create data loaders:
from torch.utils.data import DataLoader
from torch.nn.utils.rnn import pad_sequence
def collate_batch(batch):
label_list, text_list = [], []
for (_text, _label) in batch:
label_list.append(_label)
processed_text = torch.tensor(vocab(tokenizer(_text)), dtype=torch.int64)
text_list.append(processed_text)
return pad_sequence(text_list, padding_value=0.0), torch.tensor(label_list)
# Create datasets
train_dataset = list(zip(train_df['text'], train_df['label']))
dev_dataset = list(zip(dev_df['text'], dev_df['label']))
# Data loaders
train_loader = DataLoader(train_dataset, batch_size=64, shuffle=True, collate_fn=collate_batch)
dev_loader = DataLoader(dev_dataset, batch_size=64, shuffle=False, collate_fn=collate_batch)
Model Architecture
Implement LSTM-based sentiment classifier:
import torch.nn as nn
class SentimentLSTM(nn.Module):
def __init__(self, vocab_size, embedding_dim, hidden_dim, output_dim, n_layers, dropout):
super().__init__()
self.embedding = nn.Embedding(vocab_size, embedding_dim)
self.lstm = nn.LSTM(embedding_dim, hidden_dim, n_layers, dropout=dropout, batch_first=True)
self.dropout = nn.Dropout(dropout)
self.fc = nn.Linear(hidden_dim, output_dim)
def forward(self, text):
embedded = self.dropout(self.embedding(text))
output, (hidden, cell) = self.lstm(embedded)
hidden = self.dropout(hidden[-1])
return self.fc(hidden)
# Model parameters
vocab_size = len(vocab)
embedding_dim = 100
hidden_dim = 256
output_dim = 1
n_layers = 2
dropout = 0.5
model = SentimentLSTM(vocab_size, embedding_dim, hidden_dim, output_dim, n_layers, dropout)
Training
Train the model with appropriate loss function and optimizer:
import torch.optim as optim
# Loss and optimizer
criterion = nn.BCEWithLogitsLoss()
optimizer = optim.Adam(model.parameters())
# Training loop
def train(model, iterator, optimizer, criterion):
model.train()
epoch_loss = 0
for texts, labels in iterator:
optimizer.zero_grad()
predictions = model(texts).squeeze(1)
loss = criterion(predictions, labels.float())
loss.backward()
optimizer.step()
epoch_loss += loss.item()
return epoch_loss / len(iterator)
# Run training
n_epochs = 10
for epoch in range(n_epochs):
train_loss = train(model, train_loader, optimizer, criterion)
print(f'Epoch {epoch+1}: Train Loss: {train_loss:.3f}')
Evaluation
Evaluate model performance on validation and test sets:
from sklearn.metrics import classification_report, accuracy_score
def evaluate(model, iterator):
model.eval()
predictions, true_labels = [], []
with torch.no_grad():
for texts, labels in iterator:
preds = torch.sigmoid(model(texts).squeeze(1)) > 0.5
predictions.extend(preds.cpu().numpy())
true_labels.extend(labels.cpu().numpy())
return predictions, true_labels
# Evaluate on dev set
dev_preds, dev_labels = evaluate(model, dev_loader)
print(f'Accuracy: {accuracy_score(dev_labels, dev_preds):.3f}')
print(classification_report(dev_labels, dev_preds))
Model Saving and Loading
# Save model
torch.save(model.state_dict(), 'sentiment_model.pth')
# Load model
model.load_state_dict(torch.load('sentiment_model.pth'))
model.eval()
Model Details
Architecture
- Embedding Layer: Converts words to dense vectors
- LSTM Layers: Captures sequential dependencies in text (2 layers with dropout)
- Fully Connected Layer: Binary classification output
- Dropout: Prevents overfitting
Hyperparameters
- Embedding dimension: 100
- Hidden dimension: 256
- Number of LSTM layers: 2
- Dropout rate: 0.5
- Learning rate: 0.001 (Adam optimizer)
Results
Performance Metrics
- Accuracy: [X]% on test set
- Precision: [X]% for positive class
- Recall: [X]% for positive class
- F1-Score: [X]%
Training Logs
TensorBoard logs are available in the logs/ directory for monitoring training progress.
Analysis
- LSTM effectively captures long-range dependencies in reviews
- Dropout and proper regularization prevent overfitting
- Model performs well on longer reviews with clear sentiment
- Challenges with neutral or mixed sentiment reviews
Project Structure
A4-SentimentAnalysis/
├── README.md
├── LICENSE
├── .vscode/
│ └── settings.json
├── code/
│ └── report_rev1.ipynb # Implementation notebook
├── data/
│ ├── movie_dev.jsonl
│ ├── movie_test.jsonl
│ ├── movie_train.jsonl
│ └── movie.jsonl
├── logs/
│ ├── events.out.tfevents... # TensorBoard logs
│ └── [timestamp]/
│ └── events.out.tfevents...
├── report/
│ └── report.html # HTML report
└── Resources/
└── NLP_Spring1401_HW4.pdf # Assignment specification
Features
- LSTM-based sentiment classification
- Comprehensive data preprocessing pipeline
- TensorBoard integration for experiment tracking
- Evaluation metrics and confusion matrix analysis
- Model serialization for deployment
- Support for variable-length text inputs
Applications
- Movie review sentiment analysis
- Product review classification
- Social media sentiment monitoring
- Customer feedback analysis
- Text classification for various domains
Future Improvements
- Incorporate pre-trained embeddings (GloVe, Word2Vec)
- Use transformer-based models (BERT, RoBERTa)
- Add attention mechanisms
- Implement multi-class sentiment analysis
- Add model interpretability features
Contributors
- [Your Name]
- Course: Natural Language Processing, Sharif University of Technology
- Term: Spring 1401
License
This project is licensed under the MIT License - see the LICENSE file for details.
Xet Storage Details
- Size:
- 8.46 kB
- Xet hash:
- 75c7e72edb44ed9776696d3e0b48b74b69f3af1b3de831008c9e285111f58fad
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.