Author: Nicolette Mtisi
Tools: Python · NLTK · scikit-learn · pandas · Matplotlib · Seaborn
This NLP project sorts airline customer tweets into positive, neutral or negative sentiment. Airlines get thousands of these messages every day, and an automatic classifier helps customer-experience teams spot complaints and trends without reading every tweet by hand.
Dataset: Twitter US Airline Sentiment (Kaggle), 14,640 tweets about six U.S. airlines. It is included here as Tweets.csv.
- Text cleaning: lowercasing and removing URLs, @mentions, #hashtags and punctuation.
- NLP preprocessing with NLTK: tokenization, stopword removal and lemmatization.
- Features: TF-IDF vectorization, keeping the top 5,000 terms.
- Model: Logistic Regression trained on an 80/20 train/test split.
- Evaluation: accuracy, a per-class classification report and a confusion matrix.
Accuracy: 80.0% on the held-out test set (2,928 tweets).
| Sentiment | Precision | Recall | F1 | Test Tweets |
|---|---|---|---|---|
| Negative | 0.82 | 0.94 | 0.88 | 1,889 |
| Neutral | 0.67 | 0.49 | 0.57 | 580 |
| Positive | 0.82 | 0.62 | 0.71 | 459 |
| Sentiment Distribution | Confusion Matrix |
|---|---|
![]() |
![]() |
Key takeaways
- Negative tweets make up about 63% of the data, and the model detects them very reliably (94% recall).
- Neutral is the hardest class. These tweets are often questions or plain statements that share wording with the other two classes.
- Example: "I love the friendly service on this airline!" is classified as positive.
git clone https://github.com/nic-stack/Twitter-Sentiment-Analysis.git
cd Twitter-Sentiment-Analysis
pip install -r requirements.txt
jupyter notebook sentiment_analysis.ipynbThe notebook downloads the NLTK resources it needs (punkt, stopwords, wordnet) on the first run.
| File | Description |
|---|---|
sentiment_analysis.ipynb |
Full pipeline: preprocessing, training, evaluation and charts |
Tweets.csv |
Twitter US Airline Sentiment dataset |
sentiment_distribution.png |
Class distribution chart |
confusion_matrix.png |
Test-set confusion matrix |
requirements.txt |
Python dependencies |
- Handle the class imbalance with class weights or resampling to improve neutral and positive recall.
- Compare against transformer models. A follow-up project fine-tuned DistilBERT on this task (model · live demo).
- 📝 Medium write-up
- 🌐 Portfolio · LinkedIn

