This repository contains an NLP sentiment-analysis project based on a BERT fine-tuning workflow. The notebook trains a transformer sequence classifier for positive, negative, and neutral social-media sentiment.
The hosted Streamlit app is a portfolio-ready interactive dashboard. It lets visitors:
- analyze a single tweet/review-style text
- compare multiple texts in a batch view
- inspect positive and negative word signals
- understand the original BERT training workflow
- see what model artifacts are required for real BERT inference online
Run locally:
pip install -r requirements.txt
python -m streamlit run app.pyFor Streamlit Cloud:
Main file path: app.py
The original app loaded a fine-tuned BERT model from a local machine path:
models/sent_model
Those trained model/tokenizer artifacts are not included in this repository, so the hosted dashboard uses a transparent lexical analyzer instead of pretending to run unavailable BERT weights.
The original BERT app code is preserved as:
legacy_bert_app.py
.
|-- app.py # Streamlit Cloud dashboard
|-- legacy_bert_app.py # Original BERT inference app
|-- fine-tune-bert-model.ipynb # BERT fine-tuning notebook
|-- requirements.txt # Lightweight app dependencies
`-- README.md
The notebook covers:
- Loading a Twitter sentiment dataset.
- Cleaning and preparing tweet text.
- Encoding positive, negative, and neutral labels.
- Tokenizing text with
bert-base-uncased. - Fine-tuning
TFBertForSequenceClassification. - Evaluating the model on a test split.
- Saving tokenizer/model artifacts for deployment.
- Add the trained BERT tokenizer/model files or a reliable hosted model path.
- Replace the lexical analyzer with cached Hugging Face inference.
- Add test-set metrics directly to the Streamlit dashboard.
- Add confusion matrix and per-class precision/recall/F1.
- Add example social-media monitoring use cases.