Twitter Sentiment Detector (Capstone)
End-to-end ML pipeline for social media sentiment classification.
Twitter Sentiment Detector (Capstone) was designed and built by Yadnyesh Mulay, an AI-first full-stack developer based in Nashik, India. Capstone project from the Skillship Foundation Aatmanirbhar Python Program. Complete ML pipeline: data collection via Twitter API, preprocessing, feature engineering (TF-IDF, n-grams), model training (Logistic Regression, Naive Bayes, SVM), evaluation, and deployment-ready serialization.
The problem
Social media sentiment analysis requires domain-specific preprocessing and model selection, generic models fail on tweet-specific noise (hashtags, mentions, abbreviations, emojis).
The approach
Tweet-optimized preprocessing pipeline + ensemble of classical ML models with rigorous evaluation methodology.
How it was built
This capstone project demonstrates the full ML lifecycle: raw tweet collection → cleaning (URL removal, mention handling, emoji processing) → feature extraction with TF-IDF vectorization including bi-grams and tri-grams → training multiple classifiers with cross-validation → hyperparameter tuning via GridSearchCV → model comparison with ROC curves and confusion matrices → joblib serialization for production deployment. Includes Jupyter notebooks documenting each phase with visualizations.
Where the AI sits
Classical ML (Logistic Regression, Naive Bayes, SVM) with TF-IDF features, pre-LLM era approach. Valuable for understanding feature engineering fundamentals that transfer to modern embedding-based approaches.
Stack
- Python
- scikit-learn
- pandas
- NLTK
- Tweepy
- Matplotlib
- Seaborn
- Jupyter
Links
More work by Yadnyesh Mulay: all projects · about Yadnyesh Mulay · home