Bhavya
Bhavya
Data Scientist

Bhavya

Transforming complex data into actionable insights through innovative machine learning solutions.

2+Experiences
6+Projects
6+Skills
01.

About

I'm a passionate data science undergraduate skilled in Python and AI/LLM development. I thrive on extracting insights from large datasets and enjoy collaborating with teams to deliver impactful, data-driven solutions. My experience in mentoring and technical guidance fosters a collaborative environment that drives success.

02.

Experience

Jun 2026 – Jul 2026

Data Science Intern · Zappify

  • Developed and deployed comprehensive data science workflows, encompassing data preprocessing, exploratory data analysis (EDA), and visualization utilizing Python, Pandas, and Seaborn.
  • Implemented and assessed machine learning models using Scikit-learn, applying standard classification metrics such as accuracy, precision, recall, and F1-score.
Nov 2025 – Present

Creative Head · The Empirical Society

  • Led a team of 10 designers for college technical events, enhancing social media engagement by approximately 35% while coordinating end-to-end design workflows across events, marketing, and editorial teams.
03.

Selected work

R

Resume Screening: ClassifierCheck out

Python · Scikit-learn · TF-IDF · NLTK · GridSearchCV
  • Constructed an NLP pipeline to classify 962 resumes into 25 categories using text cleaning and TF-IDF with bigrams (20,000 features).
  • Benchmarked 6 classifiers under a stratified 80/20 split; Linear SVM tuned via 3-fold GridSearchCV (C {0.01–10}) achieved optimal performance — results interpreted with caution given small test set (n=193); exported predictions to test_predictions.csv for auditability.
Python · Scikit-learn · TF-IDF · NLTK · GridSearchCV
M

Message Spam FilteringCheck out

Python · Scikit-learn · TF-IDF · Streamlit · Deployed Web App Link
  • Developed and deployed an NLP classifier on 5,572 SMS messages; transitioned from MultinomialNB to ComplementNB to address 87/13 class imbalance, enhancing spam recall from 72% to 89% at 97.4% test accuracy.
  • Engineered bigram TF-IDF features (ngram_range=(1,2)) to capture phrase-level spam patterns overlooked by unigrams; applied stratified train/test split with full-data refit before deployment.
Python · Scikit-learn · TF-IDF · Streamlit · Deployed Web App Link
C

Customer Churn PredictionCheck out

Python · Scikit-learn · StandardScaler · K-Fold CV
  • Developed a classification pipeline on telecom churn data; benchmarked Logistic Regression, Random Forest, and SVM under K-Fold CV with 80/20 class imbalance managed via class_weight='balanced' on Logistic Regression.
  • Achieved best model at 81% churn recall compared to 0% for a naive majority-class baseline scoring 83% accuracy — framed recall, not accuracy, as the critical metric for retention decisions.
Python · Scikit-learn · StandardScaler · K-Fold CV
A

Airbnb Rental Price PredictionCheck out

Python · Scikit-learn · Pandas · GridSearchCV
  • Executed an end-to-end regression pipeline on 27,379 NYC listings; compared Linear Regression vs Random Forest (5-fold CV: 44.6% vs 54.5% R²); diagnosed ~100% train / 56.2% test R² gap as a feature ceiling due to missing key variables.
  • Eliminated price outliers via IQR (25,720 rows), applied LabelEncoder for tree-compatible encoding, and utilized hexbin actual-vs-predicted plots instead of scatter to manage 25K-point overplotting.
Python · Scikit-learn · Pandas · GridSearchCV
B

Barbie – Voice-Activated AI AssistantCheck out

Groq LLaMA 3.3 70B · gTTS · NewsAPI · PyWhatKit
  • Built an end-to-end agentic AI voice assistant with real-time speech-to-text and LLM-powered NLP; integrated autonomous web control (YouTube, live news), achieving sub-3s round-trip interaction latency.
  • Implemented persona-driven prompt engineering maintaining consistent AI character behaviour across 50+ distinct
  • conversation scenarios; integrated live NewsAPI and autonomous task execution — directly analogous to AI Copilot and Agentic AI architectures.
Groq LLaMA 3.3 70B · gTTS · NewsAPI · PyWhatKit
A

Automated AI ChatbotCheck out

Python · Groq API · LLaMA 3.3 · Prompt Engineering · LLM Integration
  • Engineered a production-ready conversational AI system using Groq-hosted LLaMA 3.3 70B, implementing structured prompt engineering pipelines with multi-turn dialogue management across 10+ exchange threads without memory degradation.
  • Designed async LLM inference architecture achieving sub-2s average response latency; built modular pipeline supporting plug-and-play LLM backend swapping — demonstrating production-grade system design thinking.
  • Implemented automated fallback and error-recovery logic; reduced hallucination rate ~40% through iterative output validation and structured response formatting.
Python · Groq API · LLaMA 3.3 · Prompt Engineering · LLM Integration
04.

Skills & tools

LanguagesAI/MLFrameworks/LibrariesDatabasesTools & PlatformsOther Technical Skills
05.

Education

Guru Gobind Singh Indraprastha University – GTB4CEC

B.Tech – Computer Science & Engineering (Data Science)
2024 – 2028 · CGPA: 8.90 / 10
06.

Leadership & activities

Nov 2025 – Present

Creative Head · The Empirical Society

07.

Certifications

GenAI-Powered Data Analytics

Tata iQ · 2026

Data Visualization

TCS · 2026
08.

Get in touch

Have a role or a project in mind? I'd love to hear from you.

Made with TailorCVbuild yours free