Summary
Overview
Work History
Education
Skills
Websites
Timeline
Generic

DIKSHITHA AKULA

Baltimore,MD

Summary

Data Scientist with nearly 4 years of experience, Conceptualized Python, SQL, PySpark, and machine learning. Prepared NLP, deep learning, predictive modeling, and deployed models on AWS, Azure, and GCP. Strong in data visualization using Tableau and Power BI, building ETL pipelines, and collaborating with cross-functional teams to deliver actionable, data-driven solutions to improve business decisions and operational efficiency.

Overview

6
6
years of professional experience

Work History

Data Scientist

McKinsey & Company
12.2024 - Current
  • Streamlined content-based recommendation algorithms using NLP techniques on 500k+ content records, improving personalization precision by 10%, and leading to a 2% churn reduction.
  • Automated content data cleaning processes using Python and Pandas, reducing manual processing time by 6 hours weekly and enhancing data accuracy to 99.9%, directly improving recommendation model precision by 8%.
  • Applied NLP and deep learning to classify content and summarize descriptions, improving search accuracy by 15%.
  • Built interactive dashboards in Tableau and Power BI, reducing manual reporting time by 40% for the analytics team.
  • Collaborated with product and engineering teams to implement actionable solutions supporting content strategy and user engagement growth.
  • Systematized model pipelines, decreasing runtime by 30%, and minimizing manual intervention across machine learning workflows.

Junior Data Scientist

Zoho Corporation
08.2021 - 07.2023
  • Developed predictive models for finance and supply chain, increasing forecast accuracy by 20%, enabling better operational decisions.
  • Mechanized ETL pipelines using Python and Airflow, reducing manual data handling time by 35% across client projects.
  • Conducted feature engineering and trained regression, classification, and forecasting models on datasets exceeding 200k records.
  • Created Power BI and Excel dashboards for clients, helping visualize KPIs and reducing decision-making time by 25%.
  • Coordinated Implemented alongside diverse teams to execute projects solutions aligned with client requirements, timelines, and business goals.
  • Documented model workflows and assumptions to maintain knowledge sharing, improving team efficiency for future projects.

Data Scientist Intern

Genpact
07.2020 - 07.2021
  • Mechanized five data validation checks within the preprocessing pipeline using Python, reducing data discrepancies by 20% and saving ~10 hours per week for the team.
  • Systematized reporting tasks using Python and Excel VBA, saving 10+ hours weekly and reducing manual errors.
  • Extracted, cleansed, and validated over 100,000 data records, ensuring data integrity and consistency for input into machine learning models, resulting in a 3% increase in model accuracy.
  • Co-ordinated in deploying ML models on AWS and Azure cloud platforms, improving scalability and accessibility.
  • Designed dashboards and visualizations to communicate insights clearly to business stakeholders and management teams.
  • Guided in team discussions, providing findings and recommendations to improve project results and operational efficiency.

Education

Master’s - Data Science

University of Maryland, Baltimore County (UMBC)
05.2025

Bachelor’s - Technology

Jawaharlal Nehru Technological University
07.2021

Skills

  • Programming & Data Science: Python, R, PySpark, MATLAB, Scikit-learn, TensorFlow, PyTorch, Pandas, NumPy
  • Data Warehousing & Querying: Snowflake, BigQuery, PostgreSQL, MySQL, SQL Server, MongoDB, Cassandra, Data Lake, ETL Pipelines
  • AI & Machine Learning: NLP, Deep Learning, Time Series Analysis, Supervised Learning, Unsupervised Learning, Recommendation Engines, RAG, Topic Modeling, Text Summarization
  • Generative AI & LLM Tools: OpenAI APIs, Hugging Face Transformers, LangChain, Llamalndex, Vector Databases (FAISS, Weaviate), Prompt Engineering, Fine-tuning LLMs, Llama 2
  • Cloud Computing: AWS (S3, EC2, SageMaker), Azure (Azure ML, Cognitive Services), Google Cloud (Vertex AI, BigQuery), Databricks, Serverless Computing, Cloud Infrastructure Management, IAM
  • Data Engineering & Orchestration: Hadoop, Spark, Airflow (ETL & Orchestration), Docker, Kubernetes, Kafka, Data Governance, Data Quality
  • Visualization & Reporting: Tableau, Power BI, Looker, Seaborn, Matplotlib, Data Storytelling, Interactive Dashboards
  • Statistical Analysis: Regression Analysis, Hypothesis Testing, A/B Testing, Experiment Design, Statistical Modeling, Causal Inference, Bayesian Analysis

Timeline

Data Scientist

McKinsey & Company
12.2024 - Current

Junior Data Scientist

Zoho Corporation
08.2021 - 07.2023

Data Scientist Intern

Genpact
07.2020 - 07.2021

Bachelor’s - Technology

Jawaharlal Nehru Technological University

Master’s - Data Science

University of Maryland, Baltimore County (UMBC)
DIKSHITHA AKULA