
Applied ML researcher and engineer with experience building production ML systems, LLM-based applications, retrieval-augmented pipelines, and scalable data-driven workflows. Strong hands-on background in Python, SQL, PyTorch, AWS, ranking systems, MLOps, and multimodal generative modeling, with research interests in interpretable AI, diffusion models, and enterprise GenAI systems. Currently pursuing a PhD focused on Generative AI and LLMs, following completion of a Master’s degree.
Skills
Languages: Python, SQL, Java, JavaScript, C, C#
Frameworks & Libraries: PyTorch, TensorFlow, Pandas, Scikit-learn, Flask, FastAPI, Django, React, Nodejs
Generative AI & LLMs: Retrieval-Augmented Generation (RAG), LLM prompting, HyDE query rewriting, multimodal LLM applications
Machine Learning: Diffusion models, graph neural networks (GNNs), multimodal learning, supervised learning, time-series forecasting
Data & Systems: Data preprocessing, feature engineering, distributed systems, large-scale data pipelines, search ranking (LTR, Solr)
Infrastructure & MLOps: AWS (EC2, S3, Lambda, SageMaker), Azure ML, GCP (BigQuery, Compute Engine), Docker, Kubernetes, SLURM, model serving
Big Data: Spark, Hadoop, PySpark, Databricks
Other: Git, REST APIs, microservices, system design, production ML deployment
NaturalOCEAN: Retrieval-Augmented LLMs for Personality Prediction
• Built a RAG pipeline (Llama-3.1-8B) to predict Big Five traits from text using hybrid retrieval (BM25 + dense + cross-encoder reranking) over psychologically grounded corpora.
• Evaluated prompting and retrieval strategies (HyDE rewriting, corpus fusion) across controlled baselines and external validation (n=406, leakage-free).
• Showed retrieval improves MSE over zero-shot, while frontier models outperform on classification accuracy, revealing a calibration vs decision tradeoff.
Semantic Personality Dataset and Controllable Facial Expression Generation
• Built a semantically interpretable dataset combining FACS Action Units, gaze, and head motion with Five-Factor personality annotations.
• Developed an attention-based diffusion model for controllable facial expression generation using speech and personality traits, achieving FID 0.6 and R² 0.77.
• Demonstrates strength in multimodal learning, generative AI, and quantitative model evaluation.
Multi-Modal GNN for Personality Trait Prediction
• Developed a graph neural network to fuse multimodal behavioral signals for personality trait prediction, achieving 93% accuracy.
• Applied attention mechanisms to improve interpretability across heterogeneous input features.
• Demonstrates experience in multimodal representation learning, predictive modeling, and applied AI system design.