
AI Engineer specializing in GPU-accelerated inference and production AI systems, with hands-on experience building low-latency, high-throughput model-serving platforms using NVIDIA Triton, TensorRT, and Kubernetes. Proven ability to take deep learning models from optimization to reliable, monitored, zero-downtime production deployments. IEEE-published researcher with a strong foundation in computer vision, deep learning architectures, and performance-driven AI infrastructure.