I’m an AI Engineer focused on building reliable GenAI, LLM evaluation, and multimodal AI systems that move beyond demos and work in real-world environments.
My work spans LLM-based agent evaluation, human-in-the-loop feedback pipelines, multimodal machine learning, and AI-powered product development. At Handshake, I evaluate multi-step LLM agent workflows by analyzing Plan + Code reasoning traces, identifying failure modes such as hallucinated outputs, data leakage, evaluation shortcuts, unsafe file operations, and misaligned actions. I focus on improving reasoning reliability, reproducibility, model alignment, and evaluation quality for production-scale AI research initiatives.
I’m also building an AI-powered financial decision simulator at Cheaha Infosys, designed to support human decision-making using contextual AI reasoning on real financial data. The system integrates a Machine Unlearning-based AI layer that adapts to individual user behavior and functions as a personalized AI decision assistant, currently in active MVP development.
My research background includes multimodal AI across vision, audio, video, text, and EEG. I developed an attention-based multimodal emotion recognition model combining EEG, audio, and video signals, achieving 91.2% accuracy and improving performance by 7.39% over baseline models. I have also worked on weakly supervised medical image segmentation using sparse annotations to improve segmentation robustness across medical and natural image datasets.
Alongside AI research, I bring strong engineering experience across Python, PyTorch, Hugging Face, Node.js, TypeScript, PostgreSQL, Supabase, Firebase, AWS Cognito, OAuth, SAML, Splunk, ServiceNow, and production debugging workflows.
I’m especially interested in roles where I can build trustworthy GenAI systems, evaluate LLM behavior, develop multimodal ML solutions, and turn AI research into practical products.