Keerthiram Murugesan
Research Scientist, IBM Research.
firstname.lastname@ibm.com
IBM Research AI
Hi, welcome to my personal webpage.
I am a Research Scientist at the IBM Thomas J. Watson Research Center, where I work at the intersection of Artificial Intelligence, Machine Learning, and Natural Language Understanding. My research is motivated by an increasingly important question: How can we build AI systems that remain reliable, interpretable, and trustworthy as they learn, adapt, and act in dynamic real-world environments? I received my Ph.D. from the School of Computer Science at Carnegie Mellon University. My current research spans four interconnected themes: 1) LLM safety and governance, building safeguards, detectors, and unlearning mechanisms to make LLMs safe and policy-compliant (e.g., Granite Guardian); 2) agentic AI and multi-agent systems, designing reliable agents that plan, reason, and collaborate across complex workflows; 3) cognitive architectures for reasoning — bridging fast/intuitive and slow/deliberate reasoning in AI through neuro-symbolic and metacognitive approaches; and 4) privacy and transparency in AI, ensuring AI systems protect user privacy and produce interpretable, factually grounded outputs.
My research has resulted in more than 50 publications in leading AI conferences, including NeurIPS, ICML, ICLR, ACL, EMNLP, AAAI, IJCAI, KDD, and other premier venues. My publications reflect a sustained effort to bridge learning theory, reasoning, and governance in AI systems. Several of these research contributions have been integrated into IBM products such as IBM Granite & Granite Guardian and watsonx, open-source project such as Mellea enabling research ideas to influence real-world deployments at scale.
I work with researchers across academia and industry and serve as a principal investigator or co-PI on several university partnerships. These collaborations allow me to pursue long-term scientific questions while ensuring that the resulting technologies address practical challenges faced by organizations deploying AI today. I also actively mentor students on topics spanning trustworthy AI, foundation models, reasoning, and agentic systems. I am always happy to collaborate with students and external researchers working on related problems. Feel free to reach out if you’d like to discuss research or collaborations.
Selected publications
- STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language ModelsACL 2024 Findings, 2024
- Granite Guardian: Comprehensive LLM SafeguardingIn 2025 Annual Conference of the North American Chapter of the Association for Computational Linguistics Industry Track, 2025
- Thinking Fast and Slow in Human and Machine IntelligenceCommunications of the ACM, 2025
- Language Models Coupled with Metacognition Can Outperform Reasoning ModelsarXiv preprint arXiv:2508.17959, 2025
- Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational AgentsIn Findings of the Association for Computational Linguistics: ACL 2025, 2025
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language ModelsarXiv preprint arXiv:2511.08484, 2025