Keerthiram Murugesan

Research Scientist, IBM Research.

keerti.png

firstname.lastname@ibm.com

IBM Research AI

Hi, welcome to my personal webpage.

I am a Research Scientist at the IBM Thomas J. Watson Research Center, where I work at the intersection of Artificial Intelligence, Machine Learning, and Natural Language Understanding. My research is motivated by an increasingly important question: How can we build AI systems that remain reliable, interpretable, and trustworthy as they learn, adapt, and act in dynamic real-world environments? I received my Ph.D. from the School of Computer Science at Carnegie Mellon University. My current research spans four interconnected themes: 1) LLM safety and governance, building safeguards, detectors, and unlearning mechanisms to make LLMs safe and policy-compliant (e.g., Granite Guardian); 2) agentic AI and multi-agent systems, designing reliable agents that plan, reason, and collaborate across complex workflows; 3) cognitive architectures for reasoning — bridging fast/intuitive and slow/deliberate reasoning in AI through neuro-symbolic and metacognitive approaches; and 4) privacy and transparency in AI, ensuring AI systems protect user privacy and produce interpretable, factually grounded outputs.

My research has resulted in more than 50 publications in leading AI conferences, including NeurIPS, ICML, ICLR, ACL, EMNLP, AAAI, IJCAI, KDD, and other premier venues. My publications reflect a sustained effort to bridge learning theory, reasoning, and governance in AI systems. Several of these research contributions have been integrated into IBM products such as IBM Granite & Granite Guardian and watsonx, open-source project such as Mellea enabling research ideas to influence real-world deployments at scale.

I work with researchers across academia and industry and serve as a principal investigator or co-PI on several university partnerships. These collaborations allow me to pursue long-term scientific questions while ensuring that the resulting technologies address practical challenges faced by organizations deploying AI today. I also actively mentor students on topics spanning trustworthy AI, foundation models, reasoning, and agentic systems. I am always happy to collaborate with students and external researchers working on related problems. Feel free to reach out if you’d like to discuss research or collaborations.

Selected publications

  1. Text-based rl agents with commonsense knowledge: New challenges, environments and baselines
    Keerthiram Murugesan, Mattia Atzeni, Pavan Kapanipathi, and 6 more authors
    In Proceedings of the AAAI Conference on Artificial Intelligence, 2021
  2. ACL
    STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language Models
    Shreyas Basavatia, Keerthiram Murugesan, and Shivam Ratnakar
    ACL 2024 Findings, 2024
  3. Granite Guardian: Comprehensive LLM Safeguarding
    Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia, and 19 more authors
    In 2025 Annual Conference of the North American Chapter of the Association for Computational Linguistics Industry Track, 2025
  4. Thinking Fast and Slow in Human and Machine Intelligence
    Francesco Fabiano, Marianna B Ganapini, Andrea Loreggia, and 6 more authors
    Communications of the ACM, 2025
  5. Language Models Coupled with Metacognition Can Outperform Reasoning Models
    Vedant Khandelwal, Francesca Rossi, Keerthiram Murugesan, and 4 more authors
    arXiv preprint arXiv:2508.17959, 2025
  6. ACL
    Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
    Ivoline C Ngong, Swanand Kadhe, Hao Wang, and 4 more authors
    In Findings of the Association for Computational Linguistics: ACL 2025, 2025
  7. Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
    Huzaifa Arif, Keerthiram Murugesan, Ching-Yun Ko, and 3 more authors
    arXiv preprint arXiv:2511.08484, 2025