AI Research Engineer · NPCI, Mumbai
Teaching language models to act reliably.
I work on post-training LLMs — preference optimisation and RL with verifiable rewards — building RL environments for tool-using agents, and evaluating whether models stay safe when real money is on the line.

Research
I'm on the AI Research team at the National Payments Corporation of India, the organisation behind UPI. My work sits where current methods meet a hard constraint: models that take actions on people's money must be correct, cautious and auditable. I contribute to FiMI, NPCI's 24B financial LLM, and to research on safe tool-using banking agents — and I take that work all the way to production.
Post-training
SFT, preference optimisation and RL with verifiable rewards; what each buys at the response vs. trajectory level.
Environments for agents
Deterministic simulators, verifiers and reward design for multi-turn tool use with confirmation-gated actions.
Evaluation
Trajectory-level, mostly deterministic checks for safety and reliability — beyond final-answer LLM judging — in Indic, multilingual settings.
News
- paperIndicBankBench is out — a benchmark for the safety and reliability of banking assistants, with code.
- paperFiMI Banking — preference optimisation vs. verifiable-reward RL for safe bank agents.
- launchMyUPI, built on FiMI, goes live on the BHIM app in 5 languages.
- paperFiMI technical report — a 24B domain-specific LLM for India's finance ecosystem.
- launchUPI Help, powered by FiMI, launched at Global Fintech Fest 2025.
- featureReal-time anomaly-detection work featured by NVIDIA at SuperComputing '24.
Selected results
On FiMI, domain adaptation (CPT → IFT → SFT on 68B curated English/Hindi/Hinglish tokens, 64 H100s) gave +20% on finance reasoning and +87% on domain tool calling over the Mistral Small 24B base (arXiv:2602.05794).
Papers
IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
799 cases · 20 axes · four-stage trajectory scoring (safety, tool use, adequacy, advisory) with deterministic gates · 11 models from 8 families.
@article{paul2026indicbankbench,
title = {IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking},
author = {Paul, Suvradip and Bhushan, Chandra and Sharma, Harsh and Kukreja, Nitin and Dedhia, Yatharth and Doshi, Keyur and Devadiga, Prashant},
journal = {arXiv preprint arXiv:2609.29167},
year = {2026}
}
FiMI Banking: A Sovereign Model for Indian Retail Banking
Controlled retail-banking setting from vetted documents, ground truth and tools; compares preference optimisation with verifiable-reward RL.
@article{npci2026fimibanking,
title = {FiMI Banking: A Sovereign Model for Indian Retail Banking},
author = {{NPCI AI Research Team} and Kumar, Aman and others},
journal = {arXiv preprint arXiv:2609.03960},
year = {2026}
}
FiMI: A Domain-Specific Language Model for Indian Finance Ecosystem
Mistral Small 24B adapted via continued pre-training on 68B tokens, instruction tuning and tool-centric SFT for Indian digital payments.
@article{kathar2026fimi,
title = {FiMI: A Domain-Specific Language Model for Indian Finance Ecosystem},
author = {Kathar, Aboli and Kumar, Aman and others},
journal = {arXiv preprint arXiv:2602.05794},
year = {2026}
}
Research projects
Banking Gym
A deterministic, NeMo Gym-compatible environment for training bank agents: stateful bank simulator, confirmation-gated writes, verifier-graded end-state rewards, parallel isolated rollouts.
FiMI post-training
Led post-training and agentic integration: multi-stage CPT → IFT → SFT, synthetic tool-calling data, distributed training with DeepSpeed ZeRO-3 and FlashAttention-2.
IndicBankBench
A benchmark that checks the whole trajectory — wrong account, stale context, invalid writes — not just the final answer.
From paper to production
FiMI runs as the multi-agent backend (LangGraph, MCP) behind UPI dispute resolution and mandate management in 5 languages, with RAG grounding and guardrails.
Experience
National Payments Corporation of India (NPCI)
Senior Associate Data Scientist (Apr 2025 –), AI Research team: LLM post-training, RL for agents, evaluation. Earlier, as Associate Data Scientist: GNN fraud models scoring 15,000 transactions/sec at 20 ms.
Xoriant Solutions
Software Engineer: an enterprise LLM chatbot and document-AI systems (LiLT/BERT entity extraction, NER).
R. C. Patel Institute of Technology
B.Tech, Computer Science & Engineering.