AI Research Engineer · NPCI, Mumbai

Teaching language models to act reliably.

I work on post-training LLMs — preference optimisation and RL with verifiable rewards — building RL environments for tool-using agents, and evaluating whether models stay safe when real money is on the line.

Portrait of Nitin Kukreja
3
arXiv papers (2026)
24B
params in FiMI, India's financial LLM
52→80%
out-of-scope refusal after RL post-training
799
multi-turn cases in IndicBankBench
§1

Research

I'm on the AI Research team at the National Payments Corporation of India, the organisation behind UPI. My work sits where current methods meet a hard constraint: models that take actions on people's money must be correct, cautious and auditable. I contribute to FiMI, NPCI's 24B financial LLM, and to research on safe tool-using banking agents — and I take that work all the way to production.

Post-training

SFT, preference optimisation and RL with verifiable rewards; what each buys at the response vs. trajectory level.

Environments for agents

Deterministic simulators, verifiers and reward design for multi-turn tool use with confirmation-gated actions.

Evaluation

Trajectory-level, mostly deterministic checks for safety and reliability — beyond final-answer LLM judging — in Indic, multilingual settings.

§2

News

§3

Selected results

Out-of-scope refusal52% → 80%
Edge-case score0.509 → 0.718
Order-sensitive score0.590 → 0.679
Tokens used (lower is better)100% → 71%
before post-trainingafter
Fig. 1Post-training a banking agent with preference optimisation and verifiable-reward RL improves safety and reliability while using 29% fewer tokens. Numbers from FiMI Banking (arXiv:2609.03960).

On FiMI, domain adaptation (CPT → IFT → SFT on 68B curated English/Hindi/Hinglish tokens, 64 H100s) gave +20% on finance reasoning and +87% on domain tool calling over the Mistral Small 24B base (arXiv:2602.05794).

§4

Papers

arXiv2026

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

S. Paul, C. Bhushan, H. Sharma, N. Kukreja, Y. Dedhia, K. Doshi, P. Devadiga

799 cases · 20 axes · four-stage trajectory scoring (safety, tool use, adequacy, advisory) with deterministic gates · 11 models from 8 families.

@article{paul2026indicbankbench,
  title   = {IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking},
  author  = {Paul, Suvradip and Bhushan, Chandra and Sharma, Harsh and Kukreja, Nitin and Dedhia, Yatharth and Doshi, Keyur and Devadiga, Prashant},
  journal = {arXiv preprint arXiv:2609.29167},
  year    = {2026}
}
arXiv2026

FiMI Banking: A Sovereign Model for Indian Retail Banking

NPCI AI Research Team, A. Kumar, … , N. Kukreja, … (18 authors, alphabetical)

Controlled retail-banking setting from vetted documents, ground truth and tools; compares preference optimisation with verifiable-reward RL.

@article{npci2026fimibanking,
  title   = {FiMI Banking: A Sovereign Model for Indian Retail Banking},
  author  = {{NPCI AI Research Team} and Kumar, Aman and others},
  journal = {arXiv preprint arXiv:2609.03960},
  year    = {2026}
}
arXiv2026

FiMI: A Domain-Specific Language Model for Indian Finance Ecosystem

A. Kathar, A. Kumar, … , N. Kukreja, … (40 authors, alphabetical)

Mistral Small 24B adapted via continued pre-training on 68B tokens, instruction tuning and tool-centric SFT for Indian digital payments.

@article{kathar2026fimi,
  title   = {FiMI: A Domain-Specific Language Model for Indian Finance Ecosystem},
  author  = {Kathar, Aboli and Kumar, Aman and others},
  journal = {arXiv preprint arXiv:2602.05794},
  year    = {2026}
}
§5

Research projects

Banking Gym

RL environment · NPCI

A deterministic, NeMo Gym-compatible environment for training bank agents: stateful bank simulator, confirmation-gated writes, verifier-graded end-state rewards, parallel isolated rollouts.

RLVRverifierstool use

FiMI post-training

24B LLM · NPCI

Led post-training and agentic integration: multi-stage CPT → IFT → SFT, synthetic tool-calling data, distributed training with DeepSpeed ZeRO-3 and FlashAttention-2.

64×H10068B tokensHindi/Hinglish

IndicBankBench

evaluation · open source

A benchmark that checks the whole trajectory — wrong account, stale context, invalid writes — not just the final answer.

799 cases20 axesdeterministic gates

From paper to production

UPI Help · MyUPI

FiMI runs as the multi-agent backend (LangGraph, MCP) behind UPI dispute resolution and mandate management in 5 languages, with RAG grounding and guardrails.

multi-agentRAG / Graph-RAGguardrails
§6

Experience

National Payments Corporation of India (NPCI)

Mumbai · Sep 2023 – present

Senior Associate Data Scientist (Apr 2025 –), AI Research team: LLM post-training, RL for agents, evaluation. Earlier, as Associate Data Scientist: GNN fraud models scoring 15,000 transactions/sec at 20 ms.

Xoriant Solutions

Pune · Aug 2021 – Sep 2023

Software Engineer: an enterprise LLM chatbot and document-AI systems (LiLT/BERT entity extraction, NER).

R. C. Patel Institute of Technology

Shirpur · 2017 – 2021

B.Tech, Computer Science & Engineering.

§7

Open source & teaching

BibTeX copied