Mumbai, India

Sagar Sanjay Pawar

Engineer · GATE 2026 AIR 1202

Trying to get good at the fundamentals.

I studied Electronics and Telecommunications, then spent two years building backend systems at TCS - mostly working on large-scale financial infrastructure. That experience gave me a sharper sense of what production software actually demands. Since then, I've been moving toward computer science more deliberately: taking the exams, reading more, and trying to work on problems that require real understanding rather than pattern-matching. NLP came in through curiosity - specifically around how language models handle languages they've barely seen. That question turned into a project.

Tata Consultancy Services 2023 – 2025

Systems Engineer

  • Worked on the backend of ICICI Bank's iMobile UPI system - a platform handling transactions at scale with little room for error.
  • Built and maintained REST APIs using Spring Boot, from initial design through deployment and post-production monitoring.
  • Took ownership of specific service flows end-to-end, including writing test plans, handling bug triage, and coordinating releases.
  • Responded to production issues under time pressure, which taught more about system behaviour than any documentation could.
  • Automated internal log-processing workflows, cutting down the kind of repetitive manual work that accumulates quietly in large teams.
  • Built tooling to streamline UAT cycles, reducing friction between development and validation phases.
Project Gora github ↗

Gormati - also called Banjari or Lambadi - is an Indo-Aryan language with millions of speakers across India. It belongs to the Banjara community and carries a distinct oral tradition. Despite this, it barely exists in modern NLP infrastructure.

This project looks at what happens when you run Gormati through standard multilingual models like Google's MuRIL. The tokenizer doesn't handle it well. Gormati text gets fragmented into far more tokens than equivalent text in well-represented languages.

A roughly 41.83% tokenization inefficiency gap - which translates directly into higher compute costs, lost context, and worse model performance.

The longer answer is about what this reveals structurally: which languages get built into the foundations of NLP systems, and which don't. The project documents these tokenization failures and makes a case for more deliberate language representation in multilingual model development.

Low-resource NLP Tokenization Language Representation MuRIL
View on GitHub
Log Cleaner Tool Internal Tooling

A utility that processes raw application logs and outputs structured JSON - built to reduce the manual effort involved in parsing unformatted log data during debugging and audits. Used internally during the TCS project lifecycle.

Adroit Security System Published · IJERT

An intrusion detection system built on a Raspberry Pi, using computer vision to identify and flag unauthorized access. Published in IJERT - an early experiment in applied computer vision that held together under real conditions.

GATE 2026

AIR 1202

PGEE 2025

AIR 309

TIFR 2025

Written Cleared

GATE 2025

Qualified

Languages

Python C C++ SQL

Tools

Git Postman Swagger Spring Boot

ML / NLP

Tokenization MuRIL Corpus Analysis

The work I find worth doing usually resists being finished quickly.

Understanding something properly takes longer than it looks like it should. That's not a problem to solve - it's just how it is.

I'm drawn to questions at the edges: where a system breaks, where a language disappears from the model, where the abstraction leaks.

Research and engineering feel like the same activity at a certain depth of effort.

The direction is toward research - specifically at the intersection of machine learning and language. Graduate study in CS is the immediate goal. I'm interested in low-resource language modelling and the structural questions underneath it: what does it take to represent a language well, and what goes wrong when you don't. The GATE and PGEE results are part of that path, not the end of it.

I watch a fair amount of cinema, sometimes something current worth the time. I read, though not always at the pace I'd like. History pulls me in more than most subjects: the way decisions stack up over decades and become something nobody planned for. I take photographs occasionally - nothing serious, just a habit of noticing light and framing. I go trekking when I can, which is less often than I'd prefer. Cricket is somewhere in the background, always. I move between Marathi, Hindi, and English depending on context - and each one genuinely feels a little different to think in.

Get in touch

Open to conversations about research, systems, language, and what comes next. Email is the best way.