A decade of backend engineering, now aimed at one problem: taking LLM systems from demo to production in environments where failure is expensive — payments, fintech, enterprise scale.
I'm a backend engineer who moved to the GenAI frontier about two years ago — and stayed, because the hardest problems in AI right now aren't the models, they're the systems around the models: retrieval, evaluation, observability, guardrails, scale.
At Oracle Cloud Infrastructure I helped build enterprise AI services used by over 170,000 people. Now at Visa, I work on GenAI in the payments domain — where "mostly correct" isn't a passing grade.
Off the clock: long drives, Himalayan hikes, One Piece, and building a small LLM from scratch just to know every layer of the thing I work with.
The mountains are where I reset. The name — handle @himalyanmonk everywhere — is a reminder to build with focus and without noise. Yes, the spelling is intentional.
GenAI systems in the payments domain — building AI that holds up under fintech-grade correctness, compliance, and scale requirements.
Core engineer on an enterprise GenAI platform scaled from 1,500 to 170,000+ users. Built agentic workflows with LangGraph, RAG pipelines, real-time transcription with speaker diarization, and a PII scrubbing layer — working across 5+ product teams in 3 time zones.
Backend engineering for travel-tech systems operating at global scale.
Low-level systems work in cybersecurity — where I learned to respect what happens under the abstractions.
Where it started.
Core services for an internal AI platform — RAG pipelines, agent orchestration, and the unglamorous reliability work that let it grow 100x.
1.5K → 170K usersLangGraph StateGraph with a supervisor agent for intent routing, MCP integration over Microsoft Graph, conditional edges, and SSE streaming — an agent that actually does the work.
FastAPI · LangGraph · MCPAn intelligence layer that reads the flood so you don't have to — LLM-driven filtering and summarization that cuts ~90% of the noise.
~90% noise filteredMiddleware metrics and SQL-view analytics giving product teams real visibility into API behavior — built after debugging a production connection-pool exhaustion the hard way.
Production-hardenedBuilding a small language model layer by layer — attention, KV cache, the lot — because using transformers isn't the same as understanding them.
A passion project exploring AI-based early detection of cardiac conditions from ECG signals. Long game, high stakes, worth it.