01. What I Build
I build LLM systems end to end: the agent workflows and retrieval a customer actually touches, the evals that say whether a change made things better, and the training and inference stack underneath. Master’s in Information Systems from Northeastern.
On the applied side that has meant working with a client to scope the problem, then shipping it: an LLM concierge and operations agent for a hotel, built in LangChain and LangGraph with RAG over property data, MCP connectors into their booking system and PMS so the agent can act rather than just answer, and an eval harness that runs on every prompt change.
Further down, on the training side, that has meant multi-GPU fine-tuning with FSDP, cluster orchestration on Kubernetes and Ray, and a three-cluster SkyPilot pipeline whose spot workers can be reclaimed mid-run without losing the run. On the inference side, serving with vLLM behind FastAPI, retargeting a stack from CUDA to ROCm, and writing the CUDA, HIP, and Triton kernels underneath when the framework leaves performance on the table.
The part I care most about is measurement. Whether it’s a kernel or an agent, most of the wins below started as a measurement that disagreed with someone’s assumption, and every number on this page names the baseline it was measured against.
I work in the open, upstream: a merged fix to the Mamba2 cached forward pass in HuggingFace transformers (PR #46084) and a libc refactor in LLVM (PR #175396).
The stack I work in
Agents & Applied AI
Training Stack
Inference & Kernels
Cluster & Platform
Also fluent in C++, C, Python, SQL, Shell, PySpark, and R; scikit-learn and XGBoost for the classical end.
02. Systems I’ve Built
Distributed RL Training Pipeline for GPU Kernel Generation
- Architected a three-cluster SkyPilot pipeline decoupling the trainer, the job queue (FastAPI + Redis), and a fleet of preemptible spot GPU workers, so a reclaimed node requeues its rollout instead of failing the run.
- Trained Qwen2.5-Coder-7B with GRPO via verl to optimize GPU kernels across NVIDIA and AMD targets, scoring rollouts on a speedup reward clipped at 4.0× so the policy could not farm unbounded timing gains.
- Built the timing-and-correctness harness that scores every candidate kernel, hardened against hardcoded outputs, timing-loop manipulation, and cache-warming artifacts.
Blackwell NVFP4 Tensor-Core Kernels — NVIDIA GB200 Challenge
Applied the newest hardware path the moment it shipped: block-scaled NVFP4 matrix operations on Blackwell B200, written in CuteDSL with inline PTX. 3–6× over the challenge reference kernel via SwiGLU epilogues and TMA memory pipelining — placing 42nd on GPUMODE’s NVFP4 Grouped GEMM leaderboard and 95th on Gated Dual GEMM.
Qwen600 Inference: CUDA → ROCm Port for AMD MI300X
Retargeted an inference stack to a second vendor: six transformer kernels ported from CUDA to ROCm/HIP and tuned for MI300X. Reached 12.1 ms/token decode on a single MI300X at 79–83% of peak HBM bandwidth, through bfloat16 coalesced access that cut memory traffic roughly in half.
Llama 3.1 Training & Inference in Pure C/CUDA
Owned the whole path end to end — complete forward and backward passes with attention, no framework underneath. Built on Karpathy’s llm.c, with hand-tuned RMSNorm (2.3× over the reference kernel), bfloat16 SwiGLU, and coalesced memory access.
RepositoryScaling Diffusion Training: Dask Preprocessing + AMP
- Profiled a from-scratch DDPM training run, found data preprocessing was the bottleneck rather than the GPU, and moved it onto a Dask worker pool for 1.48× CPU throughput scaling.
- Cut wall-clock training time 11% with mixed precision, and loss variance 29% with EMA weight averaging.
End-to-End SFT→DPO Pipeline on SageMaker
Carried one model the full distance from raw data to a served endpoint: ETL over 123K conversation samples, hybrid SFT-then-DPO training with sharding, and deployment on SageMaker. Evaluated at 71% macro-AUC with a 40% improvement in output consistency.
OCR Pipeline with Layout Detection
Deep-learning OCR over complex financial documents using ViT + CRNN. Reached 98% text-extraction accuracy and 42% faster processing through ONNX runtime and custom CUDA kernel optimization — the serving path that later held sub-100ms P99 in production.
Two-Stage Cardiac Arrhythmia Detection
A CNN-Transformer architecture over raw ECG signals, split into two stages: the first classifies overall rhythm at 94–98% accuracy, the second does adaptive ectopic-beat detection at 90–95% using dynamic patient-specific rate windowing rather than a fixed threshold.
Medical Knowledge-Graph RAG
A Graph RAG framework for medical retrieval: automated triplet extraction builds a knowledge graph in Neo4j dynamically, and semantic search over it through LangChain returns answers with their supporting context rather than a flat passage match.
RepositoryFintech SEC Data Platform
- Built a master financial database for US public companies with raw, transformed, and denormalized fact-table schemas, so the same data serves both audit trails and fast reads.
- Engineered automated SEC ETL pipelines in Airflow — scraped filings through S3/Snowflake staging with validation checks at each hop.
Stock Price Forecaster
Multivariate Hidden Markov Models forecasting price, keyed on opening/closing differentials and volume triggers. Deployed live with Akaike Information Criterion driving state selection, so the model is prevented from overfitting its own state count.
Repository03. Experience
GPU Kernel Engineer
Abundant
- Set up and ran evaluation and RL rollout experiments in Harbor, benchmarking coding agents (Claude Code, Codex) in containerized RL environments with novel ML-systems tasks on GPUs and TPUs that frontier models fail to solve.
- Built job orchestration for GPU verifiers: scheduling verifier runs on rented cloud GPUs and wiring them into agent harnesses so Claude's outputs are automatically tested on real hardware.
- Built persistent megakernels for four LLM architectures with ThunderKittens, collapsing the forward pass into a single launch.
GPU Kernel Engineer
Parsewave
- Designed verifiable, reward-hacking-resistant reward signals that couple numerical correctness with profiled performance, blocking exploits such as hardcoded outputs, timing-loop manipulation, and cache-warming artifacts.
- Built Harbor RL environments and synthesized novel CUDA kernel problems on Blackwell (SM100/B200) that frontier models fail to solve, targeting NVFP4 and TMA-driven workloads to expand the models' capability frontier.
AI Engineer
Community Dream Foundation
- Shipped an LLM-powered concierge and operations assistant for a hotel client, building agent workflows in LangChain and LangGraph on vLLM behind FastAPI, with RAG over property data (room inventory, policies, local guides).
- Wired up custom Skills and MCP connectors to the hotel's booking system, PMS, and internal knowledge base over REST APIs, so the agent could execute transactional operations.
- Built an evaluation harness with curated guest-query sets and regression checks on prompt changes.
Graduate Teaching and Research Assistant
Northeastern University
- Profiled and optimized a vLLM serving path, reworking CPU–GPU synchronization and pinned-memory handling to cut inter-token latency spikes from 70 ms to 48 ms and hold sub-200 ms responses — raising sustained capacity from 12 to 32+ concurrent sessions on the same hardware.
- Built a multimodal speech interface for the Northeastern University chatbot, pairing text-to-speech with FastConformer RNNT speech recognition in an end-to-end streaming pipeline.
- Served as a Teaching Assistant, mentoring students on distributed computing workflows including SLURM job scheduling and parallel processing across CPU and GPU environments.
- Implemented a knowledge distillation pipeline compressing a Qwen image-to-image model into lightweight GAN architectures, achieving a 65% parameter reduction while maintaining visual quality.
AI Engineer
BulkBeings
- Optimized PyTorch training pipelines for Mixtral and Llama models, achieving a 42% training acceleration across 8 A100 GPUs FSDP configuration through memory coalescing and kernel fusion.
- Designed and scaled ML cluster orchestration using Kubernetes and Ray for distributed training, managing GPU scheduling and resource allocation.
- Engineered high-performance GPU kernels for attention mechanisms and feedforward layers using CUDA (CUTLASS); reduced OOM errors by 85% through systematic memory profiling.
ML Engineer (Research)
BulkBeings
- Prototyped OCR model (ViT+CRNN) achieving 98% precision and 42% faster inference via ONNX/CUDA optimization, meeting sub-100ms P99 latency.
- Implemented an observability stack with Prometheus, Grafana, and OpenTelemetry to track model throughput, latency and GPU utilization.
- Developed a two-stage Conv1D-Transformer for beat-level ECG classification (89% F1-score) deployed on an L4 GPU.
ML Engineer Intern
Velozity Global Solutions Pvt
- Architected end-to-end ML pipelines in AWS, implementing predictive segments using XGBoost on behavioral patterns, resulting in a 45% increase in campaign conversions.
- Redesigned temporal feature extraction logic to prevent temporal data leakage, ensuring robust temporal train-test splits.
- Collaborated on developing retail mix optimization systems using Bayesian hierarchical models (PyMC3), increasing LTV prediction accuracy by 34%.
Education
-
MS, Information Systems
Northeastern University, Boston MA. AI/ML, data engineering, and distributed systems. Thesis: ModelOpt — Zero-Shot Computer Vision Model Optimization With Tree Search and Federated Knowledge Sharing.
2023–2025 -
B.Tech, Computer Science
Anna University, Chennai, India. Graduated with distinction.
2019–2023
04. Engineering Notes
Upstream contributions
-
Fixed the Mamba2 cached forward pass for multi-token inputs — the bug blocked chunked prefill and speculative decoding on Mamba2 models.
Merged -
Refactored
Mergedilogbf128in LLVM’s libc to a header-only implementation. -
Added an AWQ option to account for quantization of the smooth layer during scale search.
Open
How I Customized Llama 3.1 8B on a Budget
Democratizing AI: Fine-tuning Large Language Models with Limited Resources.
Fine-tuned large language models consistently outperform generic retrieval-augmented systems. In this article, I document a detailed, step-by-step roadmap to fine-tuning Meta's open-source Llama 3.1 8B parameter model using LoRA adapters, quantization techniques, and budget cloud resources.
Read Full ArticleGPU MODE NVFP4 Blackwell Challenge
Mar 13, 2026
The unfiltered worklog of tackling nvfp4_gemm, nvfp4_dual_gemm, and nvfp4_group_gemm challenges — from MLIR serialization errors to PTX-optimized SwiGLU epilogues.
AMD FlyDSL — competes with CuTe/Cutlass?
Mar 4, 2026
Comparing NVIDIA CuTile and AMD FlyDSL MLIR-based tile programming frameworks by porting a fused MoE kernel and analyzing the generated IR.
Optimizing LLMs on AMD MI300X
Jan 2025
A deep dive into porting transformer kernels from NVIDIA CUDA to AMD ROCm/HIP, achieving 12.1ms token generation with 83% memory bandwidth.
Research
- BitSkip (arXiv:2510.23766) — authored a paper on the compositional effects of quantization and early exit in large language models: what happens to quality when you stack two inference-cost reductions that are usually studied in isolation.
- ModelOpt (Master’s thesis, Northeastern, 2025) — a research framework for zero-shot computer-vision model optimization using tree search and federated knowledge sharing.
NerdingOut on AI
Field notes on LLM inference, GPU kernels, and ML systems — in the open. New posts straight to your inbox.
05. Get In Touch
I’m looking for AI engineer, forward deployed engineer, and ML systems roles — shipping LLM systems with customers, agent workflows and RAG in production, and the training and inference infrastructure underneath. Happy to talk through any of the numbers on this page.
Or reply straight to bhuvaneshwaran.r@northeastern.edu.
- Phone (857) 693-4328
- Location Boston, MA · open to remote
- Calendar Book a slot on Cal.com