I'm Ramshankar. I ship LLM systems with customers: agent workflows, RAG over real ops data, and the eval harness that keeps them honest in production. I also build the stack that gets a model from experiment to production: distributed training pipelines, high-performance serving, and the GPU kernels (CUDA, HIP, Triton) underneath. Recently: evaluation and RL infrastructure for coding agents at Abundant, persistent megakernels across four LLM architectures, NVFP4 Blackwell kernels (3–6× over reference), and 12.1 ms/token decode on a single AMD MI300X.

Open for new opportunities

Ramshankar Bhuvaneswaran - Portrait

01. What I Build

I build LLM systems end to end: the agent workflows and retrieval a customer actually touches, the evals that say whether a change made things better, and the training and inference stack underneath. Master’s in Information Systems from Northeastern.

On the applied side that has meant working with a client to scope the problem, then shipping it: an LLM concierge and operations agent for a hotel, built in LangChain and LangGraph with RAG over property data, MCP connectors into their booking system and PMS so the agent can act rather than just answer, and an eval harness that runs on every prompt change.

Further down, on the training side, that has meant multi-GPU fine-tuning with FSDP, cluster orchestration on Kubernetes and Ray, and a three-cluster SkyPilot pipeline whose spot workers can be reclaimed mid-run without losing the run. On the inference side, serving with vLLM behind FastAPI, retargeting a stack from CUDA to ROCm, and writing the CUDA, HIP, and Triton kernels underneath when the framework leaves performance on the table.

The part I care most about is measurement. Whether it’s a kernel or an agent, most of the wins below started as a measurement that disagreed with someone’s assumption, and every number on this page names the baseline it was measured against.

I work in the open, upstream: a merged fix to the Mamba2 cached forward pass in HuggingFace transformers (PR #46084) and a libc refactor in LLVM (PR #175396).

The stack I work in

Agents & Applied AI

LangChain LangGraph MCP RAG / Graph RAG Neo4j FastAPI Eval harnesses Prompt regression tests

Training Stack

PyTorch FSDP verl / GRPO TRL Unsloth Hugging Face JAX Mixed precision Knowledge distillation

Inference & Kernels

vLLM CUDA HIP / ROCm Triton CuteDSL PTX CUTLASS ThunderKittens ONNX bfloat16 / NVFP4

Cluster & Platform

SkyPilot Kubernetes Ray SLURM Docker AWS / SageMaker GCP / Vertex AI Prometheus Grafana OpenTelemetry

Also fluent in C++, C, Python, SQL, Shell, PySpark, and R; scikit-learn and XGBoost for the classical end.

02. Systems I’ve Built

Qwen600 Inference: CUDA → ROCm Port for AMD MI300X

12.1 msPer-token decode · MI300X
79–83%Peak HBM bandwidth
C++ HIP ROCm AMD MI300X bfloat16

Retargeted an inference stack to a second vendor: six transformer kernels ported from CUDA to ROCm/HIP and tuned for MI300X. Reached 12.1 ms/token decode on a single MI300X at 79–83% of peak HBM bandwidth, through bfloat16 coalesced access that cut memory traffic roughly in half.

Llama 3.1 Training & Inference in Pure C/CUDA

2.3×RMSNorm kernel speedup
C/CUDA llm.c bfloat16 RoPE

Owned the whole path end to end — complete forward and backward passes with attention, no framework underneath. Built on Karpathy’s llm.c, with hand-tuned RMSNorm (2.3× over the reference kernel), bfloat16 SwiGLU, and coalesced memory access.

Repository

Scaling Diffusion Training: Dask Preprocessing + AMP

1.48×Preprocessing throughput scaling
PyTorch Dask AMP EMA DDPM
  • Profiled a from-scratch DDPM training run, found data preprocessing was the bottleneck rather than the GPU, and moved it onto a Dask worker pool for 1.48× CPU throughput scaling.
  • Cut wall-clock training time 11% with mixed precision, and loss variance 29% with EMA weight averaging.
Repository

End-to-End SFT→DPO Pipeline on SageMaker

123KSamples through the ETL path
AWS SageMaker SFT DPO ETL Sharding

Carried one model the full distance from raw data to a served endpoint: ETL over 123K conversation samples, hybrid SFT-then-DPO training with sharding, and deployment on SageMaker. Evaluated at 71% macro-AUC with a 40% improvement in output consistency.

OCR Pipeline with Layout Detection

98%Text extraction accuracy
42%Faster than baseline
PyTorch ViT CRNN ONNX Runtime CUDA Kernels

Deep-learning OCR over complex financial documents using ViT + CRNN. Reached 98% text-extraction accuracy and 42% faster processing through ONNX runtime and custom CUDA kernel optimization — the serving path that later held sub-100ms P99 in production.

Two-Stage Cardiac Arrhythmia Detection

94–98%Rhythm classification
90–95%Ectopic beat detection
PyTorch CNN-Transformer Signal Processing ECG

A CNN-Transformer architecture over raw ECG signals, split into two stages: the first classifies overall rhythm at 94–98% accuracy, the second does adaptive ectopic-beat detection at 90–95% using dynamic patient-specific rate windowing rather than a fixed threshold.

Medical Knowledge-Graph RAG

Neo4j LangChain Graph RAG Python

A Graph RAG framework for medical retrieval: automated triplet extraction builds a knowledge graph in Neo4j dynamically, and semantic search over it through LangChain returns answers with their supporting context rather than a flat passage match.

Repository

Fintech SEC Data Platform

Python Airflow Snowflake S3 FastAPI
  • Built a master financial database for US public companies with raw, transformed, and denormalized fact-table schemas, so the same data serves both audit trails and fast reads.
  • Engineered automated SEC ETL pipelines in Airflow — scraped filings through S3/Snowflake staging with validation checks at each hop.
Repository

Stock Price Forecaster

Streamlit Hidden Markov Models Time Series Python

Multivariate Hidden Markov Models forecasting price, keyed on opening/closing differentials and volume triggers. Deployed live with Akaike Information Criterion driving state selection, so the model is prevented from overfitting its own state count.

Repository

03. Experience

GPU Kernel Engineer

Abundant

Aug 2026 – Present
  • Set up and ran evaluation and RL rollout experiments in Harbor, benchmarking coding agents (Claude Code, Codex) in containerized RL environments with novel ML-systems tasks on GPUs and TPUs that frontier models fail to solve.
  • Built job orchestration for GPU verifiers: scheduling verifier runs on rented cloud GPUs and wiring them into agent harnesses so Claude's outputs are automatically tested on real hardware.
  • Built persistent megakernels for four LLM architectures with ThunderKittens, collapsing the forward pass into a single launch.

GPU Kernel Engineer

Parsewave

Jun 2026 – Aug 2026
  • Designed verifiable, reward-hacking-resistant reward signals that couple numerical correctness with profiled performance, blocking exploits such as hardcoded outputs, timing-loop manipulation, and cache-warming artifacts.
  • Built Harbor RL environments and synthesized novel CUDA kernel problems on Blackwell (SM100/B200) that frontier models fail to solve, targeting NVFP4 and TMA-driven workloads to expand the models' capability frontier.

AI Engineer

Community Dream Foundation

Remote Apr 2026 – Jun 2026
  • Shipped an LLM-powered concierge and operations assistant for a hotel client, building agent workflows in LangChain and LangGraph on vLLM behind FastAPI, with RAG over property data (room inventory, policies, local guides).
  • Wired up custom Skills and MCP connectors to the hotel's booking system, PMS, and internal knowledge base over REST APIs, so the agent could execute transactional operations.
  • Built an evaluation harness with curated guest-query sets and regression checks on prompt changes.

Graduate Teaching and Research Assistant

Northeastern University

Boston, MA Jan 2025 – Dec 2025
  • Profiled and optimized a vLLM serving path, reworking CPU–GPU synchronization and pinned-memory handling to cut inter-token latency spikes from 70 ms to 48 ms and hold sub-200 ms responses — raising sustained capacity from 12 to 32+ concurrent sessions on the same hardware.
  • Built a multimodal speech interface for the Northeastern University chatbot, pairing text-to-speech with FastConformer RNNT speech recognition in an end-to-end streaming pipeline.
  • Served as a Teaching Assistant, mentoring students on distributed computing workflows including SLURM job scheduling and parallel processing across CPU and GPU environments.
  • Implemented a knowledge distillation pipeline compressing a Qwen image-to-image model into lightweight GAN architectures, achieving a 65% parameter reduction while maintaining visual quality.

AI Engineer

BulkBeings

Chennai, India May 2024 – Aug 2024
  • Optimized PyTorch training pipelines for Mixtral and Llama models, achieving a 42% training acceleration across 8 A100 GPUs FSDP configuration through memory coalescing and kernel fusion.
  • Designed and scaled ML cluster orchestration using Kubernetes and Ray for distributed training, managing GPU scheduling and resource allocation.
  • Engineered high-performance GPU kernels for attention mechanisms and feedforward layers using CUDA (CUTLASS); reduced OOM errors by 85% through systematic memory profiling.

ML Engineer (Research)

BulkBeings

Chennai, India May 2023 – Dec 2023
  • Prototyped OCR model (ViT+CRNN) achieving 98% precision and 42% faster inference via ONNX/CUDA optimization, meeting sub-100ms P99 latency.
  • Implemented an observability stack with Prometheus, Grafana, and OpenTelemetry to track model throughput, latency and GPU utilization.
  • Developed a two-stage Conv1D-Transformer for beat-level ECG classification (89% F1-score) deployed on an L4 GPU.

ML Engineer Intern

Velozity Global Solutions Pvt

Chennai, India Jan 2022 – Aug 2022
  • Architected end-to-end ML pipelines in AWS, implementing predictive segments using XGBoost on behavioral patterns, resulting in a 45% increase in campaign conversions.
  • Redesigned temporal feature extraction logic to prevent temporal data leakage, ensuring robust temporal train-test splits.
  • Collaborated on developing retail mix optimization systems using Bayesian hierarchical models (PyMC3), increasing LTV prediction accuracy by 34%.

Education

  • MS, Information Systems

    Northeastern University, Boston MA. AI/ML, data engineering, and distributed systems. Thesis: ModelOpt — Zero-Shot Computer Vision Model Optimization With Tree Search and Federated Knowledge Sharing.

    2023–2025
  • B.Tech, Computer Science

    Anna University, Chennai, India. Graduated with distinction.

    2019–2023

04. Engineering Notes

Upstream contributions

Featured Article

How I Customized Llama 3.1 8B on a Budget

Democratizing AI: Fine-tuning Large Language Models with Limited Resources.

Fine-tuned large language models consistently outperform generic retrieval-augmented systems. In this article, I document a detailed, step-by-step roadmap to fine-tuning Meta's open-source Llama 3.1 8B parameter model using LoRA adapters, quantization techniques, and budget cloud resources.

Read Full Article

GPU MODE NVFP4 Blackwell Challenge

Mar 13, 2026

The unfiltered worklog of tackling nvfp4_gemm, nvfp4_dual_gemm, and nvfp4_group_gemm challenges — from MLIR serialization errors to PTX-optimized SwiGLU epilogues.

Read Substack

AMD FlyDSL — competes with CuTe/Cutlass?

Mar 4, 2026

Comparing NVIDIA CuTile and AMD FlyDSL MLIR-based tile programming frameworks by porting a fused MoE kernel and analyzing the generated IR.

Read Substack

Optimizing LLMs on AMD MI300X

Jan 2025

A deep dive into porting transformer kernels from NVIDIA CUDA to AMD ROCm/HIP, achieving 12.1ms token generation with 83% memory bandwidth.

Read Substack

Research

  • BitSkip (arXiv:2510.23766) — authored a paper on the compositional effects of quantization and early exit in large language models: what happens to quality when you stack two inference-cost reductions that are usually studied in isolation.
  • ModelOpt (Master’s thesis, Northeastern, 2025) — a research framework for zero-shot computer-vision model optimization using tree search and federated knowledge sharing.

NerdingOut on AI

Field notes on LLM inference, GPU kernels, and ML systems — in the open. New posts straight to your inbox.

Subscribe on Substack

05. Get In Touch

I’m looking for AI engineer, forward deployed engineer, and ML systems roles — shipping LLM systems with customers, agent workflows and RAG in production, and the training and inference infrastructure underneath. Happy to talk through any of the numbers on this page.

Or reply straight to bhuvaneshwaran.r@northeastern.edu.