Hi, I'm Nikhil Kuniyil.
I'm an M.S. student in Electrical & Computer Engineering at UC San Diego, and this fall I'm interning at Amazon on the Echo Spatial Perception team.
Before this, I was a Machine Learning Engineer Intern at Adobe Search, optimizing our existing LLM inference serving stack. At UCSB's Vision Research Lab, I worked on small object detection in high-resolution satellite imagery, which led to RareSpot and RareSpot+ (IJCV 2026).
I also write here, and I'm always happy to chat, so feel free to reach out.
Writing
- Feb 5, 2026SFT vs SFT + DPO: A Comparison
Comparing supervised fine-tuning alone versus combining it with Direct Preference Optimization for LLM alignment.
- Sep 25, 2025Multi-GPU Training, Part 2: Fully Sharded Data Parallelism
Explains how Fully Sharded Data Parallelism shards parameters and optimizer state to unlock trillion-parameter scale.
- Sep 23, 2025Multi-GPU Training, Part 1: Data Parallelism
How data parallel training shards mini-batches, synchronizes gradients, and scales workloads across GPU clusters.
- Sep 20, 2025Streaming Multiprocessors: Scheduling and Execution
How warp schedulers, tensor cores, and instruction pipelines inside an H100 SM keep massive thread counts in flight.
- Sep 19, 2025The GPU Memory Hierarchy: L2, L1, and Registers
A tour of caching, shared memory, and register files on modern accelerators, with tips for keeping tensor cores fed.
- Sep 18, 2025High Bandwidth Memory (HBM): Why GPUs Need It for Machine Learning
Explains why large-scale training depends on terabyte-per-second HBM stacks and how to budget their bandwidth.
Publications
- Jul 2026RareSpot+: A Benchmark, Model, and Active Learning Framework for Small and Rare Wildlife in Aerial Imagery
Bowen Zhang, Jesse T. Boulerice, Charvi Mendiratta, Nikhil Kuniyil, Satish Kumar, Hila Shamon, B. S. Manjunath
International Journal of Computer Vision (IJCV), vol. 134, 2026
Extends RareSpot with geospatially guided active learning that uses spatial priors between prairie dogs and their burrows to cut redundant labeling; +35.2% mAP@50 over the baseline on a 2 km² aerial survey.
- Jun 2025RareSpot: Spotting Small and Rare Wildlife in Aerial Imagery with Multi-Scale Consistency and Context-Aware Augmentation
Bowen Zhang, Jesse T. Boulerice, Nikhil Kuniyil, Charvi Mendiratta, Satish Kumar, Hila Shamon, B. S. Manjunath
CVPR 2025 Workshop on Computer Vision for Animals (CV4Animals)
Multi-scale consistency learning and context-aware augmentation for detecting tiny, sparse animals in drone imagery, improving detection accuracy by over 35% versus baselines.
Projects
- Apr 2026Tiny-GRPO: RL Post-Training from Scratch
Python · PyTorch · Hugging Face Transformers
Implemented Group Relative Policy Optimization from scratch to post-train SmolLM2-135M through a base → SFT → GRPO pipeline with an exact-match reward, raising held-out accuracy from 9.4% to 15.6% and parse rate from 62.5% to 100%.
Experience
- Fall 2026
Amazon · Software Development Engineer Intern
Echo Spatial Perception team, working on the spatial and sensor-fusion systems behind Alexa's ambient-device experiences.
- Summer 2026
Adobe · Machine Learning Engineer Intern
ML training and inference infrastructure for Search & Discovery. Integrated NVIDIA Dynamo into the LLM serving stack with disaggregated prefill/decode and KV-cache-aware routing, cutting end-to-end response latency by 45%.
- Summer 2025
Amazon · Software Development Engineer Intern
Built an LLM-powered root-cause analysis tool on Amazon Bedrock for Alexa integration-test failures, cutting investigation time by 70% and MTTR by 50%.
- 2023 – 2025
Vision Research Lab, UCSB · Machine Learning Researcher
Co-authored RareSpot and built its end-to-end detection pipeline as the only undergraduate in the lab, plus a Kubernetes workflow that cut model execution time by 30%.
- Summer 2024
Learfield · Software Engineer Intern
Built a centralized Next.js sign-in page consolidating authentication across services, and a CI/CD pipeline for Prometheus metrics that cut deployment time by 40%.
- Summer 2023
Allstate · Data Engineer Intern
Refactored PySpark ETL workflows processing 500M+ records a week, cutting compute cost by 25%, and built automated anomaly detection for NLP training data.