Hi, I'm Nikhil Kuniyil.
I'm an M.S. student in Electrical & Computer Engineering at UC San Diego, and this fall I'm interning at Amazon on the Echo Spatial Perception team.
Before this, I was a Machine Learning Engineer Intern at Adobe Search, optimizing our existing LLM inference serving stack. At UCSB's Vision Research Lab, I worked on small object detection in high-resolution satellite imagery, which led to RareSpot and RareSpot+ (IJCV 2026).
I also write here, and I'm always happy to chat, so feel free to reach out.
Writing
- Feb 5, 2026SFT vs SFT + DPO: A Comparison
Comparing supervised fine-tuning alone versus combining it with Direct Preference Optimization for LLM alignment.
- Sep 25, 2025Multi-GPU Training, Part 2: Fully Sharded Data Parallelism
How FSDP splits weights, gradients, and optimizer states across GPUs so models too big for one GPU can still be trained.
- Sep 23, 2025Multi-GPU Training, Part 1: Data Parallelism
How data parallel training splits batches across GPUs, keeps gradients in sync, and why it matches training on one big GPU.
- Sep 20, 2025Streaming Multiprocessors: Scheduling and Execution
What's inside a streaming multiprocessor, and how it keeps thousands of threads busy at once.
- Sep 19, 2025The GPU Memory Hierarchy: L2, L1, and Registers
How the L2 cache, L1 and shared memory, and registers bring data closer to the cores, and why that matters for fast kernels.
- Sep 18, 2025High Bandwidth Memory (HBM): Why GPUs Need It for Machine Learning
Why large models depend on HBM, how it works, and what its capacity and bandwidth numbers actually mean.
Publications
- Jul 2026RareSpot+: A Benchmark, Model, and Active Learning Framework for Small and Rare Wildlife in Aerial Imagery
Bowen Zhang, Jesse T. Boulerice, Charvi Mendiratta, Nikhil Kuniyil, Satish Kumar, Hila Shamon, B. S. Manjunath
International Journal of Computer Vision (IJCV), vol. 134, 2026
Extends RareSpot with geospatially guided active learning that uses spatial priors between prairie dogs and their burrows to cut redundant labeling; +35.2% mAP@50 over the baseline on a 2 km² aerial survey.
- Jun 2025RareSpot: Spotting Small and Rare Wildlife in Aerial Imagery with Multi-Scale Consistency and Context-Aware Augmentation
Bowen Zhang, Jesse T. Boulerice, Nikhil Kuniyil, Charvi Mendiratta, Satish Kumar, Hila Shamon, B. S. Manjunath
CVPR 2025 Workshop on Computer Vision for Animals (CV4Animals)
Multi-scale consistency learning and context-aware augmentation for detecting tiny, sparse animals in drone imagery, improving detection accuracy by over 35% versus baselines.
Projects
- Apr 2026Tiny-GRPO: RL Post-Training from Scratch
Python · PyTorch · Hugging Face Transformers
Implemented Group Relative Policy Optimization from scratch to post-train SmolLM2-135M through a base → SFT → GRPO pipeline with an exact-match reward, raising held-out accuracy from 9.4% to 15.6% and parse rate from 62.5% to 100%.
Experience
- Fall 2026
Amazon · Software Development Engineer Intern
Echo Spatial Perception team, working on the spatial and sensor-fusion systems behind Alexa's ambient-device experiences.
- Summer 2026
Adobe · Machine Learning Engineer Intern
ML training and inference infrastructure for Search & Discovery. Integrated NVIDIA Dynamo into the LLM serving stack with disaggregated prefill/decode and KV-cache-aware routing, cutting end-to-end response latency by 45%.
- Summer 2025
Amazon · Software Development Engineer Intern
Built an LLM-powered root-cause analysis tool on Amazon Bedrock for Alexa integration-test failures, cutting investigation time by 70% and MTTR by 50%.
- 2023 – 2025
Vision Research Lab, UCSB · Machine Learning Researcher
Co-authored RareSpot and built its end-to-end detection pipeline as the only undergraduate in the lab, plus a Kubernetes workflow that cut model execution time by 30%.
- Summer 2024
Learfield · Software Engineer Intern
Built a centralized Next.js sign-in page consolidating authentication across services, and a CI/CD pipeline for Prometheus metrics that cut deployment time by 40%.
- Summer 2023
Allstate · Data Engineer Intern
Refactored PySpark ETL workflows processing 500M+ records a week, cutting compute cost by 25%, and built automated anomaly detection for NLP training data.