Writing
Deep dives, notes, and opinions. Mostly machine learning so far, but not only. Also available as an RSS feed.
2026
- SFT vs SFT + DPO: A Comparison
Comparing supervised fine-tuning alone versus combining it with Direct Preference Optimization for LLM alignment.
2025
GPUs for Machine Learning · Part 6
Multi-GPU Training, Part 2: Fully Sharded Data ParallelismExplains how Fully Sharded Data Parallelism shards parameters and optimizer state to unlock trillion-parameter scale.
GPUs for Machine Learning · Part 5
Multi-GPU Training, Part 1: Data ParallelismHow data parallel training shards mini-batches, synchronizes gradients, and scales workloads across GPU clusters.
GPUs for Machine Learning · Part 4
Streaming Multiprocessors: Scheduling and ExecutionHow warp schedulers, tensor cores, and instruction pipelines inside an H100 SM keep massive thread counts in flight.
GPUs for Machine Learning · Part 3
The GPU Memory Hierarchy: L2, L1, and RegistersA tour of caching, shared memory, and register files on modern accelerators, with tips for keeping tensor cores fed.
GPUs for Machine Learning · Part 2
High Bandwidth Memory (HBM): Why GPUs Need It for Machine LearningExplains why large-scale training depends on terabyte-per-second HBM stacks and how to budget their bandwidth.
GPUs for Machine Learning · Part 1
How GPUs Power Modern Machine Learning: An Introduction to GPU ArchitectureSets the stage for the GPU architecture series—compute, memory, and interconnect pillars for large-scale training.
- Hierarchical Text-Conditional Image Generation with CLIP Latents
Breaks down unCLIP’s diffusion prior, decoder, and editing workflows for high-fidelity text-to-image synthesis.
- Variational Autoencoders
Covers evidence lower bound derivations and the mechanics of VAEs for generative modeling.
- Byte Pair Encoding
Walkthrough of the BPE tokenization algorithm that powers modern language models.
- Zero-Shot Text-to-Image Generation
Analyzes early text-to-image transformers, discrete VAEs, and zero-shot generation techniques.
- Learning Transferable Visual Models From Natural Language Supervision
A deep dive into CLIP’s contrastive training on 400M image–text pairs and its zero-shot recognition capabilities.