Studio

A space for notes, tools, small apps, and things I’m figuring out in public.

Longform note
Inside Qwen3.5‑4B: Architecture & Forward Pass

A deep, illustrated walkthrough of the Qwen3.5-4B hybrid architecture and exactly how inference runs, with interactive visualizations.

Qwen Inference Optimization Architecture
Blog / tutorial Open piece →
Longform note
From Self-Attention to Grouped-Query Attention

The transformer skeleton, self-attention's math, prefill vs decode, the KV cache, and why GQA exists, built up from first principles.

Qwen Inference Optimization Attention
Blog / tutorial Open piece →
Longform note
FlashAttention: Fast, Exact Attention by Respecting the Memory Hierarchy

The GPU memory hierarchy (SRAM vs HBM), why standard attention wastes bandwidth, tiling + online softmax, IO complexity, and what changed in FlashAttention-2 and 3.

Inference Optimization Attention GPU
Blog / tutorial Open piece →
Longform note
The Story of Modern Audio Models

A longer walkthrough of how audio modeling evolved from hand-crafted features to self-supervised learning and multimodal systems.

Audio Models Longform
Blog / tutorial Open piece →
Interactive blog
Audio Model Explorer

A quick-reference lab for comparing model families, architectures, training choices, and when to reach for each one.

Audio Models Interactive Reference
Interactive blog Open tool →
Field guide
A Field Guide to Audio Features

What each handcrafted audio feature captures — from zero-crossing rate to MFCC deltas — and when to reach for which.

Audio Models Reference
Blog / reference Open guide →

No pieces match that search or tag.