A deep, illustrated walkthrough of the Qwen3.5-4B hybrid architecture and exactly how inference runs, with interactive visualizations.
The transformer skeleton, self-attention's math, prefill vs decode, the KV cache, and why GQA exists, built up from first principles.
The GPU memory hierarchy (SRAM vs HBM), why standard attention wastes bandwidth, tiling + online softmax, IO complexity, and what changed in FlashAttention-2 and 3.
A longer walkthrough of how audio modeling evolved from hand-crafted features to self-supervised learning and multimodal systems.
A quick-reference lab for comparing model families, architectures, training choices, and when to reach for each one.
What each handcrafted audio feature captures — from zero-crossing rate to MFCC deltas — and when to reach for which.
No pieces match that search or tag.