pytorcher.com Attention Lab

THE TECH BLOG / AI CO-DESIGN

AI co-design.
Step by step.

From the math inside a model to the hardware that runs it. Understand how architecture, precision, kernels and memory fit together—one design decision at a time.

Start with Attention Lab

START READING

The first deep dive

One published guide. Plenty to unpack.

PUBLISHED · INTERACTIVE GUIDE

Attention Lab

Attention architectures, one matrix at a time.

Trace a token from full causal attention to Qwen3.5’s hybrid stack. Follow every projection, inspect real checkpoint configs, and see what changing context length does to persistent memory.

  • 9 architecture chapters
  • Real model configurations
  • Live shapes & memory
Read the guide
INSIDE THE GUIDEGPT-2 · one head

Q Compare a query with the context

[1, 64] × [64, 1,024]
= [1, 1,024] scores

V Mix the values

[1, 1,024] × [1,024, 64]
= [1, 64] output

One new token. 1,024 context positions. The highlighted axis is history, not query length.

Supplementary shape preview · B=1, Q=1 · dhead=768÷12=64. See the source config and full derivation.

THE BASELINEWhy does KV memory grow?Start with full causal attention HEAD SHARINGWhat changes when heads share K/V?Understand grouped-query attention COMPRESSIONWhat does a latent cache store?Trace DeepSeek’s MLA THE HYBRIDHow does Qwen3.5 combine memory?Walk both layer types

THE LEARNING MAP

One system. Many connected decisions.

Start with attention, then connect the model to its execution. The outlines below show the planned sequence; only Attention Lab is published today.

01 / ARCHITECTUREAvailable

Attention & persistent state

Trace the math. Count what persists. Understand the design choices behind MHA, GQA, MLA and recurrent hybrids.

Explore Attention Lab
02 / WORKLOADPlanned guide

Prefill, decode & batching

Understand why processing a prompt and generating the next token put different demands on the same model.

See the steps
  1. Write down B, Q and T for each workload.
  2. Trace the changed shapes, reuse and data movement.
  3. Compare latency, throughput and memory as batching changes.
03 / NUMERICSPlanned guide

Precision & quantization

Follow a number from its stored format to its accumulation—and understand what a smaller format costs.

See the steps
  1. Separate weight, activation, cache and state precision.
  2. Work through scaling, rounding and quantization error.
  3. Count metadata and test accuracy alongside memory savings.
04 / KERNELSPlanned guide

From matmul to GPU kernel

Connect a logical tensor operation to the tiles, memory accesses and synchronization that execute it.

See the steps
  1. Start with a small matrix multiplication.
  2. Map its tiles and reuse to the memory hierarchy.
  3. Compare separate and fused operations with measured work and traffic.
05 / HARDWAREPlanned guide

Compute, bandwidth & capacity

Build a concrete explanation for a bottleneck before deciding whether more compute will help.

See the steps
  1. Count operations and bytes for a stated workload.
  2. Estimate limits from compute throughput and memory bandwidth.
  3. Compare the estimate with measurements and explain the gap.
06 / SYSTEMSPlanned guide

MoE & distributed inference

Follow weights, tokens and communication across devices. Separate active computation from resident capacity.

See the steps
  1. Distinguish total parameters from the experts used by a token.
  2. Trace routing and the placement of weights and state.
  3. Account for communication, balance and end-to-end latency.

HOW WE LEARN HERE

Make every step inspectable.

A clear explanation should let you reconstruct the result yourself.

  1. 01

    Start with something real

    Use a named model, a published configuration and a clearly defined workload.

  2. 02

    Follow the operations

    Write down the shapes and explain what each multiplication or state update does.

  3. 03

    Connect the tradeoffs

    Separate compute, persistent memory and temporary storage. Keep assumptions visible.

  4. 04

    Work the numbers

    Substitute actual dimensions, vary one input, and check whether the explanation still holds.

BEGIN WITH ONE TOKEN

Then follow everything it touches.

Open Attention Lab