THE TECH BLOG / AI CO-DESIGN
AI co-design.
Step by step.
From the math inside a model to the hardware that runs it. Understand how architecture, precision, kernels and memory fit together—one design decision at a time.
Start with Attention Lab
START READING
The first deep dive
One published guide. Plenty to unpack.
PUBLISHED · INTERACTIVE GUIDE
Attention architectures, one matrix at a time.
Trace a token from full causal attention to Qwen3.5’s hybrid stack. Follow every projection, inspect real checkpoint configs, and see what changing context length does to persistent memory.
- 9 architecture chapters
- Real model configurations
- Live shapes & memory
Read the guide
INSIDE THE GUIDEGPT-2 · one head
Q Compare a query with the context
[1, 64] × [64, 1,024]
= [1, 1,024] scores
V Mix the values
[1, 1,024] × [1,024, 64]
= [1, 64] output
One new token. 1,024 context positions. The highlighted axis is history, not query length.
Supplementary shape preview · B=1, Q=1 · dhead=768÷12=64. See the source config and full derivation.
THE LEARNING MAP
One system. Many connected decisions.
Start with attention, then connect the model to its execution. The outlines below show the planned sequence; only Attention Lab is published today.
01 / ARCHITECTUREAvailable
Attention & persistent state
Trace the math. Count what persists. Understand the design choices behind MHA, GQA, MLA and recurrent hybrids.
Explore Attention Lab
02 / WORKLOADPlanned guide
Prefill, decode & batching
Understand why processing a prompt and generating the next token put different demands on the same model.
See the steps
- Write down B, Q and T for each workload.
- Trace the changed shapes, reuse and data movement.
- Compare latency, throughput and memory as batching changes.
03 / NUMERICSPlanned guide
Precision & quantization
Follow a number from its stored format to its accumulation—and understand what a smaller format costs.
See the steps
- Separate weight, activation, cache and state precision.
- Work through scaling, rounding and quantization error.
- Count metadata and test accuracy alongside memory savings.
04 / KERNELSPlanned guide
From matmul to GPU kernel
Connect a logical tensor operation to the tiles, memory accesses and synchronization that execute it.
See the steps
- Start with a small matrix multiplication.
- Map its tiles and reuse to the memory hierarchy.
- Compare separate and fused operations with measured work and traffic.
05 / HARDWAREPlanned guide
Compute, bandwidth & capacity
Build a concrete explanation for a bottleneck before deciding whether more compute will help.
See the steps
- Count operations and bytes for a stated workload.
- Estimate limits from compute throughput and memory bandwidth.
- Compare the estimate with measurements and explain the gap.
06 / SYSTEMSPlanned guide
MoE & distributed inference
Follow weights, tokens and communication across devices. Separate active computation from resident capacity.
See the steps
- Distinguish total parameters from the experts used by a token.
- Trace routing and the placement of weights and state.
- Account for communication, balance and end-to-end latency.
HOW WE LEARN HERE
Make every step inspectable.
A clear explanation should let you reconstruct the result yourself.
- 01
Start with something real
Use a named model, a published configuration and a clearly defined workload.
- 02
Follow the operations
Write down the shapes and explain what each multiplication or state update does.
- 03
Connect the tradeoffs
Separate compute, persistent memory and temporary storage. Keep assumptions visible.
- 04
Work the numbers
Substitute actual dimensions, vary one input, and check whether the explanation still holds.