What is the difference between a GPU and TPU? – Reiner Pope
Full episode: • Chip design from the bottom up – Reiner Pope Me on twitter: https://x.com/dwarkesh_sp
- Uploaded
- Uploaded May 24, 2026
- File type
- YouTube
- Queried
- 0
Full transcript
Showing the full transcript for this video.
Title: What is the difference between a GPU and TPU? – Reiner Pope Platform: youtube Creator: Dwarkesh Clips Source: v=95w9IIUjk7g Transcript source: assemblyai Summary: Full episode: • Chip design from the bottom up – Reiner Pope Me on twitter: com/dwarkesh_sp Transcript: Speaker A: What is the difference at a high level between how a GPU and a TPU work? Speaker B: Yeah, so I mean, I think there's sort of a high-level organization principle that is different. And then there's sort of inside the cores what are different. But we'll look sort of outside the, like at the high level.
So we'll take a GPU and a TPU and what does like sort of the top-level block structure look like? If you think of this as the whole chip, in each case. The organization of the GPU is mostly a bunch of almost identical units, which are these, these are the SMs. And then they've got an L2 memory in the middle and then a bunch more of these SMs on the bottom. And so there's sort of this fairly regular grid of cores. And then if we look at a TPU in comparison, you end up with much coarser-grained units of logic.
And so you end up with something like some large number of maybe just a few matrix units. These are the big systolic arrays. And then in the middle you've got some vector unit, and then you've got your matrix units at the bottom. So now sort of like matrix units with a vector unit in the middle, sort of this is the whole TPU chip. You can sort of think of scaling this thing down into a really tiny unit with a smaller matrix unit, smaller vector unit. And that is sort of what an SM is.
So sort of at a very high-level point of view, the GPU has a lot of tiny, tiny TPUs tiled across the whole chip. Speaker A: Oh, interesting. So you're suggesting the Tensor Core within a streaming SM is analogous to an MXU? Speaker B: Yeah, it's very, very similar. Yeah. Speaker A: I see. And so if you had more lack of structure, having a bunch of tiny TPUs makes a lot of sense. Whereas if you just have huge matrix multiplications, you're like, why don't we avoid the cost of having the individual SMs with their own registers and warp schedulers and things like that?
Why don't we just make a huge thing and amortize those costs across the whole thing? Speaker B: I think this shows up in how large you can grow things. We've seen this theme, especially with the systolic array, where larger systolic array amortizes the register file costs better. Speaker A: Yeah. Speaker B: This sort of design allows you to have larger systolic arrays, whereas the sort of GPU design constrains you to having small units of everything. There is a trade-off, however. There ends up being, because of this sort of coarse-grained separation of things there, you need to move a lot of data from the vector unit to the matrix units.
And so like, you need to move a lot of data through a sort of like two lines of parameter here. Whereas if you sort of look at the equivalent thing here, you've got vector units everywhere and you need to move data through this line, through this line, through this line, through this line, through this line, through this line. So the amount of data you can move between a vector unit and a matrix unit is actually much higher in, in a GPU than in a TPU because instead of having to move all the data through these just 2 lines, you're moving all these data through 16 lines or something of wiring instead in a GPU.
Speaker A: Right. But also you might have to move across less area. Speaker B: Which I mean is also a saving, like it's an energy saver. So data ends up moving, like if you can operate entirely within an SM, the data movement is much smaller. But then the moment you want to operate across SMs, it becomes sort of more complicated and expensive. Speaker A: So you don't have to comment, but one might expect that a thing Maddox might try to do is to get the GPU-like smaller structure of systolic arrays surrounded by SRAM, but also at the same time make it so that the things you need in an SM to support the CUDA architecture but take a bunch of space you might discard.
Speaker B: Yeah, we've talked publicly about something which we call a susplittable systolic array, which is sort of, in some sense, you can think of as like big systolic arrays that can be small systolic arrays as well. Speaker A: If you enjoyed this clip, you can watch the full episode here and subscribe for more clips. Thanks.
Want to learn more?
Ask about this video