AMD Details Renoir: The Ryzen Mobile 4000 Series 7nm APU Uncovered
by Dr. Ian Cutress on March 16, 2020 11:00 AM ESTWhat’s New
AMD’s latest Ryzen mobile product is the first design the company has done that combines CPU, GPU, and IO all on a monolithic die in TSMC’s 7nm process.
The CPU part of the design is very similar to what we’ve seen on the desktop: two quad core groups each with their own L3 cache shared between the cores. Compared to the desktop design, the mobile is listed as being ‘optimized for mobile’, primarily by the smaller L3 cache – only 4 MB per quad-core group, rather than the 32 MB per quad-core group we see on the desktop. While the smaller L3 cache might mean more trips out to main memory to get data, overall AMD sees it as saving both power and die area, with this level of cache being the right balance for a power limited chip.
Compared to the previous generation of Zen mobile processors, this generation on the CPU side of the equation comes with the 15% per-core iso-frequency improvement, down to the improvements at the heart of each core. We’ve covered these in detail in our desktop analysis. However for the mobile platform, not only is there a raw performance uplift, but we’re also seeing frequency uplift as well, moving from 4.0 GHz in the prior gen up to 4.3 GHz here. Actual workload performance AMD says gets a significant uplift due to the new power features we’ll discuss in due course.
On the GPU side is where we see bigger changes. AMD does two significant things here – it has reduced the maximum number of graphics compute units from 11 to 8, but also claims a +59% improvement in graphics performance per compute unit despite using the same Vega graphics architecture as in the prior generation. Overall, AMD says, this affords a peak compute throughput of 1.79 TFLOPS (FP32), up from 1.41 TFLOPS (FP32) on the previous generation, or a +27% increase overall.
AMD manages to improve the raw performance per compute unit through a number of changes to the design of the APU. Some of this is down to using 7nm, but some is down to design decisions, but it also requires a lot of work on the physical implementation side.
For example, the 25% higher peak graphics frequency (up from 1400 MHz to 1750 MHz) comes down a lot to physical implementation of the compute units. Part of the performance uplift is also due to memory bandwidth – the new Renoir design can support LPDDR4X-4266 at 68.3 GB/s, compared to DDR4-2400 at 38.4 GB/s. Most GPU designs need more memory bandwidth, especially APUs, so this will help drastically on that front.
There are also improvements in the data fabric. For GPUs, the data fabric is twice as wide, allowing for less overhead when bulk transferring data into the compute units. This technically increases idle power a little bit compared the previous design, however the move to 7nm easily takes that onboard. With less power overhead for bulk transfer data, this makes more power available to the GPU cores, which in turn means they can run at a higher frequency.
Coming to the Infinity Fabric, AMD has made significant power improvements here. One of the main ones is decoupling the frequency of Infinity Fabric from the frequency of the memory – AMD was able to do this because of the monolithic design, whereas in the chiplet design of the desktop processors, the fix between the two values has to be in place otherwise more die area would be needed to transverse the variable clock rates. This is also primarily the reason we’re not seeing chiplet based APUs at this time. However, the decoupling means that the IF can idle at a much lower frequency, saving power, or adjust to a relevant frequency to mix power and performance when under load.
Again we see the double bus width from the graphics to the engine pop up here, giving a better power-per-bit metric. But one of the key aspects from this graph is showing that the power consumed by the fabric in the new processors is very even across a wide bandwidth range compared to the older processor, where the voltages likely had to be stepped up as bandwidth increased, and introducing additional latency factors for performance. Luckily Renoir does away with this, and AMD are claiming a 75% better fabric efficiency compared to the previous generation.
Orthogonal to the raw improvements, AMD has also improved the media capabilities, with a new HDR/WCG encode engine for HEVC, which according to AMD should give a 31% encoding speedup when used.
95 Comments
View All Comments
Namisecond - Thursday, March 26, 2020 - link
Are you sure you don't mean an XPS 15 rather than the 13? The 13's have traditionally use U-class processors rather than H-class.deksman2 - Monday, March 16, 2020 - link
The iGP should be about 30% faster than Zen+.AMD improved Vega iGP uArch in Zen 2 mobile, so that per core, that Vega is now 56% faster than Zen+.
The Zen 2 mobile Vega iGP does apparently have lower amount of compute units, but because of improved performance per core, the iGP should still be about 30% faster than previous iterations.
So, yes, I do think games set to Medium settings and 1080p should work fine... more to the point, since these APU's also apparently support LPDDR4 with high frequencies, bandwidth won't be a problem... so hypothetically at least we could see even better performance.
Those improvements that went into Zen 2 mobile Vega will also be integrated in RDNA2.
LittlePaul - Monday, March 16, 2020 - link
I am wondering on the top left of the die shot, it looks like 64-bit GDDR6 controller comparing to Navi 10 die shot (exactly the same and way different to Piccasso )I am dreaming, right?
Spunjji - Tuesday, March 17, 2020 - link
Isn't that the CPU cores? Pretty sure that's what it is.uibo - Monday, March 16, 2020 - link
The Renoir APU graph on the first page has a block called "video codec next 2". I thought the engine was called "Video core next"?edzieba - Monday, March 16, 2020 - link
Note that AMD have chosen power-under-load as their look-at-the-big-number-we-have 'slide headline' figure, whilst the killer when it came to battery life for Ryzen Mobile 2xxx/3xxx was idle power.Fataliity - Monday, March 16, 2020 - link
There are slides on other sites, like notebookcheck.com and the video from the australians (hardware unboxed i think) That show the actual slides, and the data AMD themselves got.1065G7 still wins in some idles, but renoir pulls ahead in actual use it seems. Which evens out to about equal / slightly better for Renoir, depending on how you weigh it.
Spunjji - Tuesday, March 17, 2020 - link
They mention a number of improvements to idle, too - particularly from the Infinity Fabric, which was their big weakness before.anonomouse - Monday, March 16, 2020 - link
" Compared to the desktop design, the mobile is listed as being ‘optimized for mobile’, primarily by the smaller L3 cache – only 4 MB per quad-core group, rather than the 32 MB per quad-core group we see on the desktop"Isn't it 16MB per CCX on desktop?
Farfolomew - Tuesday, March 24, 2020 - link
I saw that too and questioned its validity. 32mb per 4-cores? That seems really high (I can't remember what it was reading the Zen2 desktop review last year) and if it's the case, that's a huge cut down for the mobile line