PS3 used the Cell processor but it’s debatable how much that was an asset vs handicap. Total PlayStation sales dropped for that generation vs both PS2 and PS4. The manufacturing costs didn’t fall nearly as fast as expected and it was a poor fit in terms of cross platform development etc.
For starters, IBM did also do the CPU for the GameCube, and that was likely a part of the Wii using an upgraded PPC arch for that.
On the flipside, there is the theory (I think even Copetti brings it up in their XBox 360 Architecture breakdown [0]) that IBM using the Cell PPE for the 360's tri-cores left a sour taste in Toshiba, but more-so Sony's mouths.
I think the big 'X factor' though, was that, for as much pain as it caused AMD in the short term, (it's so easy to forget their 'malaise' era, i.e. Early Bulldozer and the GloFo split pains to their margins[1]), AMD made all the 'right' choices to let the console vendors have their cake and eat it too.
Namely, AMD was more than happy to do a custom core if there was a volume contract (similar to what IBM was willing to for the 360/GC/Wii) but also now had a capable, in-house GPU. (And thankfully had Bobcat as a stepping point towards Jaguar[2])
There's part of me that asks, if AMD had an ARM core, if we would have all Consoles powered by AMD chips now. Nintendo likely bought into Tegra because it was an ARM core, and for better or worse their mobile stuff by the time of making that choice had 15+ years of proven ARM success (GBA, DS, 3DS) behind it.
(See also, Intel in the 2010s scrambling with half-assed promises of integrating custom functionality or FPGAs with x86 cores.)
[1] - From what I recollect, the GloFo split and how the contracts were drawn up as far as their production, had a huge impact on their ability to produce due to yields and thermals, as well as the contracts for how GloFo got paid; it was at least part of them diversifying with TSMC as soon as they reasonably could.
[2] - As a Rant, I am pretty sure, if Jaguar had Desktop/Mobile versions that included a Dual channel DDR controller, they would have cleaned up on the low cost laptop market. I had one with, I think it was an A5000 or A5200, and for how tiny the battery was it could last wayyyy longer than any of the intel laptops I had for the time, but churned if you were doing memory heavy stuff.
I think Nintendo used Tegra because it was super cheap. AMD has access to the same ARM cores as everyone else (see Seattle and Sound Waves) but a semi-custom chip would have been more expensive than an overstocked Tegra.
To your point... yeah, Switch was the only 'real' big volume hit for Tegra that I remember (Although I did like my 2012 Nexus 7) with any staying power.
Hell Nvidia was so desperate they did the whole Shield thing...
> On the flipside, there is the theory (I think even Copetti brings it up in their XBox 360 Architecture breakdown [0]) that IBM using the Cell PPE for the 360's tri-cores left a sour taste in Toshiba, but more-so Sony's mouths.
It's more than a theory. It's pretty much spelled out explicitly in The Race For A New Game Machine how salty not just Sony and Toshiba were in the broad sense, but also the Sony and Toshiba engineers that the IBM team worked with felt pretty betrayed.
Intel did release a few generations of their Xeons with an Arria built in but it's been a couple years since the last release. It's a cool idea but I think all the money dried up quick once the AI boom started. Plus they spun off Altera so I doubt they'll be doing it again.
>There's part of me that asks, if AMD had an ARM core, if we would have all Consoles powered by AMD chips now. Nintendo likely bought into Tegra because it was an ARM core, and for better or worse their mobile stuff by the time of making that choice had 15+ years of proven ARM success (GBA, DS, 3DS) behind it.
AMD worked on a ARM core https://en.wikipedia.org/wiki/AMD_K12 but I'm not sure it would had work, they were never the best on the power consumption, so I'm not sure it would have been a good choice for the Switch.
The story about that AMD ARM core is that they were making it for Amazon. Performance was not good enough so it was canned. Amazon then went on to make their own chips called Graviton.
What does IBM actually do? I just can never understand their business and operating model. It seems like they just do a bunch of random stuff and sell to the most enterprisey of enterprises.
You should dive into the topic of Mainframe. There are a lot of financial transaction to other critical infrastructure are dependent on it. And not just because of backward compatibility but technical superiority.
That being said I think it's a natural consequence of the difference between a mainframe and commodity servers. A mainframe is going to be running pretty disparate workloads simultaneously, so it makes sense to steal from your neighbor if they aren't using their cache. Whereas it's more likely that a commodity server is just running the same server on each core, and if you have a different workload, you pick a different shape of server to run it on. There are pros and cons to both.
I do wonder about the spectre consequences of borrowing cache lines from other cores though.
Even though their CPUs are insanely fast, the real power of mainframes is in their IO. The amount of data you can push through those machines is absolutely mind blowing.
Historically this was because each I/O "channel" was a separate computer that handled the actual communication with the device, be it a terminal, disk, tape drive, card reader, printer, etc. and exchange data with the CPU via DMA. This allowed mainframe CPUs, which in the past weren't particularly fast, to handle huge workloads involving hundreds or thousands of users. These days, even commodity computers get blazing fast I/O to bus mastering devices. Where the mainframes win today is on reliability, built-in redundancy, hot-pluggability and expandability of components (you can just plug in CPUs, memory, disks, and network interfaces as long as you can afford them with the machine still running), and service and support. (Mainframes phone home immediately if they detect problems and an IBM service person will be on site the same day to fix it.)
A 4-rack system is about 1.5 racks of CPUs and 2.5 racks of IO. That's a lot of IO for a single computer, all accessible at bus speeds.
Mere mortals such as me, have to make do with cloud-based clusters where the IO is distributed across multiple machines connected by very fast interconnects, but they don't behave as a single machine, nor have the IO always accessible at local bus speeds.
> Mainframes phone home immediately if they detect problems and an IBM service person will be on site the same day to fix it.
The joke usually went like this: technician shows up to fix the machine, operator says "We didn't call you", to what the technician answers "You didn't. Your computer did"
Basically for Telum II, (I don't know what changed here from Telum III, the core under discussion with the ARM decoders) each CPU core has a giant 36MB L2 cache. Then, rather than a discrete L3 cache, the cores keep track of L2 residency needed for that core's working set, and allocate the rest of their L2 to a shared pool that is the L3 cache. Then the same thing with L4 being the same pools in all of the other chips on the same drawer (which you can kind of think of as close to a single server).
The z mainframe team does good technical work but then the high price cancels out all the value of that work. I'm not sure if that counts as innovation or not.
Sometimes quality of service matters more than cost.
If the consumer segment started to develop a sense that their payment cards were glitchy or unreliable, our economy could suffer real consequences. Error free, secure, instant payment experiences are essential. Go visit a local grocery store and measure how long it takes to authorize a transaction each time a card is presented to the terminal. Pay special attention to how fast the Visa and American Express networks are. At my local Kroger the terminal and network setup is so fast that you can use the chip reader almost like a mag stripe reader.
Not totally - it's really nice when you don't have to worry about complex cluster scaling and it's all nice and vertical. Everything in those machines is also hot swappable so really the only downtime will be from an application defect or a catastrophic event (fire in the datacenter).
I was one of them, but I don't think IBM hasn't innovated, my contention is more that they have de-emphasised software and hardware in favour of services, and they've moved away from consumer-facing activity to be entirely B2B, and that there is an overall feeling of slow decline.
Their hardware advances are real, Power chips are still excellent and IBM's mainframes are pretty unique, but the niche for both of those seems less relevant over time and some of this stuff looks to me like hype-work to keep the name relevant while the leadership place ever more emphasis on enterprise services and consultancy.
Maybe I'm wrong, but they aren't a company that get mentioned in the same breath as Microsoft, Google, Nvidia or Apple, not any more.
Yeah, a huge chunk of IBM's service revenue has always been effectively mainframe services, with wall street accounting spin.
I suspect that almost zero companies adopted mainframes after 1980 or so, so it's ALL legacy market. However, IBM always invests a lot of money in hardware to keep the mainframe perceptually leading edge and "sexy", so they can hold-on to those customers. So you gotta give them credit for that.
(IBM and Microsoft were always 'in the same breath' for years, Microsoft totally out-smarted them, and IBM gave up on that.)
No, this is more like any modern processor, which translates instruction codes into micro-ops. To over-simplify IBM just has two of these units per thread rather than one.
I wonder how they handle potential differences in memory barriers, instruction order scheduling and other stuff and do they run the core in one mode continuously or do the mix instruction streams from different instruction sets? Anybody got a link to an article?
Possible different micro ops for different semantics.
They also don’t mix instructions sets within the same process - the diagram I saw had ARM Linux as a guest under z/VM or KVM. For generations now no OS (not VM, not z/OS) hasn’t seen the bare metal machine, only ran under the PR/SM hypervisor, which is what does the logical partitions now.
In order to properly run OSs for the 360 and 370 generations, s390x also has instructions for setting up CPU flags to more precisely emulate older machines. From an s390x binary you can, IIRC, do a jump to an address telling it that, from the jump forward the ISA is the one of a 360 until it encounters a return, which restores 390 mode.
The diagram is the last picture, and according to it you choose the ISA at the VM level: either Linux on s390x or Linux on arm64, but not both on the same VM.
Interesting that z/VM doesn't seem to support spinning up ARM Linux VMs, at least according to this diagram, but it does support bringing up s390x Linux VMs (the LinuxONE Community Cloud creates Linux VMs under z/VM). Also, it looks like z/OS is running directly on the LPAR, which wasn't common the last time I looked (a decade ago, more or less).
I suspect that is not an inherent limitation of the dual-ISA design (the HotChips slides mention some sort of bidirectional thread state mapping between arm and z), but just about there not being z/VM release that supports that (building such a thing is probably SMoP, but another question is whether that makes business sense).
In general implementing a weaker memory model (e.g. aarch64) on a stronger memory model (e.g. x86_64 or s390x) is fairly easy, while the reverse is more difficult (see Apple's processors which have a dedicated "stronger" mode to better support execution of translated x86_64 code). It all requires some additional complexity, but starting from a complicated high-performance CISC architecture which already supports a wide range of backwards compatibility modes you are already going to have many of the building blocks on hand to support something new.
This feels like a baby step towards Arm being able to emulate z/Arch workloads, maybe with a bit of secret sauce for certain specific operations, which doesn't seem very much like IBM.
I thought about that, too. I think they’re doing this for the same reason IBM has supported Linux LPARs:
since a lot of customers who lease System/z currently probably get overprovisioned hardware, why not try a last-ditch attempt to sell the excess capacity as ARM LPARs?
IIRC, they had different licensing prices for cores that would run z/OS workloads and cores that would run Linux on s390x (and other tier for Java, I think). This looks like they’ll have one for Linux on ARM as well.
There's also the poorly advertised IBM-native S390 emulator. I worked on it over a decade ago, you could run S390 on PPC and X86 offically, and ARM and a few others were unofficial.
Not even the weirdest thing IBM did to get compatibility between their mainframes and other CPU architectures. To build the XT/370 and AT/370 expansion cards, they custom-ordered modified 68000s that decoded System/370 instructions instead of the 68k instruction set, with most of the instructions handled by the new microcode and the few stragglers software-emulated:
ARM has better software support for AI applications and half of their presentation was about their inference accelerators that can go in the mainframes (and POWER machines).
IBM mainframes are almost designed by their users. The previous generation skipped a lot of speed boost on the CPU side because their users didn’t want the machine to blow over their power delivery limits.
Now, with their architecture behind it, I’m sure these ARM Linux partitions will have the fastest ARM cores ever made. My experience with Linux on s390x is that it feels like a normal server that’s just ludicrously fast - almost as if it came from the future.
At least back in the day, there was talk about Nvidia + ppc64le. Wasn't there even a supercomputer with that setup? But I guess that has fallen by the wayside.
We have some older Power9 with NVIDIA V100. Nvidia drivers stopped a while ago (before V100 were considered “old” also for x86_64).
The main problem is software support by most machine learning libraries. While it is normal to have to compile many things from source when the binary is not provided even by some third party, in some cases the software will not compile on ppc64le and surely is not tested to work.
The funny thing is that Fujitsu has a sizable enterprise IT business, and their Primergy servers are very good, so why don't they offer their own ARM CPUs?
Mainframes are exquisitely "balanced" machines that match processing capacity and IO for a relatively small set of workloads. You won't see mainframes doing scientific number crunching or AI training, but you'll increasingly see them adding more AI inference into their transaction processing workloads for things like fraud and anomaly detection. While their CPUs are prodigiously fast (the one announced is designed to run at 5.7 GHz, and the cores themselves are huge compared to other architectures), a lot of emphasis is placed on the IO systems that connect the machine to storage and other systems so that the CPU is busy all the time. It's very normal to run a mainframe at close to 100% CPU utilization.
Because their clients want to deploy Linux AI workloads directly on the mainframe to gain lower latencies and higher troughput and ARM has better library support than s390x. Half of the last two mainframe-dedicated Hot Chips presentations were about the Spyre inference accelerator. Every newer mainframe CPU also has a big inference accelerator built-in, which is accessible to ARM binaries at instruction-level latencies (and has dedicated s390x instructions).
Sure, don't disagree, RISC-V has 10Gbit nic, 32GB, multi core machines available now. Software slower to get there but catching up quite quickly. Don't discount that too.
The article is quite short on details but I wonder if this is an ARM core that can also execute IBM Z, or the other way around; in other words, which ISA does it execute its first instruction at the reset vector? In that past I've worked on a conceptual design for an "x86+" CPU, which would start as a regular x86 but then be able to execute code of other architectures like ARM, MIPS, PPC, etc. isolated in their own code segments, similar to how V86 mode works.
IBM's mainframes do not really have a reset vector, the CPU gets initialized by some external means into whatever state the OS expects and then the clock gets enabled, so that distinction is quite moot.
Every single physical core on the chip decodes and executes both s390x and Arm AArch64 instructions, dynamically switching between the two ISA modes and convert them into micro ops. Mode switching is hypervisor-driven.
And those chips are fast. 5.7 Ghz 2nm node.
For who: When you need to run Linux programs on high security, mission critical environment. Others should not care.
There was a card just detailed on Hot Chips that used a whole bunch of RISC-V CPUs on 2nm to run at 1.8GHz, but really low power, for compute on a CXL card that could also have SSDs for cached storage.
There will be RISC-V implementions specific to many use cases of CPUs.
I have to admit we were more in the "desktop" use case, then the 7GHz (I don't know how they will manage to keep the backend reasonably fed without severe speculation though).
"AI acceleration for in transaction fraud detection", sounds interesting but also quite risky. Is it just a generic inference accelerator or something more exotic?
20 years ago, IBM had their eCLipz initiative to unify z and POWER. What actually shipped was two different CPUs sharing some design, but some of the media reports on it made it sound like it may have originally been more ambitious-one CPU with two instruction sets.
If that’s what it was, looks like they finally shipped it, just replacing POWER with ARM
97 comments
[ 0.21 ms ] story [ 4.6 ms ] threadThere was a time when IBM dominated console CPUs for a massively successful generation (PS3, Xbox 360, Nintendo Wii).
PS3 used the Cell processor but it’s debatable how much that was an asset vs handicap. Total PlayStation sales dropped for that generation vs both PS2 and PS4. The manufacturing costs didn’t fall nearly as fast as expected and it was a poor fit in terms of cross platform development etc.
PS4 moved to AMD.
IBM stopped building servers based on x86 because the margins were too thin for their tastes, but they never stopped building on top of POWER and Z.
For starters, IBM did also do the CPU for the GameCube, and that was likely a part of the Wii using an upgraded PPC arch for that.
On the flipside, there is the theory (I think even Copetti brings it up in their XBox 360 Architecture breakdown [0]) that IBM using the Cell PPE for the 360's tri-cores left a sour taste in Toshiba, but more-so Sony's mouths.
I think the big 'X factor' though, was that, for as much pain as it caused AMD in the short term, (it's so easy to forget their 'malaise' era, i.e. Early Bulldozer and the GloFo split pains to their margins[1]), AMD made all the 'right' choices to let the console vendors have their cake and eat it too.
Namely, AMD was more than happy to do a custom core if there was a volume contract (similar to what IBM was willing to for the 360/GC/Wii) but also now had a capable, in-house GPU. (And thankfully had Bobcat as a stepping point towards Jaguar[2])
There's part of me that asks, if AMD had an ARM core, if we would have all Consoles powered by AMD chips now. Nintendo likely bought into Tegra because it was an ARM core, and for better or worse their mobile stuff by the time of making that choice had 15+ years of proven ARM success (GBA, DS, 3DS) behind it.
(See also, Intel in the 2010s scrambling with half-assed promises of integrating custom functionality or FPGAs with x86 cores.)
[0] - https://www.copetti.org/writings/consoles/xbox-360/
[1] - From what I recollect, the GloFo split and how the contracts were drawn up as far as their production, had a huge impact on their ability to produce due to yields and thermals, as well as the contracts for how GloFo got paid; it was at least part of them diversifying with TSMC as soon as they reasonably could.
[2] - As a Rant, I am pretty sure, if Jaguar had Desktop/Mobile versions that included a Dual channel DDR controller, they would have cleaned up on the low cost laptop market. I had one with, I think it was an A5000 or A5200, and for how tiny the battery was it could last wayyyy longer than any of the intel laptops I had for the time, but churned if you were doing memory heavy stuff.
Hell Nvidia was so desperate they did the whole Shield thing...
It's more than a theory. It's pretty much spelled out explicitly in The Race For A New Game Machine how salty not just Sony and Toshiba were in the broad sense, but also the Sony and Toshiba engineers that the IBM team worked with felt pretty betrayed.
AMD worked on a ARM core https://en.wikipedia.org/wiki/AMD_K12 but I'm not sure it would had work, they were never the best on the power consumption, so I'm not sure it would have been a good choice for the Switch.
https://research.ibm.com/blog
You should dive into the topic of Mainframe. There are a lot of financial transaction to other critical infrastructure are dependent on it. And not just because of backward compatibility but technical superiority.
IBM has never stopped innovating. It’s just that most people can’t afford their machines.
That being said I think it's a natural consequence of the difference between a mainframe and commodity servers. A mainframe is going to be running pretty disparate workloads simultaneously, so it makes sense to steal from your neighbor if they aren't using their cache. Whereas it's more likely that a commodity server is just running the same server on each core, and if you have a different workload, you pick a different shape of server to run it on. There are pros and cons to both.
I do wonder about the spectre consequences of borrowing cache lines from other cores though.
Mere mortals such as me, have to make do with cloud-based clusters where the IO is distributed across multiple machines connected by very fast interconnects, but they don't behave as a single machine, nor have the IO always accessible at local bus speeds.
> Mainframes phone home immediately if they detect problems and an IBM service person will be on site the same day to fix it.
The joke usually went like this: technician shows up to fix the machine, operator says "We didn't call you", to what the technician answers "You didn't. Your computer did"
https://chipsandcheese.com/p/telum-ii-at-hot-chips-2024-main...
If the consumer segment started to develop a sense that their payment cards were glitchy or unreliable, our economy could suffer real consequences. Error free, secure, instant payment experiences are essential. Go visit a local grocery store and measure how long it takes to authorize a transaction each time a card is presented to the terminal. Pay special attention to how fast the Visa and American Express networks are. At my local Kroger the terminal and network setup is so fast that you can use the chip reader almost like a mag stripe reader.
https://usa.visa.com/about-visa/newsroom/press-releases.rele...
Their hardware advances are real, Power chips are still excellent and IBM's mainframes are pretty unique, but the niche for both of those seems less relevant over time and some of this stuff looks to me like hype-work to keep the name relevant while the leadership place ever more emphasis on enterprise services and consultancy.
Maybe I'm wrong, but they aren't a company that get mentioned in the same breath as Microsoft, Google, Nvidia or Apple, not any more.
I suspect that almost zero companies adopted mainframes after 1980 or so, so it's ALL legacy market. However, IBM always invests a lot of money in hardware to keep the mainframe perceptually leading edge and "sexy", so they can hold-on to those customers. So you gotta give them credit for that.
(IBM and Microsoft were always 'in the same breath' for years, Microsoft totally out-smarted them, and IBM gave up on that.)
They also don’t mix instructions sets within the same process - the diagram I saw had ARM Linux as a guest under z/VM or KVM. For generations now no OS (not VM, not z/OS) hasn’t seen the bare metal machine, only ran under the PR/SM hypervisor, which is what does the logical partitions now.
In order to properly run OSs for the 360 and 370 generations, s390x also has instructions for setting up CPU flags to more precisely emulate older machines. From an s390x binary you can, IIRC, do a jump to an address telling it that, from the jump forward the ISA is the one of a 360 until it encounters a return, which restores 390 mode.
The diagram is the last picture, and according to it you choose the ISA at the VM level: either Linux on s390x or Linux on arm64, but not both on the same VM.
(or QEMU, for partition level emulation).
https://en.wikipedia.org/w/index.php?title=PC-based_IBM_main...
https://thechipletter.substack.com/p/motorola-intel-ibm-make...
https://www.cpushack.com/2013/03/22/cpu-of-the-day-ibm-micro...
IBM mainframes are almost designed by their users. The previous generation skipped a lot of speed boost on the CPU side because their users didn’t want the machine to blow over their power delivery limits.
Now, with their architecture behind it, I’m sure these ARM Linux partitions will have the fastest ARM cores ever made. My experience with Linux on s390x is that it feels like a normal server that’s just ludicrously fast - almost as if it came from the future.
The main problem is software support by most machine learning libraries. While it is normal to have to compile many things from source when the binary is not provided even by some third party, in some cases the software will not compile on ppc64le and surely is not tested to work.
(Yes, I'm aware of Ampere Altra. The space desperately needs competition and commoditization, imho.)
If you haven't already.
Looks like this one you choose the ISA per process? Per VM?
start as a regular x86 but then be able to execute code of other architectures
https://en.wikipedia.org/wiki/Alternate_Instruction_Set
Every single physical core on the chip decodes and executes both s390x and Arm AArch64 instructions, dynamically switching between the two ISA modes and convert them into micro ops. Mode switching is hypervisor-driven.
And those chips are fast. 5.7 Ghz 2nm node.
For who: When you need to run Linux programs on high security, mission critical environment. Others should not care.
https://www.servethehome.com/ibm-z-and-linuxone-dual-isa-pro...
Why couldn't they be 6GHz? x86_64 ISA did it again?
That said IBM should add RISC-V support.
So I guess the IBM stuff is "server slow" then.
I wish RISC-V microarchitectures would get access to 2nm silicon processes, because RISC-V at 7GHz... yummy.
So, don't assume they would crank up the GHz.
I have to admit we were more in the "desktop" use case, then the 7GHz (I don't know how they will manage to keep the backend reasonably fed without severe speculation though).
AI does not always mean LLM.
If that’s what it was, looks like they finally shipped it, just replacing POWER with ARM