13 comments

[ 0.25 ms ] story [ 18.0 ms ] thread
> Consumers in the PC segment expect high performance across a wide range of applications

It seems there is a persistent rift between what hardware designers think users want and what users are asking for, only recently being addressed by major OEMs with products such as the MacBook Neo.

Nobody I personally know wants a faster CPU; they want competent, faster, native feeling software and more power efficient hardware. CPUs have been plenty fast for at least a decade now, for anything the majority of users and even gamers are doing. I'm playing through the new and notoriously demanding LEGO Batman game on an RTX 2080 and an i7-4790k, a CPU from 2013 that even the base M1 easily beats. The game runs fine. Most people I am confident do not need better than the M1.

I was under the impression that race-to-idle made "fast" a significant factor in "power efficient"? Obviously needs software to not be terrible and needs to be able to scale the back down, but still.
That may be true at iso-voltage/frequency, but there are definitely places on that curve where it is better to be at a lower voltage (and thus frequency) than to run at a higher voltage just so you can race to idle.
MacBook Neo certainly does not address my needs with a 8 GB and 512 SSD for 900 euros.

Additionally people shipping Electron crap also don't address my needs at all, everything that is Web based I already have an installed browser for it.

I think you underestimate how intensive it is to scroll LinkedIn in Edge while a YouTube video, Outlook, Adobe's updaters, antivirus, device management, and whatever incidental malware the user has picked up run in the background. That's the reality for a lot of people.
I believe that falls under competent software. However, I do see your point.
[Disclaimer: I wrote Rosetta 2, so everything I say is biased by that.]

Contemporary out-of-order CPUs are incredibly over-provisioned; microarchitects will justify a new feature by another 0.1% gain on some benchmark. The end result is a CPU that's pretty decent at handling the sort of code bloat that comes from a binary translator.

There's also a big tradeoff in adding more optimizations to a binary translator. You would like to be able to precisely handle exceptions (especially ones caused by invalid memory accesses) while presenting a userspace exception handler with an architecturally valid state for the source program. There are some optimizations that would be easy to do in principle but are painful for maintaining this mapping between source program states and translated program states. The complexity burden combined with the difficulty of debugging that added complexity (or exhaustively verifying it up-front) shapes many of your decisions when writing a production binary translator. You should always have more Cool (tm) ideas than you actually use in practice.

Would Rosetta 2 have worked as well for x86 APX i.e. 32 GPRs instead of 16?

(16x x86 regs fit in 32x Arm regs, whereas 32x in 32x is... much less of a fit)

Yeah, modern CPUs are great at executing garbage code reasonably fast. Getting binary-translated code to come close to native performance is still difficult.

Apple obviously had an advantage as they also control the hardware (and Arm helped by adding some extensions to simplify translation); significantly easing two difficult parts of translating from x86: TSO and status flags. AVX is somewhat annoying (256-bit regs -> 128-bit regs, frequent merging of scalar values), but manageable.

> You would like to be able to precisely handle exceptions (especially ones caused by invalid memory accesses) while presenting a userspace exception handler with an architecturally valid state for the source program.

This is absolutely annoying and makes many optimizations much more difficult as a lot of additional state needs to be kept around, either for real or in metadata for reconstruction (including weird status flags, fun with partially written flags (inc/dec), maybe-written flags (shift/rotate), etc.). Does Rosetta 2 always have precise status flags (including PF/AF) at every possibly-faulting memory access? (This should be rarely needed in practice, so I've never implemented flag recovery in my own binary translators (primarily for research, Instrew but also non-public).)

> TSO

yeah that one is more messy on Windows, with extensive reliance on RCpc...

> and status flags

it's part of FEAT_FlagM(2) - has been there since the Snapdragon 8cx Gen 3 on the Windows side

When I bought my M1 Macbook, I had a firsthand experience with how well binary translation worked. It was almost perfect with one MAJOR exception - anything that used JIT - so stuff like Electron apps or Java stuff (IntelliJ-based IDEs). Which was kinda ironic - JITs were designed around the idea of portability across CPU architectures, yet usually they are one of the hardest pieces of code to port, and they make this sort of binary translation approach - which has relatively long compile times but good exec time - very slow and painful.
I'd be curious to know the compiler used by the Geekbench version they tested.

MSVC2019+ now include "volatile metadata" to the output executables to make emulation on ARM faster[0][1], though it seems to be off by default since MSVC2022, I am curious to know why they disabled it.

Microsoft now also offers an "Arm64EC" ABI, that, as I understand it, helps calling into native ARM libraries from emulated x86 code by avoiding the need to translate calling conventions, but that's mostly relevant for third-party libraries in this context (I imagine that Microsoft ships all system libraries with the "emulation compatible" ABI with Prism?).

I'd be curious to see some Prism vs FEX (vs Rosetta vs Box64) benchmarks.

[0]: https://fex-emu.com/FEX-2504/#windows-pe-volatile-metadata-s...

[1]: https://learn.microsoft.com/en-us/cpp/build/reference/volati...

> though it seems to be off by default since MSVC2022

This seems to be a typo in the docs, VS2022 is 17.x and still generates volatile metadata. VS2026 is 18.x.

Last time I tested it, the penalty in Prism for running x64 code without volatile metadata was ~25% on Snapdragon X.

ARM64EC code is essentially x64 code pre-translated to ARM64. It's built against the x64 emulation ABI conventions but runs directly as native ARM64 code. Translation thunking conventions allow for cross-calling between ARM64EC and emulated x64 code.

A lot of the OS libs are shipped in Windows 11 ARM as ARM64X, so they're hybrid ARM64EC+ARM64. x86 libs like MSVCRT.DLL do still seem to be compiled with volatile metadata. DUMPBIN /LOADCONFIG reveals if volatile metadata has been included.