23 comments

[ 2.2 ms ] story [ 69.3 ms ] thread
(comment deleted)
The core guidelines library is definitely not doing the right thing here. Very odd.
Herb Sutter's comment on why it's ok is confusing to me:

> Regarding the use of UB internally: It's okay and if anyone is worried about it the use of UB is benign on the platforms we target (e.g., they don't involve hitting any hardware trap representations for these types)

Isn't the outcome of the UB (ie. whether it will "rm -rf /" or something else) dependent on both the target and the compiler? And the compiler (or future compiler) could plausibly make the assumption that the narrowing to an unrepresentable value will never occur and change behaviour because of it?

No, UB is allowed special powers for compiler and standard library implementors, which is what Herb Sutter means with internal behaviour.

Meaning MSVC is aware of these cases, so the compiler has special cases for it.

Yes, this is all true but Sutter's comment is that the specific platforms that this specific implementation of the GSL targets results in the correct output. The platforms officially supported are:

GCC 12, 13, 14

XCode 14.3.1, 15.4

Clang 16, 17, 18

Visual Studio with MSVC VS2019, VS2022

Visual Studio with LLVM VS2019, VS2022

For what it's worth, a GSL developer later reopened that GitHub issue and stated that they're going to look into fixing the UB. Sutter may have just been stating an assumption.

https://github.com/microsoft/GSL/issues/786#issuecomment-513...

> I'll raise this issue in the next internal GSL sync. I'd agree with y'all that this behavior: https://godbolt.org/z/4Tr1fe9xG is undesirable

But ... surely Sutter ought to know better than to say "because the hardware handles this conversion reasonably, it's a benign case of UB"? Surely he knows that compilers can and will optimize based on the assumption that UB never happens?

The problem isn't, "oh no what if my CPU's float->int conversion instruction traps", that's an extremely naive way to think about UB. Everyone who has thought seriously about UB in C++ for any length of time knows this. It's worrying that this was Sutter's response.

UB is bad not because it actually leads to any particular result on any particular platform or compiler, but because semantically it invalidates assumptions about a program. Rust is explicit on this, but it absolutely still applies to C/C++.
In LLVM, the result of floating-to-int conversion that is out of range of the int is a poison value, which means you get essentially the full unpredictability of UB.

That said, I'm a little hard-pressed to think of optimizations that would actually take advantage of poison, because floating-point range isn't really computed in the optimizer.

Here's (my modified version of) an example someone came up with on lobste.rs: https://godbolt.org/z/e69b4Tqbs

I don't know exactly which optimization passes do what, but a few observations:

* The 'foo(unsigned int n)' function should never return a value that's greater than 'n', since it returns 'i < n ? i : n'.

* The value printed by the 'foo' function should always be the same as the value that's returned.

Yet the value it prints is 2700624104 (which is greater than 'n', which is 10 in this case), and the returned value is 2700623376, which is different. (The exact numbers vary run to run)

If the conversion "just" resulted in a bogus value, we would have expected some number <=10 to be printed two times.

Yeah Herb's 100% wrong here. Its common when people are downplaying the memory safety issues with C++ that they say things like this, but its completely incorrect. All invoked UB is potentially equally serious, and this is exploitable memory unsafety. Compilers can and do optimise away this kind of stuff (as other people have explained here)

There's also important context in that Herb is currently one of the people leading the current memory safety approach for C++

At this point of time Herb Sutter was working for Microsoft. When he says "we" the compiler team is included.

What he means is that, it works for Microsoft as it is and zero fucks are given for other compilers and platforms.

Hopefully this will be part of UB fixes for C++29, where plenty of UB is being redefined as erroneous behaviour instead.
How could it be defined behaviour, when the result is different on ARM and x86?
Sounds like the standard should say that it results in an implementation-defined value (or wording to that effect). Saying it's UB gives the compilers way too much leeway.
> The correct fix is to bounds check before casting.

This will do wonders for speed. Actually explicitly using the safe isntr might be better. Something like this will happily compile to a single instr and cause you no grief even if the compiler had it out for you with UB. These instrs all clearly define outputs for all inputs (note that said outputs may not match across architectures)

   static inline __attribute__((always_inline)) int f2i(float myFloat) {
      int myInt;

      #if defined(__arm__)
         asm("VCVT.S32.F32 %0, %1":"=r"(myInt), "t"(myFloat));
      #elif defined (__aarch64__)
         asm("FCVTZS %0, %1":"=r"(myInt), "w"(myFloat));
      #elif defined (__x86_64__)
         asm("CVTTSS2SI %0, %1":"=r"(myInt), "x"(myFloat));
      #else
         #if 0 // be boring
            if (myFloat <= TOO_SMALL_FLOAT || myFloat => TOO_BIG_FLOAT)
               abort();
         #else
            #warning "Embrace the UB"
         #endif
         myInt = (int)myFloat;
      #endif
      return myInt;
   }
Today people think that Java's main feature was OOP, but its main selling point was "no undefined behavior" (e.g. "int" means 32-bit signed integer with overflows, on any platform, no exceptions). Today it sounds normal, but back in the day that was what made Java popular.
In C/C++ just adding two integers can lead to undefined behaviour. Your expectations are too high.