Despite the fact they're two different languages, people are often taught as though it's basically just a subset relationship and then they carry this through into the actual software they write.
Because neither of these are memory safe languages, you are required to ensure you've made no mistakes or else anything might happen.
The example included below "Why C and C++ Differ" is UB in both C and C++ (and says nothing about "why" C & C++ differ). The article isn't overall wrong but that bit is just...
(Feels a bit like LLM junk, but honestly not sure.)
Because if the comment is correct, it's useful signal for other potential readers to know that they can skip it.
I try to be as charitable as possible when reading and commenting online, but we don't have infinite attention and thanks to AI, it's easier than ever to end up wasting it on things that aren't worth reading.
(I'm not claiming that the article here is or is not an example of that.)
I just watched video[0] that argues if you can consteval something, then it must be free of UB, which seems compelling, at least for these lower level helpers.
Unfortunately, it seems memcpy is not constexpr, so cannot be used like this, but std::bit_cast actually works[1] as constexpr, so I think in C++ that would now be the most preferred use. It won't allow the union either.
Careful, only language UB is guaranteed to be detected in constant evaluation, UB in standard library functions isn't. It is a good sanity check though.
Also it seems that bit_cast in particular is going to be strengthened in this regard:
I really hope you're talking mostly about header files; for actual code this is such a horrible idea (and in a lot of cases just won't work without massive efforts.) Even for header files, arguably one ought to really know what they're doing.
For sure, a union which is used as API surface to a C library which a C++ codebase uses is the most obvious and dangerous footgun here.
The article never even discusses how C++ actually implements unions to begin with, so it's just incomplete. (Only the "active" member is meaningful, the others are considered as meaningless and shouldn't be touched.)
You're halfway there. It's a little more messed up than that.
IIRC, this is all little-endian. So assume our 8 bytes labeled A-H. For a 32 bit/4 byte integer, they will be read as DCBA. For a 64 bit/8 byte integer, HGFEDCBA.
So the struct is set up like a = DBCA, b = HGFE. So when we assign 2 and 3 to a and b, it should look like this in memory:
00000010 00000000 00000000 00000000
and
00000011 00000000 00000000 00000000
When we cast that to a uint64 and assign 4 to it, we should wind up with:
Does reinterpret_cast works in c++ for this?
Yeah, you could use it for that, but it’s got some caveats and is unsafe in some cases.
https://en.cppreference.com/cpp/language/reinterpret_cast
EDIT: If you’re able to use C++20 then clearly bit_cast is better.
reinterpret_cast<> for punning purposes can lead to UB. The recommended C++ alternative to avoid UB is bit_cast<>
https://stackoverflow.com/questions/53401654/why-was-stdbit-...
No.
What's so genuine about this?
Despite the fact they're two different languages, people are often taught as though it's basically just a subset relationship and then they carry this through into the actual software they write.
Because neither of these are memory safe languages, you are required to ensure you've made no mistakes or else anything might happen.
It continues with
This smells like Claudism/LLMish.
The example included below "Why C and C++ Differ" is UB in both C and C++ (and says nothing about "why" C & C++ differ). The article isn't overall wrong but that bit is just...
(Feels a bit like LLM junk, but honestly not sure.)
Feels like LLM to me. I looked at some of the rest of the article and... pretty much certain now.
Why leave such disparaging comments?
Because if the comment is correct, it's useful signal for other potential readers to know that they can skip it.
I try to be as charitable as possible when reading and commenting online, but we don't have infinite attention and thanks to AI, it's easier than ever to end up wasting it on things that aren't worth reading.
(I'm not claiming that the article here is or is not an example of that.)
Even if it's not from an LLM, it's still wrong! You've focused on the last important of that comment, really.
Is it necessary to use the lowest possible font weight?
I just watched video[0] that argues if you can consteval something, then it must be free of UB, which seems compelling, at least for these lower level helpers.
Unfortunately, it seems memcpy is not constexpr, so cannot be used like this, but std::bit_cast actually works[1] as constexpr, so I think in C++ that would now be the most preferred use. It won't allow the union either.
[0] https://www.youtube.com/watch?v=-LAXqqqX274 [1] https://godbolt.org/z/xGzjTMGvv
__builtin_memcpy is constexpr under clang with constraints
Careful, only language UB is guaranteed to be detected in constant evaluation, UB in standard library functions isn't. It is a good sanity check though.
Also it seems that bit_cast in particular is going to be strengthened in this regard:
https://cplusplus.github.io/LWG/issue4539
The conclusion is simply wrong. `union` should not be used for type-punning in C++, or in C code which might be compiled by a C++ compiler.
I can't downvote posts yet, but if you have the ability, consider it. This is clearly LLM slop without human review.
I really hope you're talking mostly about header files; for actual code this is such a horrible idea (and in a lot of cases just won't work without massive efforts.) Even for header files, arguably one ought to really know what they're doing.
For sure, a union which is used as API surface to a C library which a C++ codebase uses is the most obvious and dangerous footgun here.
The article never even discusses how C++ actually implements unions to begin with, so it's just incomplete. (Only the "active" member is meaningful, the others are considered as meaningless and shouldn't be touched.)
I don't understand the example in "Why C and C++ Differ".
Shouldn't the code return 3 in the base case and 4 if the check for 2 was assumed to always hold ?
You're halfway there. It's a little more messed up than that.
IIRC, this is all little-endian. So assume our 8 bytes labeled A-H. For a 32 bit/4 byte integer, they will be read as DCBA. For a 64 bit/8 byte integer, HGFEDCBA.
So the struct is set up like a = DBCA, b = HGFE. So when we assign 2 and 3 to a and b, it should look like this in memory:
00000010 00000000 00000000 00000000 and 00000011 00000000 00000000 00000000
When we cast that to a uint64 and assign 4 to it, we should wind up with:
00000100 00000000 00000000 00000000 00000000 00000000 00000000 00000000
Which effectively zeros out b.
So if the conditional is evaluated, it will evaluate to false, we return b, which is 0.
If the conditional is not evaluated, we should return the value in a, which is 4.
Oooh thanks, I missed that the cast was into a 64 bit integer
Hackers don't make Quake III Arena by trusting C rules or compiler optimizations anyway.