 |
 |
|
 |
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Okay, I guess everyone who has ever touched C or C++ has at least heard
rumors of this: The standard data types, such as "int", "short int" or
"long int", are anything but. For instance, a "long int" will typically
be 32 bits wide on a 32 bit machine, but 64 bits on a 64 bit machine -
unless you're running Windows, in which case it's still 32 bit. And
"int" will typically be 32 bits wide - unless you're running on a 64 bit
Cray machine, in which case it will be a whopping 64 bit as well. Or on
an embedded computer, in which case it may be as small as 16 bits. Hell,
there are even systems out there where the most fundamental data type,
"char", is not 8 but 16 bits wide!
I've learned this lesson about a decade ago (well, the thing about the
"char" data type, anyway; I might have heard about the other type woes
long before), and also that fortunately the header file <limits.h> (C)
or <climits> (C++) provides some relief: It provides a set of macros
that at least tell you what the minimum and maximum values of those
types are. For instance, "UINT_MAX" will tell you the highest number
that fits in /your/ particular "unsigned int", "SHRT_MAX" will tell you
the same for (signed) "short", and so forth.
Now the data type ambiguity has struck back with a vengeance, right
before my eyes, in the POV-Ray code:
Imagine you need to read a 32 bit integer from a file, and convert it to
a floating point value in the range from 0.0 (correspondng to integer
value 0) to 1.0 (corresponding to integer value 2^32-1). How do you do that?
Well, after reading 4 bytes from the file into a variable that is
supposedly large enough to hold those 4 bytes (we're using an unsigned
int there... whoops), you convert the value straight to floating point
format (giving you a value from 0.0 to 2.0^32-1.0), and then of course
divide by UINT_MAX...
... wait, *WHAT?*
Okay, I can understand how someone might be oblivious enough of the type
issues to shove a 32 bit value into an unsigned int without thinking
twice. But that /constant/ is there exactly because unsigned int is
/not/ guaranteed to be 32 bits wide - and we're seriously using that
very same constant with the /adamant/ presumption that it /is/?
*NOM!*
There. Another bite mark in my desk.
Needless to say, Imma throw this outta the window. (The code, not the
desk. That would bee a tad too heavy.)
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
clipka <ano### [at] anonymous org> wrote:
> Okay, I guess everyone who has ever touched C or C++ has at least heard
> rumors of this: The standard data types, such as "int", "short int" or
> "long int", are anything but. For instance, a "long int" will typically
> be 32 bits wide on a 32 bit machine, but 64 bits on a 64 bit machine -
> unless you're running Windows, in which case it's still 32 bit. And
> "int" will typically be 32 bits wide - unless you're running on a 64 bit
> Cray machine, in which case it will be a whopping 64 bit as well. Or on
> an embedded computer, in which case it may be as small as 16 bits. Hell,
> there are even systems out there where the most fundamental data type,
> "char", is not 8 but 16 bits wide!
>
> I've learned this lesson about a decade ago (well, the thing about the
> "char" data type, anyway; I might have heard about the other type woes
> long before), and also that fortunately the header file <limits.h> (C)
> or <climits> (C++) provides some relief: It provides a set of macros
> that at least tell you what the minimum and maximum values of those
> types are. For instance, "UINT_MAX" will tell you the highest number
> that fits in /your/ particular "unsigned int", "SHRT_MAX" will tell you
> the same for (signed) "short", and so forth.
>
> Now the data type ambiguity has struck back with a vengeance, right
> before my eyes, in the POV-Ray code:
>
> Imagine you need to read a 32 bit integer from a file, and convert it to
> a floating point value in the range from 0.0 (correspondng to integer
> value 0) to 1.0 (corresponding to integer value 2^32-1). How do you do that?
>
> Well, after reading 4 bytes from the file into a variable that is
> supposedly large enough to hold those 4 bytes (we're using an unsigned
> int there... whoops), you convert the value straight to floating point
> format (giving you a value from 0.0 to 2.0^32-1.0), and then of course
> divide by UINT_MAX...
>
> ... wait, *WHAT?*
>
> Okay, I can understand how someone might be oblivious enough of the type
> issues to shove a 32 bit value into an unsigned int without thinking
> twice. But that /constant/ is there exactly because unsigned int is
> /not/ guaranteed to be 32 bits wide - and we're seriously using that
> very same constant with the /adamant/ presumption that it /is/?
>
> *NOM!*
> There. Another bite mark in my desk.
>
> Needless to say, Imma throw this outta the window. (The code, not the
> desk. That would bee a tad too heavy.)
I've heard about this (Except for the char thing... that's just weird) but never
given it much thought. (Short sighted maybe)
I know that Qt Widgets provides it's own fixed width data types, but the C99
stdint.h had them too.
http://en.cppreference.com/w/cpp/types/integer
Regards,
A.D.B.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
That said: The fixed types with -exactly- 8, 16, 32 or 64 bits appear to be
optional, so there's no guarantee they'll be defined...
that would be aggravating.
http://www.azillionmonkeys.com/qed/pstdint.h
http://www.boost.org/doc/libs/1_38_0/libs/integer/index.html
https://en.wikipedia.org/wiki/C_data_types#Downloads
In all likelihood, you're probably aware of all of these, since you're code-fu
is vastly stronger than mine.
Regards,
A.D.B.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 8/22/2015 7:36 PM, Anthony D. Baye wrote:
> In all likelihood, you're probably aware of all of these, since you're code-fu
> is vastly stronger than mine.
I bet my unco fu is stronger than his ;-)
It is OT after all. :-)
--
Regards
Stephen
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Am 22.08.2015 um 20:16 schrieb Anthony D. Baye:
> I've heard about this (Except for the char thing... that's just weird) but never
> given it much thought. (Short sighted maybe)
Would that be signed or unsigned short sighted? :-P
> I know that Qt Widgets provides it's own fixed width data types, but the C99
> stdint.h had them too.
C++11 <cstdint> has them as well. (*)
Alas, C++03 is based on pre-99 C, so it doesn't have them. So POV-Ray
does its best to try and figure out matching types, but may have to fall
back to relying on the guaranteed minimum widths of the char, short,
int, long and long long types.
----------------------------
(*)
BTW, it should be noted that actually <stdint.h> or <cstdint> does *not*
necessarily have those standard fixed width data types "intN_t" and
"uintN_t" with N={8,16,32,64}, as they are only mandatory if the runtime
environment happens to supports integers of that exact width; in ILP64
environments, for instance, you could theoretically miss out on N=32
(presuming short is 16 bits wide).
Also, all the "intN_t" are unavailable if the runtime environment does
not use two's complement format for negative integers.
The only thing we're guaranteed to get are "intN_least_t",
"uintN_least_t", "intN_fast_t" and "uintN_fast_t", with N={8,16,32,64},
which are /at least/ N bits wide, with the added guarantee that the
"*_least_t" variants are the smallest and the "*_fast_t" variants the
fastest types that fit the bill.
Also note that "intN_least_t" and "intN_fast_t" may be unable to
represent -2^(N-1) (in contrast to "intN_t" which, if present, is
guaranteed to represent exactly the range from -2^(N-1) to 2^(N-1)-1).
So yeah, <stdint.h> alias <cstdint> helps... a tiny bit.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Am 22.08.2015 um 20:36 schrieb Anthony D. Baye:
> That said: The fixed types with -exactly- 8, 16, 32 or 64 bits appear to be
> optional, so there's no guarantee they'll be defined...
>
> that would be aggravating.
... ah, you noticed that alredy.
Did you also note the part where it says that the signed variants are
only present if negatives use two's complement format?
One more thing that bugs me is that standards /still/ don't provide a
straightforward way to detect the byte ordering of the standard integer
data types (unless I missed some recent news).
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
clipka <ano### [at] anonymous org> wrote:
> Am 22.08.2015 um 20:16 schrieb Anthony D. Baye:
>
> > I've heard about this (Except for the char thing... that's just weird) but never
> > given it much thought. (Short sighted maybe)
>
> Would that be signed or unsigned short sighted? :-P
>
Definitely unsigned. My hindsight is perfectly clear.
Regards,
A.D.B.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 21/08/2015 01:12 PM, clipka wrote:
> Okay, I guess everyone who has ever touched C or C++ has at least heard
> rumors of this: The standard data types, such as "int", "short int" or
> "long int", are anything but.
If I'm remembering my history right, C was basically invented to write
Unix in. From the very beginning, it was a programming language
*specifically designed* for system programming.
You know, the kind of programming where knowing exactly how many bits
you're dealing with is 100% critical.
And yet, this is one of the few programming languages on Earth which
doesn't guarantee how many bits are in a particular data type, and
provides no way to specify what you actually want.
Does that seem weird to anybody else??
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Am 23.08.2015 um 10:54 schrieb Orchid Win7 v1:
> On 21/08/2015 01:12 PM, clipka wrote:
>> Okay, I guess everyone who has ever touched C or C++ has at least heard
>> rumors of this: The standard data types, such as "int", "short int" or
>> "long int", are anything but.
>
> If I'm remembering my history right, C was basically invented to write
> Unix in. From the very beginning, it was a programming language
> *specifically designed* for system programming.
>
> You know, the kind of programming where knowing exactly how many bits
> you're dealing with is 100% critical.
That's only true /some/ of the time; and for those cases, C has bit
fields, which provide far more fine-grained control than any
guaranteed-exact-size type system could ever give you.
More to the point, system programming is the kind of programming where
performance is 100% critical /all/ of the time - and where you therefore
want to use the machine's native data types almost everywhere, rather
than some guaranteed-exact-size type system that might impose an
unnecessary overhead on your particular machine.
Also, it was designed back in the times when "portability" wasn't equal
to "interchangeability"; who cared whether your system used the same
inode size as anyone else - you wouldn't physically mount its hard drive
into another machine anyway. You wouldn't even physically mount your
removable storage media on any other machine. You only /had/ that one
machine.
Networking - yeah, that might have been a bit tedious; but back then
that was only an ever so tiny portion (and as mentioned before bit
fields would be your friend there; ever tried to assemble a raw IP frame
in Java?); most data transfer to the outside world would have been to
and from terminals, with links that would use character-based data
transfer, and hardware that would automatically trim your smallest
native data type ("char") to whatever bits per character the serial link
was configured to use - which more often than not would have been 7
rather than 8.
> And yet, this is one of the few programming languages on Earth which
> doesn't guarantee how many bits are in a particular data type, and
> provides no way to specify what you actually want.
>
> Does that seem weird to anybody else??
No, not really, for the above reasons.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA256
Le 23/08/2015 10:54, Orchid Win7 v1 a écrit :
> On 21/08/2015 01:12 PM, clipka wrote:
>> Okay, I guess everyone who has ever touched C or C++ has at least
>> heard rumors of this: The standard data types, such as "int",
>> "short int" or "long int", are anything but.
>
> If I'm remembering my history right, C was basically invented to
> write Unix in. From the very beginning, it was a programming
> language *specifically designed* for system programming.
>
But it had to deal with at least two machines, PDP-11 with 16 bits
capability (the registers where 16 bits, long integers took 2
registers, for 32 bits), and some fancy machines with 9 bits per byte,
providing 18 bits for native integers.
In such a time with different hardwares, adaptation of the language
was the PITA.
> You know, the kind of programming where knowing exactly how many
> bits you're dealing with is 100% critical.
C language provided a common minimal set of assertion:
at least 8 bits per char (but you can have more)
at least 16 bits per short
int is at least as large as short, and long int at least 32 bits.
and btw, signed integer value could, or not, be using the complement
to 2 (or 1, or anything else).
float and double... another story. (I know of a C compiler on a system
which use 48 bits for one of them, nothing like the usual 32 and 64
bits you can get used to). Ieee-757 can be used, or not.
>
> And yet, this is one of the few programming languages on Earth
> which doesn't guarantee how many bits are in a particular data
> type, and provides no way to specify what you actually want.
And where you have a hell of time to determine if you are on a little
endian, big endian, mixed endian... until you smash an union in the scop
e.
>
> Does that seem weird to anybody else??
You expect control... the target was easing port of bigger code,
including the compiler itself. At that time, porting code was costing
so much (in time & money).
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v2
iJwEAQEIAAYFAlXZ89YACgkQhKAm8mTpkW0CxQP/W5OTI5c0RDb4TABmOkWmbLp1
Puttcb9mydP3uudq+BEKOeesMkOEO0j/r2gBGNAJHQdDAx1KrO8AYmpB2315TMrp
pajSI316MdCiXL+DtkFyGHLqNXnIR7EWm8j6fXQomD8UKuClo3WFzoK50aYntXC9
e3bQ9x3WZBEp2o4TiH0=
=gdYK
-----END PGP SIGNATURE-----
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Am 23.08.2015 um 18:24 schrieb Le_Forgeron:
> and btw, signed integer value could, or not, be using the complement
> to 2 (or 1, or anything else).
Actually that would be 2's complement, 1's complement, or signed
magnitude. If your hardware would use any other format than that, you'd
need an abstraction layer to implement standard C.
I didn't know 9 bit machines actually existed, but yeah - just what I'm
saying. Imagine trying to get any decent performance out of your OS on
such a system if it was written in a language that mandated integer
types to be exactly 8, 16 or 32 bits in size.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
clipka <ano### [at] anonymous org> wrote:
> Okay, I guess everyone who has ever touched C or C++ has at least heard
> rumors of this: The standard data types, such as "int", "short int" or
> "long int", are anything but. For instance, a "long int" will typically
> be 32 bits wide on a 32 bit machine, but 64 bits on a 64 bit machine -
> unless you're running Windows, in which case it's still 32 bit. And
> "int" will typically be 32 bits wide - unless you're running on a 64 bit
> Cray machine, in which case it will be a whopping 64 bit as well. Or on
> an embedded computer, in which case it may be as small as 16 bits. Hell,
> there are even systems out there where the most fundamental data type,
> "char", is not 8 but 16 bits wide!
>
char, according to wikipedia, is supposed to be exactly one byte. Byte is
defined, in turn, to be large enough to carry any member of the basic execution
character set and UTF-8 code units. This implies that it must be at least 8 bits
wide.
POSIX requires it to be exactly 8 bits, but the exact number of bits can be
checked with CHAR_BIT from limits.h (or <climits>). It should be 8 bits on most
systems.
Regards,
A.D.B.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
I worked on such platform, not a PDP-11 (I am too young for that) but an
awkward embedded DSP chip:
http://roker.spamt.net/c++/datatypes_c55x.png
It was very annoying to do C++ on that, e.g. file I/O and std::string
didn't work as I expected. :-(
Lars R.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Am 24.08.2015 um 12:28 schrieb Lars Rohwedder:
> I worked on such platform, not a PDP-11 (I am too young for that) but an
> awkward embedded DSP chip:
>
> http://roker.spamt.net/c++/datatypes_c55x.png
>
> It was very annoying to do C++ on that, e.g. file I/O and std::string
> didn't work as I expected. :-(
Been there, done that. I once had to write a portable driver for some
proprietary communications protocol to run on a set of embedded systems,
and on one of the target platforms the message data somehow kept getting
misaligned in the buffers. The first hint that put me on the right track
was that this weirdo also had the baffling ability to store pointers in
char-wide variables, while providing 64k of RAM...
That thing was some obscure microcontroller with a built-in Bluetooth
stack, and yes - apparently some DSP capabilities as well. Can't
remember the name of the controller or the manufacturer though.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 24/08/2015 11:28 AM, Lars Rohwedder wrote:
> I worked on such platform, not a PDP-11 (I am too young for that) but an
> awkward embedded DSP chip:
>
> http://roker.spamt.net/c++/datatypes_c55x.png
>
> It was very annoying to do C++ on that, e.g. file I/O and std::string
> didn't work as I expected. :-(
I saw an FAQ page somewhere that had examples of systems where the size
of a pointer really does change depending on what it points to.
I imagine trying to do C on a Harvard architecture machine would be
"interesting" for exactly this reason. (And I gather that's quite
popular in DSP chips...)
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Am 24.08.2015 um 19:44 schrieb Orchid Win7 v1:
> On 24/08/2015 11:28 AM, Lars Rohwedder wrote:
>> I worked on such platform, not a PDP-11 (I am too young for that) but an
>> awkward embedded DSP chip:
>>
>> http://roker.spamt.net/c++/datatypes_c55x.png
>>
>> It was very annoying to do C++ on that, e.g. file I/O and std::string
>> didn't work as I expected. :-(
>
> I saw an FAQ page somewhere that had examples of systems where the size
> of a pointer really does change depending on what it points to.
>
> I imagine trying to do C on a Harvard architecture machine would be
> "interesting" for exactly this reason. (And I gather that's quite
> popular in DSP chips...)
Get a MCS-51 development environment and see for yourself...
It might also be interesting to see if and how the MCS-51 C compiler
supports the bit-addressable portion of data memory, external data
memory (which has a 16-bit address space separate from both core data
memory and instruction memory), and so forth...
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
> I saw an FAQ page somewhere that had examples of systems where the size
> of a pointer really does change depending on what it points to.
>
> I imagine trying to do C on a Harvard architecture machine would be
> "interesting" for exactly this reason. (And I gather that's quite
> popular in DSP chips...)
You don't need an exotic DSP platform, just use C or C++ on DOS. There
you have different "memory models" (TINY, SMALL, MEDIUM, COMPACT, LARGE,
HUGE), which result in different pointer sizes for code and/or data and
the way pointer arithmetics work at all.
In the SMALL model object and function pointers are incomparable. They
both are 16 bit but point to completely independent 64K memory segments.
In the MEDIUM and COMPACT models the size of object pointers and
function pointers differ (one is 16 the other is 32 bit, in the other
model the sizes are reverted)
I'd like to see that every C programmer has to learn C on such a
platform so they never ever learn to du non-ISO-C-compliant pointer
conversions. But I think it is too late for that. ;-(
Lars R.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 28/08/2015 02:37 PM, Lars Rohwedder wrote:
>> I saw an FAQ page somewhere that had examples of systems where the size
>> of a pointer really does change depending on what it points to.
>>
>> I imagine trying to do C on a Harvard architecture machine would be
>> "interesting" for exactly this reason. (And I gather that's quite
>> popular in DSP chips...)
>
> You don't need an exotic DSP platform, just use C or C++ on DOS. There
> you have different "memory models" (TINY, SMALL, MEDIUM, COMPACT, LARGE,
> HUGE), which result in different pointer sizes for code and/or data and
> the way pointer arithmetics work at all.
Oh, hello LPCWSTR, I didn't see you there... :-}
(In other words: Yes, 30 years later, the Win32 API still has these
obsolete type designations in it. Oh goodie.)
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
clipka <ano### [at] anonymous org> wrote:
> One more thing that bugs me is that standards /still/ don't provide a
> straightforward way to detect the byte ordering of the standard integer
> data types (unless I missed some recent news).
You can detect it at runtime, but it's impossible to detect it at
compile time. (Yes, I have researched it. It's just not possible.)
(Ok, I'm not 100% certain that the runtime trick is 100% standard
kosher, but it works in all relevant modern hardware.)
--
- Warp
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Orchid Win7 v1 <voi### [at] dev null> wrote:
> And yet, this is one of the few programming languages on Earth which
> doesn't guarantee how many bits are in a particular data type, and
> provides no way to specify what you actually want.
C doesn't want to force the compiler to generate inefficient code
behind the scenes in order to support some particular type. For example,
if the target hardware doesn't support, let's say, 32-bit integers,
C doesn't want to force the compiler to generate inefficient code
that handles 32-bit integers on that system. It allows the compiler to
refuse support.
In principle C tries to be as portable as possible in the sense that it
makes no assumptions about the target hardware. It doesn't even assume
that a byte is 8 bits, and it doesn't make any assumptions about the
bitness of the target platform. (I think that in principle it doesn't
even assume that integral variables use 2's complement representation.)
The only guarantee that it gives is that sizeof(char) <= sizeof(short)
<= sizeof(int) <= sizeof(long). (There might have been a guarantee that
sizeof(char) < sizeof(long), but I don't remember if that's true.)
--
- Warp
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Orchid Win7 v1 <voi### [at] dev null> wrote:
> I saw an FAQ page somewhere that had examples of systems where the size
> of a pointer really does change depending on what it points to.
Many C programmers assume that you can convert any pointer to void*
and back, as they assume that all pointers have the same size.
In C++ you can't always make that assumption. More precisely, a pointer
to a virtual function is (most probably) larger than a regular function
pointer. A pointer-to-virtual-function cannot be converted to void* and
back without malfunction.
--
- Warp
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|
 |