POV-Ray : Newsgroups : povray.unofficial.patches : PovRay faster Server Time
10 Oct 2026 14:11:49 EDT (-0400)
  PovRay faster (Message 51 to 89 of 89)  
<<< Previous 50 Messages Goto Initial 50 Messages
From: Mark Wagner
Subject: Re: PovRay faster
Date: 1 Feb 2001 00:03:59
Message: <3a78ee3f@news.povray.org>
Thorsten Froehlich wrote in message <3a78cef2$1@news.povray.org>...

>pvengine.exe (at least the VC 6 one) is 1622016 bytes.  On the Mac you get
>about the same size.


That includes some rather huge bitmap images.  Eliminate those, and you have
a much smaller executable.

--
Mark


Post a reply to this message

From: Ben Chambers
Subject: Re: PovRay faster
Date: 1 Feb 2001 00:40:20
Message: <3A78F6DC.886120E9@hotmail.com>
Ken wrote:

> And thus the reason Chris Cason offers two compiles for Windows. One with
> Watcom the other MSVC6++ both pentium II optimized. The price of each of
> these two compilers is not cheap and requires that the developer has access
> to equipment with archetecture the compiler is optimized for.

I beg your pardon, but there's only one official version for windows (at least on
povray.org).  If there is a faster compile available, I'd like to hear of it.
...Chambers


Post a reply to this message

From: Ben Chambers
Subject: Re: PovRay faster
Date: 1 Feb 2001 00:54:09
Message: <3A78FA1A.6B5B9F5A@hotmail.com>
Hmm... I think everyone is starting (?) to go about this the wrong way... Myself
included...
Ok, what not simply keep the official version the way it is, and _somebody_
offer a platform specific, _unofficial_ compile?  I know someone recently did a
3dNow version (sorry, I don't remember who), this would be greatly appreciated
if you can find a good way of posting it on the web...

...Chambers


Post a reply to this message

From: Ken
Subject: Re: PovRay faster
Date: 1 Feb 2001 01:51:39
Message: <3A790808.90F9A335@pacbell.net>
Ben Chambers wrote:

> I beg your pardon,

You are pardoned :)

> but there's only one official version for windows (at least on
> povray.org).  If there is a faster compile available, I'd like to hear of it.

The official distribution has the Watcom compile. There is an official
MSVC6++ compile also available that my personal testing shows a reasonable
perfomance increase (see my comparison testing report in the announcements
group). You can download the MSVC6++ version directly from the link provided
below.

ftp://povray.org/pub/povray/Official/Windows/pve-vc6.zip

-- 
Ken Tyler - 1400+ POV-Ray, Graphics, 3D Rendering, and Raytracing Links:
http://home.pacbell.net/tylereng/index.html http://www.povray.org/links/


Post a reply to this message

From: Ken
Subject: Re: PovRay faster
Date: 1 Feb 2001 01:55:33
Message: <3A7908F3.6D88A107@pacbell.net>
Ben Chambers wrote:
> 
> Hmm... I think everyone is starting (?) to go about this the wrong way... Myself
> included...
> Ok, what not simply keep the official version the way it is, and _somebody_
> offer a platform specific, _unofficial_ compile?  I know someone recently did a
> 3dNow version (sorry, I don't remember who), this would be greatly appreciated
> if you can find a good way of posting it on the web...

Historically this is the way it has been handled. There have been
many patches released using various compilers and optimizing routines
over the years. Some work better than others and it takes the pressure
off of the POV-Team from having to purchase additional software and
hardware. It also minimizes their support headaches (to a certain
degree).

-- 
Ken Tyler - 1400+ POV-Ray, Graphics, 3D Rendering, and Raytracing Links:
http://home.pacbell.net/tylereng/index.html http://www.povray.org/links/


Post a reply to this message

From: Wlodzimierz ABX Skiba
Subject: Re: PovRay faster
Date: 1 Feb 2001 03:10:37
Message: <3a7919fd$1@news.povray.org>
Mark Wagner wrote in message <3a78ee3f@news.povray.org>...
> Thorsten Froehlich wrote in message <3a78cef2$1@news.povray.org>...
> >pvengine.exe (at least the VC 6 one) is 1622016 bytes.  On the Mac you get
> >about the same size.
>
> That includes some rather huge bitmap images.  Eliminate those, and you have
> a much smaller executable.


dos version of megapov06a compiled with djgpp is about 1.2 MB

ABX


Post a reply to this message

From: Warp
Subject: Re: PovRay faster
Date: 1 Feb 2001 05:18:45
Message: <3a793804@news.povray.org>
Peter J. Holzer <hjp### [at] sikituwsracat> wrote:
: a dll with the time critical routines

  What do you mean "a dll"?
  How do you use dll's with *STANDARD* C (or C++)?
  Don't forget that povray has to be compiled for several platforms. Also
for Linux (I don't think that Linux supports Windows DLL's), not to talk
about other platforms as well (several Unix variants, MacOS, DOS, BeOS, etc).

-- 
char*i="b[7FK@`3NB6>B:b3O6>:B:b3O6><`3:;8:6f733:>::b?7B>:>^B>C73;S1";
main(_,c,m){for(m=32;c=*i++-49;c&m?puts(""):m)for(_=(
c/4)&7;putchar(m),_--?m:(_=(1<<(c&3))-1,(m^=3)&3););}    /*- Warp -*/


Post a reply to this message

From: Nikole & Joe Arcaro
Subject: Re: PovRay faster
Date: 1 Feb 2001 05:34:26
Message: <3a793bb2@news.povray.org>
Has anyone considered support for multiprocessor systems ?

ie multithreaded POV ...

I can get som free time on a quad Xeon System, but it would be wasted on
pov.

Joe ...


"Wlodzimierz ABX Skiba" <abx### [at] abxartpl> wrote in message
news:3a7919fd$1@news.povray.org...
> Mark Wagner wrote in message <3a78ee3f@news.povray.org>...
> > Thorsten Froehlich wrote in message <3a78cef2$1@news.povray.org>...
> > >pvengine.exe (at least the VC 6 one) is 1622016 bytes.  On the Mac you
get
> > >about the same size.
> >
> > That includes some rather huge bitmap images.  Eliminate those, and you
have
> > a much smaller executable.
>
>
> dos version of megapov06a compiled with djgpp is about 1.2 MB
>
> ABX
>
>


Post a reply to this message

From: Christophe Bouffartigue
Subject: Re: PovRay faster
Date: 1 Feb 2001 05:45:27
Message: <3A793E47.9846D38C@nanterre.marelli.fr>
Nikole & Joe Arcaro wrote:
> 
> Has anyone considered support for multiprocessor systems ?
> 
> ie multithreaded POV ...

http://www.wozzeck.net/images/pmp/

Bouf.


Post a reply to this message

From: A B 
Subject: Re: PovRay faster
Date: 1 Feb 2001 09:26:40
Message: <3A79721F.FA61D6C3@neuro.informatik.uni-ulm.de>
Ken wrote:
> 
> Warp wrote:
> >
> >   Is it worth the efforts to make this huge amount of work just to
> > optimize one executable binary for one CPU?
> 
> As someone who owns an Athlon machine I consider that a silly question :)
> 
> Besides there are now more than just one person using this CPU. Lots of
> people are buying them.
> 
> --
> Ken Tyler - 1400+ POV-Ray, Graphics, 3D Rendering, and Raytracing Links:
> http://home.pacbell.net/tylereng/index.html http://www.povray.org/links/

I think the whole discussion is obsolete. In the time an cpu optimized
version of POV Ray runs stable and without a few errors more the the
original version, there is a new processor under which the original POV
will run faster than the optimised. Therefor I value all the work for
cpu-optimized POV versions for lost time. 

My personal opinion.
Yours Axel Baune


Post a reply to this message

From: Ron Parker
Subject: Re: PovRay faster
Date: 1 Feb 2001 09:34:18
Message: <slrn97isvc.18h.ron.parker@fwi.com>
On Thu, 01 Feb 2001 15:26:39 +0100, A.B. wrote:
>I think the whole discussion is obsolete. In the time an cpu optimized
>version of POV Ray runs stable and without a few errors more the the
>original version, there is a new processor under which the original POV
>will run faster than the optimised. Therefor I value all the work for
>cpu-optimized POV versions for lost time. 

Except that the optimized version will run still faster under the new
processor than the original version.  There's a place for optimization,
but it has to be done carefully, with an eye to portability.


-- 
Ron Parker   http://www2.fwi.com/~parkerr/traces.html
My opinions.  Mine.  Not anyone else's.


Post a reply to this message

From: Daniel Jungmann
Subject: Re: PovRay faster
Date: 1 Feb 2001 11:10:46
Message: <3a798a86@news.povray.org>
For 1 I need not to rewrite the compiler, I need to write a lot of macros
( not exactly the same, a lot of work) and I need to analyse the source
code, because the code must bee "parallel" so the SIMD instructions can
work. This changes are "native" C and won't speed down the "normal" version
of PovRay. And a special compiler which supports MMX, 3D!Now and SSE. I just
have a beta version of this compiler, the full version cost 499.-- US
dollar.


Post a reply to this message

From: Daniel Jungmann
Subject: Re: PovRay faster
Date: 1 Feb 2001 11:23:30
Message: <3a798d82@news.povray.org>
I think every new processor will have SIMD instrctions and a
codeoptimization in that direction would speed up PovRay without processor
dependent instructions.


Post a reply to this message

From: Tony[B]
Subject: Re: PovRay faster
Date: 1 Feb 2001 12:25:11
Message: <3a799bf7@news.povray.org>
> And a special compiler which supports MMX, 3D!Now and SSE. I just
> have a beta version of this compiler, the full version cost 499.-- US
> dollar.

I am searching around the web for compilers that you can use... I've found
the following

http://www.cs.virginia.edu/~lcc-win32/
Free. Supports MMX and 3DNow.

http://developer.intel.com/software/products/compilers/c50/index.htm
Free for 30 days. Works if you have Visual Studio. Supports all the stuff in
Intel processors.

http://www.codeplay.com/
"VectorC" with support for MMX, 3DNow! and Intel SIMD. Also integrates into
Visual Studio. There doesn't seem to be a free version (that we could use to
compile POV-Ray), but it's cost ranges from $80 for the special version (no
Athlon or Pentium 4 support, slower compiling and less frequent updates), to
$750 for the pro. version. This one if listed on AMD's developer area.

That's pretty much it. I can't find any more free/sharware compilers that
will do optimizations for this stuff you're adding. Keep up the good work!


Post a reply to this message

From: Daniel Jungmann
Subject: Re: PovRay faster
Date: 1 Feb 2001 12:33:35
Message: <3a799def@news.povray.org>
Thank you, I will try try them.


Post a reply to this message

From: Alessandro Coppo
Subject: Re: PovRay faster
Date: 1 Feb 2001 18:57:41
Message: <3a79f7f5@news.povray.org>
About the DLL problem (I was already flamed to death in past for this
subject).

DLLs are binary modules which contain code loadable and callable at runtime.
Windows has DLLs. Linux, BSD, Solaris etc.etc. have shared libraries (look
into /usr/lib for all those .so files...). VMS has shared libraries. I
presume that Mac's have a similar mechanism, OSX being a Mach kernel + GNU
tools is in the U*X pack.

The only OS which I know for sure that 1) has a pov version 2) has nothing
like DLLs is MS-DOS. So, should you decide that MSDOS is dead (well,
actually it was never alive...) you can assume that you have the .dll/.so
capability in ANY case.

With respect to optimizations, you should not put too-atomic ops into
separate code, but try to create macro-ops which, even with the added
overhead of call, gain speed (I don't know if it is possible). A proposal
for a custom compile: create a multithreaded engine which can take advantage
of multi CPU boards (render a different set of lines on each CPU and merge
the result.). The speedup would be almost linear (what about a quad 1GHz
Athlon PC?)

Bye!!!

Alessandro Coppo
a.c### [at] iolit


Post a reply to this message

From: Chris Huff
Subject: Re: PovRay faster
Date: 1 Feb 2001 21:59:35
Message: <chrishuff-CFE08E.22011001022001@news.povray.org>
In article <3a79f7f5@news.povray.org>, "Alessandro Coppo" 
<a.c### [at] iolit> wrote:

> DLLs are binary modules which contain code loadable and callable at 
> runtime. Windows has DLLs. Linux, BSD, Solaris etc.etc. have shared 
> libraries (look into /usr/lib for all those .so files...). VMS has 
> shared libraries. I presume that Mac's have a similar mechanism, OSX 
> being a Mach kernel + GNU tools is in the U*X pack.

However, they all do things in their own platform/OS-specific way, so 
you still have to (re)write code for handling the libraries for each 
platform. Handling the Linux/BSD Unix/Mac OS X group might not be too 
difficult, but add in Mac OS versions before and including 9.1, and the 
various Windows versions, and other OS's, and you have a problem. There 
is no standard C mechanism for doing this kind of thing.


> A proposal for a custom compile: create a multithreaded engine which 
> can take advantage of multi CPU boards (render a different set of 
> lines on each CPU and merge the result.). The speedup would be almost 
> linear (what about a quad 1GHz Athlon PC?)

This is a different matter completely...and very difficult to do 
properly, because POV-Ray was never designed to have multiple threads. 
Even if you ignore cross-platform differences, you have problems getting 
it to share data among the different threads properly (for example, 
radiosity, maybe photon maps, etc.). If you don't do this, you end up 
with visible artifacts where parts of the image rendered on different 
threads join together. You would need to do a lot of reworking on the 
source code to get it to work at all.
I think there is a multithreaded version of POV-Ray available (PVM-POV), 
but I don't know much about it, and I don't think it works well with 
radiosity. I don't even know if it does things on a per-image basis, it 
may just distribute different frames to different processors, or give 
different chunks of the image to different processors.

-- 
Christopher James Huff
Personal: chr### [at] maccom, http://homepage.mac.com/chrishuff/
TAG: chr### [at] tagpovrayorg, http://tag.povray.org/

<><


Post a reply to this message

From: Thorsten Froehlich
Subject: Re: PovRay faster
Date: 2 Feb 2001 00:14:49
Message: <3a7a4249@news.povray.org>
In article <3a79f7f5@news.povray.org> , "Alessandro Coppo" <a.c### [at] iolit>
wrote:

> About the DLL problem (I was already flamed to death in past for this
> subject).
>
> DLLs are binary modules which contain code loadable and callable at runtime.
> Windows has DLLs. Linux, BSD, Solaris etc.etc. have shared libraries (look
> into /usr/lib for all those .so files...). VMS has shared libraries. I
> presume that Mac's have a similar mechanism, OSX being a Mach kernel + GNU
> tools is in the U*X pack.
>
> The only OS which I know for sure that 1) has a pov version 2) has nothing
> like DLLs is MS-DOS. So, should you decide that MSDOS is dead (well,
> actually it was never alive...) you can assume that you have the .dll/.so
> capability in ANY case.

Nobody said you cannot (easily) overcome the "DLL problem", but you have to
keep POVLegal in mind!

     Thorsten


Post a reply to this message

From: Peter Popov
Subject: Re: PovRay faster
Date: 2 Feb 2001 01:18:57
Message: <d5kk7tcvc6fc8vcpoi1toteiimldnjg2m3@4ax.com>
On Thu, 01 Feb 2001 22:01:10 -0500, Chris Huff <chr### [at] maccom>
wrote:

>I think there is a multithreaded version of POV-Ray available (PVM-POV), 
>but I don't know much about it, and I don't think it works well with 
>radiosity. I don't even know if it does things on a per-image basis, it 
>may just distribute different frames to different processors, or give 
>different chunks of the image to different processors.

If you mean PMP (PVM MegaPOV), it does work with radiosity. Francois
Dispot was so kind as to keep me informed about the progress while he
and the original PVM-POV author were developing it and from what I've
seen, PMP does a pretty decent job with radiosity. It shares octree
data among the slaves. However, all rad calcs have to be done on
preview stage. Luckily, there are rad options to ensure that.


Peter Popov ICQ : 15002700
Personal e-mail : pet### [at] vipbg
TAG      e-mail : pet### [at] tagpovrayorg


Post a reply to this message

From: Rk
Subject: Re: PovRay faster
Date: 2 Feb 2001 03:23:48
Message: <3a7a6e94$1@news.povray.org>
Here's some useful links for 3dNow! optimization:

Part 1 of the tutorial:
http://www.amdzone.com/articleview.cfm?articleID=282

Part 2 of the tutorial:
http://www.pcstats.com/articleview.cfm?articleID=354

Microsoft provides an MSVC++ 6 processor pack for 3DNow:
http://msdn.microsoft.com/vstudio/downloads/ppack/download.asp

Cheers
Rk


Post a reply to this message

From: Daniel Jungmann
Subject: Re: PovRay faster
Date: 2 Feb 2001 08:14:43
Message: <3a7ab2c3@news.povray.org>
> Microsoft provides an MSVC++ 6 processor pack for 3DNow:
> http://msdn.microsoft.com/vstudio/downloads/ppack/download.asp

The MSVC++ 6.0 processor pack works and now I use it.


Post a reply to this message

From: Peter J  Holzer
Subject: Re: PovRay faster
Date: 3 Feb 2001 18:02:13
Message: <slrn97p19e.fh7.hjp-usenet@teal.h.hjp.at>
On 2001-02-01 10:18, Warp <war### [at] tagpovrayorg> wrote:
>Peter J. Holzer <hjp### [at] sikituwsracat> wrote:
>: a dll with the time critical routines
>
>  What do you mean "a dll"?
>  How do you use dll's with *STANDARD* C (or C++)?

You don't have to for two reasons:

1) Creating and using a DLL is normally done with special options to the
    linker. No change to the source code is necessary, unless you need
    to decide at run time which DLL you want to load, which isn't
    necessary in this case. The compilation and linking process is
    system dependent anyway, so an additional flag for the linker won't
    hurt.

2) The optimizations Daniel are talking about are extremely system
    dependent anyway. If you start replacing major parts of the source
    code with inline assembly which will only compile with a certain
    compiler and run on a certain platform, adding a bit of code which
    handles loading a shared library on that platform (should it really
    be necessary. which I doubt) is the least of your worries.

>  Don't forget that povray has to be compiled for several platforms. Also
>for Linux (I don't think that Linux supports Windows DLL's),

No, it hasn't windows EXE files either, so why should it have Windows
DLLs? It has its own scheme for shared libraries.

In fact I am using povray on Linux (and sometimes HP-UX and Solaris),
not Windows.

	hp

-- 
   _  | Peter J. Holzer    | All Linux applications run on Solaris,
|_|_) | Sysadmin WSR       | which is our implementation of Linux.
| |   | hjp### [at] wsracat      | 
__/   | http://www.hjp.at/ |	-- Scott McNealy, Dec. 2000


Post a reply to this message

From: Alessandro Coppo
Subject: Re: PovRay faster
Date: 4 Feb 2001 03:26:56
Message: <3a7d1250@news.povray.org>
On Windows you have the LoadLibrary/GetProcAddress/FreeLibrary calls. On
Linux you have almost identical calls (dlopen and friends). I have no
experience of other OSes, but I think things cannot be much different. In a
future version of my open source C++ library (CXL, you can find it on my
site) I plan to create a wrapper for these things making them transparent to
handle either on Windows and on Linux

Just as an example, if you look at CXL-3patch1 sources, you will find that I
have a HiResTimer class which is implemented in two completely different
ways on Windows and on Linux, but when use it you write EXACTLY the same
code (and you have almost exactly the same results).

Creating a DLL requires some macro support for code (EXPORTS etc. you can
even make the source capable of being compiled into an app/lib or dll
TRANSPARENTLY) some linker options and a runtime loading mechanism (provided
by any modern OS,  EVEN Windows has it!!!).

Alessandro Coppo
a.c### [at] iolit
www.geocities.com/alexcoppo


Post a reply to this message

From: Daniel Jungmann
Subject: Re: PovRay faster
Date: 4 Feb 2001 05:56:41
Message: <3a7d3569@news.povray.org>
I think the discussion was not helpful. There are a some things which can be
optimized without any quality lost. The general things are platform
independent, the processor independent things can be put in separate source
files. PovRay would check the processor and the supported instructions (MMX,
3D!Now, SSE etc.) and the required instructions are not available PovRay
shows a message and show you where you can get the right version for your
processor. Here are a short list of things which can be optimized :

single precision which can be optimized using 3D!Now or SSE

1) Lightning
2) Mesh

general things which can be optimized

1) loops
2) float divisions (replace mutiple divisions with one division and
multiplications)
3) if then / ? :
4) integer multiplications and divisions (replace them with shift or
additions)
5) multiple integer calculations

integer arithmetic which can be otimized using MMX

1) every time multiple integers are calculate

other things which can be optimized using 3D!Now, MMX, SSE etc.

1) mathematical functions, e.g. sin, cos, tan, exp, sqr, sqrt, division etc.
2) ? :


Post a reply to this message

From: Francois Labreque
Subject: Re: PovRay faster
Date: 4 Feb 2001 09:09:39
Message: <3A7D61E8.7AC8AA0C@videotron.ca>
Daniel Jungmann wrote:
> 
> general things which can be optimized
> 
> 3) if then / ? :

Apart from "looking cooler", is there any benefit to using the ternary
operator?  Is it really faster than an "if then else" statement?  I was
under the impression that most compilers would produce identical code
underneath.

-- 
Francois Labreque | The surest sign of the existence of extra-
    flabreque     | terrestrial intelligence is that they never
        @         | bothered to come down here and visit us!
  videotron.ca                                  - Calvin


Post a reply to this message

From: Thorsten Froehlich
Subject: Re: PovRay faster
Date: 4 Feb 2001 09:50:27
Message: <3a7d6c33@news.povray.org>
In article <3a7d3569@news.povray.org> , "Daniel Jungmann" <DSJ### [at] gmxnet> 
wrote:

> general things which can be optimized
>
> 1) loops
> 2) float divisions (replace mutiple divisions with one division and
> multiplications)
> 3) if then / ? :
> 4) integer multiplications and divisions (replace them with shift or
> additions)
> 5) multiple integer calculations

It is a complete waste of time to optimize these by hand.  Every compiler
can do a much better job on these and the code stays readable on your site!

If you want to seriously speed up POV-Ray, look into the intersection
algorithms and find ways how to reduce the number of intersections further.
Bounding boxes are a start, but maybe there are additional algorithms?


     Thorsten


Post a reply to this message

From: Daniel Jungmann
Subject: Re: PovRay faster
Date: 4 Feb 2001 10:54:58
Message: <3a7d7b52$1@news.povray.org>
I don't know, I just mean "if then" and "? :", not replace "if then" with "?
:".


Post a reply to this message

From: Daniel Jungmann
Subject: Re: PovRay faster
Date: 4 Feb 2001 11:05:19
Message: <3a7d7dbf@news.povray.org>
> It is a complete waste of time to optimize these by hand.  Every compiler
> can do a much better job on these and the code stays readable on your
site!
>
> If you want to seriously speed up POV-Ray, look into the intersection
> algorithms and find ways how to reduce the number of intersections
further.
> Bounding boxes are a start, but maybe there are additional algorithms?
>
>
>      Thorsten

Which compiler does these things?


Post a reply to this message

From: Ken
Subject: Re: PovRay faster
Date: 4 Feb 2001 11:07:42
Message: <3A7D7EEB.974E7C5C@pacbell.net>
Daniel Jungmann wrote:
> 
> > It is a complete waste of time to optimize these by hand.  Every compiler
> > can do a much better job on these and the code stays readable on your
> site!
> >
> > If you want to seriously speed up POV-Ray, look into the intersection
> > algorithms and find ways how to reduce the number of intersections
> further.
> > Bounding boxes are a start, but maybe there are additional algorithms?
> >
> >
> >      Thorsten
> 
> Which compiler does these things?

Compilers do not do these things. Programmers do.

-- 
Ken Tyler


Post a reply to this message

From: Ron Parker
Subject: Re: PovRay faster
Date: 4 Feb 2001 11:59:18
Message: <slrn97r2gb.3hg.ron.parker@fwi.com>
On Sun, 04 Feb 2001 08:10:19 -0800, Ken wrote:
>Compilers do not do these things. Programmers do.

I think he meant "which compiler [optimizes integer multiplications etc.]"
and the answer is, every one I've seen since I started writing C ten years
ago.

-- 
Ron Parker   http://www2.fwi.com/~parkerr/traces.html
My opinions.  Mine.  Not anyone else's.


Post a reply to this message

From: Peter J  Holzer
Subject: Re: PovRay faster
Date: 4 Feb 2001 12:01:48
Message: <slrn97qvc2.hut.hjp-usenet@teal.h.hjp.at>
On 2001-02-01 16:12, Daniel Jungmann <DSJ### [at] gmxnet> wrote:
>For 1 I need not to rewrite the compiler, I need to write a lot of macros
>( not exactly the same, a lot of work) and I need to analyse the source
>code, because the code must bee "parallel" so the SIMD instructions can
>work.

No, this was possibility number 2, not 1. Please read what people write,
or ask if something isn't expressed clearly enough.

Number 1 was the ideal case where you don't have to change the source
code at all, because the compiler is smart enough to figure out inherent
parallelisms.

Number 2 was the less ideal case, where the compiler knows about SIMD
instructions, but can use them only if the code is arranged in a certain
way. 

In reality, early versions of optimizers will recognize only a few
constructs (situation number 2), and each successive version will
recognize a few more. At the same time, programmers will adapt their
programming style to the capabilities of the optimizers, so the
situation will asymptotically approach situation number 1.

	hp

-- 
   _  | Peter J. Holzer    | All Linux applications run on Solaris,
|_|_) | Sysadmin WSR       | which is our implementation of Linux.
| |   | hjp### [at] wsracat      | 
__/   | http://www.hjp.at/ |	-- Scott McNealy, Dec. 2000


Post a reply to this message

From: Thorsten Froehlich
Subject: Re: PovRay faster
Date: 4 Feb 2001 21:04:09
Message: <3a7e0a19@news.povray.org>
In article <3a7d7dbf@news.povray.org> , "Daniel Jungmann" <DSJ### [at] gmxnet> 
wrote:

> Which compiler does these things?

Every compiler should do basic optimisations.  Sometimes they don't tell you
want exactly they do.  What you are seeking can be called any of the terms
below, but other names or methods are possible [1].  The list of
optimizations below is not complete and not precise, but should give you an
idea what to look for in your compiler documentation.

> 1) loops

- loop-invariant code motion (takes out code that doesn't change in a loop)
- loop unrolling (on expense of code size, reduces branches for loops)

> 2) float divisions (replace mutiple divisions with one division and
> multiplications)

- Replacing division by constants with multiplaction

See important issues regarding applying other optimisations to
floating-point numbers in [1, section 12.3.2]!

> 3) if then / ? :

Taken care of by the intermediate code representation in any compiler.

> 4) integer multiplications and divisions (replace them with shift or
> additions)

- Strength reduction (this is exactly what you are suggesting)
- Algebraic simplification and reassociation (associativity, commutativity,
distributivity)
- Value numbering (eliminating one of two equivalent calculations), note
that some compilers can do this on a global level!

> 5) multiple integer calculations

- constant folding (calculates constants before they are compiled)

In addition, compilers can do things that are hard to do for a human in
assembly language, such as register coloring (reusing registers to hold
different variables that are both in scope but not used at the same time),
instruction scheduling accoring to the processor manufacture's instruction
execution time tables (this is one of the things why you can pick the
processor to optimise forī, next to different instruction sets, of course),
instruction scheduling (avoid stalls in processor pipeline).

Now, all these listed optimisations and many more are probably in every
compiler you have, but to be specific, all of them are supported by the
Intel compiler according to [1].  And I am sure most of them (if working
seems to be another matter, as I am told) are in Visual C++, too.


  Thorsten


____________________________________________________
Thorsten Froehlich
e-mail: mac### [at] povrayorg

I am a member of the POV-Ray Team.
Visit POV-Ray on the web: http://mac.povray.org



[1] Muchnick, Steven S., Advanced Compiler Design and Implementation
    ISBN 1-55860-320-4

____________________________________________________
Thorsten Froehlich, Duisburg, Germany
e-mail: tho### [at] trfde

Visit POV-Ray on the web: http://mac.povray.org


Post a reply to this message

From: Thorsten Froehlich
Subject: Re: PovRay faster
Date: 5 Feb 2001 19:56:29
Message: <3a7f4bbd$1@news.povray.org>
In article <3a7e0a19@news.povray.org> , "Thorsten Froehlich" 
<tho### [at] trfde> wrote:

> Now, all these listed optimisations and many more are probably in every
> compiler you have, but to be specific, all of them are supported by the
> Intel compiler according to [1].  And I am sure most of them (if working
> seems to be another matter, as I am told) are in Visual C++, too.

To add to this, Intel actually has a readable and up-to-date feature list on
their site at <http://www.intel.com/software/products/compilers/c50/opt.pdf>


     Thorsten


____________________________________________________
Thorsten Froehlich, Duisburg, Germany
e-mail: tho### [at] trfde

Visit POV-Ray on the web: http://mac.povray.org


Post a reply to this message

From: Warp
Subject: Re: PovRay faster
Date: 6 Feb 2001 08:17:57
Message: <3a7ff985@news.povray.org>
Thorsten Froehlich <tho### [at] trfde> wrote:
: It is a complete waste of time to optimize these by hand.  Every compiler
: can do a much better job on these and the code stays readable on your site!

  This is true, of course.

  It's important to know what kind of code does the compiler generate from
certain operations so that you can concentrate on the important parts when
optimizing and forget the obsolete parts (because the compiler handles them).
  For example, trying to optimize an operation where an integer variable
is multiplied/divided with an integer constant by replacing it with shifts
and or operations is a complete waste of time since the compiler will do
it itself anyways.
  In some platforms it can be even slower to make all those shifts and
ors instead of letting the compiler optimize the code. That is, for example
the compiler-generated code for this:

  a = b*1040;

may be in some platforms faster than the compiler-generated code for this:

  a = (b<<10)|(b<<4);

because the compiler might be able to use some platform-dependant optimizations
for the multiplication which it can't do to the shifts (for example in some
platforms integer multiplication takes just 1 clock while the shifts and the
or take 3 clocks).

  By the way, some times the compiler can go too far in this.
  This is the case with gcc. When generating the assembler code, it will
always convert the multiplication of an integer variable and an integer
constant to shifts and ors/additions/substractions, no matter what is the
constant value.
  For example something like a*123456789 generates about 10-20 assembler
operations.
  One could think that after a certain amount of operations a plain integer
multplication would be faster (specially with current processors).
  Go figure.

-- 
char*i="b[7FK@`3NB6>B:b3O6>:B:b3O6><`3:;8:6f733:>::b?7B>:>^B>C73;S1";
main(_,c,m){for(m=32;c=*i++-49;c&m?puts(""):m)for(_=(
c/4)&7;putchar(m),_--?m:(_=(1<<(c&3))-1,(m^=3)&3););}    /*- Warp -*/


Post a reply to this message

From: Peter J  Holzer
Subject: Re: PovRay faster
Date: 6 Feb 2001 18:02:40
Message: <slrn980sf6.fb6.hjp-usenet@teal.h.hjp.at>
On 2001-02-06 13:17, Warp <war### [at] tagpovrayorg> wrote:
>  By the way, some times the compiler can go too far in this.
>  This is the case with gcc. When generating the assembler code, it will
>always convert the multiplication of an integer variable and an integer
>constant to shifts and ors/additions/substractions, no matter what is the
>constant value.
>  For example something like a*123456789 generates about 10-20 assembler
>operations.

Not always. E.g., egcs-2.91.66 for Intel will compile

    int foo (int a) {
	return a*123456789;
    }

    int bar (int a) {
	return a*4;
    }


into:


	    .file	"foo.c"
	    .version	"01.01"
    gcc2_compiled.:
    .text
	    .align 4
    .globl foo
	    .type	 foo,@function
    foo:
	    pushl %ebp
	    movl %esp,%ebp
	    imull $123456789,8(%ebp),%eax
	    leave
	    ret
    .Lfe1:
	    .size	 foo,.Lfe1-foo
	    .align 4
    .globl bar
	    .type	 bar,@function
    bar:
	    pushl %ebp
	    movl %esp,%ebp
	    movl 8(%ebp),%edx
	    leal 0(,%edx,4),%eax
	    leave
	    ret
    .Lfe2:
	    .size	 bar,.Lfe2-bar
	    .ident	"GCC: (GNU) egcs-2.91.66 19990314/Linux (egcs-1.1.2 release)"

As you can see, the first multiplication is translated to an imul
instruction, while the second one is translated to a leal (which is
essentially a shift-and-add operation.

As far as I can remember, gcc has behaved like this since at least the
1.3x releases.

	hp

-- 
   _  | Peter J. Holzer    | All Linux applications run on Solaris,
|_|_) | Sysadmin WSR       | which is our implementation of Linux.
| |   | hjp### [at] wsracat      | 
__/   | http://www.hjp.at/ |	-- Scott McNealy, Dec. 2000


Post a reply to this message

From: Warp
Subject: Re: PovRay faster
Date: 7 Feb 2001 08:19:17
Message: <3a814b55@news.povray.org>
What optimization options did you use?

-- 
char*i="b[7FK@`3NB6>B:b3O6>:B:b3O6><`3:;8:6f733:>::b?7B>:>^B>C73;S1";
main(_,c,m){for(m=32;c=*i++-49;c&m?puts(""):m)for(_=(
c/4)&7;putchar(m),_--?m:(_=(1<<(c&3))-1,(m^=3)&3););}    /*- Warp -*/


Post a reply to this message

From: Peter J  Holzer
Subject: Re: PovRay faster
Date: 7 Feb 2001 20:02:23
Message: <slrn983nat.pho.hjp-usenet@teal.h.hjp.at>
On 2001-02-07 13:19, Warp <war### [at] tagpovrayorg> wrote:
>  What optimization options did you use?

-O3, but it produces the same code with -O and -O2. 

Even without any optimization the same imull and leal instructions are
generated, there are only a few superfluous moves and jmps.

	hp

-- 
   _  | Peter J. Holzer    | All Linux applications run on Solaris,
|_|_) | Sysadmin WSR       | which is our implementation of Linux.
| |   | hjp### [at] wsracat      | 
__/   | http://www.hjp.at/ |	-- Scott McNealy, Dec. 2000


Post a reply to this message

From: Warp
Subject: Re: PovRay faster
Date: 8 Feb 2001 07:39:39
Message: <3a82938a@news.povray.org>
Peter J. Holzer <hjp### [at] hjpat> wrote:
: Even without any optimization the same imull and leal instructions are
: generated, there are only a few superfluous moves and jmps.

  It may be that you have a newer version of gcc than I do (the gcc version
here is actually quite old).

-- 
char*i="b[7FK@`3NB6>B:b3O6>:B:b3O6><`3:;8:6f733:>::b?7B>:>^B>C73;S1";
main(_,c,m){for(m=32;c=*i++-49;c&m?puts(""):m)for(_=(
c/4)&7;putchar(m),_--?m:(_=(1<<(c&3))-1,(m^=3)&3););}    /*- Warp -*/


Post a reply to this message

From: Thorsten Froehlich
Subject: Re: PovRay faster
Date: 8 Feb 2001 09:46:40
Message: <3a82b150$1@news.povray.org>
In article <3a82938a@news.povray.org> , Warp <war### [at] tagpovrayorg>  wrote:

> Peter J. Holzer <hjp### [at] hjpat> wrote:
> : Even without any optimization the same imull and leal instructions are
> : generated, there are only a few superfluous moves and jmps.
>
>   It may be that you have a newer version of gcc than I do (the gcc version
> here is actually quite old).

gcc 2.95.2 is current (and has been for years).  There seems to be going to
be an update (finally) - version 3.0 within a few weeks according to the
official gcc website.


      Thorsten


____________________________________________________
Thorsten Froehlich, Duisburg, Germany
e-mail: tho### [at] trfde

Visit POV-Ray on the web: http://mac.povray.org


Post a reply to this message

<<< Previous 50 Messages Goto Initial 50 Messages

Copyright 2003-2023 Persistence of Vision Raytracer Pty. Ltd.