 |
 |
|
 |
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
I have compiled a version of pov-ray with gcc 3.1, optimized for pentium 3
cpus. I have also applied a small patch which potentially increases
performance. If you want to try it out, you can get it from
http://www.povworld.org/povray/povray.bz2
I have not had time to do too much testing but it should be quite a bit
faster than the official binary. Please someone test against windows binary.
- Micha
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Micha Riser wrote:
> I have compiled a version of pov-ray with gcc 3.1, optimized for pentium 3
> cpus. I have also applied a small patch which potentially increases
> performance. If you want to try it out, you can get it from
>
> http://www.povworld.org/povray/povray.bz2
>
> I have not had time to do too much testing but it should be quite a bit
> faster than the official binary. Please someone test against windows binary.
>
> - Micha
>
Bless you Micha! This worked like a charm. Comparison rendering times of
the balcony.pov sample scene are as follows:
Windows prebuilt binary = 505s
Linux prebuilt binary = 785s
Micha's p3 optimized binary = 444s
I'll do some more thorough testing with other scenes such as the
benchmark.pov one later today.
I don't have gcc 3.1 yet. Would the small patch you applied help
performance in a gcc 2.96 compiled version?
-Roz
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Roz wrote:
>
> Windows prebuilt binary = 505s
> Linux prebuilt binary = 785s
> Micha's p3 optimized binary = 444s
Great! What CPU do you have btw?
>
> I don't have gcc 3.1 yet. Would the small patch you applied help
> performance in a gcc 2.96 compiled version?
>
Yeah, it is an algorithmic change and therefore compile independant. I'll
post details later. It's is national holiday here in Switzerland today...
- Micha
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Micha Riser wrote:
> I have compiled a version of pov-ray with gcc 3.1, optimized for pentium 3
> cpus. I have also applied a small patch which potentially increases
> performance. If you want to try it out, you can get it from
Wow! That's great, Micha!
--
light_source{0#macro L(K,H,W)sphere{H.5}sphere{K.5}sphere{W.5}cylinder{
H,K.5}cylinder{H,W.5}#end 3}union{L(0v*-2<2,-2>)L(y*-3z-v*5z*3-y)L(-y*3
0u*3)L(y*-3v*-5-z,z*-3-y)rotate-v*clock pigment{rgb.5}translate<0,2,9>}
// +KFF200 +KF720 +W120 +H90 -F -A -GA -P
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Thanks for it
File : orient.pov
parameters : +w800 +h600 -f +a +dgt +v
Linux Mandrake 8.2 Kernel 2.4.18
Windows 2000 Pro
It is also optimized for my Duron !
AMD Duron 700 Mhz 512 Mo Memory
Official binary : 262 seconds
Optimized version : 121 seconds
Windows version : 121 seconds
But the optimised version does not work with soft.pov
I get an error
Fabien HENON
Micha Riser a écrit :
> I have compiled a version of pov-ray with gcc 3.1, optimized for pentium 3
> cpus. I have also applied a small patch which potentially increases
> performance. If you want to try it out, you can get it from
>
> http://www.povworld.org/povray/povray.bz2
>
> I have not had time to do too much testing but it should be quite a bit
> faster than the official binary. Please someone test against windows binary.
>
> - Micha
>
> --
> http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
I get an error : 'Instruction illégale'
Fabien Hénon wrote:
> Thanks for it
> File : orient.pov
> parameters : +w800 +h600 -f +a +dgt +v
> Linux Mandrake 8.2 Kernel 2.4.18
> Windows 2000 Pro
> It is also optimized for my Duron !
> AMD Duron 700 Mhz 512 Mo Memory
>
>
> Official binary : 262 seconds
> Optimized version : 121 seconds
> Windows version : 121 seconds
>
>
> But the optimised version does not work with soft.pov
> I get an error
>
>
> Fabien HENON
>
> Micha Riser a écrit :
>
>
>>I have compiled a version of pov-ray with gcc 3.1, optimized for pentium 3
>>cpus. I have also applied a small patch which potentially increases
>>performance. If you want to try it out, you can get it from
>>
>>http://www.povworld.org/povray/povray.bz2
>>
>>I have not had time to do too much testing but it should be quite a bit
>>faster than the official binary. Please someone test against windows binary.
>>
>>- Micha
>>
>>--
>>http://objects.povworld.org - the POV-Ray Objects Collection
>
>
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
I'll look into adding easier support for various optimizations to the
build, as well as multiple binaries with different optimizations.
Building from source will probably remain the best way to get better
performance. For something like POV-Ray (CPU-intensive, fairly low build
time), that's probably reasonable.
-Mark Gordon
On Thu, 01 Aug 2002 12:46:28 -0400, Micha Riser wrote:
> I have compiled a version of pov-ray with gcc 3.1, optimized for pentium
> 3 cpus. I have also applied a small patch which potentially increases
> performance. If you want to try it out, you can get it from
>
> http://www.povworld.org/povray/povray.bz2
>
> I have not had time to do too much testing but it should be quite a bit
> faster than the official binary. Please someone test against windows
> binary.
>
> - Micha
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Fabien Hénon wrote:
> Thanks for it
> File : orient.pov
> parameters : +w800 +h600 -f +a +dgt +v
> Linux Mandrake 8.2 Kernel 2.4.18
> Windows 2000 Pro
> It is also optimized for my Duron !
> AMD Duron 700 Mhz 512 Mo Memory
>
Nice to hear.
>
> But the optimised version does not work with soft.pov
> I get an error
>
Works fine here with soft.pov. This is probably because amd duron does not
support 'SSE' instruction what pentium 3 does. (Athlon XP does it). I can
try to make a version for amd duron as well - though if I don't have such a
processor I cannot test it..
- Micha
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Micha Riser wrote:
> Roz wrote:
>
>>Windows prebuilt binary = 505s
>>Linux prebuilt binary = 785s
>>Micha's p3 optimized binary = 444s
>
>
> Great! What CPU do you have btw?
>
It's an AMD Athlon XP 1900+ with 512meg DDR RAM.
I ran benchmark.pov through its paces and these are the rendering times:
Windows prebuilt binary = 25m 48s (1548s)
Linux prebuilt binary = 55m 40s (3340s)
Micha's p3 optimized binary = 25m 14s (1514s)
Not bad :)
>>I don't have gcc 3.1 yet. Would the small patch you applied help
>>performance in a gcc 2.96 compiled version?
>>
>
>
> Yeah, it is an algorithmic change and therefore compile independant. I'll
> post details later. It's is national holiday here in Switzerland today...
>
> - Micha
>
Happy Holiday and thanks for posting the p3 optimized binary :)
-Roz
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On Thu, 01 Aug 2002 12:46:28 -0400, Micha Riser wrote:
> I have compiled a version of pov-ray with gcc 3.1
Hmm... I found a gcc 3.2 in rawhide. Maybe I should give that a whirl...
-Mark Gordon
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Mark Gordon wrote:
> On Thu, 01 Aug 2002 12:46:28 -0400, Micha Riser wrote:
>
>
>> I have compiled a version of pov-ray with gcc 3.1
>
> Hmm... I found a gcc 3.2 in rawhide. Maybe I should give that a whirl...
>
gcc 3.2 is not released yet.. get 3.1 form gcc.gnu.org.
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Well I am willing to try your Duron-compiled version
Fabien
Micha Riser a écrit :
> Fabien Hénon wrote:
>
> > Thanks for it
> > File : orient.pov
> > parameters : +w800 +h600 -f +a +dgt +v
> > Linux Mandrake 8.2 Kernel 2.4.18
> > Windows 2000 Pro
> > It is also optimized for my Duron !
> > AMD Duron 700 Mhz 512 Mo Memory
> >
>
> Nice to hear.
>
> >
> > But the optimised version does not work with soft.pov
> > I get an error
> >
>
> Works fine here with soft.pov. This is probably because amd duron does not
> support 'SSE' instruction what pentium 3 does. (Athlon XP does it). I can
> try to make a version for amd duron as well - though if I don't have such a
> processor I cannot test it..
>
> - Micha
>
> --
> http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Fabien Hénon wrote:
> Well I am willing to try your Duron-compiled version
>
Can you post or email me a print of 'cat /proc/cpuinfo' on your system to
give me the exact details of your cpu?
- Micha
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
There it is as an attachment
Thanks
Fabien HENON
Micha Riser wrote:
> Fabien Hénon wrote:
>
>
>>Well I am willing to try your Duron-compiled version
>>
>
>
> Can you post or email me a print of 'cat /proc/cpuinfo' on your system to
> give me the exact details of your cpu?
>
> - Micha
>
Post a reply to this message
Attachments:
Download 'us-ascii' (1 KB)
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
I have done further trials to speed povray up. You can expect a up to 10%
faster gcc compile from me in the next few days. While GCC 3.1 uses MMX
registers it does not actually use SIMD instructions :( One would have to
add these by hand but this is quite tiring.
Currently I am testing Intel's compiler... and had to find that the POV-Ray
coders did a great deal in making it impossible to auto-vecotrise the
colour-operations for icc :/
- Micha
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
If people are continue improving algorithmic changes that are compiler and
platform independent, wouldn't it be a good thing to keep track of those
changes and work on a 3.5.1 or a 3.6 version?
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On Thu, 01 Aug 2002 21:19:53 +0200, Mark Gordon wrote:
>
> On Thu, 01 Aug 2002 12:46:28 -0400, Micha Riser wrote:
>
>
>> I have compiled a version of pov-ray with gcc 3.1
>
> Hmm... I found a gcc 3.2 in rawhide. Maybe I should give that a
whirl...
gcc 3.2 is gcc 3.1.1 with a new c++ ABI (c++ v3 multivendor) which is what
we are aiming for with our 1.4 Gentoo release :)
gcc 3.1.1 is released, and has finally made -march=pentium4 stable (it
was bad juju in 3.1 ) and should have some further speedups. avaiable from
http://gcc.gnu.org/
for those asking about optimization flags, gcc/make should automatically
rebuild whats needed when you just do :
export CFLAGS="optimziation";export CXXFLAGS="${CFLAGS}"; ./configure
--with-foo --enable-bar; make clean ;make all
//Spider
--
begin .signature
This is a .signature virus! Please copy me into your .signature!
See Microsoft KB Article Q265230 for more information.
end
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Micha Riser <mri### [at] gmx net> wrote:
> Currently I am testing Intel's compiler... and had to find that the POV-Ray
> coders did a great deal in making it impossible to auto-vecotrise the
> colour-operations for icc :/
This reminds me of something else: Seemingly patch makers have not followed
the strict programming guidelines in the POV-Ray source code, which has
caused a big problem: The code is extremely hard to optimize for RISC
processors, which means that POV-Ray is quite slow in them, even though
it could be a lot faster if the patches were coded in the right way.
--
#macro M(A,N,D,L)plane{-z,-9pigment{mandel L*9translate N color_map{[0rgb x]
[1rgb 9]}scale<D,D*3D>*1e3}rotate y*A*8}#end M(-3<1.206434.28623>70,7)M(
-1<.7438.1795>1,20)M(1<.77595.13699>30,20)M(3<.75923.07145>80,99)// - Warp -
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
In article <3d4b341f@news.povray.org>, Warp <war### [at] tag povray org>
wrote:
> This reminds me of something else: Seemingly patch makers have not followed
> the strict programming guidelines in the POV-Ray source code, which has
> caused a big problem: The code is extremely hard to optimize for RISC
> processors, which means that POV-Ray is quite slow in them, even though
> it could be a lot faster if the patches were coded in the right way.
What "strict programming guidelines" are you talking about?!?
The Source is very messy and inconsistent, badly commented, with
different coding styles in different areas and often doing things the
hard way when an existing function does the same thing. It isn't just
the patches that have this problem, some of the worst code is very old:
The leopard pattern (and several others) could be a 1-liner[1], but it
uses 4 temporary variables. The onion pattern has a (unnecessary)
variable named "noise" and the following comment:
/* The variable noise is not used as noise in this function */
The code is littered with unused variables and produces huge amounts of
warnings.
And there is no documentation other than the source code itself,
definitely no strict guidelines. And I don't think the code has ever
been good on RISC machines...it has been like that from the beginning.
[1] Specifically:
return Sqr((sin(EPoint[X]) + sin(EPoint[Y] + sin(EPoint[Z])))/3.0);
--
Christopher James Huff <chr### [at] mac com>
POV-Ray TAG e-mail: chr### [at] tag povray org
TAG web site: http://tag.povray.org/
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
> return Sqr((sin(EPoint[X]) + sin(EPoint[Y] + sin(EPoint[Z])))/3.0);
You probably mean Sqr((sin(EPoint[X])+sin(EPoint[Y])+sin(EPoint[Z]))/3);
Temporary variables are do not affect the performance mostly though. The
has to use several registers anyways. But to show the vector nature of this
calculation it should be written as:
DBL value=0;
for(int i=0; i<3; i++) value+=sin(EPoint[i]);
return Sqr(value/3.0);
Of course this assumes that X,Y,Z are 0-2 index.
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On Fri, 02 Aug 2002 20:58:33 -0400, Spider wrote:
> export CFLAGS="optimziation"
Are you quite sure on your spelling, there? ;-)
-Mark Gordon
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Apache wrote:
> If people are continue improving algorithmic changes that are compiler and
> platform independent, wouldn't it be a good thing to keep track of those
> changes and work on a 3.5.1 or a 3.6 version?
>
>
Yea and it seems like the p.programming newsgroup helps with that.
You've probably already noticed it but if not, Micha has posted the
patch he was talking about to that newsgroup.
-Roz
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
In article <3d4b8aec@news.povray.org>, Micha Riser <mri### [at] gmx net>
wrote:
> You probably mean Sqr((sin(EPoint[X])+sin(EPoint[Y])+sin(EPoint[Z]))/3);
Yeah, I screwed up the parentheses and the ".0" isn't necessary.
> Temporary variables are do not affect the performance mostly though. The
> has to use several registers anyways.
I know, and a good compiler would probably optimize it to the same
machine code, but there is no reason to use them, not even readability.
(Do modern compilers even pay any attention to "register"?)
My point was it doesn't conform to any kind of "strict guidelines" other
than maximum compatibility (it was apparently for compiling with an old
386 compiler). The "noise" variable I just don't understand...I guess it
helped some compiler with optimization or something.
> But to show the vector nature of this calculation it should be
> written as:
> DBL value=0;
> for(int i=0; i<3; i++) value+=sin(EPoint[i]);
> return Sqr(value/3.0);
>
> Of course this assumes that X,Y,Z are 0-2 index.
You mean for helping the compiler detect something that can be
vectorized and doing it automatically? It won't pick out the possibility
in the one-line version?
Would it get this version?
DBL value = 0;
value += sin(EPoint[0]);
value += sin(EPoint[1]);
value += sin(EPoint[2]);
return Sqr(value/3.0);
Or this (assuming a struct with x, y, and z components):
value += sin(EPoint.x);
value += sin(EPoint.y);
value += sin(EPoint.z);
I'm not surprised the for loop wasn't used in the existing
version...SIMD stuff wasn't even a factor, so of course the code wasn't
designed for it, and compilers probably weren't good enough at
optimizing to get rid of the for() loop.
--
Christopher James Huff <chr### [at] mac com>
POV-Ray TAG e-mail: chr### [at] tag povray org
TAG web site: http://tag.povray.org/
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Christopher James Huff wrote:
>
> You mean for helping the compiler detect something that can be
> vectorized and doing it automatically? It won't pick out the possibility
> in the one-line version?
> Would it get this version?
> DBL value = 0;
> value += sin(EPoint[0]);
> value += sin(EPoint[1]);
> value += sin(EPoint[2]);
> return Sqr(value/3.0);
No, unfortunately the Intel compiler (the only one that I have which does
vectorisation) does not recoginze this. You explicitly have to use a loop.
I have done rewriting vector.h in such a way. But I have no Pentium4 to
test it :(
> I'm not surprised the for loop wasn't used in the existing
> version...SIMD stuff wasn't even a factor, so of course the code wasn't
> designed for it, and compilers probably weren't good enough at
> optimizing to get rid of the for() loop.
Todays g++ 3.1 does a good job in loop-unrolling. But with my modified
loop-using 'vector.h' it does still produce a slightly slower code (OK,
maybe I have also made some mistakes in the converting..)
- Micha
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Micha Riser wrote:
>
> loop-using 'vector.h' it does still produce a slightly slower code (OK,
I have to correct this. The increasing number of background applications
had influenced the testing. g++ 3.1 is equally fast when using a 'looped'
vector.h.
- Micha
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
I only remembered an article written by Thorsten about the issue. I don't
remember the specifics.
--
#macro N(D)#if(D>99)cylinder{M()#local D=div(D,104);M().5,2pigment{rgb M()}}
N(D)#end#end#macro M()<mod(D,13)-6mod(div(D,13)8)-3,10>#end blob{
N(11117333955)N(4254934330)N(3900569407)N(7382340)N(3358)N(970)}// - Warp -
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
In article <3d4c2076@news.povray.org>, Micha Riser <mri### [at] gmx net>
wrote:
> No, unfortunately the Intel compiler (the only one that I have which does
> vectorisation) does not recoginze this. You explicitly have to use a loop.
> I have done rewriting vector.h in such a way. But I have no Pentium4 to
> test it :(
So it will only help operations on arrays then...seems like a stupid
limitation. There isn't some compiler directive that could tell it what
can be optimized?
> Todays g++ 3.1 does a good job in loop-unrolling. But with my modified
> loop-using 'vector.h' it does still produce a slightly slower code (OK,
> maybe I have also made some mistakes in the converting..)
Well, todays g++ 3.1 didn't exist 10 years ago. ;-)
With a modern compiler, I wouldn't expect it to be much slower, but how
could you get any improvement? The SIMD instructions can't handle double
precision math as far as I know...they would help colors, but not
vectors.
--
Christopher James Huff <chr### [at] mac com>
POV-Ray TAG e-mail: chr### [at] tag povray org
TAG web site: http://tag.povray.org/
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Christopher James Huff wrote:
> So it will only help operations on arrays then...seems like a stupid
> limitation. There isn't some compiler directive that could tell it what
> can be optimized?
SIMD only work on contignous blocks of memory.. so you're likely to use
arrays when the're useful.
>
> With a modern compiler, I wouldn't expect it to be much slower, but how
> could you get any improvement? The SIMD instructions can't handle double
> precision math as far as I know...they would help colors, but not
> vectors.
Yes, SSE can help for colour calculations. With it you can do operations on
4 floats simutanously. But SSE2 (which pentium4 and 64-bit AMD support)
will work on double precision! That's why I am looking for someone with a
pentium4...
- Micha
--
http://objects.povworld.org - the POV-Ray Objects Collection
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |