gaminggamedevprogrammingprojectsmatrixdevGameArtleveldesignTop Subs
8

The big win here is having floating point math working in a way that's stable. Huge update for Responsive OS.

Before I could sometimes get away with using floating point math inside a single function because the compiler's optimizer was, sometimes, rewriting it as fixed precision integer math. And I was finding myself temped to write my own library around that kind of math.

But I had a goal for a 3D graphic idea that got me to finally get my FPU (floating point unit) running properly as a preliminary task. What I'm about to show you isn't that complicated. Just know that when things are breaking you have to test every way of modifying it. So I'm only showing you the juicy and relatively clean results. This required me to learn some assembly, of which I know very little.

The first thing to do was to get the low level environment into the clean state we need for floating point math. Our stack needs to be exactly 16 bit aligned any place a floating point number might be in memory. We start off on the right foot by getting the top of the stack aligned.

start:
    ; More code above I'm about to show
    mov esp, stack_top
    and esp, -16 ; Line added to move the stack to the closest 16bit alignment
    push ebx 
    call kernel_main
;...Skipping a lot of lines
section .bss
    align 16
    stack_bottom:
    ;resb 16384 ; 16 KB stack
    resb 1023*1024 ; 16 KB stack
    stack_top:

We also need another compiler flag to help keep this stack growing in a 16bit aligned way. -mstackrealign

Next we need to get the FPU setup.

    ; enable fpu
    mov eax, cr0        ;Read cr0 register
    and eax, 0xFFFEFFFF ;Set low the sixteenth bit from the left 
    mov cr0, eax        ;Apply it to cr0
    ; enable sse
    mov eax, cr4        ;Read cr4 register
    or eax, 0x00000400  ;Set high the seventh bit from the right
    mov cr4, eax        ;Apply it to cr4

We haven't tested SSE instructions yet so I don't know if we are there yet, but we have enabled the flags on the CPU for both the FPU and SSE.

Then initialize the device.

    finit

We are now setup. We're about ready to do our 3D graphics. For it, while we are in assembly we just need to write one function, then we can go to C.

global sqrt

sqrt:
    fld dword [esp + 4] ;Push the first argument into the FPU stack
    fsqrt               ;Because the C ABI (application binary interface) specifies that a function returning float does so at ST(0) (top of the FPU stack), we don't need to pop the value off.
    ret

Now for C.

extern float sqrt(float); //Specify this is implemented elsewhere (boot.s)

Next I created a useful type for specifying points in 3D.

typedef struct V3 {
 float x,y,z;
} V3;

Now a few more helper math functions.

float dist(V3 p1,V3 p2) {
 float dx=p1.x-p2.x;
 float dy=p1.y-p2.y;
 float dz=p1.z-p2.z;
 return sqrt(dx*dx+dy*dy+dz*dz);
}

float mag(V3 p) {
 return sqrt(p.x*p.x+p.y*p.y+p.z*p.z);
}

float mag2(V3 p) {
 return p.x*p.x+p.y*p.y+p.z*p.z;
}

float fast_inv_sqrt(float number) {
 long i; 
 float x2, y; 
 const float threehalf = 0.5f * 3.0f;

 x2 = number * 0.5f;
 y  = number;

 // The "Magic" bit-cast and constant
 // This treats the float as an integer to perform a rough
 // approximation of the log2 of the number.
 i = *(long *) &y;                      // evil floating point bit level hacking
 i = 0x5f3759df - (i >> 1);             // what the fuck?
 y = *(float *) &i;

 // One iteration of Newton's method to refine the result
 y = y * (1.5f - (x2 * y * y));

 return y; 
}

Thank you John Romero that that juicy bit. Though I'm told he got it from someone else. The demo I ran for myself kept it simple, but we'll do this the fun way if we're going to write about it.

Now to the unique idea I wanted to try. I came up with a lazy way to calculate a mapping between a 3D point and a point on a screen that should produce a fish-eye effect. It's not 100% correct but I wanted to see it. My idea was to combine this simplification with another one. What if a game only let you view the world axis aligned or at most 45 degrees between. Combining these ideas should make the math very simple.

The first idea since I haven't described it yet is to calculate x on the screen from x in space divided by distance, plus some extra screen math. Normally this is the standard formula except your distance is meant to be the distance to the orthogonal plane to the camera view that intersects the point. Ironically, that's actually really easy to calculate when we are axis aligned. That's just the point's z value when we are looking in the z direction. But I want to see what my method looks like.

One consequence of my method is that it will always have a 180 degree view. We can't specify a different frustum from what I can tell. Imagine if x is positive, z is zero (perfectly in our periphery), and y is zero. Then distance will equal x. So our simplified screen point is just x/x, aka 1, aka the edge of the screen.

One downside of the method is that if a point has both a x value and y value (height), and near zero z-value, the point edge it will map to will be the edge of a circle connecting the top, bottom, and sides of the screen. So everything will fit in a circle. Blank corners.

Ok. Onto code.

#define overload __attribute((overloadable))

//Viewing locked in a z direction.
//You'll see why I marked this as overloadable soon
V3 overload screen_point_z(V3 p,u32 width,u32 height) {
 float d = mag(p);
 return (V3){
  .x=(1+p.x/d)*width/2,
  .y=(1-p.y/d)*height/2,
  .z=p.z>0?d:-d
 };
}

V3 overload screen_point_x(V3 p,u32 width,u32 height) {
 float d = mag(p);
 return (V3){
  .x=(1+p.z/d)*width/2,
  .y=(1-p.y/d)*height/2,
  .z=p.x>0?d:-d
 };
}

Woops. I didn't use the fast inverse squared hack. Oh well. That is where you would plug it in.

A critical thing this function is doing is instead of returning a 2D point, it makes use of the V3 struct to encode the distance of the point. And more importantly it's encoding if that's distance in front of the camera or behind, with negative distance.

Now let's get to our first sample of using it.

void kernel_main(multiboot_info *info) {
 setfbinfo(info); //Give framebuffer data to the graphics library
 makemainheap();  //One of the few programs that may not need the heap, but better safe than sorry
 V3 p1={-10.0 , -3.0 , 10.0}; //It's a house
 V3 p2={-10.0 ,  7.0 , 10.0};
 V3 p3={  0.0 , 11.0 , 10.0};
 V3 p4={ 10.0 ,  7.0 , 10.0};
 V3 p5={ 10.0 , -3.0 , 10.0};
 V3 s1=screen_point_z(p1,ref,1024,768);
 V3 s2=screen_point_z(p2,ref,1024,768);
 V3 s3=screen_point_z(p3,ref,1024,768);
 V3 s4=screen_point_z(p4,ref,1024,768);
 V3 s5=screen_point_z(p5,ref,1024,768);
 setbackground(0xFF00FF);
 setforeground(0x00FF00);
 drawbackground();
 drawseg(s1,s2);
 drawseg(s2,s3);
 drawseg(s3,s4);
 drawseg(s4,s5);
 drawseg(s5,s1);
}

This drawseg function already existed before I started this demo but we've overloaded it.
This overloaded version needs to make use of the z value in our translated vector. Let me show you it's features. First the standard method.

void overload drawseg(u32 x1, u32 y1, u32 x2, u32 y2) {
 multiboot_info *info = default_info;
 u32 *fb = (u32 *)(info->framebuffer_addr);
 i32 width = info->framebuffer_width;
 i32 height = info->framebuffer_height;
 //if(x1>=width||y1>=height||x2>=width||y2>=height) return;
 // Use signed ints for all math to allow negative values during calculation
 i32 sx1 = (i32)x1;
 i32 sy1 = (i32)y1;
 i32 sx2 = (i32)x2;
 i32 sy2 = (i32)y2;
 if(sy2 < sy1) {
  i32 temp = sx1; sx1 = sx2; sx2 = temp;
  temp = sy1; sy1 = sy2; sy2 = temp;
 }
 i32 dx = abdiff(sx1, sx2);
 i32 dy = sy2 - sy1;
 i32 steps = max(dx, dy);
 i32 idx = sx2 - sx1;
 if(steps == 0) return;
 for(i32 i = 0; i < steps; ++i) {
  // Calculate as signed integers
  i32 y = sy1 + (dy * i / steps);
  i32 x = sx1 + (idx * i / steps);
  // NOW the bounds check actually works because x and y can be negative
  //if(x < 0 || x >= width || y < 0 || y >= height) continue;
  fb[y * width + x] = foregroundColor;
 }
}

That's some good integer math from the pre-float era for our OS. Interger math is considered more correct for this anyway.

Now the overloaded method that takes the translated 3D vectors.

void overload drawseg(V3 p1,V3 p2) {
 //Both points are behind us.  Do nothing.
 if(p1.z<0 && p2.z<0) return;
 //Both points are in front of us.  Cast it to the standard function.
 if(p1.z>=0 && p1.z>=0) return drawseg((u32)p1.x,(u32)p1.y,(u32)p2.x,(u32)p2.y); //
 //One point is in front and another behind.  Use recursion to make p1 the one in front.
 if(p1.z<0 && p1.z>=0) return drawseg(p2,p1);
 // p1 is in front.  p2 is behind.
 //The most fun case.  We need to interpolate where z=0;
 float ratio=p1.z/(p1.z-p2.z);
 drawseg(p1,(V3){
  .x=p2.x*ratio+p1.x*(1-ratio),
  .y=p2.y*ratio+p1.y*(1-ratio),
  .z=0.0,
 });
}

We now have a house drawn on the screen. Or one face of a house. But in real 3D shouldn't things move? Let's make that happen.

To make our main function simple we need a variant of the screen_point_* functions.

V3 overload screen_point_z(V3 p,V3 ref,u32 width,u32 height);
V3 overload screen_point_x(V3 p,V3 ref,u32 width,u32 height);
V3 overload screen_point_z(V3 p,u32 width,u32 height);
V3 overload screen_point_x(V3 p,u32 width,u32 height);

V3 overload screen_point_z(V3 p,V3 ref,u32 width,u32 height) {
 return screen_point_z((V3){
  .x=p.x-ref.x,
  .y=p.y-ref.y,
  .z=p.z-ref.z,
 },width,height);
}

V3 overload screen_point_x(V3 p,V3 ref,u32 width,u32 height) {
 return screen_point_x((V3){
  .x=p.x-ref.x,
  .y=p.y-ref.y,
  .z=p.z-ref.z,
 },width,height);
}

Don't forget that when you do function overloading in C that you should have function signatures, even if you don't have a .h file for a demo. Boy isn't function overloading nice to have in modern C.

Now the new main function

void kernel_main(multiboot_info *info) {
 setfbinfo(info);
 makemainheap();
 V3 p1={-10.0 , -3.0 , 10.0};
 V3 p2={-10.0 ,  7.0 , 10.0};
 V3 p3={  0.0 , 11.0 , 10.0};
 V3 p4={ 10.0 ,  7.0 , 10.0};
 V3 p5={ 10.0 , -3.0 , 10.0};
 V3 s1;
 V3 s2;
 V3 s3;
 V3 s4;
 V3 s5;
 V3 ref=(V3){0.0,0.0,0.0};
 setforegroundcolor(0x00FFFF);
 setbackgroundcolor(0xFF00FF);
 char lastpress;
 for(u32 i=0;1;++i) {
  char press=check_keyboard();
  if(press) lastpress=press;
  if(lastpress=='w') ref.z+=0.01;
  if(lastpress=='a') ref.x-=0.01;
  if(lastpress=='s') ref.z-=0.01;
  if(lastpress=='d') ref.x+=0.01;
  s1=screen_point_z(p1,ref,1024,768);
  s2=screen_point_z(p2,ref,1024,768);
  s3=screen_point_z(p3,ref,1024,768);
  s4=screen_point_z(p4,ref,1024,768);
  s5=screen_point_z(p5,ref,1024,768);
  if(i%3==0) drawbackground(info);  //Motion blur effect not clearing the screen every frame
  drawseg(s1,s2);
  drawseg(s2,s3);
  drawseg(s3,s4);
  drawseg(s4,s5);
  drawseg(s5,s1);
  sleep_ms(10);
 }
}

I'll try to get a pic up later.

And there you go. A walkable minimal 3D scene running outside a standard operating system.

Comment preview