This commit is contained in:
2025-01-16 00:10:21 +00:00
parent 53b319b535
commit 74e2844025
2 changed files with 113 additions and 116 deletions
+8 -16
View File
@@ -77,10 +77,7 @@ const uint OPMinMaterialFloat=__LINE__-1; //F F //F
const uint OPMaxMaterialFloat=__LINE__-1; //F F //F
const uint OPSmoothMinMaterialFloat=__LINE__-1; //F F F //F
const uint OPSmoothMaxMaterialFloat=__LINE__-1; //F F F //F
const uint OPDupFloat=__LINE__-1; //F //F F
const uint OPDup2Float=__LINE__-1; //F //F F
const uint OPDup3Float=__LINE__-1; //F //F F
const uint OPDup4Float=__LINE__-1; //F //F F
const uint OPCopyFloat=__LINE__-1; //F //F
const uint OPAbsVec2=__LINE__-1; //V2 //V2
const uint OPSignVec2=__LINE__-1; //V2 //V2
const uint OPFloorVec2=__LINE__-1; //V2 //V2
@@ -100,10 +97,7 @@ const uint OPAcosVec2=__LINE__-1; //V2 //V2
const uint OPAtanVec2=__LINE__-1; //V2 //V2
const uint OPMinVec2=__LINE__-1; //V2 V2 //V2
const uint OPMaxVec2=__LINE__-1; //V2 V2 //V2
const uint OPDupVec2=__LINE__-1; //V2 //V2 V2
const uint OPDup2Vec2=__LINE__-1; //V2 //V2 V2
const uint OPDup3Vec2=__LINE__-1; //V2 //V2 V2
const uint OPDup4Vec2=__LINE__-1; //V2 //V2 V2
const uint OPCopyVec2=__LINE__-1; //V2 //V2
const uint OPAbsVec3=__LINE__-1; //V3 //V3
const uint OPSignVec3=__LINE__-1; //V3 //V3
const uint OPFloorVec3=__LINE__-1; //V3 //V3
@@ -123,10 +117,7 @@ const uint OPAcosVec3=__LINE__-1; //V3 //V3
const uint OPAtanVec3=__LINE__-1; //V3 //V3
const uint OPMinVec3=__LINE__-1; //V3 V3 //V3
const uint OPMaxVec3=__LINE__-1; //V3 V3 //V3
const uint OPDupVec3=__LINE__-1; //V3 //V3 V3
const uint OPDup2Vec3=__LINE__-1; //V3 //V3 V3
const uint OPDup3Vec3=__LINE__-1; //V3 //V3 V3
const uint OPDup4Vec3=__LINE__-1; //V3 //V3 V3
const uint OPCopyVec3=__LINE__-1; //V3 //V3
const uint OPAbsVec4=__LINE__-1; //V4 //V4
const uint OPSignVec4=__LINE__-1; //V4 //V4
const uint OPFloorVec4=__LINE__-1; //V4 //V4
@@ -146,17 +137,18 @@ const uint OPAcosVec4=__LINE__-1; //V4 //V4
const uint OPAtanVec4=__LINE__-1; //V4 //V4
const uint OPMinVec4=__LINE__-1; //V4 V4 //V4
const uint OPMaxVec4=__LINE__-1; //V4 V4 //V4
const uint OPDupVec4=__LINE__-1; //V4 //V4 V4
const uint OPDup2Vec4=__LINE__-1; //V4 //V4 V4
const uint OPDup3Vec4=__LINE__-1; //V4 //V4 V4
const uint OPDup4Vec4=__LINE__-1; //V4 //V4 V4
const uint OPCopyVec4=__LINE__-1; //V4 //V4
const uint OPPromoteFloatFloatVec2=__LINE__-1; //F F //V2
const uint OPPromoteFloatFloatFloatVec3=__LINE__-1; //F F F //V3
const uint OPPromoteFloatFloatFloatFloatVec4=__LINE__-1; //F F F F //V4
const uint OPPromoteVec2FloatVec3=__LINE__-1; //V2 F //V3
const uint OPPromoteFloatVec2Vec3=__LINE__-1; //F V2 //V3
const uint OPPromoteVec2FloatFloatVec4=__LINE__-1; //V2 F F //V4
const uint OPPromoteFloatVec2FloatVec4=__LINE__-1; //F V2 F //V4
const uint OPPromoteFloatFloatVec2Vec4=__LINE__-1; //F F V2 //V4
const uint OPPromoteVec2Vec2Vec4=__LINE__-1; //V2 V2 //V4
const uint OPPromoteVec3FloatVec4=__LINE__-1; //V3 F //V4
const uint OPPromoteFloatVec3Vec4=__LINE__-1; //F V3 //V4
const uint OPAcoshFloat=__LINE__-1; //F //F
const uint OPAcoshVec2=__LINE__-1; //V2 //V2
const uint OPAcoshVec3=__LINE__-1; //V3 //V3
+104 -99
View File
@@ -1,132 +1,137 @@
# Interpreter redesign
Ground up redesign of the interpreter
Maximum mesh shaders at once is 32*32*2, or 2048. For 65,536 vgprs, each mesh shader can only use 32.
Instead of a stack, use SSA and then limited registers (16?). This is a bit more compile friendly.
There are 16 available registers. Register 0 is always 0, and register 16 is the next item in the const tape.
This leaves 14 usable 32 bit float registers.
Vectors are constructed starting at the input variable and counting up - e.g. a the vec3 at 5 is made from vec3(5, 6, 7);
Instruction format:
32 bits
First 16 bits encode four 4 bit registers that are used as inputs. Next 4 bits are output register.
1 bit encodes if the instruction works up or down.
3 bits encode the length of the VecX type, where 0 means Vec1 and 7 means Vec8.
This leaves 8 bits to encode an opcode, for 256 possible opcodes.
Each SDF can have up to 8 materials associated with it. The relative weight of each material is tracked through the interpreter as an
array of 8 floats, and at the end they are evaluated once.
Steps:
1. 8x8x1 task shaders per mesh (64)
2. each task shader works out 2x2x8 blocks (32) where 8 is depth
3. each task shader spawns 4x4x2 (8x8x1?) mesh shaders per block.
4. each mesh shader outputs 4 verts and 2 triangles, for a total 128/64 per workgroup.
For a mesh that takes up 1024x1024 pixel on screen, each quad takes up 16x16 pixels.
# instruction set
## arithmetic
### component-wise types
- FloatFloat (returns float)
- Vec2Vec2 (returns vec2)
- Vec2Float (returns vec2)
- FloatVec2 (returns vec2)
- Vec3Vec3 (returns vec3)
- FloatVec3 (returns vec3)
- Vec3Float (returns vec3)
- Vec4Vec4 (returns vec4)
- FloatVec4 (returns vec4)
- Vec4Float (returns vec4)
- Mat2Mat2 (returns mat2)
- Mat2Float (returns mat2)
- FloatMat2 (returns mat2)
- Mat3Mat3 (returns mat3)
- Mat3Float (returns mat3)
- FloatMat3 (returns mat3)
- Mat4Mat4 (returns mat4)
- Mat4Float (returns mat4)
- FloatMat4 (returns mat4)
### duo types
- VecXVecX (returns vecx)
- VecXVec1 (returns vecx)
- Vec1VecX (returns vecx)
### duo instructions
- Add
- Sub
- Mul
- Div
- Mod
- Rem
- Pow
- Atan2
- Min
- Max
### trio types
- VecXVecXVecX (returns vecx)
- VecXVecXVec1 (returns vecx)
- VecXVec1VecX (returns vecx)
- VecXVec1Vec1 (returns vecx)
- Vec1VecXVecX (returns vecx)
- Vec1VecXVec1 (returns vecx)
- Vec1Vec1VecX (returns vecx)
### trio instructions
- SmoothMin
- SmoothMax
- Clamp
- Mix
- Step
- SmoothStep
- FMA
### unary types
- Float
- Vec2
- Vec3
- Vec4
- Mat2
- Mat3
- Mat4
- VecX
### unary instructions
- negate
- round
- roundeven
- trunc
- abs
- sign
- floor
- ceil
- fract
- sqrt
- inversesqrt
- exp
- exp2
- log
- log2
- radians
- degrees
- sin
- cos
- tan
- asin
- acos
- atan
- sinh
- cosh
- tanh
- asinh
- acosh
- atanh
- exp
- log
- exp2
- log2
- sqrt
- inversesqrt
- square
- cube
### Extra Matrix
- MulMat2Vec2 (returns vec2)
- MulVec2Mat2 (returns vec2)
- MulMat3Vec3 (returns vec3)
- MulVec3Mat3 (returns vec3)
- MulMat4Vec4 (returns vec4)
- MulVec4Mat4 (returns vec4)
### Extra Vector
- CrossVec3Vec3 (returns vec3)
- DotVec2Vec2 (returns vec2)
- DotVec3Vec3 (returns vec3)
- DotVec4Vec4 (returns vec4)
- LengthVec2 (returns float)
- LengthVec3 (returns float)
- LengthVec4 (returns float)
- DistanceVec2 (returns float)
- DistanceVec3 (returns float)
- DistanceVec4 (returns float)
- NormaliseVec2 (returns float)
- NormaliseVec3 (returns float)
- NormaliseVec4 (returns float)
## Matrix manipulation
- TransposeMat2
- TransposeMat3
- TransposeMat4
- InvertMat2
- InvertMat3
- InvertMat4
## Promotion and Demotion
- PromoteFloatFloatVec2
- PromoteFloatFloatFloatVec3
- PromoteFloatFloatFloatFloatVec4
- PromoteVec2FloatVec3
- PromoteVec2FloatFloatVec4
- PromoteVec2Vec2Vec4
- PromoteVec3FloatVec4
- Promote4FloatMat2
- Promote2Vec2Mat2
- PromoteVec4Mat2
- Promote3Vec3Mat3
- Promote4Vec4Mat4
- PromoteMat2Mat3
- PromoteMat2Mat4
- Promote4Mat2Mat4
- PromoteMat3Mat4
- DemoteVec2FloatFloat
- DemoteVec3FloatFloatFloat
- DemoteVec4FloatFloatFloatFloat
- DemoteMat2Float
- DemoteMat2Vec2
- DemoteMat2Vec4
- DemoteMat3Vec3
- DemoteMat4Vec4
- DotVecX (returns vecx)
- LengthVecX (returns vec1)
- DistanceVecX (returns vec1)
- NormaliseVecX (returns vec1)
- FaceForwardVecX (returns vecx)
- ReflectVecX (returns vecx)
- RefractVecX (returns vecx)
## Data manipulation
### Types
- Float
- Vec2
- Vec3
- Vec4
- Mat2
- Mat3
- Mat4
### Instructions
- Swap2 (a b -- b a)
- Swap3 (a b c -- c b a)
- [Swap4 (a b c d -- d b c a)]
- Dup (a -- a a)
- Dup2 (a b -- a b a)
- Dup3 (a b c -- a b c a)
- [Dup4 (a b c d-- a b c d a)]
- Drop (a -- )
- Drop2 (a b -- b)
- Drop3 (a b c -- b c)
- [Drop4 (a b c d -- b c d)]
- CopyVecX
- PromoteVec1Vec1Vec2
- PromoteVec1Vec1Vec1Vec3
- PromoteVec1Vec1Vec1Vec1Vec4
- PromoteVec2Vec1Vec3
- PromoteVec1Vec2Vec3
- PromoteVec2Vec1Vec1Vec4
- PromoteVec1Vec2Vec1Vec4
- PromoteVec1Vec1Vec2Vec4
- PromoteVec2Vec2Vec4
- PromoteVec3Vec1Vec4
- PromoteVec1Vec3Vec4
## SDF
- SDFSphere
- SDFBox
- SDFTorus
## Extra
- Stop
- Nop
- Position
# Possible extra instructions
- Vector extract/insert dynamic?
- OpVectorShuffle