Compute-side execution model
Xe-Core — Intel Xe compute core
TD — Thread Dispatcher
Thread Dispatcher
Arbitrates thread-initiation requests and dispatches / instantiates work onto execution resources.
→
XVE — Xe Vector Engine
SIMD execution engine
Contains hardware-thread state, register state and vector execution resources.
XVE — Xe Vector Engine
Modern Intel terminology
Thread Control
Hardware-thread control
Tracks execution state and runnable hardware-thread work inside the vector engine.
→
GRF — General Register File
Per-hardware-thread register state
Stores operands, addresses, temporaries and accumulators used by execution instructions.
FP — Floating-Point
FP32 / FP16 / BF16 and supported FP formats
INT — Integer
Integer arithmetic, indexing, address calculations
EM — Extended Math
Special / transcendental math such as reciprocal and square root; exact functions vary
FP64 — 64-bit Floating-Point
Double-precision execution where supported
Branch
Control-flow execution
Send
Message / memory-access path to other GPU functions and memory resources
XMX — Xe Matrix Extensions
Matrix acceleration engine
Specialized systolic-array style matrix / dot-product acceleration available on supported Xe architectures.
DPAS — Dot Product Accumulate Systolic
Instruction / operation that drives XMX matrix dot-product accumulation.
Send → Memory Subsystem
Execution-to-memory interface
The Send path connects execution with caches, SLM and other hardware functions. It is more faithful to Intel ISA terminology than a generic standalone “Load/Store Pipeline” block.
Legacy EU execution view: FPU0 / FPU1
Older Intel Gen / Xe-LP-oriented terminology
FPU0 — Floating-Point Unit 0
Legacy SIMD arithmetic pipeline
Intel legacy documentation describes FPU0 as handling ordinary floating-point operations and integer operations.
Conceptual role:
FP — Floating-Point + INT — Integer
FP — Floating-Point + INT — Integer
FPU1 — Floating-Point Unit 1
Legacy SIMD arithmetic / extended-math pipeline
Intel legacy documentation describes FPU1 as handling floating-point operations and Extended Math operations; it is sometimes associated with the EM unit.
Conceptual role:
FP — Floating-Point + EM — Extended Math
FP — Floating-Point + EM — Extended Math
Do not map these directly onto the modern Xe2 block diagram.
FPU0 / FPU1 are useful when reading older Intel EU and performance-counter documentation.
For modern Xe mental models, prefer XVE — Xe Vector Engine with FP / INT / EM / FP64 / Branch / Send resources.
Overview
A thread is dispatched to an execution engine. Inside the Xe Vector Engine, thread state and GRF operands feed FP, INT, Extended Math, FP64, Branch or Send execution. DPAS drives the XMX matrix engine.
Modern Xe view
Uses Intel's XVE / XMX terminology rather than treating FPU0/FPU1 as the primary modern architecture.
Instruction vs hardware
DPAS is shown as an instruction / operation; XMX is the matrix hardware engine.
Legacy terms removed
SP0 / SP1 and the prior FPU0/FPU1 sub-pipeline mapping are omitted because they are not a safe universal Xe model.