Intel Xe GPU Execution Model

Intel-document-aligned mental model for modern Xe execution: thread dispatch, XVE, GRF, ALUs, XMX, Send, SLM and cache hierarchy
Modern Xe view · generation details may vary
Compute-side execution model
Xe-Core — Intel Xe compute core
TD — Thread Dispatcher
Thread Dispatcher
Arbitrates thread-initiation requests and dispatches / instantiates work onto execution resources.
→
XVE — Xe Vector Engine
SIMD execution engine
Contains hardware-thread state, register state and vector execution resources.
XVE — Xe Vector Engine Modern Intel terminology
Thread Control
Hardware-thread control
Tracks execution state and runnable hardware-thread work inside the vector engine.
→
GRF — General Register File
Per-hardware-thread register state
Stores operands, addresses, temporaries and accumulators used by execution instructions.
FP — Floating-Point FP32 / FP16 / BF16 and supported FP formats
INT — Integer Integer arithmetic, indexing, address calculations
EM — Extended Math Special / transcendental math such as reciprocal and square root; exact functions vary
FP64 — 64-bit Floating-Point Double-precision execution where supported
Branch Control-flow execution
Send Message / memory-access path to other GPU functions and memory resources
XMX — Xe Matrix Extensions
Matrix acceleration engine
Specialized systolic-array style matrix / dot-product acceleration available on supported Xe architectures.
DPAS — Dot Product Accumulate Systolic Instruction / operation that drives XMX matrix dot-product accumulation.
Send → Memory Subsystem
Execution-to-memory interface
The Send path connects execution with caches, SLM and other hardware functions. It is more faithful to Intel ISA terminology than a generic standalone “Load/Store Pipeline” block.
Legacy EU execution view: FPU0 / FPU1 Older Intel Gen / Xe-LP-oriented terminology
FPU0 — Floating-Point Unit 0
Legacy SIMD arithmetic pipeline
Intel legacy documentation describes FPU0 as handling ordinary floating-point operations and integer operations.
Conceptual role:
FP — Floating-Point + INT — Integer
FPU1 — Floating-Point Unit 1
Legacy SIMD arithmetic / extended-math pipeline
Intel legacy documentation describes FPU1 as handling floating-point operations and Extended Math operations; it is sometimes associated with the EM unit.
Conceptual role:
FP — Floating-Point + EM — Extended Math
Do not map these directly onto the modern Xe2 block diagram. FPU0 / FPU1 are useful when reading older Intel EU and performance-counter documentation. For modern Xe mental models, prefer XVE — Xe Vector Engine with FP / INT / EM / FP64 / Branch / Send resources.
Overview
A thread is dispatched to an execution engine. Inside the Xe Vector Engine, thread state and GRF operands feed FP, INT, Extended Math, FP64, Branch or Send execution. DPAS drives the XMX matrix engine.
Modern Xe view Uses Intel's XVE / XMX terminology rather than treating FPU0/FPU1 as the primary modern architecture.
Instruction vs hardware DPAS is shown as an instruction / operation; XMX is the matrix hardware engine.
Legacy terms removed SP0 / SP1 and the prior FPU0/FPU1 sub-pipeline mapping are omitted because they are not a safe universal Xe model.
Items and full names
Item Full name Category Meaning / role
Xe-Core Xe Core Modern Xe Intel GPU compute core containing vector and matrix execution resources plus local memory resources.
TD Thread Dispatcher Modern Xe Arbitrates thread-initiation requests and dispatches / instantiates work onto execution resources.
XVE Xe Vector Engine Modern Xe SIMD execution engine containing hardware-thread state, GRF state and vector execution resources.
EU Execution Unit Legacy Older Intel terminology for the GPU execution unit used in Gen-era documentation.
GRF General Register File Modern Xe High-bandwidth register storage holding operands, addresses, temporaries and accumulators for a hardware thread.
FP Floating-Point Execution Floating-point arithmetic resource class, including supported FP formats such as FP32 / FP16 / BF16 depending on architecture.
FP64 64-bit Floating-Point Execution Double-precision floating-point execution resource where supported.
INT Integer Execution Integer arithmetic used for general integer work, indexing and address calculations.
EM Extended Math Execution Special / transcendental math resource class such as reciprocal, square root and generation-dependent special functions.
Branch Branch / Control-Flow Execution Execution Handles control-flow instructions.
Send Send Instruction / Message Path Instruction path Connects the execution engine to caches, SLM and other GPU hardware functions through message-based operations.
XMX Xe Matrix Extensions Matrix hardware Intel matrix acceleration engine for systolic / dot-product style matrix operations on supported Xe architectures.
DPAS Dot Product Accumulate Systolic Instruction Dot-product accumulate instruction / operation used to drive XMX matrix computation.
FPU0 Floating-Point Unit 0 Legacy Legacy SIMD arithmetic pipeline described as handling ordinary floating-point and integer operations.
FPU1 Floating-Point Unit 1 Legacy Legacy SIMD arithmetic pipeline described as handling floating-point and Extended Math operations.
ALU Arithmetic Logic Unit Generic / legacy Generic term for arithmetic / logical execution hardware; older Intel material may use ALU0 / ALU1 language alongside FPU terminology.
SLM Shared Local Memory Memory Programmer-managed work-group-local scratchpad memory; not a cache level.
L1 Level 1 Cache Memory Near-core hardware-managed cache in Intel Xe memory hierarchy.
L2 Level 2 Cache Memory Last-level cache for Xe-HPG / Arc and Xe2 Arc-oriented products; exact organization and naming vary by generation.
HBM High Bandwidth Memory Device memory High-bandwidth off-core device memory used by some Intel GPU products.
GDDR Graphics Double Data Rate Memory Device memory Graphics-oriented off-core device memory used by some Intel GPU products.
HP Half Precision Precision term Half-precision floating-point category, typically referring to FP16 in older Intel terminology.
SP Single Precision Precision term Single-precision floating-point category, typically FP32.
DP Double Precision Precision term Double-precision floating-point category, typically FP64.