Target AI + graphics GPU with very high matrix throughput, HBM memory, PCIe/CXL connectivity, and a defined board-power envelope.
A product requirements document that says what the finished GPU must achieve.
Choose core count, L2 size, memory-controller count, matrix units, raster/RT units, NoC topology, and HBM channels.
A top-level architecture specification and early performance/power targets.
Define warp/subgroup scheduling, issue width, register-file ports, matrix-unit tile sizes, cache pipelines, and load/store queues.
A cycle-level hardware plan detailed enough for RTL engineers to implement.
Run GEMM, attention, graphics, cache-stress, and memory-bandwidth simulations before the chip exists.
Performance projections and feedback that continuously changes architecture choices.
Implement the instruction decoder, scheduler, ALUs, register file, matrix engine, caches, memory pipeline, barriers, and atomics.
Synthesizable SystemVerilog/Verilog or equivalent hardware source.
Check matrix instructions, barriers, atomics, page faults, cache coherence, race conditions, and interactions among thousands of concurrent threads.
Coverage reports, passing regressions, formal proofs, bug closures, and verification sign-off.
Add scan chains, MBIST for SRAMs, built-in self-test, JTAG hooks, and production-test modes.
A GPU design that can be tested quickly and reliably after fabrication.
An RTL expression such as a+b becomes flip-flops, adders, multiplexers, buffers, and other standard cells.
A gate-level netlist plus timing, area, and power estimates.
Floorplan compute cores, L2 slices, HBM controllers, SerDes, clock networks, power grids, and large SRAM arrays.
A complete physical layout suitable for manufacturing.
Fit enormous compute blocks together while still closing a multi-GHz clock target.
Placed and routed geometry with clocks and interconnect.
Verify that register → datapath → register paths really sustain the target GPU frequency across process, voltage, and temperature corners.
Timing closure with acceptable setup/hold margins.
Stress the case where many matrix units switch at once and check that local voltage does not collapse.
Power-integrity and reliability sign-off.
Check wire spacing, widths, via structures, density, transistor geometry, and many process-specific constraints.
A layout with no unacceptable manufacturing-rule violations.
Confirm that the placed/routed physical GPU is electrically equivalent to the intended design.
Electrical equivalence sign-off between layout and source design.
Verification, timing, power, DFT, DRC, LVS, reliability, and release criteria all reach acceptable status.
Final release candidate for manufacturing.
The final GDSII / OASIS physical-design database is frozen and released from the chip company into the manufacturing flow. After this point, a hardware bug may require a new stepping and another fabrication cycle.
The released GPU layout is converted into manufacturing mask data for the chosen process.
A production-ready mask set / manufacturing data package.
The actual compute cores, SRAMs, interconnect, I/O, and power network are fabricated.
Finished wafers containing many GPU dies.
Check basic logic, memories, clocks, leakage, and test structures before spending money on advanced packaging.
A map of known-good and bad dies.
Integrate compute dies or chiplets with HBM stacks and an interposer or advanced substrate.
Packaged GPU samples ready for boards and lab bring-up.
The design team receives the first real manufactured GPU samples, often called an early stepping such as A0. This is where simulation becomes reality.
Check power rails, leakage/current, reset, clocks, JTAG/debug access, and whether the GPU responds safely.
A living chip that can proceed to deeper initialization.
Clocks → reset → firmware → PCIe → HBM/GDDR → GPU cores → driver → first compute kernel.
Stable enough hardware/software stack for systematic validation.
The host enumerates the GPU over PCIe, firmware runs, memory initializes, and the driver loads.
A software-visible GPU device.
Run GEMM, attention, graphics, cache tests, memory stress, fault handling, clocks, thermal limits, and long-running regressions.
A list of silicon errata, validated limits, and confidence that the chip matches its specification.
Find safe voltage/frequency curves, thermal behavior, HBM margins, power limits, and process variation.
Production voltage, frequency, power, and binning tables.
Tune GEMM and attention tile sizes, register usage, cache policy, compiler scheduling, occupancy, clocks, and firmware controls.
A software stack that gets close to the hardware's real performance potential.
A0 → A1 → B0 may fix logic bugs, timing problems, analog issues, reliability problems, or manufacturability issues.
A corrected hardware revision.
Temperature cycling, voltage stress, aging tests, package reliability, memory stress, and long-duration operation.
Approval that the part is suitable for production shipment.
Improve process yield, repair/fuse strategies, binning, voltage targets, and usable SKU coverage for a very large die.
Economically viable high-volume production.
The GPU is manufactured in volume and shipped to customers, OEMs, board partners, or data-center systems. For a large modern GPU, this may be roughly 9–18 months after tapeout, depending on complexity and respins.