Hot Chips 2026 Day 1 | CPU and GPU Systems

Hot Chips 2026 Day 1 | CPU and GPU Systems
  • Intel, AMD and NVIDIA are competing over a wider part of the computing system.
    • CPUs are being designed around their role beside accelerators.
    • GPUs are being designed around memory movement, communication and continued operation through faults.
  • Delivering more useful work depends on how the processor, memory, software and surrounding system operate together.
  • This article covers seven Day 1 sessions from Hot Chips 2026, focusing on Intel, AMD and NVIDIA.

1. Faster GPUs make CPU delays more costly

(1) NVIDIA Vera: keep the accelerator supplied with work

NVIDIA (NASDAQ: NVDA) designs Vera for the CPU work around AI agents:

  • a) preparing data,
  • b) running tools,
  • c) and coordinating tasks.

As GPU execution gets faster, a slow CPU stage can occupy more of the complete response time. A faster CPU helps when its stage delays the next dependent operation.

Vera’s cores, memory and CPU–GPU connections were covered in Vera Rubin Decoded, Part 3. The focus here is how those resources affect the CPU stages that an AI agent must wait for.

1) An agent must wait for its tool results

An agent can answer after model inference or request CPU tool work whose results feed the next model step
  • The blue path sends a tool request, such as search or code execution, to CPU services.
  • The green path returns the result to the model.
  • The next dependent model step waits for that result. A faster GPU cannot finish a CPU tool call that is still running.
  • CPU response time under load therefore affects how much time the accelerator spends waiting.
Conventional SMT and Vera Spatial Multithreading resource use
  • In the blue SMT example, two hardware threads share a core's resources and can slow one another down.
  • Vera reserves shares of selected structures for each thread, reducing that interference.
  • The center SPECint test changes the neighboring thread's workload to measure the effect.
  • It is not an agent-response benchmark. Shared cache and memory can still cause contention.

2) Memory power competes with the cores' power budget

Vera's LPDDR5X bandwidth per watt compared with the stated DDR5 memory configurations
  • The vertical measure is memory bandwidth per watt. The green Vera bar is labeled 5× against the two DDR5 configurations normalized to 1.0.
  • Supplying the needed bandwidth with less memory power leaves more of the budget for active cores. This is a memory-configuration comparison, not a fivefold gain in server efficiency.

3) Integration determines where data still has to travel

Vera's cores and shared cache connected by the scalable coherency fabric
  • Vera uses NVIDIA-designed Olympus cores implementing Arm’s instruction set. Arm (NASDAQ: ARM) supplies the architecture; NVIDIA designs the cores and their surrounding system.
  • The Core and shared-cache blocks stay on the compute die. Memory interfaces and system I/O sit on separate dies.
  • A core first benefits from data already in the shared cache.
  • Reusing that data avoids a trip off the compute die to DRAM.
  • Vera and Intel make different placement choices. Compare them by the resulting waiting time, power and package cost.

2. Packaging choices follow the product's economics

  • More elaborate packaging is worthwhile when its performance or integration benefit justifies the extra manufacturing work. Diamond Rapids uses dense connections to scale server compute; Wildcat Lake simplifies the assembly for a lower-cost client product.

(1) Intel Diamond Rapids: put dense connections where reuse matters

  • Intel (NASDAQ: INTC) separates compute, cache and shared memory / I/O services across chiplets. This lets each function use a suitable manufacturing process while keeping frequently reused data close to the cores.

1) Separate manufacturing choices without making every access longer

Compute building blocks fan out to two Fabric Hubs that provide memory, I/O, and accelerators
  • The four outer red boxes locate the compute-and-cache groups. They connect directly to both central red-outlined Fabric Hubs, which supply memory and device access. Direct paths avoid sending some requests through one hub and then another.
Diamond Rapids die counts, process choices, and package links
  • Intel’s Foveros 3D Direct hybrid bonding joins compute chiplets to cache-bearing base tiles. Lateral package links connect these groups to the hubs.
  • Compute uses Intel 18A-P, while the base tiles and hubs use other process variants. The newest process is concentrated on compute instead of being required for every function in the package.
  • That flexibility creates demand for bonding, package links and validation. Cache reuse must reduce enough external traffic to make the resulting compute scale useful; adding cores also adds demand on the shared memory system.

2) Avoid spending memory bandwidth on bookkeeping

Coherence metadata moves from DRAM to an on-chip snoop filter
  • The red annotations trace the directory: bookkeeping that helps coordinate copies of data held in different cores' caches.
  • If that directory sits in DRAM, checking it consumes off-chip traffic.
  • Moving it onto the processor avoids that lookup traffic.
  • It also frees DRAM checking bits that the directory had occupied, leaving more space for memory-error protection.
  • Reducing the traffic that must cross a data path can improve useful execution without widening the path.

(2) Intel Wildcat Lake: accept more interface silicon to simplify assembly

  • Wildcat Lake adapts an existing client design to a simpler multichip package. It illustrates why reducing total product cost can justify making one part of the chip larger.

1) Save a package component, then pay for a different connection

Foveros packaging compared with Wildcat Lake's multichip package
  • The left design uses Intel’s Foveros packaging, with the red-outlined passive base die beneath the active dies supplying dense connections. Wildcat Lake’s simpler multichip package removes that base die, saving a silicon component and assembly work. The remaining package must still connect the active dies.
Interface area increases while the simpler package saves cost
  • The red box marks the larger redesigned interface.
  • Wider connection spacing leaves fewer physical connections per area.
  • Faster signaling and packet handling carry the traffic through those connections, but need more interface circuitry.
  • The interface grows by roughly 70% in the disclosed comparison.
  • Intel still reports lower total cost because removing the base die saves silicon, assembly work and associated yield loss.
  • A smaller chip block and a cheaper finished product are different objectives. Demand for advanced packaging depends on whether its benefit justifies the cost for that product tier.

2) Simpler assembly still needs engineering work

  • Packetizing the link can change when data and control signals arrive. Validation must keep a control event from triggering before its data is ready.
  • Power-state changes must also preserve client response deadlines. These costs help explain why a package change is a product redesign even when much of the underlying silicon is reused.

3. GPU performance increasingly depends on movement and coordination

(1) NVIDIA Rubin: reduce selected work and release consumers sooner

Rubin reduces execution time and interruptions through several mechanisms:

  • a) reduce selected attention work,
  • b) shorten coordination between cooperating GPUs,
  • c) and keep the system working through power constraints and diagnostics.

The weight format and transfer controls below also reduce stored bytes and repeated setup work.

Vera Rubin Decoded, Part 2 previously covered the GPU’s compute blocks, HBM4 and package connections. The Hot Chips details below explain how Rubin reduces work and coordinates data movement within that architecture.

1) Reduce attention work after score calculation

Rubin sparse attention pipeline and the operations after attention-score calculation
  • BMM1 calculates query–key scores, which indicate how relevant the context entries are to the current query. This calculation runs for all entries.
  • The green dashed region begins after BMM1. Its sparsify block removes selected near-zero entries and records the positions of those retained, reducing the data passed to softmax and BMM2.
  • Softmax converts the retained scores into attention weights. BMM2 combines the corresponding value data using those weights to produce the attention result.
A query and cached keys produce attention weights, which combine cached values into an attention result
  • The blue query–key path produces the scores that become attention weights; the green path combines values using those weights. Rubin reduces work after the initial scoring step.
  • Consider a group of four context entries:
    • Calculate relevance scores for all four.
    • The 2:4 sparsification step retains two entries and drops the other two from subsequent processing.
    • Run the later softmax and value-combination calculations using the retained entries. The dropped entries' later calculations are skipped entirely.
  • The work is reduced because those later calculations never happen. The saving is the skipped computation, minus the extra work needed to select and encode the retained entries.
  • The initial scoring cost remains. The slide's claimed 2× acceleration applies to downstream softmax and BMM2, so it does not establish a halving of total attention time or complete-model execution time.

2) Counted writes: know when received data is ready

  • When one GPU sends data to another, the receiving GPU must know when the input is complete before calculating with it. Reading too early could mean using only part of the new data or values left from earlier work.
Counted writes combine data arrival with completion tracking
  • Read both panels from top to bottom as time passes. GPU 0 sends data across NVLink into GPU 1's memory; GPU 1's processing units then load that data to use it.
  • On the left, the illustrated Blackwell sequence sends the data and then performs a separate completion-signaling sequence:
    • The memory barrier enforces the required ordering of the earlier writes.
    • An acknowledgment returns to the sender.
    • The sender updates a flag telling the receiver that the data is ready.
    • The receiver checks that flag before loading the input. These extra steps add waiting and communication around the actual payload.
  • On the right, Rubin's counted writes connect completion tracking to the data writes themselves:
    • As the required writes complete, hardware updates a counter that the receiver can check.
    • Once the counter indicates that its required input is complete, the receiver can load it without the same separate barrier, acknowledgment and flag sequence shown on the left.
    • The orange polling arrows remain: the receiver still checks readiness, but it has a shorter completion path to wait for.
  • As a simplified example, suppose GPU 1 needs three pieces of input from GPU 0:
    • Receiving only the first two is insufficient; the calculation needs all three.
    • The completion counter tracks progress toward the required input being ready. Once all three required pieces are complete, GPU 1 can proceed.
    • This example illustrates the readiness condition, not the hardware's precise counter units.
  • The saving comes from fewer coordination steps between sending data and using it. The payload still crosses the link, and the receiver still waits for every piece its calculation requires.
  • This matters when distributed experts exchange many small inputs: a repeated completion handshake can consume a significant share of each transfer's time. Shortening it can let dependent calculations start sooner.
  • Other independent work can overlap the transfer. For example, GPU 1 can compute with an already-ready block while the next block arrives; computation on that next block starts only after its own input is complete.

3) Control abrupt changes in power demand

Power demand before and after ramp and steady-state smoothing
  • The black trace shows abrupt changes in power demand. The green trace smooths the rise, running period and fall. The axes have no numerical scales.
  • At startup, a ramp-up cap limits how quickly demand rises.
  • During operation, energy storage can absorb or supply short-term differences.
  • At shutdown, the GPU can deliberately keep some activity running to avoid an abrupt drop.
  • This makes demand easier for the surrounding electrical system to handle. The deliberate activity also uses energy, so a smoother trace does not itself establish lower energy per job.

4) Diagnose hardware while useful work continues

Rubin overlaps in-situ health checks with execution and provides repair and recovery mechanisms
  • The lower timeline inserts short health checks while useful work continues.
  • A problem can therefore be detected without first taking the node offline for the check.
  • Specific repair mechanisms address specific resources:
    • SRAM error correction and repair cover local storage.
    • HBM bank remapping replaces a faulty bank with usable memory resources.
  • Concurrent checks reduce diagnostic downtime. A fault that interrupts computation still requires recovery, potentially using saved state.

5) Reduce weight storage and repeated transfer setup

Store a smaller code for each weight

A weight is a learned number the model uses in its calculations. Rubin’s lookup-table format represents each weight with a small code that selects a value from a shared table. This reduces the amount of weight data stored and read.

How the code selects a weight

For one 8 × 64 block, containing 512 weights:

  1. Choose eight representative weight values and store them in a shared table. Each value uses FP8, an 8-bit floating-point format. Number the entries 0–7.
  2. Represent each weight with the index, or entry number, of its selected value. Choosing among eight entries requires only 3 bits.
  3. During calculation, the matrix instruction looks up the selected value and uses it directly, without first writing a separate expanded weight matrix to memory.

As an illustrative example, suppose entry 5 contains 0.5. A weight represented by 0.5 stores index 5, written as 101 in binary. The hardware reads that code, retrieves 0.5 from the table and uses it in the calculation. With an input of 4, the multiplication is 4 × 0.5 = 2.

The index identifies the weight value to use; it is not itself the weight.

Direct FP8 storage uses 512 bytes; 512 three-bit indices plus an eight-value table use 200 bytes. The example follows code 5 to weight 0.5 and then multiplies an input of 4 by that value.

Why this reduces storage

The full values are stored once in the shared table. Each of the 512 weight positions stores only a smaller code, and many positions can select the same table entry.

Reduce weight storage and repeated transfer setup
  • The indexed format uses about 61% less storage than direct FP8 in this comparison, or 3.125 bits per weight including the shared table. Additional scaling metadata, alignment and other stored model data are excluded.
  • The number of weights remains 512. Fewer bytes need to be read from memory, which can shorten decode when weight delivery is the bottleneck. KV reads, activation transfers and arithmetic still remain.
  • The tradeoff is model accuracy. Original weights may contain many different values, so restricting a block to eight representative choices introduces approximation. Model quality must be checked; an FP8 table does not give every weight the full range of independent FP8 choices.

Reuse a transfer descriptor when switching experts

The illustrated Blackwell path uses expert-specific transfer descriptors; Rubin reuses a template with an inline address override
  • A tensor memory accelerator (TMA) descriptor tells the data mover where a tensor is stored and how it is arranged.
  • The left panel shows descriptors for individual experts. The right panel reuses one template for experts with the same shape and layout, supplying the selected expert’s address with the transfer instruction.
  • Switching from expert A to expert B can therefore change the address without rewriting and synchronizing a descriptor in memory between loads. The expert weights still have to move from HBM to shared memory.
  • NVIDIA’s instruction documentation describes both the lookup operation and the transfer-metadata overrides. These mechanisms reduce different costs: weight bytes and transfer setup.

6) Read inference results at the same service target

  • SemiAnalysis’s July analysis discusses early DeepSeek R1 tests with an 8K-token prompt and 1K-token output.
  • Its September AgentX analysis uses multi-turn DeepSeek V4 Pro 0831 (FP4), a 1.6-trillion-parameter model, with substantial cached context. Those are different workloads and software snapshots.
Rubin and two GB300 software configurations compared at matched P90 interactivity, with throughput normalized to utility power
  • The first column sets P90 interactivity: the reciprocal of the 90th-percentile inter-token latency measured over complete responses. It describes token delivery speed, not the wait for the first token or the duration of the whole agent task.
  • The middle columns count millions of total tokens per second per all-in utility megawatt. Total tokens include cached inputs. The values are interpolated at matched interactivity, and the power basis includes the source’s facility assumptions.

Rubin chart 1: total traffic per utility megawatt

These September AgentX charts count cached inputs, new inputs and outputs in total tokens.

SemiAnalysis compares total agentic token throughput per all-in utility megawatt against P90 per-user token speed for Rubin, two GB300 software configurations and MI355X
  • Horizontal: P90 tokens/s/user, calculated as the reciprocal of the 90th-percentile full-response inter-token latency. Farther right means faster token delivery.
  • Vertical: Total tokens/s per modeled all-in utility MW.
Read inference results at the same service target

Rubin chart 2: total tokens per ownership dollar

SemiAnalysis’s published ownership-cost screenshot compares Rubin and GB300 TRT-LLM total tokens per dollar at matched P90 interactivity and annotates the 67.4-times operating point
  • Horizontal: The same P90 user-speed measure. Vertical: Total tokens per dollar of modeled total cost of ownership (TCO). Both axes are linear.
  • Light green: Rubin TRT-LLM. Dark green: GB300 Dynamo TRT-LLM; GB300 SGLang is not selected.
  • 67.4× near 170 tokens/s/user reflects this comparison near GB300 TRT-LLM’s low-throughput endpoint. It is not a general model-speed multiplier.
  • The large-hyperscaler ownership case depends on amortized hardware and operating costs. The screenshot retains its unofficial-hosting and Rubin-preview notices; these published results were not independently reproduced.

(2) AMD MI455X: build a large memory system and make its bandwidth usable

MI455X follows three linked priorities:

  • a) enlarge the memory system,
  • b) improve the arithmetic used by AI kernels,
  • c) and keep those arithmetic units supplied through explicit data movement.

The last two determine how much of the larger memory resource becomes useful execution.

1) Organize compute around HBM and shared cache

  • Advanced Micro Devices (NASDAQ: AMD) gives MI455X 432 GB of HBM4 and 23.3 TB/s of memory bandwidth. Its chiplet design separates compute, fabric / cache and external I/O.
MI455X compute chiplets, fabric and cache dies, I/O dies, and HBM4
  • The red outlines in the left half identify representative components:
    • The large central rectangle surrounds one fabric-and-cache die (FCD).
    • The smaller rectangle inside it marks one compute die (XCD).
    • The tall rectangle at the outer edge marks an I/O die (IOD).
    • The top rectangle marks a row of HBM stacks.
  • Across the full package, two fabric-and-cache dies connect eight compute dies to twelve HBM stacks. Two I/O dies handle external connections.
  • TSMC (NYSE: TSM) supplies two different packaging technologies for this arrangement:
    • SoIC hybrid bonding stacks the compute dies directly on the fabric-and-cache dies.
    • CoWoS-L packaging connects those assemblies with the I/O dies and HBM in the larger package.
  • AMD’s compute design therefore depends on both dense vertical bonding and the package that connects it to memory.
  • Follow a data request through the possible paths:
    • If reusable data is in cache, compute can avoid another HBM read.
    • On a cache miss, the request continues through the fabric to HBM.
    • If the data belongs to another GPU, it must also cross an external connection.
  • Each path has its own bandwidth and delay.
  • A shared cache reduces repeated reads, but conflicting updates to one shared value still need coordination.

2) Match arithmetic and representation to AI kernels

  • AMD is improving both the main AI calculations and the smaller operations around them, so the whole sequence can finish sooner.
CDNA 5 combines native wave32 execution, low-precision matrix arithmetic, and transcendental acceleration

a) Handle groups of parallel work more efficiently

  • A GPU runs many threads (small units of work) together.
  • Wave32 groups 32 threads for execution.
  • AMD's native support helps these groups execute more efficiently. This concerns how work runs inside the GPU.

b) Use fewer bits per number, while controlling the loss of accuracy

  • AI calculations do not always need highly precise numbers.
  • Formats such as FP4 represent each value with fewer bits, reducing data movement and allowing suitable hardware to perform more arithmetic.
  • The tradeoff is greater rounding error.
  • Block scaling helps manage that tradeoff: a small group of values shares a multiplier.
    • As a simplified illustration of scaling, storing 1, 2, 3 with a shared multiplier of 100 represents 100, 200, 300.
    • Giving different groups their own scale helps accommodate different numerical ranges.
    • It does not recover all the precision lost, so model quality still needs checking.

c) Speed up the operations between matrix calculations

  • Attention includes matrix multiplication, but also operations such as softmax, which turns relevance scores into attention weights.
  • If matrix multiplication becomes faster while softmax stays slow, more of the total execution time is spent waiting on softmax.
  • AMD therefore improves hardware for those supporting mathematical functions too. The slide's enhanced transcendental engine addresses this work.
  • Higher peak arithmetic improves model execution only when the model tolerates the lower precision and the surrounding operations keep pace.

3) Overlap tensor movement with computation

  • The GPU can spend less time waiting by loading its next block of data while calculating with the current one.
Tensor Data Mover beside workgroup storage and arithmetic, with excess vertical space removed

a) Move data into nearby working storage

  • A tensor is an array of data used by the model. A tile is a smaller block taken from that array.
  • The gold Tensor Data Mover (TDM) transfers tiles from global memory into local data share (LDS), the nearby working storage.
  • The SIMD blocks perform arithmetic using that nearby data.
  • TDM can load data into LDS without first passing it through arithmetic registers. This leaves those resources available for computation.

b) Load the next tile while using the current one

Compute uses buffer A while a transfer fills B, then the roles change after transfer and consumption complete
  • A buffer is a reserved area of working storage. This example uses two buffers, A and B:
    • First, place the current tile in A.
    • Compute with A while TDM loads the next tile into B.
    • Once B is ready, computation can switch to B. The green arrow marks that readiness condition.
    • Once computation has finished using A, TDM can refill A with a later tile. The blue arrow marks that release condition.
    • Repeat by alternating the buffers. This is called double buffering.
  • Completion checks prevent reading unfinished data or overwriting data still in use.
  • If the next tile arrives late, computation still waits.

c) Avoid loading the same input repeatedly

  • When several workgroups need the same input, multicast can distribute it to those groups instead of requiring separate transfers for each one.
  • Overlapping transfers reduces waiting; sharing inputs reduces repeated traffic. Software must arrange both to turn the hardware capability into faster execution.

4) Separate measured results from resource ceilings

Measured MI455X memory, compute, and networking results
  • Read the four columns as separate tests:
    • Memory reads test how quickly data reaches compute.
    • FP4 arithmetic tests calculation in a 4-bit floating-point format, subject to model-quality requirements.
    • Scale-up tests communication within a closely connected accelerator group.
    • Scale-out tests communication across the wider network.
  • The multiples compare each test with MI355X.

4. AMD Helios makes the rack part of the architecture

(1) Scale local resources across the rack

Helios extends AMD's offer from a GPU to a rack-scale computing system. Its performance approach connects three levels:

  • a) connect GPUs within the rack,
  • b) access data in another GPU’s memory,
  • c) and move data beyond the rack through programmable Ethernet.
Helios groups four GPUs per tray and connects eighteen trays through scale-up planes
  • Eighteen trays of four GPUs give the rack 72 devices.
  • The upper switch planes provide alternative routes between trays.
  • Traffic can use different routes, or surviving routes after a path fails.
  • The rack's memory remains distributed: reaching another GPU's HBM still requires a network transfer.
  • The rack holds about 31 TB of HBM in total, expanding the models and concurrent requests it can accommodate. That capacity remains distributed across the GPUs.
  • Reaching another GPU's memory uses narrower peer links and shared switches. Adding GPUs expands local resources, while effective scaling depends on how much of the workload can stay local and how efficiently the remaining exchanges run.

(2) Map and access a peer GPU buffer

An importing GPU accessing an exported peer buffer through UALoE
  • First, software authorizes access to a buffer in another GPU's memory.
  • It then maps that buffer so the importing GPU can address it.
  • In the diagram, the request crosses UALoE blocks and switches to reach the remote HBM.
  • UALoE carries these transactions over Ethernet using a specialized transport.
  • The disclosed Helios fabric uses Broadcom (NASDAQ: AVGO) Tomahawk 6 switch ASICs to carry traffic between GPUs. AMD integrates the UALoE adapters into its GPUs.
  • Celestica (NYSE: CLS) is the R&D, design and manufacturing partner for the scale-up networking switch. Broadcom supplies the switching silicon; Celestica turns that silicon into the switch hardware used by the rack.

(3) Program the network adapter for traffic beyond the rack

Vulcano combines host interfaces, programmable transport and packet engines, and an 800 Gb/s Ethernet path
  • Vulcano supplies the Ethernet path to wider systems and storage.
  • Follow the diagram from the host interfaces through the red-outlined P4DMA transport engines, P4NET packet engines, and MAC Ethernet interface.
  • Those engines handle transfers, retry missing data and regulate traffic when paths are congested.
  • Their programmable behavior can adapt to changing AI traffic without assigning every protocol operation to the host CPU.
  • Together with the scale-up fabric, Vulcano lets AMD supply both communication within the rack and networking to wider systems. Coordinating those paths can reduce the integration work needed to keep accelerators supplied with data.

5. Intel Crescent Island targets a different inference envelope

(1) Choose a capacity and power envelope

Crescent Island capacity, power, and memory configuration
  • The slide's 480 GB maximum is a family option. Intel's own card has 160 GB; larger capacities belong to partner designs.
  • More capacity can keep a larger model and its KV attention state on one card.
  • That can avoid additional devices and the transfers between them.
  • Crescent Island combines that LPDDR5X capacity with a 350 W air-cooled PCIe design. It targets inference deployments where holding the model on a manageable number of cards matters alongside arithmetic throughput.
  • The air-cooled card can fit servers designed for that power and airflow, avoiding a liquid-cooling retrofit. UBS’s August 31 review highlights this installation advantage for enterprises adding inference capacity to existing infrastructure.
  • The workload argument is that a mixture-of-experts model can store many weights while reading only selected experts for each token.
    • That can favor a system with substantial capacity relative to active bandwidth demand.
  • More conversations need more KV storage, and a batch may select more distinct experts. A faster per-user target also leaves less time to read those data. The card therefore needs both sufficient capacity and sufficient delivery speed for the intended workload.

(2) Keep intermediate results near arithmetic

Xe3 pairs vector and matrix engines around a shared register file
  • Matrix engines perform dense products. Vector engines handle surrounding operations on individual values.
  • The red boxes locate register storage inside the vector-engine detail and the XMX matrix engine on the right.
  • Attention can pass an intermediate result from matrix work to an operation such as softmax through that nearby storage.
Keeping intermediate results nearby avoids a spill and reload between dependent calculations.
  • A fused kernel combines dependent operations so their intermediate values can stay nearby.
  • If the local storage runs out, some values must be written farther away and loaded again. This is called spilling.
  • More nearby storage can avoid those extra transfers when software uses it effectively.

(3) Trade additional computation for faster token generation

  • Speculative decoding can perform more arithmetic while producing tokens faster. The aim is to reduce the time spent on sequential generation steps.
A draft model proposes candidate tokens and the target verifies them before accepted progress updates the conversation
  • First, the smaller draft model proposes several candidate tokens along the upper path.
  • The larger target model checks those candidates together along the lower path.
  • Accepted candidates let generation advance several tokens in one verification pass.
  • In the example, A and B are accepted. C and the proposed suffix D are discarded; recovery depends on the algorithm.
  • Drafting adds computation, and rejected candidates waste some verification work.
  • However, checking several candidates together can reuse loaded model weights and reduce the number of sequential target-model passes needed for the accepted output.
  • When memory delivery is the bottleneck, arithmetic units may have spare capacity.
  • Using that spare capacity to reduce repeated memory reads can shorten generation time despite doing more calculations.
  • If arithmetic resources are already busy, or too few candidates are accepted, the extra drafting and verification work can outweigh the benefit.
  • The relevant gain is less time per accepted token; fewer calculations or lower energy use are not guaranteed.
  • Speculation works on possible future tokens within one request. Batching instead groups next-token work from different requests.
Accepted tokens per verification and the cost of larger draft trees
  • The left bars measure accepted tokens per verification pass across named models, with ordinary autoregressive decode at one.
  • On the right, draft-tree size increases along the horizontal axis.
    • A larger tree tests more candidate continuations.
    • The red compute-per-pass line rises faster than the blue accepted-token line.
    • Verification work therefore grows faster than useful output in this example.
  • More speculation does not automatically mean faster or cheaper inference. The added computation must produce enough accepted tokens to justify its cost.
  • Those figures come from third-party SpecBundle tests with EAGLE-3 and SGLang, not measured Crescent Island token throughput.

(4) Connect the hardware to serving frameworks

Crescent Island software layers connect model-serving frameworks to compilers, libraries, and runtime execution
  • Serving frameworks such as vLLM and SGLang accept model work.
  • Compilers and libraries map it onto the card's local storage and supported arithmetic.
  • The runtime launches the resulting operations.
  • Deployment requires the full software and hardware path to run the customer's model at the required speed. Capacity and power specifications alone do not establish that result.

6. Compare the key architectural choices

(1) CPU design and packaging

CPU design and packaging

(2) GPU execution and networking

GPU execution and networking

Hot Chips 2026 Coverage