The useful question about Jalapeño is not whether OpenAI has built a smaller GPU. It is whether an inference operator can keep a complete request inside one predictable execution domain while the bottleneck changes underneath it. A prompt begins with compute-heavy prefill, may enter a latency-sensitive draft loop, and then spends verification time reading model state and moving expert traffic. A fleet of individually efficient devices can still waste energy and time if every phase transition also moves KV cache state, crosses a network boundary or leaves another package powered but idle.
OpenAI’s Hot Chips 2026 presentation makes locality the architectural response. Jalapeño combines a large HBM4 memory system, 64 spatial compute slices, a fast collective path, a separate general network-on-chip and an Ethernet hierarchy that reaches from a 128-ASIC local domain to a 2,048-ASIC global system[1]. OpenAI then evaluates the platform with SemiAnalysis InferenceX on three open models and reports higher throughput per kilowatt together with lower end-to-end latency than the selected NVIDIA systems[2]. Those measurements deserve attention, but the denominator matters: the published comparison normalizes by accelerator package TDP rather than matched rack input power.
We read Jalapeño as a full-stack inference proposition, not a universal accelerator ranking. Its durable idea is to change the active mix of compute, memory and communication within one chip instead of changing the physical accelerator assigned to each phase. The evidence is strongest for low-latency serving on the three disclosed models. Production qualification, software maturity, rack power, failure behavior and broad workload coverage remain the tests that will decide whether the same advantage survives outside OpenAI’s laboratory.
The metric moves from chip throughput to completed requests
Peak FLOPS answers how much arithmetic an accelerator could issue under a favorable kernel. An interactive service asks a different question: how many complete requests can the system finish within a latency target and a power budget? Agentic products make that distinction sharper because one user action can contain many dependent model calls. Saving 20 ms on one step can matter more than increasing batched throughput if the next step cannot begin until the current response arrives.
OpenAI therefore frames Jalapeño around time to last token and useful work per unit of power. InferenceX sweeps operating points rather than reporting one maximum-throughput number, allowing each system to be compared along a latency-versus-throughput frontier[2]. This is a better contract than peak tensor performance for interactive serving. It exposes whether batching gains were purchased by making each user wait longer, and whether low token latency collapses the number of simultaneous requests the system can sustain.
The metric is still incomplete without its boundary. Package TDP is available for every compared accelerator, while measured wall power for every rack configuration is not. OpenAI rates Jalapeño at 700 W and reports sustained chip power at or below 550 W on the tested workloads. The public comparisons, however, divide useful throughput by the published 700 W, 1,200 W or 1,400 W package ratings for Jalapeño, GB200 and GB300, respectively[2]. Networking, hosts, memory outside the package, cooling and idle components can change the system result. Thus the benchmark establishes a chip-level operating advantage under a common normalization rule, not a complete cost-per-request result.
Three regimes make fixed specialization expensive
One inference request does not present a fixed hardware ratio. Prefill applies the prompt to the model and tends to consume matrix compute. A speculative draft model runs at a small batch and places latency on the critical path. Verification reads more model state, stresses HBM bandwidth and can produce bursty mixture-of-experts communication. Context length, acceptance rate, model shape and traffic mix continuously change the time spent in each regime.
A heterogeneous fleet can assign a specialized device to every phase, but the boundaries become part of the service. KV cache state must either move to the next device or remain remotely accessible. The first option consumes bandwidth and adds synchronization latency. The second keeps an accelerator near its state while another specialized engine does the work, leaving paid silicon and package infrastructure underused. Fixed ratios also become wrong when the model or request mix changes.
Jalapeño chooses a balanced ASIC and changes which resources are active. Compute-heavy phases can use more matrix execution; verification can draw on local memory bandwidth; communication-heavy work activates the relevant network path. Units that are not useful for the current phase can be gated within the package. This does not remove heterogeneity. It moves heterogeneity inside one scheduling and memory domain, where the runtime can alter the resource mix without copying the request’s long-lived state across a fleet boundary.
Locality begins inside the package
The physical package supplies the capacity behind that choice. The disclosed configuration combines one compute die with six HBM4 stacks, 216 GiB of capacity and 15.4 TB/s of aggregate HBM bandwidth. OpenAI lists up to 13.4 PFLOP/s for MXFP4 matrix operations at a 700 W package rating[4]. These are architectural specifications, not achieved application rates. Their purpose is to make enough compute and memory bandwidth available that software can schedule each phase without immediately crossing a package boundary.

Physical HBM stacks and logical memory slices should not be confused. Jalapeño divides execution into 64 core slices, and each slice receives a local view of HBM rather than contending through one undifferentiated global memory system[4]. The six physical stacks provide channels and capacity; the architecture distributes that resource across many logical compute neighborhoods. Software must know where tensors live, but a common operand can remain close to the cores that repeatedly consume it.
Two communication paths preserve that locality. A specialized collective network carries common high-bandwidth patterns with predictable register-to-register movement. A more flexible general NoC handles remote memory, uncommon transfers and access to the external network. Building one fabric for both roles would require enough buffering, routing and arbitration for the worst case. The separate paths let frequent collectives avoid the contention and latency tolerance designed into the general path.

One balanced ASIC is not one uniform datapath
Calling the device fungible does not mean every block executes every operation equally well. The spatial programming model exposes local tensors, explicit communication and predictable synchronization. A compiler and runtime must place work onto the 64 slices, choose the collective or general path, schedule pipelines and maintain tensor layouts that describe physical placement. Hardware simplicity is therefore purchased with a more explicit mapping problem.
OpenAI argues that this mapping problem is suitable for AI-assisted search. The team reports using internal models during implementation and bring-up, while later systems optimized kernels for models that were not in the original production plan. Three open-weight models reached high performance within two months, according to the company. Selected GPT-OSS attention and mixture-of-experts blocks ran 1.5 to 1.8 times faster than prior human-expert implementations[2]. The scope is important: these are selected kernel blocks, not a 1.5 to 1.8 times improvement for the complete model.
The software contract is consequential for adopters. A GPU absorbs irregularity through broad programmability, mature libraries and a large developer base. Jalapeño narrows the hardware around inference patterns and depends on mapping tools to recover flexibility. OpenAI controls the model, serving stack, compiler, chip and fleet, so it can pay that software cost where it produces an operating return. A merchant accelerator would need to expose the same benefit to customers that do not share its internal workload traces or optimization models.
Ethernet scale-up works because the hierarchy is explicit
The package is only the first locality boundary. OpenAI describes a 128-ASIC local domain with 600 GB/s of network bandwidth per processor and a 2,048-ASIC global domain with 200 GB/s per processor. The larger system uses a half-flattened two-level Clos topology around Broadcom Tomahawk 6 Ethernet switching, allocating the wider path to tensor-parallel traffic and the narrower path to expert-parallel communication[4]. At full scale, the disclosed totals reach 27 EFLOP/s of MXFP4 compute, 432 TiB of HBM4 capacity and 32 PB/s of aggregate memory bandwidth.
These totals are sums of endpoints, not the bandwidth available to one transfer. A tensor-parallel collective inside the local domain sees a different path from expert traffic that crosses the global tier. Oversubscription, routing, message size and simultaneous collectives determine how much of the nominal rate becomes useful. The phrase Ethernet scale-up is therefore insufficient by itself. The design decision is the hierarchy: keep latency-sensitive, high-volume exchange within 128 ASICs and use the broader 2,048-chip domain for traffic that tolerates a narrower per-device path.
Failure scope also follows that hierarchy. A 2,048-chip system needs rerouting, degraded-mode operation, job placement and repair procedures that do not assume every endpoint remains present. The Hot Chips material gives the topology and link rates but not production failure distributions, recovery times or service-level behavior. Operators should ask how the runtime remaps a model when a local-domain link, switch or accelerator disappears, and whether the remap preserves the KV locality that motivated the architecture.
The benchmark is strongest where latency is expensive
OpenAI tested GPT-OSS 120B against GB200 and DeepSeek R1 670B plus Kimi K2.5 1T against GB300. The disclosed InferenceX setting uses nominal 8k-token input and 1k-token output sequences with single-token prediction for the headline table. At peak mixed throughput, Jalapeño reports 85,448 versus 44,960 mixed tokens/s/kW on GPT-OSS, 19,641 versus 11,781 on DeepSeek, and 18,195 versus 11,862 on Kimi. Those pairs correspond to approximately 1.9, 1.7 and 1.5 times higher peak throughput per kilowatt, respectively[2].
End-to-end latency shows the more distinctive result. The reported pairs are 1.03 versus 1.80 seconds for GPT-OSS, 1.65 versus 5.99 seconds for DeepSeek and 1.56 versus 5.31 seconds for Kimi. Jalapeño also reports minimum time between tokens of 0.69, 1.43 and 1.44 ms, compared with 1.87, 5.90 and 5.48 ms for the respective NVIDIA systems. These numbers support the locality thesis because the advantage grows where waiting on memory and communication is visible to the user.

The largest multipliers require a narrower reading. At the previous system’s minimum time-between-token point, OpenAI reports 53.7 times, 104.3 times and 56.1 times more throughput per kilowatt for the three models. Those are comparisons at extreme low-latency coordinates where the baseline sustains little aggregate traffic, not average gains across the operating range. They show that Jalapeño extends the Pareto frontier. They do not mean a datacenter will finish every inference workload 104.3 times faster.
Speculative decoding also changes the comparison. The Hot Chips analysis notes that the disclosed Jalapeño headline uses single-token prediction while some NVIDIA operating points use multi-token prediction[4]. OpenAI separately reports a single-token Jalapeño comparison against multi-token GB300 on DeepSeek with about 1.5 times higher peak throughput per kilowatt and 2.2 times lower end-to-end latency. That narrower comparison is useful because it keeps the algorithmic advantage on the baseline side, but full reproducibility still needs software versions, batch policies, parallel layouts and complete system configurations.
What the numbers do not yet buy
First, package efficiency is not facility efficiency. Hosts, switches, optical modules, storage, power conversion and cooling remain outside the TDP-normalized denominator. A lower-power accelerator can also require more devices or a different network to host the same model. Matched rack input power and completed requests per rack-hour would reveal whether the chip result survives those components.
Second, the presented systems are not a neutral merchant benchmark. InferenceX is public, and the three model choices broaden the evidence beyond an OpenAI-only workload. However, OpenAI ran Jalapeño in its own laboratory and controls the compiler, kernels and operating points. The results should be reproduced by an independent operator on released hardware before they become a general procurement ranking. OpenAI also states that production qualification, software maturation and validation across more models are still underway, with initial internal deployment planned by the end of 2026[2].
Third, the architecture spends scarce HBM4 capacity to obtain locality. Six stacks and 216 GiB give one package unusual bandwidth and model capacity, but availability, yield and cost can constrain fleet scale. The 128- and 2,048-chip figures describe an intended system organization; they are not a published reliability study of a long-running production fleet. Supply, repair and degraded operation will decide how much of the theoretical domain can be offered continuously.
Lastly, Jalapeño is an inference ASIC. The public material does not establish training performance, general scientific-computing coverage or broad compatibility with GPU software. OpenAI explicitly expects to continue deploying accelerators from NVIDIA and other partners[2]. The relevant conclusion is specialization of a growing inference lane, not immediate replacement of every accelerator in the datacenter.
Nine-month tapeout shifts risk into continuous verification
OpenAI and Broadcom state that the design moved from initial work to tapeout in nine months[3]. AI models explored implementations, shortened measurement and verification loops, and optimized arithmetic structures. The Hot Chips presentation also describes substantial changes continuing close to RTL freeze. A short schedule is valuable only if the verification process preserves correctness while the design keeps moving.
The architectural style helps by making placement and synchronization explicit. Local tensors and predictable communication give both engineers and optimization agents a constrained search space. Yet the same explicitness creates a verification obligation across compiler, runtime and silicon. A mapping that is fast for one tensor shape can deadlock, overflow local storage or expose a long communication path for another. Model-specific kernel work did not disappear; OpenAI says each new family still needs optimization[2].
For other chip teams, the lesson is not that nine months is now a safe universal schedule. OpenAI combined a narrow first-generation target, Broadcom implementation experience, Celestica system integration and direct access to production workload traces. Organizations without those inputs may accelerate code generation while extending architecture exploration, signoff or software bring-up. The reusable principle is continuous convergence between workload measurement, architecture, RTL, verification and physical design, not the calendar number alone.
The acceptance test belongs at request and rack scale
A buyer or internal infrastructure team should qualify Jalapeño with measurements that preserve the request chain. Per-phase counters need to show prefill compute occupancy, draft-loop latency, verification bandwidth, collective utilization and KV-cache movement. End-to-end distributions should include p50, p95 and p99 time to last token under a stated arrival process, context-length distribution and speculative acceptance rate. Average token throughput cannot reveal a small group of users waiting behind a saturated communication path.
Energy should be recorded at the rack input. The report should separate accelerator packages, hosts, switching, optics and cooling, then divide the total by completed requests that met the latency objective. Local-domain and global-domain experiments should increase simultaneous tensor- and expert-parallel traffic until useful throughput stops scaling. Fault injection should remove one accelerator, link and switch at a time and measure state reconstruction, model remapping and latency recovery.
Software portability needs its own budget. Teams should record engineering time for a new model, kernel coverage before and after automated optimization, correctness escapes, compiler fallback paths and performance retained after a model revision. OpenAI’s selected 1.5 to 1.8 times kernel gains show that the toolchain can find valuable mappings. The procurement question is whether those gains arrive repeatedly without making every model update dependent on a small internal hardware team.
The durable insight is controlled movement
Jalapeño’s most important contribution is not a claim that one 700 W package defeats every GPU. It is a system argument that inference efficiency is lost at phase boundaries, where state moves, synchronization expands and the active hardware ratio no longer matches the request. The architecture responds at three levels: local HBM views inside 64 slices, separate fast and general communication paths on the die, and a 128-to-2,048-chip Ethernet hierarchy outside it.
That argument changes what should be measured. Peak FLOPS and HBM bandwidth remain necessary capacity indicators, but they do not show whether a request keeps its data close, whether low latency survives useful load, or whether an idle package still consumes the savings. Jalapeño’s first silicon and public model results make locality-first inference credible. Rack power, failure recovery, independent reproduction and the cost of continually mapping new models will determine whether it becomes economical infrastructure.
Source and copyright notice
This article is an independent editorial digest prepared by Silicon & Systems from OpenAI’s Hot Chips 2026 presentation, the official OpenAI performance report and launch announcement, and the cited session reporting. We rewrote the architecture, measurements and limitations in our own words. No source sentence, table, slide or figure is reproduced. The material package view is an original conceptual rendering rather than a product photograph or manufacturing drawing; all labels, topology and charts were created deterministically for this review. The Hot Chips presentation and OpenAI materials remain copyright their respective owners (2026).