ixteen months — that is all it took for OpenAI from hiring its first hardware engineers in mid-2024 to getting finished silicon in hand. Aiming to break free from Nvidia's exorbitantly priced server accelerators, the company developed its own custom chip, dubbed Jalapeño, in partnership with Broadcom. Unveiled at the Hot Chips conference, the hardware was tested hands-on by SemiAnalysis analysts in the developers' lab using the InferenceX benchmark suite. The biggest surprise: the ChatGPT creator's very first silicon outperformed not only offerings from AMD and Google, but also Nvidia's flagship Blackwell architecture in token throughput per megawatt.
Why build custom silicon
The rationale is strictly practical. Standard Nvidia GPUs are general-purpose workhorses designed to handle everything from 3D gaming graphics to training massive neural networks from scratch. Jalapeño was architected as an application-specific integrated circuit (ASIC) solely for inference — generating finished answers for end users. According to SemiAnalysis, the chip incorporates high-speed HBM4 memory and avoids proprietary lock-in: it comfortably ran third-party open-weight models including DeepSeek R1, Kimi-K2.5, and GPT-OSS. To demonstrate its versatility, engineers even used Codex prompts to get the classic game Doom running on the silicon.

In hands-on testing, the performance gains are striking. Running DeepSeek R1 in single-user mode, Jalapeño delivers over 700 tokens per second, climbing to roughly 1,400 tokens per second per user on Kimi-K2.5 and GPT-OSS. Notably, the chip sustains these speeds on standard baseline algorithms without software shortcuts, speculative decoding, or complex compute-phase splitting.
Jalapeño outperforms competing chips in perf/W, delivering superior token throughput per megawatt across the entire system.
Meanwhile, GSM8k benchmark evaluations confirmed that Jalapeño matches Nvidia's flagship silicon in output precision and reasoning quality.
What changes on user screens
Translating engineering specs into everyday terms, this silicon rivalry yields two clear benefits for anyone using AI tools daily. First, it eliminates irritating latency. During peak hours, when millions of users simultaneously generate text or debug code, servers powered by dedicated ASICs avoid bottleneck slowdowns. That familiar lag where an answer painstakingly renders character by character stems directly from server bandwidth limits.
The second benefit is cost. Running massive data centers packed with Nvidia accelerators costs developers billions, and that monopoly premium ultimately gets passed down to consumers through rising subscription fees. An energy-efficient processor that consumes less electricity per generated word directly lowers infrastructure operating costs. SemiAnalysis analysts note one caveat: current figures come from OpenAI's test benches, and independent benchmarks handling long context windows are still pending.

OpenAI has demonstrated that Nvidia's market dominance is vulnerable. If expensive green-team hardware loses its status as the default data-center standard, daily interactions with ChatGPT will become noticeably snappier while subscription prices stabilize.
