OpenAI’s first custom AI processor is beginning to produce measurable results, signaling that the competition for artificial intelligence may increasingly be fought across the entire technology stack — from models and software to networking, data centers and the chips underneath them.

WHAT’S HAPPENING

OpenAI has released the first performance results for Jalapeño, the custom AI inference processor it developed with Broadcom.

The chip was first unveiled in June. Now OpenAI says testing shows Jalapeño can perform substantially more AI work while using less power and returning answers faster than the commercial systems it tested against.

Across three large public AI models — GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T — OpenAI reported that Jalapeño delivered roughly 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems.

Those results were produced using InferenceX, a public inference benchmark from SemiAnalysis, although the testing and results were reported by OpenAI and should therefore be viewed as company performance claims until the hardware receives broader independent testing.

Jalapeño is specifically designed for inference — the enormous amount of computing required after an AI model has already been trained and begins answering questions, generating content or performing agentic tasks for users.

That distinction matters.

Training creates the intelligence.

Inference is what people consume every time they use it.

And as AI usage grows, inference is becoming one of the industry’s biggest cost, energy and infrastructure problems.

WHY IT MATTERS

OpenAI’s move into custom chips is not simply about making ChatGPT respond a little faster.

It is about gaining greater control over the economics of delivering intelligence.

Every AI response consumes computing capacity.

Every additional user creates demand.

Every agent that completes multiple steps can generate far more inference work than a simple chatbot request.

And as models become more capable, the infrastructure required to operate them becomes increasingly important.

That creates a strategic problem for AI companies that depend heavily on somebody else’s hardware.

OpenAI remains a major Nvidia customer and says it will continue deploying accelerators from Nvidia and other partners. But Jalapeño gives the company another option — one designed specifically around the workloads OpenAI expects its own systems to perform.

The bigger picture is emerging.

OpenAI increasingly operates across:

AI models

consumer products

enterprise software

developer platforms

AI agents

data-center infrastructure

networking and memory architecture

and now

custom processors.

OpenAI CFO Sarah Friar described the company’s compute strategy this week as an integrated system stretching from chips and data centers through frontier models, developer platforms and products.

That means the AI race may no longer be simply:

Who has the best model?

It may increasingly become:

Who can design the entire system around the model?

WHO BENEFITS

OpenAI stands to gain greater control over performance, cost and capacity.

If its own processors can perform more useful AI work from the same amount of electricity and infrastructure, OpenAI could potentially serve more users without increasing costs at the same rate.

That becomes particularly important as AI agents grow.

Agents may perform dozens or hundreds of model interactions while completing a single assignment. Small delays and costs that barely matter during one chatbot response can multiply dramatically across an autonomous workflow.

OpenAI says Jalapeño’s architecture was built partly around these increasingly interactive workloads.

Customers and developers could eventually benefit if greater efficiency translates into faster AI systems, more available computing capacity or lower costs.

Broadcom also benefits from becoming an important partner in OpenAI’s attempt to develop an alternative source of specialized AI computing hardware. OpenAI and Broadcom jointly unveiled Jalapeño in June as the first processor in what they described as a multigeneration platform.

And the broader AI industry could benefit from increased competition in accelerator design.

For years, Nvidia’s GPUs have been the central computing engine behind the generative AI boom.

More viable architectures mean more choices.

WHO LOSES

Nvidia does not suddenly lose because OpenAI built a chip.

That would be an exaggeration.

OpenAI still expects to use large quantities of Nvidia hardware, and Jalapeño will enter a market where Nvidia, Google and other semiconductor companies continue advancing quickly.

There is also a timing issue.

Jalapeño’s reported results compare favorably against today’s commercially available systems. But TechCrunch notes that the competitive landscape may have advanced again by the time Jalapeño reaches large-scale deployment.

The more important long-term risk to Nvidia is different.

Its largest customers increasingly have an incentive to design processors optimized around their own workloads.

Google already has TPUs.

Amazon has custom AI silicon.

Microsoft has developed its own accelerators.

And OpenAI now has Jalapeño.

None of those developments eliminates Nvidia’s position.

But together they show that the companies spending the most money on AI infrastructure do not necessarily want to remain permanently dependent on a single supplier.

There is also pressure on smaller AI companies.

Building a frontier model is already extraordinarily expensive.

If the largest AI companies increasingly gain advantages by co-designing their own models, software, networks and silicon, competing at the frontier could require capabilities far beyond simply developing better algorithms.

WHAT HAPPENS NEXT

OpenAI plans to begin deploying Jalapeño inside its own computing infrastructure before the end of 2026.

And this is not intended to be a one-off experiment.

OpenAI says Jalapeño Gen 2 is already deep in development and Gen 3 is beginning to take shape.

That may be the most important detail in the entire announcement.

OpenAI is not merely testing whether it can design one useful chip.

It is building a processor roadmap.

The company also revealed another unusual part of the development process: OpenAI models themselves helped engineers design and optimize Jalapeño.

For selected parts of GPT-OSS workloads, OpenAI says AI-generated implementations eventually ran 1.5 to 1.8 times faster than existing versions written by human experts. Those figures apply only to selected components rather than the entire model, but they point toward a potentially powerful feedback loop.

AI helps engineers design better chips.

Those chips run AI more efficiently.

More efficient AI helps design the next generation of hardware.

And the cycle repeats.

The first generation of generative AI competition was largely judged by model capability.

Who had the smartest model?

Who scored highest on the benchmark?

Who could generate the best answer?

The next phase may be harder to see from the outside.

It could increasingly depend on who controls the infrastructure required to deliver that intelligence at massive scale.

Jalapeño is only one chip.

But it represents something much larger.

OpenAI is no longer competing only to build the intelligence. It is increasingly trying to build the machine that delivers it.

Stay Sharp

Subscribe to follow the Trend newsletter and more.

Have a tip or idea?

Pass along insights or story ideas on AI, startups, and business. Focused on signal over noise, impact over headlines. Facts. Trends. Consequences. Always.

Support Independent AI Journalism

Buy Grey Ghost a Coffee

Related Deep Signals