OpenAI is working with Samsung Electronics on next-generation artificial-intelligence chips, according to comments from OpenAI Korea general manager Harrison Kim at a press conference in Seoul on Wednesday. Kim described joint production and research with Samsung as one of the areas where the two companies had made the most progress, but he did not disclose a chip name, design timeline, manufacturing plan, or commercial terms (Reuters).
The announcement matters because OpenAI is moving beyond buying compute and toward shaping more of the hardware stack that runs its products. For operators, this is not a reason to change vendors today. It is a signal that inference capacity, memory supply, and the cost per useful response are becoming strategic parts of the AI product roadmap.
What happened
OpenAI has already disclosed one custom chip effort. In June, the company introduced Jalapeno, an inference chip designed with Broadcom. OpenAI said at the time that Taiwan Semiconductor Manufacturing Company would manufacture that chip. Inference is the part of an AI service that processes a user request and produces an answer, so an inference-focused chip is aimed at serving workloads after a model has been trained rather than building the model from scratch (Reuters).
The Samsung relationship appears broader than a simple foundry order. Reuters reported that OpenAI is expanding cooperation with the South Korean technology company across semiconductors and enterprise AI services. Samsung and SK Hynix also signed letters of intent last year to supply memory chips for OpenAI’s Stargate data-center project. Kim said demand for memory will continue to grow as more advanced chips are needed for faster and more complex computing.
That distinction is important. A modern AI server is not just a processor. It also needs high-bandwidth memory, networking, power delivery, cooling, storage, and software that keeps those parts busy. A bottleneck in any one of those layers can raise the cost of serving a model or limit how quickly a provider can add capacity.
Why it matters for operators
Custom silicon can improve the economics of a large, predictable workload, but it takes time and money to design. The benefit grows when a company serves enough traffic to spread engineering and manufacturing costs across many requests. OpenAI’s scale makes that calculation more plausible than it is for a small AI application.
The Samsung work also highlights a second issue: memory is becoming a constraint alongside GPU and accelerator availability. If model providers compete for the same advanced memory packages, API capacity can be affected even when a company has secured accelerator designs. That can show up as higher prices, slower access to a new model, or limits on peak usage rather than as a headline about a chip shortage.
OpenAI’s enterprise footprint adds another reason to watch the development. Reuters said the number of ChatGPT Enterprise users at South Korean companies and institutions had risen about 28-fold by the end of August from a year earlier. That is a company-reported metric, not an independent audit, but it shows why OpenAI is investing in the supply chain behind business workloads rather than treating hardware as a background purchase.
The practical takeaway is that model capability and model economics are increasingly linked. A better chip does not automatically make an API cheaper, because providers still have to pay for memory, data centers, networking, staff, and capital. It can, however, give a provider more control over throughput and cost when the workload is large enough.
What operators should do
- Do not rebuild your stack around an unannounced chip. No product specification, launch date, or customer availability was disclosed. Keep current API contracts and fallbacks in place.
- Track cost per completed task, not only token price. Measure latency, retries, output quality, and the amount of human review required. Hardware gains matter only if the delivered workflow improves.
- Add a capacity-risk line to your vendor review. Ask providers how they handle peak demand, model migrations, rate limits, and regional outages. A second model endpoint can be more useful than a theoretical discount.
- Separate model choice from application design. Keep prompts, evaluation sets, routing, and logging portable so a future price or capacity change does not force a rushed rewrite.
- Watch memory and inference news together. A provider that controls more of the hardware stack may eventually offer better capacity planning, but the evidence will be lower latency, stable quotas, or lower total task cost—not the chip announcement alone.
For now, the story is about direction rather than a product launch. OpenAI is designing custom inference hardware with Broadcom and is deepening work with Samsung while demand for enterprise AI and memory continues to grow. That direction should make infrastructure a standing item in an operator’s quarterly AI review.
Related links
- Learn: /learn/ai-cost-projection-budgeting
- Learn: /learn/cheapest-ai-models-2026
- Tool: /tools/chatgpt
Primary source: Reuters’ report on OpenAI and Samsung’s chip work.