주메뉴바로가기본문바로가기
비즈한국 비즈한국

"Has the 'GPU=AI' formula collapsed?" What is behind the rise of next-generation AI processors?

This article was automatically translated by AI. There may be errors compared to the original Korean article.  Read original in Korean →

[비즈한국] Nvidia, which has effectively dominated the AI hardware market for the past several years, recently signed a $20 billion (approx. 29 trillion KRW) deal with Groq, an inference-specialized chip startup. This is being received not just as a large transaction, but as a signal that the currents in the AI semiconductor industry are shifting. While the market views this as a 'big deal' essentially equivalent to a merger and acquisition, Nvidia emphasizes that it is a contract centered on technology licensing and securing key talent rather than a traditional M&A. Groq will also continue to operate as an independent entity. Nevertheless, the reason this contract is drawing significant attention is clear: it symbolically demonstrates that the central formula of the AI hardware market is moving away from 'GPU=AI'.

Specialized chips for different purposes are emerging, such as LPUs optimized for inference, TPUs tailored for big tech internal strategies, and NPUs aimed at mobile environments, leading to a restructuring of AI hardware into a multi-layered system. Photo=Generative AI
Specialized chips for different purposes are emerging, such as LPUs optimized for inference, TPUs tailored for big tech internal strategies, and NPUs aimed at mobile environments, leading to a restructuring of AI hardware into a multi-layered system. Photo=Generative AI

Why the Nvidia-Groq 'Collaboration' is Drawing Attention

The starting point of the AI race was training. Computational resources were needed to train larger models faster, and the GPU was the tool that best met this demand. Nvidia rose to become the biggest beneficiary of the AI era by leveraging this trend. However, as AI transitioned from the research stage to a full-fledged service phase, industry interest shifted to post-training stages. How fast, reliably, and cost-effectively pre-trained models can be used has emerged as a new competitive factor. In this process, issues such as power consumption, latency, and operational costs have come to the fore.

It is at this point that inference-specialized chips have begun to gain attention. The 'Language Processing Unit (LPU)' developed by Groq is an Application-Specific Integrated Circuit (ASIC) designed with language model inference as a prerequisite. Unlike GPUs, which prioritize versatility, it adopts a structure that pre-defines computation paths to reduce unnecessary branching. This is advantageous for reducing latency and increasing power efficiency. In environments such as chatbots or enterprise AI services where real-time response is crucial, these characteristics directly translate into cost competitiveness. This is why Nvidia has reached out to external technology rather than sticking solely to its own GPU designs. Analysts suggest this implies a recognition that the inference domain is no longer the exclusive stage of the GPU.

However, this choice does not mean a reduction in the role of the GPU. GPUs remain a key pillar for large-scale training and general-purpose AI acceleration. Nvidia's strategy is closer to supplementing areas where GPUs are relatively inefficient rather than replacing them. Therefore, the collaboration with Groq is interpreted as a strategic move to add inference-specialized technology onto the GPU-centric structure.

What is the difference between LPU, TPU, and NPU?

This trend is not a dilemma faced solely by Nvidia. Google has long been developing its own accelerator, the 'Tensor Processing Unit (TPU)'. The TPU is an ASIC designed for Google's internal workloads, demonstrating high efficiency in core services such as search, translation, and recommendations. This is a strategic choice to reduce dependence on external GPUs and keep the cost structure of AI services under its own control. Rather than targeting the general market, the TPU has evolved with the clear goal of optimization within the Google ecosystem.

The 'Neural Processing Unit (NPU)' showcases its presence on another stage: in the palm of an individual's hand, rather than in a data center. Apple chose to embed a neural engine in the iPhone to process AI calculations on the device itself. This decision considers battery efficiency, personal privacy, and user experience simultaneously. Smartphone manufacturers, including Samsung Electronics005930, are also enhancing NPU performance to implement features like real-time translation, photo editing, and voice recognition on-device. While the NPU does not replace data-center-grade inference, it is expected to play a crucial role in spreading AI into daily life.

As such, the LPU, TPU, and NPU each started from different perspectives. Realistic demands such as data center cost burdens, big tech supply chain strategies, and the power constraints of mobile devices have differentiated the forms of these chips. As a result, AI hardware is rapidly changing from a structure where a single chip handles all roles to one where chips optimized for specific purposes coexist.

Nvidia also appears not to deny these changes. Rather, through technologies like NVLink Fusion, it is adopting a strategy to connect external ASICs to its own systems, maintaining its role as the central platform even in environments where various chips are mixed. The calculation is to maintain its influence through connection and coordination of the ecosystem rather than through hardware monopolization.

Therefore, the cooperation between Nvidia and Groq is interpreted not as a declaration of the end of the GPU empire, but as a symbolic scene showing that the AI hardware industry has entered a stage of maturity. As the industry moves from competition at the center to a phase that also considers inference and economic feasibility, a single answer is disappearing. While GPUs still hold a key position in the AI industry, they are no longer the sole performer. The AI hardware war is now expected to intensify within a multi-layered structure where each has a differentiated role.

This article was automatically translated by AI. There may be errors compared to the original Korean article.
봉성창 기자

기업이 말하는 성장의 언어와 그 뒤에 놓인 현실의 간극을 집요하게 들여다보고 있습니다. 산업 현장의 변화는 숫자만으로 설명되지 않습니다. 투자와 고용, 기술과 규제, 혁신과 책임이 충돌하는 지점에서 비로소 기업의 진짜 얼굴이 드러납니다. 그 균열을 놓치지 않고, 복잡한 산업 이슈를 독자가 납득할 수 있는 맥락으로 풀어내는 일을 해왔습니다. 빠르게 흘러가는 시장의 소음 속에서도 끝까지 물어야 할 질문을 붙들고, 비즈한국 산업팀만의 날카롭고 균형 잡힌 시선으로 산업의 현재와 다음을 기록하겠습니다.

bong@bizhankook.com
저작권자 ⓒ 비즈한국 무단전재 및 재배포 금지