[비즈한국] Nvidia, which has effectively dominated the AI hardware market for the past several years, recently signed a $20 billion (approx. 29 trillion KRW) deal with Groq, an inference-specialized chip startup. This is being received not just as a large transaction, but as a signal that the currents in the AI semiconductor industry are shifting. While the market views this as a 'big deal' essentially equivalent to a merger and acquisition, Nvidia emphasizes that it is a contract centered on technology licensing and securing key talent rather than a traditional M&A. Groq will also continue to operate as an independent entity. Nevertheless, the reason this contract is drawing significant attention is clear: it symbolically demonstrates that the central formula of the AI hardware market is moving away from 'GPU=AI'.

Why the Nvidia-Groq 'Collaboration' is Drawing Attention
The starting point of the AI race was training. Computational resources were needed to train larger models faster, and the GPU was the tool that best met this demand. Nvidia rose to become the biggest beneficiary of the AI era by leveraging this trend. However, as AI transitioned from the research stage to a full-fledged service phase, industry interest shifted to post-training stages. How fast, reliably, and cost-effectively pre-trained models can be used has emerged as a new competitive factor. In this process, issues such as power consumption, latency, and operational costs have come to the fore.
It is at this point that inference-specialized chips have begun to gain attention. The 'Language Processing Unit (LPU)' developed by Groq is an Application-Specific Integrated Circuit (ASIC) designed with language model inference as a prerequisite. Unlike GPUs, which prioritize versatility, it adopts a structure that pre-defines computation paths to reduce unnecessary branching. This is advantageous for reducing latency and increasing power efficiency. In environments such as chatbots or enterprise AI services where real-time response is crucial, these characteristics directly translate into cost competitiveness. This is why Nvidia has reached out to external technology rather than sticking solely to its own GPU designs. Analysts suggest this implies a recognition that the inference domain is no longer the exclusive stage of the GPU.
However, this choice does not mean a reduction in the role of the GPU. GPUs remain a key pillar for large-scale training and general-purpose AI acceleration. Nvidia's strategy is closer to supplementing areas where GPUs are relatively inefficient rather than replacing them. Therefore, the collaboration with Groq is interpreted as a strategic move to add inference-specialized technology onto the GPU-centric structure.
What is the difference between LPU, TPU, and NPU?
This trend is not a dilemma faced solely by Nvidia. Google has long been developing its own accelerator, the 'Tensor Processing Unit (TPU)'. The TPU is an ASIC designed for Google's internal workloads, demonstrating high efficiency in core services such as search, translation, and recommendations. This is a strategic choice to reduce dependence on external GPUs and keep the cost structure of AI services under its own control. Rather than targeting the general market, the TPU has evolved with the clear goal of optimization within the Google ecosystem.
The 'Neural Processing Unit (NPU)' showcases its presence on another stage: in the palm of an individual's hand, rather than in a data center. Apple chose to embed a neural engine in the iPhone to process AI calculations on the device itself. This decision considers battery efficiency, personal privacy, and user experience simultaneously. Smartphone manufacturers, including Samsung Electronics005930, are also enhancing NPU performance to implement features like real-time translation, photo editing, and voice recognition on-device. While the NPU does not replace data-center-grade inference, it is expected to play a crucial role in spreading AI into daily life.
As such, the LPU, TPU, and NPU each started from different perspectives. Realistic demands such as data center cost burdens, big tech supply chain strategies, and the power constraints of mobile devices have differentiated the forms of these chips. As a result, AI hardware is rapidly changing from a structure where a single chip handles all roles to one where chips optimized for specific purposes coexist.
Nvidia also appears not to deny these changes. Rather, through technologies like NVLink Fusion, it is adopting a strategy to connect external ASICs to its own systems, maintaining its role as the central platform even in environments where various chips are mixed. The calculation is to maintain its influence through connection and coordination of the ecosystem rather than through hardware monopolization.
Therefore, the cooperation between Nvidia and Groq is interpreted not as a declaration of the end of the GPU empire, but as a symbolic scene showing that the AI hardware industry has entered a stage of maturity. As the industry moves from competition at the center to a phase that also considers inference and economic feasibility, a single answer is disappearing. While GPUs still hold a key position in the AI industry, they are no longer the sole performer. The AI hardware war is now expected to intensify within a multi-layered structure where each has a differentiated role.