Samsung and SK hynix’s zHBM and HBF: Everything You Need to Know About the Next-Generation Technologies Tackling AI Memory Bottlenecks
\n
AI’s Real Limitation Isn’t the GPU—It’s Memory: The Tech Industry’s Next Bottleneck
Even if hundreds of thousands of GPUs are connected, an AI system will grind to a halt if it cannot deliver the data required for computation on time. GPUs sit idle waiting for data, power continues to be consumed, and expensive server resources fail to deliver their full performance. The fiercest competition in the tech industry is now shifting away from securing more compute units and toward building memory architectures that keep those compute units from going hungry.
Data Delivery Speed Matters More Than FLOPS
Large language models (LLMs) are not simply technologies that require enormous amounts of computation. They are memory-intensive workloads that repeatedly read billions to trillions of parameters, maintain a KV cache to process long contexts, and handle requests from multiple users simultaneously.
Throughout this process, GPUs constantly demand data. But if memory bandwidth is insufficient or storage is located too far away, the GPU’s computational performance cannot be fully utilized. This is commonly known as the memory wall.
Put simply, the performance of an AI server is determined by the following flow:
Storage → System Memory → HBM → GPU Compute
If any one stage is slow, the entire system becomes constrained by that speed. Particularly with massive models, no matter how high the GPU’s FLOPS may be, practical throughput declines if parameters and cache cannot be delivered from HBM quickly enough.
HBM’s Success Has Created a New Bottleneck
HBM (High Bandwidth Memory) emerged as an ultra-fast memory technology designed to solve this problem. By vertically stacking multiple DRAM dies and connecting them with TSVs (through-silicon vias), it provides a far wider data pathway than conventional DDR memory. Because it is positioned close to the GPU, it can also reduce latency.
HBM has been a core technology driving the performance of AI accelerators higher. But the pace at which AI models are growing is now putting pressure even on HBM’s capacity and supply.
- Model parameters are getting larger,
- Context windows are getting longer,
- More inference requests are arriving simultaneously, and
- The scale of KV caches and vector data is surging.
As a result, the existing architecture is finding it increasingly difficult to handle both the “fastest data” and the “largest volume of data.” HBM is fast but expensive and limited in capacity, while SSDs offer large capacity but lack the latency and bandwidth required for direct use by GPUs.
AI infrastructure now needs a new tier to bridge the gap between the two.
Why Samsung and SK hynix Are Focusing on zHBM and HBF
The next-generation AI memory architectures proposed by Samsung and SK hynix—zHBM (vertically stacked HBM) and HBF (High-Bandwidth Flash)—target precisely this gap.
zHBM takes the advantages of conventional HBM and pushes them even further. Its core objective is to secure greater capacity and bandwidth within a limited footprint through higher stacking density and packaging optimization. For AI accelerators, this opens the door to receiving more data, faster, within the same physical space.
HBF, on the other hand, reinterprets NAND flash as a high-bandwidth tier designed for AI. Rather than serving as a simple storage device like a conventional SSD, it aims to become a “warm memory” tier that rapidly supplies large datasets, model parameters, vector indexes, and other data from a level below HBM.
Structurally, it can be organized as follows:
| Memory Tier | Primary Role | Characteristics | |---|---|---| | HBM·zHBM | Active parameters, KV cache, data for immediate computation | Fastest, but subject to significant cost and capacity constraints | | HBF | Large-scale model data, vector indexes, rapid offloading | Slower than HBM, but expected to offer far greater capacity | | SSD·Object Storage | Long-term data storage, original training data, backups | Most affordable and high-capacity, but with high latency |
If this tiered architecture becomes a reality, AI systems will no longer need to keep all data in expensive HBM. Frequently accessed data can be placed in zHBM, data that will soon be needed in HBF, and long-term data in SSDs or object storage—enabling far more precise resource utilization.
The AI Race Is Shifting from “Bigger Chips” to “Smarter Memory”
This is not simply a competition between memory products. In the future, the performance of AI servers will be difficult to evaluate based solely on the number of GPUs they contain. Even with the same GPUs, actual throughput could vary significantly depending on which HBM is connected, how efficiently the flash tier is configured, and how effectively software controls data movement.
Fields such as LLMs handling long contexts, RAG-based search systems, large-scale vector databases, and AI-agent memory are particularly sensitive to memory architecture. As models grow larger, where data is stored and how quickly it can be retrieved will increasingly determine service quality and cost—more than raw computational power alone.
That is why zHBM and HBF are not merely “faster memory.” They represent a redesign of the infrastructure needed for AI to read longer contexts, serve more users simultaneously, and operate much larger models at a realistic cost.
The real battleground in the AI era is not the GPU itself, but the memory architecture that keeps the GPU from ever sitting idle.
zHBM from a Technology Perspective: Pushing the Limits of Vertical Stacking Once Again
What would happen if we could stack memory dies higher and more densely within a package of the same footprint? Instead of waiting for data to arrive from distant memory or storage, AI accelerators could receive data through a much wider pathway right next to the compute chip. zHBM is a next-generation AI memory architecture designed to answer precisely this question.
Conventional HBM also stacks multiple DRAM dies vertically and connects them with TSVs (through-silicon vias) to increase bandwidth. By placing the GPU and HBM close together through 2.5D packaging, it enables much faster and higher-volume data transfers than standard DDR memory. However, as AI models grow larger, it is becoming increasingly difficult to balance capacity, bandwidth, power consumption, and package area using existing approaches alone.
zHBM takes this balance to a new level. Its core idea is not simply to “stack more.” It is about reliably implementing taller stacking structures while optimizing the internal data paths, power delivery, and thermal design of the package as an integrated whole.
- Higher density: More memory dies can be stacked within the same package footprint, expanding memory capacity.
- Shorter data-movement distances: Reducing the physical distance between the compute chip and memory can improve bandwidth efficiency and lower the energy required to move data.
- Design tailored to AI workloads: Large volumes of data that move repeatedly—such as the parameters, KV cache, and activation data of large language models—can be supplied more efficiently.
- Stronger packaging competitiveness: The performance gap at the system level could widen further when zHBM is combined with interposers, hybrid bonding, and 2.5D or 3D packaging technologies.
In AI servers, bottlenecks can no longer be explained by GPU compute performance alone. Models with hundreds of billions or even trillions of parameters must fetch the data they need from memory before computation can begin. No matter how fast the compute units are, the GPU will remain idle if data delivery falls behind. This is commonly referred to as the memory wall.
zHBM matters because it offers a way to lower that wall. Higher bandwidth can support larger batch sizes and longer-context processing, while greater memory capacity reduces the need to move model parameters and caches out to external memory. As a result, zHBM could improve accelerator utilization during training, as well as the number of concurrent requests and response stability during inference.
Of course, vertical stacking is far from easy. The taller the stack, the more difficult it becomes to dissipate heat, while TSV connection reliability, power delivery, yield, and package costs also become more challenging. AI accelerators are especially demanding because they continuously consume large amounts of power. That means the success of zHBM will depend not only on peak bandwidth figures, but also on bandwidth per watt, thermal management, and high-volume manufacturing yields.
Ultimately, zHBM goes beyond simply placing memory next to the GPU. It brings memory itself to the center of AI accelerator design. In today’s technology race, vertical stacking is no longer just a packaging option—it is emerging as a critical server-performance variable that could make larger models and faster AI services possible.
The Second Memory Layer HBF Unlocks: A New Balance for AI Tech Servers
Storing all AI data in HBM is the fastest approach. However, HBM is expensive, and there are physical limits to how much capacity can be integrated into a GPU package. SSDs, by contrast, are far more affordable and offer much higher capacity, but they lack the latency and bandwidth needed to supply an LLM with data immediately.
The concept that has emerged to bridge this gap is HBF (High-Bandwidth Flash). HBF aims to become a “second memory layer” that leverages NAND flash’s high storage density while delivering far greater bandwidth from a location much closer to the AI server than a conventional SSD.
A New Layer Between HBM and SSD
The data path in a conventional AI server has been relatively simple:
- HBM: Active parameters, KV caches, and intermediate results that the GPU needs for immediate computation
- DDR memory: CPU-centric auxiliary tasks and data buffers
- SSD·Storage: Model checkpoints, training data, documents, and vector databases
The problem is that as large models grow, the proportion of data that can be kept entirely in HBM becomes smaller. This is especially true for long-context inference, large-scale RAG searches, agent memory, and vector index operations, where data that is “not in HBM right now but will soon be needed” is constantly being generated.
HBF is designed to handle precisely this data. It is not the topmost layer that responds instantly to every request like HBM, but it can serve as a warm data layer that supplies large volumes of data much faster than a conventional SSD.
If HBM handles the hottest data, HBF is the layer that rapidly prepares large volumes of data waiting for the next computation.
HBF’s Role in AI Workloads
Once HBF becomes practical, AI servers will no longer divide data simply into “memory or storage.” Data can be placed more precisely according to its importance and access frequency.
| Data Type | Suitable Location | Reason | |---|---|---| | Tensors currently being computed, active KV caches | HBM·zHBM | Require the lowest latency and highest bandwidth | | Frequently referenced model blocks, large vector indexes | HBF | Balances high capacity with fast data delivery | | Checkpoints, original datasets, long-term archival data | SSD·Object Storage | Cost efficiency and storage capacity take priority |
For example, RAG-based services search through countless documents and vectors every time a user submits a question. Keeping all this data in HBM would be fast, but costs would rise sharply. Reading it from an SSD every time, on the other hand, could increase response latency.
This is where HBF can serve as an intermediate layer. Frequently queried vector indexes or recently used document sets can be placed in HBF to accelerate search. HBM can focus on actual inference and computation, while HBF proactively supplies the data that will be needed.
Flash Evolves from ‘Storage’ into Quasi-Memory
The core significance of HBF lies in redefining flash—not as a simple storage device, but as quasi-memory that supports AI accelerators. While conventional NVMe SSDs have been designed primarily around capacity and cost, HBF is likely to place much greater emphasis on AI data movement and parallel access.
Hardware changes alone, however, will not be enough. AI frameworks and server software will also need to support functions such as:
- Prefetching data required for the next computation
- Intelligent offloading that moves data to HBF when HBM capacity is insufficient
- Automatic tier placement based on data importance
- Memory scheduling that reduces transfers between GPUs, accelerators, and storage
- Workload-specific policies that distinguish between model parameters, KV caches, and vector databases
Ultimately, HBF is not merely a problem involving a single memory chip. It is a technology architecture capable of reshaping the entire data flow of AI infrastructure.
AI Server Cost Structures Will Change Too
HBM is central to AI performance, but it is also one of the major factors driving up server prices. Keeping all data in HBM may be ideal from a performance standpoint, but it can be less cost-effective when operating services at scale.
If HBF takes its place as a high-speed, high-capacity layer beneath HBM, server operators will gain options such as:
- Operating larger models and more data with the same HBM capacity
- Partially easing the burden of increasing HBM capacity per GPU
- Improving response performance for large-scale RAG and vector search services
- Optimizing operating costs by keeping only frequently used data in HBM
- Expanding the number of concurrent requests handled by inference servers
In other words, HBF is not a technology that replaces HBM. It is closer to a technology that enables more efficient use of HBM. Going forward, AI servers are likely to evolve into multi-layer architectures that combine HBM’s speed, HBF’s capacity-to-performance advantages, and SSDs’ cost efficiency.
Tech: Memory Technology Is Changing the Size of AI Models—and How They Are Served
The race to build larger AI models is no longer a contest driven solely by algorithms and GPU computing performance. In real-world service environments, performance is determined by how quickly a model can read its parameters, how many user requests it can handle simultaneously, and how quickly it can retrieve the external knowledge it needs.
This is why zHBM and HBF are attracting attention—not merely as new memory products, but as technologies capable of reshaping the very design of AI servers.
zHBM Pushes Back the Limits of Context Length
For an LLM to read lengthy documents and maintain a conversation, it must continuously store and retrieve not only input tokens but also the results of previous computations—the KV cache (Key-Value Cache). The problem is that as the context grows longer, the KV cache expands rapidly, forcing the GPU to spend more time reading and writing memory than performing computations.
zHBM is a vertically stacked memory architecture designed to achieve higher density and bandwidth. When applied to AI accelerators, it could enable the following changes:
- Greater memory capacity within the same GPU footprint
- More model weights and KV cache retained in local memory
- Faster token generation resulting from reduced memory wait times
- Greater stability for long-document analysis, multi-turn conversations, and agent tasks
For example, services handling contexts of hundreds of thousands of tokens or more cannot be solved simply by making the model larger. They require a memory architecture capable of rapidly managing massive KV caches. zHBM could ease this bottleneck, providing the foundation for operating long-context AI with lower latency.
The Number of Concurrent Users Could Be Determined by Memory, Not the Number of GPUs
One of the most important metrics in AI inference services is throughput under concurrent workloads. When users flood a service, it must maintain the model state and KV cache for each request. If memory space runs short, the system must reduce the cache, place requests in a queue, or move data to slower system memory or storage.
That process inevitably leads to increased response latency.
By placing more high-speed memory close to the GPU or AI accelerator, a zHBM-based design could increase the number of active requests a single device can handle. For service providers, this could mean the following:
The scalability of AI services could shift from “How many GPUs have been purchased?” to “How much memory bandwidth and capacity has been secured per GPU?”
Especially for services such as real-time chatbots, code-generation tools, and automated customer support—which require both short response times and high concurrency—the differences in memory architecture could even reshape the cost structure.
HBF Changes How Data Is Supplied to RAG
RAG (Retrieval-Augmented Generation) is an approach in which a model searches documents, databases, and vector indexes for relevant information before generating an answer. In enterprise AI, RAG has effectively become core infrastructure because systems must answer based on the latest internal documents, policies, and product information.
However, RAG performance is not determined by the model alone. As the search space expands to billions of vectors and vast document collections, the speed at which data is retrieved from storage can dominate the total response time.
HBF targets this very problem. It aims to provide a high-bandwidth data-delivery layer better suited to AI workloads than conventional SSDs, while leveraging the large capacity and cost efficiency of NAND flash. If HBM handles the most frequently accessed data and the model states currently being processed, HBF could serve as an intermediate memory layer that delivers large vector indexes and long-term document repositories more rapidly.
If this architecture becomes a reality, RAG services could evolve in the following ways:
- Placing larger document repositories and vector indexes closer to the service
- Reducing the time required to move search results into GPU memory
- Handling search requests from multiple users at higher throughput
- Easing bottlenecks between knowledge retrieval and response generation
In other words, HBF is an attempt to elevate flash beyond a simple storage device and turn it into a near-memory layer that AI can reference in real time.
What Matters More Than Model Size Is Where the Data Sits
Future AI servers are likely to place data according to three questions:
Is it needed for computation right now?
Model weights and active KV caches are placed in the highest-tier layers, such as zHBM and HBM.Is it likely to be read again soon?
Frequently referenced vectors, checkpoints, and intermediate results can be stored in high-speed flash layers such as HBF.Is it intended for long-term storage?
Original documents, logs, and older data are sent to conventional SSDs or object storage.
By dividing memory tiers according to data temperature and access frequency, AI services can operate larger models and broader knowledge bases without simply adding more expensive HBM. That is why the tech industry is paying attention to zHBM and HBF. Ultimately, the next stage of AI competition is likely to be fought not only over the intelligence of models, but also over how quickly and economically that intelligence can retrieve and use the data it needs.
Tech: The Winner of the AI Memory War Will Be Determined by Ecosystems, Not Technology
The emergence of zHBM and HBF is clearly a signal that could reshape the AI infrastructure landscape. However, the mere announcement of a new memory architecture does not mean a memory revolution has been completed. The real competition begins when every factor—product specifications, advanced packaging yields, software support, supply-chain stability, and geopolitical variables—comes together.
The first thing to verify is real-world performance. How much higher bandwidth and capacity does zHBM provide compared with existing HBM3E and HBM4? And at what levels of power consumption and heat can that performance be sustained? In AI servers, simply posting a high GB/s figure is not enough. Actual throughput is determined by how efficiently the GPU or AI accelerator utilizes memory bandwidth and how reliably it can supply the KV cache and parameters of large-scale models.
HBF must be evaluated by the same standards. If flash-based tiers are not intended to completely replace HBM, the key question is how effectively they can fill the gap between HBM and SSDs. If data that requires large capacity but not consistently top-tier speed—such as large vector databases, RAG document repositories, and long-term agent memory—can be moved to HBF, the cost structure of AI systems could change dramatically. However, latency, data-movement policies, cache coherency, and failure-recovery methods must first be thoroughly validated.
Packaging Yield Is the Hidden Battleground
Vertically stacked memory delivers higher density, but it also increases manufacturing complexity. As the number of stacked dies grows and TSV connections and thermal management become more complicated, even a minor defect can reduce the yield of the entire package. In particular, memory for AI accelerators is tied to the GPU, interposer, substrate, and power-delivery structure as part of a single packaging ecosystem. Producing the memory well alone is not enough.
Ultimately, the areas in which Samsung and SK hynix must compete go beyond memory-cell performance. They must secure stable advanced-packaging production capabilities, validated customer certifications, and the ability to control quality variation during mass production. The tech industry has already seen many cases in which the market standard was set not by “the fastest product,” but by “the product that could be supplied in the greatest volume and with the highest reliability.”
Software Must Understand the Memory Hierarchy
The true value of zHBM and HBF will emerge when software can recognize and utilize them. Future AI frameworks will need to move beyond simply dividing data between GPU memory and SSDs and perform more sophisticated hierarchical placement, such as the following:
- zHBM/HBM: Parameters, activation values, and KV caches needed immediately for current computations
- HBF: Model weights, vector indexes, and intermediate data that are accessed frequently but are too large to fit entirely in HBM
- SSD·Object Storage: Long-term archival data, original documents, and training checkpoints
If this structure becomes a reality, software stacks such as PyTorch, JAX, and TensorRT will also need to redesign their approaches to memory offloading and prefetching, cache policies, and checkpointing. No matter how fast the hardware becomes, its performance advantage will disappear if data is not placed in the right tier at the moment it is needed.
Supply Chains and Geopolitics Matter as Much as Performance
The AI memory market is now a supply-chain competition as much as it is a technology competition. HBM and advanced packaging require sophisticated equipment, materials, and testing infrastructure. The more heavily supply chains are concentrated in particular countries or companies, the more likely production disruptions, export restrictions, and equipment controls are to immediately translate into shortages of AI servers.
As a result, the success of zHBM and HBF will also be tied to policy changes across major semiconductor regions, including the United States, South Korea, Taiwan, Japan, and China. For major cloud companies and GPU manufacturers, the considerations will inevitably extend beyond peak performance to include production capacity capable of supporting long-term contracts, regional supply stability, and regulatory risks.
In conclusion, the winner of the AI memory war is unlikely to be the company that unveils the flashiest technology name first. The companies that complete performance, yield, packaging, software, customer collaboration, and supply-chain stability as a single ecosystem will claim the standard for next-generation AI infrastructure. zHBM and HBF are the starting points of that competition—and the real war begins now.
Comments
Post a Comment