THE rapid expansion of artificial intelligence (AI) and generative AI (GenAI) is creating a storage burden that extends well beyond individual computing workloads, with organisations increasingly forced to retain, retrieve and manage growing volumes of data.
A global study by International Data Corporation (IDC), sponsored by Western Digital (WD), found that 94.7 per cent of surveyed organisations had increased the amount of data they stored over the past 12 months because of AI and GenAI adoption.
The IDC White Paper entitled ‘Built for Scale: The Enduring Role of HDDs in the AI Era’, found that 61 per cent of organisations experienced data growth of at least 25 per cent over the past year due to AI, while 74 per cent expected their data volumes to grow by 25 per cent or more over the next three years.
Unlike computing workloads that end when processing cycles are completed, much of the data generated by AI persists and accumulates, with organisations increasingly finding new uses for information that was previously archived.
WD chief executive officer Irving Tan said the focus of AI infrastructure discussions had largely been on computing power, even as the growing volume and longevity of data presented another major infrastructure challenge.
"For the last few years, the AI infrastructure conversation has centered on compute. But AI runs on data," said Irving Tan, CEO, WD. "Organizations are generating more data, keeping it longer, and finding new ways to create value from the information they already have. Compute requirements will evolve over time, but the need to store, manage and access data at scale is only growing. That foundation will play a critical role in determining how far AI can go."
The study found that 85.4 per cent of organisations had recorded growth in their data-lake volumes over the past year, with 59.4 per cent identifying AI-generated data — including synthetic data, inference outputs and model logs — as the leading driver.
The value attached to stored information is also rising. Nearly 95 per cent of respondents said the value of their organisations' data had increased as a result of AI and GenAI adoption.
As a result, 74.3 per cent said AI and GenAI had led them to retain data for longer, while 75.9 per cent reported bringing increasing volumes of archived cold-tier data back online to support AI workloads.
The trend is also raising expectations for archive accessibility, with 96 per cent of respondents anticipating a need for faster retrieval of archived data to support AI inference and retrieval-augmented generation (RAG) applications.
IDC found that 74.6 per cent of enterprise data among the surveyed organisations was held in warm, cool and cold storage tiers, while more than 60 per cent of data-lake volume comprised cold or infrequently accessed data.
Cost has emerged as another key consideration, with 98.2 per cent of respondents saying total cost of ownership per terabyte was important or very important when making storage decisions.
The findings suggest that the traditional divide between active and archived data is becoming less distinct as historical information is increasingly brought back into the AI data lifecycle.
IDC said organisations therefore needed to design storage architectures around the full AI data lifecycle, balancing performance, capacity, accessibility and long-term costs according to workload requirements.
The white paper was based on a quantitative survey of 763 IT and business decision-makers at manager level and above across seven countries, all with direct responsibility for AI infrastructure or data-storage decisions. IDC also conducted in-depth interviews with three senior storage industry leaders. - September 11, 2026