Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Storage
    Tape
    VAST Data
    NetApp
    Pure Storage
    Dell

    The 120PB Storage Question That Makes Open Source Look Expensive

    January 8, 2026
    15 min read

    Some storage questions make everyone in the room sit up straighter, and one hundred and twenty petabytes for HPC and AI in support of quantitative research is one of them. Nobody is asking which NAS to buy. It isn't a homelab fantasy with too many hard drives and not enough USB ports either, though people made those jokes, because storage people cope with terror through sarcasm. This is a program with a physical footprint, a support contract, a procurement war and roughly £30 million behind a three-year bet. If it goes wrong, it won't be a small migration. It'll be a career-defining incident with racks.

    The core question sounds simple. Go with DDN because an NVIDIA contact recommended it as cost-effective at this scale, or try to save money with open-source Lustre and maybe DeepSeek's 3FS, which colleagues in Hong Kong and Germany had recommended for AI workloads? On paper that's reasonable. At 120PB, though, "reasonable" gets eaten alive by operations. The choice underneath is whether the organization wants to buy a supported outcome, build a storage engineering team, or accidentally pretend those are the same thing.

    At 120PB, the organization matters more than the storage platform

    The most important number in the whole discussion is three years, even more than 120PB. That timeframe changes the texture of the decision. Anyone can draw a grand architecture on a whiteboard. Fewer teams can keep it healthy through firmware, failures, capacity expansion, performance tuning, metadata pressure, user behavior, network changes, security reviews, audits, vendor escalations, and the slow grind of researchers who only care that their jobs finish faster. At this scale, storage stops being a box and becomes an operating model.

    Open-source Lustre can be very powerful, and it runs in serious HPC environments for a reason. But "save a bit with OS Lustre" is one of those phrases that needs a warning label. Saving on licensing or vendor packaging can shift cost into people, process, testing and risk. That can be a great trade if the organization has the team for it, and a disaster if leadership thinks "open source" means "cheap enterprise storage without enterprise staffing."

    One commenter boiled the cynical view of scale-out file systems down to a joke: it's just taping together NAS servers. That's unfair to real distributed storage engineering, but the joke lands because bad scale-out designs often feel exactly like that, with a bunch of nodes, a brave diagram, and a prayer that the metadata layer doesn't turn into a haunted house.

    Lustre works. The uncomfortable part is whether this specific organization is ready to own Lustre at 120PB with AI and quantitative research users breathing down its neck.

    DDN is the obvious HPC answer, which makes the horror stories hit harder

    DDN's name naturally comes up in this kind of conversation. HPC storage, AI training pipelines, large scientific workloads and big parallel file systems are the world where DDN has long been part of the vocabulary, so an NVIDIA contact recommending it is not surprising. At that scale, a packaged, supported platform that understands GPU-fed workloads and keeps storage from becoming the bottleneck has a clean argument: spend enough on storage to make the compute useful, but not so much that the budget gets robbed from processing and networking.

    That's the optimistic case. Then the comments kicked the door open.

    One person described DDN as horribly unreliable and told a grim story about a professional services person being onsite for days after an outage, only to be laid off mid-repair. Another said every interaction they'd had with DDN was catastrophic, including business-critical SLA breaches. An anecdote like that is no benchmark or lab result, but storage buyers ignore such stories at their peril, because support experience is part of the product.

    At 120PB, nobody buys pure technology. They buy escalation paths, spare-part logistics and field expertise, and the confidence that when something weird happens on a Friday night, the vendor does not become another system that needs debugging.

    To be fair, every major storage vendor has horror stories (NetApp, Dell, IBM, Pure and VAST included), and open-source deployments definitely have them too. What hurt DDN in this thread went beyond technical skepticism to damaged trust. If your shortlist starts with "NVIDIA recommended them" and the first replies include "we ripped it out after contract expiry," the evaluation has to get much sharper.

    NetApp walked into the debate wearing a crown and a question mark

    An interesting side argument formed around NetApp. Some people pushed it as the safer enterprise choice. One commenter said "NetApp is the way" after agreeing with the DDN criticism. Another defended NetApp hard, arguing that ONTAP has decades of development behind it and that premium enterprise storage can be worth paying for when security, governance and research data matter. They framed the alternative as buying cheaper systems that need constant tinkering and bring "patch Tuesdays" into the storage world.

    That's the classic NetApp argument. Maturity, governance, security and operational polish all count, and not everything should be a race to the bottom when the data is valuable.

    The counterargument was sharp too. Someone questioned whether NetApp really fits 120PB HPC/AI scale, especially around NFS 4.2, pNFS, FlexFiles and metadata services. Another flatly said, "Not with 120PB." Others debated whether NetApp had true scale-out NAS, whether AFX changed the picture, and whether a new product can be called mature just because the company behind it is mature.

    That last point matters, because vendor maturity and product maturity are different things. A company can have 30 years of engineering history and still ship a new architecture that deserves careful proof, and a product can inherit ideas from mature systems and still behave differently at scale. Buyers love brand comfort, but physics and metadata do not care about it.

    NetApp may be a serious option for enterprise data management, and even the right option for certain parts of a 120PB estate. For HPC/AI scratch, training data and parallel access at this scale, though, "NetApp is enterprise" is not enough. The proof has to be workload-specific.

    The startup problem is really a trust problem

    The thread also had a strong anti-startup tone. Some commenters treated VAST and Hammerspace with suspicion. One said they would trust NetApp over a startup running on white-box systems any day, and another predicted Hammerspace would be gone or acquired within five years. The comments may sound dismissive, but they point to a real procurement fear: what happens if the vendor story changes before the data lifecycle ends?

    At 120PB, vendor survivability matters. Migration is hard and exit costs are painful. Data gravity stops being a metaphor and becomes a physical and economic force. If a platform turns out to be strategically wrong, moving away from it can take years, temporary duplicate capacity, network planning, downtime windows, application changes, and a level of project discipline that makes everyone miserable.

    Dismissing newer platforms simply because they are newer can be lazy, though. Some startup-era architectures exist because older systems were not built for AI-era access patterns. VAST gets attention in big-data and AI conversations because it tries to collapse some traditional storage tradeoffs around performance, capacity and namespace design. Hammerspace gets attention because global data orchestration and metadata-driven access are real problems. Newer vendors aren't automatically reckless; they still have to prove support, roadmap, economics and failure behavior at the required scale.

    For a 120PB research environment, the safest answer may be neither the oldest vendor nor the newest architecture. It may be the one that can survive the organization's actual workload and still answer the phone intelligently when something breaks.

    DeepSeek's 3FS is exciting, but excitement is not a storage strategy

    The mention of DeepSeek's 3FS gives the debate a newer AI flavor. AI teams love fresh infrastructure ideas because model training and data pipelines expose pain fast. If colleagues in Hong Kong and Germany recommend 3FS, it is worth investigating. At 120PB, though, "colleagues recommended it" should start a lab evaluation rather than end a procurement process.

    Emerging file systems can look incredible in the environment they were built for. Then they meet a different organization's security model, user base, scheduler, network, compliance process, support expectations, upgrade cadence and failure modes. Suddenly the elegant architecture needs documentation, tooling, backup integration, observability, quotas, lifecycle controls, recovery procedures, and people who understand it deeply enough to fix it when the original authors are asleep or unreachable.

    3FS may still earn a place, but production research data shouldn't be treated like a science project unless the organization explicitly wants to become part of that science project.

    Open-source and emerging systems work best when the team has enough internal engineering strength to be a real participant instead of a consumer. That means reading code, understanding failure modes, contributing fixes, building automation, designing monitoring, and accepting that support may not look like a traditional vendor escalation. If the organization wants that level of control, great. If it wants an appliance-like experience at 120PB, reality will be less kind, and the cheapest software can become the most expensive system if the team has to learn it during an outage.

    The jokes were absurd because the scale is absurd

    The comment thread also did what technical forums always do when a number gets ridiculous and turned into comedy. Someone asked whether 120PB would be good for an adult collection. Someone joked about fitting many Linux ISOs. Another imagined thousands of USB-to-SATA adapters and the challenge of finding 4,000 USB ports. It's silly, but jokes are how people mentally process a storage request big enough to feel unreal.

    Underneath them is a useful point: 120PB means power, cooling, floor space, networking, rebuild time, failure rates, spares, firmware domains, data protection strategy, namespace design, client behavior, monitoring and budget politics. At this size even tiny percentages become huge. One percent of 120PB is 1.2PB. A migration mistake stops being a folder problem, a performance bottleneck stops being one angry user, and a bad procurement assumption can waste millions.

    That is why consumer-style thinking collapses immediately. You cannot homelab your way into this class of system, or "just add drives" without understanding rebuild math, rack density, power, network oversubscription and operational staffing. You also cannot choose based on a single vendor recommendation, a single horror story, or a single favorite open-source project. The scale itself is the enemy, and every design choice gets heavier.

    The real battle is support versus control

    The most useful way to frame the decision is support versus control, rather than DDN versus Lustre versus 3FS versus NetApp versus VAST.

    A vendor platform gives you a throat to choke, a tested configuration, support contracts, reference architectures, and usually some lifecycle discipline. It also gives you vendor lock-in, pricing pressure, roadmap dependency, and the chance that support quality disappoints exactly when you need it most.

    An open-source or self-built approach gives you control, transparency, flexibility, and potentially better economics at hardware scale. It also makes your own team the escalation path. You own the integration, the weird bugs, the performance tuning and the consequences of under-staffing.

    Neither model is automatically smarter, but confusing them is deadly. If the organization has elite HPC storage engineers, strong Linux and parallel filesystem expertise, proper test environments, and the appetite to operate like a storage vendor internally, open-source Lustre or even emerging options may be viable. If it wants to focus on quantitative research and AI instead of becoming a storage product company, buying a supported platform makes more sense, even at a premium.

    The expensive solution may be cheaper if it keeps researchers productive, and the cheap solution may be better if the team can actually own it. The disaster is buying cheap while staffing as if expensive support exists.

    The proof of concept needs to be cruel

    At 120PB, a normal proof of concept is not enough. Vendors are good at demos and file systems are good at happy paths. Research workloads are anything but happy paths: metadata storms, huge sequential reads, random access patterns, checkpoint bursts, small-file misery, model training pipelines, scratch cleanup disasters, and users who will absolutely find the weirdest possible way to abuse the namespace.

    So the evaluation needs to be mean. Test the real workload, GPU starvation, metadata-heavy directories and checkpoint storms. Test mixed read/write pressure, client failures, rack loss and network congestion, and run rebuilds while users keep running jobs. Test software upgrades, quota behavior, snapshots or data protection if required, audit and governance, and restore. Then test what happens when the vendor's first-line support gives the wrong answer.

    Ask every vendor for three real references at similar scale and similar workload, since a slide of logos proves little. Ask about failed deployments and what customers hate. Ask what happens when capacity doubles, and how pricing changes when the system needs more metadata performance instead of raw capacity. Ask what data mobility looks like if the platform has to be replaced, and how long a full migration would really take. A serious vendor should survive hard questions, while a weak one will retreat into architecture diagrams.

    The budget is large, but not infinite

    Thirty million pounds over three years sounds huge until it meets 120PB of HPC/AI storage, networking, support, spares, power, facilities and staffing. The storage purchase is only one slice. If the NVIDIA contact framed DDN as cost-effective because it frees more budget for processing and networks, that is a valid systems-level concern. AI infrastructure is a balance. Overspend on storage and the GPUs wait in a smaller cluster; underspend and the GPUs starve. Either way, expensive silicon sits around judging everyone.

    This is why cost per petabyte alone is too crude, and cost per useful research outcome is the better measure. If a more expensive storage platform keeps accelerators fed, reduces outages, simplifies operations and avoids hiring a small army of specialists, it may be cheaper in practice. If a cheaper open-source platform performs well and the team can operate it confidently, the vendor premium may be wasted. The only wrong move is optimizing purchase price while ignoring operational cost.

    At 120PB, every inefficiency becomes a line item, every outage gets expensive, every migration becomes political, and every under-designed network link becomes a bottleneck with a name and a meeting attached. The budget needs to buy capacity, yes, but even more it needs to buy confidence.

    There may not be one platform to rule it all

    The thread naturally treats "best storage" as a single answer, but the real architecture may be layered. HPC/AI environments often have several storage personalities: hot training data, scratch space, checkpoint storage, long-term research datasets, governance-controlled enterprise data, archive, backup and replication. One product may not suit all of them.

    DDN or Lustre-like systems may make sense for performance-heavy parallel workloads. NetApp may make sense for governed enterprise NAS or research data with heavy security and policy needs. Object storage may fit some pipelines, and tape or cold archive may still matter for long-term retention. VAST may fit certain high-performance unstructured use cases, and 3FS may deserve a lab if AI-specific workloads match its strengths.

    Forcing every workload into one platform can simplify procurement while complicating life. Splitting platforms can improve fit while creating data movement and management pain. There is no free lunch, just different menus of regret.

    Data classification decides a lot here. What must be fast, governed, cheap, retained or shared globally? What can be regenerated, and what is irreplaceable? What access patterns exist today, and which ones are likely to appear once users discover the system is faster? Storage architecture should follow data behavior instead of vendor slogans.

    The sane answer is uncomfortable: buy expertise before buying petabytes

    The most responsible recommendation for a 120PB HPC/AI project is neither "choose DDN" nor "choose NetApp" nor "build Lustre." It is to make sure the organization has independent expertise before committing. Hire or contract people who have operated storage at this scale and dealt with failed OSTs, metadata bottlenecks, client storms, vendor escalations, bad firmware, and expensive systems behaving badly under real users, rather than people who have only seen big numbers in slides.

    The organization should run a structured bake-off, and a staffing bake-off alongside it. Who will operate this, tune it and patch it? Who will own user complaints? Who will understand whether a performance problem comes from storage, network, scheduler, GPU, application or filesystem behavior? Who will manage capacity planning, defend the architecture to finance, and design the exit path? If those answers are vague, the technology choice is premature.

    A £30 million storage program cannot be steered by vendor recommendations alone, even from NVIDIA. Forum horror stories alone can't steer it either, even when they sound terrifying. It needs workload evidence, reference customers, operational design, financial modeling, and a clear decision about whether the team wants a product or a platform they effectively co-own.

    The final decision is really about what kind of pain feels survivable

    DDN may still be the right answer if the workload is classic HPC/AI, the references are strong, the support contract has teeth, and the proof of concept shows real performance under ugly conditions. The horror stories mean the support model needs intense scrutiny, though they don't justify automatic rejection.

    Open-source Lustre may be the right answer if the organization has the talent and appetite to run it seriously, since the savings are real only if the operational model is real. DeepSeek's 3FS may be worth testing, especially for AI-specific workflows, but it should earn production trust slowly. NetApp may be excellent for governed enterprise data and some scale-out needs, but any claim about 120PB HPC/AI performance has to be proven the hard way instead of assumed from brand maturity. VAST and other newer architectures may deserve a look, but vendor survivability, support and exit cost need as much attention as benchmark numbers.

    120PB doesn't care about anyone's favorite vendor, open-source ideology, brand loyalty, startup energy, or who had a bad support experience in 2021. It cares about failure domains, metadata, throughput, latency, rebuilds, people, power, cooling, networking, and whether the platform still makes sense when everyone is tired. At this scale, "what storage should we buy?" turns into "what storage failure mode are we willing to live with?", and that choice is hidden inside every petabyte.