• 'A disk in a planet-scale computer': Meta has so many expensive G

    From TechnologyDaily@1337:1/100 to All on Sat Aug 1 18:45:25 2026
    'A disk in a planet-scale computer': Meta has so many expensive GPUs that
    it's buying SSDs to kill idle time

    Date:
    Sat, 01 Aug 2026 17:35:00 +0000

    Description:
    Meta redesigned its storage architecture with SSD caching and faster metadata access, dramatically reducing AI loading times while improving GPU utilisation.

    FULL STORY ======================================================================Copy link Facebook X Whatsapp Reddit Pinterest Flipboard Threads Email Share this article 0 Join the conversation Follow us Add us as a preferred source on Google Newsletter Subscribe to our newsletter Meta rebuilt storage systems after slow data repeatedly stalled expensive AI GPUs SSD caching dramatically reduced AI dataset loading times from hours to minutes Meta replaced complex metadata lookups with a faster unified storage architecture Meta says storage systems have failed to keep pace with AI computing power, creating delays
    that leave costly GPUs waiting instead of processing workloads efficiently.

    According to the company's engineers, storage bottlenecks remain a major
    cause of GPU stalls, increasing operating costs while slowing research progress and extending development timelines. To address those delays, Meta redesigned its storage architecture, arguing that faster movement of data can unlock greater value from expensive AI hardware investments. Latest Videos From TechRadar Watch full video here: Meta rebuilds storage architecture to keep GPUs working Meta's engineers explained that the company operates hundreds of exabyte-scale storage clusters supporting Facebook, Instagram, Reality Labs, Meta AI, advertising systems, databases, and internal data warehouses.

    Those services rely on a foundational storage layer called Tectonic, which manages object storage, file systems, block devices, and data placement
    across HDDs and SSDs . You may like This AI SSD tech makes 8 RTX 5090s
    perform like 46 GPUs in inference Why 95% of enterprise GPUs sit idle while
    AI startups can't get compute Meta will reuse terabytes worth of DDR4 memory using CXL tech and avoid paying the RAM tax

    The company has increasingly shifted from traditional file storage toward
    BLOB storage because massive AI datasets require unified access methods with significantly higher performance.

    When GPU servers request information from storage, repeated metadata lookups across several layers create latency that interrupts AI training pipelines
    and delays overall processing. Are you a pro? Subscribe to our newsletter
    Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed! Contact me with news
    and offers from other Future brands Receive email from us on behalf of our trusted partners or sponsors By submitting your information you agree to the Terms & Conditions and Privacy Policy and are aged 16 or over.

    Meta responded by rebuilding the metadata subsystem into a unified schema backed by ZippyDB while allowing clients to retrieve data directly from storage servers.

    Instead of routing every transfer through application servers, the redesigned software uses an embedded client capable of streaming information directly from the Tectonic storage layer.

    To support distributed AI deployments, Meta placed regional BLOB-storage systems beside GPU clusters, reducing delays associated with transferring information across distant infrastructure. What to read next How cloud architecture is reinventing primary storage for the AI era Stop thinking of
    AI data centers as compute systems Inference needs memory: how context is becoming AI infrastructure

    It introduced distributed caching using unused GPU host memory, producing an average cache hit rate of 80%, while metadata became accessible within 1 to 2 ms.

    The engineers further added hedged reads alongside dynamic concurrency controls, stating that "the new BLOB-storage stack is now capable of serving AI workloads without causing GPU stalls." SSD caching cuts hours from AI data ingestion Meta examined lengthy delays experienced before AI training even began, when researchers transferred enormous dataset snapshots from BLOB storage into regional GPU facilities.

    Instead of loading every dataset directly from slower storage drives, Meta created multiple cache layers that keep frequently used data much closer to GPU servers.

    That approach introduced multiple caching layers, using GPU memory as L1,
    SSDs inside GPU hosts as L2, and regional flash-backed BLOB storage as L3.

    Traditional HDD storage remained the authoritative data source, while faster cache layers supplied frequently requested information before slower disks became necessary during processing.

    Metas redesign produced sharp gains in ingestion speed, cutting a previous 150-minute loading process down to just 10 minutes.

    A separate job that once required 89 hours to complete now finishes in just over three hours from start to finish.

    Citrini analyst Jukan notes that high-capacity storage has traditionally been judged mainly by its cost per terabyte of capacity alone.

    He argues that idle GPU time can outweigh any savings gained from cheaper storage if data arrives too slowly each day.

    This suggests that the cost of flash storage is economically rational
    whenever eliminating idle time costs more than the flash itself each hour.

    Via Blocks and Files Follow TechRadar on Google News and add us as a
    preferred source to get our expert news, reviews, and opinion in your feeds.



    ======================================================================
    Link to news story: https://www.techradar.com/pro/a-disk-in-a-planet-scale-computer-meta-has-so-ma ny-expensive-gpus-that-its-buying-ssds-to-kill-idle-time


    --- Mystic BBS v1.12 A49 (Linux/64)
    * Origin: tqwNet Technology News (1337:1/100)