Palo Alto, Calif. – October 05, 2026 -- Clockwork.io has raised $31 million in new funding and secured production deployments at LinkedIn and Together AI, bringing its total funding to $73 million as enterprises seek to curb GPU downtime during large-scale AI training.
LinkedIn prevents tens of thousands of GPU-hours of downtime monthly
LinkedIn has deployed Clockwork.io's LinkPass network fault-tolerance software across its AI infrastructure fleet. Raghu Hiremagalur, SVP, CTO Infrastructure at LinkedIn, said a single InfiniBand NIC flap previously could remove an eight-GPU server from service, while a switch port flap could double the impact to 16 GPUs. LinkPass now automatically reroutes traffic onto healthy paths, letting jobs continue uninterrupted while faults are repaired.
Meta reported an infrastructure failure roughly every three hours during Llama 3 training
Large distributed AI jobs spanning thousands of GPUs must stay synchronized, and a single failed GPU, dropped link, or frozen server can stall an entire run. Meta reported unexpected interruptions averaging about one every three hours during a 54-day period of Llama 3 training on 16,384 GPUs. The standard recovery method -- reloading a checkpoint -- can take up to 90 minutes, leaving healthy GPUs idle and forcing jobs to repeat completed work.
Together AI will bring TorchPass to market as a service on its GPU Clusters
Together AI is integrating Clockwork.io's TorchPass into its GPU Clusters offering. Pavneet Ahluwalia, Product Lead at Together AI, said the company already detects faults and provisions replacement capacity automatically, and is adding TorchPass and LinkPass as a further resilience layer. The two companies plan to demonstrate a live multi-node training job continuing through injected network and GPU failures at the PyTorch Conference.
WhiteFiber expands Clockwork.io deployment across its GPU-as-a-service clusters
WhiteFiber (NASDAQ: WYFI), an existing customer, is broadening its use of Clockwork.io software across its global GPU-as-a-service footprint. Tom Sanfilippo, Chief Technology Officer at WhiteFiber, said the company's automated fleet audit validates every link and node at once and localizes faults in minutes, allowing corrections before cluster acceptance.
New TorchPass capabilities capture whole-job state without code changes
Clockwork.io introduced two TorchPass capabilities. Multi-node platform snapshots -- described as an industry first for training -- save an entire running job across every node without requiring changes to training code, enabling restoration after an interruption too large to migrate around. Fast, asynchronous application checkpoints run in the background to deliver updated model weights to inference replicas generating rollouts in reinforcement learning, reducing time replicas spend idle or working from stale models.
SemiAnalysis measured training goodput loss falling from 14% to under 3%
Dylan Patel, Founder and CEO of SemiAnalysis, said TorchPass cut training goodput loss from 14% to under 3% for a gold-rated neocloud in the firm's ClusterMAX benchmark. Patel said reinforcement learning ties training and inference together, since inference replicas generate rollouts that trainers learn from, and updated weights must return to those replicas without stalling either direction.
Premji Invest, Wing Venture Capital and Seligman Ventures co-led the funding round
The $31 million round was co-led by Premji Invest, Wing Venture Capital and Seligman Ventures, with participation from existing investors NEA and e& Capital. Clockwork.io will use the capital to accelerate rollout of its fault-tolerance suite across training, inference and reinforcement learning, expand enterprise adoption, and scale delivery through cloud partners. Greg Papadopoulos, Venture Partner at NEA, said the firm first backed Clockwork.io in 2021 and that the company's technology is already saving tens of thousands of GPU-hours a month for customers.