Microsoft reports: “There is no ground truth available … machines labelled ‘faulty’ often outnumber healthy ones in the hot buffer.”Xiong et al. (2026), “SuperBench,” ACM TOCS 44 (2).
AI-factory yield management
Infrastructure software that turns quarantined GPU nodes into decisions.
After failures that affect hundreds of machines, infrastructure teams find that standard diagnostics often pass while machines are still unsafe to return.
Investigating these “grey failures” is a high-stakes race to recover valuable assets without risking more failures.
Zenobia preserves each incident and actively investigates the isolated hardware inside the operator’s own infrastructure, one intelligently chosen experiment at a time, until the evidence justifies action.
The published operating problem
The people operating this hardware at scale already describe it, in print.
Published systems research from the largest fleets documents the key constraint: machines are flagged as “issue detected” much faster than anyone can determine what should happen to the hardware.
ByteDance reports: “Verifying whether all evicted machines are faulty is challenging … a complicated and time-consuming process.”Deng et al. (2025), “Minder,” NSDI ’25.
Google reports: “Most swapped chips aren’t fully diagnosed … the economics of managing and triaging is prohibitive.”S. Mitra et al. (2025), IEEE Design & Test 42 (6).
The consequence is organizational. Without clarity, Microsoft’s infrastructure teams write, “customers, cloud production teams, and vendors engage in unproductive blame-shifting.”
Correlation has hit its ceiling. Alibaba’s root-cause studies find that extending correlation methods “only leads to an increase in unacceptable accuracy loss.”
Organizations cited are sources of published research and are not presented as Zenobia customers or endorsers.
The method, once
The set of possible causes gets cut down with each tailored experiment.
Diagnostics asks fixed pass/fail questions in a set order. Zenobia chooses each experiment for what it can actually rule out in this incident, then stops when the evidence is strong enough to act.
Incident: accelerator resets after thermal soak under a memory-heavy workload · illustrative synthetic trace
- 01
Hold the workload constant; vary temperature.
The phenomenon appears only above 78 °C inlet—a thermal gate distinct from generic throttling.
workload software6 → 4 causes - 02
Hold temperature; vary the memory access pattern.
One synthetic row-stress pattern raises event frequency five-fold. The mechanism is workload-shaped.
power transient4 → 3 causes - 03
Run the same pattern on matched healthy nodes.
Twelve controls stay clean. The response diverges only on the suspect device, so the issue follows specific hardware.
fleet config · interconnect3 → 1 cause - 04
Decide confidently without transistor-level certainty.
Disposition: return with limits—inlet capped at 75 °C—plus one named vendor-only test to localize the device memory path.
1 mechanism · 1 named test
Stop condition: enough evidence to take a safe, commercially useful action—not an impressive pile of telemetry.
The case passport
Not a report.
A resumable investigation.
The passport carries the live hypotheses, experiments already run, the smallest workload or context that reproduces the fault, and the next suggested test. A vendor engineer can continue from that state using their own specialized tools.
Raw workloads and proprietary internals stay where they belong. The investigation travels.
Where Zenobia stands
Explicit uncertainty.
Ours included.
Where we are now
- Bounded experiment runtime on Blackwell-class hardware in quarantine, with hard thermal, power, and timing stops.
- Incident capture and matched controls against a selected healthy group for each investigation.
- Case passports for closed investigations, including the experiment ledger and reproducer.
- A major Fortune 500 design-partner deployment is currently underway.
Our product roadmap
- Vendor-side replay environments are not yet supported. The passport format is ready; product-engineering integrations are conversations we want to have.
- Supplier child cases are designed, but have not yet been exercised end to end across a real organizational boundary.
- Additional hardware families are in development, particularly custom silicon.
Walk through the example case with us.
Forty-five minutes, your quarantine workflow, one hardware family. Or start smaller.