AI-factory yield management

Infrastructure software that turns quarantined GPU nodes into decisions.

After failures that affect hundreds of machines, infrastructure teams find that standard diagnostics often pass while machines are still unsafe to return.

Investigating these “grey failures” is a high-stakes race to recover valuable assets without risking more failures.

Zenobia preserves each incident and actively investigates the isolated hardware inside the operator’s own infrastructure, one intelligently chosen experiment at a time, until the evidence justifies action.

The published operating problem

The people operating this hardware at scale already describe it, in print.

Published systems research from the largest fleets documents the key constraint: machines are flagged as “issue detected” much faster than anyone can determine what should happen to the hardware.

01 · Microsoft · systems research
Microsoft reports: “There is no ground truth available … machines labelled ‘faulty’ often outnumber healthy ones in the hot buffer.”
Xiong et al. (2026), “SuperBench,” ACM TOCS 44 (2).
02 · ByteDance · systems research
ByteDance reports: “Verifying whether all evicted machines are faulty is challenging … a complicated and time-consuming process.”
Deng et al. (2025), “Minder,” NSDI ’25.
03 · Google + Stanford · research
Google reports: “Most swapped chips aren’t fully diagnosed … the economics of managing and triaging is prohibitive.”
S. Mitra et al. (2025), IEEE Design & Test 42 (6).

The consequence is organizational. Without clarity, Microsoft’s infrastructure teams write, “customers, cloud production teams, and vendors engage in unproductive blame-shifting.”

Correlation has hit its ceiling. Alibaba’s root-cause studies find that extending correlation methods “only leads to an increase in unacceptable accuracy loss.”

Organizations cited are sources of published research and are not presented as Zenobia customers or endorsers.

The method, once

The set of possible causes gets cut down with each tailored experiment.

Diagnostics asks fixed pass/fail questions in a set order. Zenobia chooses each experiment for what it can actually rule out in this incident, then stops when the evidence is strong enough to act.

Incident: accelerator resets after thermal soak under a memory-heavy workload · illustrative synthetic trace

  1. 01

    Hold the workload constant; vary temperature.

    The phenomenon appears only above 78 °C inlet—a thermal gate distinct from generic throttling.

    workload software6 → 4 causes
  2. 02

    Hold temperature; vary the memory access pattern.

    One synthetic row-stress pattern raises event frequency five-fold. The mechanism is workload-shaped.

    power transient4 → 3 causes
  3. 03

    Run the same pattern on matched healthy nodes.

    Twelve controls stay clean. The response diverges only on the suspect device, so the issue follows specific hardware.

    fleet config · interconnect3 → 1 cause
  4. 04

    Decide confidently without transistor-level certainty.

    Disposition: return with limits—inlet capped at 75 °C—plus one named vendor-only test to localize the device memory path.

    1 mechanism · 1 named test

Stop condition: enough evidence to take a safe, commercially useful action—not an impressive pile of telemetry.

The case passport

Not a report.
A resumable investigation.

The passport carries the live hypotheses, experiments already run, the smallest workload or context that reproduces the fault, and the next suggested test. A vendor engineer can continue from that state using their own specialized tools.

Raw workloads and proprietary internals stay where they belong. The investigation travels.

Not an archiveA resumable investigation state.
Case ZN-2026-1842Worked example · synthetic case
Open · actionable
PhenomenonNVLink CRC burst above 78 °C inletEscalates to link drop within 40–90 seconds under row-stress.
Reproduction11 of 15 runs (73%) · reproducer R-3Synthetic pattern; conditions and variance recorded.
EliminatedWorkload · power · fleet config · thermalEach exclusion is linked to its experiment.
Evidence53 experiments · signed hashesAppend-only ledger; integrity without broad export.
Next named testVendor-only HBM per-channel retention scan on the suspect. The test runs inside the vendor’s environment.

Where Zenobia stands

Explicit uncertainty.
Ours included.

Where we are now

  • Bounded experiment runtime on Blackwell-class hardware in quarantine, with hard thermal, power, and timing stops.
  • Incident capture and matched controls against a selected healthy group for each investigation.
  • Case passports for closed investigations, including the experiment ledger and reproducer.
  • A major Fortune 500 design-partner deployment is currently underway.

Our product roadmap

  • Vendor-side replay environments are not yet supported. The passport format is ready; product-engineering integrations are conversations we want to have.
  • Supplier child cases are designed, but have not yet been exercised end to end across a real organizational boundary.
  • Additional hardware families are in development, particularly custom silicon.

Walk through the example case with us.

Forty-five minutes, your quarantine workflow, one hardware family. Or start smaller.