Fleet reliability · Physics · Survival ML
Catch GPU failures thirty days early. Physics-grounded fleet telemetry, survival models, and an LLM analyst that refuses to invent numbers.
Trained on a 10,000-device synthetic fleet across 12 sites. Metrics from the model card — PR-AUC, not ROC, because failures are rare.
Highest 30-day failure probability devices, ranked for RMA vs monitor. Full ops console (war room, SHAP deep dive, sim→real lab) runs locally via Streamlit.
Hot-humid sites dominate corrosion; hot-dry sites drive Arrhenius wearout — a non-obvious interaction the physics simulator is built to recover.
Local stack: FastAPI + Streamlit ops console + Pydantic AI analyst. This Vercel surface is the live product page with seeded fleet metrics.