Fleet reliability · Physics · Survival ML

SiliconSentinel

Catch GPU failures thirty days early. Physics-grounded fleet telemetry, survival models, and an LLM analyst that refuses to invent numbers.

Holdout results

Trained on a 10,000-device synthetic fleet across 12 sites. Metrics from the model card — PR-AUC, not ROC, because failures are rare.

Recall @ 80% precision
Holdout PR-AUC
Cox concordance
Walk-forward PR-AUC
d
RUL MAE (failed devices)

Decision queue

Highest 30-day failure probability devices, ranked for RMA vs monitor. Full ops console (war room, SHAP deep dive, sim→real lab) runs locally via Streamlit.

Fleet stress by site

Hot-humid sites dominate corrosion; hot-dry sites drive Arrhenius wearout — a non-obvious interaction the physics simulator is built to recover.

What ships in the repo

Local stack: FastAPI + Streamlit ops console + Pydantic AI analyst. This Vercel surface is the live product page with seeded fleet metrics.