Skip to content

Predictive maintenance, powered by AI

Predict failures.
Prevent downtime.

In our capstone project, we built a predictive-maintenance system that catches machine failures before they happen. Running on simulated CNC sensor data, it shows how manufacturing teams could maximize uptime and efficiency.

82%
Failure Recall "Recall" measures how many of the machine failures that actually happened were correctly flagged in advance. A higher number means fewer failures slip through unnoticed.
live · updated hourly
92%
Failure Precision "Precision" measures how often a failure warning turns out to be a real one. A higher number means fewer false alarms — you’re not chasing problems that were never really there.
live · updated hourly
Retrains Automatically The model isn’t trained once and left alone. It automatically retrains itself as new sensor data comes in, so its predictions keep improving instead of slowly going stale.
Last retrained 41 days ago

Live Factory.
Real Predictions.

The system can watch CNC milling machines around the clock. The moment tool wear, or process temperature starts to shift, it catches the change — often hours before that could cause a failure.

Isometric view of a CNC milling factory floor with five machines, glowing teal sensor lights, and a catwalk structure.
Drift detected CNC-01 Tool Wear Shift 2.3 hrs to failure See it in action →

The AI + MLOps loop behind continuous reliability

One system. One loop. Always learning.

  1. 1. Collect Sensors on the machines send real-time readings — temperature, vibration, pressure — every few seconds. Nothing else in this loop works without this: it’s the raw material every prediction is built on. Think of it as the system’s eyes and ears on the factory floor.

    Stream sensor data from the machines in real time.

  2. 2. Predict AI models look at those incoming readings and continuously estimate how likely each machine is to fail soon. Instead of waiting for a scheduled inspection, you get an early, ongoing read on machine health. This is what turns raw sensor data into a warning you can actually act on.

    AI models assess failure probability continuously.

  3. 3. Detect The system watches for two kinds of change: early warning signs of failure, and a shift in the incoming data itself (called "drift") that could make predictions less reliable. Catching drift automatically means problems get flagged before predictions quietly get worse. This step is what keeps the whole loop trustworthy.

    Anomalies and drift are caught early and automatically.

  4. 4. Retrain When drift or new patterns are detected, the model is automatically retrained on the freshest data — no engineer has to notice a problem and kick it off by hand. This keeps predictions accurate as machines age or conditions change. It’s what makes this a living system instead of a one-time model that slowly goes stale.

    Models retrain on fresh data through the pipeline.

  5. 5. Improve With an up-to-date model back in place, predictions get sharper instead of drifting downward over time. That means fewer surprise breakdowns and less time and money spent on unnecessary maintenance. This is where the loop pays off — every trip around Collect → Predict → Detect → Retrain leaves the system a little better than before.

    Predictions get better over time — cutting downtime and cost.

  6. Business Impact All of the above adds up to fewer unplanned stoppages, lower maintenance costs, and equipment that lasts longer before it needs replacing. This is the reason the whole loop exists — every technical step earns its place by producing this outcome. The results then feed back into Collect, so the improvement never stops.

    Increase uptime. Lower costs. Extend asset life.

feeds back into Collect — always learning

What the system can help with

Fewer surprises. Lower costs. Longer-lasting machines.

Reduce Downtime

Prevent unexpected failures and unplanned stops.

Emergency repairs are wasteful, not just costly — a failure caught too late often means an urgent replacement part rushed in on next-day freight instead of standard shipping. Preventing the failure means normal, planned shipping instead of a carbon-heavy scramble.

Extend Asset Life

Catch issues early and reduce wear and tear.

Every machine carries an environmental cost baked in before it’s even switched on — the energy and raw materials used to build it. Keeping equipment running longer instead of replacing it early spreads that upfront footprint over more years of useful life.

Lower Maintenance Costs

Replace parts based on actual condition, not guesswork.

Calendar-based maintenance often replaces parts on a fixed schedule whether they need it or not — swapping out something that still has useful life left. Replacing only what’s actually worn means less scrap metal, oil, and packaging.

Increase Operational Efficiency

More uptime. Smoother operations. Happier teams.

A machine running with an undetected fault — vibration, misalignment, friction — typically draws more energy than one running smoothly, wasting it as heat and wear. Catching problems early keeps machines running at their efficient best, not just their functional minimum.

While the dayshift sleep…

The system is able to monitor the machines all night.

Illustrative example
  1. 02:16 AM

    Drift Detected

    Tool-wear shift spotted on CNC-01 by Evidently AI.

  2. 02:17 AM

    Pipeline Triggered

    The GitHub Actions pipeline kicks off automatically.

  3. 02:28 AM

    Model Retrained

    A fresh model trains on the latest factory data.

  4. 02:45 AM

    Model Promoted

    The new version beats the old one and goes live automatically.

  5. 06:47 AM

    System Healthy

    The maintenance lead opens the dashboard. No calls. No incidents.

Overnight Impact

Illustrative example
  • 1 potential failure prevented
  • ~8 hours of downtime avoided
  • €24,300 saved
  • Production stayed on track

A walkthrough of one night — not aggregated customer data.

Try your own numbers in the calculator below ↓.

A maintenance lead, seen from behind with no identifiable face, checking an abstract glowing dashboard on a tablet on a factory floor at night.

Morning Summary

No incidents. Model upgraded overnight.

See the dashboard →

One system. Always learning. Always protecting.

Scenario: You are in charge of a factory

What could this save you?

Drag the slider to roughly what unplanned downtime costs you each month.

Illustrative example
€15,000
€1,000 €50,000

You may have seen bigger downtime numbers elsewhere — here’s where four widely-cited figures actually sit.

$2.3M/hr Automotive plants Siemens’ 2024 downtime report: automotive plants can lose up to this much per hour of idle production — a different scale of operation than a single production line like the one above. Source: Siemens, True Cost of Downtime 2024 . $125k/hr Global average ABB’s 2023 survey of 3,215 plant-maintenance leaders worldwide found this to be the global average cost per hour of unplanned downtime — many individual businesses report well below it. Source: ABB, Value of Reliability survey . €147k/hr Germany average The same ABB survey’s German breakout — an average across industrial companies of every size, not just large plants, so plenty of smaller operations sit well under it. Source: ABB, Value of Reliability survey (Germany) . <£500/hr Smallest UK firms The smallest UK manufacturers — the scale closest to a single production line — are the ones RS Industria’s survey found reporting under £500/hour. Source: RS Industria, UK manufacturing survey .

If nothing changes

€15,000

With predictive maintenance

€2,700

Estimated savings

€12,300 / month

Trees saved (equivalent) A rough, illustrative comparison, not a carbon calculation — this project’s model predicts failures, not emissions. Unplanned downtime commonly wastes energy (idle machines, restart surges, scrapped material); we compare the scale of that avoided waste to a mature tree’s yearly CO2 absorption (~21kg), using €1,500 saved per year ≈ one tree. Meant to give a feel for the scale, not a precise figure.

98 / year

live · updated hourly · This model currently catches 82% of failures before they happen. We use the model’s real recall rate — how many of the real failures it catches in advance. For example, since it currently catches 82% of failures, we assume 82% of your downtime cost is avoided too. It’s one real, live number, not a marketing multiplier — but it’s still an estimate. Your actual savings depend on your own failure costs, response times, and operations. See the live number →

For the technical reader

The MLOps System Behind the Predictions

The same loop from How it works — now with the tools that run it.

Built with

  • Python

    Data processing & modeling

    The programming language the whole system is written in — one of the most common languages for data science and machine learning.
  • Pandas

    Data analysis

    A Python library for loading, cleaning, and reshaping tabular data — think of it as a scriptable, programmable spreadsheet.
  • Scikit-learn

    ML modeling

    A Python library of ready-made machine learning algorithms. This is what the failure-prediction model itself is built and trained with.
  • NumPy

    Numerical computing

    The fast, low-level number-crunching library that Pandas and Scikit-learn are both built on top of — rarely used directly, but doing most of the heavy lifting underneath.

1. Sensor Data

2. Predict API

3. Drift Check

4. CI/CD

5. DVC

6. MLflow

Powered by

  • Evidently AI

    Data & model drift monitoring

    The open-source library that watches for drift — comparing live production data against the training data on an ongoing basis, not just once at launch.
  • GitHub Actions

    Automated CI/CD for ML

    A free automation service built into GitHub. It’s the engine that runs the retraining pipeline on a trigger or schedule — no separate server to babysit.
  • DVC

    Data version control

    Short for "Data Version Control" — same tool as in the diagram above. It’s what makes "which data trained this exact model" an answerable question instead of a guess.
  • MLflow

    Experiment tracking & model registry

    Experiment tracking and model registry, in one tool. Every training run gets logged and compared, and whichever model is currently "promoted" is the one serving live predictions.

Automated

No manual retraining

Reproducible

DVC-versioned data & pipeline

Monitored

Evidently drift checks

Open-source stack

Runs on your own machine

We don't track you

We don't believe in tracking people without their permission — so this site has no tracking, no pixels, and no zombie cookies. Nothing about your visit is collected, so there's nothing to ask your consent for. We back that up with real choices too: self-hosted fonts, no third-party scripts, nothing quietly calling home to anyone else.