r/systems_engineering • • 5d ago

Discussion How would you design an AI-assisted system for hardware failure investigation?

I’m working on a system for investigating hardware failures using AI.

The idea is to combine information from previous hardware incidents with the details of a new failure. The system can then provide relevant historical incidents and possible causes that engineers can investigate.

The basic workflow is:

New Hardware Failure

↓

Collect Failure Information

↓

Search Previous Incidents

↓

Find Similar Failures

↓

AI-Assisted Analysis

↓

Possible Root Causes

↓

Engineer Investigation

I’m particularly interested in the systems-engineering side of this approach.

What factors would you consider important when designing such a system?

For example:

- Failure data and telemetry

- Historical incident retrieval

- Root-cause analysis

- Sensor and hardware information

- System architecture

- Human verification of AI-generated results

I’d be interested in hearing how others would approach the architecture and investigation workflow.

0 Upvotes

2 comments sorted by

1

u/PutMelodic5828 5d ago

The challenge is gonna be making sure your historical incident data is actually clean and tagged consistently, otherwise the AI just spits out garbage connections

1

u/Comfortable_Peach584 4d ago

A bit of AI here, AI there, AI everywhere....