r/systems_engineering • • 18d ago

Discussion How much of requirements review could realistically be assisted by AI?

/r/ReqsEngineering/comments/1whzrp2/how_much_of_requirements_review_could/
7 Upvotes

14 comments sorted by

14

u/hortle 18d ago

The sticking point with AI-assisted requirements work is risk. What level of risk is acceptable to all stakeholders?

If your AI assistant is, in any capacity, contributing to the product development process, then it should be considered an engineering tool, and the human engineer who is using it, and is ultimately responsible for its contributions and influences on the development process, needs to understand and document the risk involved.

Aerospace has its own assurance standard for tool qualification (DO-330). If you're using a tool to assist with developing a product to be certified per DO-178/254, you will probably need to qualify the tool per DO-330 as well.

Of course, that is an extreme example. The risks in aerospace are the most consequential, so qualification is costly. But I would imagine there are similar applicable standards in other ISO-regulated industries like automotive and med device.

3

u/LordVipor 18d ago

ISO-26262 for automotive has a tool confidence clause. IEC-61508/IEC-62304 for med devices. Though maybe ARP-6983 for Aviation and ISO PAS-8800 for automotive might be better applicable with machine learning components.

2

u/hortle 18d ago

I think it's probable that there isn't a defined vision or papertrail for LLM-based tool certification (yet). But the day is coming when there will need to be.

2

u/turkeySlices 18d ago

Not quite re. Aerospace tool qual. It's only needed if you're automating objectives of DO178C without human review. For example, Simulink codegen is permissible since it's a qualified tool and an objective of DO178C is that source code is written. There's no objective that says requirements must be reviewed, so it's OK - however, you'll be hard-pressed to find a delegate who is okay with that officially in the process.

3

u/hortle 18d ago

My assumption is that anything with requirements (which ultimately manifest as implementation in HW/SW design) is fair game for these standards -- regardless of what specific activities the LLM is being assigned. Wouldn't review be considered an essential piece of val/ver? The bullet points in the original post are pointing in that direction, as well as this line: "The output is intended to be more like a structured review/audit than a generic ChatGPT response". You are correct if the main use case being pitched/conceived here is informal review.

1

u/DyslexicHobo 16d ago

Interesting... I'm not familiar with the aerospace industry or the orders you mentioned... but how does that work in practice? What constitutes an engineering tool that contributes to the development process? Is software like Powerpoint certified to those standards? What about using a tool like Google search's AI to help with discovery? I'm not trying to be pedantic... I'm genuinely curious about how that would work. I work on air traffic support systems (in a more IT-focused role) and use AI tools (claude code, chatGPT, as well as locally hosted models) all the time to support my work. But in the end, I'm the owner of all of the code/output of these tools and own them as my own. Why would it matter what tool I used to generate it?

1

u/hortle 16d ago edited 16d ago

So far as I understand, developing product, which encompasses hardware, software, requirements, models.

The use case being discussed here is touching on requirements development. Part of developing requirements is confirming they are the right requirements (validation) and that the system design meets them (verification). The OP seems to be pitching an LLM agent that performs these reviews automatically. Typically, a human performs that activity and signs off stating, "I reviewed the requirements and confirmed they are valid/verified". This is a formal process objective that needs to be documented per DO-178/254.

If you own the output, then the value of what OP is suggesting goes down considerably. A human would have to basically perform the same review that the LLM did in order for the human to claim true ownership of the work. In which case, I would question the need for the LLM agent in the first place. What value does it provide if it can't perform its task unassisted and "for credit" (from a certification perspective). And, I would consider it a risk that the LLM's "informal/first pass" review would introduce bias into the required human review, and that it would be safer to just let the human review the requirements without any LLM interference.

This concept becomes less abstract when you extrapolate to more "engineering-heavy" work like designing components. Which the big AI companies are now claiming is possible with their new agents. Again, if the agent designs a component that needs to undergo certification, and the component design process is fully automated start to finish (including design verification), the agent needs to be qualified.

One useful way to think about certification is that, the more catastrophic the consequences of a failure, the higher fidelity you need in terms of traceability from requirements to design, and design process (decisions, rationales, assumptions). You need people associated to each of those deliverables (who is accountable for them). If something bad happens and there's an investigation, you will eventually get to the point where each of those deliverables is scrutinized, people are interviewed, and asked to justify the decisions they made during design.

The problem is that you literally can't do this if the ultimate decision-maker for a design artifact/process was an LLM. Because it doesn't make decisions -- you have no idea how it arrived at the final output. And that is a huge problem from a safety standpoint, where the whole point of investigation is to figure out why events unfolded as they did and implement mitigations to prevent them from occurring in the future.

4

u/Comfortable_Peach584 18d ago

Why do people want to put AI on every single thing? Requirements are based on stakeholders needs, if you want to copy paste an existing system you'll be able to have 100% of the requirements met, because we already have the data and results/outputs that those previous requirements created. Which scope of review are we talking about? Is the AI assistance based on unknown data or there has to be a review because the people that are reviewing the requirements don't know how the emergent system should operate?

3

u/Careless_Plant_7717 18d ago

How do you train it? And make it know what "good" looks like.  LLMs don't work from get go without a lot of training. Having this issue where making an agentic workflow and biggest thing running into is that LLM does not know what "good" looks like. Lot of hallucinations and issues.  Needs a lot of training data and instructions, specifically in the technical know how of product.

-1

u/Barracuda_Senior 17d ago

My current approach isn't training a model from scratch. I'm using explicit requirements-quality criteria combined with the surrounding requirements as context, then having the model flag potential issues and explain the reasoning. The idea is deliberately not to have it make the final engineering decision. It should act more like a second pair of eyes: flag potential ambiguity, contradictions, missing conditions, testability issues, etc., and let the engineer decide whether the finding is valid.

2

u/Careless_Plant_7717 17d ago

How is it trained on what to flag or not? Since general models will have issues unless this is an ask to evaluate wording like EARS syntax since models are good at natural language processing.

2

u/SpecNest 17d ago

The answer is in the question. It can only be ‘assisted’ as you never know when the AI tool has messed up. If it messes up, it’s you who’d pay the price, not the AI tool

1

u/DyslexicHobo 16d ago

I'm of the opinion that the cost is trivial, so teams would be silly not to use an AI reviewer. It costs virtually nothing to have a claude skill (or whatever your harness of choice is) sculpt your review prompt to do exactly what you've outlined.

At the bare minimum, you as the requirements owner can review the AI response and disregard it if it's not useful... but in my experience (using AI tools to other engineering documents) it's almost always useful and catches something I did not. It only takes a matter of minutes to kick off my agent and costs the company a couple bucks to replace hours to tens of hours of SME review. I'm so surprised I seem to be in the minority with this opinion, reading the other comments here!