Glossary · Automation software engineering and architecture
Defect triage
Also known as: Bug triage, Issue triage, Anomaly triage, Triage meeting
German: Fehlertriage
In software engineering, defect triage is the process of reviewing incoming defect reports and deciding for each one whether it is valid and reproducible, how severe its impact is and how urgent the work on it is, which component and team it belongs to, and when it will be addressed. The decision is recorded in the issue tracker as a set of fields — state, severity, priority, component, owner, target release — so that later work and reporting can rely on it. Triage is a management decision about work, not a technical fix: it neither changes the product nor explains the cause. The term itself is not defined in a standard; the attributes it assigns, such as severity and priority, correspond to the anomaly classification attributes described in IEEE Std 1044.
- Software engineering
- Technical documentation
- Validation
In one sentence
Defect triage reviews incoming defect reports and decides validity, severity, priority, component, owner and target release.
Example
In the weekly triage call for a bottling line project, the team closed two reports as duplicates of an existing HMI defect, raised one report to blocking severity because the fault also appeared in manual mode, routed a suspected interlock problem to the safety change process instead of the sprint backlog, and deferred the remaining cosmetic findings to the next minor release.
How it applies
- The four questions: A triage decision answers whether the report describes a real deviation from specified behavior, how bad the impact is, who owns the affected part of the system, and what happens next (fix now, schedule, defer, reject, duplicate). If your tracker cannot express all four, triage results will be reconstructed from comments later — badly.
- Severity and priority are different fields: Severity describes the impact of the software defect on operation; priority ranks the work against everything else in the backlog. IEEE Std 1044 keeps such attributes separate, and so should your workflow: a high-severity fault in a rarely used option can legitimately get low priority, and that trade-off should be visible rather than hidden in a single number.
- Write the rules down: Document a severity ladder with observable criteria (production stop, safety function affected, data loss, workaround available, cosmetic), the response times you actually commit to, and who decides in case of disagreement. A severity definition that says "critical = very important" produces inconsistent data and unusable software metric reporting.
- Intake quality decides triage quality: Provide a report template that asks for reproduction steps, expected and observed behavior, build or firmware version, and configuration. In automation projects, add the project and software configuration version, controller firmware, operating mode, alarm and I/O state, and whether the issue appeared in virtual commissioning, during a factory acceptance test (FAT) or on site.
- Route safety-relevant findings out of the backlog: A report affecting an interlock, a protective stop or another safety function belongs in the safety-related change and assessment process, not only in the development queue. Triage should have an explicit exit for this, plus an exit for security findings that need coordinated disclosure.
- Link triage to the rest of engineering: Connect defect IDs to the branch or pull request that fixes them, to the regression test case that covers the fault afterwards, and to the release note entry. Without those links, "fixed" is an assertion, not evidence.
- Where AI helps: Assistants and classifiers are used to detect probable duplicates, suggest the component or team, propose a severity, cluster crash reports, and summarize long logs or attached traces into a readable report. This is a recommendation into a human decision rather than new logic in the shipped product, so the failure mode is a misrouted or mislabeled ticket — recoverable — not a latent fault in the control software.
- Keep the decision human and traceable: Record who accepted, changed or rejected a suggestion, and keep a reopen path for automatically closed duplicates. Suggestion quality depends on historical labels, so systematically mislabeled history will be reproduced; re-check the model's routing accuracy after component or team reorganizations.
- Measure the process, not the tool: Useful indicators are time to first triage decision, share of reports reopened after being closed, misroute rate, share of "cannot reproduce", and backlog age per severity. Naming an AI feature in a tooling report says nothing about whether triage got better.
- Regulated environments: Where defect handling is part of configuration and change control — process industry projects, medical or pharmaceutical software — the triage record is part of the evidence trail: assessment of impact, justification for deferring a known defect, and the link to the validation documentation. Check with your quality function whether AI-generated classifications may appear in that record and how they must be marked.
- Employment-law and AI Act caution: Triage data also describes people (who reported, who fixed, how fast). If AI-supported tooling is used to evaluate the performance or behavior of workers, the employment-related uses in Annex III of the EU AI Act and, in Germany, works council co-determination come into view; that requires a case-by-case assessment, and using an AI feature marketed as compliant does not answer it.
Defect triage vs. root-cause analysis
Triage happens before the investigation and decides whether and when someone investigates: it classifies, assigns and schedules. Root-cause analysis happens inside the fix and explains why the behavior occurs — which change, condition or design decision produces it — and usually spans several reports. Confusing them creates two typical failures: triage meetings that turn into debugging sessions and stall the queue, and fixes shipped after triage "assigned" a cause that nobody verified. Keep the outputs separate: triage produces fields in the tracker, root-cause analysis produces an explanation, a covering test and, where the cause is systemic, a refactoring or process action.