结论:设备在一个部件失效时仍应安全,这是考核的基本逻辑
电气安全的一个基本思路是:设备不仅在正常状态下要安全,在任意一个安全相关部件失效的情况下,也不应当产生危险。
这就是单一故障状态的考核逻辑。注意是「单一」 ——同时考虑两个独立故障通常不在考核范围内(除非一个故障必然导致另一个)。
这类项目在送检时不合格的比例不低,而其中相当一部分问题,企业在送检前用简单的方法就能自查出来。 自查发现比测试发现要省时省钱得多。
常见的故障项
| 故障状态 | 考察什么 |
|---|---|
| 保护接地断开 | 漏电流是否仍在限值内 |
| 中性线断开 | 设备行为与漏电流 |
| 任一相线断开 | 设备行为 |
| 元器件短路或开路 | 是否引起过温、起火、漏电 |
| 风扇停转 | 温升是否超限 |
| 散热受阻 | 温升 |
| 电机堵转 | 温升与过流保护 |
| 限制器失效 | 温度控制装置失效后的温升 |
| 输出短路 | 电源部分的保护 |
| 电池反接或短路 | 保护动作 |
其中温升相关的项目占比较大,因为很多故障的最终表现都是温度升高。所以温升测试在单一故障考核中占的分量很重。
自查的方法与顺序
第一步,梳理安全相关的元器件清单。 哪些元件失效会影响安全——变压器、光耦、限流元件、温控元件、保护元件。
第二步,对每个元件判断失效模式。 短路、开路,哪种更可能、后果如何。
第三步,分析后果。 会导致什么——过温、过流、漏电、机械危险。
第四步,确认保护措施。 有没有保护、保护是否可靠。
第五步,实测关键项。 有条件的话实际做几个关键的故障模拟。
第六步,记录分析过程。 这份记录本身在送检和审查时都有用。
第一到第四步是纸面分析,不需要设备,但能发现相当一部分问题。比如分析发现某个温控元件失效后没有二次保护,这个问题在纸上就能看出来。
实测时的简易方法
有条件做实测的,几个相对容易做的项目:
接地断开后的漏电流。 用漏电流测试仪,断开接地后测量。这个测试不复杂,但能发现明显超标。
风扇停转的温升。 把风扇停掉,运行至稳定,测关键点温度。用热电偶或红外测温都可以做初步判断。
散热受阻。 按可能的遮挡方式堵住通风口,测温升。
电机堵转。 堵住运动部件,看保护是否动作、多久动作、温度如何。
输出短路。 短路输出端,看保护动作。
这几项用基本的测量工具就能做初步验证,虽然精度和条件控制不如正式测试,但足以发现明显问题。
不合格的典型原因
温升超限。 通常是散热设计余量不足,正常状态下刚好合格,故障状态就超了。
缺少二次保护。 温控元件失效后没有备份保护。
保护元件选型不当。 保护动作的阈值或速度不合适。
接地断开后漏电流超限。 Y电容容值过大是常见原因。
元件规格余量不足。 故障状态下元件承受的应力超出规格。
布线与间距问题。 故障时可能发生的击穿或短路。
「正常状态刚好合格」是一个危险信号。 如果正常运行时温升已经接近限值,那么任何降低散热能力的故障都会导致超限。设计时应当留出余量,而不是卡着限值做。
与风险管理的衔接
单一故障分析与风险管理文件中的内容是相关的:
故障模式的识别。 与风险分析中的失效模式识别重合。
后果的评估。 与风险评价对应。
保护措施。 与风险控制措施对应。
验证。 测试结果验证了控制措施的有效性。
建议把单一故障分析的结果直接用于风险文件,两边共用一套分析,既减少重复也保证一致。分开做的话,两份文件的内容容易对不上,而审查时会核对。
送检准备的建议
提供电路图与元件清单。 实验室需要据此确定故障项。
说明安全相关元件。 哪些元件是安全相关的,有什么规格要求。
提供自查记录。 自查过的项目和结果,有助于实验室聚焦。
提供软件相关说明。 如果有软件控制的保护,要说明其逻辑。
准备备用样品。 某些故障模拟可能损坏样品。
最后一条很实际。 单一故障测试中有些项目会导致元件损坏甚至样品报废,送检时应当准备足够的样品,避免测到一半没样品了。
软件相关保护的考虑
现在很多保护由软件实现,这带来额外的考虑:
软件失效也是一种故障。 软件跑飞、死机、逻辑错误,都应当在故障分析中考虑。
软件不应当是仅有的一道保护。 关键的安全功能建议有硬件层面的后备,比如独立的温度熔断、机械限位。
看门狗的有效性。 如果依赖看门狗复位,要验证其确实能在异常时动作。
复位后的行为。 软件复位后设备处于什么状态,是否安全。
软件与硬件保护的配合。 两者的动作阈值和顺序要合理。
这一条值得强调。 软件的失效模式难以穷举,把安全完全寄托在软件上风险较高。硬件后备虽然增加成本,但可靠性层次更高。
自查记录的价值
做过的自查应当形成记录,它的价值在几方面:
送检时供实验室参考。 帮助聚焦,减少沟通。
作为风险控制的证据。 在风险文件中引用。
内部知识积累。 下一个产品设计时可以参考。
变更时的基线。 设计变更后可以对照原分析判断影响。
建议用统一的格式记录 ——元件、失效模式、后果、保护措施、验证方式。格式统一之后,跨产品的经验就能积累起来。
我们的做法
承接医用电气设备的委托时,我们会先看电路图和元件清单,与委托方确认故障项的范围。如果发现设计上明显缺少某项保护,我们会在正式测试前提出,因为在纸面上改比测试失败后改成本低得多。
样品方面,我们会根据故障项的数量建议准备的样品数,避免因样品损坏导致测试中断。
如果你在准备送检,想先自查一遍单一故障相关的项目,可以把电路图和元件清单发过来一起看,或者直接联系:132 4819 8029。能力范围见服务介绍,送检要求见送检要求,其他问题见常见问题。
English version
Conclusion. A basic principle of electrical safety is that equipment must be safe not only in normal condition but also when any one safety-related component fails. That is the logic of single fault condition assessment. Note the word single: two independent faults are not normally considered together unless one necessarily causes the other. Failure rates on these items at submission are not low, and a substantial share of the problems could be found by the manufacturer beforehand using simple methods. Finding them in-house is far cheaper and quicker than finding them in the laboratory.
Common fault conditions. Interruption of protective earth examines whether leakage remains within limits. Interruption of the neutral examines equipment behaviour and leakage. Interruption of a line conductor examines behaviour. Short or open circuit of components examines whether overheating, fire or leakage results. Fan stall examines whether temperature rise exceeds limits. Obstructed ventilation examines temperature rise. Motor stall examines temperature rise and overcurrent protection. Failure of a limiting device examines temperature rise once temperature control is lost. Output short circuit examines protection in the supply section. And battery reverse connection or short circuit examines protective operation. Temperature-related items make up a large share, because many faults ultimately manifest as rising temperature, which gives temperature rise testing considerable weight in this assessment.
Self-check method and order. List the safety-related components, meaning those whose failure affects safety: transformers, optocouplers, current-limiting components, thermal control components and protective devices. Determine the failure modes of each, short or open, which is more likely and what follows. Analyse the consequences, whether overheating, overcurrent, leakage or mechanical hazard. Confirm the protective measures, whether they exist and whether they are reliable. Measure the key items where facilities allow, actually simulating a few important faults. And record the analysis, which is itself useful at submission and review. The first four steps are paper analysis requiring no equipment, yet they reveal a substantial share of problems: finding on paper that a thermal control component has no secondary protection behind it needs no measurement at all.
Simple measurements. Where measurement is possible, several items are relatively easy. Leakage with earth interrupted can be measured with a leakage tester after disconnecting earth; the test is not complex and reveals clear excursions. Temperature rise with the fan stalled can be checked by stopping the fan, running to steady state and measuring key points with thermocouples or infrared for a preliminary judgement. Obstructed ventilation can be simulated by blocking the vents as they might realistically be blocked. Motor stall can be simulated by blocking the moving part and observing whether protection operates, how quickly, and what temperature results. And output short circuit can be applied to observe protective operation. These give preliminary verification with basic instruments; precision and condition control fall short of formal testing but suffice to reveal obvious problems.
Typical causes of failure. Temperature rise exceeding limits, usually from insufficient thermal design margin, passing in normal condition and failing under fault. Missing secondary protection behind a thermal control component. Inappropriate protective device selection in threshold or speed of operation. Leakage exceeding limits with earth interrupted, often from excessive Y capacitance. Insufficient component rating margin, with stress under fault exceeding the rating. And wiring and spacing problems allowing breakdown or short circuit under fault. Passing narrowly in normal condition is a warning sign: if temperature rise already approaches the limit in normal running, any fault reducing cooling will exceed it. Design with margin rather than to the limit.
Connection to risk management. Single fault analysis overlaps with the risk management file. Identifying fault modes overlaps with failure mode identification in risk analysis. Assessing consequences corresponds to risk evaluation. Protective measures correspond to risk controls. And testing verifies the effectiveness of those controls. Use the single fault analysis directly in the risk file, sharing one analysis between both, which reduces duplication and ensures consistency. Done separately, the two documents drift apart, and reviewers check them against one another.
Preparing for submission. Provide circuit diagrams and component lists, which the laboratory needs to determine the fault items. Identify the safety-related components and their rating requirements. Provide your self-check records, which help the laboratory focus. Explain any software-controlled protection and its logic. And prepare spare samples, since some fault simulations damage them. The last is practical advice: certain single fault tests damage components or write off samples entirely, so submit enough samples to avoid running out mid-programme.
How we handle it. For medical electrical equipment we review the circuit diagram and component list first and agree the scope of fault items with the client. Where a protection is clearly missing from the design we raise it before formal testing starts, because changing it on paper costs far less than changing it after a failure. On samples, we advise how many to prepare based on the number of fault items, so that testing is not interrupted by damaged samples.
Send us the circuit diagram and component list and we will review the fault items with you. Phone or WeChat: +86 132 4819 8029.