I manage 72 Dell Precision 5820 workstations that were purchased together and are about four and a half years old, just beyond their four-year warranties. Within a week, 26 systems stopped completing POST. Several show a 3-1 power-button flash code, which points toward the motherboard, memory, or CPU. Reseating memory, draining power, replacing CMOS batteries, attempting BIOS recovery, and checking components did not resolve the failures. Moving the CPU, memory, GPU, and power supply into a known-good motherboard brought one system back, and paid motherboard replacements have fixed others, strongly suggesting the boards are failing. The failures happened both during normal use and immediately after a building power shutdown for electrical testing. There are no obvious bulging capacitors or burn marks, and the systems were managed with manual driver updates. Could this be a defective motherboard batch, a power-quality problem, or another common failure mode? I need a credible root cause before the remaining systems fail too.
4 Answers
The obvious software theories do not fit very well. A Windows or graphics-driver update should not normally stop the fans and prevent the system from reaching POST, and a BIOS update cannot be installed on a board that is already completely dead. BIOS recovery, CMOS battery replacement, memory reseating, and checking the drives are still reasonable tests, but the evidence points below the operating-system level.
I would send one failed board to a competent electronics repair specialist for microscope inspection and electrical testing. A component can fail without a visibly swollen capacitor or scorch mark; power-rail regulators, MOSFETs, solder joints, and other components may need to be measured under load. A thermal camera or cold-start test might reveal a marginal component, but avoid repeatedly powering a board that may have a short. In parallel, ask Dell for escalation through the original account team and provide a formal fleet-level failure report rather than opening isolated repair cases.
The CMOS batteries were replaced in several systems and made no difference. The useful next step is testing the failed boards and correlating their serial numbers and manufacturing batches.
A failure rate of roughly one-third in a short period is far beyond normal aging. Since replacing the motherboard fixes the machines and moving the other components over does not reproduce the fault, the motherboard is the leading suspect. Preserve a few failed boards and document serial numbers, purchase dates, failure dates, POST codes, and repair results. That evidence may help push Dell toward an engineering or product-quality investigation rather than treating every case as an unrelated out-of-warranty repair.
That is exactly the concern. I can pay for repairs, but I need to know whether the remaining systems are likely to fail as well.
Do not completely dismiss the building power event. A poor shutdown, switching transient, missing surge protection, neutral issue, or other power-quality problem could damage systems when power is removed or restored. Compare the failure times with the electrician's work and have the logger checked for voltage excursions, brownouts, transients, and neutral-to-ground problems. Also verify that the workstations and UPS or surge equipment are properly grounded and that the power supplies are receiving the expected voltage.
The fact that many machines failed after the room was powered back on makes power quality worth investigating, even if the motherboard design is also unusually vulnerable.

That matches what I found. The latest failures happened at power-on with no operating-system logs, and recovery attempts did not change anything.