r/homelab • u/GeekBrownBear 720TB (raw) • 5h ago
Help Seagate Exos Drive RMA DoA. Twice???
I'm rather confused here because for 2 drives to come direct from manufacturer and be DoA is wild. Posting here to make sure I'm not going insane.
I'm using 4x 16TB Exos X16 drives (ST16000NM001G) in my Unifi NVR. One of them failed so I ran some tests on it using SeaTools and it confirmed failure. It's under warranty though so RMA it is!
Seagate makes you ship the drive to them first. Annoying but it's fine. A couple weeks later I get the replacement drive (ST16000NM002H) and it doesn't work.
I pop it into the NVR, and it makes a power up sound but doesn't spin up. The sound is like the heads unparking but I'm not sure . I take the drive out and put it in a dual bay SATA dock, same sound. Click/unpark sound, then silence, no spinning, and 3-5 seconds later it happens again.
Okay weird. It's DoA so I go through the RMA process again. Seagate refuses to do an advance replacement. So I send it back and get a new drive just today.
Same. Exact. Behavior. I'm flabbergasted, the odds of 2 DoAs is so low. I check everything now. The drive exhibits the same behavior in the NVR, Dock, my desktop, and a molex/sata adapter that you use with those USB SATA/IDE adapters.
I tried the pin3 tape trick, same behavior in all locations.
Made sure everything is working by using an older drive, spins up and goes green in all devices. Try the 2nd RMA again and same click, silence, repeat.
So now I'm at my wits end. Seagate support is less than useful so I'm hoping someone can shed some more insight into something to check.
I don't like how frustrated I am as I am a rather skilled IT person so the imposter syndrome is hitting me hard on this >.< My NVR has been in a degraded state since August 27. I'm just hoping the other 3 drives hold up and nothing surfaces before this one is fixed.
3
u/mm2kay 4h ago
I moved to a new place, and the placement for the unifi equipment wasn't ideal, it ran a bit hot, even though I know the instant nvr ran hot in general but previously it sat on my tv console in the living room. Good temps and airflow there since the room was always cold.
3 days into this new place, and the temps went up to 60C and it had some sector errors, and it didn't want to bootup. I had to do some SSH shenanigans [With the help of ChatGPT] to get it to force the NVR to reuse the drive, since it was protecting itself. [I wish it shut down if it crossed a certain temp] Drives are too expensive right now, and it's been running flawless now since I did some movement and kept the temps under control, but NVR remembers the read errors and bad sectors [51] I think 13 is the limit. I got this drive when prices were reasonable, and it was just a backup it came in a pack of two from serverpartsdeals, both tested fine. No errrors, or issues, until this heat mistake. Again wish it had protection against that by shutting down since the security cameras aren't critical, but hardware isn't cheap.
1
u/GeekBrownBear 720TB (raw) 3h ago
Drive temps are usually always around 45C and never above 50C. I have it in a wall rack in a well ventilated area.
I tested the original drive in my desktop as well, that confirmed the failure. And both RMAs have exhibited the exact same behavior in multiple environments so I'm pretty confident to rule out the equipment.
•
u/Gorbashsan 54m ago
I had batches from Western digital where I nearly had the same experience. 20 total 4tb drives, 5 went bad within week 1, RMA on them was quick, serials were nearly identical, same model and all that. Then bam three of the new RMA ones also died on me. They all had very similar serials on them and it turned out to be a small batch that had a flaw that was only kicking out errors after it been powered on for about 100 hours so they were making it past spot checks then failing.
They were passing unraid preclear running for a full 2 days. Zero issues on the smart long test as well. But then after about a 100 to 120 hours all of a sudden it was throwing multiple Smart errors and reallocating sectors at an alarming rate.
I ended up filing a trouble ticket and outlining the RMA process and the testing and the results and thankfully Western digital was pretty cool about it and they sent me a full replacement batch not just for those drives that have failed but everything that was in that serial range that I had bought which was the entire set of 20. Didn't even have to pay shipping. I got a whole new set that was from a more recent serial number so a newer batch or possibly from a different line.
At least in the end I didn't have any data loss as the entire thing was in raid 10 because it was supposed to be the new secure backup at the time for critical files. They sent the new drives out first that included the return shipping labels so I was able to swap them out one at a time and let the array rebuild each time and then just kept changing them until all the new ones were in and the old ones were out. Didn't even have any down time so that was kind of cool. And then I repackaged all of the old ones in the same box and shipping foam and anti-static bags, slapped on the shipping label they sent me over the old one and taped it back up and sent it right back.
3
u/Professional_Cup1298 4h ago
does the label on the replacement say SAS or SATA? your original is a SATA model and the ST16000NM002H looks like it could be a different interface, and a SAS drive on a SATA port can sit there and never spin up.