44 points by lateatdesk 4 days ago | 13 comments | View on ycombinator
in_absentia 2 days ago |
mrlambchop 2 days ago |
_whiteCaps_ 2 days ago |
We had a new batch of hardware come in that had a failure rate of about 50%. Normally it was on the order of < 1%. Our CM was very good at troubleshooting.
Nothing had changed in the BOM so we were left scratching our heads.
Hooked up the JTAG debugger to see where it was failing to start, and the CPU wasn't even coming up. Power rails looked good, but the CPU just wasn't booting.
Eventually we discovered that the supplier had given us a batch of crystals that were slightly more sensitive to the capacitance, and our design was just on the margin of working.
Lowering the caps to the proper values according to the xtal's datasheet got everything working again.
a96 about 2 hours ago |
Cadwhisker 1 day ago |
1. Check the power rails are correct and stable (at the target devices, not the PSU source)
2. Check the resets are correct (polarity, level, sequencing) and reaching where they're needed
3. Check the clock is toggling cleanly (no jitter, has clean monotonic waveform)
4. General signal integrity and setup/hold of signals that are related to the issue
That catches a lot of basic issues. If those are all clean, then you go deeper.
To the article author's suspicion of crystals, I have seen crystal oscillators fail (stopping toggling) as ambient temperature ramps up and down; that's a nasty one to catch and prove, but it can happen. Changing vendor was the only solution there.
kosma 2 days ago |
Neywiny 2 days ago |
XRG 2 days ago |
“Some exhibited “haunted” behavior, seemingly jumping to random sections of the mcu program code, outputting messages on the display that made no sense given the context. One of them appeared to work in slow motion, with LED blinking and display updates noticeably more sluggish than normal,”
my first thought was that it smelled like a clock issue.
Some of the nastier issues I have had the pleasure to debug included (a) traces that had microcracks which affected analog readings when the PCB heated up after prolonged usage (QC issue from the PCB fab) and (b) a (suspected) ESD strike that gradually took out several components in the weeks following as I was investigating the device while new problems kept popping up. Marginally stable composite amplifiers have also caused some headaches over the years.
Hardware really is hard.
And in the comments, they note:
"I replaced the 18pF capacitors on one of the non-working boards with 10pF capacitors, but it still doesn’t boot or respond to the debugger. That surprised me – I thought that was the answer! It was a rush rework job and I made a bit of a mess of it, including accidentally desoldering and reinstalling the crystal, so I’ll try again later with another board. But it appears that the capacitor value may not have been the issue after all. Either 10pF is a bad value, or it’s the crystal itself that’s at fault, or I’ve failed somewhere in my troubleshooting reasoning. Hmm."
In my experience, mystery stability problems are often caused by capacitors, but of a different kind: decoupling capacitors on the power supply pins. If there's not enough of them to keep up with the noise originating from motors or digital switching, I'd expect that exact issue. Intermittent "impossible" CPU states on some boards, no rhyme or reason (because sometimes, that +/- 10% saves you and sometimes it does not). I'd try more caps and possibly some ferrite beads.