r/embedded • u/Grubzer • 23h ago
Is race-to-halt viable on MCUs?
Is it usually more power efficient to run at lower clocks for longer and then deep sleep, or keep clocks high for shorter and then go to same deep sleep? Or it heavily depends on the hardware and exact use case?
13
u/Xenoamor 23h ago
I've always found it's better to run with a fast clock and sleep for as long as possible. This is assuming your code is using interrupts properly though and isn't sat with its thumb up its arse polling
4
u/Miss_Giorgia 23h ago
Fast compute than deep sleep with good interrupt routines and never pooling. Also be sure you max current draw when active at full power, even in short bursts, is below the max your battery can deliver.
2
u/dmills_00 23h ago
Depends on the behavior of the chip, the surrounding circuitry and the use case.
In particular, does the power supply switcher support lower voltage for slower clocks, that trick sometimes moves this decision.
Really, if you are doing micropower design, you need to measure.
7
u/cad_and_caffeine 22h ago
Flash wait states and PLL lock time are the other two big ones that can flip the math on short wakeups. On a lot of STM32L/U parts, waking up on the internal RC oscillator at 16 or 24 MHz lets you stay in the lower Vcore range with 0 flash wait states and a couple microseconds of wakeup time. If you crank the clock to 80+ MHz just to check a threshold and go back to sleep, you end up burning more energy stepping Vcore up, waiting for the PLL to lock, and eating 3-4 flash wait states than the faster execution saves.
1
u/dmills_00 20h ago
Even worse if you are using a crystal which can take significant time to start before the PLL lock can even begin.
2
u/ComradeGibbon 12h ago
Yeah how fast can you wake up is really the important metric.
Xtals and PLL's take too long to start up. Sleep modes that require reset and re-initialization take too long.
Important thing to consider because a lot of uP's do a terrible jobs of this is you need a fast to read monotonic time base that continues to run in sleep mode.
(insert rant about hardware designers that think keeping time in hours minutes, seconds, and low resolution sub seconds is super duper)
1
u/flatfinger 12h ago
One thing I've often wondered is why so many chips require a PLL to operate at high speed rather than having an RC oscillator which is specified to operate at a speed that the rest of the chip will be able to handle given process variations and temperature (meaning that if a CPU was specified as being able to operate at a clock speed of 100MHz, there would be no guarantee that the RC oscillator wouldn't run faster than that, but that there would be a guarantee that it would only do so under conditions that would allow the rest of the chip to operate correctly at whatever speed it happened to run).
While there are some cases where a PWM would be needed, in many situations a PWM would seem like overkill.
1
u/cad_and_caffeine 8h ago
I think the biggest catch on MCUs is that embedded Flash access time doesn't track CMOS gate delay across temperature and process corners. A free-running ring oscillator made of standard cells would happily speed up 25% on a cold fast-corner die while the core logic keeps up, but the Flash sense amps and bitline RC don't speed up at the same slope, so you'd blow past your fixed wait-state budget (plus you'd need async CDC bridges to keep UART/SPI/timer clocks from drifting all over the place).
2
u/dragonnnnnnnnnn 23h ago
It depends on the hardware and exact use case. Because they can be two cases:
- you wakeup, do some pure calculation of read something over some high band-witch link, process it and go to sleep. Then it make sense to run high clocks and "race to halt"
- you wakeup and have to receive small amount of data over some slow link or slow responding device (lets say i2c, or a mbus heatmeter etc.). In that case it doesn't make sense to stay at high clock because you are idling anyway.
1
u/mustbeset 23h ago
Didn't need to calculate that often but like others say "run fast & sleep often" is the general go to.
"current in run" is current_offset + f_cpu * current_per_MHz.
For my current controller it's 22.5mA in Run @180MHz, and 31mA@250MHz, in stop it's 0.15mA. in that configuration it's nearly always a win.
All vendors mention current consumption in the datasheets, you may have to add it up. some vendors give you a tool to calculate energy consumption, Based on 'every' setting.
1
u/felixnavid 20h ago
All modules of an MCU have a static power consumption ( if the module is ON) and a dynamic power consumption ( scales with frequency).
If you run deep sleep (with only an internal 32kHz oscillator or without even that), than you should also take into account how much time it takes to start the oscillator and/or the PLL. Newer MCUs have a very fast startup for osicllators/PLLs, but not always.
If you want very fast clocks, the PLL/support cirucit might consume more power then just a linear scale of the frequency.
Also some MCUs have disproportionate CPU performance to Flash reading speed. A lot of slightly older STM32s have cache accelerators because the CPU is much faster than the program flash. If your program doesn't fit into the cache, the CPU might just wait after the flash, consuming more power.
1
u/Plastic_Fig9225 20h ago edited 20h ago
Or it heavily depends on the hardware and exact use case?
Yes, it does. As others have noted, (CPU) clock speed may not be liniarly related to time required. Waiting for memory or, god forbid, for that I2C sensor to send a response is better spent at a lower clock.
But nice to see someone bring up the question, because "slower clock = less energy per task" is not generally true.
1
u/Panometric 17h ago
In general running the clock at Max wins because power scales with speed, but the base consumption is constant. The exception would be if you are waiting for a peripheral, basically wasting fast clocks.
1
u/ElSalyerFan 16h ago
Profile it. Depending on how much power your connected peripherals need, how much processing without peripherals you do, and how much time you spend on deep sleep, you get to wildly different answers. In my current project i started at 2Mhz because "slow clock = low current" but after tuning it it went all the way up to 12Mhz because i spent less time in power hungry routines.
1
u/duane11583 14h ago
If you can change power levels (voltages) race often wins
You need to measure your idle leakage current and integrate that over time
Verses you slow leakage current over time
Remember that power is (V2) over R
So if you can cut voltages power goes down by the square and power is what counts for batteries
Ie you power down and stop clocks on everything you can then enter WFI and after the WFI you turn clocks and voltages
There are often two or three voltage levels ie vdd retention (nothing is lost but you cannot operate even at a slow clock) vdd slow clock and then vdd nominal and vdd performance for each of these you voltage goes up and the power goes up by the square
This is why cellphone chips have a 20-40 separate power domains.
Example: if the screen is off then turn the voltage to the GPU off thus the GPU has no voltage thus no leak and 0 volts squared is still 0 watts no matter what you do
I doubt you will have that level of detail for your chips but you do have the details of your design if not go measure it and create a big ass excellent spread sheet to model it
1
u/duane11583 14h ago
Oh and do not forget your turn on times
Power supplies do not turn on in micro seconds
24
u/Ordinary-Lifeguard47 23h ago edited 23h ago
The standard modern strategy is to sleep deep, wake up extremely fast and get your task done very quick, extremely fast back to deep sleep again.
Lower power uCs (STM32U0/3/etc.) are optimized for this.