Wearable app development: constraints nobody warns you about
The first version of a wearable app usually works. The demo runs, the watch pairs, the chart on the phone moves when you wave your arm.
Then the device goes home with a patient for a week, and almost everything that matters turns out to live in the parts the demo never touched.
We spent 2.5 years on the movement disorder algorithms at PKG Health, which is now part of Empatica. Their system is FDA 510(k) cleared, CE marked and TGA approved, and our job was to refactor and validate the algorithms inside it, then make them run on other wearables. Most of what follows we learned there. None of it is exotic, but most of it only shows up once people are wearing the device at home.
The device is the smallest computer in the system
A wearable app is really three programs: firmware on the device, an app on the phone, and whatever processes the data afterwards. The first decision is which work happens where, and the device gets the least room.
Memory is tight, and every calculation the device does comes out of the battery. So for each feature, ask whether it has to run on the wrist at all, or whether the phone or a server could do it later.
When an algorithm does need to run on the device, it usually has to be rewritten for it. We ported PKG's bradykinesia, dyskinesia and tremor algorithms from Python to C for Empatica's wrist device. That took 6 months, and the slow part was the checking. Every output from the C version had to match the reference implementation before anything moved forward.
Battery life is the real spec
Sampling rate, radio use and on-device processing all draw from the same small battery. Push any one of them up and the device needs charging sooner.
It also decides how much data you get. A PKG session is about a week of continuous wrist data, and every charge is a hole in that week. If the device has to come off every night, you lose the overnight data, and with it things like sleep scoring. Work out what sampling rate the algorithm actually needs before you decide anything else, because that number sets the battery budget for the whole product.
Data arrives late, in pieces, or not at all
On a phone app, data is there when you ask for it. On a wearable, it shows up whenever the device and phone last managed to sync.
Phones limit what apps can do in the background, Bluetooth connections drop, and people leave their phone in another room. So most of the companion app's work is sync: picking up where it left off, never duplicating a chunk, and keeping timestamps straight across the device clock and the phone clock. Build the pipeline expecting gaps and late batches from day one. Retrofitting it later means reprocessing everything you've already stored.
People take the watch off
They take it off to shower, and plenty of them leave it on the bedside table overnight to charge.
A wrist that isn't moving can mean a calm patient, a patient with severe bradykinesia, or a watch sitting on a desk. If the software can't tell those apart, the off-wrist time gets scored as if someone were wearing it, and the results are wrong in a way nobody notices.
So off-wrist detection has to run before any scoring does. Our refactored version for PKG gets above 90% accuracy. We also built a machine learning version as a proof of concept, and it came in around 75 to 78%, so it never went to production. The cleaned-up original beat the newer model by a wide margin.
Two devices, two different answers
This one catches teams who already have a working product.
Say the algorithm was validated on one watch, and now a study wants to use a second one. The new device samples at a different rate, its axes point a different way, it reports acceleration in different units, and it marks gaps and charging in its own format. Feed that straight into the algorithm and it will produce scores without complaint. They'll be wrong, and nothing in the output tells you so.
For PKG we put a preprocessing layer between each device and the algorithm. It has a reader for each platform's raw format and converts everything into the single input the validated algorithm was built for, so the algorithm code itself never changes. We ran data from Apple Watch, Samsung, Sony, Empatica and ActiGraph devices through it and compared the results platform by platform against the reference device. That layer is why PKG could offer its algorithms to other device makers, Empatica and ActiGraph among them.
Processing time catches up with you
We processed thousands of week-long sessions from hospitals, clinics and research groups in Asia, Europe, the US and Australia. With one session, nobody cares whether processing takes 5 minutes or 30 seconds. With a few thousand, it decides when the results come back. Our tremor algorithm took about 5 minutes per session when we started on it. After optimisation it runs in under 30 seconds and gives the same output.
If the output is regulated, it can't drift
Consumer wearables can ship an update that nudges a step count and nobody minds. Medical ones can't.
Once a doctor or a trial sponsor is making decisions from the output, any code change has to reproduce the validated version's numbers. On the PKG work, the FDA clearance was tied to the legacy implementation, so a single deviation in our refactored outputs would have put it at risk. The clearance belongs to PKG Health, and we don't hold any ourselves. What we kept from that work is the habit: formal change control, and every output compared numerically against a reference before release. We cover what that means for software teams in what FDA 510(k) clearance means for software, and the testing side in our software testing and QA work.
Questions to settle before you write code
Most of this is cheap to plan for and expensive to fix later. These are the questions we ask first:
- What sampling rate does the algorithm actually need, and what does that do to battery life?
- What runs on the device, what runs on the phone, and what runs later on a server?
- How will the system know the device is off the wrist or charging?
- Will this ever need to run on a second device? If so, where does the conversion layer go?
- Will anyone make a clinical or research decision from the output?
If the answer to the last one is yes, plan the validation now, alongside the architecture.
We don't design wearable hardware. We come in once there's a device, an SDK or a dataset, and we write the software around it: firmware, companion apps, the data pipeline and the algorithms. Our wearable app development services page covers how a project with us runs. For the on-device side, see embedded firmware development. The PKG engagement is written up in full in the PKG Health algorithm refactoring case study, and the algorithm work itself sits under clinical algorithm development. If you have a device and a deadline, get in touch.