A handheld setup that pairs a 60 GHz mmWave radar with an RGB camera, for the cases where camera-only vital sensing fails.
Every heartbeat changes the optical absorbance of your skin by about one percent. With a good camera and a still face, that is enough for a phone to estimate your pulse. The algorithms are public. The product is compelling.
But it does not always work. Liao et al. [10] measured per-video success rates of passive smartphone rPPG across three skin tone groups, drawing on real-world data from a clinical trial. The success rate drops fast as skin tone gets darker. The algorithm is the same. The optical signal is not.
Per-video success rate of passive smartphone rPPG, from Liao et al. [10]. Same algorithm, different SNR.
Low light, head motion, walking, typing, bad camera angle, backlight. Pick almost any real-world condition and the SNR collapses [1, 2]. rPPG is not wrong. It is a thin signal that needs good conditions.
An FMCW radar is a ranging machine. At 60 GHz it resolves sub-millimeter changes in distance. That is fine enough to see a heartbeat, which moves the chest wall by a few millimeters, and breathing, which moves it by a few centimeters. The idea goes back to Adib et al. [11] on RF-based vital sensing at home, with later work showing that deep models on radar IQ can even reconstruct the seismocardiogram waveform [12]. Light in the room does not matter. Neither does skin tone.
There is a cost. The same sensitivity that lets the radar see your heartbeat also lets it see everything else. A hand twitch. A shoulder roll. The phone wobbling in your grip. Stationary mmWave vital sensing is mature, with 90th percentile error under 6 bpm [3, 6, 7]. Handheld mmWave is the open question [8]. That is where the camera comes back in. The motion that destroys the radar which needs to be corrected.
Almost every prior mmWave vital-sensing paper assumes a fixed setup and the subject staying still [3, 4, 6, 7, 11, 12]. The moment the sensor itself moves, the chest wall signal sits on top of a much larger motion variations. Remember that the heart beat only induce sub-mm displacement and the frequency is about the jitter caused by hand tremor.
Question 1: With camera and mmWave sensor sensitive to different motion patterns, can we improve the accuracy of handheld vital sensing?
Question 2: Can we do it with the same form factor a phone has, without external sensors strapped to the body? Literature mostly skip the form factor, with the closest prior work using high-power research-grade radar without pairing it with a camera [8].
Since the mmWave sensor on Pixel 4 is not accessible for customization, we built our own setup that has common sensors a phone would have. Every component is off-the-shelf and roughly consumer-priced. The point is not new hardware. The point is a faithful, synchronous capture of every modality a future phone could plausibly carry, plus the ground truth to train against.
A FastAPI + WebSocket sensor hub runs on the Pi. Each sensor lives in its own process. The BLE devices share one thread so they can share the D-Bus connection.
Sync. Every sample is stamped at the Pi when it arrives.
Latency. Each client gets a latest-wins queue, so a slow browser drops old frames instead of buffering. Frames are compressed on the Pi and decoded off the main thread in the browser. The result: a 20 Hz stream renders smoothly with end-to-end latency under 100 ms and no drift over 10 minutes.
Three dashboards run on this hub. A multi-modal capture view. A radar-only vital-signs view with browser-side HR and RR estimation. An rPPG-only view that runs POS [14] on the webcam's JPEG frames (cross-checked against the rPPG-Toolbox [9]). All three open in any browser on the same Wi-Fi.
Next: collect 5 to 10 hours of synchronized data from 10 to 15 subjects across the conditions that stress rPPG (low light, dark skin, walking, hand motion). Train sensor fusion models [7] that learns from real ground truth when to trust which sensor. Build a mmWave + rPPG benchmark.