Why Does Device Data Matter in Population Health Analytics?

Apple and Aetna have been in talksto provide Aetna’s book of members with free or subsidized Apple Watches. Although there is plenty of speculation about the details of the agreement, it’s certainly no secret that both healthcare and tech industries are showing great interest in getting consumers to track their health with mobile devices, and more specifically, wearables; market for which is projected to double by 2021. There is a lot of optimism (not necessarily well founded) that devices that track health measures can become important tools in population health management, as consumers will have easier access to a greater number of metrics around their personal health status.
But there is a perhaps less obvious drive for wearable adoption from the healthcare industry: a deluge of data collection. And while consumer health behavior change from wearables has shown at best mixed results, the value of the data that wearables monitor is indisputable. In short order, I’ll discuss: why healthcare companies (insurers, hospitals, pharma) care about device data; how and what kinds of data can be collected from a device; and then how those data are collected, aggregated, and turned into information that can be used for modeling consumer behavior.
In the interest of full disclosure, I am an Aetna employee. However, my writing is based on my experience and opinions in data analysis; not any company-specific information or “insider knowledge”. This is meant to be a general overview, not speculation on any agreements or plans between Aetna and Apple.
We now live thoroughly in the age of “Big Data”. To most analysts, that term doesn’t mean much in practice. There either are useful data or there aren’t. But there are four, very loosely defined criteria that determine whether an organization has Big Data*:
The one criteria that may be hardest to achieve, at least in healthcare, is Variety. Variety refers to data derived from different sources. In the context of a health insurer, Variety might include claims data from medical providers, EHR data, lab and pharmacy data, publicly available demographic data (e.g. Census data), and behavioral data (which is where I view the value of the Apple Watch and other devices in the healthcare space). The general goal of Data Variety is to collect as much different information as possible, in order to figure out which of those factors drives an outcome. And this is why finding sources of behavioral data is so important to population health applications.
When trying to predict health events, there’s a problem with most of those sources I listed above: they’re largely outcomes. Hospital visits, costs, even lab results, are the results of underlying conditions and events. The actual drivers of health (i.e. behaviors — what you eat, how much you exercise & how well you sleep) are vastly more difficult to derive. Consequently, modeling healthcare data isn’t so different from quantitative stock trading: essentially “backtesting” with a few controlling covariates.

