Part 4 · Why Biohacking Is Hard: The Three Challenges
To show a biohack works it must clear three hurdles: real evidence behind it, trustworthy measurement, and a result untangled from an individual body.
Part 4 of The Evidence-Based Guide to Biohacking. Part 1 asks whether biohacking is evidence-based. Part 2 maps the five main types of biohacking. Part 3 asks how far biohacking can actually be personalized.
Biohacking has always promised more than it could prove, and the reason is difficulty, not mystery. To show that a practice genuinely works, it has to clear three hurdles. There must be evidence behind it. We must be able to measure its effect in a trustworthy way. And finally, we must untangle that effect from everything else happening in a body where no two people are alike.
Each hurdle is real. Together, they explain why rigorous personalized biohacking is so difficult, and where the work now lies.
The three challenges in one sentence
Evidence tells us whether an intervention may work. Trustworthy measurement tells us whether something changed and whether we can trust the way that change was measured. Personalization asks whether we can isolate what worked for one individual. These are three different problems, and solving one does not automatically solve the next. Together, they form the Astrela Framework for Evidence-Based Biohacking.
Challenge one: is there evidence?
The first question is whether a practice works at all. As we saw in Part 1 of this guide, biohacking runs from well-evidenced fundamentals to speculative gadgets. Some methods have strong science behind them. Many have only testimonials and a good story. This distinction matters because a personal experiment cannot rescue an intervention with no credible evidence behind it.
If someone sleeps better after taking a new supplement, that observation may be real. But it does not automatically prove the supplement caused the improvement. Expectation, natural variation and other changes may all contribute. The strongest starting point is therefore population evidence: clinical trials, systematic reviews and other well-designed research. Population evidence does not tell us exactly what will happen to every individual, but it tells us where it is reasonable to begin. Separating the evidence-based from the merely popular is the first hurdle.
Challenge two: can we measure the effect in a trustworthy way?
Before we can trust a measurement, we have to decide what we are measuring. This sounds obvious, yet it is the step most often skipped. What exactly are we optimizing for? What is the number at the end that tells us the intervention worked? Until that is defined, “it’s working” is a feeling, not a finding.
There are broadly three things people measure, and each answers a different question.
A composite health or wellness score. Many apps and devices now condense everything into a single figure: a readiness score, a sleep score, a longevity or “body age” number. The appeal is obvious. One number is easy to follow and easy to feel good about when it rises. The difficulty is that most of these scores are proprietary. We rarely know which inputs they use, how they are weighted, or whether the score has ever been validated against a real health outcome. A number that goes up is satisfying. It is not the same as knowing our health improved.
A subjective self-rating. The simplest measure is to ask how we feel and mark it, say, from one to ten. This is not naive. Energy, pain, mood and sleep quality are real, and in serious research many of the most important outcomes are patient-reported precisely because the person’s experience is what matters. The weakness is that a self-rating is highly sensitive to expectation, memory, a particularly good week or the simple wish to have spent our money well. It captures the thing we care about most and is the easiest of all to fool.
Specific biomarkers. We can also measure something concrete: resting heart rate, heart rate variability, blood glucose, blood pressure, a blood lipid panel, an inflammatory marker. These are more objective, and this is often where the most useful signal lives. But two problems follow. First, a biomarker that moves is not always a health outcome that matters. A number can shift without our life getting longer or better, and mistaking a moving marker for a meaningful result is one of the oldest traps in medicine. Second, the measurement itself has to be trustworthy, and that is where the real difficulty begins.
Choosing what to measure is only half the problem. Once we have picked an outcome, especially an objective one, we still have to trust the way it is captured. Here three problems stack on top of each other.
First, the tools have to be accurate. Consumer wearables are improving quickly, but accuracy varies considerably depending on what is being measured. A 2024 living umbrella review examined 24 systematic reviews covering 11 biometric outcomes. Heart rate measurements generally performed relatively well, while errors were substantially greater for some other measures, including energy expenditure and estimates of aerobic capacity.¹ The review also highlighted a larger problem: the available validation studies represented only around 3.5 percent of the evaluations that would be needed to comprehensively assess the biometric measures offered across the devices studied.¹ In other words, a device may measure one thing reasonably well and another poorly. A wearable is not simply accurate or inaccurate. Accuracy is specific to the device, the metric and the context.
Second, the data has to be standardized. Different devices use different sensors, algorithms and sampling methods. The number called sleep score on one platform is not necessarily calculated in the same way as the sleep score on another. This makes comparison difficult, particularly when companies use proprietary algorithms that change over time. A number can be precise within one system and still be difficult to compare with another. If we want to read biological data across devices and over many years, we need the numbers to speak the same language. Often, they do not.
Third, a wellness number is not automatically a medical number. There is an important difference between a product that helps you track general wellness and a medical device intended to diagnose, monitor or guide the management of disease. In January 2026, the United States Food and Drug Administration updated its guidance for low-risk general wellness products. The guidance distinguishes general wellness functions from products making medical claims.² The distinction is visible in glucose monitoring: the FDA has explicitly warned that it has not authorized, cleared or approved any smartwatch or smart ring intended to independently measure or estimate blood glucose.³ That does not make consumer wearable data useless. It means we have to be precise about what a number can support. A number on your wrist and a number used to make a clinical decision do not automatically carry the same evidentiary weight.
So the second challenge is not simply whether we can measure something. It is whether we are measuring the right thing, and whether we can trust the way we measured it.
Challenge three: the individual equation
Suppose a method is evidence-based. Its effect is measurable and trustworthy. One hurdle still remains, and it is the hardest.
A human being is not a controlled experiment. You cannot live the same week twice, once with a change and once without. And because every person is different, no one else is a perfect stand-in for you either. Many systems interact at once. Sleep affects hunger. Stress affects sleep. Exercise affects glucose. Hormones may affect appetite, energy and weight. Any medications or treatments a person is taking act on these same systems, sometimes in ways that mask or mimic the effect we are trying to observe. And none of these systems holds still.
Formal N-of-1 trials try to address part of this problem through repeated periods, crossover designs and sometimes randomization, allowing an individual to be compared against their own responses over time.⁴ But everyday life is much harder to control. Isolating the effect of one change inside a moving, interconnected biological system is a genuinely difficult problem. It is a problem in mathematics and data science as much as in biology. And it may be the challenge that most defines the future of personalized health and personalized longevity.
Where artificial intelligence fits
The individual equation is not only a biological problem. It is a data and mathematics problem, and this is where artificial intelligence becomes interesting. Personalizing health means holding a very large number of variables together at once: sleep, activity, glucose, hormones, nutrition, stress, medications and treatments and more, all interacting and changing over time. For a long time this was simply too complex to compute. A clinician cannot hold it all in their head, and older tools could not model that many moving, interconnected factors.
Modern AI can begin to. It can handle very complex, multivariable relationships, the kind that defeated earlier approaches, and make sense of many signals together. Used well, AI could be what finally makes personalized biohacking tractable, turning scattered data into recommendations that actually fit one particular person.
But AI does not remove the first two challenges. If the evidence is weak, AI may simply give us a more sophisticated explanation of a weak claim. If the measurement is inaccurate, AI receives inaccurate data. AI can improve interpretation. It cannot manufacture biological truth from poor evidence or poor measurement. It is a powerful tool for the third challenge, not a way around the first two.
Why this matters
None of these hurdles is a reason to dismiss biohacking. They are the reason to take it seriously as a data and technology problem rather than only a lifestyle trend. The first challenge asks us to separate evidence from hype. The second asks whether we can measure an effect in a trustworthy way. The third asks whether we can isolate cause inside one complex, ever-changing person.
Clearing these hurdles requires better evidence, honest measurement, validated tools, more interoperable data and models capable of handling multivariable relationships over time. This is why we treat biohacking at Astrela as a data discipline, and why the harder parts call for building, not just advice.
What to take from this
This guide opened with a question: not whether biohacking works, but whether a given practice can be trusted, measured and measured well. The three challenges are that question unfolded. Can we trust it? That is the evidence problem. Can we measure its effect in a trustworthy way? That is the measurement problem. Can we isolate what worked inside one individual? That is the personalization problem, or what I think of as the individual equation.
Biohacking is not hard because the idea is wrong. It is hard because proving anything about one living person means solving this whole stack of problems at once. Naming the hurdles is the first step to clearing them. And clearing them is how biohacking grows from a promise into a practice we can trust.
A closing thought for the whole guide
For me, meeting these challenges is not abstract. It is a personal mission. I am driven not only by the wish to understand and improve my own health, measuring, adjusting and learning on myself, but by a larger aim: to help millions of people feel better in body and mind by uniting the best of modern data with the wisdom of ancient traditions.
That is why I want to close the guide here, with how old this instinct really is. The Rambam, Maimonides, a working physician in the twelfth century, taught in the Mishneh Torah that maintaining a healthy body is “among the ways of God,” and that a person should avoid what harms the body and become accustomed to what strengthens it.⁵ A different philosophical language appears in the Lurianic Kabbalistic tradition. The Ari’s system developed the ideas of tikkun and birur, a process of restoration and refinement associated with the elevation of dispersed sparks. Later Lurianic teaching applied this language explicitly to ordinary physical acts. Tanya, Chapter 7, for example, describes eating and drinking with a higher purpose as a process through which the vitality within the physical act can be elevated.⁶
This is not clinical evidence, and we do not offer it as proof of any protocol. But philosophically, the idea feels surprisingly relevant. How we act in the physical world matters. Ordinary physical actions can carry attention, intention and meaning. It is a philosophical lineage of a simple idea that modern health science approaches through very different methods: the body is worth observing, tending and respecting, deliberately and every day. The tools have changed. The responsibility has not.
This article is educational and is not a substitute for individual medical advice. It describes practices for general understanding only. Astrela does not recommend, endorse or advise against any specific practice, product or intervention. Before starting anything new, especially anything invasive or unregulated, speak with a qualified clinician who knows your situation.
Frequently asked
Why is it so hard to prove biohacking works?
Because proving anything about a single living person means clearing three hurdles. The practice has to be supported by evidence. Its effect has to be measured in a trustworthy way. And the effect has to be isolated inside a body where many systems interact and no two people are exactly the same. Each hurdle solves a different problem. The last one, understanding cause in one individual, may be the hardest.
Are consumer wearables accurate enough to rely on?
Sometimes, and it depends heavily on the device and the measure. Heart rate often performs relatively well, while measures such as energy expenditure and estimates of aerobic capacity can show much larger errors. A wearable is not simply accurate or inaccurate. Accuracy depends on the device, the metric and the context. Wearables can be useful for spotting trends. That does not automatically make every number suitable for clinical decision-making.
What is the hardest problem in personalized health?
Isolating cause in a single person. In a clinical trial, researchers can compare groups and use controls. In everyday life, you cannot live the same week twice, once with an intervention and once without it. Many biological systems interact at the same time and none stay constant. That makes proving that one specific change caused one specific result extremely difficult.
References
- Doherty C, et al. Keeping Pace with Wearables: A Living Umbrella Review of Systematic Reviews Evaluating the Accuracy of Consumer Wearable Technologies in Health Measurement · Sports Medicine, 2024;54:2907-2926
- General Wellness: Policy for Low-Risk Devices (Final Guidance, January 2026) · United States Food and Drug Administration
- Do Not Use Smartwatches or Smart Rings to Measure Blood Glucose Levels (FDA Safety Communication) · United States Food and Drug Administration, 2024
- Lillie EO, et al. The N-of-1 Clinical Trial: The Ultimate Strategy for Individualizing Medicine? · Personalized Medicine, 2011
- Maimonides. Mishneh Torah, Human Dispositions (Hilchot De'ot) 4:1 · Via Sefaria
- Tanya, Likutei Amarim, Chapter 7, on the elevation of the vitality in food and drink when used for a higher purpose; within the later Lurianic tradition of birur and tikkun · R. Schneur Zalman of Liadi; Lurianic and Hasidic Kabbalah