A Multi-Layer Architecture for Metabolic State Inference

A Multi-Layer Architecture for Metabolic State Inference

Scientific Framework and Initial Methodological Verification.

This paper sets out the scientific framework and an initial methodological verification. No clinical validation has been conducted.

Reviewed by the Platos Health Scientific Advisory Board: Hanno Pijl — Founding Science Advisor, Platos Health; Endocrinology, Leiden University Medical Center. Itopa Jimoh — Founding Science Advisor, Platos Health; Researcher, University of Alabama

Abstract

Standard laboratory panels sample metabolic biomarkers at intervals of months to years, producing point estimates that are informative but temporally sparse. Between measurements, sustained behavioural patterns and physiologic signals continue to shape underlying metabolic state, yet remain largely unmodelled in current preventive care. This paper describes the scientific framework underpinning Platos, a metabolic intelligence platform that infers the direction of continuous cardiometabolic state through a multi-layer architecture in which layer assignment follows the biological latency of the variable. We describe a four-stage verification methodology and report an initial methodological verification using nine synthetic clinical archetypes spanning European, South Asian, Sub-Saharan African, and mixed-heritage profiles. Because these synthetic archetypes are generated from the same physiological literature that underlies the engine, this exercise tests internal consistency and clinical plausibility rather than independent ground truth. The engine's directional inference was concordant with independent expert reasoning across the primary metabolic markers (glucose, insulin, lipids, liver) and less concordant on the inflammation and stress markers, which we identify as refinement priorities. No clinical validation has been conducted; verification at this stage is simulation-based, and empirical validation against real user data is the explicit next stage.

1. Introduction

Standard laboratory panels remain the dominant approach to cardiometabolic assessment. They are well-standardised, clinically actionable, and widely available. They are also temporally sparse: a patient with a lipid panel measured annually receives one snapshot per year. Between panels, dietary patterns, physical activity, sleep quality, and psychological stress load — all of which materially influence the trajectory of the measured biomarkers — continue to accumulate.

This measurement gap is well-recognised in preventive cardiovascular and metabolic care¹⁻³. What has been less well developed is a system-level framework capable of reasoning about the physiologic state between measurements, using the continuous data now available from wearables, self-report, and periodic device-derived measurements. Platos is a metabolic intelligence platform that fills this gap. It infers the direction of metabolic change between measurements rather than predicting biomarker values; the distinction matters, because prediction implies a claim about a future measured value, whereas the system makes claims about current latent state. This paper describes the scientific framework of the system and reports an initial methodological verification.

2. The scientific framework

Platos operates on a multi-layer taxonomy of biological variables. Each layer represents a class of variables sharing a characteristic biological latency — the timescale over which the variable meaningfully changes. The layers span behavioural inputs (levers under user control), short-latency physiologic signals (minutes to days), integrated biomarker states (weeks to months), sustained organ-level phenotypes and pre-states (months to years), and formal clinical states with established diagnostic thresholds (years to decades).

The layers are populated from data sources that vary by layer. Behavioural inputs are self-reported by the user or, where possible, inferred from wearable and connected-device signals. Short-latency physiologic signals are derived from wearables and connected devices, and — where the user has uploaded them — from short-window laboratory measurements. Integrated biomarker states are anchored by uploaded laboratory panels where available, and estimated between panels from sustained behavioural and physiologic patterns. Sustained phenotypes and clinical states provide the wider biological context in which the biomarker layer sits.

The biomarker layer

The biomarker layer is where Platos produces its outputs. The system generates rolling directional trend estimates across seven metabolic markers, each treated as a neutral, observable variable — not a diagnosis — and reported as a trend rather than a fixed value.

These seven markers span systems that conventional laboratory panels typically distribute across separate medical specialties — glucose and insulin handling, lipid transport, systemic inflammation, hepatic workload, urate balance, and adrenal rhythm. Platos models them together because the underlying physiology is interconnected: the markers influence one another, and a change in one frequently propagates to others. Each marker relationship is grounded in peer-reviewed literature and graded internally by the strength of its supporting evidence, from guideline-level consensus and meta-analyses at the highest grade down to mechanistic evidence and expert opinion at the lowest. Evidence grade directly constrains the confidence the system can assign, so that a relationship resting on weaker evidence cannot reach the highest confidence band regardless of how complete the underlying data is.

The seven markers were selected on three criteria: clinical importance in cardiometabolic disease, the strength of the supporting literature linking behavioural and physiologic inputs to each marker, and coverage of the interconnected systems that drive the majority of preventable chronic disease. Together they span the systems conventional panels distribute across separate specialties, while remaining few enough that each can be grounded in defensible evidence and reasoned about in relation to the others. Markers with weaker mechanistic links to modifiable behaviour, or with insufficient literature to support directional inference, fall outside the current scope.

Directional inference

Rather than predicting point estimates of biomarker values, Platos infers the direction of each biomarker state: whether it is stable, showing early emerging pressure, or under sustained pressure. Three directional bands are used, each paired with an explicit confidence assessment that reflects data availability, the recency of any biomarker anchoring measurements, and the internal consistency of the contributing signals.

The choice of directional bands over point estimates reflects the preventive orientation of the platform. Cardiometabolic patterns develop continuously; behavioural drivers accumulate influence over weeks and months before any diagnostic threshold is crossed. A point-estimate biomarker tells the user what the number is today. A directional band tells the user which way the trajectory is going and whether the accumulated behavioural pattern is protective or corrosive. This framing supports intervention: users can respond to an emerging directional shift by modifying the behaviours that produced it, rather than waiting for a threshold to be crossed.

Confidence is a first-class output alongside the directional inference, so that lower-data users are not presented with false precision. Confidence is computed on principled grounds rather than as a single blended score: it reflects the completeness of the inputs a marker requires, the length and internal consistency of the observation window, whether an active confounder is distorting the signal, and the strength of the clinical evidence behind the pathway being used. These principles combine non-compensatorily — a weakness in any one cannot be offset by strength in another — so that the reported confidence reflects the weakest link in the inference chain rather than an average that could conceal a critical gap. The specific computation is proprietary; the principles are described here so the confidence rating can be interpreted.

3. Verification methodology

We use a four-stage verification methodology: scientific grounding, internal verification, internal consistency verification against pre-specified expected bands, and independent expert review. The exercise tests whether the framework produces sound directional inference on defined inputs. It does not establish clinical performance, which would require prospective real-patient outcomes. No clinical validation has been conducted to date.

The synthetic archetypes, and the expected directional bands used to check the engine against them, are constructed internally from the same physiological literature that underlies the engine's own reasoning. The chain runs from literature to engine to synthetic archetype to expected band to comparison. The one point external to this loop is the independent expert review (Stage 4), in which a reviewer generates directional expectations from the archetype inputs alone, blinded to engine outputs. Absent that stage, the exercise would verify only internal consistency. It is not, and does not claim to be, validation against independent biological ground truth, which requires real user data — the subject of the verification roadmap in Section 6.

Stages 1 and 2 — Scientific grounding and internal verification

The biological plausibility of each biomarker–behaviour relationship is anchored to peer-reviewed literature and graded by evidence strength; the mechanisms underlying each biomarker state are documented in internal source-of-truth documents available to institutional partners under confidentiality. The system is then tested against synthetically generated input profiles designed to probe engine behaviour: gradient responses to single-driver elevation, correct suppression under confounder gate activation (acute illness, recent medication change), and appropriate handling of missing versus null inputs. Outcomes of both stages are documented in the internal verification log; identified issues were resolved or scoped as future workstreams prior to the archetype review.

Stage 3 — Internal consistency verification

Nine synthetic clinical archetypes were constructed internally to represent realistic user profiles across five product goal categories and across ethnic profiles (European, South Asian representing UK ICP populations, Sub-Saharan and West African representing Nigerian ICP populations, and mixed-heritage). The archetypes are synthetic profiles, not real patient data; their inputs represent clinically plausible combinations across the layers, subjected to sustained multi-week behavioural stress scenarios that mirror established metabolic challenges. Four of the nine archetypes probe specific engine behaviours: sustained improvement trajectory reading, recent medication change reshape, acute illness confounder handling, and anti-pathologising guardrails.

An expected directional band was pre-specified internally for each marker in each archetype. Engine outputs were then verified against these expected bands and mismatches flagged. This stage is an internal consistency check: both the expected bands and the engine are grounded in the same evidence base, so agreement here confirms the engine behaves as intended, not that the intended behaviour matches independent clinical ground truth.

Stage 4 — Independent expert review

An independent expert scientific reviewer generated directional predictions and confidence estimates for each of the seven markers across each archetype, working from the archetype inputs alone and blinded to engine outputs and to the internal verification. The independent predictions were then compared against engine outputs, and divergences examined for the underlying cause. The blinding — the reviewer reasons from inputs, not from the engine's output — is what protects the independence of this review, and it is against this comparison that the concordance in Section 4 is reported.

What counts as concordance

A marker-level comparison is treated as concordant when the inferred direction aligns with the independent expert's expectation, the confidence band reflects the data conditions simulated, and no gating rule is bypassed. Partial or divergent results are treated as findings to investigate, not as failures.

4. Findings

This section reports the concordance between the engine and the independent expert review (Stage 4), which is the exercise's only independent comparison. Findings are organised as the aggregate directional concordance and the confidence behaviour observed. Specific engine outputs and internal scores are not reported. The figures below summarise a small exercise — nine personas, sixty-three marker-level comparisons — and are reported without confidence intervals, inter-rater statistics, or sensitivity analyses. They should be read as indicators of directional concordance and as a map of where the reasoning aligns well and where it needs refinement, not as measured performance rates.

Aggregate directional concordance

Nine synthetic personas were compared across the seven markers, producing sixty-three marker-level comparisons between independent expert expectation and engine inference. The exercise identified areas where the reasoning architecture was strongly concordant with expert expectation — glucose, insulin, lipids, and liver — and areas requiring further refinement — inflammation and stress. Uric acid sat between the two groups. The primary metabolic markers, which rest on the strongest evidence base and the most direct behaviour-to-biomarker pathways, showed the highest concordance; the two markers with the most complex, non-linear physiology showed the lowest. The pattern locates the boundary of the current reasoning rather than presenting a single headline rate.

Confidence behaviour

The independent expert and the engine each produced confidence ratings independently from the same aggregate persona summary, and the difference between them is informative. The expert's confidence is a holistic judgment drawing on pattern recognition, spread across High and Medium. The engine's confidence is a principled computation, and the pattern is specific: across seven of the nine personas, engine confidence settled at Medium for every marker, dropping to Low only in the two personas where a confounder gate was active and capped the output.

The engine returns a Medium confidence when working from a single aggregate snapshot, and drops to Low specifically when a confounder gate fires. The gap between engine and expert confidence reflects a difference in testing conditions, not reasoning quality: the engine's confidence model is designed to operate over a continuous longitudinal data stream, and a single aggregate snapshot represents only a fraction of the observation window the model expects, which mechanically constrains it toward Medium. The expert, reasoning holistically, is not bound by that temporal constraint. Convergence between the two is expected once both work from the same kind of continuous data.

5. Scope limitations

No clinical or laboratory validation has been conducted yet. Everything reported here comes from simulation, not real patients or verified laboratory draws, and closing this gap is the priority before stronger claims are made. Simulation approximates but does not replace real physiology: the archetypes are built from the same literature the reasoning encodes, which bounds how independent this validation can be. The nine-persona sample is sufficient to confirm directional behaviour and catch obvious faults, but not to treat the 65.1% figure as a stable accuracy rate; that figure will be updated as more personas are added.

The independent expert and the engine both worked from a single aggregate data snapshot per persona, not a simulated day-by-day journey. The confidence model in particular is designed to operate over continuous longitudinal data, and a single snapshot does not exercise it on the terms it was built for.

Ethnicity is represented in the synthetic archetypes but is not yet implemented as a variable in the engine. The archetypes seen by the independent reviewer include ethnicity as part of the clinical picture, but the engine does not currently use ethnicity to modify its outputs — a known divergence between the reviewer's inputs and the current engine. Ethnicity-specific priors, together with medication-change handling and positive modelling of illness-mediated mechanisms, are scoped future workstreams.

Two specific modelling limitations account for much of the lower concordance on the stress and inflammation markers, and both are scoped for post-beta refinement. The stress marker currently models adrenal activation on a single progressive scale, higher inputs producing a higher score. Hypothalamic-pituitary-adrenal physiology is biphasic: sustained stress can progress from a high-output alarm state to a low-output exhaustion state in which cortisol is suppressed. A single-scale model can therefore read late-stage burnout, where measured cortisol is low, as low risk — the opposite of the clinical interpretation. A future biphasic model will distinguish high-output from low-output stress states.

The inflammation marker is currently weighted heavily toward adiposity-driven inflammatory signalling. Systemic inflammation can, however, be elevated in lean, active individuals through pathways not yet fully modelled — for example gut-permeability effects associated with ultra-processed food or alcohol intake. The current model can under-read inflammation in these individuals. Broadening the inflammatory inputs beyond adiposity is a scoped refinement.

The system is non-diagnostic by design. It estimates lifestyle-driven trends; it does not detect disease or guide treatment. Existing conditions such as prediabetes, hypertension, or fatty liver are factored into the reasoning as baseline priors — they adjust how strongly a marker responds to behaviour, so the output reflects the condition rather than ignoring it.

6. Verification roadmap

The verification described in this paper is the first stage of a longer programme that moves progressively from internal plausibility toward independent clinical evidence. We set it out explicitly so the boundary of the present claims is clear.

Phase 1 — retrospective comparison. Compare engine inference against verified laboratory results already held for consenting users, over real observation windows rather than synthetic snapshots. This is the first test against independent ground truth.

Phase 2 — prospective longitudinal users. Follow users forward with continuous behavioural and physiologic data and scheduled laboratory draws, evaluating whether the engine's directional inference anticipates measured biomarker movement. This is also where the confidence model is exercised on the continuous data stream it was designed for.

Phase 3 — external clinical collaborators. Independent clinical partners evaluate the inference against their own patient data, removing any residual dependence on Platos-generated inputs.

Phase 4 — outcome association. Assess whether sustained directional inference is associated with downstream clinical outcomes over longer horizons. This report will be revised as each phase produces evidence.

7. Advisory review and endorsement

The scientific framework and verification described in this paper have been reviewed and endorsed by the Platos Health Scientific Advisory Board. Continuous empirical validation — ongoing calibration of engine inference against user-uploaded biomarker measurements over time — is a design feature of the system and not a one-time exercise. Future iterations of this paper will update the findings as empirical validation accumulates.

Scientific Advisory Board

Hanno Pijl — Founding Science Advisor, Platos Health. Professor of Endocrinology at Leiden University Medical Center (LUMC), the Netherlands. Role in this work: independent expert review.

Itopa Jimoh — Founding Science Advisor, Platos Health; Researcher at the University of Alabama, with expertise in public health and community medicine and a particular focus on Sub-Saharan African cardiometabolic health patterns. Role in this work: co-development of the scientific framework and internal verification of engine outputs against the pre-specified expected bands.

References

The scientific framework draws on the peer-reviewed evidence base that grounds each biomarker–behaviour relationship in the system. A representative selection is given below; the complete pathway-by-pathway evidence grading is maintained internally and available to institutional partners under confidentiality.

1. American Diabetes Association. Standards of Care in Diabetes—2026. Diabetes Care 2026;49(Suppl).

2. World Health Organization. Guideline: Sugars intake for adults and children. Geneva: WHO.

3. World Health Organization. Guidelines on Physical Activity and Sedentary Behaviour. Geneva: WHO.

4. International Diabetes Federation. The IDF consensus worldwide definition of the metabolic syndrome.

5. Reynolds A, Mann J, Cummings J, et al. Carbohydrate quality and human health: a series of systematic reviews and meta-analyses. Lancet 2019;393:434–445.

6. Tasali E, Leproult R, Ehrmann DA, Van Cauter E. Slow-wave sleep and the risk of type 2 diabetes in humans. PNAS 2008;105:1044–1049.

7. Cappuccio FP, D’Elia L, Strazzullo P, Miller MA. Quantity and quality of sleep and incidence of type 2 diabetes: a systematic review and meta-analysis. Diabetes Care 2010;33:414–420.

8. Buxton OM, Pavlova M, Reid EW, et al. Sleep restriction for 1 week reduces insulin sensitivity in healthy men. Diabetes 2010;59:2126–2133.

9. Després JP. Body fat distribution and risk of cardiovascular disease: an update. Circulation 2012;126:1301–1313.

10. Meal timing, circadian rhythm and metabolic outcomes: evidence on eating window and glycaemic regulation. Int J Obes 2015.

11. Alcohol intake and systemic inflammatory markers: cohort evidence. 2010.

12. Purine intake, hydration and serum urate: epidemiologic and mechanistic evidence.

13. Hepatic enzyme markers and metabolic dysfunction-associated steatotic liver disease (MASLD): Rinella ME et al. multi-society nomenclature. J Hepatol 2023.

14. Physical activity, autonomic regulation and hypothalamic-pituitary-adrenal axis function: mechanistic evidence.

15. Principles of non-compensatory scoring and confidence calibration in multi-input inference systems.

Platos is a lifestyle and well-being product. It is not a medical device and does not diagnose, treat, monitor, or predict any disease. This paper describes a scientific framework and an initial methodological verification; it does not report clinical validation. Findings derive from synthetic archetypes, not patients. Version 0.9 · August 2026 · methodology in development, subject to revision. © Platos Health.

Certain methods and systems described herein are the subject of pending patent applications.