Nexus · Human Foundation Model
Nexus is the Human Foundation Model.
Mainstream world models predict the physical environment. Nexus predicts the person: their state, intent and boundaries. It is what makes an android able to live and work alongside people.
Human Foundation Model architecture
Nexus models the whole of humanity.
Built on a complete view of human health and potential, Nexus forms cross-modal, cross-context and long-horizon representations of people — a general foundation that can improve continuously.
NexCap multimodal inputs

HFM
Unified human-state model
Temporal modelling and prediction engine
Spatiotemporal encoder plus state-space model, multimodal fusion Transformer, long-sequence prediction and uncertainty modelling.
Unified human-state representation
Pose / contact / load / balance · Emotion / arousal / fatigue / comfort · Intent / social state / cognitive state
Six core predictions
Body-state prediction
Pose, balance, load, motion trend
Affect-change prediction
Emotional intensity and arousal shift
Behaviour prediction
Next action and behavioural path
Intent inference
Approach, refusal, request, collaboration
Comfort and pain boundary
Comfort boundaries and pain thresholds
Social-response prediction
Social distance, acceptance, mode of reply
Data foundation
- 01NexCap full-body multimodal capture
- 02Human Biofidelity Dataset
- 03High-quality annotation and validation
- 04Nexus training and evaluation
Application loop
- 01
Perceptual input
NexCap real-time multimodal data
- 02
World-model inference
State understanding and future prediction
- 03
Decision and generation
Motion planning / language generation
- 04
Android execution
Natural interaction and motion output
- 05
Feedback and learning
Outcome evaluation and model update
Core value
- Predictive interaction
- From executing actions to predicting what comes next — the difference between reacting and understanding.
- Safety
- Predicting contact comfort and risk raises both safety and controllability.
- Human likeness
- Behaviour, affect and social response land closer to a real person.
- Continuous evolution
- A closed data loop keeps the model improving and individualising.
Turning multimodal human signals into one predictable, complete state — so machines can finally understand and anticipate people.
NexCap data foundation
The whole human state, on one map.
NexCap records 200+ synchronised channels across more than 60 modular modalities. Nexus aligns every signal to one clock and coordinate frame, then learns their relationships as a shared human-state representation.
- Vision & motion
- Optical mocap / depth / gaze / skeleton / IMU / velocity
- Force & contact
- Six-axis force / pressure array / force plate / torque
- Muscle & neural
- sEMG / activation / muscle synergy / neural stimulation
- Physiology & affect
- ECG / HRV / respiration / EDA / SpO₂ / temperature / arousal
- Voice & audition
- Speech features / microphone array / source localisation
- Cognition & behaviour
- Gaze / attention / cognitive load / response time / social context
Nexus is trained on data nobody else has
The model is only as good as the capture system behind it — which is why NexCap is a product and not an internal tool.
