Full-body muscle-actuated humanoid control benchmark
MSK-Bench: Benchmarking Full-Body Musculoskeletal Motor Control Across Tasks, Control Paradigms, and Physiological Metrics MSK-Bench: Benchmarking Full-Body Musculoskeletal Motor Control Across Tasks, Control Paradigms, and Physiological Metrics
Abstract
Musculoskeletal (MSK) humanoids provide a physiologically grounded embodiment for studying full-body motor control, but their high-dimensional muscle actuation, delayed activation dynamics, and redundant muscle-tendon structures make learning substantially harder than torque-driven humanoid control. Existing MSK benchmarks remain fragmented across gait, prosthetics, dexterous hands, or challenge-specific tracks, leaving full-body muscle-actuated control insufficiently evaluated under standardized tasks, methods, and metrics. We introduce MSK-Bench, a benchmark of 22 full-body motor-control tasks organized into three progressively challenging categories: postural stabilization, common locomotor behaviors, and contact-rich environmental interaction. Under unified task protocols and robustness perturbations, MSK-Bench evaluates 5 representative control paradigms, including reward-based RL, agentic reward tuning, latent-action RL, imitation-prior control, and residual adaptation over imitation priors. Beyond task success and reward, MSK-Bench further reports robustness analysis and physiology-oriented diagnostics, including activation cost, joint smoothness, and EMG-envelope similarity. Our empirical study shows that embodiment-aware exploration and structured action representations improve task coverage in high-dimensional muscle spaces, imitation priors enhance reference-compatible stabilization and locomotion but degrade under contact-rich terrain mismatch, and residual adaptation can recover successful behaviors when fixed references fail. We further find that improved task success does not necessarily imply improved physiological agreement, highlighting the importance of evaluating task performance, robustness, and physiological behavior jointly. MSK-Bench provides a task-method-metric testbed for full-body muscle-actuated humanoid control.
| Benchmark / Framework | Embodiment | Environment | Task Design | Evaluation | Main Focus | ||||
|---|---|---|---|---|---|---|---|---|---|
| Actuation (Dim.) |
Full Body |
3D Scene / Terrain |
Contact Rich |
Task Hierarchy |
Tasks / Data Scale |
Physio. Eval. Metrics |
Control Families |
||
| Torque-Driven Humanoid Benchmarks | |||||||||
| HumanoidBench [87] | Torque (19/61) | ✓ | ✓ | ✓ | ✓ | 27 tasks | ✗ | 1 | Loco-manipulation |
| SkillBench / SkillBlender [49] | Torque (19) | ✓ | ● | ✓ | ✓ | 4 skills / 8 tasks | ✗ | 1 | Skill Blending |
| Mimicking-Bench [60] | Torque (19) | ✓ | ✓ | ✓ | ✗ | 6 tasks | ✗ | 1 | Scene Interaction |
| Musculoskeletal Benchmarks | |||||||||
| LocoMuJoCo [2] | Mixed | ✓ | ✗ | ✗ | ● | 12 envs / 27 tasks | ✗ | 2 | Locomotion Imitation |
| MyoSuite [12] | Muscle (≤ 80) | ✗ | ✗ | ✓ | ● | 204 tasks | ✗ | 1 | Dexterity and Agility |
| MyoDex [13] | Muscle (39) | ✗ | ✗ | ✓ | ✗ | 14 train / 34 eval. | ✗ | 1 | Hand Manipulation |
| MS-Human-700 / MsGym [126] | Muscle (700) | ✓ | ✗ | ● | ✗ | 3 tasks | ✓ | 1 | MSK Locomotion |
| MyoChallenge 2024 [107] | Muscle (≤ 80) | ● | ✓ | ✓ | ● | 2 tracks | ✗ | 1 | Prosthetics |
| MyoChallenge 2025 [15] | Muscle (≤ 80) | ● | ● | ✓ | ● | 2 tasks / 4 tracks | ✗ | 1 | Athletic Control |
| MSK-Bench (ours) | Muscle (416) | ✓ | ✓ | ✓ | ✓ | 22 tasks | ✓ | 5 | Full-body Terrain MSK |
Evaluation Protocol
One task suite, five control paradigms, seven metric families.
All 22 tasks use MuJoCo and the same 416-muscle elastic-tendon full-body model. Success requires task completion, survival, and safety under fixed task-specific criteria. Within each task, the four full-suite baselines share observations, rewards, termination conditions, and evaluation.
Control paradigms & study scope
- Reward-based RL: PPO, SAC, DepRL, and DynSyn-SAC; each is trained and evaluated on all 22 tasks.
- Agentic reward tuning: focused DepRL studies with bounded reward-coefficient updates using DeepSeek-V4-Flash and GPT-6 Astra.
- Latent-action RL: a focused study of an expert-trained, state-conditioned encoder–decoder; compression changes the command representation, not the executed muscle-action space.
- Imitation-prior control: pretrained MuscleMimic evaluated on stand, jump, walk, run, and stairs.
- Residual adaptation: a frozen imitation prior with bounded corrections and task-specific objective/timing changes on walk, run, and stairs.
Focused studies differ in task coverage and may change reward coefficients or introduce reference objectives. They are reported separately from the full-suite ranking.
What the seven metric families measure
- Success Rate: successful evaluation episodes in the restored training environment, including native noise. Family and full-suite means weight tasks equally.
- Cumulative Reward: task-specific evaluation return, never pooled across tasks.
- Peak-Efficiency Steps: the evaluated training step attaining the highest mean return.
- Perturbation Robustness: normalized trapezoidal area under success versus perturbation scale, averaged equally over tasks and action, observation, and muscle-dynamics sweeps.
- Activation Cost: mean squared muscle activation, used as an effort proxy.
- Joint Smoothness: mean squared angular jerk from finite differences of joint velocities, displayed logarithmically.
- EMG-Envelope Similarity: mean per-muscle Pearson correlation after 101-point cycle resampling, per-muscle min–max normalization, cycle averaging, and correlation-maximizing cyclic alignment.
Action/dynamics scales: 0, 0.05, 0.10, 0.15, 0.20; observation scales: 0, 0.02, 0.05, 0.08, 0.10. Activation cost and smoothness are interpreted only for successful, behaviorally comparable policies. EMG correlations assess phase-optimized waveform shape, not absolute amplitude or timing.
Task Families
From postural regulation to contact-rich environmental interaction.
Leaderboard
Benchmark Performance and Robustness Evaluation
22 tasks across three families, comparing DepRL, SAC, PPO, and DynSyn-SAC.
Train-env. SR retains native training noise; robustness columns report separate perturbation sweeps. See the paper appendix for evaluation details. Bold SR values mark the highest training-environment success rate within each task, including ties. Activation cost and smoothness should be interpreted for successful, behaviorally comparable policies.
Stabilization
6 tasks · 4 algorithmsScroll horizontally for all metrics and vertically for all tasks. Task and algorithm columns stay visible.
| Task | Algorithm | Activation Cost | Robustness (%) | Max Steps (× 107) | Train-env. SR (%) ↑ | log10 smooth (rad2/s6) | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Action | Observation | Dynamic | ||||||||||||||||||
| Action perturbation 0 | Action perturbation 0.05 | Action perturbation 0.1 | Action perturbation 0.15 | Action perturbation 0.2 | Observation perturbation 0 | Observation perturbation 0.02 | Observation perturbation 0.05 | Observation perturbation 0.08 | Observation perturbation 0.10 | Dynamic perturbation 0 | Dynamic perturbation 0.05 | Dynamic perturbation 0.1 | Dynamic perturbation 0.15 | Dynamic perturbation 0.2 | ||||||
| stand | DepRL | 0.4488 | 94 | 94 | 90 | 96 | 98 | 100 | 12 | 0 | 0 | 0 | 98 | 48 | 28 | 12 | 4 | 3.74 | 86 | 3.152 |
| SAC | 0.0484 | 100 | 100 | 100 | 100 | 96 | 100 | 0 | 0 | 0 | 0 | 100 | 100 | 100 | 100 | 100 | 7.13 | 100 | 1.875 | |
| PPO | 0.3691 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.97 | 0 | 5.479 | |
| DynSyn-SAC | 0.2859 | 52 | 42 | 56 | 46 | 36 | 46 | 46 | 22 | 30 | 14 | 40 | 32 | 40 | 30 | 30 | 5.44 | 52 | 2.104 | |
| powerlift | DepRL | 0.3356 | 68 | 68 | 64 | 56 | 52 | 62 | 62 | 38 | 4 | 0 | 66 | 70 | 52 | 42 | 38 | 4.14 | 38 | 2.631 |
| SAC | 0.2106 | 98 | 94 | 94 | 88 | 80 | 96 | 0 | 0 | 0 | 0 | 98 | 94 | 88 | 64 | 2 | 6.78 | 90 | 2.825 | |
| PPO | 0.3955 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.76 | 0 | 5.042 | |
| DynSyn-SAC | 0.3631 | 100 | 100 | 100 | 98 | 92 | 100 | 100 | 96 | 94 | 90 | 100 | 98 | 94 | 90 | 82 | 8.50 | 100 | 5.781 | |
| squat | DepRL | 0.3781 | 44 | 46 | 38 | 32 | 30 | 24 | 56 | 34 | 62 | 26 | 46 | 38 | 12 | 20 | 10 | 14.02 | 42 | 2.984 |
| SAC | 0.1184 | 0 | 0 | 0 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 4 | 0 | 0 | 0 | 0 | 4.79 | 0 | 1.708 | |
| PPO | 0.3455 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.01 | 0 | 5.390 | |
| DynSyn-SAC | 0.1955 | 48 | 38 | 32 | 26 | 22 | 50 | 44 | 38 | 36 | 30 | 48 | 42 | 32 | 26 | 18 | 7.00 | 42 | 6.503 | |
| sit | DepRL | 0.3327 | 96 | 100 | 100 | 98 | 92 | 92 | 100 | 100 | 100 | 94 | 100 | 92 | 96 | 100 | 90 | 4.88 | 96 | 2.872 |
| SAC | 0.1164 | 80 | 80 | 74 | 70 | 64 | 84 | 0 | 0 | 0 | 0 | 78 | 100 | 76 | 70 | 52 | 7.39 | 64 | 3.204 | |
| PPO | 0.4517 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.37 | 0 | 5.350 | |
| DynSyn-SAC | 0.3302 | 80 | 72 | 70 | 68 | 60 | 80 | 80 | 84 | 82 | 78 | 92 | 84 | 78 | 60 | 26 | 9.27 | 82 | 3.145 | |
| singlestand | DepRL | 0.3977 | 14 | 16 | 28 | 18 | 10 | 24 | 30 | 8 | 0 | 0 | 12 | 28 | 6 | 8 | 12 | 2.10 | 14 | 2.854 |
| SAC | 0.0849 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 4.33 | 0 | 1.954 | |
| PPO | 0.3981 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.84 | 0 | 5.436 | |
| DynSyn-SAC | 0.3845 | 24 | 28 | 24 | 20 | 18 | 26 | 22 | 16 | 0 | 0 | 22 | 18 | 8 | 0 | 0 | 10.00 | 24 | 2.931 | |
| balance | DepRL | 0.3775 | 60 | 40 | 26 | 20 | 14 | 56 | 0 | 0 | 0 | 0 | 46 | 12 | 36 | 22 | 18 | 8.86 | 26 | 2.822 |
| SAC | 0.1459 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 8.11 | 0 | 2.547 | |
| PPO | 0.3859 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.25 | 0 | 4.861 | |
| DynSyn-SAC | 0.2735 | 40 | 38 | 20 | 16 | 14 | 54 | 40 | 8 | 2 | 0 | 52 | 8 | 0 | 0 | 0 | 10.50 | 44 | 6.066 | |
Locomotion
6 tasks · 4 algorithmsScroll horizontally for all metrics and vertically for all tasks. Task and algorithm columns stay visible.
| Task | Algorithm | Activation Cost | Robustness (%) | Max Steps (× 107) | Train-env. SR (%) ↑ | log10 smooth (rad2/s6) | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Action | Observation | Dynamic | ||||||||||||||||||
| Action perturbation 0 | Action perturbation 0.05 | Action perturbation 0.1 | Action perturbation 0.15 | Action perturbation 0.2 | Observation perturbation 0 | Observation perturbation 0.02 | Observation perturbation 0.05 | Observation perturbation 0.08 | Observation perturbation 0.10 | Dynamic perturbation 0 | Dynamic perturbation 0.05 | Dynamic perturbation 0.1 | Dynamic perturbation 0.15 | Dynamic perturbation 0.2 | ||||||
| walk forward | DepRL | 0.4206 | 28 | 28 | 32 | 32 | 44 | 34 | 32 | 46 | 24 | 30 | 40 | 20 | 12 | 8 | 6 | 3.32 | 64 | 3.151 |
| SAC | 0.1427 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 9.88 | 0 | 3.262 | |
| PPO | 0.3940 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.42 | 0 | 5.558 | |
| DynSyn-SAC | 0.2771 | 82 | 88 | 68 | 60 | 48 | 82 | 82 | 80 | 76 | 76 | 82 | 76 | 76 | 64 | 8 | 5.00 | 76 | 3.417 | |
| run | DepRL | 0.3354 | 30 | 22 | 32 | 36 | 42 | 18 | 18 | 30 | 32 | 26 | 32 | 16 | 12 | 8 | 2 | 6.72 | 18 | 3.343 |
| SAC | 0.0940 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.13 | 0 | 2.854 | |
| PPO | 0.3696 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.16 | 0 | 5.524 | |
| DynSyn-SAC | 0.3221 | 96 | 94 | 84 | 70 | 68 | 96 | 94 | 90 | 86 | 80 | 96 | 88 | 2 | 2 | 0 | 10.00 | 96 | 3.435 | |
| walk turn | DepRL | 0.3385 | 46 | 60 | 62 | 66 | 68 | 34 | 40 | 58 | 40 | 26 | 48 | 46 | 40 | 36 | 36 | 6.38 | 42 | 3.057 |
| SAC | 0.1571 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.99 | 0 | 2.610 | |
| PPO | 0.3747 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.05 | 0 | 5.331 | |
| DynSyn-SAC | 0.3112 | 96 | 88 | 84 | 78 | 74 | 94 | 90 | 90 | 88 | 84 | 92 | 24 | 0 | 0 | 6 | 8.00 | 86 | 3.248 | |
| jump | DepRL | 0.3616 | 8 | 24 | 20 | 12 | 10 | 18 | 18 | 8 | 4 | 0 | 14 | 10 | 12 | 6 | 0 | 6.89 | 10 | 3.188 |
| SAC | 0.1342 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 3.92 | 0 | 2.853 | |
| PPO | 0.3880 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.07 | 0 | 5.640 | |
| DynSyn-SAC | 0.2700 | 70 | 40 | 8 | 0 | 0 | 62 | 54 | 50 | 48 | 42 | 42 | 30 | 2 | 2 | 0 | 8.00 | 54 | 3.449 | |
| sidestep | DepRL | 0.3888 | 40 | 38 | 62 | 56 | 32 | 34 | 12 | 2 | 0 | 0 | 24 | 22 | 18 | 10 | 6 | 4.38 | 28 | 3.349 |
| SAC | 0.1045 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.000032 | 0 | 3.166 | |
| PPO | 0.3729 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 4.92 | 0 | 5.408 | |
| DynSyn-SAC | 0.3630 | 74 | 66 | 54 | 42 | 30 | 74 | 78 | 74 | 78 | 54 | 64 | 72 | 60 | 40 | 40 | 5.58 | 76 | 3.223 | |
| crawl | DepRL | 0.3540 | 100 | 100 | 100 | 100 | 98 | 96 | 100 | 100 | 100 | 100 | 100 | 100 | 98 | 100 | 100 | 5.12 | 86 | 3.158 |
| SAC | 0.0338 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.87 | 0 | 2.516 | |
| PPO | 0.4036 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.34 | 0 | 5.674 | |
| DynSyn-SAC | 0.2766 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 10.00 | 100 | 3.444 | |
Interaction with Environment
10 tasks · 4 algorithmsScroll horizontally for all metrics and vertically for all tasks. Task and algorithm columns stay visible.
| Task | Algorithm | Activation Cost | Robustness (%) | Max Steps (× 107) | Train-env. SR (%) ↑ | log10 smooth (rad2/s6) | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Action | Observation | Dynamic | ||||||||||||||||||
| Action perturbation 0 | Action perturbation 0.05 | Action perturbation 0.1 | Action perturbation 0.15 | Action perturbation 0.2 | Observation perturbation 0 | Observation perturbation 0.02 | Observation perturbation 0.05 | Observation perturbation 0.08 | Observation perturbation 0.10 | Dynamic perturbation 0 | Dynamic perturbation 0.05 | Dynamic perturbation 0.1 | Dynamic perturbation 0.15 | Dynamic perturbation 0.2 | ||||||
| stair | DepRL | 0.3499 | 52 | 54 | 38 | 42 | 40 | 34 | 38 | 2 | 0 | 0 | 32 | 40 | 24 | 14 | 4 | 1.42 | 36 | 2.982 |
| SAC | 0.1993 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 6.98 | 0 | 2.459 | |
| PPO | 0.3854 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.57 | 0 | 5.303 | |
| DynSyn-SAC | 0.2759 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 5.92 | 0 | 2.854 | |
| catch | DepRL | 0.3247 | 80 | 80 | 76 | 76 | 84 | 62 | 48 | 80 | 72 | 46 | 78 | 40 | 42 | 38 | 56 | 7.23 | 28 | 2.919 |
| SAC | 0.1384 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.37 | 0 | 2.899 | |
| PPO | 0.3964 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 3.12 | 0 | 5.300 | |
| DynSyn-SAC | 0.3109 | 70 | 64 | 60 | 60 | 58 | 60 | 40 | 34 | 26 | 22 | 68 | 68 | 60 | 54 | 50 | 7.38 | 66 | 3.134 | |
| hurdle | DepRL | 0.3715 | 48 | 48 | 56 | 50 | 40 | 64 | 44 | 48 | 28 | 0 | 62 | 44 | 40 | 32 | 20 | 1.26 | 50 | 3.212 |
| SAC | 0.1852 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 4.12 | 0 | 2.816 | |
| PPO | 0.3668 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.46 | 0 | 5.398 | |
| DynSyn-SAC | 0.3717 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 5.50 | 0 | 7.102 | |
| slide | DepRL | 0.3330 | 100 | 22 | 22 | 16 | 12 | 100 | 16 | 4 | 4 | 0 | 100 | 8 | 14 | 4 | 6 | 11.32 | 100 | 3.106 |
| SAC | 0.0925 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 4.66 | 0 | 2.752 | |
| PPO | 0.3934 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 4.13 | 0 | 5.345 | |
| DynSyn-SAC | 0.2649 | 0 | 2 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2 | 2 | 0 | 0 | 3.83 | 0 | 2.864 | |
| steppingstones | DepRL | 0.3606 | 42 | 46 | 44 | 40 | 38 | 46 | 42 | 10 | 20 | 2 | 54 | 38 | 24 | 18 | 18 | 8.32 | 58 | 2.985 |
| SAC | 0.1647 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 7.33 | 0 | 2.140 | |
| PPO | 0.4020 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.27 | 0 | 5.211 | |
| DynSyn-SAC | 0.3569 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 10.00 | 0 | 6.764 | |
| polewalk | DepRL | 0.4080 | 22 | 38 | 16 | 40 | 20 | 42 | 26 | 6 | 0 | 0 | 30 | 26 | 20 | 16 | 16 | 5.14 | 12 | 2.876 |
| SAC | 0.1504 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 6.10 | 0 | 1.724 | |
| PPO | 0.4274 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.87 | 0 | 5.377 | |
| DynSyn-SAC | 0.2909 | 98 | 96 | 98 | 92 | 90 | 98 | 0 | 0 | 0 | 0 | 96 | 94 | 94 | 76 | 16 | 9.00 | 88 | 3.242 | |
| reach | DepRL | 0.3410 | 84 | 86 | 84 | 94 | 82 | 94 | 82 | 90 | 62 | 34 | 92 | 88 | 82 | 82 | 74 | 9.34 | 84 | 2.975 |
| SAC | 0.1705 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 5.85 | 0 | 2.677 | |
| PPO | 0.3965 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2.87 | 0 | 5.152 | |
| DynSyn-SAC | 0.2556 | 4 | 14 | 0 | 4 | 4 | 2 | 6 | 6 | 2 | 4 | 10 | 8 | 6 | 4 | 0 | 2.10 | 0 | 2.840 | |
| walkandsit | DepRL | 0.3648 | 30 | 26 | 28 | 14 | 10 | 28 | 36 | 32 | 18 | 10 | 22 | 28 | 22 | 14 | 12 | 2.68 | 16 | 3.365 |
| SAC | 0.1180 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 6.54 | 0 | 2.410 | |
| PPO | 0.3951 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 3.03 | 0 | 5.636 | |
| DynSyn-SAC | 0.3391 | 48 | 58 | 66 | 60 | 52 | 48 | 54 | 40 | 36 | 34 | 42 | 48 | 40 | 36 | 18 | 0.45 | 58 | 7.145 | |
| chinup | DepRL | 0.3998 | 50 | 59 | 50 | 62 | 54 | 46 | 72 | 64 | 62 | 56 | 70 | 62 | 64 | 60 | 52 | 5.46 | 32 | 2.862 |
| SAC | 0.4618 | 64 | 86 | 76 | 66 | 58 | 64 | 50 | 74 | 72 | 66 | 70 | 70 | 68 | 64 | 60 | 1.37 | 28 | 2.703 | |
| PPO | 0.3620 | 62 | 74 | 72 | 66 | 60 | 64 | 66 | 82 | 80 | 72 | 74 | 66 | 56 | 60 | 76 | 1.97 | 74 | 4.710 | |
| DynSyn-SAC | 0.2797 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 10.00 | 0 | 6.989 | |
| open door | DepRL | 0.3304 | 96 | 100 | 100 | 100 | 98 | 100 | 94 | 74 | 62 | 42 | 100 | 96 | 96 | 78 | 66 | 4.71 | 98 | 3.030 |
| SAC | 0.0962 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.97 | 0 | 2.836 | |
| PPO | 0.4112 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1.77 | 0 | 5.327 | |
| DynSyn-SAC | 0.2784 | 100 | 98 | 98 | 96 | 92 | 100 | 100 | 100 | 98 | 98 | 100 | 100 | 14 | 4 | 0 | 9.50 | 100 | 6.461 | |
Results & Analysis
Performance, robustness, and physiological behavior.
Controller rankings depend on task family and metric.
DynSyn-SAC has the highest full-suite mean success rate (52.00%), driven primarily by locomotion. DepRL performs better on interaction tasks, has nonzero observed success on 22/22 tasks, and scores higher on action and dynamics robustness; DynSyn-SAC scores higher on observation robustness. These are separately trained task-specific policies. Focused studies reveal trade-offs in reward adaptation and action representation, rather than a uniformly superior design.
| Metric ↑ | DynSyn-SAC | DepRL | SAC | PPO |
|---|---|---|---|---|
| Stabilization SR (%) · 6 tasks | 57.33 | 50.33 | 42.33 | 0.00 |
| Locomotion SR (%) · 6 tasks | 81.33 | 41.33 | 0.00 | 0.00 |
| Interaction SR (%) · 10 tasks | 31.20 | 51.40 | 2.80 | 7.40 |
| Mean SR (%) · 22 tasks | 52.00 | 48.36 | 12.82 | 3.36 |
| Action robustness | 46.66 | 53.06 | 15.36 | 3.10 |
| Observation robustness | 43.37 | 38.10 | 4.26 | 3.40 |
| Dynamics robustness | 34.75 | 39.84 | 14.50 | 2.92 |
| Mean robustness | 41.59 | 43.67 | 11.37 | 3.14 |
| Coverage · SR > 0 | 16/22 | 22/22 | 4/22 | 1/22 |
SR retains training noise. Means weight tasks equally. Robustness areas are normalized to 0–100, with the average weighting perturbation types equally. Coverage counts tasks with nonzero observed success, not universal success. Bold indicates the best point estimate in each row; these are not claims of statistical significance.
Task videos above are grouped by DepRL; the static chart shows all four reward-based methods.
| Setting | Survival (steps) ↑ | Tracking RMSE ↓ | Activation ↓ |
|---|---|---|---|
| Fixed reward | 417 | 0.4446 | 0.4769 |
| DeepSeek | 246.6 | 0.4355 | 0.3999 |
| GPT-6 Astra | 360 | 0.4324 | 0.3879 |
DeepSeek and GPT-6 Astra use the same DepRL learner, architecture, training budget, and evaluation, with an identical reward-coefficient constraint ‖w‖₁ = 40. GPT-6 Astra has better point estimates than DeepSeek on all three measures, but shorter survival than fixed-reward DepRL despite lower RMSE and activation.
| Task | Duration (s) ↑ | Divergence (s) ↑ | Tracking RMSE ↓ | Outcome |
|---|---|---|---|---|
| Walk | 0.75 | 0.53 | 0.158 | Fall |
| Run | 2.02 | 0.94 | 0.653 | Fall |
| Stairs | 0.80 | N/A | 0.173 | Fall |
GPT-6 Astra directly outputs 354-dimensional muscle excitations at 5 Hz in a 100-Hz simulation, holding each action for 0.2 s, without training or a learned low-level controller. One seed-0 rollout per task uses a 3-s horizon. Divergence denotes the first sustained tracking divergence; N/A means the fixed criterion was not met. These outcomes characterize only this tested configuration.
Task-specific tests characterize a stair-control failure.
Under the tested dynamics perturbations, MuscleMimic generally has higher measured success than DepRL on stand, jump, walk, and run. On stairs, MuscleMimic has no observed successes at any tested scale, including zero dynamics perturbation. Tracking errors characterize divergence from the reference without establishing a timing mismatch or identifying the cause of failure. This task-specific limitation motivates the combined residual adaptation study below.
Task success does not characterize muscle use.
Walking retains 100% success after adaptation while selected muscle correlations improve; running improves in both success and mean EMG correlation. Stairs recovers from 0% to 100% success, but mean correlation changes only from 0.51 to 0.52 and remains below DepRL’s 0.56. Muscle-wise EMG comparisons, task-dependent joint jerk, load-specific recruitment, and anatomical resolution reveal differences that aggregate task scores do not capture.
Residual Adaptation
Task recovery and EMG agreement must be assessed separately.
A case study connects stair-failure diagnosis to multi-metric validation. A frozen MuscleMimic prior supplies base actions and a trainable three-layer MLP supplies bounded corrections. The pipeline also changes task objectives and reference timing; its results do not isolate the effect of residual actions alone.
Combined adaptation pipeline
afull = clip(abase + α clip(ares, −1, 1), −1, 1)
- Stairs: reference progression is normal below 0.4 m planar root error and slows by 50% otherwise. Vertical root displacement is excluded from tracking; progress targets the reference state 1 s ahead.
- Walking: residual authority is 0.6 in swing and 0.2 in stance, with activation magnitude and temporal-change penalties.
- Running: flight-phase tracking penalties are halved, toe-off is rewarded, and vertical ground-reaction forces above 2.5 times body weight receive a soft penalty.
A small L2 residual penalty (cres = −0.02) replaces the generic action-rate and activation-cost terms. The 354-dimensional actions in the controller diagram exclude hand actuation.
| Task | MM SR (%) ↑ | Residual SR (%) ↑ | MM EMG r ↑ | DepRL EMG r ↑ | Residual EMG r ↑ |
|---|---|---|---|---|---|
| Walking | 100 | 100 | 0.71 / 0.69 / 0.63 | — | 0.75 / 0.80 / 0.67 |
| Running | 90 | 100 | 0.43 | 0.70 | 0.77 |
| Stairs | 0 | 100 | 0.51 | 0.56 | 0.52 |
MM = MuscleMimic. SR retains training noise. Walking reports soleus medialis / biceps femoris / semitendinosus correlations; running reports an 11-channel phase-aligned mean, and stairs reports a 12-muscle phase-aligned mean. The dash means no DepRL walking value is reported in this table. Bold marks the best point estimate in each comparison. Task recovery does not imply uniformly better human EMG agreement.
Additional diagnostics: stair height rises from 0.94 to 1.21 m and tracking error falls from 0.1968 to 0.1079 rad. Running tracking error falls from 0.2267 to 0.1338 rad and peak ground-reaction force from 6.40 to 6.07 times body weight.
Limitations & Future Work
Scope of the benchmark and next steps.
The reported results describe the evaluated tasks and perturbations; they do not establish statistical dominance or direct transfer to unseen embodiments. The shared 416-muscle humanoid controls for body design, but conclusions may not generalize to other anatomical models.
Human-reference EMG evaluation covers three tasks with task-dependent lower-limb signals: three selected muscles for walking in Table V, 11 independent channels for running, and 12 matched muscles for stairs. The correlations measure phase-optimized activation-envelope shape rather than complete biological similarity. Residual adaptation changes actions, rewards, and reference timing together, so its gains characterize the complete pipeline.
Future extensions include more dexterous manipulation tasks, longer-horizon behaviors, real-world validation, and broader physiological measurements.
Citation
BibTeX
@misc{ou2026mskbenchbenchmarkingfullbodymusculoskeletal,
title={MSK-Bench: Benchmarking Full-Body Musculoskeletal Motor Control Across Tasks, Control Paradigms, and Physiological Metrics},
author={Mengtao Ou and Zongzheng Zhang and Zhenghao Xiao and Yixuan Pan and Ziwen Zhuang and Hang Zhao and Hongyang Li and Yanan Sui and Libin Liu and Hao Zhao},
year={2026},
eprint={2609.26872},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.26872},
}