Full-body muscle-actuated humanoid control benchmark

MSK-Bench: Benchmarking Full-Body Musculoskeletal Motor Control Across Tasks, Control Paradigms, and Physiological Metrics MSK-Bench: Benchmarking Full-Body Musculoskeletal Motor Control Across Tasks, Control Paradigms, and Physiological Metrics

Mengtao Ou*,1, Zongzheng Zhang*,1, Zhenghao Xiao1, Yixuan Pan2, Ziwen Zhuang3
Hang Zhao3, Hongyang Li2, Yanan Sui4, Libin Liu5, Hao Zhao†,1

1 Institute for AI Industry Research (AIR), Tsinghua University 2 The University of Hong Kong (HKU) 3 Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University 4 School of Aerospace Engineering, Tsinghua University 5 Peking University (PKU)

* Equal Contribution † Corresponding author

416muscles
22motor tasks
3task families
5control paradigms
7metric families
stand
powerlift
squat
sit
singleleg stand
balance
walk forward
run
walk turn
jump
sidestep
crawl
stairs
catch
hurdle
slide
stepping stones
pole walk
reach
walk and sit
chin up
open door

Abstract

Musculoskeletal (MSK) humanoids provide a physiologically grounded embodiment for studying full-body motor control, but their high-dimensional muscle actuation, delayed activation dynamics, and redundant muscle-tendon structures make learning substantially harder than torque-driven humanoid control. Existing MSK benchmarks remain fragmented across gait, prosthetics, dexterous hands, or challenge-specific tracks, leaving full-body muscle-actuated control insufficiently evaluated under standardized tasks, methods, and metrics. We introduce MSK-Bench, a benchmark of 22 full-body motor-control tasks organized into three progressively challenging categories: postural stabilization, common locomotor behaviors, and contact-rich environmental interaction. Under unified task protocols and robustness perturbations, MSK-Bench evaluates 5 representative control paradigms, including reward-based RL, agentic reward tuning, latent-action RL, imitation-prior control, and residual adaptation over imitation priors. Beyond task success and reward, MSK-Bench further reports robustness analysis and physiology-oriented diagnostics, including activation cost, joint smoothness, and EMG-envelope similarity. Our empirical study shows that embodiment-aware exploration and structured action representations improve task coverage in high-dimensional muscle spaces, imitation priors enhance reference-compatible stabilization and locomotion but degrade under contact-rich terrain mismatch, and residual adaptation can recover successful behaviors when fixed references fail. We further find that improved task success does not necessarily imply improved physiological agreement, highlighting the importance of evaluating task performance, robustness, and physiological behavior jointly. MSK-Bench provides a task-method-metric testbed for full-body muscle-actuated humanoid control.

Overview of MSK-Bench task families and method evaluation radar charts
Overview of MSK-Bench (Fig. 1). An open-source suite of 22 tasks across stabilization, locomotion, and interaction, evaluated through 5 control paradigms and 7 metric families. Radar charts use eight reported axes derived from these metrics, averaged over applicable tasks in each family. EMG-envelope similarity is evaluated only where matched human reference data are available; cumulative reward is reported separately in the learning curves.
Benchmark / Framework Embodiment Environment Task Design Evaluation Main Focus
Actuation
(Dim.)
Full
Body
3D Scene /
Terrain
Contact
Rich
Task
Hierarchy
Tasks /
Data Scale
Physio. Eval.
Metrics
Control
Families
Torque-Driven Humanoid Benchmarks
HumanoidBench [87]Torque (19/61)✓✓✓✓27 tasks✗1Loco-manipulation
SkillBench / SkillBlender [49]Torque (19)✓●✓✓4 skills /
8 tasks
✗1Skill Blending
Mimicking-Bench [60]Torque (19)✓✓✓✗6 tasks✗1Scene Interaction
Musculoskeletal Benchmarks
LocoMuJoCo [2]Mixed✓✗✗●12 envs /
27 tasks
✗2Locomotion Imitation
MyoSuite [12]Muscle (≤ 80)✗✗✓●204 tasks✗1Dexterity and Agility
MyoDex [13]Muscle (39)✗✗✓✗14 train /
34 eval.
✗1Hand Manipulation
MS-Human-700 / MsGym [126]Muscle (700)✓✗●✗3 tasks✓1MSK Locomotion
MyoChallenge 2024 [107]Muscle (≤ 80)●✓✓●2 tracks✗1Prosthetics
MyoChallenge 2025 [15]Muscle (≤ 80)●●✓●2 tasks /
4 tracks
✗1Athletic Control
MSK-Bench (ours)Muscle (416)✓✓✓✓22 tasks✓5Full-body Terrain MSK
Benchmark comparison. ✓/●/✗ denote full/partial/no support in the cited evaluations. We compare actuation, full-body embodiment, scene/terrain support, contact-rich interaction, task hierarchy, benchmark scale, physiological evaluation, and evaluated control families. Reference numbers follow the paper.

Evaluation Protocol

One task suite, five control paradigms, seven metric families.

All 22 tasks use MuJoCo and the same 416-muscle elastic-tendon full-body model. Success requires task completion, survival, and safety under fixed task-specific criteria. Within each task, the four full-suite baselines share observations, rewards, termination conditions, and evaluation.

Control paradigms & study scope

  1. Reward-based RL: PPO, SAC, DepRL, and DynSyn-SAC; each is trained and evaluated on all 22 tasks.
  2. Agentic reward tuning: focused DepRL studies with bounded reward-coefficient updates using DeepSeek-V4-Flash and GPT-6 Astra.
  3. Latent-action RL: a focused study of an expert-trained, state-conditioned encoder–decoder; compression changes the command representation, not the executed muscle-action space.
  4. Imitation-prior control: pretrained MuscleMimic evaluated on stand, jump, walk, run, and stairs.
  5. Residual adaptation: a frozen imitation prior with bounded corrections and task-specific objective/timing changes on walk, run, and stairs.

Focused studies differ in task coverage and may change reward coefficients or introduce reference objectives. They are reported separately from the full-suite ranking.

What the seven metric families measure

  1. Success Rate: successful evaluation episodes in the restored training environment, including native noise. Family and full-suite means weight tasks equally.
  2. Cumulative Reward: task-specific evaluation return, never pooled across tasks.
  3. Peak-Efficiency Steps: the evaluated training step attaining the highest mean return.
  4. Perturbation Robustness: normalized trapezoidal area under success versus perturbation scale, averaged equally over tasks and action, observation, and muscle-dynamics sweeps.
  5. Activation Cost: mean squared muscle activation, used as an effort proxy.
  6. Joint Smoothness: mean squared angular jerk from finite differences of joint velocities, displayed logarithmically.
  7. EMG-Envelope Similarity: mean per-muscle Pearson correlation after 101-point cycle resampling, per-muscle min–max normalization, cycle averaging, and correlation-maximizing cyclic alignment.

Action/dynamics scales: 0, 0.05, 0.10, 0.15, 0.20; observation scales: 0, 0.02, 0.05, 0.08, 0.10. Activation cost and smoothness are interpreted only for successful, behaviorally comparable policies. EMG correlations assess phase-optimized waveform shape, not absolute amplitude or timing.

Task Families

From postural regulation to contact-rich environmental interaction.

Leaderboard

Benchmark Performance and Robustness Evaluation

22 tasks across three families, comparing DepRL, SAC, PPO, and DynSyn-SAC.

Train-env. SR retains native training noise; robustness columns report separate perturbation sweeps. See the paper appendix for evaluation details. Bold SR values mark the highest training-environment success rate within each task, including ties. Activation cost and smoothness should be interpreted for successful, behaviorally comparable policies.

Stabilization

6 tasks · 4 algorithms

Scroll horizontally for all metrics and vertically for all tasks. Task and algorithm columns stay visible.

TaskAlgorithmActivation
Cost
Robustness (%)Max Steps
(× 107)
Train-env.
SR (%) ↑
log10 smooth
(rad2/s6)
ActionObservationDynamic
Action perturbation 0Action perturbation 0.05Action perturbation 0.1Action perturbation 0.15Action perturbation 0.2Observation perturbation 0Observation perturbation 0.02Observation perturbation 0.05Observation perturbation 0.08Observation perturbation 0.10Dynamic perturbation 0Dynamic perturbation 0.05Dynamic perturbation 0.1Dynamic perturbation 0.15Dynamic perturbation 0.2
standDepRL0.44889494909698100120009848281243.74863.152
SAC0.04841001001001009610000001001001001001007.131001.875
PPO0.36910000000000000001.9705.479
DynSyn-SAC0.28595242564636464622301440324030305.44522.104
powerliftDepRL0.335668686456526262384066705242384.14382.631
SAC0.210698949488809600009894886426.78902.825
PPO0.39550000000000000001.7605.042
DynSyn-SAC0.36311001001009892100100969490100989490828.501005.781
squatDepRL0.378144463832302456346226463812201014.02422.984
SAC0.11840004000000400004.7901.708
PPO0.34550000000000000002.0105.390
DynSyn-SAC0.19554838322622504438363048423226187.00426.503
sitDepRL0.332796100100989292100100100941009296100904.88962.872
SAC0.11648080747064840000781007670527.39643.204
PPO0.45170000000000000002.3705.350
DynSyn-SAC0.33028072706860808084827892847860269.27823.145
singlestandDepRL0.397714162818102430800122868122.10142.854
SAC0.08490000000000000004.3301.954
PPO0.39810000000000000001.8405.436
DynSyn-SAC0.3845242824201826221600221880010.00242.931
balanceDepRL0.3775604026201456000046123622188.86262.822
SAC0.14590000000000000008.1102.547
PPO0.38590000000000000001.2504.861
DynSyn-SAC0.27354038201614544082052800010.50446.066

Locomotion

6 tasks · 4 algorithms

Scroll horizontally for all metrics and vertically for all tasks. Task and algorithm columns stay visible.

TaskAlgorithmActivation
Cost
Robustness (%)Max Steps
(× 107)
Train-env.
SR (%) ↑
log10 smooth
(rad2/s6)
ActionObservationDynamic
Action perturbation 0Action perturbation 0.05Action perturbation 0.1Action perturbation 0.15Action perturbation 0.2Observation perturbation 0Observation perturbation 0.02Observation perturbation 0.05Observation perturbation 0.08Observation perturbation 0.10Dynamic perturbation 0Dynamic perturbation 0.05Dynamic perturbation 0.1Dynamic perturbation 0.15Dynamic perturbation 0.2
walk forwardDepRL0.420628283232443432462430402012863.32643.151
SAC0.14270000000000000009.8803.262
PPO0.39400000000000000002.4205.558
DynSyn-SAC0.2771828868604882828076768276766485.00763.417
runDepRL0.335430223236421818303226321612826.72183.343
SAC0.09400000000000000002.1302.854
PPO0.36960000000000000002.1605.524
DynSyn-SAC0.322196948470689694908680968822010.00963.435
walk turnDepRL0.33854660626668344058402648464036366.38423.057
SAC0.15710000000000000000.9902.610
PPO0.37470000000000000001.0505.331
DynSyn-SAC0.31129688847874949090888492240068.00863.248
jumpDepRL0.36168242012101818840141012606.89103.188
SAC0.13420000000000000003.9202.853
PPO0.38800000000000000002.0705.640
DynSyn-SAC0.27007040800625450484242302208.00543.449
sidestepDepRL0.3888403862563234122002422181064.38283.349
SAC0.10450000000000000000.00003203.166
PPO0.37290000000000000004.9205.408
DynSyn-SAC0.36307466544230747874785464726040405.58763.223
crawlDepRL0.35401001001001009896100100100100100100981001005.12863.158
SAC0.03380000000000000002.8702.516
PPO0.40360000000000000001.3405.674
DynSyn-SAC0.276610010010010010010010010010010010010010010010010.001003.444

Interaction with Environment

10 tasks · 4 algorithms

Scroll horizontally for all metrics and vertically for all tasks. Task and algorithm columns stay visible.

TaskAlgorithmActivation
Cost
Robustness (%)Max Steps
(× 107)
Train-env.
SR (%) ↑
log10 smooth
(rad2/s6)
ActionObservationDynamic
Action perturbation 0Action perturbation 0.05Action perturbation 0.1Action perturbation 0.15Action perturbation 0.2Observation perturbation 0Observation perturbation 0.02Observation perturbation 0.05Observation perturbation 0.08Observation perturbation 0.10Dynamic perturbation 0Dynamic perturbation 0.05Dynamic perturbation 0.1Dynamic perturbation 0.15Dynamic perturbation 0.2
stairDepRL0.3499525438424034382003240241441.42362.982
SAC0.19930000000000000006.9802.459
PPO0.38540000000000000001.5705.303
DynSyn-SAC0.27590000000000000005.9202.854
catchDepRL0.32478080767684624880724678404238567.23282.919
SAC0.13840000000000000001.3702.899
PPO0.39640000000000000003.1205.300
DynSyn-SAC0.31097064606058604034262268686054507.38663.134
hurdleDepRL0.3715484856504064444828062444032201.26503.212
SAC0.18520000000000000004.1202.816
PPO0.36680000000000000002.4605.398
DynSyn-SAC0.37170000000000000005.5007.102
slideDepRL0.333010022221612100164401008144611.321003.106
SAC0.09250000000000000004.6602.752
PPO0.39340000000000000004.1305.345
DynSyn-SAC0.26490220000000022003.8302.864
steppingstonesDepRL0.3606424644403846421020254382418188.32582.985
SAC0.16470000000000000007.3302.140
PPO0.40200000000000000001.2705.211
DynSyn-SAC0.356900000000000000010.0006.764
polewalkDepRL0.40802238164020422660030262016165.14122.876
SAC0.15040000000000000006.1001.724
PPO0.42740000000000000002.8705.377
DynSyn-SAC0.2909989698929098000096949476169.00883.242
reachDepRL0.34108486849482948290623492888282749.34842.975
SAC0.17050000000000000005.8502.677
PPO0.39650000000000000002.8705.152
DynSyn-SAC0.2556414044266241086402.1002.840
walkandsitDepRL0.36483026281410283632181022282214122.68163.365
SAC0.11800000000000000006.5402.410
PPO0.39510000000000000003.0305.636
DynSyn-SAC0.33914858666052485440363442484036180.45587.145
chinupDepRL0.39985059506254467264625670626460525.46322.862
SAC0.46186486766658645074726670706864601.37282.703
PPO0.36206274726660646682807274665660761.97744.710
DynSyn-SAC0.279700000000000000010.0006.989
open doorDepRL0.3304961001001009810094746242100969678664.71983.030
SAC0.09620000000000000001.9702.836
PPO0.41120000000000000001.7705.327
DynSyn-SAC0.278410098989692100100100989810010014409.501006.461

Results & Analysis

Performance, robustness, and physiological behavior.

T1

Controller rankings depend on task family and metric.

DynSyn-SAC has the highest full-suite mean success rate (52.00%), driven primarily by locomotion. DepRL performs better on interaction tasks, has nonzero observed success on 22/22 tasks, and scores higher on action and dynamics robustness; DynSyn-SAC scores higher on observation robustness. These are separately trained task-specific policies. Focused studies reveal trade-offs in reward adaptation and action representation, rather than a uniformly superior design.

Full-suite point estimates · Table II
Metric ↑DynSyn-SACDepRLSACPPO
Stabilization SR (%) · 6 tasks57.3350.3342.330.00
Locomotion SR (%) · 6 tasks81.3341.330.000.00
Interaction SR (%) · 10 tasks31.2051.402.807.40
Mean SR (%) · 22 tasks52.0048.3612.823.36
Action robustness46.6653.0615.363.10
Observation robustness43.3738.104.263.40
Dynamics robustness34.7539.8414.502.92
Mean robustness41.5943.6711.373.14
Coverage · SR > 016/2222/224/221/22

SR retains training noise. Means weight tasks equally. Robustness areas are normalized to 0–100, with the average weighting perturbation types equally. Coverage counts tasks with nonzero observed success, not universal success. Bold indicates the best point estimate in each row; these are not claims of statistical significance.

Task-specific reward-learning curves for 22 tasks, comparing DepRL, DynSyn-SAC, SAC, and PPO
Task-specific reward-learning curves for 22 tasks (Fig. 2). DynSyn-SAC attains higher returns than DepRL on jump and hurdle over the displayed overlapping ranges, while DepRL leads on walk-turn and catch; SAC remains competitive on powerlift. Shading shows across-run variation. The full-suite table separately reports success, nonzero-success coverage, and robustness over all 22 tasks.
Agentic reward tuning reward weights and robustness perturbation plots
Reward adaptation and robustness (Fig. 3). In the original Agentic-DepRL setting, coefficients are updated every 20,000 steps and each nonzero-weight change is limited to 20%; unnormalized coefficient magnitudes still grow. Fixed-reward DepRL maintains higher survival under the tested action and observation perturbations, while agentic tuning leads at larger dynamics-perturbation scales. Both share learner, architecture, and budget. Each point averages 50 held-out seeds; survival is capped at 1,000 steps.
Matched reasoner comparison · Table III
SettingSurvival (steps) ↑Tracking RMSE ↓Activation ↓
Fixed reward4170.44460.4769
DeepSeek246.60.43550.3999
GPT-6 Astra3600.43240.3879

DeepSeek and GPT-6 Astra use the same DepRL learner, architecture, training budget, and evaluation, with an identical reward-coefficient constraint ‖w‖₁ = 40. GPT-6 Astra has better point estimates than DeepSeek on all three measures, but shorter survival than fixed-reward DepRL despite lower RMSE and activation.

Direct muscle-control diagnostic · Table IV
TaskDuration (s) ↑Divergence (s) ↑Tracking RMSE ↓Outcome
Walk0.750.530.158Fall
Run2.020.940.653Fall
Stairs0.80N/A0.173Fall

GPT-6 Astra directly outputs 354-dimensional muscle excitations at 5 Hz in a 100-Hz simulation, holding each action for 0.2 s, without training or a learned low-level controller. One seed-0 rollout per task uses a 3-s horizon. Divergence denotes the first sustained tracking divergence; N/A means the fixed criterion was not met. These outcomes characterize only this tested configuration.

Latent action compression reward and robustness plots
128Dlatent dimension
64Dlatent dimension
32Dlatent dimension
Action compression in a focused study (Fig. 4). A 128-dimensional state-conditioned representation improves returns over the displayed overlapping training interval and success at several perturbation scales relative to the 416-dimensional baseline. Stronger compression to 64, 32, or 16 dimensions degrades performance and reconstruction statistics. The executed muscle-action space remains unchanged; this is not a full-suite comparison.
T2

Task-specific tests characterize a stair-control failure.

Under the tested dynamics perturbations, MuscleMimic generally has higher measured success than DepRL on stand, jump, walk, and run. On stairs, MuscleMimic has no observed successes at any tested scale, including zero dynamics perturbation. Tracking errors characterize divergence from the reference without establishing a timing mismatch or identifying the cause of failure. This task-specific limitation motivates the combined residual adaptation study below.

Robustness comparison and tracking error analysis between imitation prior and DepRL
Task-dependent dynamics robustness and stair-tracking errors (Fig. 5). (a–e) Success rates of MuscleMimic/MM (PPO) and DepRL on stand, jump, walk, run, and stairs. MuscleMimic has no observed stair successes, including at zero perturbation. (f) From step 150 to 300, absolute-site error rises from 0.35 to 1.31 and root-XYZ error from 0.25 to 0.77. These errors describe reference divergence; they do not establish its cause. Both methods share success conditions and perturbation settings within each task. This focused protocol differs from the full-suite benchmark; its values are excluded from aggregate robustness scores.
T3

Task success does not characterize muscle use.

Walking retains 100% success after adaptation while selected muscle correlations improve; running improves in both success and mean EMG correlation. Stairs recovers from 0% to 100% success, but mean correlation changes only from 0.51 to 0.52 and remains below DepRL’s 0.56. Muscle-wise EMG comparisons, task-dependent joint jerk, load-specific recruitment, and anatomical resolution reveal differences that aggregate task scores do not capture.

Walking EMG-envelope similarity across policies
Walking EMG-envelope comparison (Fig. 6). Human recordings from Gait120 versus simulated activations for reward-based RL, latent-action RL, agentic tuning, and imitation-prior control. Signals are resampled to 101 gait-cycle points, min–max normalized per muscle, cycle-averaged, and cyclically aligned to maximize Pearson correlation. Correlations assess phase-optimized waveform shape, not absolute amplitude or timing. Agreement varies by muscle: latent-action RL improves biceps femoris correlation (0.95 vs. 0.84) but lowers vastus lateralis correlation (0.21 vs. 0.73) relative to reward-based RL. Walking and stair references use Gait120; running uses the 3-m/s recordings of Van Hooren and Meijer.
Task-wise joint smoothness across benchmark tasks
Task-wise joint smoothness (Fig. 7). Bars report log10 mean squared angular jerk computed from finite differences of joint velocities. Levels depend on the task, with overlapping ranges across task families. This characterizes movement fluctuations, not a cross-task ranking of success, control quality, or human similarity.
Load-dependent muscle-family activation during powerlift
0.05kgbarbell weight
0.5kgbarbell weight
5kgbarbell weight
10kgbarbell weight
Load-dependent recruitment in powerlift (Fig. 8). Heatmap values sum the time-averaged activations of the actuators in each muscle family. Trunk/back totals are 30.1, 50.4, and 46.0 at 0.05, 5, and 10 kg. These describe load-specific muscle use, not a uniform activation increase, instantaneous activation, or agreement with human recordings.
Effect of anatomical resolution on powerlift muscle recruitment
416-muscle modelshared muscle families
700-muscle modeladditional axial stabilizers
Anatomical resolution and powerlift recruitment (Fig. 9). Activation mass for the 416- and 700-muscle models. The 700-muscle model adds neck, deep-trunk, and deep-scapular groups, expanding the available recruitment pathways and diagnostics. These model-specific profiles do not establish greater human similarity; only this anatomy study uses the 700-muscle model.

Residual Adaptation

Task recovery and EMG agreement must be assessed separately.

A case study connects stair-failure diagnosis to multi-metric validation. A frozen MuscleMimic prior supplies base actions and a trainable three-layer MLP supplies bounded corrections. The pipeline also changes task objectives and reference timing; its results do not isolate the effect of residual actions alone.

Combined adaptation pipeline

afull = clip(abase + α clip(ares, −1, 1), −1, 1)

  • Stairs: reference progression is normal below 0.4 m planar root error and slows by 50% otherwise. Vertical root displacement is excluded from tracking; progress targets the reference state 1 s ahead.
  • Walking: residual authority is 0.6 in swing and 0.2 in stance, with activation magnitude and temporal-change penalties.
  • Running: flight-phase tracking penalties are halved, toe-off is rewarded, and vertical ground-reaction forces above 2.5 times body weight receive a soft penalty.

A small L2 residual penalty (cres = −0.02) replaces the generic action-rate and activation-cost terms. The 354-dimensional actions in the controller diagram exclude hand actuation.

Frozen MuscleMimic prior plus trainable residual controller, with example stair, run, and walk motions
Residual architecture and example motions (Fig. 10). The frozen prior and trainable correction are combined into the executed action. The lower branch supplies the residual action in the equation above. Blue skeletons are references. Rewards and reference timing also change.
MuscleMimic
Stairs
Run
Walk
Residual adaptation
Stairs
Run
Walk
MuscleMimic and the combined residual adaptation pipeline. The top row shows imitation-prior rollouts for stairs, running, and walking, with adapted rollouts below. Videos illustrate behavior; task success and EMG-envelope agreement are reported separately in Table V.
Residual adaptation results · Table V
TaskMM SR (%) ↑Residual SR (%) ↑MM EMG r ↑DepRL EMG r ↑Residual EMG r ↑
Walking1001000.71 / 0.69 / 0.63—0.75 / 0.80 / 0.67
Running901000.430.700.77
Stairs01000.510.560.52

MM = MuscleMimic. SR retains training noise. Walking reports soleus medialis / biceps femoris / semitendinosus correlations; running reports an 11-channel phase-aligned mean, and stairs reports a 12-muscle phase-aligned mean. The dash means no DepRL walking value is reported in this table. Bold marks the best point estimate in each comparison. Task recovery does not imply uniformly better human EMG agreement.

Additional diagnostics: stair height rises from 0.94 to 1.21 m and tracking error falls from 0.1968 to 0.1079 rad. Running tracking error falls from 0.2267 to 0.1338 rad and peak ground-reaction force from 6.40 to 6.07 times body weight.

Limitations & Future Work

Scope of the benchmark and next steps.

The reported results describe the evaluated tasks and perturbations; they do not establish statistical dominance or direct transfer to unseen embodiments. The shared 416-muscle humanoid controls for body design, but conclusions may not generalize to other anatomical models.

Human-reference EMG evaluation covers three tasks with task-dependent lower-limb signals: three selected muscles for walking in Table V, 11 independent channels for running, and 12 matched muscles for stairs. The correlations measure phase-optimized activation-envelope shape rather than complete biological similarity. Residual adaptation changes actions, rewards, and reference timing together, so its gains characterize the complete pipeline.

Future extensions include more dexterous manipulation tasks, longer-horizon behaviors, real-world validation, and broader physiological measurements.

Citation

BibTeX

@misc{ou2026mskbenchbenchmarkingfullbodymusculoskeletal,
      title={MSK-Bench: Benchmarking Full-Body Musculoskeletal Motor Control Across Tasks, Control Paradigms, and Physiological Metrics},
      author={Mengtao Ou and Zongzheng Zhang and Zhenghao Xiao and Yixuan Pan and Ziwen Zhuang and Hang Zhao and Hongyang Li and Yanan Sui and Libin Liu and Hao Zhao},
      year={2026},
      eprint={2609.26872},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.26872},
}