RefOpt TMR-G1 evaluator
Frozen G1 motion encoder used for G1-FID in Optimizing the Reference Motion for Human-to-Robot Retargeting. This is the existing evaluator checkpoint used by RefOpt, not a new training run.
Source code and experiment setup
Download
Use the v1 revision for the paper's evaluator. Download both student.pt
and stats/; substituting newly trained weights or new normalization
statistics changes the metric.
hf download Akihisa-Watanabe/RefOpt-TMR-G1-evaluator --revision v1 --local-dir data/tmr_g1
In an updated RefOpt checkout, python scripts/download_data.py --what evaluator
performs this download without requiring login or fetching BONES-SEED.
The BONES-SEED dataset itself remains gated and is not redistributed here.
Loading and inputs
Run from the RefOpt repository root in its Python 3.10 environment (uv sync).
The repository pins Kimodo, which provides the G1 skeleton and TMR architecture.
import numpy as np
from refopt.eval.embedding import G1Encoder, fid
encoder = G1Encoder.load(
"data/tmr_g1/student.pt", "data/tmr_g1/stats", device="cpu",
)
# Each qpos is a [T, 36] array at 30 fps in the repository's G1 joint order:
# root translation (metres), root quaternion (wxyz), 29 joint angles (radians).
# method_qpos and clean_qpos are lists of trajectories for the evaluation set.
method_embeddings = np.stack([encoder.embed(qpos) for qpos in method_qpos])
clean_embeddings = np.stack([encoder.embed(qpos) for qpos in clean_qpos])
score = fid(method_embeddings, clean_embeddings)
For the paper, Experiment 1 evaluates all 300 configured test clips, at 30 fps, with at most 300 frames per clip. Its evaluation code applies the repository's trajectory conventions. See metric definitions.
Model and training
The student is an ACTOR-style transformer motion encoder on Kimodo's 34-joint
G1 skeleton: 4 layers, 4 attention heads, feed-forward size 1024, latent size
256, and training dropout 0.1. Inference uses the VAE mean token and unit-normalizes
the resulting 256-dimensional embedding. student.pt contains the encoder's
state dictionary, architecture configuration and frame rate.
The frozen teacher is NVIDIA TMR-SOMA-RP-v1. The student was trained on paired BONES-SEED SOMA and G1 motions using cosine alignment and symmetric in-batch InfoNCE (temperature 0.1), with AdamW for 20 epochs. The actor-disjoint split is in configs/source/bones_seed.json. Only training actors contribute to weight updates and feature normalization; the checkpoint is selected by validation cosine similarity.
stats/ contains the training feature mean and standard deviation for the
body, global-root and local-root channels in Kimodo's layout. These are
normalization arrays, not clean-test embedding statistics. No per-clip
training arrays, raw motions, teacher embeddings, optimizer states or logs
are included. Both compared motion sets must pass through this same encoder.
Intended use and limitations
This encoder is for G1-FID evaluation of Unitree G1 motion in RefOpt, not a retargeting solver, a controller, or a human-motion quality score. FID measures distributional similarity and does not certify physical feasibility, safety, or frame-wise reconstruction accuracy. Do not compare its scores with FID from another encoder or a different evaluation set.
License and attribution
This distilled G1 adaptation and its weights are distributed under the NVIDIA Open Model License, separately from RefOpt's MIT-licensed code. See Notice for the upstream attribution. NVIDIA's original teacher checkpoint is not included.
Training data includes Motion Data by Bones Studio. Use of the underlying dataset is subject to the BONES Motion Capture Dataset License Agreement. Applicable dataset restrictions on Results remain in effect; this release does not grant access to or redistribution rights for the raw dataset.
Model tree for Akihisa-Watanabe/RefOpt-TMR-G1-evaluator
Base model
nvidia/TMR-SOMA-RP-v1