microduck_rl revision 2b25a48. No checkpoint export, GPU run or physical deployment was performed by this site. The commands below operate on your own checkpoint and write local files.Know what each file represents.
| Artifact | Purpose |
|---|---|
| Training checkpoint (.pt) | Saved training state used by the matching exporter. Keep the task and training configuration with it. |
| Exported model (.onnx) | Inference graph. The official export path includes the observation normalization learned during training. |
| Policy package | A model, manifest and README describing how the runtime should use it. Passing package checks does not establish walking quality or hardware safety. |
A policy takes an observation—the expected robot state and command values—and returns actions. The current shared contract has 61 observation values and 14 actions. Their order and meaning matter as much as their count. MicroDuck RL quick start and task registry ↗.
Export the checkpoint with its matching task.
Use the configured environment from your training run. Replace the example checkpoint path with the actual file you saved. The following task ID is for the walking example; use the task that produced your checkpoint. Keep an existing first-walk.onnx elsewhere if you need to preserve it.
uv run scripts/export.py Mjlab-Velocity-Flat-MicroDuck \
--checkpoint-file /absolute/path/to/model_3000.pt \
--num-envs 1 \
--onnx-file first-walk.onnxThis is the project exporter, not a generic conversion command. It builds the training environment and exports the actor together with its normalizer. Seeing plausible motion in play does not verify that a hand-converted model contains that normalizer. Check the exporter’s completion message and actual output file. Official ONNX exporter and observation normalization ↗.
Rehearse the exported file with the right layout.
uv run scripts/infer_policy.py \
--walking first-walk.onnx \
--new-cmd-obsThe viewer prints the walking model’s input shape. The flag above selects the current command layout. Legacy 51-input policies need it off. Do not pad an arbitrary model to make the dimensions match: that does not repair different observation meanings, joint ordering, timing or action scaling. The local script uses CPU ONNX inference, but still needs its Python dependencies, model assets and a working viewer. Local playback flags and keyboard controls ↗.
| Action | Browser simulator | Local infer_policy.py |
|---|---|---|
| Sit / stand | R in feet mode | Y with the matching sit/stand policy loaded |
| Left / right kick | Q / E in feet mode | K / L with kick policies loaded |
| Forward roll | R is assigned to sit/stand | R only when a roulade policy is supplied |
Optional actions need their matching model files. The walking-only command above does not load sitting, kicking or roll policies. For the current browser mapping, use the simulator guide.
Inspect a package before sharing.
The publisher has a dry-run mode that stages a local package. Replace YOUR_USERNAME with the intended Hugging Face namespace. This example describes a perpetual walking policy; do not use those semantics for an unrelated trick.
uv run publish --onnx first-walk.onnx \
--repo YOUR_USERNAME/microduck-first-walk \
--kind perpetual --slot walk --dry-runThe publisher checks the expected shape, executes a basic output check and validates the manifest before staging policy.onnx, manifest.json and README.md. Dry-run stops before the Hub upload. It recreates its local publish-first-walk/ staging folder, so keep manual notes outside that folder. Policy packaging, validation and dry-run implementation ↗.
A finite, nonconstant model output is only a basic validity check. Before any hardware attempt, document what happened in simulation when starting, stopping, turning and recovering. Keep the robot’s runtime documentation and your exact configuration beside those observations.
Keep provenance attached to the result.
Original repository and full revision:
Task ID and training configuration:
Checkpoint file or run URL:
Exact export command:
ONNX filename and SHA-256:
Input/output shapes and command layout:
Playback command and robot model:
Action scale and control frequency:
Observed simulation behavior:
Known failures and untested cases:
Hardware validation, if any:
Original author, source and license references:For a custom policy, record its source and permissions separately from robot meshes, music or other bundled assets. A result from simulation should remain labeled as simulation when you share it.
Continue with your result.
What would you like to understand next?
Keep the source, your environment and your observations together.