Skip to content
TroubleshootingSource reviewed

MicroDuck Troubleshooting: Loading, CUDA & ONNX

First identify the failing stage: browser loading, dependency setup, GPU discovery, training, or policy replay. Save the first error and the source revision. Change one thing at a time, then repeat the same check to see whether the symptom changed.

Start at the first failing stage.

Use the table matching your environment. A browser loading problem, a missing Python dependency and an ONNX shape mismatch happen at different stages. Keep the original error before changing versions; otherwise you may lose the clue that identifies the failing layer.

Browser simulator problems.

Browser symptoms and the next check
SymptomTry thisWhat to record
Stuck on loadingOpen the direct simulator. Try the official boot diagnostic view ↗. Wait for its asset or policy step to finish.Browser version, the last loading step and any SYSTEM HALTED error. A transient network failure is not evidence that your GPU cannot train.
The robot ignores the arrowsEnter or resume the scene, click it to give it focus, then hold an arrow. Try the direct page if the wrapper is capturing keys.Whether the welcome overlay closed, whether the page scrolled, and whether the speed readout changed.
R sits instead of rollingUse the current browser key map. R now controls sit/stand in feet mode.The tutorial URL and simulator revision. Do not substitute shortcuts from the local Python viewer.
Kicks fail after switching modeReturn to Feet mode. Kicks and ground pick are legs-only actions; the first Rollers switch also has extra loading.Selected mode and whether its assets finished loading.

These checks combine our desktop entry/walking/reset check with the Simulator README: modes, loading and controls ↗ and Browser keyboard implementation ↗. No browser benchmark or mobile compatibility claim is implied.

Installation and CUDA problems.

Dependency and GPU symptoms
SymptomNext check
Python version rejectedRun uv run python --version from the checkout. This revision requires 3.12; use the pinned setup steps.
GPU exists, but CUDA is False or GPU count is 0Run the Python/CUDA probe from the same environment as training. Record PyTorch’s version and CUDA build, OS and architecture. On Linux ARM, inspect the project’s selected wheel source before changing drivers.
ARM dependency download times outThe official quick start suggests a longer uv HTTP timeout for the first large CUDA-wheel download. Try the one-command setting below. It only addresses download timeouts.
Jetson Thor: missing NVPL or cuDSS libraryRead issue #38 with your exact JetPack, driver and architecture. It documents a Thor-specific package problem; a GB10 package recipe is not automatically a Jetson recipe. The report was still open at review.
An unrelated train command runs, or train is missingUse uv run train inside the synced checkout. The project intentionally owns this entry point; its dependency file documents a previous executable collision. Re-sync the matching revision before altering your system PATH.
CUDA out of memoryRecord the allocation error, GPU memory and environment count. Close only your own unused GPU jobs, then retry the same task with fewer environments. This is a diagnostic reduction, not a guaranteed cure for every allocation failure.
Retry a timed-out dependency download
UV_HTTP_TIMEOUT=600 uv sync --locked --python 3.12

Run that one-line assignment in a POSIX shell such as bash or zsh. The underlying constraints and executable collision are documented in Python, dependency pins and ARM package sources ↗; GPU discovery is described by PyTorch: checking CUDA availability ↗. Jetson Thor report #38 (community report, open at review) ↗ is a user report, not a fix tested by this site.

Export and replay problems.

Policy symptoms and matching checks
SymptomNext check
ONNX expects 61 inputs but receives an older layoutFor the reviewed current policy contract, use --new-cmd-obs in the local viewer. Leave it off for legacy 51-input policies. Confirm the actual input shape before choosing.
A .pt file will not load as ONNXUse the official checkpoint exporter. The file types serve different purposes; renaming a checkpoint does not convert it.
play looks plausible, but exported replay behaves badlyCheck that the official exporter included observation normalization. Then compare task, checkpoint, robot model, action scale and command layout. This is an investigation order, not a diagnosis of your specific policy.
A remote job continues after Ctrl-COpen the saved job URL, inspect its status and explicitly cancel if appropriate. Interrupting the log stream does not stop the job.

Inspect Local playback flags and keyboard controls ↗ and Official ONNX exporter and observation normalization ↗ for the layout and normalization behavior. The cloud log behavior is documented in MicroDuck cloud job submission guide ↗; cancellation is covered in Hugging Face: managing and cancelling jobs ↗.

Send a report someone can act on.

Replace the fields below with what you observed. Include a short, relevant error excerpt; remove tokens and personal paths before sharing. “What worked last” is often more useful than a screenshot of the final exception alone.

Troubleshooting report template
Goal:
Failing stage (browser / setup / CUDA / training / export / replay):
Page URL or repository + full commit:
OS and architecture:
Browser OR Python / uv / PyTorch / CUDA versions:
GPU and driver, if relevant:
Task ID / policy source / input shape:
Exact command or sequence of actions:
Expected behavior:
Actual behavior:
First error (short excerpt, secrets removed):
Last successful step:
One change tried and its result:

Start a related discussion →. For a suspected upstream bug, include the same minimal reproduction at the original project’s available support channel.

Continue with your result.

CONTINUE THE CONVERSATION

What would you like to understand next?

Keep the source, your environment and your observations together.