Latest article
microduck_rl Architecture: Environment, mjlab Physics, Rewards, and Policy Runner
Trace observations through actions, physics, reward shaping, termination, and policy updates in the Microduck loop.
A simulation-first series on Microduck reinforcement-learning environments and sim-to-real gates.

Latest article
Trace observations through actions, physics, reward shaping, termination, and policy updates in the Microduck loop.
By publication date
01 → 09
Trace observations through actions, physics, reward shaping, termination, and policy updates in the Microduck loop.
Choose a robotics RL stack by fidelity, throughput, assets, hardware path, reproducibility, and safety burden.
Deploy training and evaluation workers with pinned simulators, GPU isolation, artifact storage, and a guarded hardware handoff.
Start with a pinned mjlab environment, inspect observations and rewards, and save a reproducible rollout.
Design a Microduck RL project with task manifests, deterministic fixtures, policy cards, telemetry, and staged hardware gates.
A source-backed overview of Microduck RL environments, the mjlab loop, and sim-to-real safety gates.
Benchmark simulation throughput, GPU memory, parallel environments, checkpoint storage, and accepted-policy review.
Govern robot assets, checkpoints, telemetry, control interfaces, operator approvals, and incident recovery.
Use deterministic fixtures to read environment registration, mjlab adapters, rewards, resets, runners, and checkpoint export.