Job Description
Getting AI to reason correctly about the physical world — locomotion, manipulation, contact dynamics, and reinforcement learning in simulation — is one of the genuinely hard unsolved problems in the field. Chegg is seeking practitioners with deep simulation and robot-learning experience to build evaluation environments, train RL policies, and assess embodied AI performance. If you’ve spent time tuning reward functions and debugging contact models in MuJoCo, this role was designed for you.
Core Responsibilities
- Build and iterate on physics-based simulation environments for locomotion, dexterous manipulation, and multi-agent tasks used as AI evaluation benchmarks
- Implement and tune reinforcement learning training pipelines — PPO, SAC, TD3, and related algorithms — to produce stable, generalisable policies
- Define reward structures, observation specifications, and action interfaces that result in robust policies transferable across task variations
- Diagnose unstable training dynamics, contact discontinuities, and simulation artefacts; document root causes and corrective actions clearly
- Evaluate trained policies for task success, physical plausibility, and sim-to-real transfer potential; produce structured assessment reports
Key Qualifications
- Substantial hands-on simulation experience with MuJoCo, dm_control, Gymnasium-Robotics, or a directly comparable physics engine
- Solid theoretical and applied understanding of reinforcement learning — policy gradient methods, actor-critic architectures, reward shaping
- Proficient in Python and familiar with standard robotics or ML tooling
- Experienced resolving unstable training dynamics and reward function pathologies in practice, not just in theory
- Clear technical writer; dependable on independent asynchronous work
Nice to Have
- Practical experience with ROS, Isaac Lab, or other robotics middleware
- Research background in motion planning, computer vision, or human-robot interaction
- Published work or open-source contributions in robot learning
Why Chegg
- Fully remote and flexible
- Task-based commitment — typically 10–40 hours per week
- First-mover opportunity in a fast-expanding AI frontier
- Growing programme with expanded ongoing engagement for top contributors