AI & ML interests
RL Environments at Scale
Recent Activity
View all activity
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
HuggingEnvs/watercolour-grpo-hps-only
Reinforcement Learning β’ Updated β’ 29 β’ 1 -
HuggingEnvs/watercolour-reference-pool
Viewer β’ Updated β’ 178 β’ 523 β’ 1 -
HuggingEnvs/watercolour-rollouts-hps-only
Viewer β’ Updated β’ 470 β’ 378 -
HuggingEnvs/watercolour-grpo-judge-led
Reinforcement Learning β’ Updated β’ 15 β’ 1
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
HuggingEnvs/watercolour-grpo-hps-only
Reinforcement Learning β’ Updated β’ 29 β’ 1 -
HuggingEnvs/watercolour-reference-pool
Viewer β’ Updated β’ 178 β’ 523 β’ 1 -
HuggingEnvs/watercolour-rollouts-hps-only
Viewer β’ Updated β’ 470 β’ 378 -
HuggingEnvs/watercolour-grpo-judge-led
Reinforcement Learning β’ Updated β’ 15 β’ 1
Deterministic data-analysis agent tasks from the jupyter-agent dataset β verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents