213 checkpoints from four async multi-turn agentic-env GSPO runs on Qwen3-8B / Qwen3.5-9B; one subfolder per iteration.
will
willamazon1
·
AI & ML interests
None yet
Recent Activity
updated a collection 13 minutes ago
Qwen3-8B TMax AENV v39b updated a collection 13 minutes ago
Qwen3-8B TMax AENV v39b updated a collection 13 minutes ago
Qwen3-8B TMax AENV v39bOrganizations
None yet
TMax / Smith agentic RL
213 checkpoints from four async multi-turn agentic-env GSPO runs on Qwen3-8B / Qwen3.5-9B; one subfolder per iteration.
Qwen3-8B TMax AENV v39b
Qwen3-8B agentic RL run tmax_aenv_v39b: the SFT init (tmax_sft_v3 iter353) plus RL checkpoints at iterations 69-219, every 10 steps.
models 36
willamazon1/Qwen3-8B-tmax-aenv-r1v8
Text Generation • Updated
willamazon1/qwen3-8b-tmax-aenv-v39b-iter139
Text Generation • 8B • Updated
willamazon1/qwen3-8b-tmax-aenv-v39b-iter219
Text Generation • 8B • Updated
willamazon1/qwen3-8b-tmax-aenv-v39b-iter209
Text Generation • 8B • Updated
willamazon1/qwen3-8b-tmax-aenv-v39b-iter199
Text Generation • 8B • Updated
willamazon1/qwen3-8b-tmax-aenv-v39b-iter189
Text Generation • 8B • Updated
willamazon1/qwen3-8b-tmax-aenv-v39b-iter179
Text Generation • 8B • Updated
willamazon1/qwen3-8b-tmax-aenv-v39b-iter169
Text Generation • 8B • Updated
willamazon1/qwen3-8b-tmax-aenv-v39b-iter159
Text Generation • 8B • Updated
willamazon1/qwen3-8b-tmax-aenv-v39b-iter149
Text Generation • 8B • Updated