The datasets behind our Qwen3-VL-8B video RLVR runs: the base 24f/100k mixture and the HopChain v5 multi-hop corpora.
Nguyen Quang Trung
ngqtrung
AI & ML interests
None yet
Recent Activity
upvoted a paper 4 days ago
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal
Training updated a collection 5 days ago
Video RLVR — final training data updated a collection 5 days ago
Video RLVR — final training dataOrganizations
Qwen3-VL-8B RLVR — Models (v1)
Qwen3-VL-8B GRPO RLVR checkpoints from a token-dropout exploration study. OMR ppexplore=winner (0.714); video ~0.485 dead-heat.
VMAR — Raw & Source
Raw multi-style distilled traces, the pre-distillation template seed, and the never-trained real-audio eval.
videorl
Qwen3-VL-8B video RLVR - GRPO checkpoint ladder
Every published checkpoint of the 8B video-QA GRPO campaign, ranked by core-3 mean_accuracy (5,645 rows). Untrained base = 0.4426.
Qwen3-VL-8B RLVR — Datasets (v1)
Curated SFT + GRPO RL datasets (video MC-QA, OMR math-image, OpenMMReasoner-RL, Vero) for Qwen3-VL-8B post-training.
VMAR — Train-Ready (SFT + RL + Eval)
Train-ready VMAR datasets: teacher-distilled SFT corpus, RL prompt set, and the curated in-loop eval benchmark.
Video RLVR — final training data
The datasets behind our Qwen3-VL-8B video RLVR runs: the base 24f/100k mixture and the HopChain v5 multi-hop corpora.
Qwen3-VL-8B video RLVR - GRPO checkpoint ladder
Every published checkpoint of the 8B video-QA GRPO campaign, ranked by core-3 mean_accuracy (5,645 rows). Untrained base = 0.4426.
Qwen3-VL-8B RLVR — Models (v1)
Qwen3-VL-8B GRPO RLVR checkpoints from a token-dropout exploration study. OMR ppexplore=winner (0.714); video ~0.485 dead-heat.
Qwen3-VL-8B RLVR — Datasets (v1)
Curated SFT + GRPO RL datasets (video MC-QA, OMR math-image, OpenMMReasoner-RL, Vero) for Qwen3-VL-8B post-training.
VMAR — Raw & Source
Raw multi-style distilled traces, the pre-distillation template seed, and the never-trained real-audio eval.
VMAR — Train-Ready (SFT + RL + Eval)
Train-ready VMAR datasets: teacher-distilled SFT corpus, RL prompt set, and the curated in-loop eval benchmark.
videorl