Learned Routing for Specialist–VLM Collaboration

Artifacts for When Should a Vision Specialist Defer? Budgeted Collaboration with Vision–Language Models.

Code and reproduction instructions · Latest paper · Supplement

Latest release

Use releases/20260913_lr.

  • 80 main-table router checkpoints: 16 model pairs × five seeds, in full_budget_20260913_weights.tar.gz.
  • Experiment code, configurations, histories and per-example outputs, including the other objectives and scoring experiments. Only the main-table router weights are included in this release.
  • Eight specialist input caches covering train, validation and test splits, with an input manifest.
  • Paper sources and numerical summaries in the code repository; latest PDFs and result data here.

To reconstruct Tables 1 and 3 without training, download paper_replay_runs.tar.gz and paired_outcomes.tar.gz, then follow the reproduction instructions. The replay checks 80 main runs, 320 Table 3 runs and 19,392 curve points.

Main checkpoints use validation selection across 0–100% calls. Table 3 preserves its earlier checkpoint selection settings, detailed in the supplement.

Root-level archives are retained for earlier releases. The latest release does not add supplementary router checkpoints.

Data and licenses

Original code is MIT licensed. Upstream model and dataset licenses remain unchanged. This release provides derived experiment features, labels and outputs; obtain original images from their providers. ConstructionSite-derived data retain CC BY-NC 4.0 terms. See model and dataset sources.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support