Imaginator-8B

The Imaginator from Beyond Thinking: Imagining in 360ยฐ for Humanoid Visual Search.

An agent searching a 360ยฐ scene sees only a narrow field of view at a time. The Imaginator looks at the views seen so far and says where in the full panorama the target probably is โ€” including in the parts nobody has looked at yet. It does not act. A separate, frozen search policy decides what to do with the guess, so the same Imaginator plugs into any policy without retraining it.

Trained from Qwen3-VL-8B-Instruct.

Path in this repo What it is
Imagine-8B/ this model
HVS-3B/ the frozen search policy used for the numbers below, from Yu et al.

Input and output

Input โ€” every narrow view seen so far, each labelled with the camera direction it was taken at, plus the instruction:

[View 1] viewing direction: (0,0)
<image>
[View 2] viewing direction: (0,-30)
<image>
Human Instruction: look for the black and white striped blanket
Decide your next action.

viewing direction is (yaw, pitch) in degrees: yaw in [0,360) measured clockwise, pitch in [-90,90] with positive up. Views accumulate across the episode; there is one <image> per view, in the order listed.

Output โ€” a reasoning block that separates what is visible from what is inferred, then a single predicted target location:

<think>[Observed]
- rolled rugs: (342,-4)
- ceiling light: (7,34)
- shopping cart handle: (40,-32)
[Imagined]
- rug sample shelves: (126,-9)
- bedding section: (200,-5)
</think><answer>suggest check(126,-9)</answer>

[Observed] are landmarks the model can see in the given views, with their absolute panorama coordinates. [Imagined] are landmarks it infers lie outside them โ€” this is the part that carries the spatial prior. The <answer> is one absolute (yaw, pitch) guess at where the target is.

The coordinate is a proposal, not a detection. On HOS its top-1 hit rate under the benchmark tolerance is 39.04%, and the system still reaches 62.75, because the policy is free to reject a bad guess and keep searching.

Downstream, the harness converts check(yaw,pitch) into the relative rotate(dyaw,dpitch) or submit(yaw,pitch) form the policy was trained on, depending on whether the target is already within tolerance of the current view, and appends it to the policy's turn.

Results

H*Bench success rate, from the paper. The policy is identical in both rows and frozen; the only difference is whether it receives the Imaginator's guess.

HOS HPS
HVS-3B 48.04 24.12
HVS-3B + Imaginator-8B 62.75 39.38

Usage

vLLM cannot serve a model from a subdirectory of a repo, so fetch it first:

hf download jdzhang0929/Imaginator --include "Imagine-8B/*" --local-dir ./checkpoints
vllm serve ./checkpoints/Imagine-8B --port 8001 --served-model-name imaginator

With transformers the subfolder is addressable directly:

from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained(
    "jdzhang0929/Imaginator", subfolder="Imagine-8B",
    dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained(
    "jdzhang0929/Imaginator", subfolder="Imagine-8B")

Greedy decoding, max_tokens=4096. The full two-model loop โ€” view accumulation, coordinate conversion, and the multi-hypothesis injection used in the paper โ€” is in the code repo.

Training

Two stages. Stage 1 pretrains on 1.92M pseudo-labelled panorama samples (10,000 steps at batch 192) drawn from Sun360, Matterport3D and DiT360 renders. Stage 2 is a clean SFT on 4,223 H*Bench trajectories, 2 epochs at batch 64, lr 1e-5 cosine.

Stage 2 is reproducible from public data: the trajectories are at jdzhang0929/Imagine-in-360-Dataset and the images they reference ship with H*Bench.

Data separation

Evaluation panoramas were compared against the stage-2 SFT panoramas exhaustively at the pixel level โ€” 857 ร— 382 pairs, 36 yaw rotations each, under two criteria (MAE < 2.0 grey levels for "same photo", Pearson r โ‰ฅ 0.90 for "same viewpoint"). Two overlaps surfaced, both inherited from the original H*Bench split, and both are removed from the released SFT trajectories. Against the 40,453 stage-1 pseudo-label panoramas the same sweep found nothing at either threshold; the global maximum correlation was 0.8965.

Limitations

  • English instructions only.
  • Top-1 coordinate accuracy is 39.04% (HOS) / 26.12% (HPS). Treat the output as a prior over where to look, not as a localization result.
  • Rotation-only search from a fixed viewpoint; no translation.
  • Coordinates assume the equirectangular convention above. A different yaw origin or pitch sign will silently produce plausible but wrong guesses.

Citation

@article{imagining360,
  title  = {Beyond Thinking: Imagining in 360{\deg} for Humanoid Visual Search},
  author = {Zhang, Jingdong and others},
  year   = {2026}
}

Built on H*Bench and the HVS models from Yu et al.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jdzhang0929/Imaginator

Finetuned
(521)
this model