Instructions to use jdzhang0929/Imaginator with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jdzhang0929/Imaginator with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="jdzhang0929/Imaginator")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("jdzhang0929/Imaginator", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jdzhang0929/Imaginator with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jdzhang0929/Imaginator" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jdzhang0929/Imaginator", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/jdzhang0929/Imaginator
- SGLang
How to use jdzhang0929/Imaginator with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jdzhang0929/Imaginator" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jdzhang0929/Imaginator", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jdzhang0929/Imaginator" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jdzhang0929/Imaginator", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use jdzhang0929/Imaginator with Docker Model Runner:
docker model run hf.co/jdzhang0929/Imaginator
Imaginator-8B
The Imaginator from Beyond Thinking: Imagining in 360ยฐ for Humanoid Visual Search.
An agent searching a 360ยฐ scene sees only a narrow field of view at a time. The Imaginator looks at the views seen so far and says where in the full panorama the target probably is โ including in the parts nobody has looked at yet. It does not act. A separate, frozen search policy decides what to do with the guess, so the same Imaginator plugs into any policy without retraining it.
Trained from Qwen3-VL-8B-Instruct.
| Path in this repo | What it is |
|---|---|
Imagine-8B/ |
this model |
HVS-3B/ |
the frozen search policy used for the numbers below, from Yu et al. |
Input and output
Input โ every narrow view seen so far, each labelled with the camera direction it was taken at, plus the instruction:
[View 1] viewing direction: (0,0)
<image>
[View 2] viewing direction: (0,-30)
<image>
Human Instruction: look for the black and white striped blanket
Decide your next action.
viewing direction is (yaw, pitch) in degrees: yaw in [0,360) measured
clockwise, pitch in [-90,90] with positive up. Views accumulate across the
episode; there is one <image> per view, in the order listed.
Output โ a reasoning block that separates what is visible from what is inferred, then a single predicted target location:
<think>[Observed]
- rolled rugs: (342,-4)
- ceiling light: (7,34)
- shopping cart handle: (40,-32)
[Imagined]
- rug sample shelves: (126,-9)
- bedding section: (200,-5)
</think><answer>suggest check(126,-9)</answer>
[Observed] are landmarks the model can see in the given views, with their
absolute panorama coordinates. [Imagined] are landmarks it infers lie outside
them โ this is the part that carries the spatial prior. The <answer> is one
absolute (yaw, pitch) guess at where the target is.
The coordinate is a proposal, not a detection. On HOS its top-1 hit rate under the benchmark tolerance is 39.04%, and the system still reaches 62.75, because the policy is free to reject a bad guess and keep searching.
Downstream, the harness converts check(yaw,pitch) into the relative
rotate(dyaw,dpitch) or submit(yaw,pitch) form the policy was trained on,
depending on whether the target is already within tolerance of the current
view, and appends it to the policy's turn.
Results
H*Bench success rate, from the paper. The policy is identical in both rows and frozen; the only difference is whether it receives the Imaginator's guess.
| HOS | HPS | |
|---|---|---|
| HVS-3B | 48.04 | 24.12 |
| HVS-3B + Imaginator-8B | 62.75 | 39.38 |
Usage
vLLM cannot serve a model from a subdirectory of a repo, so fetch it first:
hf download jdzhang0929/Imaginator --include "Imagine-8B/*" --local-dir ./checkpoints
vllm serve ./checkpoints/Imagine-8B --port 8001 --served-model-name imaginator
With transformers the subfolder is addressable directly:
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained(
"jdzhang0929/Imaginator", subfolder="Imagine-8B",
dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained(
"jdzhang0929/Imaginator", subfolder="Imagine-8B")
Greedy decoding, max_tokens=4096. The full two-model loop โ view accumulation,
coordinate conversion, and the multi-hypothesis injection used in the paper โ
is in the code repo.
Training
Two stages. Stage 1 pretrains on 1.92M pseudo-labelled panorama samples (10,000 steps at batch 192) drawn from Sun360, Matterport3D and DiT360 renders. Stage 2 is a clean SFT on 4,223 H*Bench trajectories, 2 epochs at batch 64, lr 1e-5 cosine.
Stage 2 is reproducible from public data: the trajectories are at jdzhang0929/Imagine-in-360-Dataset and the images they reference ship with H*Bench.
Data separation
Evaluation panoramas were compared against the stage-2 SFT panoramas exhaustively at the pixel level โ 857 ร 382 pairs, 36 yaw rotations each, under two criteria (MAE < 2.0 grey levels for "same photo", Pearson r โฅ 0.90 for "same viewpoint"). Two overlaps surfaced, both inherited from the original H*Bench split, and both are removed from the released SFT trajectories. Against the 40,453 stage-1 pseudo-label panoramas the same sweep found nothing at either threshold; the global maximum correlation was 0.8965.
Limitations
- English instructions only.
- Top-1 coordinate accuracy is 39.04% (HOS) / 26.12% (HPS). Treat the output as a prior over where to look, not as a localization result.
- Rotation-only search from a fixed viewpoint; no translation.
- Coordinates assume the equirectangular convention above. A different yaw origin or pitch sign will silently produce plausible but wrong guesses.
Citation
@article{imagining360,
title = {Beyond Thinking: Imagining in 360{\deg} for Humanoid Visual Search},
author = {Zhang, Jingdong and others},
year = {2026}
}
Built on H*Bench and the HVS models from Yu et al.
Model tree for jdzhang0929/Imaginator
Base model
Qwen/Qwen3-VL-8B-Instruct