Mobile-OV: V11 Image Bridge + Exp1 Video Bridge
One downloadable inference-only bundle for
neo_mobileov-infer-clear.
This release selects V11-balanced (not V11-control), Exp1-64K and the released
NeoDragon Hybrid. It is not an experimental Joint V2/V3 or DMD checkpoint.
The older mobile_ov_135k_full.pt in this repository is a separate legacy file
and is not needed for this release.
Contents
- One shared SmolVLM2-500M backbone including its vision encoder and connector.
- V11-balanced image bridge: job 17291461, step 160000.
- Exp1 video bridge: job 17108893, step 64000.
- DreamLite-mobile image generator and image VAE.
- Released NeoDragon Hybrid video DiT and its video VAE encoder/decoder.
- Required configs, scheduler, tokenizers, processors and checksum manifest.
No optimizer, gradient scaler, RNG state, training history, teacher model, duplicate Smol backbone, native Qwen/T5/CLIP neural encoder, SSD1B or QuickSR weights are included. The small Qwen tokenizer remains for the trained image condition-length rule. All seven safetensors weight files retain their original bytes/dtypes; no quantization or lossy precision conversion was applied.
The .tar.gz is a single distribution file, not a pickle checkpoint to pass
to torch.load or a standard Transformers AutoModel directory. Its component
layout preserves lazy loading, shared weights and separate conversion targets.
It contains everything the clean inference runtime needs except Python packages
and the runtime code itself. No sibling research repo is needed.
Download And Run
Clone the clean code repo and install its tested dependencies first. Run the following from that repo root, using an empty destination for the new assets:
hf download leduy99/Mobile-OV \
v11-exp1/mobile-ov-v11-exp1-inference.tar.gz \
v11-exp1/mobile-ov-v11-exp1-inference.tar.gz.sha256 \
--local-dir downloads
(cd downloads/v11-exp1 && sha256sum -c mobile-ov-v11-exp1-inference.tar.gz.sha256)
mkdir -p assets
tar -xzf downloads/v11-exp1/mobile-ov-v11-exp1-inference.tar.gz -C assets
python generate.py --assets assets/v11-exp1 --device cpu --verify-assets \
--prompt 'A red toy car drives along a sunny coastal road.' \
--output output/example.mp4
The command generates the first frame with DreamLite and the continuation with NeoDragon. Default video: 49 frames at 24 fps, 320x512; image anchor: 640x1024, four steps. Logical DreamLite time IDs remain 800x1280, as in the validated recipe. The two VAEs are not interchangeable: the RGB image is encoded by the video VAE.
For GPU inference on a SLURM-managed machine, obtain an allocation before using
--device cuda, or use the clean repo's scripts/generate_local.sbatch.
python understand.py --assets assets/v11-exp1 --device cpu \
--image output/example.anchor.png --prompt 'Describe this image.'
Text understanding and sampled-frame video understanding are also supported. Instruction-based image/video editing is not validated.
Verification And Limitations
v11-exp1/bundle_report.json records archive size, checksum, per-component tensor
counts and dtypes. manifest.json inside the archive verifies every bundled
asset. Preparation checks all input hashes and detects changes while packing.
Source checkpoint hashes are preserved, but machine-local paths are omitted.
The clean runtime's recorded CPU parity checks cover conditions, three images, three videos and both bridge export graphs. Packaging does not improve the known prompt-following or motion weaknesses. This is a Linux/PyTorch reference, not a finished iPhone/iPad/CoreML model; device latency, memory and quantization quality are not established by this release. See the code repo's validation log.
Licenses And Attribution
Read NOTICES.md and the notices inside the archive before use. In particular, DreamLite is CC BY-NC 4.0 (non-commercial). NeoDragon carries BSD-3-Clause-Clear plus the Qualcomm Responsible AI License; SmolVLM2 is Apache-2.0. Packaging does not replace these terms or grant patent rights. Use this bundle for non-commercial research consistent with all component terms.
Original generators and backbone are the work of their respective authors. The Mobile-OV research contributes the trained bridges and their integration; no claim of training these foundational generators from scratch is made.
Model tree for leduy99/Mobile-OV
Base model
HuggingFaceTB/SmolLM2-360M