Thanks for the report, this is now fixed! :)
Hannes von Essen
AI & ML interests
Deep Learning Researcher at Embedl | Creator of hfviewer.com
Recent Activity
liked a model 3 days ago
IvmeLabs/Ivme-Conversate-v2-Base liked a model 3 days ago
Quazim0t0/Chimera-64M repliedto their post 5 days ago
๐ฃ Understand any Hugging Face model!
๐ฌ You can now hover a node to get a nice animated explanation of how the operation works and which paper introduced it! https://hfviewer.com
๐๏ธ Also, embedded HF Viewer graphs now default to being animated! Create and customize your animation at
https://hfviewer.com/model-card-embed
๐ฅ Finally, feel free to join our discord to influence the future direction of HF Viewer! https://discord.gg/a5eEmtTTPVOrganizations
replied to their post 5 days ago
posted an update 6 days ago
Post
1187
๐ฃ Countdown Kimi K3 with us!
Read our deep dive into the architecture of Kimi K3 and get notified on July 27 when it gets released and is viewable on hfviewer.com!
https://hfviewer.com/moonshotai/kimi-k3
Read our deep dive into the architecture of Kimi K3 and get notified on July 27 when it gets released and is viewable on hfviewer.com!
https://hfviewer.com/moonshotai/kimi-k3
posted an update 10 days ago
Post
109
๐ฃ Understand any Hugging Face model!
๐ฌ You can now hover a node to get a nice animated explanation of how the operation works and which paper introduced it! https://hfviewer.com
๐๏ธ Also, embedded HF Viewer graphs now default to being animated! Create and customize your animation at
https://hfviewer.com/model-card-embed
๐ฅ Finally, feel free to join our discord to influence the future direction of HF Viewer! https://discord.gg/a5eEmtTTPV
๐ฌ You can now hover a node to get a nice animated explanation of how the operation works and which paper introduced it! https://hfviewer.com
๐๏ธ Also, embedded HF Viewer graphs now default to being animated! Create and customize your animation at
https://hfviewer.com/model-card-embed
๐ฅ Finally, feel free to join our discord to influence the future direction of HF Viewer! https://discord.gg/a5eEmtTTPV
posted an update 18 days ago
Post
156
๐ฃ HF Viewer now supports Hugging Face login! ๐ค
โก Generate visualizations for all your models at once!
๐ Feature your models in the community showcase and set up your own profile/org collection page on hfviewer!
โ๏ธ Write your own model articles on hfviewer in the same novel interactive style as our "Gemma 4 family" article - with linking between the graph nodes and article text!
๐ We are also now rolling out support for tensor shapes, FLOPs and param counts per layer! :)
Thanks for all your positive feedback and suggestions! โค๏ธ
Try the logged in experience here: https://hfviewer.com/
Here are some really cool articles already written by our users:
๐ LFM2.5-Audio: edge-first speech inference by Anna Piunova at Liquid AI
https://hfviewer.com/LiquidAI/LFM2.5-Audio-1.5B
๐ DeepSeek V4 mHC Explained by Shakti Wadekar
https://hfviewer.com/deepseek-ai/DeepSeek-V4-Pro
๐ Borealis - open recipe for training Audio LLMs by Alex Wortega
https://hfviewer.com/Vikhrmodels/Borealis-5b-it
โก Generate visualizations for all your models at once!
๐ Feature your models in the community showcase and set up your own profile/org collection page on hfviewer!
โ๏ธ Write your own model articles on hfviewer in the same novel interactive style as our "Gemma 4 family" article - with linking between the graph nodes and article text!
๐ We are also now rolling out support for tensor shapes, FLOPs and param counts per layer! :)
Thanks for all your positive feedback and suggestions! โค๏ธ
Try the logged in experience here: https://hfviewer.com/
Here are some really cool articles already written by our users:
๐ LFM2.5-Audio: edge-first speech inference by Anna Piunova at Liquid AI
https://hfviewer.com/LiquidAI/LFM2.5-Audio-1.5B
๐ DeepSeek V4 mHC Explained by Shakti Wadekar
https://hfviewer.com/deepseek-ai/DeepSeek-V4-Pro
๐ Borealis - open recipe for training Audio LLMs by Alex Wortega
https://hfviewer.com/Vikhrmodels/Borealis-5b-it
posted an update 2 months ago
Post
4932
๐ฃ Add architecture visualization to model card!
๐ For all creators out there: add a model visualization to your model card to capture your audience's attention!
๐ฑ๏ธ When clicked, it opens an interactive view with multiple levels of granularity!
1๏ธโฃ Paste url at https://hfviewer.com/model-card-embed
2๏ธโฃ Paste generated code in your README.md!
3๏ธโฃ โจ
๐ For all creators out there: add a model visualization to your model card to capture your audience's attention!
๐ฑ๏ธ When clicked, it opens an interactive view with multiple levels of granularity!
1๏ธโฃ Paste url at https://hfviewer.com/model-card-embed
2๏ธโฃ Paste generated code in your README.md!
3๏ธโฃ โจ
replied to their post 3 months ago
Hi! Simply paste the link to the Hugging Face model you want to view on https://hfviewer.com, or install the extension from the post to have it automatically on all model pages you visit! :)
posted an update 3 months ago
Post
11674
๐ฃ Hugging Face Visualizer, now as Chrome extension!
https://hfviewer.com
โจ After installing, Hugging Face model pages will have an architecture visualization on the model page itself!
๐ Link:
https://chromewebstore.google.com/detail/hugging-face-viewer/mmadlggmpkpiockpjfepaohcllbnakej
Thanks for all the nice feedback so far! โค๏ธ
https://hfviewer.com
โจ After installing, Hugging Face model pages will have an architecture visualization on the model page itself!
๐ Link:
https://chromewebstore.google.com/detail/hugging-face-viewer/mmadlggmpkpiockpjfepaohcllbnakej
Thanks for all the nice feedback so far! โค๏ธ
reacted to unmodeled-tyler's post with ๐ 3 months ago
Post
4130
Hey Hugging Face!
Repo: https://github.com/unmodeled-tyler/vessel-browser
I wanted to share a cool feature from my open source AI native web browser, Vessel: Persistent highlights!
You can highlight anything on the page and the context is provided to the agent. It's kind of a fun way to learn about new stuff, synthesize info, or just deepen your comprehension/understanding.
Since highlights are persistent, you can close the page, come back later - and your highlights will be exactly where you left them. I've found this particularly useful when reviewing technical blogs, model cards, etc.
Check it out!
Repo: https://github.com/unmodeled-tyler/vessel-browser
I wanted to share a cool feature from my open source AI native web browser, Vessel: Persistent highlights!
You can highlight anything on the page and the context is provided to the agent. It's kind of a fun way to learn about new stuff, synthesize info, or just deepen your comprehension/understanding.
Since highlights are persistent, you can close the page, come back later - and your highlights will be exactly where you left them. I've found this particularly useful when reviewing technical blogs, model cards, etc.
Check it out!
posted an update 3 months ago
Post
276
๐ฃ I made a visualizer for Hugging Face models: https://hfviewer.com
โจ Simply paste a Hugging Face URL to get an interactive visualization of the architecture!
๐ The recent Qwen3.6-27B model as an example: https://hfviewer.com/Qwen/Qwen3.6-27B
Feel free to try it out and give me feedback on how it can be improved! โค๏ธ
โจ Simply paste a Hugging Face URL to get an interactive visualization of the architecture!
๐ The recent Qwen3.6-27B model as an example: https://hfviewer.com/Qwen/Qwen3.6-27B
Feel free to try it out and give me feedback on how it can be improved! โค๏ธ
reacted to JonnaMat's post with ๐ฅ๐ 5 months ago
Post
1090
Qwen3.5 on-device benchmarks on the Nvidia Jetson lineup are now live ๐
We've added the latest Qwen3.5 models (0
8B - 9B) to our on-device inference benchmarks (Nvidia Jetson Orin Nano Super, AGX Orin, AGX Thor).
๐ Explore TPS, TTFT, E2E latency, and TPOT. Measured on real hardware: embedl/Edge-Inference-Benchmarks
๐ Stay tuned for additional benchmarks and Embedl-optimized models: Enabling models run faster and on less expensive hardware.
If you're working on edge LLM deployment, we'd love to discuss your use case.
We've added the latest Qwen3.5 models (0
8B - 9B) to our on-device inference benchmarks (Nvidia Jetson Orin Nano Super, AGX Orin, AGX Thor).
๐ Explore TPS, TTFT, E2E latency, and TPOT. Measured on real hardware: embedl/Edge-Inference-Benchmarks
๐ Stay tuned for additional benchmarks and Embedl-optimized models: Enabling models run faster and on less expensive hardware.
If you're working on edge LLM deployment, we'd love to discuss your use case.
reacted to JonnaMat's post with ๐ฅ 5 months ago
Post
1620
๐คฏ Edge-Grade Vision Reasoning. Now Practically Lossless. ๐คฏ
Introducing
๐ embedl/Cosmos-Reason2-2B-W4A16-Edge2
Optimized for Jetson Orin Nano Super and AGX Orin
nvidia .
๐ Try it out on Jetson (image+video+text):
๐ค What is Edge2? Most weights โ INT4 | Activations โ FP16 | Select sensitive layers โ kept in FP16.
Edge2 preserves precision where it matters most; while keeping the model small and fast enough for edge GPUs. ๐
Introducing
๐ embedl/Cosmos-Reason2-2B-W4A16-Edge2
Optimized for Jetson Orin Nano Super and AGX Orin
๐ Try it out on Jetson (image+video+text):
docker run --rm -it \
--network host \
--shm-size=8g \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
--runtime=nvidia \
--name=vllm-serve \
-e HF_TOKEN=hf_*** \
-e HF_HOME=/root/.cache/huggingface \
ghcr.io/nvidia-ai-iot/vllm:latest-jetson-orin \
vllm serve "embedl/Cosmos-Reason2-2B-W4A16-Edge2" \
--max-model-len 8192 \
--gpu-memory-utilization 0.75 \
--max-num-seqs 2๐ค What is Edge2? Most weights โ INT4 | Activations โ FP16 | Select sensitive layers โ kept in FP16.
Edge2 preserves precision where it matters most; while keeping the model small and fast enough for edge GPUs. ๐
posted an update 5 months ago
Post
1373
Cosmos-Reason2-2B on Nano Super
Hi! Today, me and my team is releasing a version of Cosmos-Reason2-2B that is quantized so that it fits on the NVIDIA Jetson Orin Nano Super.
We managed to find a mixed precision configuration such that it maintains virtually the same accuracy as the unquantized model while being able to run really efficiently on the Nano Super and other edge devices :)
embedl/Cosmos-Reason2-2B-W4A16-Edge2
Hi! Today, me and my team is releasing a version of Cosmos-Reason2-2B that is quantized so that it fits on the NVIDIA Jetson Orin Nano Super.
We managed to find a mixed precision configuration such that it maintains virtually the same accuracy as the unquantized model while being able to run really efficiently on the Nano Super and other edge devices :)
embedl/Cosmos-Reason2-2B-W4A16-Edge2
reacted to JonnaMat's post with ๐ฅ 5 months ago
Post
2569
โก Blackwell-native Vision Reasoning at the edge โก
Released a NVFP4A16-variant of nvidia/Cosmos-Reason2-2B:
embedl/Cosmos-Reason2-2B-NVFP4A16
๐ Optimized for Blackwell with minimal accuracy drop compared to its FP16 counterpart.
Thorough on-device benchmarks on AGX Thor in the modelcard. ๐ค ๐
Try it out:
Released a NVFP4A16-variant of nvidia/Cosmos-Reason2-2B:
embedl/Cosmos-Reason2-2B-NVFP4A16
๐ Optimized for Blackwell with minimal accuracy drop compared to its FP16 counterpart.
Thorough on-device benchmarks on AGX Thor in the modelcard. ๐ค ๐
Try it out:
docker run --rm -it \
--network host \
--shm-size=8g \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
--runtime=nvidia \
--name=vllm-serve \
-e HF_TOKEN=hf_*** \
-e HF_HOME=/root/.cache/huggingface \
nvcr.io/nvidia/vllm:26.01-py3 \
vllm serve "embedl/Cosmos-Reason2-2B-NVFP4A16" \
--host 0.0.0.0 \
--port 8000 \
--tensor-parallel-size 1 \
--max-model-len 16384 \
--gpu-memory-utilization 0.9