Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Jinki Jeong
PRO
Anserwise
1
23
175
Follow
mayafree's profile picture
webxos's profile picture
ginipick's profile picture
26 followers
ยท
33 following
AI & ML interests
None yet
Recent Activity
liked
a Space
about 4 hours ago
FINAL-Bench/POCKET-Image-Studio
liked
a model
about 4 hours ago
FINAL-Bench/POCKET-Image-Zimage
reacted
to
SeaWolf-AI
's
post
with ๐ฅ
1 day ago
POCKET now speaks Gemma 4 โ a 26B model that loads in every app, and runs on your PC with no GPU We're adding a Gemma-4 sibling to POCKET: POCKET-26B, built from Google's Gemma-4-26B-A4B (Apache-2.0). Our flagship POCKET-35B is a Qwen-family MoE and needs a recent llama.cpp; POCKET-26B trades a little size for the thing people kept asking for โ it just loads, everywhere, today: Ollama, LM Studio, PocketPal, MLX, any stock llama.cpp. No fork, no bleeding-edge runtime, no CUDA, no cloud. It's a sparse Mixture-of-Experts (25.2B total, ~4B active per token), so the work per token stays small โ a real 26B that generates on a CPU with no graphics card. Two things make it stand out: 1) Universal compatibility. Gemma 4 is a standard, widely-supported architecture, so POCKET-26B runs on the tools you already have โ no waiting for your app to add a new model type. 2) Quality that survives compression. Measured GPQA-Diamond (198 q, greedy): โข Full base: 67.7% โข POCKET-26B Q4_K_M (17 GB): 67.7% โ lossless โข POCKET-26B Q2_K (11 GB): 67.2% โ near-lossless, at 11 GB Live, on a CPU-only box (our demo Space โ POCKET-26B vs Bonsai-27B, same machine, same stock llama.cpp): POCKET-26B โ 19 tok/s vs Bonsai โ 6 tok/s โ about 3ร faster generation, no GPU. (Honest notes: shared CPU box, sequential race; a dedicated machine is faster.) Where it fits in the family: โข POCKET-35B (Qwen MoE) โ bigger, top-tier, needs a recent llama.cpp. โข POCKET-26B (Gemma 4) โ loads in any app, quality-robust when compressed. The demo runs the Q4_K_M build; Q2_K (11 GB) is the smallest footprint. For a true โค8 GB phone, the 5 GB POCKET-KR (Qwen) is still the pick. Try it and grab it: ๐ฅ๏ธ Live demo (Gemma4-based, answering on a CPU, no GPU): https://huggingface.co/spaces/FINAL-Bench/POCKET-26B-CPU ๐ฆ POCKET-26B-GGUF (Q4_K_M 17 GB ยท Q2_K 11 GB): https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF ๐ POCKET collection: https://huggingface.co/collections/FINAL-Bench/pocket-models
View all activity
Organizations
None yet
models
5
Sort:ย Recently updated
Anserwise/AWAXIS-Hybrid-28B
Text Generation
โข
28B
โข
Updated
Jun 8
โข
13
โข
2
Anserwise/AWAXIS-KR-31B
Image-Text-to-Text
โข
52B
โข
Updated
Jun 8
โข
27
โข
4
Anserwise/AWAXIS-Think-31B
Text Generation
โข
31B
โข
Updated
Jun 4
โข
74
โข
2
Anserwise/AWAXIS-Think-28B
Text Generation
โข
28B
โข
Updated
Apr 24
โข
65
โข
17
Anserwise/AWAXIS-Think-27b
Text Generation
โข
27B
โข
Updated
Apr 23
โข
8
โข
1
datasets
0
None public yet